Retrieval-Augmented Generation · AI Agents · LLM Reasoning
GrepSeek vs FORT-Searcher vs SAAS: Three Fixes for the Search-Agent Loop
GrepSeek changes what the agent searches with (grep, no vector index), FORT-Searcher changes the training data (shortcut-resistant tasks), and SAAS changes when it stops (half the retrievals, equal accuracy).
The search-agent loop, and where each fix lands
A trained search agent runs the same loop as an untrained one: decide what to look for, issue a search, read what came back, decide whether to search again, then answer. Three 2026 papers attack three different links in that chain, and almost nothing they report is directly comparable, which is itself the most useful fact in this comparison.
GrepSeek replaces the search tool. Instead of querying a pre-built embedding index, the agent runs shell commands (grep, pipes, regex) directly against the raw corpus. Its claim is about retrieval fidelity and infrastructure: no index to build, no index to go stale, exact where embeddings are fuzzy. FORT-Searcher leaves the tool alone and rebuilds the training data: it synthesizes deep-search tasks engineered so that no single clue exposes the answer early, then SFTs a small agent on them. Its claim is about capability per active parameter. SAAS leaves both alone and retrains the stopping decision: it uses reinforcement learning to teach a small model where its own knowledge boundary sits, so it stops firing retrievals it does not need. Its claim is about cost at equal accuracy.
So the honest framing is a division of labor, not a horse race. A team can want all three at once: FORT-style data to teach real evidence acquisition, a GrepSeek-style tool if the corpus suits lexical access, and SAAS-style RL to keep the inference bill down.
GrepSeek: change what the agent searches with
Standard RAG embeds every document once and answers by nearest-neighbor lookup. That index is a frozen, lossy summary of the corpus: you see only what the embedding model decided to encode, and every corpus update forces a re-embedding pass. GrepSeek’s direct corpus interaction (DCI) drops the index. The agent searches the actual bytes the way a developer searches a codebase: literal or regex patterns, refinement after reading hits, leads followed across files.
Training is two-stage. First, a cold-start set of verified trajectories built by pairing an answer-aware Tutor (which steers queries toward evidence) with an answer-blind Planner (which keeps the trajectories realistic, since the deployed agent will not know the answer either). Only trajectories that actually surface the supporting text survive. Second, GRPO (the same critic-free group-relative RL popularized by DeepSeek-R1) refines the cold-started policy on whether trajectories led to correct answers.
The engineering piece that makes DCI trainable at all is a sharded-parallel execution engine: it splits each grep across shards, runs them concurrently, and guarantees byte-exact equivalence to serial execution. A parallel search that silently reordered matches would corrupt the RL reward signal, so the byte-exactness clause is what converts a 7.6x speedup from a demo trick into a training-loop component.
The paper’s own ceiling is lexical: queries with heavy surface-form variation (synonyms, paraphrase, morphology) are exactly where dense retrieval wins, and the authors conclude DCI works best alongside embedding retrieval, not instead of it.
FORT-Searcher: change what the agent is trained on
FORT starts from an uncomfortable measurement: a multi-hop question can look hard on paper yet collapse in practice, because one rare clue identifies the answer in the first query, or the model already knows the entity from parametric memory. Training on such tasks teaches recognition of overexposed clues, not search. FORT defines four shortcut risks: evidence co-coverage, single-clue selectivity, exposed constants, and prior-knowledge binding. It then synthesizes tasks that control all four, with an adversarial refinement step in which a strong search agent tries to break each draft task. Shortcut-prone drafts get repaired; over-fuzzed drafts get narrowed until they are solvable but still search-heavy.
The resulting FORT-Searcher is deliberately modest in scale: Qwen3-30B-A3B-Thinking with about 3B active parameters, trained with supervised fine-tuning only. That restraint is what makes the data-quality claim legible: the gains cannot be credited to a bigger model.
The most transferable contribution is diagnostic rather than the checkpoint. FORT measures when the answer first appears in a trajectory (answer hit time) and how much solving actually costs. FORT data pushes average answer hit time from 18.7 to 46.9 and average solving cost from 92.1 to 141.0 versus REDSearcher data under the same diagnostic setting: the agent must acquire evidence before it can answer. And the ablation is the cleanest causality evidence in this set: strip the shortcut controls and synthetic-question accuracy jumps from 29.0 to 81.6 while answer hit time collapses from 46.5 to 11.8. The tasks became easier in exactly the way the pipeline was designed to prevent.
SAAS: change when the agent stops
SAAS targets over-search: retrieval calls the agent makes when its parametric knowledge already covered the question (unnecessary search), or after it has already gathered enough evidence (redundant search). The measured baseline is stark: efficiency-focused HiPRAG still over-searches on effectively every question, a 100% question-level over-search ratio.
The method has three parts. A search boundary is estimated per training question by contrasting two rollouts, one with search disabled and one enabled, under the current policy, and re-estimated as the policy changes. A boundary-aware reward penalizes unnecessary searches with zero tolerance but only charges redundant searches past a minimum-sufficient threshold, which avoids the flat-penalty failure mode where the agent learns to be search-shy and hard multi-hop questions collapse. And a two-stage curriculum trains reasoning capability first and turns on the efficiency regularization only afterward, so the model cannot game the reward by refusing to search before it knows how to search.
The result on Qwen2.5-7B-Instruct: average search count falls from 2.19 to 0.97 per question while average accuracy holds at 48.7% versus the HiPRAG baseline’s 49.8%. The question-level over-search ratio drops from 100% to 45.9%, and the step-level redundant-search ratio falls from 19.5% to 6.3%. Results are also reported on Qwen2.5-3B-Instruct (1.13 searches) and Qwen3-4B-Instruct, across seven QA benchmarks: TriviaQA, PopQA, NQ, HotpotQA, 2WikiMultiHopQA, MuSiQue, and Bamboogle. The authors are explicit that accuracy is traded for cost, not improved; if the bottleneck is answer quality, this is the wrong tool.
Key numbers
| Measurement | GrepSeek | FORT-Searcher | SAAS | Setting and source | Same harness? |
|---|---|---|---|---|---|
| Agent scale | Not emphasized | 3B active (Qwen3-30B-A3B) | 7B / 3B / 4B instruct | Each paper’s own setup | No |
| Training signal | Cold-start SFT, then GRPO | SFT only on FORT data | RL with boundary-aware reward | Each paper’s own setup | No |
| Overall deep-search score | Not reported | 66.2 overall | Not reported | FORT: five benchmarks with complete comparable-size results | Within FORT row only |
| Named comparators | Best F1/EM aggregate vs retrieval and agent baselines | MiroThinker-1.7-mini 64.6, Qwen3.5-35B-A3B 59.9 | HiPRAG accuracy 49.8% | Each paper’s own baselines | No |
| BrowseComp | Not reported | 72.2 (55.9 without context management) | Not reported | FORT paper | No |
| BrowseComp-ZH | Not reported | 75.0 (62.1 without context management) | Not reported | FORT paper | No |
| xbench-DeepSearch-2505 / 2510 | Not reported | 80.8 / 57.2 | Not reported | FORT paper | No |
| Seal-0 | Not reported | 46.0 | Not reported | FORT paper | No |
| Trajectory diagnostic | Not reported | Answer hit 46.9, solving cost 141.0 (REDSearcher: 18.7 / 92.1) | Step-level redundant search 6.3% (baseline 19.5%) | Each paper’s own diagnostic | No |
| Shortcut-control ablation | Not applicable | Accuracy 29.0 → 81.6, hit time 46.5 → 11.8 when controls removed | Not applicable | FORT paper | Yes, ablation only |
| Avg searches per question | Not reported | Not reported | 0.97 vs HiPRAG 2.19 (7B); 1.13 (3B) | SAAS paper, seven QA benchmarks | Within SAAS row only |
| Accuracy vs efficiency baseline | Best F1/EM on seven QA benchmarks (aggregate) | Not measured | 48.7% vs HiPRAG 49.8% | SAAS paper, Qwen2.5-7B | Yes, SAAS vs HiPRAG |
| Over-search ratio | Not reported | Not reported | 45.9% vs HiPRAG 100% | SAAS paper | Yes, SAAS vs HiPRAG |
| Infrastructure claim | 7.6x sharded-parallel grep, byte-exact | Not applicable | Not applicable | GrepSeek paper | n/a |
Read the table for what it does not contain. No paper here reports a head-to-head against either of the other two. FORT’s 66.2 and SAAS’s 48.7% are different metrics on different benchmarks with different agents; GrepSeek’s “best F1/EM aggregate” is not on the same seven benchmarks as SAAS’s seven QA benchmarks, and FORT’s BrowseComp numbers share no harness with either. The only strictly same-harness comparisons are internal: FORT versus its named baseline agents, SAAS versus HiPRAG, and each paper’s own ablations.
What the three share, and what they don’t
All three treat search as a trainable behavior rather than a prompt-engineering problem, and all three report their agents at modest scale (FORT explicitly so, to isolate the data effect). All three also ship something more reusable than a checkpoint: GrepSeek’s verified cold-start recipe, FORT’s trajectory diagnostics (log when the answer first appears and whether the model guessed before evidence), and SAAS’s staged reward design are each portable to other agents.
The differences matter for adoption. GrepSeek changes your infrastructure: you need shell-level access to the corpus and a tolerance for lexical failure modes. FORT changes your data pipeline: you need a task-synthesis loop and a strong adversary agent, but the deployed system is ordinary SFT. SAAS changes your reward: you need RL infrastructure and boundary-estimation rollouts at training time, but the deployment is a smaller inference bill. And their failure modes are disjoint: GrepSeek loses on paraphrase-heavy corpora, FORT’s gains are unproven outside verifiable short-answer search tasks, and SAAS explicitly gives up nothing on accuracy but gains nothing either.
When to use which
- Your corpus is a codebase, a filesystem, or any text you can shell into, and building an embedding index is expensive or stale-prone: GrepSeek’s DCI is the natural fit, run hybrid with dense retrieval for paraphrase-heavy queries.
- Your search agent answers before it searches, or one clue collapses your multi-hop tasks: the FORT diagnostics come first. Log answer hit time on your own trajectories; if answers arrive early, no model scale will fix a data problem that FORT-style synthesis addresses directly.
- Your retrieval bill (latency, API cost, context tokens) dwarfs your model bill: SAAS is the only one of the three that attacks cost directly, and its 2.19-to-0.97 result comes at a one-point accuracy delta on a 7B model.
- You can only do one: fix the data. FORT’s ablation (accuracy 29.0 to 81.6 and hit time 46.5 to 11.8 when shortcut controls are removed) is the single largest measured effect in this set, and bad training data silently caps everything downstream.
- You can stack them: FORT data teaches evidence acquisition, GrepSeek’s tool decides how evidence is fetched, SAAS’s RL trims the fetch count. Nothing in the three papers contradicts; they operate on different links of the same loop.
Limits and open questions
The central gap is the absence of a shared harness: nobody trained one agent on FORT data, armed it with GrepSeek’s tool, added SAAS’s reward, and reported the combination against single-fix ablations. Until that exists, the division of labor above is an inference from three independent papers, not a measured decomposition. Scale is also unexplored: FORT’s agent is 3B active, SAAS’s largest is 7B, and whether shortcut-resistant data or boundary-aware rewards still matter at frontier scale is an open question. GrepSeek’s evaluation is lexical-access-friendly corpora, which biases it away from the paraphrase-heavy open-web regime where BrowseComp-style agents live; FORT and SAAS use web-search or fixed-corpus QA setups that GrepSeek’s byte-exact engine does not address. And all three evaluate verifiable short answers; none of them measures search quality for open-ended research, enterprise knowledge bases, or judgment-graded tasks. Finally, both FORT’s diagnostics and SAAS’s boundary estimation rely on strong-model judges or contrastive rollouts at training time, so their numbers are operational proxies, not ground truth.
FAQ
How do the three methods compare on the same benchmark?
They do not. FORT-Searcher reports deep-search benchmarks (66.2 overall, BrowseComp 72.2, xbench-DeepSearch-2505 80.8, Seal-0 46.0), SAAS reports QA accuracy plus search counts on seven QA benchmarks (48.7% accuracy at 0.97 searches versus HiPRAG’s 49.8% at 2.19), and GrepSeek reports an aggregate F1/Exact-Match lead across its own seven benchmarks. Any table that ranks the three agents against each other is combining numbers that never shared a harness.
What results does FORT-Searcher actually deliver?
Among comparable-size open agents with complete scores, 66.2 overall across five benchmarks (ahead of MiroThinker-1.7-mini at 64.6 and Qwen3.5-35B-A3B at 59.9), with 72.2 on BrowseComp, 75.0 on BrowseComp-ZH, 80.8 on xbench-DeepSearch-2505, 57.2 on xbench-DeepSearch-2510, and 46.0 on Seal-0. Context management alone moves BrowseComp from 55.9 to 72.2.
Does SAAS improve QA accuracy compared with normal search agents?
No, and it does not claim to. On Qwen2.5-7B-Instruct, average accuracy is 48.7% versus HiPRAG’s 49.8%, within about a point, while average searches per question drop from 2.19 to 0.97 and the over-search ratio falls from 100% to 45.9%. It is a cost method, not an accuracy method.
What is different about GrepSeek’s method compared with RAG?
GrepSeek removes the embedding index entirely: the agent runs shell commands like grep and regex directly on the corpus bytes, trained first on verified cold-start trajectories from a paired Tutor and Planner, then with GRPO. Its sharded-parallel engine makes that 7.6x faster than serial execution while staying byte-exact. The measured cost is weakness on paraphrase and synonym queries, which is why the authors position it as a complement to dense retrieval rather than a replacement.
Which method should a search-agent builder try first?
Diagnose before choosing. If trajectories show the answer appearing before real evidence gathering, start with FORT-style data synthesis and its hit-time logging. If the bill is the problem and accuracy is acceptable, SAAS. If the index is the problem (stale, expensive, or missing shell-level access), GrepSeek. The wrong first move is picking the highest headline number, because the three numbers measure three different things.