Pre-Fetching and Predictive Retrieval in Agentic Systems
Predictive signals in language models enable retrieval before information demands become critical.

An agent running standard retrieval-augmented generation hits the same wall over and over: it generates text, reaches a point where it needs outside information, stops, sends a retrieval call, waits for the result, and only then resumes. That stall is a structural property of synchronous retrieval itself, baked into the architecture that couples generation to retrieval round-trip time, not a bug that faster databases or leaner embeddings could optimize away. Wuyang Zhang and Shichao Pei, in the paper behind the ICML 2026 poster "Predictive Prefetching for Retrieval-Augmented Generation" (arXiv:2605.17989), frame this coupling as the core limitation of synchronous RAG: because retrieval and generation are sequential and dependent, the speed of generation can never exceed the speed of the slowest retrieval call in the chain. Complex agentic tasks compound the problem, since each one requires many such pauses strung together, and every pause adds a latency tax that stacks on the one before it.
Some asynchronous approaches already try to decouple the two processes, but Zhang and Pei point out that these rely on heuristic coordination and assume information demands stay stable while the model decodes. A retrieval trigger built for a fixed information target cannot track a target that moves. That mismatch, between static triggering logic and dynamic information need, is the condition predictive prefetching was built to address.
Generation dynamics signal retrieval needs before the model knows it
A language model's generation process gives off signals of what it is about to need before it consciously registers that need. Zhang and Pei show that internal representations, token-level entropy, attention patterns, and value states, carry signs of rising uncertainty well before that uncertainty becomes critical. The paper's own figures put this window at roughly 8 to 16 tokens before the moment the model would otherwise stall and wait for a retrieval call. The model's forward pass, in other words, leaks information about its own coming ignorance before the ignorance arrives.
Zhang and Pei call these leaks semantic precursors in generation dynamics. Generation behavior becomes a kind of forecast of its own future needs, readable in the model's internals several tokens before the model would otherwise be forced to pause.
The practical consequence is significant. If uncertainty can be forecast rather than merely detected, retrieval does not have to wait for the moment of need. It can begin during the window between precursor and peak, while the model is still generating text it already knows how to produce, absorbing the retrieval latency inside generation time.
This is a meaningfully different claim than the one behind simple entropy thresholding, an older and more common technique. Zhang and Pei's retrieval predictor component outperforms entropy thresholding on prediction accuracy by a measurable margin, because it is built to catch the earlier signal. A system that acts on the precursor is actually forecasting.
The three-component architecture that operationalizes semantic precursors
Turning that forecasting insight into a working system requires three components running in parallel with generation. Zhang and Pei's framework names them the retrieval predictor, the context monitor, and the query generator, and the coordination between the three is what makes the system work as prefetching rather than as a faster version of the old reactive loop.
The retrieval predictor is the component that watches the signals described above. That probability becomes a trigger.
A trigger alone is not enough, so the context monitor earns its place in the architecture. The context monitor exists to prevent exactly that kind of premature, underspecified retrieval, and it is the piece of the design that most clearly separates this architecture from naive parallel retrieval schemes that fire as soon as any signal appears.
The goal is for results to arrive timed to when the model actually needs them, not simply as fast as technically possible.
All three components share a Result Cache, where documents and embeddings are stored, and all three improve through online learning that adapts them based on retrieval outcomes. That means the predictor is not a fixed rule shipped once and left alone. It tunes itself to the domain and the specific model it runs alongside.
One of the more counterintuitive findings in the paper concerns timing. There is an optimal prefetch window, and it sits a few tokens later than the earliest possible trigger, not at the earliest point technically available.
Zhang and Pei evaluate the architecture across multiple model families, including Llama, GPT-OSS, and Qwen, and find consistent improvements in time-to-first-token across all of them. The breadth of that result matters less for any single number than for what it implies: the mechanism is a property of how generation itself behaves under uncertainty, not an artifact tuned to one architecture.
Query timing and formulation matter as much as trigger timing
Knowing when to retrieve solves only half the problem. A system that fires at exactly the right moment but asks the wrong question squanders the latency it just saved and can hand the model content that actively misleads it.
Earlier asynchronous systems ran into this directly. PipeRAG, an earlier approach Zhang and Pei cite by name, runs retrieval in parallel with generation but builds its queries from context that is already out of date by the time the fetch resolves. Zhang and Pei identify this query staleness as the central failure mode their framework is built to correct.
That is also why the 3 to 4 token waiting window discussed above is not an arbitrary tuning choice. Pushed too late, the query stays well-formed but the latency savings shrink.
A separate but related question is how many rounds of retrieval a task actually needs. Predictive prefetching does not remove the need for that iterative structure where it is warranted. It changes the timing of each round within that structure. The query generator has to formulate a query for a context the model has not reached yet, anticipating where the reasoning is headed rather than only describing where it currently stands.
How well the performance evidence holds up
The results reported at ICML 2026 are real and substantial. Across four question-answering benchmarks, Zhang and Pei report up to 43.5% reduction in end-to-end latency and up to 62.4% improvement in time-to-first-token, alongside a reduction in the number of retrieval calls issued, all while holding answer quality comparable to synchronous RAG baselines.
The retrieval call reduction deserves attention on its own terms. A system that gets faster while also making fewer calls is learning which fetches to skip, so the gains come from selectivity as well as speed.
The benchmark profile, though, deserves precision. These are question-answering tasks: structured, bounded, with a defined factual target. Real agentic work is often open-ended and multi-step, with information needs that shift within a single task rather than staying fixed from start to finish, and that open-ended regime is closer to the one where Zhang and Pei themselves note that assumptions of stable information demand break down. But the four QA benchmarks used to demonstrate the gains do not themselves stress-test the harder multi-domain case the paper is reasoning about. The results establish that predictive prefetching works well on bounded factual retrieval. They do not yet establish how far the same gains extend into open-ended, shifting-target agentic work.
The objection that predictive prefetching cannot solve: dynamic context in long-horizon tasks
The strongest challenge to this architecture is that predictive prefetching assumes a retrieval target already exists at the moment the precursor signal fires, and in long-horizon agentic tasks that assumption can simply be false.
Retrieval-augmented generation, in its standard form, is built around a fixed corpus: a defined set of documents exists, and retrieval selects from within it. There is no corpus to select from in advance because the scope of relevant information is still being constructed by the task. No amount of predictive accuracy solves a problem where the target does not yet exist to be predicted.
A related failure runs in the opposite direction. A system that prefetches aggressively can therefore make a model's context window larger and its effective use of that window worse at the same time.
There is also a coordination cost baked into the predictor itself. Its efficiency depends entirely on getting the prediction right. None of this makes prefetching the wrong approach. It marks the edge of what prefetching alone can do, and it is the reason predictive triggering has to sit inside a broader retrieval strategy that also includes active context management and the iterative query refinement described earlier, rather than standing in as a complete solution by itself.
Production agentic retrieval system design around these constraints
Production teams are converging on a layered design that treats retrieval triggering, context assembly, and grounding as three distinct engineering problems. That separation is a direct response to the limits just described: no single mechanism, prefetching included, can carry the whole burden of making agentic retrieval reliable.
Azure AI Search's agentic retrieval offers a concrete look at where enterprise infrastructure stands on this. As of the 2026-04-01 REST API version, select features are generally available, while more advanced capabilities, such as query planning and answer synthesis, remain in preview. What matters architecturally is that retrieval decisions are embedded directly into the reasoning flow rather than treated as a pre-step that happens before reasoning starts or a post-step that happens after it ends. That is the research pattern described earlier appearing as supported production infrastructure, a pattern visible in this vendor's product but not a verdict on it alone.
The pattern driving all of this predates predictive prefetching and will outlast any single implementation of it. Retrieval strategy research (arXiv:2509.04820) reinforces this finding: systems built around iterative retrieval show mixed but directionally clear results, helping complex tasks while occasionally costing simpler ones a little performance, regardless of how quickly any individual fetch resolves. Speed inside a bad loop is still a bad loop. Teams building these systems must decide not only whether they can prefetch, but what their retrieval loop looks like in the first place, and where prefetching fits inside it.
Context engineering as the discipline that determines whether prefetched content helps or hurts
Fetching content faster does not automatically make an agent reason better with it. Whether prefetched material helps or hurts depends on a separate discipline: context engineering, the set of decisions that govern what enters the model's context window, in what form, and where.
The common failure mode in production systems is treating the context window as a staging area, dumping retrieved content into it and trusting the model to sort out what matters. That trust is frequently misplaced. Longer context does not reliably produce better answers, and position-dependent biases mean models can systematically underuse material buried in the middle of a long input, regardless of how relevant that material actually is.
Context engineering addresses this through three active decisions. Positioning determines where within the context retrieved material sits, and that placement is an empirically consequential choice.
The link back to prefetching is direct. Content fetched against a predicted future state can be entirely accurate and still arrive misaligned with where generation actually went. Without that reconciliation step, the latency gains from prefetching do not reliably turn into gains in output quality.
The retrieval layer as the attack surface prefetching expands
A retrieval system that runs continuously and autonomously, reaching out to live content ahead of any explicit user request, is a larger and more persistent attack surface than one that only retrieves when asked.
The structural vulnerability here is indirect prompt injection: instructions hidden inside retrieved content bypass user-facing safeguards because the agent treats retrieved content as trusted context. That is a property of how retrieval-augmented generation handles the material it brings in, not a flaw specific to any one system.
The GeminiJack incident shows what this looks like in practice. Once that content is indexed by a RAG system, any employee running an ordinary search can trigger those instructions, which then prompt the agent to search across connected data sources and exfiltrate information through an image URL. The attack scales precisely because it requires no targeting of individual users; it only requires the content to be indexed once.
CVE-2025-53773, affecting GitHub Copilot, shows the same category of weakness in a different context. The retrieval of PR content was the actual attack vector, not some separate exploit layered on top of it.
Prefetching widens this exposure. The EU AI Act's Article 12 requires traceability for high-risk AI decisions, specifically logging that covers inputs, decision points, and risk-relevant events, and output logs alone are unlikely to satisfy that requirement. OWASP ranks prompt injection as the top vulnerability facing LLM-based applications, and in a prefetching architecture, that injection surface extends to anything the system anticipates the agent might need, not only to what a user explicitly requests.
Where the infrastructure metaphor for agentic retrieval is heading
Agents that retrieve continuously and anticipatorily are better understood through the model of a content delivery network than through the model of a search index. A search index assumes a human typing a query and waiting for a ranked list. A content delivery network assumes constant, anticipatory movement of content toward wherever it will be needed next, which is a much closer description of what predictive prefetching is actually doing.
That shift carries real consequences for how retrieval infrastructure gets built. Teams positioned to build production-grade agentic systems treat retrieval as a first-class engineering concern in its own right, owning the full pipeline from crawl through context delivery instead of wrapping a search endpoint designed for human typing patterns. Closing that gap takes infrastructure designed from the start around how agents actually retrieve, not infrastructure adapted after the fact from tools built for people.
Sources
- ICML Poster Predictive Prefetching for Retrieval-Augmented Generation
- [2605.17989] Predictive Prefetching for Retrieval-Augmented Generation
- Fishing for Answers: Exploring One-shot vs. Iterative Retrieval Strategies for Retrieval Augmented Generation
- Predictive Prefetching for Retrieval-Augmented Generation


