How We Rebuilt Citations for an Agentic Pipeline
As our agents gained more retrieval paths, source numbers could drift from the evidence. We moved citation binding into retrieval.
Contents
In one production conversation, our agent found the right products and gave the right specifications. Every citation on those claims opened an unrelated datasheet. The answer looked verified, but the links pointed elsewhere.
We introduced inline citations earlier this year so readers could check an answer against its source. The first design fit a retrieval-augmented generation (RAG) pipeline: retrieve the documents, number them, then generate an answer from that complete set.
Our agent now retrieves information throughout a turn: product data, documentation, fallback searches, sometimes in parallel. Citation numbers still arrived at the end, in a separate source list.
When the content and number took different paths
Our first agentic citation system stored retrieved sources in a per-turn registry. Tools returned content to the agent; the registry produced a numbered source list when the agent was ready to answer. The model had to match earlier tool results to that later list.
In the reported case, a catalog search hit its per-turn limit. Another retrieval path returned the correct products, but recorded them only for grounding checks. They never received citation numbers. The final source list contained documents from an earlier lookup, so the agent used those numbers for its product claims. The resolver then linked the claims to the unrelated documents those numbers actually named.
Put the number beside the evidence
Now a retrieval tool registers each citable source as it returns it. The registry gives that source a number, and the tool places the number next to the content the agent reads.
A product table has a Cite column. A structured result has a cite field. A text block has a [cite: N] heading. The agent sees each number with its source, even when several tools contribute to the answer. At answer time it still receives citation rules, but no separate source list.
The per-turn registry remains: we use its record of retrieved documents to evaluate whether the answer is grounded in the evidence. The main agent already has those documents in its tool results, so we skip the duplicate source list, saving tokens and reducing the context it has to process.
The citation processor keeps the same contract: <citation>7</citation> resolves to source 7 in that turn’s registry. It then adds a source link and, where possible, a deep link to the quoted passage. Channel renderers use that result for web, Slack, email, Discord, Discourse, and the Chrome extension, wherever a Rapidflare agent can connect.
Some evidence has no public citation
The registry keeps two views. all_docs contains the evidence the agent retrieved and feeds our grounding checks. docs contains sources that may receive citation numbers. A document without a URL, a private or institutional source, or a glossary definition can inform an answer without becoming a public link.
The tool leaves the number off those results. Our citation rules tell the agent to state an uncitable result without a citation tag. It must not borrow a number from another item.
Parallel retrieval needs stable numbering
Tools can run concurrently, and one catalog query can search several catalogs at once. The registry therefore assigns numbers under a lock. Repeated retrieval of the same citable source returns its existing number. A new source gets the next one. The number beside a result always names the source the registry will resolve later.
Cached citation numbers can point to the wrong source
Numbers restart each turn. A cached tool result could replay cite: 4 when source 4 in the new turn is a different document. The cache would also skip the tool body, leaving the cached source out of the new registry entirely.
We disabled tool-result caching for tools that write to the registry. Other tools can still use the cache. Caching these retrieval results would require registering the sources and assigning numbers again in the current turn. The cost of losing cache hits on repeated production questions remains unmeasured.
What we measured
We ran 293 paired cases from three production datasets before shipping the change. Correctness, faithfulness, and completeness showed no detected regression. Claims found in the document they cited rose from 66.7% to 68.6%. We also saw 70 more grounded citations and 13 more answers with a citation.
Those citation changes are directional evidence, not a guaranteed lift. On one dataset, two runs of unchanged code differed by 32% in the number of citations emitted. That variation was larger than the gap between the old and new designs.
Removing the final source list cut citation-numbering prompt content from about 354,000 to 28,600 tokens across the 293 turns. The old list repeated source identities and URLs after the model had already read the associated tool results. Median latency changed by +1.4, +0.4, and -0.9 seconds across the three datasets, on turns that typically took 45 to 90 seconds.
Our production-case regression suite made this comparison possible. We have written more about building source-grounded answer keys and testing the LLM judge itself.
The model still chooses which citation number to emit. The resolver confirms that the number names a source, but does not prove that source supports the claim. This change also does not verify that the agent searched every source before saying a product is absent. The number now travels with each result; claim-level verification and search coverage remain separate problems.
About the author
Founding Engineer at Rapidflare working on generative AI. Computer Science graduate from Arizona State University (Dean's List), where he researched large language models in the ARC Lab and built computer-vision and OCR systems for logistics automation.