Where an agent earns its tokens

Four pipelines over the same 150 Olympic questions, on one TigerGraph graph, with Qwen3.8-27B serving every pipeline including the planner and the verifier. Only LLM tokens are counted; GSQL and Python cost nothing.

The result in one chart

Accuracy against LLM tokens per question. The interesting property is not that the top-left corner is occupied, it is how little horizontal distance separates the agentic pipeline from the zero-token control.

Plain retrieval sits far right and low. The graph removes most of the gap. The agent removes the rest, and lands left of both — it is the cheapest of the three pipelines that use a model at all.

The four pipelines

P4 is the control, not a competitor: it reaches full marks at zero tokens and therefore reports which questions never needed an agent. Every fallback run is labelled deterministic-fallback so it can never be mistaken for a routing decision.

Where each pipeline fails

The 150 questions come from five templates, and the split between them is the argument. Chunk retrieval collapses on exactly the two templates that need arithmetic over a set, and is respectable on the single-document one.

Tests

The regression suite is stdlib unittest and needs neither TigerGraph nor the LLM endpoint. It was executed while this page was built, so the panel reflects a real run rather than a claim.

It encodes the data traps this corpus hides — singular/plural infobox twins, fused multi-venue strings, eight date formats, near-duplicate event names, multi-athlete medal fields — and the scoring rules the organiser stated in prose.

It has already caught diacritics not folded when scoring, unranked supersession clauses never extracted, abbreviated month names yielding no date tokens, and a value extractor that returned a whole sentence where a bare number was due.

Per-question record

Every figure above traces back to these records. The hidden split is the one the organiser scores server-side; it is included in full.