Four pipelines over the same 150 Olympic questions, on one TigerGraph graph, with Qwen3.8-27B serving every pipeline including the planner and the verifier. Only LLM tokens are counted; GSQL and Python cost nothing.
Accuracy against LLM tokens per question. The interesting property is not that the top-left corner is occupied, it is how little horizontal distance separates the agentic pipeline from the zero-token control.
Plain retrieval sits far right and low. The graph removes most of the gap. The agent removes the rest, and lands left of both — it is the cheapest of the three pipelines that use a model at all.
P4 is the control, not a competitor: it reaches full marks at zero
tokens and therefore reports which questions never needed an agent. Every fallback
run is labelled deterministic-fallback so it can never be mistaken for a
routing decision.
The 150 questions come from five templates, and the split between them is the argument. Chunk retrieval collapses on exactly the two templates that need arithmetic over a set, and is respectable on the single-document one.
The regression suite is stdlib unittest and needs
neither TigerGraph nor the LLM endpoint. It was executed while this page was
built, so the panel reflects a real run rather than a claim.
Every figure above traces back to these records. The hidden split is the one the organiser scores server-side; it is included in full.