AI agents have become the dominant source of token traffic on OpenRouter. In the last week of September, OpenRouter says agentic API keys used about seven times as many tokens as human-classified API keys—and nearly nine out of every ten agentic tokens were cached prompt reads.
The important part is not simply that “AI is using more AI.” The new OpenRouter telemetry published October 8 shows how differently autonomous agents use models: they repeatedly resend system prompts, tool definitions, files, conversation history, and prior tool results while taking many small steps toward a goal.

Agents keep rereading the same context
A normal chat often looks like one human question followed by one model answer. An agent behaves more like a loop. It reads the goal, checks tools, takes an action, reads the result, updates its state, and calls the model again. Each new model call can carry most of the previous context with it.
That is why raw token volume can be misleading. OpenRouter says cached prompt tokens were just under 90% of all agentic tokens in September. Novel prompt tokens made up most of the remainder, while output plus reasoning represented only about 1%.

Seven times the tokens does not mean seven times the bill
Cached tokens are much cheaper than fresh input or generated output because the provider can reuse prior computation instead of processing the same prompt prefix from scratch. OpenRouter’s example using current Claude Sonnet 5 pricing puts one million agentic tokens with an 87% cache share at about $0.51, versus roughly $2.08 if none of that context were cached.
So the headline token count is not a spending chart. It is better understood as a picture of workload shape. Agents are creating enormous volumes of repeated context, and the economics increasingly depend on whether providers can recognize and reuse that context efficiently.
That distinction matters alongside BitcoinVersus’ earlier look at model routing. Choosing a model is only part of the cost equation; routing a long-running agent between providers can also affect cache reuse, latency, and effective price.
The hardware story is memory, not only compute
KV cache stores the attention keys and values produced while a model processes previous tokens. Keeping that state available allows the model to continue generating without recomputing the entire sequence. The tradeoff is memory capacity and bandwidth: longer contexts, more simultaneous agents, and more active sessions all increase the amount of state that inference infrastructure has to manage.
That makes the trend relevant to the memory supply chain. BitcoinVersus recently covered Micron’s surge in AI-memory demand. Agentic workloads add another reason HBM capacity and memory bandwidth can matter just as much as raw accelerator FLOPS.
Cache performance has improved fast
OpenRouter’s September data also shows how aggressively inference providers have optimized for repeated context. Models launched in Q3 2026 averaged an 88.2% agentic cache rate, compared with 54.8% for models launched in Q2 2025.
Across providers, the median agentic cache rate rose from 44% in January to 85% in September. Eleven of 54 providers were above 90% in September, with some near 96%. Those numbers help explain why agent use can grow so quickly without token spending rising at exactly the same rate.
Agent activity is not the same thing as human adoption
There is an important limitation in the data. OpenRouter is one routing platform, not the entire AI market. It classifies API keys as agentic, human, or mixed using behavioral signals such as tool-call rate, turn count, and timing. The figures therefore describe traffic on OpenRouter, not a census of every AI request worldwide.
Tom’s Hardware’s earlier analysis used August data showing roughly 7.3 trillion agentic tokens versus 1.4 trillion human tokens and correctly noted that token volume should not be confused with spending. OpenRouter’s newer September update strengthens the direction of the trend while also making the caching explanation clearer.
Why this matters for software teams
The useful metric for an agent may be shifting from “how much does this model cost per million tokens?” toward a bundle of measurements: cache-hit rate, tokens per completed task, number of model turns, latency, tool-call success, and cost per useful outcome.
That also helps explain why more generated work does not automatically equal more finished work. BitcoinVersus recently reported that AI coding agents can produce more code while human review becomes the bottleneck. Agent infrastructure has a similar lesson: enormous token throughput is only valuable when it turns into completed, correct work.
The next AI bottleneck may be remembering efficiently
The newest OpenRouter data suggests the agent era is changing the shape of inference. Human chat emphasizes fresh questions and readable answers. Agents emphasize loops, tool calls, persistent context, and repeated reads.
That makes caching less of a background optimization and more of a core systems feature. If autonomous agents continue taking a larger share of model traffic, the winners may not simply be the systems with the most compute. They may be the ones that can keep enormous amounts of context close to the model, reuse it reliably, and turn repeated memory access into useful work at the lowest cost.
BitcoinVersus.Tech Editor’s Note: The 7× figure applies to OpenRouter’s classified weekly token traffic in the final week of September 2026. It does not mean agents globally use seven times more AI than humans, nor does seven times the token volume imply seven times the cost. Cached prompts account for nearly 90% of OpenRouter’s agentic traffic.
Follow BitcoinVersus.Tech on X for AI infrastructure, semiconductors, memory, networking, Bitcoin mining, and data-center reporting.
Support independent technology reporting: Bitcoin donations help fund BitcoinVersus.Tech research and publishing.
Disclaimer: BitcoinVersus.Tech provides technology news and analysis for informational purposes only.

Leave a Reply