AI agents are now consuming roughly five times as many tokens as human-driven traffic on OpenRouter, turning autonomous software into a rapidly growing source of inference demand—and making memory, not just compute, a bigger infrastructure constraint.
According to OpenRouter’s own aggregated usage data, agentic workloads overtook human traffic earlier this year and have continued accelerating. A current analysis highlighted by Tom’s Hardware puts agent traffic at about 7.3 trillion tokens versus 1.4 trillion for human-driven use in the measured period.
The headline is not simply that agents generate more text. The deeper infrastructure issue is that agents repeatedly reread long contexts, pushing cached-token volume and KV-cache memory requirements upward even when those tokens are cheaper to serve than uncached prompts.
Agents crossed humans and kept accelerating
OpenRouter classifies API keys as agentic, mixed, or human using a weighted set of behavioral signals that includes tool-call rates, turn counts, timing gaps, and other usage patterns. Its data therefore measures activity on one major AI routing platform—not every AI interaction on the internet—but it offers a large real-world view of how production workloads are changing.
Futurum Group CEO Daniel Newman summarized the trend in a widely shared Twitter post, arguing that AI systems are increasingly becoming consumers of AI themselves as agents call models, tools, and other agents in loops.
The trend fits BitcoinVersus.tech’s recent coverage of OpenAI’s always-on Dots agents. Software that keeps working after the user stops typing naturally creates more model calls, more context replay, and more background inference than a conventional one-question chatbot.
More than 85% of agent tokens are cached prompts
The most important number may be the composition of that traffic. Independent reporting on the OpenRouter figures says more than 85% of agent tokens are coming from cached prompts.
That means agents are spending much of their token volume rereading context they have already seen: conversation history, instructions, tool outputs, code, documents, system state, and other persistent information required to continue a long-running task.
Cached tokens can dramatically reduce repeated computation, but they do not disappear from the hardware problem. During inference, transformer systems maintain key-value data for previous tokens so the model does not recompute the entire attention history every time it generates the next token. That retained state is the KV cache.
KV cache turns context into a memory-capacity problem
The longer an agent runs, the more state it may need to preserve or repeatedly reintroduce. Multiple concurrent agents multiply that pressure. A server can therefore have enough raw accelerator compute for more requests while still running into memory-capacity or memory-bandwidth limits.
That helps explain why AI infrastructure increasingly revolves around the whole serving stack rather than just GPU arithmetic. BitcoinVersus.tech recently covered Cloudflare Clef’s push for 38.8-millisecond agent decisions, another example of inference architecture being optimized around latency, routing, and repeated machine-to-machine calls.
OpenRouter’s Head of Insights Peter Walker captured the crossover earlier in the year, noting that February 6 may have been the last day on the platform when human-driven traffic consumed more tokens than agents.
Token volume is not the same thing as spending
The 5× comparison needs an important caveat: token volume is not the same as cost or revenue. Cached-token reads are generally cheaper than processing the same input from scratch, and OpenRouter’s figures describe its own platform rather than the entire AI market.
Still, lower token prices do not erase physical resource use. Context has to be stored, moved, scheduled, and served. As agent concurrency rises, the infrastructure challenge shifts toward fitting more active contexts into memory while keeping latency predictable.
That efficiency race is also visible in BitcoinVersus.tech’s coverage of GPT-6.1 Sol’s push toward cheaper high-performance inference. Lower per-request economics can stimulate even more usage, especially when software agents can generate requests continuously without waiting for a person.
The next AI user may be another AI
Human adoption remains important, but autonomous software changes the scaling equation. A single person can supervise multiple agents, and each agent can produce a chain of model calls far larger than the user’s original prompt.
That means future AI capacity planning may need to forecast not only how many people use a service, but how many machine-generated calls each human request can trigger, how much context those agents preserve, and how long that state remains resident in accelerator memory.
If OpenRouter’s trajectory holds, AI infrastructure is moving into a phase where software agents become major customers of the compute stack themselves—and the race for memory capacity may become just as consequential as the race for more FLOPS.
BitcoinVersus.Tech
Advertisement
BitcoinVersus.Tech Editor’s Note:
We volunteer daily to ensure the credibility of the information on this platform is Verifiably True. If you would like to support to help further secure the integrity of our research initiatives, please donate here: 3C9o19EH5HSiwEPyCTmEKzxhNCbo2X6TTb
BitcoinVersus.tech is not a financial advisor. This media platform reports on financial subjects purely for informational purposes.

Leave a comment