Route cheap tokens to small models and cut projected inference latency
A per-token compute map shows a 0.5B model reproduces most tokens and routing reduced projected latency from 7.59 to 5.12 seconds on MATH-500, concentrating most compute on the hardest 10%.
AI generated — machine-made illustration, not a photograph of the event.
Routing tokens by how much compute they actually need can reduce inference latency and concentrate cloud spend on a small fraction of expensive tokens, lowering time and cost risk for teams that can run multi-model routing.
What actually changed
Researchers measured how much inference compute individual tokens really need using a Mixture-of-Agents (MoA) setup: a panel of 15 language models from three families (Qwen, OLMo and R1-distilled). Each model attempted to reproduce a reference sequence token by token, conditioned on the correct prefix. They defined a token's sufficient compute as the smallest agent that succeeds at reproducing it. The measurement found that a 0.5B agent reproduces between 92% and 95% of reference tokens on three core benchmarks. Across the model panels, the top 10% most expensive tokens account for roughly 64–80% of estimated FLOPs. On all 500 items in the MATH-500 benchmark, using the MoA-derived map for model routing reduced projected latency from 7.59 seconds to 5.12 seconds and slightly improved accuracy compared with the best confidence-routing baseline stated in the paper.
Who it affects
- Teams running latency-sensitive inference where a single large model currently handles every token.
- Teams using or evaluating Qwen-family models, since Qwen is one of the families included in the panel.
- Teams that already experiment with model routing or speculative decoding: the paper provides a measured per-token signal that those systems can use to make routing decisions.
This most directly matters for setups where you can host multiple model sizes and add a routing decision rather than a static single-model deployment.
What it costs or what it replaces
The announcement does not state pricing. It does, however, quantify what would be replaced or reduced operationally:
- Replaces a uniform, single-model inference approach (equal compute per token) with a routing strategy that runs smaller models for most tokens and routes only harder tokens to larger models.
- Replaces or improves on a confidence-routing baseline used in the paper: the MoA-derived map reduced projected latency from 7.59s to 5.12s on MATH-500 and
- arXiv cs.AI — original reporting
Links above go to the original publisher. Signalcraft states the consequence; it does not reproduce their text.
0 comments
No comments yet. If you have run any of this, that is the most useful thing you could add.