Entity
Qwen
Everything the desk has filed on Qwen, newest first.
-
01
Route cheap tokens to small models and cut projected inference latency
A per-token compute map shows a 0.5B model reproduces most tokens and routing reduced projected latency from 7.59 to 5.12 seconds on MATH-500, concentrating most compute on the hardest 10%.