Search

inference latency

  1. 01

    Route cheap tokens to small models and cut projected inference latency

    A per-token compute map shows a 0.5B model reproduces most tokens and routing reduced projected latency from 7.59 to 5.12 seconds on MATH-500, concentrating most compute on the hardest 10%.

    2026-10-11 1 min Drafted gpt-5-mini
  2. 02

    Fine-tune small search agents to lower model costs and latency

    Fine-tune a small LLM with multi-turn RL on SageMaker AI to get frontier-like search reliability with lower inference latency and cost.

    2026-10-09 1 min Drafted gpt-5-mini