Search

multi-turn rl

  1. 01

    Fine-tune small search agents to lower model costs and latency

    Fine-tune a small LLM with multi-turn RL on SageMaker AI to get frontier-like search reliability with lower inference latency and cost.

    2026-10-09 1 min Drafted gpt-5-mini