Amazon SageMaker AI now supports multi-turn reinforcement learning (RL), a serverless model customization technique for fine-tuning models on multi-step agentic tasks. It extends existing fine-tuning options (RLVR, RLAIF) by training models against custom agent environments and rewarding full decision sequences across a task. This enables smaller, lower-cost models to match or exceed larger general-purpose models on targeted workloads. SageMaker manages the full training loop including rollout orchestration, trajectory collection, and checkpoint management. It integrates with Amazon Bedrock AgentCore Runtime, EKS, EC2, Fargate, and other infrastructure. Built-in MLflow tracking and evaluation metrics (reward, pass@k, trajectory) are included. The feature is fully serverless with pay-per-token pricing, available via SageMaker Studio and the Python SDK, supporting models like Qwen 3.6 27B, Nova Lite 2.0, GPT-OSS-20B, and Gemma 31B.

2m read timeFrom aws.amazon.com
Post cover image
117 Impressions