Researchers introduce Avengers-Pro, a test-time routing framework that dynamically assigns queries to different LLMs based on performance-efficiency scores. The system embeds and clusters incoming queries, then routes each to the most suitable model from an ensemble of varying capacities. Testing across 6 benchmarks and 8 leading models including GPT-5-medium, Gemini-2.5-pro, and Claude-opus-4.1, Avengers-Pro achieves 7% higher accuracy than the strongest single model while reducing costs by up to 63% for comparable performance.
218 Impressions