<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/getting-more-from-each-token-how-copilot-improves-context-handling-and-model-routing-sh4tcvdez" -->

---
title: Getting more from each token: How Copilot improves...
description: GitHub Copilot is improving token efficiency through two main mechanisms: prompt caching and deferred tool loading in VS Code, and Auto model selection that...
canonical: https://daily.dev/posts/getting-more-from-each-token-how-copilot-improves-context-handling-and-model-routing-sh4tcvdez
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Getting more from each token: How Copilot improves context handling and model routing | daily.dev
og:description: GitHub Copilot is improving token efficiency through two main mechanisms: prompt caching and deferred tool loading in VS Code, and Auto model selection that...
og:url: https://daily.dev/posts/getting-more-from-each-token-how-copilot-improves-context-handling-and-model-routing-sh4tcvdez
og:image: https://api.daily.dev/og/posts/Sh4TcvdeZ.png
og:image:alt: Getting more from each token: How Copilot improves context handling and model routing
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Getting more from each token: How Copilot improves context handling and model routing

**[GitHub Blog](https://daily.dev/sources/ghblog)** · 8 min read · 0 upvotes · 0 comments

## Summary

GitHub Copilot is improving token efficiency through two main mechanisms: prompt caching and deferred tool loading in VS Code, and Auto model selection that routes tasks to the most appropriate model. Prompt caching reuses model state for repeated prompt prefixes, while tool search loads tool definitions on demand rather than sending all schemas upfront. The Auto feature uses HyDRA, a routing model that considers task complexity, reasoning depth, and real-time model health to pick the best-fit model without requiring manual selection. Auto is cache-aware, avoiding mid-conversation model switches that would break cached prefixes. It supports 16 language families with routing accuracy within four points of the English baseline. Auto is expanding to Copilot CLI, GitHub App, and additional IDEs, and will become the only model option for Free and Student plans. Practical tips include keeping context focused, avoiding mid-session model changes, planning before using parallel agents, and limiting enabled tools to what's needed.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://github.blog/ai-and-ml/github-copilot/getting-more-from-each-token-how-copilot-improves-context-handling-and-model-routing>

## Questions this post answers

### How does GitHub Copilot's Auto model selection decide which model to use for a task?

Auto combines two signals: real-time model health (availability, utilization, speed, error rates, cost) and task-aware routing via a model called HyDRA that evaluates reasoning depth, code complexity, debugging difficulty, and tool orchestration needs. HyDRA identifies models meeting the quality bar for a task, then picks the best fit among them, rather than always defaulting to the biggest or cheapest model.

_Anyone tuning AI coding workflows for cost and quality can track routing changes like this via daily.dev._

### Why does switching AI models mid-conversation in GitHub Copilot hurt performance or cost more?

Switching models mid-conversation breaks the cached prompt prefix, forcing Copilot to recompute context that would otherwise be reused across turns, which can cost more than any savings from the routing change. Copilot's Auto feature instead routes only at natural cache boundaries: on the first turn and after conversation compaction, keeping the same model in place between those points so the cache keeps building.

_Developers watching their AI credit spend can follow efficiency changes like this through daily.dev._

### How does GitHub Copilot's routing model handle non-English language conversations?

The routing model was trained on conversations across 16 language families, including CJK and European languages. In evaluations, routing accuracy stayed within four points of the English baseline across language groups, with no statistically significant quality gap, meaning model selection quality holds up regardless of the developer's working language.

_Teams evaluating multilingual AI coding tool reliability can track updates like this on daily.dev._

## Similar posts on daily.dev

- [Auto model selection now routes based on your task in VS Code](https://daily.dev/posts/auto-model-selection-now-routes-based-on-your-task-in-vs-code-gzeadlpau) · GitHub Changelog · 0 upvotes · 0 comments
- [Github Copilot–Auto model select](https://daily.dev/posts/github-copilot-auto-model-select-rsl2szxzb) · The Art of Simplicity · 2 upvotes · 0 comments

---

Tags: [#github](https://daily.dev/tags/github), [#ai-agents](https://daily.dev/tags/ai-agents), [#ai-inference](https://daily.dev/tags/ai-inference), [#context-engineering](https://daily.dev/tags/context-engineering), [#ai-gateway](https://daily.dev/tags/ai-gateway)

[View this post on daily.dev](https://daily.dev/posts/getting-more-from-each-token-how-copilot-improves-context-handling-and-model-routing-sh4tcvdez)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Getting more from each token: How Copilot improves context handling and model routing","url":"https://daily.dev/posts/getting-more-from-each-token-how-copilot-improves-context-handling-and-model-routing-sh4tcvdez","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/getting-more-from-each-token-how-copilot-improves-context-handling-and-model-routing-sh4tcvdez"},"datePublished":"2026-06-17T19:42:03.962Z","dateModified":"2026-09-13T20:50:07.589Z","description":"GitHub Copilot is improving token efficiency through two main mechanisms: prompt caching and deferred tool loading in VS Code, and Auto model selection that...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/01832fd6da8a7628e6ac776f390f661b?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/01832fd6da8a7628e6ac776f390f661b?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"GitHub Blog","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"GitHub Blog","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/106cf162b88840808484d4b5429b59b1","url":"https://daily.dev/sources/ghblog"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/getting-more-from-each-token-how-copilot-improves-context-handling-and-model-routing-sh4tcvdez","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"github,ai-agents,ai-inference,context-engineering,ai-gateway","timeRequired":"PT8M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"GitHub Blog","item":"https://daily.dev/sources/ghblog"},{"@type":"ListItem","position":3,"name":"Getting more from each token: How Copilot improves context handling and model routing"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/getting-more-from-each-token-how-copilot-improves-context-handling-and-model-routing-sh4tcvdez#faq","mainEntity":[{"@type":"Question","name":"How does GitHub Copilot's Auto model selection decide which model to use for a task?","acceptedAnswer":{"@type":"Answer","text":"Auto combines two signals: real-time model health (availability, utilization, speed, error rates, cost) and task-aware routing via a model called HyDRA that evaluates reasoning depth, code complexity, debugging difficulty, and tool orchestration needs. HyDRA identifies models meeting the quality bar for a task, then picks the best fit among them, rather than always defaulting to the biggest or cheapest model. Anyone tuning AI coding workflows for cost and quality can track routing changes like this via daily.dev."}},{"@type":"Question","name":"Why does switching AI models mid-conversation in GitHub Copilot hurt performance or cost more?","acceptedAnswer":{"@type":"Answer","text":"Switching models mid-conversation breaks the cached prompt prefix, forcing Copilot to recompute context that would otherwise be reused across turns, which can cost more than any savings from the routing change. Copilot's Auto feature instead routes only at natural cache boundaries: on the first turn and after conversation compaction, keeping the same model in place between those points so the cache keeps building. Developers watching their AI credit spend can follow efficiency changes like this through daily.dev."}},{"@type":"Question","name":"How does GitHub Copilot's routing model handle non-English language conversations?","acceptedAnswer":{"@type":"Answer","text":"The routing model was trained on conversations across 16 language families, including CJK and European languages. In evaluations, routing accuracy stayed within four points of the English baseline across language groups, with no statistically significant quality gap, meaning model selection quality holds up regardless of the developer's working language. Teams evaluating multilingual AI coding tool reliability can track updates like this on daily.dev."}}]}
```

