<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/impactful-scheduling-for-gpu-clusters-0wgrntsww" -->

---
title: Impactful scheduling for GPU clusters | daily.dev
description: Ai2's AI Infrastructure team describes replacing a priority-based GPU scheduler with a new system combining GPU time budgets, hierarchical fair-share...
canonical: https://daily.dev/posts/impactful-scheduling-for-gpu-clusters-0wgrntsww
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Impactful scheduling for GPU clusters | daily.dev
og:description: Ai2's AI Infrastructure team describes replacing a priority-based GPU scheduler with a new system combining GPU time budgets, hierarchical fair-share...
og:url: https://daily.dev/posts/impactful-scheduling-for-gpu-clusters-0wgrntsww
og:image: https://api.daily.dev/og/posts/0wGrnTsww.png
og:image:alt: Impactful scheduling for GPU clusters
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Impactful scheduling for GPU clusters

**[Hugging Face](https://daily.dev/sources/huggingface)** · 17 min read · 2 upvotes · 1 comments

## Summary

Ai2's AI Infrastructure team describes replacing a priority-based GPU scheduler with a new system combining GPU time budgets, hierarchical fair-share allocation, and a time-slicing scheduling contract across thousands of H100, B200, and B300 GPUs. The old system caused squatting, priority inflation, and heavy on-call toil from negotiating shutdowns of non-preemptible jobs. The new approach treats GPU time like a budget allocated top-down by managers, paired with a fair-share scheduler (in the lineage of Hadoop Fair Scheduler, SLURM Fair Tree, YARN) and a minimum-runtime contract that lets workloads be safely preempted and requeued. After rollout, teams received 98% of owed GPU hours, occupancy held at 98%, debug workload p90 queue wait dropped from 2 hours to 30 seconds, and human-in-the-loop repairs fell 74%. Remaining challenges include interactive dev sessions suffering from preemption and emerging capacity fragmentation for large jobs, which are being addressed with a planned CPU-only cluster and restorable sessions.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://huggingface.co/blog/allenai/impactful-scheduling>

## Questions this post answers

### How did Ai2 reduce GPU squatting and priority inflation on their training clusters?

Ai2 replaced its priority-based GPU scheduler with a budget-based hierarchical fair-share system where every workload must be funded by a GPU time budget to be protected from preemption. Squatting now drains a team's own budget, so it provides no benefit, and priority only affects sorting within a team rather than across the whole cluster, eliminating the incentive for universal high-priority flagging.

_Teams weighing scheduler redesigns for GPU clusters can find engineering deep dives like this on daily.dev._

### What results did Ai2 see after rolling out its new GPU scheduling contract with minimum runtime guarantees?

Over a 30-day test period, teams received 98% of their owed GPU hours, cluster occupancy held steady at 98%, and debug workload p90 queue wait time fell from 2 hours to 30 seconds. Human-in-the-loop repairs dropped by 74% because unhealthy hosts could automatically drain workloads once they reached their 8-hour minimum runtime cap.

_Developers evaluating scheduling tradeoffs for infrastructure projects can follow results-driven writeups like this on daily.dev._

### What is the tradeoff of applying minimum runtime protection to interactive GPU dev sessions?

Capping protected runtime at 8 hours made interactive data-analysis and debugging sessions preemptible after that window, causing researchers to lose volatile session state and have to rebuild it by hand, a regression from the old system where sessions could run uninterrupted for up to a week. Ai2 is addressing this by building a separate CPU-only cluster with restorable sessions.

_Infrastructure teams balancing preemption policy against developer experience can track lessons like this on daily.dev._

## Community discussion

Top comments from developers on daily.dev.

**@kartiknvj** · 0 upvotes

> Priority inflation is the failure mode nobody plans for: give teams a priority knob and every job becomes P0 within a quarter. Time budgets plus hierarchical fair-share is the right fix, and the minimum-runtime contract is what makes preemption humane. Where does the fair-share tree sit, per team, per project, or per user, and who arbitrates when two orgs both sit under budget?

## Similar posts on daily.dev

- [Ensuring Balanced GPU Allocation in Kubernetes Clusters with Time-Based Fairshare](https://daily.dev/posts/ensuring-balanced-gpu-allocation-in-kubernetes-clusters-with-time-based-fairshare-t7sb4ftnr) · NVIDIA Developer · 0 upvotes · 0 comments
- [Red Hat AI 3.5 tackles the GPU queue that can stall AI pilots](https://daily.dev/posts/red-hat-ai-3-5-tackles-the-gpu-queue-that-can-stall-ai-pilots-zw7qaelqv) · The New Stack · 0 upvotes · 0 comments
- [Same Cluster, 33 Points More Utilization: What Changed Was the Order](https://daily.dev/posts/same-cluster-33-points-more-utilization-what-changed-was-the-order-ksmb8mjvo) · Hugging Face · 0 upvotes · 0 comments
- [GPU Scheduling and Bin-Packing in Kubernetes: Pack More AI onto Every GPU](https://daily.dev/posts/gpu-scheduling-and-bin-packing-in-kubernetes-pack-more-ai-onto-every-gpu-6nuaafnff) · Cast AI · 0 upvotes · 0 comments
- [You’re probably underutilizing your GPUs](https://daily.dev/posts/you-re-probably-underutilizing-your-gpus-ei9uypxwp) · Stack Overflow Blog · 0 upvotes · 0 comments

---

Tags: [#machine-learning](https://daily.dev/tags/machine-learning), [#kubernetes](https://daily.dev/tags/kubernetes)

[View this post on daily.dev](https://daily.dev/posts/impactful-scheduling-for-gpu-clusters-0wgrntsww)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Impactful scheduling for GPU clusters","url":"https://daily.dev/posts/impactful-scheduling-for-gpu-clusters-0wgrntsww","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/impactful-scheduling-for-gpu-clusters-0wgrntsww"},"datePublished":"2026-10-09T15:25:15.625Z","dateModified":"2026-10-09T15:35:35.301Z","description":"Ai2's AI Infrastructure team describes replacing a priority-based GPU scheduler with a new system combining GPU time budgets, hierarchical fair-share...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/c1682fccca5731ced50f18015144133a?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/c1682fccca5731ced50f18015144133a?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Hugging Face","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Hugging Face","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/f1f55c67d81a4330acf5b90b26b0c8e1","url":"https://daily.dev/sources/huggingface"},"commentCount":1,"discussionUrl":"https://daily.dev/posts/impactful-scheduling-for-gpu-clusters-0wgrntsww","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":2},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":1}],"keywords":"machine-learning,kubernetes","timeRequired":"PT17M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Hugging Face","item":"https://daily.dev/sources/huggingface"},{"@type":"ListItem","position":3,"name":"Impactful scheduling for GPU clusters"}]}
{"@context":"https://schema.org","@type":"WebPage","@id":"https://daily.dev/posts/impactful-scheduling-for-gpu-clusters-0wgrntsww","comment":[{"@type":"Comment","text":"Priority inflation is the failure mode nobody plans for: give teams a priority knob and every job becomes P0 within a quarter. Time budgets plus hierarchical fair-share is the right fix, and the minimum-runtime contract is what makes preemption humane. Where does the fair-share tree sit, per team, per project, or per user, and who arbitrates when two orgs both sit under budget?","datePublished":"2026-10-09T19:11:04.629Z","url":"https://daily.dev/posts/0wGrnTsww#c-2l3N3GISM","author":{"@type":"Person","name":"kartik-nvjk","url":"https://daily.dev/kartiknvj","image":"https://media.daily.dev/image/upload/s--3gGgsVCw--/f_auto/v1781456774/avatars/avatar_TvTVeiMdkRCqWUDullFmy?_a=BAMAMiWQ0"}}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/impactful-scheduling-for-gpu-clusters-0wgrntsww#faq","mainEntity":[{"@type":"Question","name":"How did Ai2 reduce GPU squatting and priority inflation on their training clusters?","acceptedAnswer":{"@type":"Answer","text":"Ai2 replaced its priority-based GPU scheduler with a budget-based hierarchical fair-share system where every workload must be funded by a GPU time budget to be protected from preemption. Squatting now drains a team's own budget, so it provides no benefit, and priority only affects sorting within a team rather than across the whole cluster, eliminating the incentive for universal high-priority flagging. Teams weighing scheduler redesigns for GPU clusters can find engineering deep dives like this on daily.dev."}},{"@type":"Question","name":"What results did Ai2 see after rolling out its new GPU scheduling contract with minimum runtime guarantees?","acceptedAnswer":{"@type":"Answer","text":"Over a 30-day test period, teams received 98% of their owed GPU hours, cluster occupancy held steady at 98%, and debug workload p90 queue wait time fell from 2 hours to 30 seconds. Human-in-the-loop repairs dropped by 74% because unhealthy hosts could automatically drain workloads once they reached their 8-hour minimum runtime cap. Developers evaluating scheduling tradeoffs for infrastructure projects can follow results-driven writeups like this on daily.dev."}},{"@type":"Question","name":"What is the tradeoff of applying minimum runtime protection to interactive GPU dev sessions?","acceptedAnswer":{"@type":"Answer","text":"Capping protected runtime at 8 hours made interactive data-analysis and debugging sessions preemptible after that window, causing researchers to lose volatile session state and have to rebuild it by hand, a regression from the old system where sessions could run uninterrupted for up to a week. Ai2 is addressing this by building a separate CPU-only cluster with restorable sessions. Infrastructure teams balancing preemption policy against developer experience can track lessons like this on daily.dev."}}]}
```

