<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/ai21-trains-its-models-on-ai-hypercomputer-skzkrwjlk" -->

---
title: AI21 trains its models on AI Hypercomputer | daily.dev
description: AI21, the lab behind the Jamba model family, describes how it moved from manually negotiating GPU capacity in Slack to using Kueue on GKE for automated...
canonical: https://daily.dev/posts/ai21-trains-its-models-on-ai-hypercomputer-skzkrwjlk
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: AI21 trains its models on AI Hypercomputer | daily.dev
og:description: AI21, the lab behind the Jamba model family, describes how it moved from manually negotiating GPU capacity in Slack to using Kueue on GKE for automated...
og:url: https://daily.dev/posts/ai21-trains-its-models-on-ai-hypercomputer-skzkrwjlk
og:image: https://api.daily.dev/og/posts/SKZkrwJLk.png
og:image:alt: AI21 trains its models on AI Hypercomputer
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# AI21 trains its models on AI Hypercomputer

**[Google Cloud](https://daily.dev/sources/gcp)** · 6 min read · 0 upvotes · 0 comments

## Summary

AI21, the lab behind the Jamba model family, describes how it moved from manually negotiating GPU capacity in Slack to using Kueue on GKE for automated scheduling across its shared fleet of Google Cloud A3 and A3 Ultra (NVIDIA H100/H200) instances. The switch cut high-priority job wait times from 72 to 12 hours, eliminated roughly 20 weekly manual interventions, and reduced GPU fragmentation from 15% to 8%. AI21 evaluated Apache YuniKorn and Volcano before choosing Kueue for its simplicity and native Kubernetes integration, and became a design partner that helped shape new Kueue features like Admission Fair Sharing and Topology Aware Scheduling.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://cloud.google.com/blog/products/containers-kubernetes/ai21-trains-its-models-on-ai-hypercomputer>

## Questions this post answers

### How does Kueue's Admission Fair Sharing feature work on Kubernetes?

Admission Fair Sharing reorders the Kueue admission queue to prioritize teams that have historically consumed less cluster capacity, without preempting jobs already running. It was built after AI21 submitted production requirements to the Kueue team as a design partner, specifically to achieve fairness across multi-node jobs without the disruption that preemption-based fairness usually causes.

_Teams fine-tuning multi-tenant GPU scheduling can follow Kueue developments like this one on daily.dev._

### What GPU scheduler did AI21 choose for training Jamba models on Google Cloud and why?

AI21 chose Kueue over Apache YuniKorn and Volcano to schedule GPU workloads on GKE across pooled Google Cloud A3 and A3 Ultra instances. YuniKorn didn't cover all required use cases, and Volcano required replacing core Kubernetes scheduler components, while Kueue worked with standard Kubernetes and didn't force rewriting existing job specs.

_Developers comparing Kubernetes batch schedulers can track these tradeoffs on daily.dev._

### How much did Kueue reduce GPU job wait times and fragmentation for AI21's training cluster?

High-priority multi-node training jobs that previously waited up to 72 hours for contiguous capacity now wait about 12 hours after adopting Kueue on GKE. Manual scheduling interventions dropped from roughly 20 per week to zero, and GPU fragmentation across the cluster fell from 15% to 8%, while overall compute spend stayed unchanged since the fleet already ran near 100% utilization.

_Infrastructure engineers optimizing GPU cluster utilization can follow real-world results like this on daily.dev._

## Similar posts on daily.dev

- [Kubernetes Did Not Miss the AI Wave. It Absorbed It](https://daily.dev/posts/kubernetes-did-not-miss-the-ai-wave-it-absorbed-it-d4xqnmisz) · Cloud Native Now · 5 upvotes · 4 comments
- [GPU Job Queueing with Kueue and DRA: Scheduling AI Workloads Without Idle Capacity](https://daily.dev/posts/gpu-job-queueing-with-kueue-and-dra-scheduling-ai-workloads-without-idle-capacity-g1csd6qkw) · Cast AI · 0 upvotes · 0 comments
- [Tame Ray workloads on OpenShift AI with KubeRay and Kueue](https://daily.dev/posts/tame-ray-workloads-on-openshift-ai-with-kuberay-and-kueue-ebyhfxbjb) · Red Hat Developer · 1 upvotes · 0 comments
- [Building an AI factory on Kubernetes](https://daily.dev/posts/building-an-ai-factory-on-kubernetes-m4brxz3ty) · CNCF · 1 upvotes · 0 comments
- [Netflix Adopts Cloud-Native Job Queueing System Kueue to Replace an In-House Solution](https://daily.dev/posts/netflix-adopts-cloud-native-job-queueing-system-kueue-to-replace-an-in-house-solution-7gbkjcg1w) · InfoQ · 0 upvotes · 0 comments

---

Tags: [#kubernetes](https://daily.dev/tags/kubernetes), [#ai-infrastructure](https://daily.dev/tags/ai-infrastructure), [#gke](https://daily.dev/tags/gke)

[View this post on daily.dev](https://daily.dev/posts/ai21-trains-its-models-on-ai-hypercomputer-skzkrwjlk)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"AI21 trains its models on AI Hypercomputer","url":"https://daily.dev/posts/ai21-trains-its-models-on-ai-hypercomputer-skzkrwjlk","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/ai21-trains-its-models-on-ai-hypercomputer-skzkrwjlk"},"datePublished":"2026-10-02T19:23:18.608Z","dateModified":"2026-10-02T19:28:06.441Z","description":"AI21, the lab behind the Jamba model family, describes how it moved from manually negotiating GPU capacity in Slack to using Kueue on GKE for automated...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/71a924f2b121958fb9f73a189cb1663d?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/71a924f2b121958fb9f73a189cb1663d?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Google Cloud","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Google Cloud","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/gcp","url":"https://daily.dev/sources/gcp"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/ai21-trains-its-models-on-ai-hypercomputer-skzkrwjlk","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"kubernetes,ai-infrastructure,gke","timeRequired":"PT6M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Google Cloud","item":"https://daily.dev/sources/gcp"},{"@type":"ListItem","position":3,"name":"AI21 trains its models on AI Hypercomputer"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/ai21-trains-its-models-on-ai-hypercomputer-skzkrwjlk#faq","mainEntity":[{"@type":"Question","name":"How does Kueue's Admission Fair Sharing feature work on Kubernetes?","acceptedAnswer":{"@type":"Answer","text":"Admission Fair Sharing reorders the Kueue admission queue to prioritize teams that have historically consumed less cluster capacity, without preempting jobs already running. It was built after AI21 submitted production requirements to the Kueue team as a design partner, specifically to achieve fairness across multi-node jobs without the disruption that preemption-based fairness usually causes. Teams fine-tuning multi-tenant GPU scheduling can follow Kueue developments like this one on daily.dev."}},{"@type":"Question","name":"What GPU scheduler did AI21 choose for training Jamba models on Google Cloud and why?","acceptedAnswer":{"@type":"Answer","text":"AI21 chose Kueue over Apache YuniKorn and Volcano to schedule GPU workloads on GKE across pooled Google Cloud A3 and A3 Ultra instances. YuniKorn didn't cover all required use cases, and Volcano required replacing core Kubernetes scheduler components, while Kueue worked with standard Kubernetes and didn't force rewriting existing job specs. Developers comparing Kubernetes batch schedulers can track these tradeoffs on daily.dev."}},{"@type":"Question","name":"How much did Kueue reduce GPU job wait times and fragmentation for AI21's training cluster?","acceptedAnswer":{"@type":"Answer","text":"High-priority multi-node training jobs that previously waited up to 72 hours for contiguous capacity now wait about 12 hours after adopting Kueue on GKE. Manual scheduling interventions dropped from roughly 20 per week to zero, and GPU fragmentation across the cluster fell from 15% to 8%, while overall compute spend stayed unchanged since the fleet already ran near 100% utilization. Infrastructure engineers optimizing GPU cluster utilization can follow real-world results like this on daily.dev."}}]}
```

