Google Developers
Read post

Run Ray on TPU, Part 2: Ray AI libraries

Part 2 of a guide on running Ray on Google Cloud TPUs covers the three main AI libraries: Ray Serve for LLM inference (including multi-host tensor-parallel models via the topology field), Ray Data's iter_jax_batches() for device-sharded JAX input pipelines, and JaxTrainer for distributed JAX training with checkpointing and fault tolerance. Key gotchas include importing JAX inside the worker function, using topology instead of chip counts, and setting accelerator_config.topology in Serve to avoid silent multi-host deployment failures. Official rayproject/ray:-tpu Docker images and TPU utilization metrics in the Ray Dashboard are also now available.

    #machine-learning#kubernetes
Jul 24•7m read time•From developers.googleblog.com
Post cover image
Table of contents
RecapRay Serve on TPURay Data on TPU: feeding the accelerators with iter_jax_batchesTwo final extras: TPU Docker images and dashboard metricsIn SummaryWhat's nextAdditional resources
117 Impressions
Google Developers's image
Google Developers

GoogleDevs' platform is a central hub for developers interested in Google technologies, APIs, and de...

817 Followers

•

1.8K Upvotes

Would you recommend this post?

Copy link
WhatsApp
Facebook
X
New Squad
  • © 2026 Daily Dev Ltd.
  • Guidelines
  • Explore
  • Tags
  • Sources
  • Squads
  • Leaderboard