<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/accelerating-spatio-temporal-attention-for-video-diffusion-on-tpus-by8crmpum" -->

---
title: Accelerating Spatio-Temporal Attention for Video...
description: Video diffusion models spend a growing share of latency in self-attention as resolution and sequence length increase, reaching 88.2% of per-layer latency at...
canonical: https://daily.dev/posts/accelerating-spatio-temporal-attention-for-video-diffusion-on-tpus-by8crmpum
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Accelerating Spatio-Temporal Attention for Video Diffusion on TPUs | daily.dev
og:description: Video diffusion models spend a growing share of latency in self-attention as resolution and sequence length increase, reaching 88.2% of per-layer latency at...
og:url: https://daily.dev/posts/accelerating-spatio-temporal-attention-for-video-diffusion-on-tpus-by8crmpum
og:image: https://api.daily.dev/og/posts/By8crMpuM.png
og:image:alt: Accelerating Spatio-Temporal Attention for Video Diffusion on TPUs
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Accelerating Spatio-Temporal Attention for Video Diffusion on TPUs

**[Google Developers](https://daily.dev/sources/googledevs)** · 11 min read · 0 upvotes · 0 comments

## Summary

Video diffusion models spend a growing share of latency in self-attention as resolution and sequence length increase, reaching 88.2% of per-layer latency at 1440p. Building on Sparse VideoGen's spatial/temporal head routing, engineers optimized TPU Pallas Splash Attention kernels through three iterations: naive sparse traversal (actually slower than dense), full/boundary tile specialization (31% faster than dense), and tile-aligned sparse traversal that rounds mask boundaries to tile edges, reaching a 2.40x speedup over dense attention. They also addressed distributed inference overhead by relocating token permutation for temporal heads to after the head-local all-to-all exchange, cutting attention latency by 38% in one experiment. Combined, these optimizations produced end-to-end denoising speedups of 1.28x at 720p, 1.49x at 1080p, and 1.69x at 1440p, saving over 16 minutes per generated video at the highest resolution while maintaining comparable PSNR quality.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://developers.googleblog.com/accelerating-spatio-temporal-attention-for-video-diffusion-on-tpus>

## Questions this post answers

### How much speedup can tile-aligned sparse attention give over dense Splash attention on TPU v6e?

A tile-aligned sparse traversal implementation achieved a 2.40x speedup over dense Splash attention, reducing latency to 32.76 ms compared to 78.70 ms for dense attention. This was achieved by aligning the sparse mask boundaries with tile boundaries, reducing the fraction of tiles requiring intra-tile masking from 27.55% to just 2.22%, with masking needed only for padding at sequence edges.

_Teams tuning TPU inference kernels can track sparse attention optimization techniques like this on daily.dev._

### Why does naive sparse attention sometimes run slower than dense attention on TPUs?

A naive sparse attention kernel that retained only 39% of query-key pairs actually ran 22% slower than dense attention (96.37 ms versus 78.70 ms), because it still evaluated the exact mask inside every visited tile, including boundary tiles containing both retained and excluded pairs. This elementwise masking bottlenecks the Vector Processing Unit and stalls the Matrix Multiply Unit even when a tile is fully valid.

_Engineers debugging unexpected sparse-kernel slowdowns can follow hardware-level attention optimization writeups on daily.dev._

### How much does sparse attention speed up video diffusion denoising at 1440p versus 720p resolution?

At 1440p (302K tokens), an aggressive sparse attention schedule cut denoising time from 2,471 seconds to 1,461 seconds, a 1.69x speedup, saving over 16 minutes per video while maintaining 24.05 dB PSNR. At 720p the same approach gave only a 1.28x speedup, because attention's share of per-layer latency grows from 55.5% at 720p to 88.2% at 1440p.

_Developers choosing resolution and latency trade-offs for video generation can follow these benchmarks on daily.dev._

## Similar posts on daily.dev

- [thu-ml/TurboDiffusion: TurboDiffusion: 100–200× Acceleration for Video Diffusion Models](https://daily.dev/posts/thu-ml-turbodiffusion-turbodiffusion-100-200-acceleration-for-video-diffusion-models-jfiqr69ii) · Hacker News · 2 upvotes · 0 comments
- [Tuning Flash Attention for Peak Performance in NVIDIA CUDA Tile](https://daily.dev/posts/tuning-flash-attention-for-peak-performance-in-nvidia-cuda-tile-ytpdwf7wh) · NVIDIA Developer · 0 upvotes · 0 comments
- [HeyGen x Google Cloud: Bringing Avatar IV to TPUs](https://daily.dev/posts/heygen-x-google-cloud-bringing-avatar-iv-to-tpus-qzakruxj5) · Google Developers · 0 upvotes · 0 comments
- [Generalized Dot-Product Attention: Tackling Real-World Challenges in GPU Training Kernels – PyTorch](https://daily.dev/posts/generalized-dot-product-attention-tackling-real-world-challenges-in-gpu-training-kernels-pytorch-phuahr4yf) · PyTorch · 0 upvotes · 0 comments
- [MiniMax H3 on vLLM-Omni: From System-Wide Optimization to Real-Time Serving with FastVideo’s FastH3](https://daily.dev/posts/minimax-h3-on-vllm-omni-from-system-wide-optimization-to-real-time-serving-with-fastvideo-s-fasth3-mhkh8zrjp) · vLLM · 0 upvotes · 0 comments

---

[View this post on daily.dev](https://daily.dev/posts/accelerating-spatio-temporal-attention-for-video-diffusion-on-tpus-by8crmpum)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Accelerating Spatio-Temporal Attention for Video Diffusion on TPUs","url":"https://daily.dev/posts/accelerating-spatio-temporal-attention-for-video-diffusion-on-tpus-by8crmpum","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/accelerating-spatio-temporal-attention-for-video-diffusion-on-tpus-by8crmpum"},"datePublished":"2026-09-30T16:22:44.723Z","dateModified":"2026-09-30T16:23:09.081Z","description":"Video diffusion models spend a growing share of latency in self-attention as resolution and sequence length increase, reaching 88.2% of per-layer latency at...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/30e5cc890b745b0383f60ac3a001e5db?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/30e5cc890b745b0383f60ac3a001e5db?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Google Developers","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Google Developers","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/ea4bbded4e7f45ccb82e121a3156535f","url":"https://daily.dev/sources/googledevs"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/accelerating-spatio-temporal-attention-for-video-diffusion-on-tpus-by8crmpum","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"","timeRequired":"PT11M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Google Developers","item":"https://daily.dev/sources/googledevs"},{"@type":"ListItem","position":3,"name":"Accelerating Spatio-Temporal Attention for Video Diffusion on TPUs"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/accelerating-spatio-temporal-attention-for-video-diffusion-on-tpus-by8crmpum#faq","mainEntity":[{"@type":"Question","name":"How much speedup can tile-aligned sparse attention give over dense Splash attention on TPU v6e?","acceptedAnswer":{"@type":"Answer","text":"A tile-aligned sparse traversal implementation achieved a 2.40x speedup over dense Splash attention, reducing latency to 32.76 ms compared to 78.70 ms for dense attention. This was achieved by aligning the sparse mask boundaries with tile boundaries, reducing the fraction of tiles requiring intra-tile masking from 27.55% to just 2.22%, with masking needed only for padding at sequence edges. Teams tuning TPU inference kernels can track sparse attention optimization techniques like this on daily.dev."}},{"@type":"Question","name":"Why does naive sparse attention sometimes run slower than dense attention on TPUs?","acceptedAnswer":{"@type":"Answer","text":"A naive sparse attention kernel that retained only 39% of query-key pairs actually ran 22% slower than dense attention (96.37 ms versus 78.70 ms), because it still evaluated the exact mask inside every visited tile, including boundary tiles containing both retained and excluded pairs. This elementwise masking bottlenecks the Vector Processing Unit and stalls the Matrix Multiply Unit even when a tile is fully valid. Engineers debugging unexpected sparse-kernel slowdowns can follow hardware-level attention optimization writeups on daily.dev."}},{"@type":"Question","name":"How much does sparse attention speed up video diffusion denoising at 1440p versus 720p resolution?","acceptedAnswer":{"@type":"Answer","text":"At 1440p (302K tokens), an aggressive sparse attention schedule cut denoising time from 2,471 seconds to 1,461 seconds, a 1.69x speedup, saving over 16 minutes per video while maintaining 24.05 dB PSNR. At 720p the same approach gave only a 1.28x speedup, because attention's share of per-layer latency grows from 55.5% at 720p to 88.2% at 1440p. Developers choosing resolution and latency trade-offs for video generation can follow these benchmarks on daily.dev."}}]}
```

