<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/nvidia-tensorrt-accelerates-stable-diffusion-nearly-2x-faster-with-8-bit-post-training-quantization-hylvx9m1t" -->

---
title: NVIDIA TensorRT Accelerates Stable Diffusion Nearly 2x...
description: NVIDIA TensorRT has developed an 8-bit post-training quantization toolkit to speed up diffusion deployment on NVIDIA hardware while preserving image quality....
canonical: https://daily.dev/posts/nvidia-tensorrt-accelerates-stable-diffusion-nearly-2x-faster-with-8-bit-post-training-quantization-hylvx9m1t
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: NVIDIA TensorRT Accelerates Stable Diffusion Nearly 2x Faster with 8-bit Post-Training Quantization | daily.dev
og:description: NVIDIA TensorRT has developed an 8-bit post-training quantization toolkit to speed up diffusion deployment on NVIDIA hardware while preserving image quality....
og:url: https://daily.dev/posts/nvidia-tensorrt-accelerates-stable-diffusion-nearly-2x-faster-with-8-bit-post-training-quantization-hylvx9m1t
og:image: https://api.daily.dev/og/posts/hyLVX9M1t.png
og:image:alt: NVIDIA TensorRT Accelerates Stable Diffusion Nearly 2x Faster with 8-bit Post-Training Quantization
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# NVIDIA TensorRT Accelerates Stable Diffusion Nearly 2x Faster with 8-bit Post-Training Quantization

**[NVIDIA Developer](https://daily.dev/sources/nvidiadev)** · 5 min read · 1 upvotes · 0 comments

## Summary

NVIDIA TensorRT has developed an 8-bit post-training quantization toolkit to speed up diffusion deployment on NVIDIA hardware while preserving image quality. The performance of TensorRT INT8 and FP8 quantization recipes for diffusion models achieve significant speedups on NVIDIA RTX 6000 Ada GPUs. SmoothQuant is a popular PTQ method for diffusion models, but it has limitations. TensorRT has developed a fine-grained tuning pipeline called SmoothQuant to address these limitations. TensorRT 8-bit quantization can be used to accelerate diffusion models by calibrating, exporting ONNX, and building the TensorRT engine.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://developer.nvidia.com/blog/tensorrt-accelerates-stable-diffusion-nearly-2x-faster-with-8-bit-post-training-quantization/>

## Similar posts on daily.dev

- [Don’t just attend KubeCon \+ CloudNativeCon, Merge Forward your experience\!](https://daily.dev/posts/don-t-just-attend-kubecon-cloudnativecon-merge-forward-your-experience--l0rpp73x8) · CNCF · 1 upvotes · 0 comments
- [Announcing H2 2026 KCDs](https://daily.dev/posts/announcing-h2-2026-kcds-m96goajm1) · CNCF · 1 upvotes · 0 comments
- [Two months of Open Community Groups](https://daily.dev/posts/two-months-of-open-community-groups-asf52zhbs) · CNCF · 0 upvotes · 0 comments
- [CNCF Unveils Schedule for KubeCon \+ CloudNativeCon Europe 2026](https://daily.dev/posts/cncf-unveils-schedule-for-kubecon-cloudnativecon-europe-2026-ikhcoa5cb) · CNCF · 2 upvotes · 0 comments
- [CNCF Debuts KubeCon \+ CloudNativeCon Japan 2026 Schedule](https://daily.dev/posts/cncf-debuts-kubecon-cloudnativecon-japan-2026-schedule-xp5pyudub) · CNCF · 1 upvotes · 0 comments

---

Tags: [#data-science](https://daily.dev/tags/data-science), [#genai](https://daily.dev/tags/genai), [#nvidia](https://daily.dev/tags/nvidia), [#stable-diffusion](https://daily.dev/tags/stable-diffusion), [#diffusion-models](https://daily.dev/tags/diffusion-models)

[View this post on daily.dev](https://daily.dev/posts/nvidia-tensorrt-accelerates-stable-diffusion-nearly-2x-faster-with-8-bit-post-training-quantization-hylvx9m1t)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"NVIDIA TensorRT Accelerates Stable Diffusion Nearly 2x Faster with 8-bit Post-Training Quantization","url":"https://daily.dev/posts/nvidia-tensorrt-accelerates-stable-diffusion-nearly-2x-faster-with-8-bit-post-training-quantization-hylvx9m1t","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/nvidia-tensorrt-accelerates-stable-diffusion-nearly-2x-faster-with-8-bit-post-training-quantization-hylvx9m1t"},"datePublished":"2024-03-08T01:35:40.883Z","dateModified":"2024-05-09T09:28:05.588Z","description":"NVIDIA TensorRT has developed an 8-bit post-training quantization toolkit to speed up diffusion deployment on NVIDIA hardware while preserving image quality....","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/d11962b797b1125e53cf2aa47eafe4fe?_a=AQAEufR","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/d11962b797b1125e53cf2aa47eafe4fe?_a=AQAEufR","isAccessibleForFree":true,"articleSection":"NVIDIA Developer","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"NVIDIA Developer","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/86e45aab42ba48ce83103d01b1119910","url":"https://daily.dev/sources/nvidiadev"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/nvidia-tensorrt-accelerates-stable-diffusion-nearly-2x-faster-with-8-bit-post-training-quantization-hylvx9m1t","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"data-science,genai,nvidia,stable-diffusion,diffusion-models","timeRequired":"PT5M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"NVIDIA Developer","item":"https://daily.dev/sources/nvidiadev"},{"@type":"ListItem","position":3,"name":"NVIDIA TensorRT Accelerates Stable Diffusion Nearly 2x Faster with 8-bit Post-Training Quantization"}]}
```

