<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/introducing-gigagpt-gpt-3-sized-models-in-565-lines-of-code-wlzrs6dyh" -->

---
title: Introducing gigaGPT: GPT-3 sized models in 565 lines of code
description: gigaGPT is a codebase by Cerebras that allows for training GPT models with over 100 billion parameters. It utilizes the memory and compute capacity of Cerebras...
canonical: https://daily.dev/posts/introducing-gigagpt-gpt-3-sized-models-in-565-lines-of-code-wlzrs6dyh
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Introducing gigaGPT: GPT-3 sized models in 565 lines of code | daily.dev
og:description: gigaGPT is a codebase by Cerebras that allows for training GPT models with over 100 billion parameters. It utilizes the memory and compute capacity of Cerebras...
og:url: https://daily.dev/posts/introducing-gigagpt-gpt-3-sized-models-in-565-lines-of-code-wlzrs6dyh
og:image: https://api.daily.dev/og/posts/WLzrs6DyH.png
og:image:alt: Introducing gigaGPT: GPT-3 sized models in 565 lines of code
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Introducing gigaGPT: GPT-3 sized models in 565 lines of code

**[Hacker News](https://daily.dev/sources/hn)** · 6 min read · 0 upvotes · 0 comments

## Summary

gigaGPT is a codebase by Cerebras that allows for training GPT models with over 100 billion parameters. It utilizes the memory and compute capacity of Cerebras hardware to enable large scale training on vanilla torch.nn code. The models trained with gigaGPT show promising results and can scale to models in excess of 1 trillion parameters.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://www.cerebras.net/blog/introducing-gigagpt-gpt-3-sized-models-in-565-lines-of-code>

## Similar posts on daily.dev

- [Don’t just attend KubeCon \+ CloudNativeCon, Merge Forward your experience\!](https://daily.dev/posts/don-t-just-attend-kubecon-cloudnativecon-merge-forward-your-experience--l0rpp73x8) · CNCF · 1 upvotes · 0 comments
- [Announcing H2 2026 KCDs](https://daily.dev/posts/announcing-h2-2026-kcds-m96goajm1) · CNCF · 1 upvotes · 0 comments
- [Two months of Open Community Groups](https://daily.dev/posts/two-months-of-open-community-groups-asf52zhbs) · CNCF · 0 upvotes · 0 comments
- [CNCF Unveils Schedule for KubeCon \+ CloudNativeCon Europe 2026](https://daily.dev/posts/cncf-unveils-schedule-for-kubecon-cloudnativecon-europe-2026-ikhcoa5cb) · CNCF · 2 upvotes · 0 comments
- [CNCF Debuts KubeCon \+ CloudNativeCon Japan 2026 Schedule](https://daily.dev/posts/cncf-debuts-kubecon-cloudnativecon-japan-2026-schedule-xp5pyudub) · CNCF · 1 upvotes · 0 comments

---

Tags: [#machine-learning](https://daily.dev/tags/machine-learning)

[View this post on daily.dev](https://daily.dev/posts/introducing-gigagpt-gpt-3-sized-models-in-565-lines-of-code-wlzrs6dyh)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Introducing gigaGPT: GPT-3 sized models in 565 lines of code","url":"https://daily.dev/posts/introducing-gigagpt-gpt-3-sized-models-in-565-lines-of-code-wlzrs6dyh","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/introducing-gigagpt-gpt-3-sized-models-in-565-lines-of-code-wlzrs6dyh"},"datePublished":"2023-12-11T19:46:37.051Z","dateModified":"2023-12-11T19:46:34.964Z","description":"gigaGPT is a codebase by Cerebras that allows for training GPT models with over 100 billion parameters. It utilizes the memory and compute capacity of Cerebras...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/b14f04f9065b0b4e1950c93486e0974c?_a=AQAEufR","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/b14f04f9065b0b4e1950c93486e0974c?_a=AQAEufR","isAccessibleForFree":true,"articleSection":"Hacker News","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Hacker News","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/hn","url":"https://daily.dev/sources/hn"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/introducing-gigagpt-gpt-3-sized-models-in-565-lines-of-code-wlzrs6dyh","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"machine-learning","timeRequired":"PT6M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Hacker News","item":"https://daily.dev/sources/hn"},{"@type":"ListItem","position":3,"name":"Introducing gigaGPT: GPT-3 sized models in 565 lines of code"}]}
```

