<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/pewdiepie-didn-t-trust-big-ai-so-he-built-his-own-tryjlkofy" -->

---
title: PewDiePie didn't trust Big AI... so he built his own
description: PewDiePie's open-source project Odysius (nearly 90,000 GitHub stars) has a new fine-tuned, uncensored model called Ajax, built on top of Alibaba's Qwen 3.5B...
canonical: https://daily.dev/posts/pewdiepie-didn-t-trust-big-ai-so-he-built-his-own-tryjlkofy
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: PewDiePie didn't trust Big AI... so he built his own | daily.dev
og:description: PewDiePie's open-source project Odysius (nearly 90,000 GitHub stars) has a new fine-tuned, uncensored model called Ajax, built on top of Alibaba's Qwen 3.5B...
og:url: https://daily.dev/posts/pewdiepie-didn-t-trust-big-ai-so-he-built-his-own-tryjlkofy
og:image: https://api.daily.dev/og/posts/TRyjlkOFY.png
og:image:alt: PewDiePie didn't trust Big AI... so he built his own
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# PewDiePie didn't trust Big AI... so he built his own

**[Fireship](https://daily.dev/sources/fireship)** · 5 min read · 43 upvotes · 9 comments

## Summary

PewDiePie's open-source project Odysius (nearly 90,000 GitHub stars) has a new fine-tuned, uncensored model called Ajax, built on top of Alibaba's Qwen 3.5B parameter model. The video explains how he attempted to distill knowledge from OpenAI's models (resulting in his account being banned twice), then resorted to supervised fine-tuning with a small hand-collected and synthetic dataset, reinforcement learning via GRPO (DeepSeek's technique), and finally used the tool Heretic to strip safety refusals from the model. The piece covers distillation's mechanics, its role in Chinese model development, and OpenAI's countermeasures like encrypting chain-of-thought outputs. The video is sponsored by Namespace, a GitHub Actions runner replacement.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://www.youtube.com/watch?v=_5p1_TNSWqQ>

## Questions this post answers

### What is the Ajax AI model that PewDiePie created for the Odysius project?

Ajax is a fine-tuned, uncensored language model built on top of Alibaba's Qwen 3.5 billion parameter model, created to power the open-source Odysius project (nearly 90,000 GitHub stars). It was trained using supervised fine-tuning on a small custom dataset, reinforcement learning with GRPO, and then desensored using the tool Heretic to remove refusal behavior.

_Developers fine-tuning their own open models can follow projects like this one on daily.dev._

### What is GRPO and how does it differ from traditional reinforcement learning with a critic model?

GRPO, or group relative policy optimization, is a reinforcement learning technique introduced by DeepSeek where a model attempts the same task multiple times, scores the results, and learns to favor attempts that outperform the group's average. Unlike traditional RL methods, it doesn't require a separate critic or teacher model, while still improving the model's performance in a specific domain.

_daily.dev surfaces technique explainers like GRPO for engineers building domain-tuned models._

### Why did OpenAI ban an account for distilling from its models?

OpenAI banned PewDiePie's account twice for distillation, the practice of training a smaller student model to match the output probability distribution of a larger teacher model, which violates OpenAI's terms of service. OpenAI began encrypting its chain-of-thought reasoning outputs as an API response starting in 2024 specifically to prevent this kind of distillation from competitors.

_Engineers weighing API terms of service risk follow stories like this on daily.dev._

## Community discussion

Top comments from developers on daily.dev.

**@seyoungcho** · 9 upvotes

> We individuals all need to do this. Like building our own bespoke model

**@harshitmanepally** · 2 upvotes

> That is why he ruled yt so long

**@f4nt0** · 1 upvotes

> Great information, it's good that someone is fighting for free AI to everyone

**@itsbrianmckinley** · 1 upvotes

> I never though that pewdipie will be geeking so hard!

**@ristotoldsep** · 0 upvotes

> Getting banned twice and still shipping the model is very on-brand tbh.

## Similar posts on daily.dev

- [I tried PewDiePie's open-source AI workspace, and it's weirdly great](https://daily.dev/posts/i-tried-pewdiepie-s-open-source-ai-workspace-and-it-s-weirdly-great-ovxaw8qm3) · XDA Developers · 21 upvotes · 2 comments
- [PewDiePie Releases His Own Self-Hosted AI Workspace for Free](https://daily.dev/posts/pewdiepie-releases-his-own-self-hosted-ai-workspace-for-free-jahriaalb) · 80 LEVEL · 157 upvotes · 52 comments

---

Tags: [#machine-learning](https://daily.dev/tags/machine-learning), [#deep-learning](https://daily.dev/tags/deep-learning), [#openai](https://daily.dev/tags/openai), [#reinforcement-learning](https://daily.dev/tags/reinforcement-learning)

[View this post on daily.dev](https://daily.dev/posts/pewdiepie-didn-t-trust-big-ai-so-he-built-his-own-tryjlkofy)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"PewDiePie didn't trust Big AI... so he built his own","url":"https://daily.dev/posts/pewdiepie-didn-t-trust-big-ai-so-he-built-his-own-tryjlkofy","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/pewdiepie-didn-t-trust-big-ai-so-he-built-his-own-tryjlkofy"},"datePublished":"2026-10-05T20:39:21.761Z","dateModified":"2026-10-05T20:40:24.344Z","description":"PewDiePie's open-source project Odysius (nearly 90,000 GitHub stars) has a new fine-tuned, uncensored model called Ajax, built on top of Alibaba's Qwen 3.5B...","image":"https://i.ytimg.com/vi/_5p1_TNSWqQ/sddefault.jpg","thumbnailUrl":"https://i.ytimg.com/vi/_5p1_TNSWqQ/sddefault.jpg","isAccessibleForFree":true,"articleSection":"Fireship","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Fireship","logo":"https://media.daily.dev/image/upload/s--75ndqrkr--/f_auto,t_logo/v1702882094/logos/fireship.jpg","url":"https://daily.dev/sources/fireship"},"commentCount":9,"discussionUrl":"https://daily.dev/posts/pewdiepie-didn-t-trust-big-ai-so-he-built-his-own-tryjlkofy","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":43},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":9}],"keywords":"machine-learning,deep-learning,openai,reinforcement-learning","timeRequired":"PT5M","video":{"@type":"VideoObject","name":"PewDiePie didn't trust Big AI... so he built his own","description":"PewDiePie's open-source project Odysius (nearly 90,000 GitHub stars) has a new fine-tuned, uncensored model called Ajax, built on top of Alibaba's Qwen 3.5B...","thumbnailUrl":"https://i.ytimg.com/vi/_5p1_TNSWqQ/sddefault.jpg","uploadDate":"2026-10-05T20:39:21.761Z","duration":"PT5M","url":"https://api.daily.dev/r/TRyjlkOFY","embedUrl":"https://www.youtube.com/embed/_5p1_TNSWqQ"}}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Fireship","item":"https://daily.dev/sources/fireship"},{"@type":"ListItem","position":3,"name":"PewDiePie didn't trust Big AI... so he built his own"}]}
{"@context":"https://schema.org","@type":"WebPage","@id":"https://daily.dev/posts/pewdiepie-didn-t-trust-big-ai-so-he-built-his-own-tryjlkofy","comment":[{"@type":"Comment","text":"We individuals all need to do this. Like building our own bespoke model","datePublished":"2026-10-06T01:38:15.784Z","url":"https://daily.dev/posts/TRyjlkOFY#c-CwU3A2ufj","author":{"@type":"Person","name":"SeyoungCho","url":"https://daily.dev/seyoungcho","image":"https://media.daily.dev/image/upload/s---QniJFEK--/f_auto/v1724901542/avatars/avatar_O3ucZSadxc0iLADSOH3rX"},"interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":9}},{"@type":"Comment","text":"That is why he ruled yt so long","datePublished":"2026-10-06T19:05:05.403Z","url":"https://daily.dev/posts/TRyjlkOFY#c-vHDGxnHhY","author":{"@type":"Person","name":"HARSHIT MANEPALLY","url":"https://daily.dev/harshitmanepally","image":"https://media.daily.dev/image/upload/s--O0TOmw4y--/f_auto/v1715772965/public/noProfile"},"interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":2}},{"@type":"Comment","text":"Great information, it’s good that someone is fighting for free AI to everyone","datePublished":"2026-10-06T12:41:54.430Z","url":"https://daily.dev/posts/TRyjlkOFY#c-9b8uhdOND","author":{"@type":"Person","name":"Gabriel Stundner","url":"https://daily.dev/f4nt0","image":"https://media.daily.dev/image/upload/s--UpSzkOxs--/f_auto/v1753488097/avatars/avatar_whNNg0ajcmne1B7hPHb0U?_a=BAMClqZW0"},"interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1}},{"@type":"Comment","text":"I never though that pewdipie will be geeking so hard!","datePublished":"2026-10-06T17:48:19.061Z","url":"https://daily.dev/posts/TRyjlkOFY#c-ByGXAhGCV","author":{"@type":"Person","name":"Brian McKinley","url":"https://daily.dev/itsbrianmckinley"},"interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1}},{"@type":"Comment","text":"Getting banned twice and still shipping the model is very on-brand tbh.","datePublished":"2026-10-07T14:15:02.124Z","url":"https://daily.dev/posts/TRyjlkOFY#c-mQLoBwJJJ","author":{"@type":"Person","name":"Risto Tõldsep","url":"https://daily.dev/ristotoldsep","image":"https://lh3.googleusercontent.com/a/ACg8ocLDWc6mZn0JwNmXXw6WY0L_HJ6pegzRttooC5VgtXESHSj3MNYx=s96-c"}}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/pewdiepie-didn-t-trust-big-ai-so-he-built-his-own-tryjlkofy#faq","mainEntity":[{"@type":"Question","name":"What is the Ajax AI model that PewDiePie created for the Odysius project?","acceptedAnswer":{"@type":"Answer","text":"Ajax is a fine-tuned, uncensored language model built on top of Alibaba's Qwen 3.5 billion parameter model, created to power the open-source Odysius project (nearly 90,000 GitHub stars). It was trained using supervised fine-tuning on a small custom dataset, reinforcement learning with GRPO, and then desensored using the tool Heretic to remove refusal behavior. Developers fine-tuning their own open models can follow projects like this one on daily.dev."}},{"@type":"Question","name":"What is GRPO and how does it differ from traditional reinforcement learning with a critic model?","acceptedAnswer":{"@type":"Answer","text":"GRPO, or group relative policy optimization, is a reinforcement learning technique introduced by DeepSeek where a model attempts the same task multiple times, scores the results, and learns to favor attempts that outperform the group's average. Unlike traditional RL methods, it doesn't require a separate critic or teacher model, while still improving the model's performance in a specific domain. daily.dev surfaces technique explainers like GRPO for engineers building domain-tuned models."}},{"@type":"Question","name":"Why did OpenAI ban an account for distilling from its models?","acceptedAnswer":{"@type":"Answer","text":"OpenAI banned PewDiePie's account twice for distillation, the practice of training a smaller student model to match the output probability distribution of a larger teacher model, which violates OpenAI's terms of service. OpenAI began encrypting its chain-of-thought reasoning outputs as an API response starting in 2024 specifically to prevent this kind of distillation from competitors. Engineers weighing API terms of service risk follow stories like this on daily.dev."}}]}
```

