<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/github---danau5tin-ai-trains-ai-rl-training-an-ai-agent-to-rl-train-ai-agents--gqshn7as2" -->

---
title: GitHub - Danau5tin/ai-trains-ai: RL-training an AI agent...
description: A developer built a recursive RL system where an AI agent (Qwen3.6-35B-A3B with LoRA) is trained via reinforcement learning to write complete RL training jobs...
canonical: https://daily.dev/posts/github---danau5tin-ai-trains-ai-rl-training-an-ai-agent-to-rl-train-ai-agents--gqshn7as2
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: GitHub - Danau5tin/ai-trains-ai: RL-training an AI agent to RL-train AI agents. | daily.dev
og:description: A developer built a recursive RL system where an AI agent (Qwen3.6-35B-A3B with LoRA) is trained via reinforcement learning to write complete RL training jobs...
og:url: https://daily.dev/posts/github---danau5tin-ai-trains-ai-rl-training-an-ai-agent-to-rl-train-ai-agents--gqshn7as2
og:image: https://api.daily.dev/og/posts/gqsHn7aS2.png
og:image:alt: GitHub - Danau5tin/ai-trains-ai: RL-training an AI agent to RL-train AI agents.
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# GitHub - Danau5tin/ai-trains-ai: RL-training an AI agent to RL-train AI agents.

**[Hacker News](https://daily.dev/sources/hn)** · 13 min read · 0 upvotes · 0 comments

## Summary

A developer built a recursive RL system where an AI agent (Qwen3.6-35B-A3B with LoRA) is trained via reinforcement learning to write complete RL training jobs for smaller models. The outer loop uses Tinker (GRPO) to train the trainer agent, while the inner loop runs the agent-written jobs on Runpod GPUs using prime-rl. Over 54 training steps, reward climbed from ~0.0 to ~0.63, and the skill transferred to a held-out task family the agent never trained on. The agent also learned to select better base models and hyperparameters autonomously. The full pipeline, trained weights (LoRA adapter on HuggingFace), reward code, and retrospectives are open-sourced. Total cost for the headline arc was ~$1,275.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://github.com/Danau5tin/ai-trains-ai>

## Similar posts on daily.dev

- [Amazon SageMaker AI launches multi-turn reinforcement learning for AI agent model customization](https://daily.dev/posts/amazon-sagemaker-ai-launches-multi-turn-reinforcement-learning-for-ai-agent-model-customization-zoa9pgyuq) · AWS · 0 upvotes · 0 comments
- [How to Customize an LLM for AI Agents using SFT and QLoRA](https://daily.dev/posts/how-to-customize-an-llm-for-ai-agents-using-sft-and-qlora-jeancs7na) · freeCodeCamp · 0 upvotes · 0 comments

---

Tags: [#python](https://daily.dev/tags/python), [#ai-agents](https://daily.dev/tags/ai-agents), [#reinforcement-learning](https://daily.dev/tags/reinforcement-learning), [#lora](https://daily.dev/tags/lora)

[View this post on daily.dev](https://daily.dev/posts/github---danau5tin-ai-trains-ai-rl-training-an-ai-agent-to-rl-train-ai-agents--gqshn7as2)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"GitHub - Danau5tin/ai-trains-ai: RL-training an AI agent to RL-train AI agents.","url":"https://daily.dev/posts/github---danau5tin-ai-trains-ai-rl-training-an-ai-agent-to-rl-train-ai-agents--gqshn7as2","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/github---danau5tin-ai-trains-ai-rl-training-an-ai-agent-to-rl-train-ai-agents--gqshn7as2"},"datePublished":"2026-07-14T14:17:34.627Z","dateModified":"2026-07-14T14:25:32.229Z","description":"A developer built a recursive RL system where an AI agent (Qwen3.6-35B-A3B with LoRA) is trained via reinforcement learning to write complete RL training jobs...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/2e698921a7335a65c4586bdddcd4982e?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/2e698921a7335a65c4586bdddcd4982e?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Hacker News","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Hacker News","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/hn","url":"https://daily.dev/sources/hn"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/github---danau5tin-ai-trains-ai-rl-training-an-ai-agent-to-rl-train-ai-agents--gqshn7as2","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"python,ai-agents,reinforcement-learning,lora","timeRequired":"PT13M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Hacker News","item":"https://daily.dev/sources/hn"},{"@type":"ListItem","position":3,"name":"GitHub - Danau5tin/ai-trains-ai: RL-training an AI agent to RL-train AI agents."}]}
```

