<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/google-s-scientisttwo-automates-ml-experimentation-with-recursive-self-improvement-foxy29kyn" -->

---
title: Google&#x27;s ScientistTwo automates ML experimentation with...
description: Google Research introduced ScientistTwo, a multi-agent system that autonomously runs the full ML research loop: proposing ideas, running experiments,...
canonical: https://daily.dev/posts/google-s-scientisttwo-automates-ml-experimentation-with-recursive-self-improvement-foxy29kyn
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Google&#x27;s ScientistTwo automates ML experimentation with recursive self-improvement | daily.dev
og:description: Google Research introduced ScientistTwo, a multi-agent system that autonomously runs the full ML research loop: proposing ideas, running experiments,...
og:url: https://daily.dev/posts/google-s-scientisttwo-automates-ml-experimentation-with-recursive-self-improvement-foxy29kyn
og:image: https://api.daily.dev/og/posts/FoxY29KYn.png
og:image:alt: Google&#x27;s ScientistTwo automates ML experimentation with recursive self-improvement
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Google's ScientistTwo automates ML experimentation with recursive self-improvement

**[Collections](https://daily.dev/sources/collections)** · 2 min read · 1 upvotes · 0 comments

## Summary

Google Research introduced ScientistTwo, a multi-agent system that autonomously runs the full ML research loop: proposing ideas, running experiments, incorporating reviewer feedback, and then iterating on its own best results rather than only human baselines. Tested on 107 ML problems drawn from ICLR, ICML, and NeurIPS papers, it beat the human baseline in 86 cases (80.4% success rate) with an average relative improvement of 25.2%. The piece frames this as evidence that the mechanical grind of experimentation can now be largely automated, while scientific judgment about which problems matter remains an open, harder question.

## Content

Google Research published a paper on ScientistTwo, a many-agent harness designed for long-horizon research tasks. The core idea is recursive self-improvement: the system improves a human-defined method, then uses its own discovery as the new baseline and tries to improve that too.

The loop works like this - propose ideas, run experiments, discard what doesn't help, incorporate reviewer feedback, and repeat. It's less a research assistant and more a research loop that runs without human intervention between cycles.

The results are hard to dismiss. Across 107 ML problems drawn from ICLR, ICML, and NeurIPS papers, ScientistTwo improved on the human baseline in 86 cases - an 80.4% success rate with an average relative improvement of 25.2%.

One thing worth sitting with: the paper suggests that the *experimentation* part of research can be automated well before the *judgment* part. ScientistTwo can run trials and optimize metrics, but it's still working within problem definitions that humans set. The loop is automated; the question of what's worth looping on isn't.

There's also an interesting implication for agent evaluation more broadly - if you auto-optimize the agent harness itself, eval scores can improve significantly, which raises questions about what those scores actually measure.

The full paper is from Google Cloud AI Research and is worth reading if you follow autonomous research agents or self-improving systems.

## Questions this post answers

### What is Google's ScientistTwo and how does it improve on ML research baselines?

ScientistTwo is a many-agent system from Google Research that autonomously runs the ML research loop of proposing ideas, running experiments, incorporating reviewer feedback, and iterating using its own best result as the new baseline. Tested on 107 ML problems drawn from ICLR, ICML, and NeurIPS papers, it beat the human baseline in 86 cases, an 80.4% success rate, with an average relative improvement of 25.2%.

_Following autonomous research agents like this helps teams gauge how much experimentation work can be offloaded, a topic tracked on daily.dev._

## Community take

How the wider developer community reacted, aggregated from 2 discussions and 28 comments across x (as of 2026-09-24).

**TL;DR:** Reactions are split between excitement about a closed research loop and skepticism that beating published baselines on 86/107 problems is really 'new' science versus re-derivation; many raise concerns about validation, provenance, and what happens once self-generated results become the next baseline.

**Sentiment:** 20% positive · 45% mixed · 35% skeptical

**The case for**

- Some see the closed hypothesis-experiment-review-iterate loop as a genuine step toward automating experimentation.
- A few argue that building on prior published work is simply how science normally progresses, so re-derivation criticism is overstated.
- One commenter suggests the real achievement is proving the research loop can be closed at all, regardless of product longevity.

**The pushback**

- Several point out that solving 86/107 problems from published papers is re-derivation of known answers, not frontier discovery.
- Multiple commenters worry that once a self-generated result becomes the next baseline, errors or bad assumptions could compound unnoticed through the research tree.
- Some question whether the improvement is driven by the model itself or mostly by the reviewer-feedback/evaluation loop.
- There's concern about self-validation via simulated peer review, asking who checks the checker once the loop runs on its own outputs.
- A few call for reproducible lineage, logging, and rollback mechanisms to track which version produced which result.
- One commenter argues the paper mislabels recursive configuration within known problem-spaces as autonomous scientific pioneering.

**By community**

- x (mixed): Replies split between calling the closed research loop a meaningful step and dismissing the 86/107 result as mere re-derivation of published answers, with recurring worries about validation and baseline drift.

**Hottest debate:** Whether beating human baselines on problems drawn from already-published papers counts as genuine new discovery or just re-derivation.

**Open questions**

- Has the system produced any result that no human baseline ever reached, beyond re-deriving published answers?
- How much of the improvement comes from the reviewer-feedback loop versus the model's own capability?
- How should a reproducible lineage and rollback mechanism be built so a bad discovery doesn't compound once it becomes the next baseline?
- Who or what validates results once the system's peer review is itself part of the closed loop?

**Highlights**

> @rohanpaul_ai 86 of 107 is re-derivation though, not frontier work: every problem comes from published papers where humans already found the answer. has anything in the paper shown scientisttwo producing a result no human baseline ever reached?
> — [Raccoon679 on x · 1 comments](https://x.com/Raccoon679/status/2101451945013858693)

> @rohanpaul_ai recursive improvement needs lineage and rollback more than another loop. once an accepted result becomes the next baseline, one bad discovery can compound through the whole research tree.
> — [johnroodepic on x · 1 comments](https://x.com/johnroodepic/status/2101456039715746166)

> @rohanpaul_ai It validates itself through its own simulated peer review. A closed loop checking its own closed loop isn't verification. Who checks the checker once the loop runs on its own results?
> — [PatrickFarrelAI on x](https://x.com/PatrickFarrelAI/status/2101638860530692549)

> @websterweby @rohanpaul_ai That is the right comparison to run. If reviewer feedback is carrying most of the lift, the value may be in the evaluation loop and task curation rather than raw model autonomy.
> — [hars\_7086 on x](https://x.com/hars_7086/status/2101619393314578610)

> @rohanpaul_ai The interesting question is not only whether the system can improve a method, but whether each improvement survives independent validation. A recursive research loop can compound useful discoveries, but it can also compound an unnoticed assumption. Separating generation,
> — [CedricDoutres on x · 1 points](https://x.com/CedricDoutres/status/2101491342144946300)

**Source threads**

- [x](https://x.com/rohanpaul_ai/status/2101449041091657746) · 0 points · 28 comments
- [x](https://x.com/rohanpaul_ai/status/2101756032607506835) · 0 points · 0 comments

## Similar posts on daily.dev

- [Andrej Karpathy’s 630-line Python script ran 50 experiments overnight without any human input](https://daily.dev/posts/andrej-karpathy-s-630-line-python-script-ran-50-experiments-overnight-without-any-human-input-pjqwmdwbv) · The New Stack · 45 upvotes · 0 comments

---

Tags: [#machine-learning](https://daily.dev/tags/machine-learning), [#ai-agents](https://daily.dev/tags/ai-agents)

[View this post on daily.dev](https://daily.dev/posts/google-s-scientisttwo-automates-ml-experimentation-with-recursive-self-improvement-foxy29kyn)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Google's ScientistTwo automates ML experimentation with recursive self-improvement","url":"https://daily.dev/posts/google-s-scientisttwo-automates-ml-experimentation-with-recursive-self-improvement-foxy29kyn","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/google-s-scientisttwo-automates-ml-experimentation-with-recursive-self-improvement-foxy29kyn"},"datePublished":"2026-09-20T19:31:22.811Z","dateModified":"2026-09-24T08:20:42.333Z","description":"Google Research introduced ScientistTwo, a multi-agent system that autonomously runs the full ML research loop: proposing ideas, running experiments,...","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/google-s-scientisttwo-automates-ml-experimentation-with-recursive-self-improvement-foxy29kyn","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"machine-learning,ai-agents","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"Google's ScientistTwo automates ML experimentation with recursive self-improvement"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/google-s-scientisttwo-automates-ml-experimentation-with-recursive-self-improvement-foxy29kyn#faq","mainEntity":[{"@type":"Question","name":"What is Google's ScientistTwo and how does it improve on ML research baselines?","acceptedAnswer":{"@type":"Answer","text":"ScientistTwo is a many-agent system from Google Research that autonomously runs the ML research loop of proposing ideas, running experiments, incorporating reviewer feedback, and iterating using its own best result as the new baseline. Tested on 107 ML problems drawn from ICLR, ICML, and NeurIPS papers, it beat the human baseline in 86 cases, an 80.4% success rate, with an average relative improvement of 25.2%. Following autonomous research agents like this helps teams gauge how much experimentation work can be offloaded, a topic tracked on daily.dev."}}]}
```

