<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/swe-sweep-a-benchmark-where-agents-hunt-for-bugs-unprompted-bfwgolqrg" -->

---
title: SWE-sweep: a benchmark where agents hunt for bugs unprompted
description: A new open-ended coding benchmark called SWE-sweep, created by Ofir Press, asks AI agents to find and fix bugs in a repository with no hints, then checks their...
canonical: https://daily.dev/posts/swe-sweep-a-benchmark-where-agents-hunt-for-bugs-unprompted-bfwgolqrg
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: SWE-sweep: a benchmark where agents hunt for bugs unprompted | daily.dev
og:description: A new open-ended coding benchmark called SWE-sweep, created by Ofir Press, asks AI agents to find and fix bugs in a repository with no hints, then checks their...
og:url: https://daily.dev/posts/swe-sweep-a-benchmark-where-agents-hunt-for-bugs-unprompted-bfwgolqrg
og:image: https://api.daily.dev/og/posts/bfwgOlQrg.png
og:image:alt: SWE-sweep: a benchmark where agents hunt for bugs unprompted
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# SWE-sweep: a benchmark where agents hunt for bugs unprompted

**[Collections](https://daily.dev/sources/collections)** · 2 min read · 0 upvotes · 0 comments

## Summary

A new open-ended coding benchmark called SWE-sweep, created by Ofir Press, asks AI agents to find and fix bugs in a repository with no hints, then checks their fixes against real bugfixes from later commits using unit tests. Top models currently score under 5%, making it one of the harder benchmarks for measuring agentic coding progress. A commentator (giffmana) raises design concerns: scoring can punish models that find different-but-equally-valid bugs, the construction method is guessable so labs could train directly on the test, and future commit history must be properly hidden to avoid leakage. Despite these flaws, the verdict is that the benchmark is still useful as a near-term north star, with a caution against over-reading small score differences once models improve.

## Content

SWE-sweep is a new coding benchmark from @OfirPress. The task is open-ended: an agent gets no hints and no supervision, and has to find and fix bugs on its own. Top models score under 5%, so labs haven't really started climbing this one yet. @OfirPress describes it as one of the group's most challenging benchmarks, the kind meant to set a north star for future AI.

## How it works

As @giffmana describes it, the agent receives a repo at a commit from the past and is told to find and fix all bugs. Its fixes are then tested against real bugfixes from later commits, using their unit tests, to see whether it caught them.

## Limitations

@giffmana doesn't think the design is perfect, and points to a few problems:

- **Scoring can punish useful work.** A model might find 8 real bugs, but if the test covers 8 different ones, it scores zero. It would be just as useful as a model that found exactly the 8 tested bugs.
- **Training on the test is an obvious move.** Now that the construction is known, the recipe for it is fairly clear. Even so, if model providers do this more widely, the models should still get better at the underlying skill.
- **Future history has to be hidden.** The eval environments need to prune later commits so the agent can't read the answers. @giffmana says he didn't check, but hopes the authors learned that lesson from earlier benchmarks.

## Takeaway

@giffmana's verdict: perfect is the enemy of good, and this is a very useful new SWE-style benchmark for the near future. His one caution is that once models reach high scores, small ranking differences shouldn't be over-read.

## Questions this post answers

### What is the SWE-sweep benchmark and how does it test AI coding agents?

SWE-sweep is a coding benchmark created by Ofir Press where an AI agent receives a repository at a past commit with no hints and must find and fix all bugs on its own. Its fixes are tested against real bugfixes from later commits using their unit tests to check if it caught them. Top models currently score under 5%, making it one of the hardest open-ended agentic coding benchmarks so far.

_Developers tracking how coding agents are evaluated can follow benchmark breakdowns like this on daily.dev._

### What are the main design flaws in agentic bug-finding benchmarks like SWE-sweep?

Scoring can unfairly punish useful work: a model that finds 8 real but different bugs than the ones tested scores zero despite being just as useful. The benchmark construction method is also fairly guessable, meaning labs could train directly on the test. Additionally, future commit history must be carefully hidden from agents, or they could simply read the answers ahead of time.

_Anyone weighing benchmark results when choosing a coding agent can dig into these caveats on daily.dev._

## Community take

How the wider developer community reacted, aggregated from 2 discussions and 34 comments across x (as of 2026-10-02).

**TL;DR:** Commenters find the unprompted bug-hunting idea genuinely novel and see the sub-5% scores as a meaningful signal that bug detection is the hard part of maintenance, but several raise sharp methodological concerns about how fixes are scored and whether the benchmark construction can be gamed.

**Sentiment:** 30% positive · 45% mixed · 25% skeptical

**The case for**

- Testing whether agents can notice a bug exists, not just fix a given one, is seen as the real test of agentic capability.
- The very low scores (under 5%) are viewed as an honest, demoralizing-but-useful signal versus benchmarks like SWE-bench Verified where models already score near-perfect.
- Using real future bugfixes makes contamination easier to check for.

**The pushback**

- Scoring against the exact original patch means a model that fixes the bug a different valid way can still score zero.
- Labels only cover bugs that were eventually filed and fixed, so genuinely found-but-unreported bugs score zero too.
- The construction method is seen as guessable, raising concern that labs could train directly on the test.
- Aggregate pass rate is criticized for hiding which specific bug classes or failure modes a model systematically misses.

**By community**

- x (mixed): Replies mix genuine enthusiasm for the no-hints bug-hunting concept with pointed critiques of the scoring methodology and risk of test-set leakage, with the benchmark authors actively responding to concerns.

**Hottest debate:** Whether scoring a fix only against the original future patch unfairly zeroes out models that find and fix real bugs in a different valid way.

**Open questions**

- Does the benchmark reveal which specific bugs a model found, or only an aggregate pass count?
- At the current ~4.7% score, do models mostly fail by flagging non-bugs or by finding nothing at all?
- Does the benchmark's bug set include security vulnerabilities, not just functional bugs?

**Highlights**

> @giffmana The labels are only bugs somebody eventually fixed. Find a real one nobody ever filed and you score zero for it.
> — [AIQuanting on x](https://x.com/AIQuanting/status/2105913128324305128)

> @giffmana that mismatch matters most when a valid fix is outside the future test suite. a second lane could run broader property tests and manually adjudicate a sample of disputed patches. exact future-test matches are useful, but a narrow scoring target once agents find different bugs.
> — [CodeWex on x](https://x.com/CodeWex/status/2105913440996937929)

> @giffmana The "found 8 but the wrong 8" problem is the whole game in production. Aggregate score hides which bug class a model systematically walks past. Building closing agents, we track which failure keeps recurring, not pass rate, since that's the one that bites.
> — [vikasmalpani on x](https://x.com/vikasmalpani/status/2106002947654267060)

> @giffmana the <5% is the honest number. swe-bench verified hands the model a known bug and opus 5 clears 97%, this one makes it find the bug first. detection was always the hard half of maintenance, it just wasn't measurable until now.
> — [thebasedcapital on x](https://x.com/thebasedcapital/status/2105921388305399994)

> @OfirPress Finding vs fixing is the interesting split.  SWE-bench hands the agent an issue, here it has to notice something is wrong first.  At 4.7%, do models mostly fail by flagging non-bugs, or by reporting nothing at all?
> — [hi\_soouu on x](https://x.com/hi_soouu/status/2106066753684144421)

**Source threads**

- [x](https://x.com/giffmana/status/2105910335391596928) · 0 points · 25 comments
- [x](https://x.com/OfirPress/status/2105671012956135511) · 0 points · 9 comments

## Similar posts on daily.dev

- [Separating signal from noise in coding evaluations](https://daily.dev/posts/separating-signal-from-noise-in-coding-evaluations-swoud4dtk) · Hacker News · 0 upvotes · 0 comments

---

Tags: [#ai](https://daily.dev/tags/ai), [#llm](https://daily.dev/tags/llm), [#ai-agents](https://daily.dev/tags/ai-agents)

[View this post on daily.dev](https://daily.dev/posts/swe-sweep-a-benchmark-where-agents-hunt-for-bugs-unprompted-bfwgolqrg)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"SWE-sweep: a benchmark where agents hunt for bugs unprompted","url":"https://daily.dev/posts/swe-sweep-a-benchmark-where-agents-hunt-for-bugs-unprompted-bfwgolqrg","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/swe-sweep-a-benchmark-where-agents-hunt-for-bugs-unprompted-bfwgolqrg"},"datePublished":"2026-10-02T06:39:11.233Z","dateModified":"2026-10-02T17:10:12.259Z","description":"A new open-ended coding benchmark called SWE-sweep, created by Ofir Press, asks AI agents to find and fix bugs in a repository with no hints, then checks their...","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/swe-sweep-a-benchmark-where-agents-hunt-for-bugs-unprompted-bfwgolqrg","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"ai,llm,ai-agents","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"SWE-sweep: a benchmark where agents hunt for bugs unprompted"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/swe-sweep-a-benchmark-where-agents-hunt-for-bugs-unprompted-bfwgolqrg#faq","mainEntity":[{"@type":"Question","name":"What is the SWE-sweep benchmark and how does it test AI coding agents?","acceptedAnswer":{"@type":"Answer","text":"SWE-sweep is a coding benchmark created by Ofir Press where an AI agent receives a repository at a past commit with no hints and must find and fix all bugs on its own. Its fixes are tested against real bugfixes from later commits using their unit tests to check if it caught them. Top models currently score under 5%, making it one of the hardest open-ended agentic coding benchmarks so far. Developers tracking how coding agents are evaluated can follow benchmark breakdowns like this on daily.dev."}},{"@type":"Question","name":"What are the main design flaws in agentic bug-finding benchmarks like SWE-sweep?","acceptedAnswer":{"@type":"Answer","text":"Scoring can unfairly punish useful work: a model that finds 8 real but different bugs than the ones tested scores zero despite being just as useful. The benchmark construction method is also fairly guessable, meaning labs could train directly on the test. Additionally, future commit history must be carefully hidden from agents, or they could simply read the answers ahead of time. Anyone weighing benchmark results when choosing a coding agent can dig into these caveats on daily.dev."}}]}
```

