<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/2404-08144-llm-agents-can-autonomously-exploit-one-day-vulnerabilities-tttgishuo" -->

---
title: [2404.08144] LLM Agents can Autonomously Exploit One-day...
description: LLM agents can autonomously exploit one-day vulnerabilities in real-world systems, as shown in this work. GPT-4 has a high performance in exploiting these...
canonical: https://daily.dev/posts/2404-08144-llm-agents-can-autonomously-exploit-one-day-vulnerabilities-tttgishuo
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: [2404.08144] LLM Agents can Autonomously Exploit One-day Vulnerabilities | daily.dev
og:description: LLM agents can autonomously exploit one-day vulnerabilities in real-world systems, as shown in this work. GPT-4 has a high performance in exploiting these...
og:url: https://daily.dev/posts/2404-08144-llm-agents-can-autonomously-exploit-one-day-vulnerabilities-tttgishuo
og:image: https://api.daily.dev/og/posts/TTtGIsHuo.png
og:image:alt: [2404.08144] LLM Agents can Autonomously Exploit One-day Vulnerabilities
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# [2404.08144] LLM Agents can Autonomously Exploit One-day Vulnerabilities

**[Lobsters](https://daily.dev/sources/lobsters)** · 1 min read · 0 upvotes · 1 comments

## Summary

LLM agents can autonomously exploit one-day vulnerabilities in real-world systems, as shown in this work. GPT-4 has a high performance in exploiting these vulnerabilities when provided with CVE descriptions.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://arxiv.org/abs/2404.08144>

## Community discussion

Top comments from developers on daily.dev.

**@gabibeyo** · 0 upvotes

> Paper shows agents chaining public CVEs without a human in the room.
>
>
> That is the opposite of a jail problem. It is a "what did it call, in what order, against what target" problem.
>
>
> If your SOC cannot reconstruct the tool trail, the exploit writeup becomes folklore.Autonomous exploit demos keep proving the same thing.
>
>
> The model is capable. The missing piece is whether anyone can reconstruct the path after it moves.
>
>
> Monitor what agents do. Do not sleep on the jail.

## Similar posts on daily.dev

- [Don’t just attend KubeCon \+ CloudNativeCon, Merge Forward your experience\!](https://daily.dev/posts/don-t-just-attend-kubecon-cloudnativecon-merge-forward-your-experience--l0rpp73x8) · CNCF · 0 upvotes · 0 comments
- [Announcing H2 2026 KCDs](https://daily.dev/posts/announcing-h2-2026-kcds-m96goajm1) · CNCF · 1 upvotes · 0 comments
- [Two months of Open Community Groups](https://daily.dev/posts/two-months-of-open-community-groups-asf52zhbs) · CNCF · 0 upvotes · 0 comments

---

Tags: [#cyber](https://daily.dev/tags/cyber)

[View this post on daily.dev](https://daily.dev/posts/2404-08144-llm-agents-can-autonomously-exploit-one-day-vulnerabilities-tttgishuo)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"[2404.08144] LLM Agents can Autonomously Exploit One-day Vulnerabilities","url":"https://daily.dev/posts/2404-08144-llm-agents-can-autonomously-exploit-one-day-vulnerabilities-tttgishuo","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/2404-08144-llm-agents-can-autonomously-exploit-one-day-vulnerabilities-tttgishuo"},"datePublished":"2024-04-24T17:36:10.529Z","dateModified":"2024-05-09T09:23:16.455Z","description":"LLM agents can autonomously exploit one-day vulnerabilities in real-world systems, as shown in this work. GPT-4 has a high performance in exploiting these...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/1fed0de190485cbbb37a86fdf656d71b?_a=AQAEufR","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/1fed0de190485cbbb37a86fdf656d71b?_a=AQAEufR","isAccessibleForFree":true,"articleSection":"Lobsters","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Lobsters","logo":"https://media.daily.dev/image/upload/s--tl8v_Fku--/f_auto,t_logo/v1698841318/logos/lobste.jpg","url":"https://daily.dev/sources/lobsters"},"commentCount":1,"discussionUrl":"https://daily.dev/posts/2404-08144-llm-agents-can-autonomously-exploit-one-day-vulnerabilities-tttgishuo","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":1}],"keywords":"cyber","timeRequired":"PT1M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Lobsters","item":"https://daily.dev/sources/lobsters"},{"@type":"ListItem","position":3,"name":"[2404.08144] LLM Agents can Autonomously Exploit One-day Vulnerabilities"}]}
{"@context":"https://schema.org","@type":"WebPage","@id":"https://daily.dev/posts/2404-08144-llm-agents-can-autonomously-exploit-one-day-vulnerabilities-tttgishuo","comment":[{"@type":"Comment","text":"Paper shows agents chaining public CVEs without a human in the room.\nThat is the opposite of a jail problem. It is a “what did it call, in what order, against what target” problem.\nIf your SOC cannot reconstruct the tool trail, the exploit writeup becomes folklore.Autonomous exploit demos keep proving the same thing.\nThe model is capable. The missing piece is whether anyone can reconstruct the path after it moves.\nMonitor what agents do. Do not sleep on the jail.","datePublished":"2026-08-26T13:01:48.342Z","url":"https://daily.dev/posts/TTtGIsHuo#c-0tcMMm9kJ","author":{"@type":"Person","name":"Gabi Beyo","url":"https://daily.dev/gabibeyo","image":"https://lh3.googleusercontent.com/a/ACg8ocLemxROLDZ6RYrC2zzv0zJ9ezzj62VOCNwxNCdVNa_M9tqvU4JI=s96-c"}}]}
```

