<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/llms-full.txt" -->

---
---

\# daily.dev - Full Context > This is the expanded llms.txt for daily.dev: the full markdown of every page indexed in https://daily.dev/llms.txt, for AI/LLM consumption with larger context windows. Blog posts are listed as links only. Individual daily.dev posts are not included; fetch them via \`.md\` URLs or the Public API described in llms.txt. --- ## Core Pages --- title: "daily.dev Blog | Developer news, deep-dives, and editorial" url: https://daily.dev/blog/ description: "Developer news, deep-dives, and opinion from the daily.dev community. Browse all posts at /blog/ or fetch each post as markdown via /blog/\[slug\].md/." lastUpdated: "2026-04-30" llmsTxt: https://daily.dev/llms.txt --- ## daily.dev Blog The daily.dev blog publishes original deep-dives on developer tools and frameworks, breaking news on the engineering ecosystem, and editorial from the daily.dev community. ### Browse - \[Blog index\](https://daily.dev/blog/): every post, newest-first, with category filters and search - \[RSS feed\](https://daily.dev/rss.xml): subscribe to new posts - Category archives at \`https://daily.dev/categories/\[slug\]/\` ### Read individual posts as markdown Every blog post is also available as plain markdown for LLM ingestion. Append \`.md/\` to any post URL: - \`https://daily.dev/blog/\[slug\]/\` returns the rendered HTML page - \`https://daily.dev/blog/\[slug\].md/\` returns the same content as plain markdown Use the \[RSS feed\](https://daily.dev/rss.xml) or \[sitemap\](https://daily.dev/sitemap-blog.xml) to enumerate available posts. --- --- title: "daily.dev apps | Browser extension, iOS, Android, web app" url: https://daily.dev/apps/ description: "Get daily.dev as a browser extension (Chrome / Edge), an iOS app, an Android app, or a web app. Free forever, no signup required." lastUpdated: "2026-04-30" llmsTxt: https://daily.dev/llms.txt --- ## Get daily.dev wherever you work daily.dev ships across every surface engineers use day-to-day. Pick your platform. Your feed, squads, and history sync between all of them. ### Available platforms - \[Browser extension\](https://api.daily.dev/get): turns every new tab into a personalized dev feed (Chrome, Edge) - \[iOS app\](https://apps.apple.com/app/daily-dev/id6740634400): read on the go, save for later, native push for must-read posts - \[Android app\](https://play.google.com/store/apps/details?id=dev.daily): same feed, same squads, native experience - \[Web app\](https://daily.dev): works in any modern browser, no install required ### Why developers use the apps daily - Every surface is free, forever. No paywall, no signup required to read - Personalization gets sharper the more you use it. A relevant feed within 7 days - Squads sync across mobile and desktop so you can keep the conversation going - All apps are fully open source; the codebase is on \[GitHub\](https://github.com/dailydotdev/daily) \[See the rendered apps page\](https://daily.dev/apps/) for download badges, screenshots, and the full feature breakdown. --- --- title: "Features | daily.dev" url: https://daily.dev/features/ description: "Every daily.dev feature in one place: the personalized feed, squads, AI search, bookmarks, streaks, the browser extension, and the agent surfaces. Free forever, with an optional Plus upgrade." lastUpdated: "2026-08-31" llmsTxt: https://daily.dev/llms.txt --- ## What daily.dev does daily.dev pulls thousands of sources into one ranked feed that opens with every new tab. Millions of developers use it to keep up without keeping tabs open. It is free forever. Plus is for power users who want more on top. ### The feed - \*\*Personalized feed\*\*: ranked from what you read, upvote, and follow, and it sharpens over time - \*\*New tab feed\*\*: the browser extension replaces the blank new-tab page with your feed - \*\*Presidential briefing\*\*: an AI brief compressing everything that mattered into a two-minute read (unlimited on Plus) - \*\*Breaking news\*\*: the stories blowing up right now, ranked by how the community is reacting - \*\*Follow tags and sources\*\*: languages, frameworks, tools, and publications you choose - \*\*Advanced custom feeds\*\* (Plus): extra feeds scoped to one topic each - \*\*Keyword filters\*\* (Plus): mute the buzzwords you are sick of hearing - \*\*AI-powered clean titles\*\* (Plus): clickbait headlines rewritten into what the post actually says - \*\*Auto-translate\*\* (Plus): titles and summaries in your language - \*\*Ad-free everywhere\*\* (Plus): the free plan carries developer-relevant ads; Plus turns them off on every surface ### Reading - \*\*Built-in reader\*\*: articles open inside daily.dev with the discussion beside them - \*\*Bookmarks and reading reminders\*\*: save anything, set a time to come back to it - \*\*Spotlight search\*\*: Cmd+K for posts, squads, people, tags, and actions - \*\*AI search\*\*: plain-language questions answered from community-vetted posts, with sources linked - \*\*Reading history\*\*: everything you have opened, in order - \*\*Bookmark folders\*\* (Plus): organize a long reading list into folders ### Community - \*\*Squads\*\*: communities built around the tools you use - \*\*Discussions, upvotes, and downvotes\*\*: every post arrives with its conversation attached - \*\*Create and share posts\*\*: publish a link, a question, or a write-up into a squad - \*\*Follow developers\*\*: see what the people whose taste you trust are surfacing - \*\*Members-only Squad\*\* (Plus): a space for Plus members, feedback, and priority support ### Your profile - \*\*Reading streaks\*\*, \*\*reputation\*\*, a \*\*public profile\*\*, a shareable \*\*DevCard\*\*, and \*\*notifications\*\* you control ### Everywhere - \*\*Browser extension\*\* (Chrome, Edge), \*\*iOS and Android apps\*\*, and the \*\*web app\*\* - \*\*Cross-device sync\*\* for bookmarks, history, squads, follows, and your streak - \*\*Dark and light themes\*\* on every surface ### For agents and automation - \*\*Markdown twins\*\*: every marketing, blog, and post page answers on a \`.md\` URL and honours \`Accept: text/markdown\` - \*\*Agent Skills\*\*: discoverable at \[/.well-known/agent-skills/index.json\](https://daily.dev/.well-known/agent-skills/index.json) - \*\*Claude Code plugin\*\*: the same skills, packaged - \*\*Public API\*\* (Plus): REST access to feeds, posts, comments, search, bookmarks, and custom feeds ### Getting the most out of it 1\. Make it your new tab and then forget about it 2\. Follow eight to twelve tags on day one 3\. Downvote early; the ranking learns fastest from what you reject 4\. Bookmark with a reminder, not with hope 5\. Join two squads, not ten 6\. Read the replies before the article \[See the rendered features page\](https://daily.dev/features/) for the full breakdown, and \[pricing\](https://daily.dev/pricing/) for what free covers. --- --- title: "Pricing | daily.dev" url: https://daily.dev/pricing/ description: "daily.dev is free forever. The feed, the extension, the mobile apps, squads, AI search, and bookmarks cost nothing, with no trial and no credit card. Plus is $14.99/month or $89.99/year." lastUpdated: "2026-08-31" llmsTxt: https://daily.dev/llms.txt --- ## daily.dev is free forever The free plan is the real product, and it stays this rich because daily.dev is an ad-supported business. Millions of developers use daily.dev every day without paying anything. Plus was built for the die-hard fans who want more. ### What free includes - Personalized feed, and the new-tab feed via the browser extension - iOS and Android apps, and the web app, all synced to one account - Squads, discussions, and upvotes - AI search - Bookmarks and reading reminders - Streaks, reputation, and your public profile - Reading history, and following tags, sources, and people - Over a hundred other features ### What free does not mean here - \*\*No trial that ends\*\*: there is no day 14 and no downgrade email - \*\*No credit card\*\*: nothing on the free plan asks for payment details, and you do not need an account to start reading - \*\*No seat limits\*\*: no team plan to outgrow and no per-seat maths - \*\*No paywalled content\*\*: free readers see the same feed, squads, and discussions as everyone else daily.dev is an ad-supported business, which is why the free plan can stay this rich. Plus is for the people who want more. ### daily.dev Plus (optional) \*\*$14.99/month\*\* or \*\*$89.99/year\*\*, in USD. Cancel any time. Checkout on \[the Plus page\](https://daily.dev/plus) may show the amount in your own currency. Plus was designed for the die-hard fans of the platform: people who already live in daily.dev and want sharper control over what reaches them. It is optional. The free plan stays the real product. It adds: - Advanced custom feeds - Unlimited presidential briefings - AI-powered clean titles - Keyword filters - Auto-translate your feed - Bookmark folders - Ad-free everywhere: no ads on the feed, on mobile, or on the web - Members-only Squad - Public API access ### If you cancel You keep Plus until the end of the period you paid for, then continue on the free plan with your feed, bookmarks, squads, and streak intact. \[See the rendered pricing page\](https://daily.dev/pricing/) for the full comparison, and \[features\](https://daily.dev/features/) for everything the product does. --- --- title: "About daily.dev | The developer community" url: https://daily.dev/about-us/ description: "daily.dev is the developer-first community where engineers come to discover what is next. Read about the mission, the team, and the open-source values." lastUpdated: "2026-04-30" llmsTxt: https://daily.dev/llms.txt --- ## What daily.dev is daily.dev is an editorial home for engineers. We started in 2017 with a simple browser extension that replaced your new-tab page with a feed of posts every developer should know about. Today, millions of engineers across iOS, Android, web, and four browser extensions use daily.dev to keep their finger on the pulse. ### The mission To build the developer network you wish existed: open, relevant, and aligned with the people who actually write the code. ### What we believe - The best learning happens between peers, not from gurus - Quality should always beat quantity; algorithms should serve the reader, not the advertiser - Open source is a value, not a marketing line. Most of our codebase is on \[GitHub\](https://github.com/dailydotdev/daily) and we credit every contributor who ships work into production - Distributed-by-default is a feature: the team works async across 8 countries and we hire for output, not hours ### Where to read more - \[Careers\](https://daily.dev/careers/): open roles - \[Recruitment process\](https://daily.dev/recruitment-process/): how we hire - \[Brand assets\](https://brand.daily.dev/): logos, screenshots, press materials - \[Blog\](https://daily.dev/blog/): what we publish \[See the rendered About us page\](https://daily.dev/about-us/) for the full story. --- --- title: "Contact daily.dev | Press, partnerships, and product" url: https://daily.dev/contact/ description: "Reach the daily.dev team for press, partnerships, business, advertising, content, or product feedback. Email and form-based options for every audience." lastUpdated: "2026-04-30" llmsTxt: https://daily.dev/llms.txt --- ## How to reach us Pick the channel that matches what you need. We route through dedicated inboxes so the right person gets back to you fast. ### Channels - \*\*Press & PR\*\*: \`hello@daily.dev\` (or grab assets from \[brand.daily.dev\](https://brand.daily.dev/)) - \*\*Partnerships\*\*: \`partnerships@daily.dev\` - \*\*Recruiter sourcing\*\*: see the \[recruiter site\](https://recruiter.daily.dev/) - \*\*Native ads & sponsorships\*\*: see the \[business site\](https://business.daily.dev/) - \*\*Product feedback\*\*: \[feedback page\](https://daily.dev/feedback/) or upvote what is already on the roadmap - \*\*Bug reports\*\*: open an issue on \[GitHub\](https://github.com/dailydotdev/daily) - \*\*Privacy & data requests\*\*: \`privacy@daily.dev\` (see the \[privacy center\](https://daily.dev/privacy/)) \[See the rendered contact page\](https://daily.dev/contact/) for the form-based options. --- --- title: "Careers at daily.dev | Remote-first engineering, design, content" url: https://daily.dev/careers/ description: "Build daily.dev with us. Remote-first, async-by-default, equity-aligned. No open roles right now, but you can hear about the next one the day it opens." lastUpdated: "2026-08-09" llmsTxt: https://daily.dev/llms.txt --- ## Build daily.dev with us We are a small, remote-first team building the place engineers actually trust to learn, connect, and grow careers. ### How we work - Async-by-default. Output over hours - Equity-aligned for every full-time hire: generous grants, real outcomes - Remote-first, forever. No mandatory office. Two retreats a year somewhere lovely - Flexible schedule. Work the hours that match your time zone and your focus ### Where we are Eight countries across European, African, Middle East, and North American time zones: Israel, United Kingdom, Croatia, Australia, United Arab Emirates, Portugal, South Africa, United States. ### Open roles We have no open roles right now. We are a small team and we open roles a few times a year, so rather than leave a posting up to collect applications we cannot act on, we say so plainly. Two things you can do in the meantime: - Turn on \[job preferences\](https://daily.dev/settings/job-preferences) on daily.dev. We post our roles inside the product before anywhere else, so you will hear about ours the day it opens alongside every other role that fits you. - Write to \`hi@daily.dev\` anyway. We have hired people who wrote to us before a role existed. A human reads every one of these. \[See the rendered careers page\](https://daily.dev/careers/) for the full culture pitch and perks list. --- --- title: "Recruitment process | How daily.dev hires" url: https://daily.dev/recruitment-process/ description: "How daily.dev hires: the steps, the timing, what we evaluate, and how we keep the process humane and predictable for every candidate." lastUpdated: "2026-04-30" llmsTxt: https://daily.dev/llms.txt --- ## How we hire We optimize the recruitment process for two things: respect for the candidate's time, and quality we can actually trust. ### The steps 1\. \*\*Application\*\*: short form, no cover letter required. Apply via the role on the \[careers page\](https://daily.dev/careers/), or email \`hi@daily.dev\` if nothing open fits you 2\. \*\*Intro chat\*\* (30 min): with the hiring manager. Two-way: we tell you about the work, you tell us what you are looking for 3\. \*\*Take-home or live exercise\*\*: scoped to about 2 hours, paid where local rules allow, designed around real work the role does 4\. \*\*Team interviews\*\* (90 to 120 min total): usually two conversations: technical depth and craft / collaboration 5\. \*\*Founder chat\*\* (30 min): with one of the founders, focused on values, ambition, and questions you have 6\. \*\*Offer\*\*: within 48 hours of the founder chat. We share the full comp range up front ### What we evaluate - \*\*Craft\*\*: taste, rigor, attention to the work - \*\*Output bias\*\*: ship-vs-discuss ratio - \*\*Async fluency\*\*: written communication, decision documentation - \*\*Open-source instincts\*\*: how you handle being wrong in public ### What we do not evaluate - Pedigree. We do not care where you went to school or which name brand you worked at - Brain teasers, whiteboard puzzles, leetcode trivia - Hours worked. We hire grown-ups; you decide when your brain is on \[See the rendered recruitment process page\](https://daily.dev/recruitment-process/) for the full story. --- --- title: "Wall of Love | daily.dev" url: https://daily.dev/wall-of-love/ description: "Real, consented testimonials from developers in the daily.dev community — engineers, founders, CTOs, and students in their own words. Every quote links to a verifiable daily.dev profile." lastUpdated: "2026-05-28" llmsTxt: https://daily.dev/llms.txt --- ## What developers say about daily.dev Every quote on the \[Wall of Love\](https://daily.dev/wall-of-love/) was written by a real developer in the daily.dev community and published with their explicit permission. No edits, no paid endorsements. Each card links to that developer's daily.dev profile so anyone can verify the human behind the quote. ### The three things developers tell us most - \*\*Stay current\*\* — the largest share of testimonials describe daily.dev as how they stay up to date with tech without browsing all over the internet. \*"It allows me to be up to date with tech without having to browse all over the internet."\* — Ross, Tech Lead. - \*\*Learn faster\*\* — developers credit the curated feed with surfacing tools, articles, and ideas they would never have searched for. \*"I often discover useful articles and resources that help me grow and stay motivated."\* — Kristijan Kelić, Full-stack engineer. - \*\*Find your people\*\* — the community (squads, comments, discussions) turns a news feed into a place developers feel part of. \*"I've met other devs through squads and comments — feels like a real community."\* — Claudette, Software Developer. The voices span every career stage — Senior engineers, Tech Leads, Principal Architects, CTOs, Founders, and Students — across the daily.dev community of millions of developers. \[Read the full wall\](https://daily.dev/wall-of-love/), or \[share your own story\](https://it057218.typeform.com/to/uQihIeNC). --- --- title: "Product feedback | daily.dev" url: https://daily.dev/feedback/ description: "Tell daily.dev what to build next. Upvote ideas, file bugs, or write a long-form proposal. Every input goes into the public roadmap." lastUpdated: "2026-04-30" llmsTxt: https://daily.dev/llms.txt --- ## Help us build the right thing Every product decision at daily.dev passes the "would I want this in my new tab?" test. Your feedback is the strongest input we have. ### How to leave feedback - \*\*Upvote / suggest\*\*: the in-app feedback panel surfaces requests across the community - \*\*File a bug\*\*: open an issue at \[github.com/dailydotdev/daily\](https://github.com/dailydotdev/daily) - \*\*Long-form proposal\*\*: email \`product@daily.dev\` with the rough scope and the user pain you are solving for - \*\*Public Discord\*\*: drop into the daily.dev community Discord for real-time conversations ### What happens after you submit - Every request is read by a human within 48 hours on weekdays - Bug reports get triaged into the next sprint or the public backlog - Larger proposals get a written response from the relevant product owner \[See the rendered feedback page\](https://daily.dev/feedback/) for the upvote panel and the Discord invite. --- --- title: "Uninstall confirmation | daily.dev" url: https://daily.dev/uninstall/ description: "Confirmation page after uninstalling the daily.dev browser extension. Tell us why so we can do better next time, or reinstall in one click." lastUpdated: "2026-04-30" llmsTxt: https://daily.dev/llms.txt --- ## Sorry to see you go This page is the post-uninstall confirmation for the daily.dev browser extension. ### What we ask A two-question feedback form: why you uninstalled, and what would have made you stay. Both optional. ### What we offer - A one-click \[reinstall link\](https://api.daily.dev/get) if you changed your mind - A pointer at our \[iOS\](https://apps.apple.com/app/daily-dev/id6740634400) and \[Android\](https://play.google.com/store/apps/details?id=dev.daily) apps if a different surface fits better - A direct line to \`product@daily.dev\` if there is something specific you want fixed \[See the rendered uninstall page\](https://daily.dev/uninstall/) for the form. --- ## Legal --- title: "Terms of Service | daily.dev" url: https://daily.dev/tos/ description: "The terms governing your use of daily.dev. Accounts, content, intellectual property, liability, governing law, and how disputes are resolved." lastUpdated: "2026-04-30" llmsTxt: https://daily.dev/llms.txt --- ## Terms of Service These are the rules that govern your use of daily.dev. Reading them in full at the \[canonical Terms of Service page\](https://daily.dev/tos/) is the only safe summary; this is a navigation aid, not legal advice. ### What the full document covers - \*\*Account creation and termination\*\*: eligibility, login, and what gets you off the platform - \*\*Acceptable use\*\*: what you can and cannot do with the product - \*\*Content & licensing\*\*: what you grant us when you submit content, and what we grant you when you read others' content - \*\*Intellectual property\*\*: daily.dev marks, third-party marks, and DMCA / takedown procedures - \*\*Disclaimers and limitation of liability\*\*: service is provided "as is"; cap on damages - \*\*Governing law and dispute resolution\*\*: jurisdiction, arbitration, class-action waiver - \*\*Changes\*\*: how we communicate updates and what counts as accepting new terms ### Related documents - \[Privacy center\](https://daily.dev/privacy/): how we handle your data - \[Creators terms\](https://daily.dev/creators-terms/): additional terms for content creators publishing on daily.dev - \[Cookies policy\](https://www.iubenda.com/privacy-policy/14695236/cookie-policy): what we set and why \[See the rendered Terms of Service page\](https://daily.dev/tos/) for the binding legal text. --- --- title: "Refund Policy | daily.dev" url: https://daily.dev/refunds/ description: "The 30-day refund policy for daily.dev Plus. Eligibility, how to request, cancellations, and what is and is not covered." lastUpdated: "2025-02-18" llmsTxt: https://daily.dev/llms.txt --- ## Refund Policy for daily.dev Plus Every daily.dev Plus plan is covered by a 30-day hassle-free refund policy. If daily.dev Plus isn't the right fit, you can request a full refund within 30 days of your initial purchase. No questions asked. ### Eligibility - 30-day window from initial purchase, full refund to the original payment method - First-purchase only: renewal payments (monthly or annual) are non-refundable - After the 30-day window: cancel anytime to stop future charges, but no partial refunds for unused time ### How to request - Email \`support@daily.dev\`, or contact Paddle (our payment processor) at \`help@paddle.com\` - Include your account details so we can process the request quickly - Approved refunds are credited back within 5 to 10 business days ### Cancellation vs refund - Cancelling stops future renewals and keeps access until the end of the current billing period - A cancellation by itself does NOT trigger a refund unless the request is made inside the 30-day window - Manage subscription from \`daily.dev\` account settings ### Refunds are not issued for - Requests after the 30-day window - Subscription renewals - Partial refunds for unused time after cancellation - Requests tied to a policy violation (Terms of Service breach, platform misuse) ### Related documents - \[Terms of Service\](https://daily.dev/tos/): base terms governing daily.dev usage - \[Privacy center\](https://daily.dev/privacy/): how we handle your data \[See the rendered Refund Policy page\](https://daily.dev/refunds/) for the binding policy text. --- --- title: "Creators Terms | daily.dev" url: https://daily.dev/creators-terms/ description: "Additional terms for creators publishing content on daily.dev. Licensing, attribution, content guidelines, monetization, and takedowns." lastUpdated: "2026-04-30" llmsTxt: https://daily.dev/llms.txt --- ## Creators Terms The Creators Terms supplement the \[main Terms of Service\](https://daily.dev/tos/) for anyone publishing content on daily.dev: sources, squads, contributors, and partner publications. ### What the full document covers - \*\*Licensing\*\*: the rights you grant daily.dev when you submit content, and what readers can do with it - \*\*Attribution\*\*: how authors and original sources are credited - \*\*Content guidelines\*\*: quality bar, prohibited content, AI-generated disclosure rules - \*\*Monetization\*\*: revenue share for partners and contributors where applicable - \*\*Takedowns\*\*: how we handle requests to remove content, both yours and third parties ### Related documents - \[Terms of Service\](https://daily.dev/tos/): base terms (these creators terms add to those) - \[Privacy center\](https://daily.dev/privacy/): how we handle data, including from contributors \[See the rendered Creators Terms page\](https://daily.dev/creators-terms/) for the binding legal text. --- --- title: "Privacy Center | daily.dev" url: https://daily.dev/privacy/ description: "Privacy hub for daily.dev: the privacy policy and cookie policy covering the website, web app, browser extensions and mobile apps, plus data requests. Authoritative copies hosted on iubenda." lastUpdated: "2026-08-18" llmsTxt: https://daily.dev/llms.txt --- ## daily.dev Privacy Center A single hub for every privacy and data document we publish. The authoritative copies live on iubenda; we link out so you always read the latest counsel-reviewed version. ### Policies One privacy policy and one cookie policy cover every daily.dev surface: the website, the web app, the browser extensions, and the mobile apps. - \[Privacy Policy\](https://www.iubenda.com/privacy-policy/14695236): what we collect, why, who we share it with, and your rights - \[Cookie Policy\](https://www.iubenda.com/privacy-policy/14695236/cookie-policy): analytics, marketing and feature-flag cookies, and cookies set inside the apps - \[Permissions Explained\](https://daily.dev/blog/permissions-explained-daily-dev-browser-extension/): plain-English breakdown of every browser-extension permission we request ### Data subject requests Email \`privacy@daily.dev\` for any GDPR / CCPA request: access, deletion, portability, or correction. ### Related documents - \[Terms of Service\](https://daily.dev/tos/) - \[Creators Terms\](https://daily.dev/creators-terms/) \[See the rendered privacy hub\](https://daily.dev/privacy/) for the full overview. --- ## AI Handbook — The Agentic AI Hub # The Agentic AI Hub > A software engineer's complete reference to AI, from foundations to agents. Version 1.1 · Last updated 2026-07-22 · License CC BY 4.0 · Published by daily.dev Every page below is available as LLM-friendly markdown via its \`.md/\` URL. For rendered HTML, omit \`.md\`. ## Part I. Foundations - \[How LLMs Actually Work\](https://daily.dev/agentic-ai-hub/how-llms-actually-work.md/): A large language model predicts the next token given the tokens so far, and that is essentially all it does. - \[Core Concepts\](https://daily.dev/agentic-ai-hub/core-concepts.md/): Quick definitions for the vocabulary you'll run into everywhere. Each one is short, and several get a fuller treatment later in the handbook. - \[Prompting & Context Engineering\](https://daily.dev/agentic-ai-hub/prompting-context-engineering.md/): Prompting is the highest-leverage, lowest-cost skill in this whole handbook. Small changes in how you frame a request move quality more than most people expect. - \[Safety, Alignment & Governance\](https://daily.dev/agentic-ai-hub/safety-alignment-governance.md/): Building with AI means owning its failure modes. This section is deliberately neutral and tight, enough to reason about risk without ideology. ## Part II. Models, Labs & Benchmarks - \[The Labs: Who's Who\](https://daily.dev/agentic-ai-hub/the-labs-whos-who.md/): Snapshot as of July 19, 2026\. Valuations and funding are dated and sourced. - \[Frontier Model Comparison\](https://daily.dev/agentic-ai-hub/frontier-model-comparison.md/): A dated snapshot of the current top-tier, closed ("frontier") models, meaning the API-only models from the major labs that set the capability bar. - \[Open-Weight Models\](https://daily.dev/agentic-ai-hub/open-weight-models.md/): The models you can download, self-host, and fine-tune. As of July 2026\. Two things to know: (1) "open weight" ≠ open source. - \[Specialized Models\](https://daily.dev/agentic-ai-hub/specialized-models.md/): Beyond the general-purpose flagships, each modality and job has its own leaders, and the "best" one shifts constantly, so treat these as categories to shop in with current examples. - \[Benchmarks & Leaderboards\](https://daily.dev/agentic-ai-hub/benchmarks-leaderboards.md/): The single most useful mental model here is that benchmarks tell you what a model can do on someone else's task. - \[Pricing & Cost Reference\](https://daily.dev/agentic-ai-hub/pricing-cost-reference.md/): Pricing is a Frontier Model Comparison snapshot. This section is the durable part: how to reason about cost so you're not surprised by a bill. - \[How to Choose a Model\](https://daily.dev/agentic-ai-hub/how-to-choose-a-model.md/): The honest answer is to shortlist from benchmarks and reputation, then run a small eval on your actual task. ## Part III. Tools, Products & Agents - \[Chat Assistants & Apps\](https://daily.dev/agentic-ai-hub/chat-assistants-apps.md/): The consumer/prosumer chat apps are the front door to each lab's frontier model. - \[Coding Tools & CLI Agents\](https://daily.dev/agentic-ai-hub/coding-tools-cli-agents.md/): Two families here: IDE assistants (autocomplete + chat + in-editor agents living in your editor) and CLI/terminal agents (you talk to an agent in the shell. - \[Agent Harnesses & Frameworks\](https://daily.dev/agentic-ai-hub/agent-harnesses-frameworks.md/): Two very different things get called "agent tooling." - \[No-Code & App Builders\](https://daily.dev/agentic-ai-hub/no-code-app-builders.md/): Prompt-to-app tools turn a description into a running web app with UI, backend, and deploy. - \[Voice & Real-Time Agents\](https://daily.dev/agentic-ai-hub/voice-real-time-agents.md/): Real-time voice agents = STT (speech-to-text) → LLM → TTS (text-to-speech) stitched into a low-latency loop (or, increasingly, a single speech-to-speech model). - \[Browser & Computer-Use Agents\](https://daily.dev/agentic-ai-hub/browser-computer-use-agents.md/): These agents perceive a screen or DOM and act (click, type, scroll, navigate) to complete tasks a human would do in a browser or OS. - \[Multimodal Creative Tools\](https://daily.dev/agentic-ai-hub/multimodal-creative-tools.md/): Image, video, and audio generation. Quality is now high across the board. Choose by control, licensing, and API availability more than raw "wow." - \[MCP & the Tool/Context Ecosystem\](https://daily.dev/agentic-ai-hub/mcp-tool-context-ecosystem.md/): Models are only half the story. The other half is the plumbing that gives models hands: how an assistant reaches your files, databases, SaaS tools, and the web. ## Part IV. Building with AI: The Engineering Stack - \[RAG & Retrieval\](https://daily.dev/agentic-ai-hub/rag-retrieval.md/): Retrieval-Augmented Generation (RAG) is the practice of fetching relevant text (or other data) at query time and injecting it into the model's context so the model can answer from information it wasn't trained on. - \[Vector Databases & Memory\](https://daily.dev/agentic-ai-hub/vector-databases-memory.md/): A vector database stores embeddings and serves approximate nearest-neighbor (ANN) search, which finds the closest vectors to a query vector quickly using index structures like HNSW (graph-based, the most common) or IVF/IVF-PQ (cluster + quantize, memory-efficient at scale). - \[Fine-Tuning & Post-Training\](https://daily.dev/agentic-ai-hub/fine-tuning-post-training.md/): Reach for the cheapest tool that works, in this order: - \[Dataset Engineering\](https://daily.dev/agentic-ai-hub/dataset-engineering.md/): Data quality, not model choice, is usually the ceiling on a fine-tune. - \[Inference, Serving & Optimization\](https://daily.dev/agentic-ai-hub/inference-serving-optimization.md/): TTFT (Time To First Token). Latency until the first token appears. Dominated by the prefill (prompt-processing) phase and prompt length. - \[AI Hardware & Accelerators\](https://daily.dev/agentic-ai-hub/ai-hardware-accelerators.md/): The 2026 picture is that NVIDIA still dominates training and general inference. - \[APIs, Inference Platforms & Gateways\](https://daily.dev/agentic-ai-hub/apis-inference-platforms-gateways.md/): Four rough tiers, each with different trade-offs on model access, speed, price, and control. - \[Local & Self-Hosting Stack\](https://daily.dev/agentic-ai-hub/local-self-hosting-stack.md/): Running models on your own machine or servers. - \[LLMOps: Evals, Observability & Guardrails\](https://daily.dev/agentic-ai-hub/llmops-evals-observability-guardrails.md/): Shipping LLM features without evals and observability is flying blind: outputs are non-deterministic, quality is subjective, and regressions are silent. ## Part V. Working with AI: Techniques, Workflows & Adoption - \[Levels of AI Adoption\](https://daily.dev/agentic-ai-hub/levels-of-ai-adoption.md/): Before you can improve how your team works with agents, you need a shared way to say where you are. - \[Agentic Engineering: Core Ideas\](https://daily.dev/agentic-ai-hub/agentic-engineering-core-ideas.md/): "Agentic coding" is an overloaded term. - \[Orchestration Patterns\](https://daily.dev/agentic-ai-hub/orchestration-patterns.md/): Once one agent works, the obvious move is to run several. This is where the biggest gains and the biggest self-inflicted wounds both live. - \[Context, Skills & Memory\](https://daily.dev/agentic-ai-hub/context-skills-memory.md/): Agents are only as good as what they know when they start, and by default they start knowing nothing about your repo, your conventions, or last week's decisions. - \[Verification & Testing for Agents\](https://daily.dev/agentic-ai-hub/verification-testing-for-agents.md/): If generation is cheap and verification is the bottleneck, then verification infrastructure is your leverage. - \[Human Factors & the Agent-Era Career\](https://daily.dev/agentic-ai-hub/human-factors-agent-era-career.md/): The techniques above make you faster. This chapter is about what they can quietly cost you, and how not to pay it. ## Part VI. Ecosystem & Staying Current - \[People to Follow (X/Twitter)\](https://daily.dev/agentic-ai-hub/people-to-follow.md/): X/Twitter is where much of AI moves in real time. Papers get discussed hours before the press notices, and model launches often break here first. - \[Newsletters\](https://daily.dev/agentic-ai-hub/newsletters.md/): Email survives because it's async and curated. Tagged by focus and cadence. - \[Blogs & Publications\](https://daily.dev/agentic-ai-hub/blogs-publications.md/): OpenAI Blog: openai.com/news. Launches and research. The announcement of record. - \[Podcasts & YouTube\](https://daily.dev/agentic-ai-hub/podcasts-youtube.md/): Latent Space (swyx & Alessio): latent.space/podcast. Best for: the AI-engineering discipline. - \[Communities\](https://daily.dev/agentic-ai-hub/communities.md/): Where real-time, tacit knowledge lives, the kind that never makes it into a blog post. - \[Key Papers & Reading List\](https://daily.dev/agentic-ai-hub/key-papers-reading-list.md/): A chronological spine of the field. The ones tagged Start here or Landmark are the priority reads; the rest fill in context. - \[Platforms\](https://daily.dev/agentic-ai-hub/platforms.md/): The infrastructure you'll actually live in. - \[Events & Conferences\](https://daily.dev/agentic-ai-hub/events-conferences.md/): 🔄 Dates and formats change yearly. Always confirm on the official site. - \[Learning Paths, Courses & Books\](https://daily.dev/agentic-ai-hub/learning-paths-courses-books.md/): Organized beginner → advanced. Pick a lane and go deep. Breadth comes with time. ## Part VII. Appendices - \[AI Adoption Stats & Trends\](https://daily.dev/agentic-ai-hub/ai-adoption-stats-trends.md/): 🔄 Fast-moving section. All figures are dated and sourced; surveys are annual and market forecasts get revised, so check the live sources for the latest. - \[Glossary of Key Terms\](https://daily.dev/agentic-ai-hub/glossary.md/): A jargon-buster for the terms used throughout this handbook. Where a full section exists, terms link to it. --- --- title: "How LLMs Actually Work" url: https://daily.dev/agentic-ai-hub/how-llms-actually-work/ description: "A large language model predicts the next token given the tokens so far, and that is essentially all it does." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- A large language model predicts the next token given the tokens so far, and that is essentially all it does. Everything else, including writing code, planning a trip, and passing the bar exam, is emergent behavior that falls out of getting extremely good at that single objective over a large enough model and corpus. \*\*Next-token prediction.\*\* The model consumes a sequence of tokens and outputs a probability distribution over what comes next across its entire vocabulary. Sampling from that distribution (greedily, or with temperature/top-p randomness) gives you one token. Append it, feed the whole thing back in, and repeat. Fluent paragraphs are just this loop run hundreds of times. The important intuition is that the model has no plan for the sentence it's about to write. It commits one token at a time, and its apparent foresight comes from having learned, during training, what well-formed continuations look like. This is also why telling a model to "think step by step" used to help so much. You were handing it more tokens to reason over before it had to commit. Today's reasoning models do that internally by default, so the explicit instruction matters much less than it once did (more in \[Prompting & Context Engineering\](/agentic-ai-hub/prompting-context-engineering/)). \*\*Training vs. inference.\*\* These are two different regimes. \*Training\* is where the weights change. The model sees text, predicts the next token, is scored against the actual next token, and its billions of parameters are nudged via gradient descent to be a little less wrong. This happens once (well, in expensive batches) and costs enormous compute. \*Inference\* is what happens every time you send a prompt. The weights are frozen, and the model just runs the forward pass to produce tokens. When people say a model "learned" something from your conversation, they're usually wrong. Nothing about your chat changed the weights. The model only "knows" your conversation because it's sitting in the context window (more on that below). \*\*Pre-training → post-training.\*\* A modern assistant is built in stages: - \*\*Pre-training\*\* is the giant next-token-prediction run over a broad slice of the internet, books, and code. It produces a \*base model\*, a raw next-token predictor that's shockingly knowledgeable but not helpful. Ask it a question and it might continue with more questions, because that's a plausible continuation of text on the web. - \*\*Supervised fine-tuning (SFT)\*\* teaches the base model to behave like an assistant by training it on curated examples of instructions paired with good responses. This is where "answer the question" becomes the default behavior. - \*\*RLHF (Reinforcement Learning from Human Feedback)\*\* refines it further using human preference judgments. Humans (or a model trained to imitate them) rank candidate responses. A reward model learns those preferences, and the LLM is optimized to produce responses the reward model scores highly. RLHF is why models are polite, hedge appropriately, and refuse harmful requests. \*\*RLAIF (Reinforcement Learning from AI Feedback)\*\* swaps human labelers for an AI labeler (often guided by a written "constitution" of principles) to scale up the preference data cheaply. \[Anthropic's Constitutional AI work\](https://www.anthropic.com/research/constitutional-ai-harmlessness-from-ai-feedback) is the canonical example. - \*\*DPO (Direct Preference Optimization)\*\* is a more recent, simpler alternative to the full RLHF pipeline: it optimizes the model directly on preference pairs without training a separate reward model or running RL, which makes it cheaper and more stable to tune. See the \[original DPO paper\](https://arxiv.org/abs/2305.18290). - \*\*GRPO and RLVR (the 2026 reasoning wave).\*\* The newest post-training leans on reinforcement learning against \*verifiable\* rewards instead of human preference. \*\*RLVR\*\* (RL from Verifiable Rewards) trains on signals a computer can check automatically, like whether the tests pass or the math is correct, and \*\*GRPO\*\* (Group Relative Policy Optimization, popularized by DeepSeek) is a cheaper RL algorithm that drops the separate value model older methods needed. Together they are the engine behind today's reasoning models. See \[this post-training overview\](https://huggingface.co/blog/karina-zadorozhny/guide-to-llm-post-training-algorithms). Collectively, SFT + preference optimization is \*\*post-training\*\*, the phase that turns a knowledgeable predictor into a usable assistant. A huge fraction of a model's "personality" and instruction-following quality is decided here, not in pre-training. \*\*What "reasoning" / test-time-compute models do differently.\*\* A newer class of models (OpenAI's o-series, DeepSeek-R1, Claude's extended-thinking modes, Gemini's thinking variants) is trained to spend more compute \*at inference time\* by generating a long internal chain of thought before answering. Rather than committing to an answer immediately, they produce (often hidden) intermediate reasoning tokens that explore, check, and backtrack, and this measurably improves performance on math, coding, and multi-step logic. The key shift is economic. Instead of only scaling training compute, you scale \*test-time\* compute, trading latency and tokens for accuracy. These models are typically trained with reinforcement learning against verifiable rewards (did the code pass? is the math correct?), which is why they excel where answers can be checked. The practical takeaway is to use them when a problem genuinely decomposes into steps, and to expect them to be slower and pricier per query. See \[Core Concepts\](/agentic-ai-hub/core-concepts/) for when reasoning is worth it. \*\*Why context ≠ memory.\*\* This trips up nearly everyone. The context window is working memory that exists \*only for the duration of a single request\*. When you have a long chat, the interface re-sends the accumulated transcript on every turn. The model isn't remembering, it's re-reading. Close the session, or overflow the window, and it's gone. Persistent memory (a model that recalls you across sessions) is an \*engineering layer\* on top. Your app stores facts in a database or vector store and re-injects them into context. The model itself is stateless between calls. Internalize this and a lot of behavior stops being mysterious: why the model "forgets" something from 50 messages ago (it fell off the front of the window), why it can't recall your last conversation (nothing persisted it), and why repeating context helps. \*\*Tokenization intuition.\*\* Models don't see characters or words. They see \*tokens\*, subword chunks produced by a tokenizer (typically a byte-pair-encoding scheme). Common English words are often one token. Rarer words split into pieces ("tokenization" → \`token\` + \`ization\`), and whitespace and punctuation carry tokens too. Rough rule of thumb for English: \*\*\~4 characters or \~0.75 words per token\*\*, so 1,000 tokens ≈ 750 words. This has real consequences. Token boundaries are why models historically fumbled "how many r's in strawberry" (the word is a couple of opaque tokens, not individual letters). It's why code and non-English languages can be less token-efficient (more tokens per character means more cost and less effective context). And it's why every pricing and context-limit number is quoted in tokens, not words. When you reason about cost and context budgets, think in tokens. --- --- title: "Core Concepts" url: https://daily.dev/agentic-ai-hub/core-concepts/ description: "Quick definitions for the vocabulary you'll run into everywhere. Each one is short, and several get a fuller treatment later in the handbook." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- Quick definitions for the vocabulary you'll run into everywhere. Each one is short, and several get a fuller treatment later in the handbook. \*\*Context windows & tokenization.\*\* The context window is the maximum number of tokens (prompt + generated output) a model can attend to at once. Frontier models in 2026 commonly offer 100K to 1M+ token windows 🔄. But bigger isn't automatically better. Attention cost grows with length, and models exhibit a well-documented \["lost in the middle" effect\](https://arxiv.org/abs/2307.03172) where information buried in the center of a long context is recalled less reliably than material at the start or end. Treat context as a scarce, curated resource, not a dumping ground. \*\*Parameters vs. active params / MoE.\*\* "Parameters" are the learned weights, the crude proxy for model capacity. A \*\*Mixture-of-Experts (MoE)\*\* model has many parameters but only \*activates a subset\* (a few "experts") per token, so a model with, say, hundreds of billions of total parameters might use only tens of billions per forward pass. So MoE decouples total capacity (which helps quality) from per-token compute (which drives cost and speed). This is why you'll see "total parameters" and "active parameters" quoted separately. The latter is what you actually pay for at inference. \*\*Quantization & precision.\*\* Weights are usually trained in 16-bit precision but can be \*quantized\* down to 8-bit, 4-bit, or lower for inference. This shrinks memory footprint and speeds things up, at some cost to quality. A 4-bit quant of a big model often beats a full-precision small model at the same memory budget. If you're running models locally or self-hosting, quantization is the single biggest lever on what fits your hardware. \*\*Reasoning / test-time compute.\*\* Covered in \[How LLMs Actually Work\](/agentic-ai-hub/how-llms-actually-work/). The mental model is to spend more tokens thinking to buy accuracy on hard, decomposable problems. Not worth it for simple lookups, formatting, or classification, and you'll pay latency for nothing. \*\*Multimodality.\*\* Modern models increasingly ingest (and sometimes emit) more than text, including images, audio, and video, by encoding each modality into the same token stream the language model already understands. In practice, you can hand a model a screenshot, a chart, or a PDF page and ask about it. Native multimodality (trained in from the start) tends to reason across modalities better than bolted-on adapters. \*\*Embeddings & vector search.\*\* An \*embedding\* is a fixed-length vector that captures the meaning of a piece of text (or image), such that semantically similar things land near each other in vector space. Store a corpus of embeddings in a \*\*vector database\*\* and you can retrieve by \*meaning\* rather than keywords, embedding the query and finding its nearest neighbors. This is the retrieval engine under most RAG systems. Full treatment in \[RAG & Retrieval\](/agentic-ai-hub/rag-retrieval/). \*\*RAG.\*\* Retrieval-Augmented Generation grounds a model's answer in documents you fetch at query time (usually via embedding search) and inject into the context window, instead of relying on what it memorized during training. It's the standard way to make a model use your private or current data without retraining it. Full treatment in \[RAG & Retrieval\](/agentic-ai-hub/rag-retrieval/). \*\*Fine-tuning vs. prompting vs. tools.\*\* Three ways to change model behavior, in rough order of effort: - \*Prompting.\* Change what you say. Zero cost, instant, the right first move for most problems. - \*Tools.\* Give the model functions to call (search, code execution, APIs) so it can fetch facts and take actions instead of hallucinating them. The backbone of agents. - \*Fine-tuning.\* Change the weights on your own data. Powerful for style, format, and narrow domains, but costly, slow to iterate, and often unnecessary once prompting + RAG + tools are exhausted. See \[Fine-Tuning & Post-Training\](/agentic-ai-hub/fine-tuning-post-training/). The practical heuristic: reach for fine-tuning last, not first. \*\*Prompt caching.\*\* Many providers let you cache a large, static prefix of your prompt (a system prompt, a big document, a tool schema) so that reusing it across requests is much cheaper and faster than re-processing it every time. If your app sends the same 10K-token preamble on every call, caching it can cut cost and latency dramatically. Design your prompts so the stable parts come first. That's what gets cached. \*\*Structured output.\*\* Instead of free-form prose, you can constrain a model to emit valid JSON matching a schema (via JSON mode, function/tool schemas, or grammar-constrained decoding). This is what makes LLMs usable as reliable components in software. You get parseable data, not a paragraph you have to regex. When an LLM output feeds another program, prefer structured output. Patterns in \[Prompting & Context Engineering\](/agentic-ai-hub/prompting-context-engineering/). --- --- title: "Prompting & Context Engineering" url: https://daily.dev/agentic-ai-hub/prompting-context-engineering/ description: "Prompting is the highest-leverage, lowest-cost skill in this whole handbook. Small changes in how you frame a request move quality more than most people expect." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- Prompting is the highest-leverage, lowest-cost skill in this whole handbook. Small changes in how you frame a request move quality more than most people expect. This section is the \*non-agent\* version, covering single-shot and few-turn interactions. The agentic version, where prompts become skills, memory, and dynamically assembled context, is \[Context, Skills & Memory\](/agentic-ai-hub/context-skills-memory/). For worked examples, the provider prompt-engineering guides are the best reference (\[Anthropic\](https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/overview) and \[OpenAI\](https://platform.openai.com/docs/guides/prompt-engineering)). \*\*System vs. user prompts.\*\* The \*system prompt\* sets durable context: who the model is, what rules it follows, what format to output, what tools it has. The \*user prompt\* is the specific request. Keep role/constraints/format in the system prompt and the actual task in the user prompt. This separation makes behavior more consistent and, conveniently, lets the stable system prefix be cached (see prompt caching, \[Core Concepts\](/agentic-ai-hub/core-concepts/)). Put the most important instructions where they won't get lost, early and restated near the end for long inputs. \*\*Zero-shot vs. few-shot.\*\* \*Zero-shot\* is just asking. \*Few-shot\* means showing a handful of input→output examples before the real query, which anchors format and style far more reliably than describing them in words. If you find yourself writing paragraphs explaining the exact output shape you want, stop and show two or three examples instead. Few-shot is the fastest fix for "the model \*almost\* does what I want." \*\*Chain-of-thought (CoT).\*\* Asking the model to reason step by step before answering improves accuracy on math, logic, and multi-step tasks. This is the \[foundational result\](https://arxiv.org/abs/2201.11903) that reasoning models later industrialized. For non-reasoning models, an explicit "think through this step by step" still helps. For dedicated reasoning models, it's largely automatic and you can often just state the problem. One caveat is that the visible chain of thought is a \*performance aid\*, not a faithful audit log of the model's internal computation, so don't treat it as a guaranteed explanation of \*why\* it answered. \*\*Structured output & tool-use prompting.\*\* To get clean JSON, use the provider's structured-output / tool-calling features rather than begging in prose, and supply the schema. For tool use, describe each tool's purpose and parameters precisely. The model decides when to call based largely on your descriptions, so a vague tool description is a bug. Give it an explicit "if you don't know, call the search tool rather than guessing" and you'll cut hallucinations. \*\*Context engineering: what goes in the window and why.\*\* As tasks get serious, prompting becomes \*context engineering\*: deciding what information occupies the finite, quality-sensitive context window. The discipline is curation, not accumulation. Include what the model needs for \*this\* task (relevant retrieved documents, the necessary examples, current state) and ruthlessly exclude noise. Order matters (put critical material at the edges, not buried in the middle; recall the "lost in the middle" effect from \[Core Concepts\](/agentic-ai-hub/core-concepts/)). More context is not more better. Irrelevant filler measurably degrades performance and costs tokens. \*\*Common failure modes.\*\* - \*Hallucination.\* Confident fabrication. Mitigate with retrieval/tools and by explicitly permitting "I don't know." - \*Instruction drift.\* In long conversations, early instructions fade. Restate the important ones. - \*Overstuffed context.\* Dumping everything in and hoping. Curate instead. - \*Ambiguity.\* The model fills gaps with assumptions. Be specific about format, audience, and constraints. - \*Sycophancy.\* Models tend to agree with a confidently stated premise. Don't lead the witness if you want an honest check. \*\*Prompting cheat-sheet.\*\* | Goal | Technique | Quick note | | :---- | :---- | :---- | | Reliable format | Few-shot examples | Show 2-3, don't just describe | | Better hard reasoning | Chain-of-thought / reasoning model | Let it think before answering | | Parseable output | Structured output / JSON schema | Use the API feature, not prose | | Fewer hallucinations | Tools + RAG + "say I don't know" | Ground it in sources | | Consistent persona/rules | System prompt | Stable prefix, also cacheable | | Steer style | Few-shot + explicit constraints | Examples beat adjectives | | Complex task | Decompose into steps | One job per call when you can | | Cheaper repeated calls | Prompt caching | Stable content first | The meta-skill: \*\*iterate\*\*. Write a prompt, look at failures, fix the specific failure, repeat. Prompting is empirical, not theoretical. --- --- title: "Safety, Alignment & Governance" url: https://daily.dev/agentic-ai-hub/safety-alignment-governance/ description: "Building with AI means owning its failure modes. This section is deliberately neutral and tight, enough to reason about risk without ideology." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- Building with AI means owning its failure modes. This section is deliberately neutral and tight, enough to reason about risk without ideology. \*\*What alignment / RLHF targets.\*\* Alignment is the effort to make a model's behavior track human intent and values: helpful, honest, and harmless, roughly. RLHF and its relatives (RLAIF, DPO, see \[How LLMs Actually Work\](/agentic-ai-hub/how-llms-actually-work/)) are the primary technical levers. They shape the model to prefer responses people judge as good and to refuse ones they judge as harmful. But alignment is imperfect and shallow in important ways. A model can be helpful and honest in ordinary use yet still be steered off the rails by adversarial input. That gap is why the next paragraph matters so much. \*\*Jailbreaks & prompt injection: the critical one for builders.\*\* Two related but distinct threats: - A \*jailbreak\* is a user crafting input that coaxes the model past its safety training (role-play framings, obfuscation, "ignore previous instructions"). - \*Prompt injection\* is more insidious and is \*\*the\*\* security problem for agents and tool use. When your app feeds the model untrusted content (a web page, an email, a document, a tool's output), that content can contain instructions the model follows \*as if they came from you\*. An attacker who controls a web page your agent browses can tell it to exfiltrate data, misuse a connected tool, or take destructive actions. The \[OWASP Top 10 for LLM Applications\](https://genai.owasp.org/llm-top-10/) ranks prompt injection as the number-one risk. The core danger is that \*\*the model doesn't reliably distinguish trusted instructions from untrusted data in its context\*\*. It's all just tokens. The more capable your agent (tools, memory, autonomy, access to real systems), the higher the stakes. Defenses in depth: constrain tool permissions and scopes, keep a human in the loop for consequential actions, sandbox execution, sanitize and clearly delimit untrusted content, and never assume the model will "know better." Treat any model handling untrusted input as a partially-compromised component and design blast radius accordingly. \*\*Red-teaming & evals basics.\*\* \*Red-teaming\* is adversarially probing a system for failures (harmful outputs, jailbreak susceptibility, injection vulnerabilities) before attackers or users find them. \*Evals\* are systematic, repeatable tests of behavior: correctness, safety, refusal rates, regressions between versions. The discipline mirrors software testing. You can't manage what you don't measure, and vibes don't survive contact with production. Build a small eval set early, automate it, and gate releases on it. This is covered practically in the agents and production parts of the handbook. \*\*The regulatory & governance landscape as of mid-2026\*\* 🔄. This moves fast, so verify the current status before relying on any specific date. The short version: if you ship AI, three separate things can constrain you, a law in the EU, a patchwork of rules in the US, and your own company's internal policies. - \*\*EU AI Act (Europe's main AI law).\*\* It sorts AI systems by how risky they are and puts the heaviest duties on the riskiest uses, and it is rolling out in phases. The worst uses (like government social scoring) were banned in February 2025\. Rules for providers of general-purpose models began in \[August 2025\](https://ai-act-service-desk.ec.europa.eu/en/ai-act/timeline/timeline-implementation-eu-ai-act). From \*\*August 2, 2026\*\*, you have to tell users when they are talking to an AI and label AI-generated images, audio, and video (\[summary\](https://www.digitalapplied.com/blog/eu-ai-act-august-2026-transparency-obligations-agency-checklist)). The heaviest "high-risk" requirements were pushed back by a mid-2026 simplification package, to \[December 2027 and August 2028\](https://www.insideglobaltech.com/2026/05/28/eu-ai-act-update-timeline-relief-targeted-simplification-and-new-prohibitions/). \*\*What it means for you:\*\* if you have EU users, plan for the transparency and labeling duties now. The heavier compliance paperwork for high-risk uses comes later, but it is coming. - \*\*United States (no single federal law).\*\* The US picture is unsettled. In December 2025 the White House issued an \[executive order\](https://www.whitehouse.gov/fact-sheets/2025/12/fact-sheet-president-donald-j-trump-ensures-a-national-policy-framework-for-artificial-intelligence/) pushing a light-touch, pro-industry approach and moving to challenge stricter state laws (\[analysis\](https://www.orrick.com/en/Insights/2025/12/5-Things-to-Know-About-Trumps-AI-Executive-Order)). But individual states still have their own AI laws on the books (California's SB 53, Colorado's SB 205), and many state attorneys general are pushing back. \*\*What it means for you:\*\* there is no one national rulebook. Requirements vary by state and are likely to keep shifting through 2026\. 🔄 - \*\*Enterprise governance (your own company's rules).\*\* Separate from any law, most companies set their own internal AI policies: which tools are allowed, how models and vendors get reviewed, how customer data and personal information are handled, when a human has to sign off, and what gets logged. Frameworks like the \[NIST AI Risk Management Framework\](https://www.nist.gov/itl/ai-risk-management-framework) are common starting points. \*\*What it means for you:\*\* inside a company, this internal governance is usually the real gate that blocks or clears a launch, so plan for it early instead of treating it as an afterthought. The throughline of this section is that capability and risk scale together. The same autonomy that makes an agent useful makes it dangerous when fed hostile input or pointed at real systems. Design for that from the start. --- --- title: "The Labs: Who's Who" url: https://daily.dev/agentic-ai-hub/the-labs-whos-who/ description: "Snapshot as of July 19, 2026\. Valuations and funding are dated and sourced." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- \*Snapshot as of July 19, 2026\. Valuations and funding are dated and sourced. 🔄\* \*\*How this is organized:\*\* grouped by what developers actually reach for each lab for. The \*\*Open/closed\*\* column tells you whether you can download the weights and run them yourself (open) or only reach the model through an API (closed). Several labs do both, so it reflects their main stance. | Lab | HQ | Open/closed | Flagship (Jul 2026) | Website | Backing / latest valuation | | :---- | :---- | :---- | :---- | :---- | :---- | | \*\*OpenAI\*\* | USA | Closed | GPT-5.6 family | \[openai.com\](https://openai.com) | \~$852B (Mar 2026) | | \*\*Anthropic\*\* | USA | Closed | Claude Fable 5 (+ Opus 4.8 / Sonnet 5) | \[anthropic.com\](https://anthropic.com) | \~$965B (May 2026) | | \*\*Google DeepMind\*\* | UK / USA | Mixed | Gemini 3.1 Pro (+ 3.6 Flash, 3.5 Flash-Lite, Gemma) | \[deepmind.google\](https://deepmind.google) | Alphabet subsidiary | | \*\*xAI\*\* | USA | Closed | Grok 4.5 | \[x.ai\](https://x.ai) | Part of SpaceX (\~$1.25T combined) | | \*\*Microsoft AI\*\* | USA | Mixed | MAI-Thinking-1 (+ Phi) | \[microsoft.com/ai\](https://www.microsoft.com/ai) | Microsoft division | | \*\*Amazon\*\* | USA | Closed | Nova 2 | \[aws.amazon.com/nova\](https://aws.amazon.com/nova/) | Amazon division | | \*\*Meta (MSL)\*\* | USA | Open → closed | Muse Spark 1.1 (closed) + open Llama 4 | \[ai.meta.com\](https://ai.meta.com) | Meta division | | \*\*Mistral\*\* | France | Open + closed | Medium 3.5 + Large 3 (both open) | \[mistral.ai\](https://mistral.ai) | \~€11.7B (2025) | | \*\*DeepSeek\*\* | China | Open (MIT) | DeepSeek V4 | \[deepseek.com\](https://www.deepseek.com) | >$50B (Jun 2026) | | \*\*Alibaba (Qwen)\*\* | China | Open + closed | Qwen3.8-Max (open line: Qwen3.6) | \[qwen.ai\](https://qwen.ai) | Alibaba division | | \*\*Moonshot (Kimi)\*\* | China | Open | Kimi K3 (2.8T) | \[moonshot.ai\](https://www.moonshot.ai) | \~$20B (May 2026) | | \*\*Zhipu / Z.ai\*\* | China | Open | GLM-5.2 | \[z.ai\](https://z.ai) | Public, HK IPO \~$6.6B (Jan 2026) | | \*\*Cohere\*\* | Canada | Mixed | Command A+ / Embed 4 | \[cohere.com\](https://cohere.com) | \~$6.8B (2025) | | \*\*AI21 Labs\*\* | Israel | Hybrid | Jamba | \[ai21.com\](https://www.ai21.com) | \~$1.4B (2023) | | \*\*Reka\*\* | USA | Hybrid | Reka Flash / Core | \[reka.ai\](https://reka.ai) | \~$1B (2025) | ## Frontier labs - \*\*OpenAI.\*\* The default general-purpose API and the largest tooling ecosystem. - \*\*Anthropic.\*\* A leading model for coding and agentic/tool-use workloads; safety-focused. - \*\*Google DeepMind.\*\* The largest context windows and deep integration across Cloud, Android, and Workspace. - \*\*xAI.\*\* Real-time, X-connected models; now part of publicly-traded SpaceX (branded "SpaceXAI") after the 2026 merger. - \*\*Microsoft AI.\*\* First-party "MAI" models on Azure/Foundry, alongside its OpenAI partnership. It also ships the open Phi small models. - \*\*Amazon (Nova).\*\* Price-performance models tightly integrated with AWS and Bedrock. ## Open-weight leaders - \*\*Meta (MSL).\*\* Llama is still a self-hosting/fine-tuning default, but the lab is pivoting toward closed frontier models (Muse Spark 1.1). - \*\*Mistral AI.\*\* The go-to European, permissively-licensed, self-hostable option. - \*\*DeepSeek.\*\* Frontier-adjacent quality at rock-bottom cost under a permissive MIT license. - \*\*Alibaba (Qwen).\*\* The most common open base for fine-tuning and edge deployment, with the broadest range of sizes. - \*\*Moonshot (Kimi).\*\* Frontier-competitive open weights for coding and agents (Kimi K3). - \*\*Zhipu / Z.ai (GLM).\*\* Strong open coding and agentic models from a now-public, well-capitalized vendor. ## Enterprise & specialist - \*\*Cohere.\*\* Strong retrieval/embedding models and on-prem deployment for regulated industries. - \*\*AI21 Labs.\*\* Efficient long-context (Jamba) for enterprise RAG. - \*\*Reka.\*\* Deployable, efficient multimodal models plus an openly downloadable reasoning model. --- --- title: "Frontier Model Comparison" url: https://daily.dev/agentic-ai-hub/frontier-model-comparison/ description: "A dated snapshot of the current top-tier, closed (\\"frontier\\") models, meaning the API-only models from the major labs that set the capability bar." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- > A dated snapshot of the current top-tier, \*\*closed\*\* ("frontier") models, meaning the API-only models from the major labs that set the capability bar. Open-weight models (DeepSeek, Kimi, Qwen, Mistral, GLM) are in \[Open-Weight Models\](/agentic-ai-hub/open-weight-models/). For their hosted API prices, check \[OpenRouter\](https://openrouter.ai/models). It's here for the at-a-glance shape of the field, not as a live price sheet: prices and versions move weekly, so use the \[leaderboards\](/agentic-ai-hub/benchmarks-leaderboards/) for current specs, speed, and prices. \*\*As of July 19, 2026.\*\* Prices are standard-tier API, USD per 1M tokens, short-context base rate (long-context tiers typically cost \~2×). Output speed (tok/s) is essentially never an official spec, so check \[Artificial Analysis\](https://artificialanalysis.ai/) for measured throughput. Update (July 21, 2026): Google shipped Gemini 3.6 Flash (reflected below) alongside Gemini 3.5 Flash-Lite (cheaper, faster) and a security-only 3.5 Flash Cyber; Gemini 3.5 Pro is still in partner testing and behind schedule, and Gemini 4 pre-training has begun. | Model | Provider | Released | Context | Modalities (in) | Reasoning | In $/1M | Out $/1M | | :---- | :---- | :---- | :---- | :---- | :---- | :---- | :---- | | GPT-5.6 "Sol" (flagship) | OpenAI | Jul 2026 | 1M | text, image | Yes | $5.00 | $30.00 | | GPT-5.6 "Terra" (balanced) | OpenAI | Jul 2026 | \~1M | text, image | Yes | $2.50 | $15.00 | | GPT-5.6 "Luna" (fast) | OpenAI | Jul 2026 | \~1M | text, image | Yes | $1.00 | $6.00 | | Claude Fable 5 (flagship) | Anthropic | 2026 | 1M | text, image | Yes | $10.00 | $50.00 | | Claude Opus 4.8 | Anthropic | 2026 | 1M | text, image | Yes | $5.00 | $25.00 | | Claude Sonnet 5 | Anthropic | 2026 | 1M | text, image | Yes | $2.00 ¹ | $10.00 ¹ | | Claude Haiku 4.5 | Anthropic | 2026 | 1M | text, image | Yes | $1.00 | $5.00 | | Gemini 3.1 Pro (flagship) | Google | Feb 2026 | \~1M | text, image, audio, video | Yes | \~$2.00 ² | \~$12.00 ² | | Gemini 3.6 Flash | Google | Jul 2026 | \~1M | text, image, audio, video | Yes | \~$1.50 | \~$7.50 | | Grok 4.5 | xAI | 2026 | 500K | text, image | Yes | $2.00 | $6.00 | | Command A Plus | Cohere | May 2026 | 128K | text | No | $2.50 | $10.00 | | Amazon Nova 2 Pro | Amazon | 2026 | 1M | multimodal | n/a | \~$2.50 | \~$12.50 | ¹ Sonnet 5 price rises to $3 / $15 on Sept 1, 2026\. ² Gemini 3.1 Pro is tiered by context length (≤200K shown). Verify on the \[Gemini pricing page\](https://ai.google.dev/gemini-api/docs/pricing). Most flagships are now hybrid reasoning models (a toggleable "thinking" mode). Sources: \[OpenAI\](https://developers.openai.com/api/docs/pricing) · \[Anthropic\](https://platform.claude.com/docs/en/about-claude/pricing) · \[Google\](https://ai.google.dev/gemini-api/docs/pricing) · \[xAI\](https://x.ai/api). \*\*Live sources (bookmark these instead of trusting the table for exact numbers):\*\* \[Artificial Analysis\](https://artificialanalysis.ai/) (intelligence + speed + price + latency), \[OpenRouter models\](https://openrouter.ai/models) (live per-provider pricing & real throughput), \[llm-stats.com\](https://llm-stats.com/) (spec sheets with pricing baked in). --- --- title: "Open-Weight Models" url: https://daily.dev/agentic-ai-hub/open-weight-models/ description: "The models you can download, self-host, and fine-tune. As of July 2026\. Two things to know: (1) \\"open weight\\" ≠ open source." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- > The models you can download, self-host, and fine-tune. \*\*As of July 2026.\*\* Two things to know: (1) \*"open weight" ≠ open source\*. Some ship under permissive licenses (Apache-2.0 / MIT), others under custom, use-restricted ones, so the \*\*License\*\* column matters. Check it before you build on a model. (2) The open frontier is now \*\*led by Chinese labs\*\* (DeepSeek, Qwen, Kimi, GLM), while Meta has pivoted toward closed models. | Model (latest open) | Released | Params (total / active) | License | Context | Best at | Run with | | :---- | :---- | :---- | :---- | :---- | :---- | :---- | | \*\*DeepSeek V4\*\* (Pro / Flash) | Apr 2026 | Pro 1.6T/49B · Flash 284B/13B | MIT ✅ | 1M | Reasoning, coding, cheap inference | vLLM, SGLang | | \*\*Qwen3.6\*\* (35B-A3B) | Apr 2026 | 35B / 3B active | Apache-2.0 ✅ | 262K (\~1M YaRN) | Agentic coding, repo-level reasoning | Ollama, vLLM, SGLang | | \*\*Kimi K3\*\* (Moonshot) | Jul 2026 | 2.8T / \~50B active | Modified MIT ✅ | 1M | Coding/agentic, near-Opus-4.8 | vLLM, SGLang | | \*\*GLM-5.2\*\* (Zhipu / Z.ai) | 2026 | \~744B / 40B active | MIT ✅ | \~1M | Coding/agentic/reasoning | vLLM, SGLang | | \*\*Llama 4 Scout\*\* (Meta) | Apr 2025 | 109B / 17B active | Llama 4 Community ⚠️ | 10M | Very long context, single-GPU | Ollama, vLLM | | \*\*Llama 4 Maverick\*\* (Meta) | Apr 2025 | 400B / 17B active | Llama 4 Community ⚠️ | 1M | Multimodal, general chat | vLLM | | \*\*Mistral Large 3\*\* | Dec 2025 | 675B / 41B active | Apache-2.0 ✅ | Large | Multimodal, multilingual | vLLM | | \*\*Ministral 3\*\* (3/8/14B) | Dec 2025 | dense | Apache-2.0 ✅ | n/a | Best small cost/perf, edge | Ollama, llama.cpp | | \*\*Gemma 4\*\* (Google) | Apr 2026 | 12B (+ MoE 26B-A4B) | Apache-2.0 ✅ | 256K | On-device → server, multimodal | Ollama, llama.cpp, MLX | | \*\*Qwen3\*\* (0.6B-235B) | 2025 | dense + MoE | Apache-2.0 ✅ | 32K-128K+ | Widest size range, 119 languages | Ollama, vLLM | | \*\*Phi-4\*\* family (Microsoft) | 2024-26 | 3.8B-\~15B | MIT ✅ | 16K-128K | Reasoning-per-parameter, on-device | Ollama, llama.cpp | ✅ permissive · ⚠️ source-available with use restrictions (check the model card). Rough VRAM rule of thumb: a dense model needs \~(params × bytes-per-param) of VRAM, e.g. a 27B model at 4-bit quantization ≈ 16 GB. MoE models load all experts into memory but only compute the active set, so they're fast but still memory-hungry. Sources: \[Qwen\](https://huggingface.co/Qwen) · \[DeepSeek\](https://github.com/deepseek-ai/deepseek-v3) · \[Mistral\](https://mistral.ai/news/mistral-3/) · \[Gemma\](https://ai.google.dev/gemma/docs/releases) · \[Kimi\](https://huggingface.co/moonshotai) · \[GLM\](https://huggingface.co/zai-org/GLM-5) · \[Llama\](https://ai.meta.com/blog/). \*\*Track open models live:\*\* \[Hugging Face trending\](https://huggingface.co/models) (the canonical release feed; the original Open LLM Leaderboard was archived in 2025), \[LMArena\](https://lmarena.ai/) (includes open models), \[llm-stats.com\](https://llm-stats.com/llm-updates). --- --- title: "Specialized Models" url: https://daily.dev/agentic-ai-hub/specialized-models/ description: "Beyond the general-purpose flagships, each modality and job has its own leaders, and the \\"best\\" one shifts constantly, so treat these as categories to shop in with current examples." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- > Beyond the general-purpose flagships, each modality and job has its own leaders, and the "best" one shifts constantly, so treat these as \*categories to shop in\* with current examples. The tooling for building on top of them lives in \[Part IV\](/agentic-ai-hub/#engineering-stack). Creative tools get their own deep-dive in \[Multimodal Creative Tools\](/agentic-ai-hub/multimodal-creative-tools/). \*\*As of July 2026.\*\* ## Coding models "Best coding model" is now mostly a general-flagship question, but the harness matters as much as the weights (\[Agent Harnesses & Frameworks\](/agentic-ai-hub/agent-harnesses-frameworks/)), and open models are genuinely competitive. - \*\*Closed leaders:\*\* Claude Opus 4.8 / Sonnet 5 (agentic coding, tool use), GPT-5.6 (Codex variant), Claude Fable 5 (premium). Gemini is capable but less commonly reached for by developers for hands-on coding. - \*\*Open leaders:\*\* the open-weight models in \[Open-Weight Models\](/agentic-ai-hub/open-weight-models/), notably Kimi K3, Qwen3.6-Coder, DeepSeek V4, and GLM-5.2, are genuinely competitive on real-repo tasks (see that table for licenses and sizes). - \*\*Specialized:\*\* Cursor's in-house model (tuned for fast in-editor edits), xAI \`grok-build\` (early), and fill-in-the-middle models for autocomplete inside IDE tools. - \*\*Judge on\*\* \[SWE-bench Verified\](https://www.swebench.com/verified.html), \[Terminal-Bench\](https://www.tbench.ai/), and LMArena's WebDev Arena, \*\*not\*\* HumanEval (saturated). ## Embedding & reranking models The backbone of \[RAG\](/agentic-ai-hub/rag-retrieval/) and search. Pick on retrieval quality \*for your domain\* (run a small eval), then on dimensions (storage/latency cost) and max input length. | Model | Type | Open? | Notes | | :---- | :---- | :---- | :---- | | OpenAI (text-embedding-3-large) | Embedding | No | Strong default, adjustable dimensions | | Cohere Embed 4 | Embedding | No | Multilingual, multimodal, long inputs | | Voyage-3 (voyage-3-large) | Embedding | No | Top retrieval quality, domain variants (code, finance, law) | | Gemini Embedding | Embedding | No | Google-ecosystem, strong MTEB scores | | BGE-M3 / GTE-Qwen / Nomic-Embed / E5 | Embedding | ✅ | Best open options; self-host to cut cost | | Cohere Rerank 3.5 / Voyage Rerank / BGE-reranker / Jina | Reranker | mixed | Big precision boost as a cheap second stage | ## Image & video models Fast-moving, and covered with leaders, licensing, and API notes in \[Multimodal Creative Tools\](/agentic-ai-hub/multimodal-creative-tools/). In brief: image leaders span Midjourney (aesthetics), FLUX.2 (open weight), Ideogram 3 (text-in-image), Adobe Firefly (commercially indemnified), and native GPT/Gemini generation; video leaders span Sora 2, Veo 3 (native audio), Runway, and Kling, with open Alibaba Wan. Judge on control, licensing, and API access rather than raw "wow," and watch clip length and resolution for video. ## Audio & voice models The model layer under the \[voice agents\](/agentic-ai-hub/voice-real-time-agents/) in \[Voice & Real-Time Agents\](/agentic-ai-hub/voice-real-time-agents/) and the \[music tools\](/agentic-ai-hub/multimodal-creative-tools/) in \[Multimodal Creative Tools\](/agentic-ai-hub/multimodal-creative-tools/). - \*\*TTS:\*\* ElevenLabs (quality/voice cloning), Cartesia Sonic (low latency), Deepgram Aura, plus provider TTS (OpenAI, Google). - \*\*STT:\*\* Deepgram, AssemblyAI, OpenAI Whisper (open, self-hostable). - \*\*Speech-to-speech / realtime:\*\* provider realtime APIs, Amazon Nova Sonic, the basis for \[voice agents\](/agentic-ai-hub/voice-real-time-agents/). - \*\*Music:\*\* Suno, Udio. → \[Multimodal Creative Tools\](/agentic-ai-hub/multimodal-creative-tools/) ## Small & on-device models For phones, laptops, and edge: Gemma 3n (mobile-optimized), Phi-4-mini, Ministral 3 (3B/8B/14B), Qwen3 0.6-4B, Llama 1B/3B, SmolLM3\. Quantize to GGUF/MLX and run via Ollama, llama.cpp, MLX, or LM Studio (\[Local & Self-Hosting Stack\](/agentic-ai-hub/local-self-hosting-stack/)). ## Frontier-specialized (research-adjacent) Worth knowing even if not day-to-day dev APIs: \*\*world models\*\* (Google's "Omni" / Genie-class, for simulation and embodied agents), \*\*science models\*\* (AlphaFold-class for biology, materials), and \*\*time-series/tabular\*\* foundation models. Mostly research or niche-API today, but moving toward productization. --- --- title: "Benchmarks & Leaderboards" url: https://daily.dev/agentic-ai-hub/benchmarks-leaderboards/ description: "The single most useful mental model here is that benchmarks tell you what a model can do on someone else's task." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- > The single most useful mental model here is that \*\*benchmarks tell you what a model \*can\* do on someone else's task. Only your own eval tells you what it does on \*yours\*.\*\* Read benchmarks to shortlist, then test on your workload (\[LLMOps: Evals, Observability & Guardrails\](/agentic-ai-hub/llmops-evals-observability-guardrails/)). ## The benchmarks worth knowing (and their catch) \*\*LMArena / Chatbot Arena.\*\* \*Human preference\*: users vote between two anonymous answers, and the votes become \*\*Elo\*\* ratings (the chess rating system, where beating a stronger opponent raises your score more). \*\*Good for:\*\* a real-world read on open-ended chat quality and which model people actually prefer. \*\*Catch:\*\* rewards length and formatting (use the style-controlled board), has no ground truth, and small rank gaps are noise. → \[lmarena.ai\](https://lmarena.ai/) \*\*MMLU / MMLU-Pro.\*\* \*Broad multiple-choice knowledge\* across dozens of subjects. \*\*Good for:\*\* a rough general-knowledge floor. \*\*Catch:\*\* the original is \*\*saturated\*\* (frontier models cluster in the low 90s) and contaminated, so use MMLU-Pro (harder, 10 options). → \[MMLU-Pro board\](https://artificialanalysis.ai/evaluations/mmlu-pro) \*\*GPQA Diamond.\*\* \*Expert science reasoning\*, 198 PhD-level "Google-proof" questions. \*\*Good for:\*\* telling genuine scientific reasoning apart from recall. \*\*Catch:\*\* approaching saturation, and with only 198 items a few questions swing the score. \*\*SWE-bench Verified.\*\* \*Real coding\*: fix actual GitHub issues so the project's hidden tests pass (quote the 500-item verified subset). \*\*Good for:\*\* the most credible single "can it fix real bugs" number. \*\*Catch:\*\* heavily \*\*scaffold/harness-dependent\*\* (the same model varies widely by agent), Python-centric, and older instances risk contamination (see \*\*SWE-bench Pro\*\*). → \[swebench.com\](https://www.swebench.com/verified.html) \*\*Terminal-Bench.\*\* \*Agentic CLI competence\*: multi-step tasks in a real sandboxed shell. \*\*Good for:\*\* how well an agent actually drives a terminal, close to real agent work. \*\*Catch:\*\* small task count, and the harness matters as much as the model. → \[tbench.ai\](https://www.tbench.ai/) \*\*τ²-bench (tau-bench).\*\* \*Tool use across multi-turn conversations\* against a stateful database under policy rules. \*\*Good for:\*\* agent \*reliability\*, since its \*\*pass^k\*\* metric (solved on all k runs) exposes brittleness that single runs hide. \*\*Catch:\*\* narrow domains, and results depend on the simulated user. → \[github\](https://github.com/sierra-research/tau2-bench) \*\*AIME / FrontierMath.\*\* \*Math reasoning.\* \*\*Good for:\*\* hard, checkable step-by-step reasoning. \*\*FrontierMath\*\* (Epoch AI) keeps its problems private and still has real headroom. \*\*Catch:\*\* AIME problems leak every year, so always check which exam year is quoted. → \[FrontierMath\](https://epoch.ai/frontiermath) \*\*Humanity's Last Exam (HLE).\*\* \*Frontier knowledge\*, \~2,500 expert questions built to stay hard. \*\*Good for:\*\* separating the very top models (scores are still only in the low-mid 20s%). \*\*Catch:\*\* pure knowledge and reasoning, no agentic ability or tool use. \*\*HumanEval.\*\* \*Legacy code-gen\* (164 Python functions). \*\*Good for:\*\* almost nothing now, it's saturated (96-98%) and contaminated. Use \*\*LiveCodeBench\*\* instead (problems released after training cutoffs). Also worth knowing: \*\*GDPval\*\* (professional real-work tasks), \*\*Aider Polyglot\*\* (multi-language edits), \*\*GAIA\*\* (general assistant), and \*\*SimpleBench\*\* (common-sense traps where humans beat models). ## Live leaderboards to bookmark - \[\*\*Artificial Analysis\*\*\](https://artificialanalysis.ai/leaderboards/models). Intelligence index + live speed, latency, price, context in one view. \*Best for cost/performance trade-offs.\* - \[\*\*LMArena\*\*\](https://lmarena.ai/). Human-preference Elo with category/style boards. \*Best for subjective chat quality.\* - \[\*\*llm-stats.com\*\*\](https://llm-stats.com/). 300+ models, composite score with price baked in. \*Best for fast head-to-head selection.\* - \[\*\*LM Council\*\*\](https://lmcouncil.ai/benchmarks). Aggregates \~18 \*independently-run\* benchmarks. \*Best for avoiding vendor self-reports.\* - \[\*\*Vals AI\*\*\](https://www.vals.ai/benchmarks). Industry-specific evals (legal, finance, healthcare). \*Best for domain buyers.\* - \[\*\*OpenRouter Rankings\*\*\](https://openrouter.ai/rankings). Real token-usage by model/app across a large routing marketplace. \*Best for "what devs actually ship."\* - \[\*\*Vercel AI Gateway leaderboards\*\*\](https://vercel.com/ai-gateway/leaderboards/models). Models ranked by real usage through Vercel's gateway. \*Best for the app-builder / TypeScript perspective.\* - \[\*\*Scale SEAL / Scale Labs\*\*\](https://labs.scale.com/leaderboard). Private, contamination-resistant expert evals (incl. SWE-bench Pro). \*Best for trust when public benchmarks are gamed.\* - \[\*\*Vellum LLM Leaderboard\*\*\](https://www.vellum.ai/llm-leaderboard). Clean side-by-side benchmark + price/context tables. \*Best for a quick reference sheet.\* - \[\*\*SWE-bench\*\*\](https://www.swebench.com/) / \[\*\*Terminal-Bench\*\*\](https://www.tbench.ai/leaderboard/terminal-bench/2.0). The canonical coding-agent boards. \*\*Task-specific arenas\*\* (LMArena spin-offs and others, when you care about one skill): \[\*\*WebDev Arena\*\*\](https://arena.ai/blog/webdev-arena/) (web-app building), \*\*Copilot/Code Arena\*\* (in-editor coding), \*\*Design Arena\*\* (UI/design generation), \*\*Search Arena\*\* (retrieval-grounded answers), \*\*Vision/Text-to-Image Arena\*\* (multimodal). Each isolates one capability the general boards blur together. ## Why leaderboards mislead (keep these in mind) - \*\*Contamination.\*\* Public benchmarks leak into training data, so a high score can be memorization. Prefer contamination-resistant or refreshed evals. - \*\*Saturation.\*\* Once models cluster near the ceiling, the gaps between them are noise. - \*\*Style and gaming.\*\* Preference boards reward length and formatting, and vendors can tune to a benchmark's format. - \*\*Harness dependence.\*\* Agentic scores depend as much on the scaffolding as the model, so compare same-harness numbers. \*(\[Papers with Code\](https://hyper.ai/en/news/42900) shut down in 2025, so don't rely on it for current SOTA.)\* --- --- title: "Pricing & Cost Reference" url: https://daily.dev/agentic-ai-hub/pricing-cost-reference/ description: "Pricing is a Frontier Model Comparison snapshot. This section is the durable part: how to reason about cost so you're not surprised by a bill." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- > Pricing is a \[Frontier Model Comparison\](/agentic-ai-hub/frontier-model-comparison/) snapshot. This section is the \*durable\* part: how to reason about cost so you're not surprised by a bill. \*\*As of July 2026.\*\* \*\*How LLM pricing works.\*\* You pay per token, and \*\*input and output are priced separately\*\* (output is typically 3 to 8× input, because it's generated sequentially). "Tokens" ≈ ¾ of a word in English. A back-of-envelope estimate: \`cost ≈ (input\_tokens × in\_price + output\_tokens × out\_price) / 1,000,000\`, multiplied by your request volume. \*\*The four levers that cut cost the most:\*\* 1\. \*\*Prompt caching.\*\* Reuse of a repeated prefix (system prompt, big context) at a large discount (often \~90% off input, sometimes \~99% for cache hits, e.g. DeepSeek). The single biggest win for agents and RAG that resend the same context. Structure prompts so the stable part comes first. 2\. \*\*Batch APIs.\*\* Non-urgent jobs (evals, bulk generation) at \~50% off, with a delayed SLA. Most providers offer it. 3\. \*\*Right-sizing the model.\*\* Use a mini/flash/haiku tier for routing, extraction, and classification. Reserve the flagship for genuinely hard steps. A \[router\](/agentic-ai-hub/apis-inference-platforms-gateways/) automates this. 4\. \*\*Reasoning budget.\*\* Hybrid models bill their hidden "thinking" tokens as output. Cap the reasoning effort where you don't need it. It's often the largest line item on agentic workloads. \*\*Watch-outs.\*\* Long-context tiers often cost \~2× above a threshold (e.g. >200K tokens). Newer models may use a \*\*different tokenizer\*\*. Anthropic's newest generation reportedly emits \~30% more tokens for the same text, so compare \*cost per task\*, not just the headline per-token price. Vision/audio inputs are often billed as a large fixed token count per image/second. \*\*Ballpark tiers (per 1M output tokens, July 2026):\*\* budget/open ≈ $0.30 to $3 (DeepSeek, Flash/Haiku/mini, self-host), mid ≈ $5 to $15 (Sonnet, Gemini Pro, GPT balanced), premium ≈ $25 to $90 (Opus, GPT Pro tiers). For live cost-per-task comparisons across providers, use \[Artificial Analysis\](https://artificialanalysis.ai/) and \[OpenRouter\](https://openrouter.ai/models). --- --- title: "How to Choose a Model" url: https://daily.dev/agentic-ai-hub/how-to-choose-a-model/ description: "The honest answer is to shortlist from benchmarks and reputation, then run a small eval on your actual task." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- > The honest answer is to shortlist from benchmarks and reputation, then \*\*run a small eval on your actual task\*\*. Model rankings flip depending on the workload. That said, here are good starting defaults by need (July 2026). Swap specific names for whatever tops the \[live boards\](/agentic-ai-hub/benchmarks-leaderboards/) when you read this. \*\*By task:\*\* - \*\*Coding & agents\*\* → a frontier flagship (Claude Opus/Sonnet, GPT-5.6); open option: Qwen3.6-Coder, Kimi K3, DeepSeek V4\. Weight \[SWE-bench Verified\](https://www.swebench.com/verified.html) + Terminal-Bench, and remember the \*harness\* matters as much as the model. - \*\*High-volume / cheap\*\* (classification, extraction, routing) → a mini/flash/haiku tier, or self-hosted open (DeepSeek, Qwen). Don't pay flagship prices for structured extraction. - \*\*Long context\*\* (whole codebases, long documents) → Gemini (\~2M), or 1M-context GPT/Claude. Beware: quality degrades toward the end of very long windows ("lost in the middle"), so test retrieval accuracy, don't assume it. - \*\*On-device / private\*\* → Gemma 3n, Phi-4-mini, Ministral 3, small Qwen, quantized via Ollama/llama.cpp (\[Local & Self-Hosting Stack\](/agentic-ai-hub/local-self-hosting-stack/)). - \*\*Vision / multimodal\*\* → Gemini (strong native multimodal incl. video/audio), GPT-5.6, Claude; open: Qwen-VL, Llama 4\. - \*\*Reasoning-heavy\*\* (math, hard logic, research) → flagships with "thinking"/extended-reasoning mode enabled; open: DeepSeek R-series, GLM. - \*\*Data residency / sovereignty\*\* → Mistral (EU), Cohere (on-prem/enterprise), or self-hosted open weights. \*\*Decision shortcuts:\*\* - \*Default to a flagship until cost or latency forces you down a tier\*: it's cheaper than debugging quality problems from an under-powered model. - \*Prefer an \[aggregator/gateway\](/agentic-ai-hub/apis-inference-platforms-gateways/)\* (OpenRouter, a router) early on so you can swap models with a config change, not a rewrite. - \*Decouple your code from the model\*: program against a provider-agnostic client so "choosing a model" stays a one-line decision you can revisit weekly. - \*Re-evaluate quarterly.\* At the current pace, the best model for your task probably changes a few times a year. --- --- title: "Chat Assistants & Apps" url: https://daily.dev/agentic-ai-hub/chat-assistants-apps/ description: "The consumer/prosumer chat apps are the front door to each lab's frontier model." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- The consumer/prosumer chat apps are the front door to each lab's frontier model. They differ less in raw IQ now than in \*\*surface features\*\*: memory, connectors, agent modes, voice, and how aggressively they route you to a smaller model when you hit limits. Pick based on ecosystem lock-in (docs, cloud, IDE) more than benchmark deltas. ## The assistants at a glance 🔄 (as of July 2026) | Assistant | Maker | Best for | Watch out | Pricing | | :---- | :---- | :---- | :---- | :---- | | \[ChatGPT\](https://chatgpt.com) | OpenAI | Widest feature surface: agent mode, code interpreter, deep research, image gen, connectors, voice. \*\*ChatGPT Work\*\* (Jul 2026) folds Codex in to rival Claude for agentic and coding work. | Free/Plus users are routed to smaller models and capped on the strongest reasoning | Free · Plus \~$20 · Pro \~$200 | | \[Claude\](https://claude.ai) | Anthropic | Coding, long-document and codebase reasoning, careful writing, Artifacts (live-rendered apps/docs) | Smaller free-tier limits; no native image generation | Free · Pro \~$20 · Max \~$100-200 | | \[Gemini\](https://gemini.google.com) | Google | Very large context, native multimodality (image + video), Workspace integration (Gmail/Docs/Sheets) | The standalone app and "Gemini in Workspace" differ in features and data handling | Free · AI Pro \~$20 · AI Ultra \~$250 | | \[Grok\](https://grok.com) | xAI | Real-time X data, looser content policy, "Heavy" multi-agent mode | Fewer guardrails cut both ways; thinner ecosystem than Google/Microsoft | Free on X · SuperGrok \~$30 · Heavy \~$300 | | \[Copilot\](https://copilot.microsoft.com) | Microsoft | Living inside Office (Word/Excel/Outlook/Teams) and enterprise governance | Consumer Copilot and paid Microsoft 365 Copilot are different products | Free · Pro \~$20 · M365 \~$30/user | | \[Perplexity\](https://www.perplexity.ai) | Perplexity | Cited, up-to-date web answers with inline sources; Deep Research | Tuned for retrieval and synthesis, not open-ended creativity or long coding | Free · Pro \~$20 · Max \~$200 | \*Others worth knowing:\* \[Meta AI\](https://www.meta.ai) (free, in WhatsApp/Instagram/Ray-Ban glasses), \[DeepSeek\](https://chat.deepseek.com) (free app, cheap strong reasoning), \[Mistral Le Chat\](https://chat.mistral.ai) (European, privacy-forward), \[Poe\](https://poe.com) (many models on one subscription). Prices move, so verify on each site. --- --- title: "Coding Tools & CLI Agents" url: https://daily.dev/agentic-ai-hub/coding-tools-cli-agents/ description: "Two families here: IDE assistants (autocomplete + chat + in-editor agents living in your editor) and CLI/terminal agents (you talk to an agent in the shell." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- Two families here: \*\*IDE assistants\*\* (autocomplete + chat + in-editor agents living in your editor) and \*\*CLI/terminal agents\*\* (you talk to an agent in the shell. It reads, edits, runs, and tests across your repo). The 2025→2026 shift was from "smarter autocomplete" to "agentic". The tool plans, edits many files, runs commands, and iterates. Pricing has largely moved to \*\*usage/credits on top of a seat\*\*, so watch token burn. ## IDE assistants | Tool | Best for | Watch out | Price | | :---- | :---- | :---- | :---- | | \[Cursor\](https://cursor.com) | Whole-repo agentic edits (Composer), fast multi-file changes, predictive tab-completion | Agent runs burn credits; the pricing model has changed more than once | Free · Pro \~$20 · Business/usage 🔄 | | \[GitHub Copilot\](https://github.com/features/copilot) | Ubiquity (VS Code, JetBrains, Neovim, Visual Studio); agent mode; choose Anthropic/OpenAI/Google models | Historically more conservative than Cursor; premium requests metered | Free · Pro \~$10 · Pro+ \~$39 · Business \~$19/user | | \[Windsurf\](https://windsurf.com) | Its Cascade agent flow and clean, agent-driven UX | 2025 ownership saga (now Cognition); watch roadmap/pricing continuity | Free · paid \~$15+/mo, credit-based 🔄 | | \[Zed\](https://zed.dev) | Speed (Rust), multiplayer editing, native agent panel with external agents via ACP | Leaner extension ecosystem; newer AI layer | Open source; AI free with BYO key + hosted plan 🔄 | \*Also:\* \[JetBrains AI Assistant / Junie\](https://www.jetbrains.com/ai/), \[Kiro\](https://kiro.dev) (AWS spec-driven IDE), \[Trae\](https://www.trae.ai) (ByteDance). ## CLI / terminal agents These live in your shell, see the whole repo, and can run commands. Increasingly the \*primary\* interface for serious agentic coding. | Tool | Best for | Watch out | Price / open | | :---- | :---- | :---- | :---- | | \[Claude Code\](https://claude.com/claude-code) | Deep multi-step work across big codebases (plan, edit, test, git); subagents, hooks, skills, MCP. A de-facto reference harness | Will run commands and burn tokens; use permission modes and review | Claude Pro/Max or API; CLI free to install | | \[OpenAI Codex CLI\](https://github.com/openai/codex) | Local agentic coding on OpenAI models, with local-to-cloud handoff | "Codex" spans a CLI, IDE extension, and cloud agent, so know which you're using | Open source (Apache-2.0); ChatGPT plan or API | | \[Aider\](https://aider.chat) | Tight, git-native editing loops (auto-commit each change); model-agnostic, scriptable | A precise pair-programmer more than an autonomous agent; bare-bones UI | Free (MIT); pay only model API | | \[Gemini CLI\](https://github.com/google-gemini/gemini-cli) | Generous free Gemini access, large context, web/Google tools; MCP + ACP | Newer than Claude Code/Aider; tracks whichever Gemini you're allotted | Free tier via Google account; open source | | \[Amp\](https://ampcode.com) | High-autonomy "let the agent cook" runs, subagents/threads, team sharing (Sourcegraph) | Token-metered; can get pricey on big autonomous runs | Usage-based credits, free to start 🔄 | \*Also:\* \[OpenCode\](https://opencode.ai), \[Cline\](https://cline.bot), \[Charm Crush\](https://github.com/charmbracelet/crush) (all open source). ## PR / review & background bots | Tool | Best for | Link | | :---- | :---- | :---- | | GitHub Copilot code review | Assign an issue and it opens a PR, or reviews yours inline, inside GitHub's flow | \[copilot\](https://github.com/features/copilot) | | Cursor Bugbot | Automated PR bug and security review with fast turnaround; can auto-review agent-written PRs | \[cursor.com/bugbot\](https://cursor.com/bugbot) | | CodeRabbit | Line-by-line PR feedback and summaries; free for open source | \[coderabbit.ai\](https://coderabbit.ai) | | Greptile | PR review that reasons over your whole repo graph | \[greptile.com\](https://greptile.com) | | Devin (Cognition) | An autonomous engineer you assign well-scoped tickets; runs in its own cloud env (supervise) | \[devin.ai\](https://devin.ai) | | Graphite Diamond / Qodo | Review + test generation | \[graphite\](https://graphite.dev) · \[qodo\](https://www.qodo.ai) | \*\*Rule of thumb:\*\* IDE assistant for \*flow\* (you're editing, it helps), CLI agent for \*delegation\* (you describe, it executes across the repo), PR bot for \*gatekeeping\* (it reviews what humans and agents produced). --- --- title: "Agent Harnesses & Frameworks" url: https://daily.dev/agentic-ai-hub/agent-harnesses-frameworks/ description: "Two very different things get called \\"agent tooling.\\"" lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- Two very different things get called "agent tooling." Keep them separate: \*\*(a) Harnesses\*\* are a finished agent \*loop\* you drive (a REPL/CLI that manages context, tools, permissions, and the model call). You configure and prompt it. You don't write the loop. \*\*(b) Build frameworks\*\* are libraries you use to \*write your own\* agent/workflow in code (orchestration, state, tool-calling, memory). ## (a) Harnesses you drive an agent through | Harness | Best at | Watch out | Key fact | | :---- | :---- | :---- | :---- | | \[Claude Code\](https://claude.com/claude-code) | Long multi-step work across large codebases (plan → edit → run tests → commit); subagents, hooks, skills, MCP, permission modes; scriptable via the Claude Agent SDK | Runs commands and burns tokens; use permission modes and review | Anthropic's terminal-first harness, a de-facto reference implementation; included with Pro/Max or via API | | \[OpenAI Codex\](https://github.com/openai/codex) | Delegating a task to run locally or hand off to the cloud, on OpenAI models; also inside ChatGPT Work | "Codex" spans a CLI, IDE extension, and cloud agent, so know which you're using | Open-source CLI; use with a ChatGPT plan or API | | \[Pi\](https://pi.dev) | Being small enough to read end-to-end and bend to your own workflow; model-agnostic | A toolkit for building your own agent, not a batteries-included consumer app | Minimal, hackable terminal harness by Mario Zechner; open source, npm \`@mariozechner/pi-coding-agent\` (\[writeup\](https://mariozechner.at/posts/2025-11-30-pi-coding-agent/)) | | \[OpenClaw\](https://openclaw.ai) | A self-hosted, pluggable, cross-platform assistant that can wrap external coding harnesses via ACP | Community project, so capabilities and stability vary; verify current state before depending on it | Open source (aka Clawdbot); \[GitHub\](https://github.com/openclaw/openclaw) | | \[Amp\](https://ampcode.com) | Team-shared, aggressive autonomous runs (also in \[Coding Tools & CLI Agents\](/agentic-ai-hub/coding-tools-cli-agents/)) | Token-metered; can get pricey on big runs | Sourcegraph's high-autonomy harness; usage-based | \*Emerging interop standard:\* \*\*ACP (Agent Client Protocol)\*\*, an open protocol from the Zed team (think "LSP for coding agents"), lets editors talk to any compliant agent harness, so Claude Code, Gemini CLI, and others can appear inside different front-ends. It's editor-to-agent, complementing MCP's agent-to-tool role (\[MCP & the Tool/Context Ecosystem\](/agentic-ai-hub/mcp-tool-context-ecosystem/)). ## (b) Build frameworks (roll your own agent) | Framework | Language | Model | Best at | Note | | :---- | :---- | :---- | :---- | :---- | | LangGraph | Py/JS | Any | Stateful, graph-based agent workflows | From LangChain; production-oriented | | LangChain | Py/JS | Any | Broad integrations, RAG glue | Large surface; some see it as heavy | | LlamaIndex | Py/TS | Any | RAG / data-connected agents | Best-in-class retrieval | | CrewAI | Python | Any | Role-based multi-agent "crews" | Simple mental model | | AutoGen / AG2 | Python | Any | Multi-agent conversations | Microsoft origin; forked into AG2 | | OpenAI Agents SDK | Py/JS | OpenAI-first | Lightweight handoffs + guardrails | Successor to Swarm | | Claude Agent SDK | Py/TS | Claude | Production agents on the Claude Code engine | Same harness that powers Claude Code | | Pydantic AI | Python | Any | Type-safe, validated agent I/O | Great DX for Python teams | Links: \[LangChain/LangGraph\](https://www.langchain.com) · \[LlamaIndex\](https://www.llamaindex.ai) · \[CrewAI\](https://www.crewai.com) · \[AutoGen\](https://microsoft.github.io/autogen/) / \[AG2\](https://ag2.ai) · \[OpenAI Agents SDK\](https://openai.github.io/openai-agents-python/) · \[Claude Agent SDK\](https://docs.claude.com/en/api/agent-sdk/overview) · \[Pydantic AI\](https://ai.pydantic.dev). \*\*Framework vs. roll-your-own.\*\* For a single agent with a handful of tools, a plain provider SDK + a loop (or a thin harness) is often \*simpler and more debuggable\* than a framework. The model does the reasoning. You mostly need a good tool layer and context management. Reach for a framework when you need \*\*durable state, multi-agent coordination, human-in-the-loop checkpoints, or heavy retrieval\*\*, i.e., orchestration you'd otherwise reinvent. When in doubt, start minimal (many teams found frameworks added indirection they later removed). --- --- title: "No-Code & App Builders" url: https://daily.dev/agentic-ai-hub/no-code-app-builders/ description: "Prompt-to-app tools turn a description into a running web app with UI, backend, and deploy." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- Prompt-to-app tools turn a description into a running web app with UI, backend, and deploy. Good for prototypes, internal tools, landing pages, and MVPs. The ceiling shows up on complex state, custom infra, and long-term maintainability. Most now let you export or push to GitHub, so you're less locked in than the "no-code" label suggests. | Tool | Best at | Watch out | Price | | :---- | :---- | :---- | :---- | | \[v0\](https://v0.app) (Vercel) | Production-grade \*frontend\* components and pages that drop into a Vercel/Next codebase (React/Next + Tailwind + shadcn/ui); strong design output | Opinionated toward the Vercel/React stack; full-stack is less its strength than UI | Free tier + paid credits (\~$20/mo and up) 🔄 | | \[Bolt\](https://bolt.new) (StackBlitz) | Prompt to a \*running\* full-stack app entirely in the browser (npm, dev server, one-click deploy); framework-flexible | Token-metered, so iterating on errors eats credits fast; complex apps strain it | Free tier + paid token plans 🔄 | | \[Lovable\](https://lovable.dev) | Clean full-stack apps with \*\*Supabase\*\* (auth/DB) wired in, polished UI; strong for founders/MVPs | The more custom the logic, the more you fight the abstraction | Free tier + paid credits (\~$25/mo and up) 🔄 | | \[Replit Agent\](https://replit.com) | End-to-end build-\*and-run\*: scaffolds, codes, installs deps, provisions a DB, and deploys in one hosted environment; good for beginners | Usage/"checkpoint" billing can surprise you; you're in Replit's ecosystem | Free tier + Core \~$20/mo + usage 🔄 | \*Also:\* \*\*Gemini Canvas / Google Opal / AI Studio "Build"\*\* (Google's prompt-to-app surfaces), \*\*Claude Artifacts\*\* (shareable mini-apps inside Claude), \*\*Base44\*\* (Wix-owned app builder), \*\*Softr / Glide\*\* (no-code on top of your data), \[\*\*Bubble\*\*\](https://bubble.io) (mature visual full-stack, now AI-assisted). \*\*Limits to plan for:\*\* generated code can be verbose or insecure. State management and auth are where prototypes break. Always review before shipping anything handling real user data, and export to a real repo once past the prototype stage. --- --- title: "Voice & Real-Time Agents" url: https://daily.dev/agentic-ai-hub/voice-real-time-agents/ description: "Real-time voice agents = STT (speech-to-text) → LLM → TTS (text-to-speech) stitched into a low-latency loop (or, increasingly, a single speech-to-speech model)." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- Real-time voice agents = \*\*STT (speech-to-text) → LLM → TTS (text-to-speech)\*\* stitched into a low-latency loop (or, increasingly, a single speech-to-speech model). The "platform" layer (Vapi/Retell/Bland) handles telephony, turn-taking, and orchestration; the "infra" layer (LiveKit/Pipecat) gives you the transport/pipeline to build your own; and the "model" layer (ElevenLabs/Deepgram/Cartesia + realtime APIs) supplies the ears and voice. All 🔄. Pricing here is per-minute and shifts often. ## Platforms (build a phone/voice agent fast) | Platform | Best at | Watch out | Price | | :---- | :---- | :---- | :---- | | \[Vapi\](https://vapi.ai) | Fast time-to-production with pluggable STT/LLM/TTS, function calling, and telephony; big ecosystem | Per-minute cost stacks (platform + each model + telephony), so model your unit economics | Usage-based, \~$0.05+/min plus provider costs 🔄 | | \[Retell AI\](https://www.retellai.com) | Natural turn-taking and call reliability for support/sales | None | Usage-based per minute 🔄 | | \[Bland AI\](https://www.bland.ai) | High-volume outbound/inbound calling with an all-in-one stack | More closed than Vapi or Retell | Usage-based 🔄 | \*Also:\* \[\*\*Synthflow\*\*\](https://synthflow.ai), \[\*\*Vocode\*\*\](https://www.vocode.dev) (open source). ## Infrastructure (build your own pipeline) | Tool | Best at | Key fact | | :---- | :---- | :---- | | \[LiveKit\](https://livekit.io) | Scalable real-time audio/video transport and building custom voice agents (self-host or cloud); the transport backbone many voice apps and some labs' voice features run on | Open-source WebRTC platform + an \*\*Agents\*\* framework; LiveKit Cloud is usage-priced | | \[Pipecat\](https://www.pipecat.ai) | Composing the STT→LLM→TTS pipeline with full control and vendor-swappable components | Open-source Python framework (BSD), originated at Daily | | \[Ultravox\](https://www.ultravox.ai) | Lower-latency, more natural voice by feeding audio straight into an LLM-derived model | Open-weight \*\*speech-to-speech\*\* that skips separate STT; open weights + hosted API, per-minute 🔄 | ## Voices & ears (TTS / STT / realtime models) | Tool | Best at | Watch out | Price | | :---- | :---- | :---- | :---- | | \[ElevenLabs\](https://elevenlabs.io) | Natural, expressive \*\*TTS\*\*, voice cloning, dubbing, and a full \*\*Agents\*\* platform; broad language coverage | Premium pricing; voice-cloning consent/ethics matter | Free tier + paid from \~$5/mo (character-metered) 🔄 | | \[Deepgram\](https://deepgram.com) | Fast, accurate, cheap streaming \*\*STT\*\* (and now TTS) at scale (Nova-series) | None | Usage-based per minute/hour; free credits to start 🔄 | | \[Cartesia\](https://cartesia.ai) | The fastest realtime \*\*TTS\*\* (Sonic, state-space models) with tiny time-to-first-audio | None | Usage-based, free tier 🔄 | \*Realtime APIs:\* \[\*\*OpenAI Realtime API\*\*\](https://platform.openai.com/docs/guides/realtime) (speech-to-speech via GPT-4o-realtime/successors), \[\*\*Google Gemini Live API\*\*\](https://ai.google.dev/gemini-api/docs/live), \[\*\*AssemblyAI\*\*\](https://www.assemblyai.com) (STT with speech understanding). \*\*Latency budget rule:\*\* conversational agents feel natural under \~800ms round-trip. Speech-to-speech models (Ultravox, OpenAI Realtime) cut the STT/TTS hops. Classic pipelines win on component choice and cost control. --- --- title: "Browser & Computer-Use Agents" url: https://daily.dev/agentic-ai-hub/browser-computer-use-agents/ description: "These agents perceive a screen or DOM and act (click, type, scroll, navigate) to complete tasks a human would do in a browser or OS." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- These agents perceive a screen or DOM and act (click, type, scroll, navigate) to complete tasks a human would do in a browser or OS. Useful for automation over apps that lack APIs; still the \*\*least reliable and highest-risk\*\* agent category. Treat them as an intern who can be \*\*prompt-injected by any web page\*\*: sandbox, scope, and supervise. ## Models / agents | Agent | Best at | Watch out | Key fact | | :---- | :---- | :---- | :---- | | \[Claude for Chrome / computer use\](https://www.claude.com) | Operating your browser (Chrome extension) or a virtual desktop via the API's computer-use tool, with permission prompts on sensitive actions | Anthropic explicitly warns about \*\*prompt-injection\*\*; access gated/rolled out to subscribers | Available to Max subscribers (Chrome) / via API (computer use) 🔄 | | \[OpenAI computer-using agents\](https://openai.com) | Multi-step web tasks in a hosted virtual browser, via ChatGPT agent and the computer-use tool / Responses API | Standalone "Operator" was consolidated into ChatGPT agent, so check current naming; supervise checkouts/logins | Pro/Plus feature + API 🔄 | | \[Google Project Mariner / Gemini\](https://deepmind.google) | Agentic browser tasks surfacing in Gemini and Chrome | Research-turned-product; capability still maturing | 🔄 | \*Also:\* \[\*\*Perplexity Comet\*\*\](https://www.perplexity.ai/comet) (agentic browser), \*\*The Browser Company's Dia\*\*, and other AI browsers. ## Browser-automation infrastructure (for developers) | Tool | Best at | Key fact | | :---- | :---- | :---- | | \[Browserbase\](https://www.browserbase.com) | Running many concurrent, stealthy browser sessions with auth/proxy/captcha handling; the infra layer under browser agents | Managed headless-browser cloud; usage-based, free tier to start 🔄 | | \[Stagehand\](https://www.stagehand.dev) | Mixing deterministic Playwright with natural-language \`act\`/\`extract\`/\`observe\` steps, for reliability of code plus flexibility of an agent | Browserbase's open-source (MIT) SDK on Playwright; pairs with Browserbase | \*Also:\* \[\*\*Playwright\*\*\](https://playwright.dev) (the deterministic automation baseline, now with AI/MCP add-ons), \[\*\*Browser Use\*\*\](https://browser-use.com) (popular open-source browser-agent library), \[\*\*Skyvern\*\*\](https://www.skyvern.com) (vision-based automation), \*\*Notte\*\*, \*\*Steel.dev\*\* (agent browser infra). \*\*Reliability & safety caveats (do not skip):\*\* - \*\*Prompt injection is the core threat.\*\* A malicious page can hijack an agent that reads it. Never give a browser agent standing access to email, banking, or admin without human confirmation on each sensitive action. - \*\*Reliability is task-dependent.\*\* Success rates on clean, well-known flows are high. Long-tail UIs, popups, and captchas still break agents. Build retries and human fallbacks. - \*\*Prefer APIs when they exist.\*\* Computer-use is a fallback for apps with no API, not a first choice. Prefer deterministic (Playwright/Stagehand) steps for known flows and let the model handle only the fuzzy parts. - \*\*Isolate.\*\* Run in a sandboxed/ephemeral browser (Browserbase-style), scope credentials, and log everything. --- --- title: "Multimodal Creative Tools" url: https://daily.dev/agentic-ai-hub/multimodal-creative-tools/ description: "Image, video, and audio generation. Quality is now high across the board. Choose by control, licensing, and API availability more than raw \\"wow.\\"" lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- Image, video, and audio generation. Quality is now high across the board. Choose by \*\*control, licensing, and API availability\*\* more than raw "wow." \*\*Commercial-use rights vary\*\*, so check the license before shipping generated media. ## Image | Tool | Best at | Watch out | Key fact | | :---- | :---- | :---- | :---- | | \[Midjourney\](https://www.midjourney.com) | The most striking, art-directed stills; strong style control and consistency (web app + Discord) | No official public API (third-party wrappers are unofficial); subscription-gated | Aesthetic leader; paid from \~$10/mo, latest V8-series 🔄 | | \[FLUX\](https://bfl.ai) (Black Forest Labs) | Developer control: run it yourself or via API, strong prompt adherence and editing (FLUX Kontext); the base for many products | Variants carry different licenses (some non-commercial), so read them | Leading \*\*open-weight\*\* image family; open weights + hosted API (per-image) 🔄 | | \[Ideogram\](https://ideogram.ai) | Posters, logos, and typography where legible in-image \*\*text\*\* matters | None | Best-in-class text rendering; free tier + paid, API available 🔄 | | \[Adobe Firefly\](https://firefly.adobe.com) | Enterprise use with \*\*indemnification\*\* and native Photoshop/Creative Cloud integration; now aggregates partner models | None | Trained on licensed/Adobe Stock data; free credits + Firefly plans, API for business 🔄 | | Native (\[GPT image\](https://platform.openai.com) · \[Gemini\](https://ai.google.dev)) | The fastest path if you're already calling those APIs (gpt-image's strong instruction-following, Gemini "Nano Banana") | None | Generation built into the frontier chat APIs 🔄 | ## Video | Tool | Maker | Best at | API? | | :---- | :---- | :---- | :---- | | Sora | OpenAI | Cinematic realism, physics, audio | Yes 🔄 | | Veo | Google DeepMind | Realism + native audio, long shots | Yes (Vertex/Gemini) | | Runway | Runway | Filmmaker control, editing tools | Yes | | Kling | Kuaishou | Strong motion/realism, value | Yes 🔄 | | Pika | Pika | Fun effects, fast social clips | Limited | | Luma Dream Machine | Luma | Fast, smooth, cheap iteration | Yes | Links & latest generations 🔄: \[Sora 2\](https://sora.com) (OpenAI) · \[Veo 3-series\](https://deepmind.google/models/veo/) (Google, via Vertex AI) · \[Runway Gen-series\](https://runwayml.com) · \[Kling 2.x/3.x\](https://klingai.com) · \[Pika\](https://pika.art) · \[Luma Ray-series\](https://lumalabs.ai/dream-machine). ## Audio & music | Tool | Best at | Watch out | Key fact | | :---- | :---- | :---- | :---- | | \[Suno\](https://suno.com) | Full, catchy text-to-\*\*song\*\* tracks (vocals + instrumentation) from a prompt | Commercial rights depend on tier; the licensing/legal landscape is contested | Category leader; latest v4/v5-series; free tier + paid \~$10/mo and up 🔄 | | \[Udio\](https://www.udio.com) | High audio fidelity and fine-grained editing/extend controls | None | Suno's main rival; free tier + paid 🔄 | | \[ElevenLabs\](https://elevenlabs.io) | Voice/audio production for developers with a real API: \*\*music\*\*, sound effects, dubbing (see \[Voice & Real-Time Agents\](/agentic-ai-hub/voice-real-time-agents/)) | None | Beyond TTS; the audio house for builders | \*\*Dev note:\*\* if you need generation \*in your product\*, prefer tools with real APIs (Sora, Veo via Vertex, Runway, Kling, Luma, FLUX, gpt-image, Ideogram, ElevenLabs). Midjourney and most music tools remain app-first. Always confirm \*\*commercial licensing\*\* per tool and per tier. --- --- title: "MCP & the Tool/Context Ecosystem" url: https://daily.dev/agentic-ai-hub/mcp-tool-context-ecosystem/ description: "Models are only half the story. The other half is the plumbing that gives models hands: how an assistant reaches your files, databases, SaaS tools, and the web." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- Models are only half the story. The other half is the \*\*plumbing that gives models hands\*\*: how an assistant reaches your files, databases, SaaS tools, and the web. That layer is becoming standardized fast. ## What MCP is The \*\*Model Context Protocol (MCP)\*\* is an open standard, introduced by Anthropic in late 2024 and since adopted across the industry, for connecting AI applications to external tools and data. The tagline that stuck: \*\*"a USB-C port for AI."\*\* Instead of every app hand-coding a bespoke integration for every tool, an app speaks MCP once and can talk to any MCP \*\*server\*\*. → \[modelcontextprotocol.io\](https://modelcontextprotocol.io) The architecture is a simple client/server split: - \*\*MCP host / client\*\* is the AI app (Claude Code, ChatGPT, Cursor, an IDE, your own agent) that wants capabilities. - \*\*MCP server\*\* is a small program exposing three kinds of things to the client: - \*\*Tools\*\* are actions the model can call (query a DB, send a Slack message, open a PR). - \*\*Resources\*\* are data the model can read (files, records, docs). - \*\*Prompts\*\* are reusable templated instructions/workflows the server offers. - \*\*Transport\*\* is local (stdio) or remote (HTTP/SSE, with an auth story via OAuth for hosted servers). ## Why it matters - \*\*Write once, connect everywhere, vendor-neutrally.\*\* Before MCP it was an N×M problem: every AI app times every tool. Now you build one server for your tool and every MCP-speaking client can use it, whether that's Claude, a Cursor agent, or your own app on the OpenAI Agents SDK. - \*\*Auth is part of the standard, but authorization is on you.\*\* Remote MCP servers authenticate via OAuth. Exposing a tool still means deciding \*what the agent may do with it\*, so scope credentials tightly and treat tool access as a security boundary (see \[Safety, Alignment & Governance\](/agentic-ai-hub/safety-alignment-governance/) and \[Browser & Computer-Use Agents\](/agentic-ai-hub/browser-computer-use-agents/)). - \*\*Broad adoption.\*\* What started at Anthropic is now supported across major agent runtimes and IDEs, with OpenAI, Google, and Microsoft tooling interoperating with it. That cross-vendor buy-in is what turned MCP from a nice idea into the default integration layer. ## The ecosystem around it - \*\*Server registries & directories.\*\* Official and community registries catalog thousands of MCP servers (databases, SaaS apps, dev tools, browsers, search). Reference servers exist for the common cases (filesystem, git, fetch, Postgres, etc.). - \*\*First-party servers.\*\* Many vendors now ship their \*own\* MCP server (GitHub, Linear, Stripe, Sentry, Notion, Cloudflare, and more), so you get maintained, authenticated access rather than a scraped hack. - \*\*Hosted/remote servers + connectors.\*\* The "connectors" you toggle inside Claude, ChatGPT, and others are frequently MCP servers under the hood, with OAuth so a user grants scoped access to their accounts. - \*\*Skills & plugins.\*\* Layered on top: packaged instructions/workflows (e.g., Claude \*\*Skills\*\*, plugins) that \*use\* MCP tools but add domain know-how and multi-step procedure. MCP gives the \*hands\*. Skills give the \*playbook\*. ## How this composes with function calling Under the hood, MCP tools surface to the model through the \*\*same function-/tool-calling mechanism\*\* every frontier API already exposes: the model emits a structured call, the host executes it (here, by routing to an MCP server), and the result returns to the model. So MCP doesn't replace function calling. It \*\*standardizes where the functions come from and how they're described\*\*, so the same tool works across providers instead of being re-declared per app. Provider tool-calling formats still differ slightly (OpenAI, Anthropic, Google each have their own JSON shape), but MCP sits above that, and adapters/SDKs bridge the gaps. ## Practical guidance for builders - \*\*Consuming tools?\*\* Prefer an existing, maintained MCP server over hand-rolling an API client, so you inherit auth, schemas, and updates. - \*\*Exposing your product to AI?\*\* Ship an MCP server. It's becoming the expected way to be "AI-accessible," the way a public API became table stakes a decade ago. - \*\*Security is on you.\*\* MCP servers can read data and take actions, so the same cautions as agents apply: scope credentials, review tool permissions, be wary of \*\*prompt injection\*\* via tool \*outputs\* (a malicious record can carry instructions), and don't auto-approve destructive tools. Only install servers you trust. - \*\*Keep tool surfaces lean.\*\* Exposing 200 tools to a model degrades tool selection. Curate the set per task (many hosts let you enable/disable servers or tools per session). \*\*Bottom line:\*\* MCP is the connective tissue of the agent era, the standard that lets any model, in any app, use any tool. Build agents and you'll produce and consume MCP servers constantly. --- --- title: "RAG & Retrieval" url: https://daily.dev/agentic-ai-hub/rag-retrieval/ description: "Retrieval-Augmented Generation (RAG) is the practice of fetching relevant text (or other data) at query time and injecting it into the model's context so the model can answer from information it wasn't trained on." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- \*\*Retrieval-Augmented Generation (RAG)\*\* is the practice of fetching relevant text (or other data) at query time and injecting it into the model's context so the model can answer from information it wasn't trained on. It exists because model weights are frozen at training time, context windows are finite and expensive, and hallucination drops sharply when the model is handed the actual source material. The canonical pipeline is: \*\*ingest → chunk → embed → index → retrieve → (rerank) → assemble prompt → generate\*\*, usually with a citation step so answers are auditable. ## Chunking Chunking splits documents into retrievable units. The tension is always the same. Chunks too large dilute the embedding signal and waste context budget. Chunks too small lose the surrounding meaning a passage needs to be interpretable. Common strategies: - \*\*Fixed-size with overlap.\*\* Split every N tokens (commonly 200 to 500) with a 10 to 20% overlap so ideas straddling a boundary aren't orphaned. A simple and robust default. - \*\*Recursive / structural.\*\* Split on natural boundaries (headings, paragraphs, sentences) before falling back to size limits. It respects document structure. - \*\*Semantic chunking.\*\* Place boundaries where the embedding similarity between adjacent sentences drops, so each chunk is topically coherent. It costs more compute at ingest and often improves recall. - \*\*Late chunking / contextual retrieval.\*\* Embed with document-level context in view, or prepend a short LLM-generated summary of the parent document to each chunk (Anthropic popularized "contextual retrieval," which prepends chunk-situating context before embedding). This reduces the "chunk makes no sense out of context" failure. There is no universal best chunk size. It is an empirical knob you tune against your own eval set and your embedding model's optimal input length. ## Embeddings An \*\*embedding\*\* maps text to a dense vector such that semantically similar text lands nearby in vector space, measured by cosine similarity or dot product. Retrieval quality is bounded by embedding quality, so model choice matters. Practical considerations: dimensionality (higher can capture more nuance but costs storage and compute. Many 2026 models support \*\*Matryoshka\*\* truncation so you can trade dimensions for accuracy at query time), maximum input length (must exceed your chunk size), and domain fit (code, legal, and multilingual corpora often need specialized models). The \*\*MTEB\*\* leaderboard on Hugging Face is the standard starting point for comparing embedding models, though you should always re-benchmark on your own data. Leading options as of July 2026 include OpenAI's \`text-embedding-3\` family, Cohere Embed, Voyage AI (now owned by MongoDB), Google's Gemini embeddings, and strong open models (Qwen3-Embedding, which sits at/near the top of MTEB, plus BGE, E5, Nomic, Jina). \*\*The same embedding model must be used for both indexing and querying.\*\* ## BM25 vs vector vs hybrid - \*\*BM25 (lexical / sparse).\*\* A decades-old bag-of-words ranking function that scores documents on exact term overlap weighted by term frequency and inverse document frequency. It works well for rare tokens, product codes, names, acronyms, and exact-match queries, but it is blind to synonyms and paraphrase. - \*\*Vector (dense) search.\*\* Nearest-neighbor over embeddings. It captures meaning and paraphrase and can retrieve a passage that shares no words with the query, but it is weaker on exact identifiers and rare out-of-distribution terms. - \*\*Hybrid.\*\* Run both and fuse the result lists, typically with \*\*Reciprocal Rank Fusion (RRF)\*\*, which combines rankings without needing to normalize incompatible score scales. Hybrid is the common default for production RAG because lexical and semantic retrieval fail in different, complementary ways. As of 2026 this is the mainstream recommendation across most production RAG guides. ## Rerankers First-stage retrieval optimizes for recall: get the right passage somewhere in the top 50 to 100 cheaply. A \*\*reranker\*\* (usually a cross-encoder that reads the query and each candidate \*together\*, rather than comparing pre-computed vectors) then reorders those candidates for precision, so the top 3 to 8 you actually put in the prompt are the best ones. This two-stage "retrieve-then-rerank" pattern is one of the higher-leverage quality upgrades in RAG. It adds latency and per-call cost, so rerank a shortlist rather than the whole corpus. Common choices as of July 2026: \*\*Cohere Rerank\*\*, \*\*Voyage rerank\*\*, \*\*Jina reranker\*\*, and open cross-encoders like \*\*BGE reranker\*\*. \*\*ColBERT\*\*-style late-interaction models sit in between (token-level matching, more precise than single-vector, cheaper than full cross-encoding). ## Retrieval evaluation RAG has two failure surfaces, retrieval and generation, that must be evaluated separately. - \*\*Retrieval metrics.\*\* precision@k, recall@k, MRR (mean reciprocal rank), and nDCG measure whether the right chunks were fetched and ranked highly. These need labeled query→relevant-doc pairs (a "golden set"). - \*\*Generation/answer metrics.\*\* faithfulness (is the answer grounded in the retrieved context, i.e. not hallucinated?), answer relevancy, and context precision/recall. \[Ragas\](https://docs.ragas.io/) is a widely used open-source framework for these RAG-specific metrics. \[DeepEval\](https://www.deepeval.com/), \[TruLens\](https://www.trulens.org/), and \[Arize Phoenix\](https://phoenix.arize.com/) are common alternatives, and general LLM-as-judge harnesses (see \[LLMOps: Evals, Observability & Guardrails\](/agentic-ai-hub/llmops-evals-observability-guardrails/)) also apply. Build a golden eval set of real questions early. Even 50 to 100 hand-labeled examples catch most regressions and turn "it feels better" into a number. ## Common failure modes - \*\*Bad chunking.\*\* The answer is split across two chunks, and neither retrieved chunk is sufficient. - \*\*Retriever/generator mismatch.\*\* The right chunk is retrieved but buried at rank 40\. Without reranking it never reaches the prompt. - \*\*Lost in the middle.\*\* Models attend less to content in the middle of long contexts, so stuffing 50 chunks can perform worse than 5 well-ranked ones. Fewer, better chunks usually win. - \*\*Stale or duplicated index.\*\* Near-duplicate chunks crowd out diverse evidence, and without a re-index pipeline answers drift from the source of truth. - \*\*Query/document asymmetry.\*\* Short keyword queries embed poorly against long passages. Techniques like \*\*HyDE\*\* (embed a hypothetical answer) or query rewriting/expansion help. - \*\*No citations / no grounding check.\*\* You can't tell hallucination from fact. ## Frameworks & pointers | Tool | What it is | | :---- | :---- | | \[LlamaIndex\](https://www.llamaindex.ai/) | Data framework centered on RAG/indexing, with connectors and retrieval abstractions | | \[LangChain / LangGraph\](https://www.langchain.com/) | General orchestration with many retrieval integrations; LangGraph handles stateful/agentic retrieval | | \[Haystack\](https://haystack.deepset.ai/) (deepset) | Production-oriented, pipeline-first framework | | \[Ragas\](https://docs.ragas.io/) | RAG evaluation | | \[DSPy\](https://dspy.ai/) | Programmatic prompt/pipeline optimization, increasingly used to tune RAG pipelines | > Rule of thumb: reach for RAG when the knowledge is large, changes often, or must be cited. Reach for fine-tuning (\[Fine-Tuning & Post-Training\](/agentic-ai-hub/fine-tuning-post-training/)) when you need to change \*behavior, format, or style\*, not to inject facts. --- --- title: "Vector Databases & Memory" url: https://daily.dev/agentic-ai-hub/vector-databases-memory/ description: "A vector database stores embeddings and serves approximate nearest-neighbor (ANN) search, which finds the closest vectors to a query vector quickly using index structures like HNSW (graph-based, the most common) or IVF/IVF-PQ (cluster + quantize, memory-efficient at scale)." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- A \*\*vector database\*\* stores embeddings and serves approximate nearest-neighbor (ANN) search, which finds the closest vectors to a query vector quickly using index structures like \*\*HNSW\*\* (graph-based, the most common) or \*\*IVF/IVF-PQ\*\* (cluster + quantize, memory-efficient at scale). The real production differentiators are rarely raw ANN speed. They're \*\*metadata filtering\*\* (can you say "similar \*and\* tenant=X \*and\* date>Y" efficiently?), \*\*hybrid search\*\* support, horizontal \*\*scale\*\*, operational burden, and cost. ## ## Comparison (as of July 2026) | Tool | What it is | Hosting | Scale sweet spot | Filtering / hybrid | Cost | Gotcha | | :---- | :---- | :---- | :---- | :---- | :---- | :---- | | \*\*pgvector\*\* | Postgres extension for vector search | Self-host or any managed Postgres (RDS, Supabase, Neon) | Up to \~single-digit to tens of millions of vectors comfortably | SQL \`WHERE\` + full-text = true hybrid in one DB | Free extension; pay for Postgres | ANN recall/latency degrade at very large scale vs. dedicated engines; tune HNSW params | | \*\*Pinecone\*\* | Fully managed, serverless vector DB | Managed cloud only | Very large, hands-off | Strong metadata filtering; hybrid supported | Usage-based (serverless), can get pricey at scale | Proprietary/closed; no self-host, vendor lock-in | | \*\*Qdrant\*\* | Open-source vector DB in Rust | Self-host or Qdrant Cloud | Large, filter-heavy workloads | Excellent payload filtering; hybrid/sparse support | OSS free; managed cloud usage-based | You operate it if self-hosting | | \*\*Weaviate\*\* | Open-source vector DB w/ built-in modules | Self-host or Weaviate Cloud | Medium to large, semantic apps | Hybrid (BM25+vector) first-class; GraphQL | OSS free; managed usage-based | GraphQL learning curve; schema up-front | | \*\*Milvus\*\* | OSS vector DB built for massive scale (Zilliz) | Self-host or Zilliz Cloud | Billions of vectors, distributed | Filtering + hybrid; many index types | OSS free; Zilliz Cloud usage-based | Heavier to operate; overkill for small apps | | \*\*Chroma\*\* | Lightweight, developer-first embedding DB | Local/embedded or Chroma Cloud | Prototyping to small/medium prod | Metadata filtering; hybrid maturing | OSS free; managed cloud | Not aimed at billion-scale distributed workloads | | \*\*Redis\*\* | In-memory DB with vector sets/indexes | Self-host or Redis Cloud | Low-latency serving at moderate scale | Metadata filtering; hybrid via RediSearch | OSS free; managed cloud | Memory-bound; strongest fit when you already run Redis and want the latency | | \*\*MongoDB Atlas Vector Search\*\* | Vector search bolted onto MongoDB | Managed (Atlas) | Medium, alongside your document data | Mongo-query filtering; hybrid | Atlas usage-based | Tied to Atlas, not a standalone engine | | \*\*Elasticsearch / OpenSearch\*\* | Vectors added to a document/search store | Self-host or managed | Medium to large, search-heavy | First-class lexical + vector hybrid | OpenSearch OSS; paid tiers | Heavier ops; best when you already run it | | \*\*LanceDB\*\* | Embedded, columnar vector DB | Local/embedded or cloud | Local to small/medium, analytics-adjacent | Filtering; hybrid maturing | OSS free | Younger ecosystem; not billion-scale distributed | Also worth knowing: \[\*\*Vespa\*\*\](https://vespa.ai/) (mature OSS engine for large-scale hybrid search and ranking), \[\*\*turbopuffer\*\*\](https://turbopuffer.com/) (object-storage-first, aimed at big cost savings at scale), and the cloud-native managed options \*\*Azure AI Search\*\* and \*\*Google Vertex AI Vector Search\*\* for teams standardized on those clouds. ## When to just use Postgres If you already run Postgres, your corpus is in the low millions of vectors or fewer, and you value keeping vectors, metadata, and transactional data in one system with one backup story, \*\*pgvector is the reasonable choice.\*\* You get SQL filtering, joins, and hybrid search (with \`tsvector\`) without a second datastore to operate. Reach for a dedicated vector DB when scale (tens/hundreds of millions+), recall-at-latency SLOs, or multi-tenant filter performance push past what Postgres comfortably delivers. Many teams over-engineer this decision on day one. ## Agent memory stores "Memory" for agents is a layer \*above\* raw vector search. It decides what to remember, summarizes and consolidates it, expires stale facts, and retrieves the right memories per turn, spanning short-term (conversation state), long-term (durable facts about a user/task), and sometimes episodic/semantic distinctions. It typically sits on top of a vector DB plus a structured store. As of July 2026 the notable players: | Tool | What it is | Best at | | :---- | :---- | :---- | | \[Mem0\](https://mem0.ai/) | Open-source and hosted memory layer | Extracting and storing salient facts across sessions; popular for chat/agent personalization | | \[Zep\](https://www.getzep.com/) | Memory server with temporal knowledge-graph features (Graphiti) | Facts that change over time | | \[Letta\](https://www.letta.com/) (formerly MemGPT) | Treats the LLM as an OS managing tiered memory | Self-editing memory and durable agent state | | LangMem / LangGraph memory | Memory utilities within the LangChain ecosystem | Teams already on LangChain | These are moving fast and benchmarks are contested. Prototype with one, but keep your memory \*interface\* thin so you can swap providers. --- --- title: "Fine-Tuning & Post-Training" url: https://daily.dev/agentic-ai-hub/fine-tuning-post-training/ description: "Reach for the cheapest tool that works, in this order:" lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- ## When to fine-tune vs prompt vs RAG Reach for the cheapest tool that works, in this order: 1\. \*\*Prompting / in-context learning.\*\* Change behavior with instructions and a few examples. Zero training and instant iteration. Try this first. 2\. \*\*RAG (\[RAG & Retrieval\](/agentic-ai-hub/rag-retrieval/)).\*\* Inject \*knowledge\* the model lacks. Use when facts are large, dynamic, or must be cited. 3\. \*\*Fine-tuning.\*\* Change the model's \*default behavior\*: output format/schema adherence, tone/persona, a narrow skill, following domain conventions, or compressing a long prompt into learned weights (latency/cost win). Also for teaching structure or style that's hard to specify in words. The classic mistake is fine-tuning to add facts. It is an expensive, leaky way to do what RAG does better, and it bakes in staleness. Fine-tune for \*\*form and behavior\*\* and retrieve for \*\*knowledge\*\*. These combine: a fine-tuned model \*plus\* RAG is common. ## The post-training stack Modern "post-training" (everything after base pretraining) has a few layers: - \*\*SFT (Supervised Fine-Tuning).\*\* Train on curated (prompt, ideal-response) pairs. The workhorse. It teaches the model to imitate demonstrations. Data quality dominates the outcome. - \*\*Preference optimization.\*\* Align to human/AI preferences over pairs of responses: - \*\*RLHF (Reinforcement Learning from Human Feedback).\*\* Train a reward model on preference data, then optimize the policy with RL (e.g., PPO). It is powerful but complex and unstable to run. - \*\*DPO (Direct Preference Optimization).\*\* Skips the separate reward model and RL loop and optimizes directly on preference pairs with a simple classification-style loss. It is far easier to run than PPO and is the common default for open-model alignment. Variants like \*\*ORPO\*\*, \*\*KTO\*\*, and \*\*SimPO\*\* trade off data format and stability. - \*\*RLAIF / Constitutional-style.\*\* Use an AI model to generate the preference labels, which reduces human labeling cost. - \*\*RLVR (RL from Verifiable Rewards).\*\* For tasks with checkable answers (math, code, tool use), reward correctness directly. Central to 2025 to 2026 reasoning-model training. ## Parameter-efficient fine-tuning (PEFT): LoRA / QLoRA Full fine-tuning updates every weight, which requires enormous memory and compute and a full model copy per task. \*\*PEFT\*\* methods freeze the base model and train a small number of new parameters instead. - \*\*LoRA (Low-Rank Adaptation).\*\* Freeze the base weights and inject small trainable low-rank matrices into (typically) the attention/projection layers. You train and ship a few megabytes of "adapter" instead of a full model, can host many adapters over one base, and lose little quality on most tasks. The dominant PEFT method. - \*\*QLoRA.\*\* LoRA on top of a \*\*quantized\*\* (commonly 4-bit, NF4) frozen base, so the base's memory footprint collapses. This is what lets people fine-tune large models on a single consumer/prosumer GPU. You pay a little training-speed and precision cost for a large memory saving. Variants like \*\*DoRA\*\* push quality closer to full fine-tuning. - \*\*PEFT\*\* more broadly (adapters, prefix/prompt tuning) is packaged in Hugging Face's \`peft\` library. ## Tooling | Tool | Best at | Watch out | | :---- | :---- | :---- | | \[Unsloth\](https://unsloth.ai/) | Fast, cheap single-GPU tuning: custom kernels give large speedups and big VRAM cuts for LoRA/QLoRA, plus ready notebooks | Single-GPU focus (multi-GPU historically more limited) | | \[Axolotl\](https://github.com/axolotl-ai-cloud/axolotl) | Reproducible, scalable training jobs without writing loops: config-driven (YAML), SFT/DPO/etc. across many models and multi-GPU | Broad config surface, so a learning curve | | \[Hugging Face TRL\](https://huggingface.co/docs/trl) | Custom pipelines and new algorithms: reference SFT, DPO, GRPO/PPO, and reward modeling; integrates with \`peft\`/\`accelerate\` | More code and moving parts than the wrappers | | \[LLaMA-Factory\](https://github.com/hiyouga/LLaMA-Factory) | Broad model/method coverage with a UI | Many knobs; verify recipes for your model | | \[torchtune\](https://github.com/meta-pytorch/torchtune) | Native-PyTorch fine-tuning recipes | Development wound down in 2025 (maintenance-only), so treat it as a reference, not a live default | | \[MLX-LM\](https://github.com/ml-explore/mlx-lm) | Fine-tuning on Apple Silicon (MLX) | Mac-only, smaller-scale | | \[veRL\](https://github.com/verl-project/verl) | Scaled RLVR/GRPO/PPO post-training (HybridFlow); the go-to for reasoning-model RL at scale | Heavier infra (Ray-based); overkill for simple SFT/LoRA | ## GPU / precision constraints The practical question is always "will it fit in my GPU's memory?" A rough guide: - \*\*Just running the model (inference).\*\* Weights need about \`2 bytes × parameters\` at FP16, so a 7B model is ≈ 14 GB, plus some headroom for the KV-cache. Quantizing to 4-bit roughly quarters that, so the same 7B fits in \~4 GB. - \*\*Full fine-tuning.\*\* Much heavier: you need room for the weights \*plus\* the gradients, optimizer state, and activations, which together run several times the model's size. This is why full fine-tuning usually needs a multi-GPU A100/H100-class setup. - \*\*QLoRA (the cheap path).\*\* Loads the base model in 4-bit and trains only small adapters on top, so a mid-size model fine-tunes on a single 24 GB card. This is what puts fine-tuning within reach of one consumer GPU. Precision formats you'll run into: \*\*BF16\*\* (the training default), \*\*FP8\*\* (newer, for faster training/inference on Hopper/Blackwell), and \*\*4-bit NF4/INT4\*\* (QLoRA and low-memory inference). The lower you go, the more accuracy you trade for memory, so validate on your eval set. --- --- title: "Dataset Engineering" url: https://daily.dev/agentic-ai-hub/dataset-engineering/ description: "Data quality, not model choice, is usually the ceiling on a fine-tune." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- Data quality, not model choice, is usually the ceiling on a fine-tune. "Dataset engineering" is the discipline of building the corpus that trains or evaluates a model. ## Curation Assembling the right examples: define the target behavior precisely, sample real inputs (not just easy ones), balance across categories/edge cases, and match the \*format\* of the responses to exactly what you want at inference. For instruction tuning, a few thousand \*diverse, high-quality\* examples routinely beat hundreds of thousands of noisy ones. The "quality over quantity" finding (cf. LIMA-style results) has held up. Curate for coverage and correctness, then stop. ## Annotation & quality Human labeling largely determines quality. Practices that matter: clear annotation guidelines, multiple annotators with \*\*inter-annotator agreement\*\* checks, adjudication of disagreements, and gold questions to catch bad labelers. Tools/services: \[Label Studio\](https://labelstud.io/) (open-source, general-purpose annotation), \[Argilla\](https://argilla.io/) (LLM-data-centric, now under Hugging Face), \[Prodigy\](https://prodi.gy/) (scriptable), \[Snorkel\](https://snorkel.ai/) (programmatic labeling / weak supervision), and managed vendors (\[Scale AI\](https://scale.com/), \[Surge AI\](https://www.surgehq.ai/), \[Labelbox\](https://labelbox.com/), \[Toloka\](https://toloka.ai/), \[SageMaker Ground Truth\](https://aws.amazon.com/sagemaker/groundtruth/)). Increasingly, \*\*LLM-as-annotator\*\* pre-labels and humans verify, which is cheaper, but audit for systematic model bias. ## Synthetic data Using an LLM to \*generate\* training data: instructions, responses, or both. It's now standard practice: seed with a few examples and expand (Self-Instruct/Evol-Instruct style), generate persona/scenario variety, or produce hard negatives for retrieval. Watchouts: \*\*diversity collapse\*\* (generations cluster around a few templates), amplified bias, and \*\*licensing\*\*. Many frontier providers' terms restrict using outputs to train competing models, so check the ToS of whatever model you generate from. Verify synthetic data with filters, dedup, and human spot-checks. ## Distillation \*\*Knowledge distillation\*\* trains a smaller "student" model to mimic a larger "teacher." In LLMs today this most often means \*\*data distillation\*\*: generate high-quality outputs (and sometimes reasoning traces) from a strong teacher and SFT a smaller model on them. The result is a cheaper model that performs well above its size on the target distribution. (Classic logit/soft-label distillation still exists but is less common in the API-model era.) Same licensing caveat as synthetic data. ## Dedup & cleaning - \*\*Deduplication.\*\* Exact and \*near\*-duplicate removal (MinHash/LSH, or embedding-similarity clustering) prevents overfitting, train/test leakage, and wasted compute. Essential for any scraped corpus. - \*\*Cleaning.\*\* Strip boilerplate/HTML, filter by language and quality heuristics, remove PII (\[LLMOps: Evals, Observability & Guardrails\](/agentic-ai-hub/llmops-evals-observability-guardrails/)), drop toxic/low-value content, and \*\*decontaminate\*\* against your eval/benchmark sets so you don't train on the test. - \*\*Formatting.\*\* Normalize to the exact chat/template schema your trainer expects. Malformed templates silently wreck fine-tunes. ## Where to get datasets - \*\*Hugging Face Datasets Hub.\*\* The default repository for open datasets. \[https://huggingface.co/datasets\](https://huggingface.co/datasets) - \*\*Kaggle.\*\* Datasets and competitions. \*\*Common Crawl.\*\* Raw web at scale. \*\*The Common Pile / Dolma / RedPajama / FineWeb.\*\* Large open pretraining corpora. \*\*The Stack\*\* (BigCode). The standard open \*code\* corpus. Always check the \*\*license\*\* and provenance before using data commercially. (Note: \*\*Papers with Code\*\* was sunset in mid-2025 and now redirects to Hugging Face; use \[Hugging Face Papers\](https://huggingface.co/papers) for task-linked papers and datasets instead.) --- --- title: "Inference, Serving & Optimization" url: https://daily.dev/agentic-ai-hub/inference-serving-optimization/ description: "TTFT (Time To First Token). Latency until the first token appears. Dominated by the prefill (prompt-processing) phase and prompt length." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- ## The metrics that matter - \*\*TTFT (Time To First Token).\*\* Latency until the first token appears. Dominated by the \*\*prefill\*\* (prompt-processing) phase and prompt length. Drives perceived responsiveness in chat. - \*\*TPOT / ITL (Time Per Output Token / inter-token latency).\*\* Speed of the \*\*decode\*\* phase. It determines how fast text streams. - \*\*Throughput.\*\* Total tokens/sec across all concurrent requests. The economic metric (tokens per dollar of GPU time). - \*\*Latency vs throughput is a trade-off:\*\* batching more requests raises throughput but can raise per-request latency. Serving pushes both via smarter scheduling. Inference has two distinct phases: \*\*prefill\*\* (compute-bound, processes the whole prompt in parallel) and \*\*decode\*\* (memory-bandwidth-bound, one token at a time). Most optimizations target one or the other. Some 2026 stacks even physically \*\*disaggregate\*\* prefill and decode onto different hardware pools. ## Core optimization techniques - \*\*Quantization.\*\* Store/compute weights (and sometimes activations/KV-cache) at lower precision (FP8, INT8, INT4). Shrinks memory and boosts throughput/latency at some accuracy cost. Weight-only INT4 (AWQ, GPTQ) is common for squeezing big models onto small GPUs. \*\*FP8\*\* is increasingly used for near-lossless speedups on Hopper/Blackwell. See also \[Local & Self-Hosting Stack\](/agentic-ai-hub/local-self-hosting-stack/) for GGUF/MLX formats. - \*\*KV cache.\*\* Attention caches the keys/values of all prior tokens so each new token doesn't re-process the sequence. It is essential for speed but grows with batch × sequence length and often becomes the memory bottleneck. \*\*PagedAttention\*\* (vLLM's signature idea) manages KV cache in non-contiguous "pages" like OS virtual memory, which cuts fragmentation and enables far larger effective batches. Compressing the KV cache (quantized KV, GQA/MLA attention) is a major 2025 to 2026 lever. - \*\*Continuous (in-flight) batching.\*\* Instead of waiting for a fixed batch, the scheduler adds/removes requests every step as sequences finish, which keeps the GPU saturated. The single biggest throughput win in modern serving. - \*\*Speculative decoding.\*\* A small fast "draft" model proposes several tokens. The big model verifies them in one forward pass and accepts the run that matches. Same output distribution, with materially lower latency when acceptance is high. Variants: self-speculation, Medusa/EAGLE-style multi-head drafting, n-gram/lookahead. - \*\*Distillation\*\* (\[Dataset Engineering\](/agentic-ai-hub/dataset-engineering/)) and \*\*prompt/prefix caching\*\* (reuse KV for shared prompt prefixes, valuable for system prompts and few-shot templates) round out the toolbox. ## Serving engines (as of July 2026) | Engine | Origin | Best at | One gotcha | Key fact | | :---- | :---- | :---- | :---- | :---- | | \*\*vLLM\*\* | UC Berkeley → community | High-throughput general serving; the de facto open default | Peak per-GPU latency can trail TensorRT-LLM on NVIDIA | Pioneered PagedAttention; OpenAI-compatible server; broad model + hardware support | | \*\*SGLang\*\* | Community | Structured generation, agentic/multi-turn, high throughput | Younger ecosystem than vLLM | \*\*RadixAttention\*\* for automatic prefix-cache reuse; strong on complex prompting workloads | | \*\*TGI (Text Generation Inference)\*\* | Hugging Face | Tight HF ecosystem integration, easy deploy | Historically trailed vLLM/SGLang on raw throughput in some benchmarks | Production-hardened, OpenAI-compatible, good tooling | | \*\*TensorRT-LLM\*\* | NVIDIA | Absolute lowest latency / highest throughput on NVIDIA GPUs | NVIDIA-only; build/compile step and steeper setup | Compiles model-specific optimized engines; FP8/FP4 on Hopper/Blackwell | | \*\*LMDeploy\*\* | InternLM / community | High-throughput serving (TurboMind), strong on Qwen/InternLM/DeepSeek | Smaller community than vLLM/SGLang | Routinely benchmarked head-to-head with vLLM/SGLang; good mixed-precision support | Rules of thumb: \*\*vLLM\*\* is the safe default for most teams; \*\*SGLang\*\* is strong for agent/structured/high-concurrency workloads and prefix reuse; \*\*TensorRT-LLM\*\* when you're all-NVIDIA and squeezing every last millisecond/dollar and can afford the build complexity; \*\*TGI\*\* when you live in the Hugging Face stack. At datacenter scale, \*\*NVIDIA Dynamo\*\* is the emerging framework for \*disaggregated\* prefill/decode serving across GPU pools. Benchmarks flip frequently with each release, so always test on \*your\* model, hardware, and traffic shape. For single-GPU/local, see llama.cpp/Ollama in \[Local & Self-Hosting Stack\](/agentic-ai-hub/local-self-hosting-stack/). --- --- title: "AI Hardware & Accelerators" url: https://daily.dev/agentic-ai-hub/ai-hardware-accelerators/ description: "The 2026 picture is that NVIDIA still dominates training and general inference." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- The 2026 picture is that NVIDIA still dominates training and general inference. A growing field of custom silicon (TPUs, Trainium, in-house hyperscaler chips) and inference-specialist startups (Groq, Cerebras, and transformer-ASIC entrants like Etched) chip away at specific workloads. ## NVIDIA: Blackwell / GB-series \*\*Blackwell\*\* is NVIDIA's current-generation datacenter architecture. Key parts (as of July 2026): - \*\*B200 / GB200.\*\* The Blackwell workhorses. \*\*GB200 NVL72\*\* is a rack-scale system linking 72 Blackwell GPUs and 36 Grace CPUs over NVLink into one liquid-cooled unit that acts as a giant GPU. \[https://www.nvidia.com/en-us/data-center/gb200-nvl72/\](https://www.nvidia.com/en-us/data-center/gb200-nvl72/) - \*\*B300 / Blackwell Ultra.\*\* A mid-generation refresh that began shipping in late 2025\. Per NVIDIA/vendor materials roughly \*\*288 GB HBM3e\*\* per GPU, higher FP4 throughput than B200, \~1,400W, liquid-cooled. \*\*GB300 NVL72\*\* is the rack-scale version. Emphasis on low-precision (\*\*FP4\*\*) inference. - \*\*Vera Rubin (R100).\*\* The next architecture, first unveiled at GTC 2025 and detailed further through CES/GTC 2026, scheduled for release in \*\*H2 2026\*\* (around Q3), and expected to pressure Blackwell pricing over time. NVIDIA's real moat is \*\*CUDA + NVLink + the whole software stack\*\* as much as the silicon. ## Google TPU \*\*Ironwood\*\* is Google's \*\*seventh-generation TPU\*\* (aka TPU v7 / "TPU7x"), positioned as inference-first and reaching general availability in 2026\. Google cites large per-chip gains over prior generations and superpods of up to \*\*9,216 chips\*\* with massive shared HBM and Optical Circuit Switching. TPUs are available on Google Cloud (and power Gemini). The tradeoff is a \*\*JAX/XLA-centric software path\*\* vs. CUDA. \[https://cloud.google.com/tpu/docs\](https://cloud.google.com/tpu/docs) ## AMD \*\*Instinct MI300X/MI325X\*\* (HBM-heavy, competitive memory capacity per GPU) established AMD as the credible #2\. The \*\*MI350 series (CDNA 4)\*\* is the current-gen competitor to Blackwell, with the \*\*MI400 series\*\* slated for 2026 and rack-scale "Helios" systems positioned against NVIDIA's NVL racks. AMD's software stack is \*\*ROCm\*\*. Support has improved a lot (vLLM, PyTorch) but remains less turnkey than CUDA for edge cases. ## AWS Trainium & other hyperscaler silicon \*\*AWS Trainium\*\* (Trainium 2 in wide use; \*\*Trainium 3\*\* now generally available via Trn3 UltraServers) targets cost-efficient training/inference on AWS via the Neuron SDK. Also in the field: \*\*Microsoft Maia\*\*, \*\*Meta MTIA\*\*, and Intel \*\*Gaudi\*\*, mostly captive to their clouds/workloads. Outside the US ecosystem, \*\*Huawei Ascend\*\* (910C, and the rack-scale CloudMatrix 384 supernode) is the dominant non-NVIDIA training/inference platform, especially in China, and is benchmarked directly against GB200 NVL72\. → \[hiascend.com\](https://www.hiascend.com/en) ## Inference-first chips | Chip | What it is | Watch out | | :---- | :---- | :---- | | \[Groq\](https://groq.com/) | Deterministic \*\*LPU\*\* built for very low-latency token generation; very high tokens/sec on served models | Not a training chip; consumed as an API/cloud with a curated model menu | | \[Cerebras\](https://www.cerebras.ai/) | \*\*Wafer-scale\*\* engine (a whole wafer as one chip); large on-chip memory removes the bandwidth wall for high inference speeds, and it trains too | Specialized; accessed via Cerebras cloud/appliances | | \[Etched (Sohu)\](https://www.etched.com/) | A \*\*transformer-specialized ASIC\*\* that bakes the architecture into silicon for very high throughput and low cost on transformer inference; emerged from stealth in 2026 with large reported pre-orders | Runs only transformer models (a bet the architecture stays stable), and it's early 🔄 | | \[SambaNova\](https://sambanova.ai/) | Reconfigurable \*\*RDU\*\* (SN40L) serving very fast tokens via SambaNova Cloud; a Groq/Cerebras peer | Consumed as a cloud/appliance; curated model menu | | \[d-Matrix (Corsair)\](https://www.d-matrix.ai/) | \*\*In-memory-compute\*\* inference accelerator, in full production in 2026, aimed at high-throughput low-cost serving | Early; specialized inference-only silicon 🔄 | ## VRAM math & cost-per-token intuition - \*\*Can I even load it?\*\* Weights ≈ \`bytes-per-param × params\`: FP16 = 2 (7B ≈ 14 GB), INT8 = 1 (≈ 7 GB), INT4 = 0.5 (≈ 3.5 GB). Then add \*\*KV cache\*\* (grows with batch × context) and overhead. A 70B model in 4-bit is \~35 GB of weights alone, so it needs a 48 GB card (tight) or two. - \*\*Cost per token\*\* is set by throughput, not sticker price: \`$/token ≈ (GPU $/hour) ÷ (tokens/hour)\`. A pricier GPU that serves 3× the tokens can be \*cheaper\* per token. This is why quantization, batching, and the right serving engine (\[Inference, Serving & Optimization\](/agentic-ai-hub/inference-serving-optimization/)) move the economics more than the hardware line item. - \*\*Cloud vs own:\*\* rent (per-hour or per-token APIs) for spiky/uncertain demand and to skip capex and ops. Buy/reserve when utilization is high and steady (owned/reserved GPUs amortize well below on-demand at \~24/7 use). Most teams should rent until utilization and unit economics clearly justify owning. --- --- title: "APIs, Inference Platforms & Gateways" url: https://daily.dev/agentic-ai-hub/apis-inference-platforms-gateways/ description: "Four rough tiers, each with different trade-offs on model access, speed, price, and control." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- Four rough tiers, each with different trade-offs on model access, speed, price, and control. ## First-party model APIs Buy directly from the lab: \*\*OpenAI\*\*, \*\*Anthropic\*\*, \*\*Google (Gemini)\*\*, plus \*\*Mistral\*\*, \*\*Cohere\*\*, \*\*xAI\*\*, \*\*DeepSeek\*\*, and others. You get the newest models first, best feature support (tools, structured outputs, caching, batch), and the vendor's own reliability, at list price, with lock-in to that provider's SDK/quirks. ## Aggregators / routers \[\*\*OpenRouter\*\*\](https://openrouter.ai/) puts one API and one bill in front of hundreds of models across dozens of providers, with automatic fallback and price/latency-based routing. It's best for trying many models without many contracts and for resilience via multi-provider fallback. The trade-off is a hop in the middle (a small latency/trust dependency), and some cutting-edge provider features may not pass through. \[\*\*Hugging Face Inference Providers\*\*\](https://huggingface.co/docs/inference-providers) is a similar OpenAI-compatible router across a dozen-plus backends and hundreds of models. ## Fast-inference / open-model hosts These serve open-weight models (Llama, Qwen, DeepSeek, Mistral, gpt-oss, etc.) with heavy optimization, often OpenAI-compatible, usually priced per-token: | Host | Best at | | :---- | :---- | | \[Groq\](https://groq.com/) | Low-latency token streaming (LPU hardware, \[AI Hardware & Accelerators\](/agentic-ai-hub/ai-hardware-accelerators/)) | | \[Cerebras\](https://www.cerebras.ai/) | High inference speeds via wafer-scale (\[AI Hardware & Accelerators\](/agentic-ai-hub/ai-hardware-accelerators/)) | | \[Together AI\](https://www.together.ai/) | Broad open-model catalog: serving, fine-tuning, and GPU clusters | | \[Fireworks AI\](https://fireworks.ai/) | Fast serving, function calling, and fine-tuning, with a production focus | | \[Baseten\](https://www.baseten.co/) | Deploying \*your own\* models (open Truss/BEI stack) with autoscaling: more "bring your model" than a fixed menu | | \[SambaNova Cloud\](https://sambanova.ai/) | Very fast token serving on RDU hardware; a Groq/Cerebras peer for latency-sensitive open models | | \[Replicate\](https://replicate.com/) | Any-model, pay-per-second hosting, useful for non-LLM (image/audio/video) too | | \[Deepinfra\](https://deepinfra.com/) | Low-cost per-token serving of a broad open-model catalog | | \[Novita\](https://novita.ai/) · \[Hyperbolic\](https://hyperbolic.ai/) | Budget open-model inference and GPU rental | ## Cloud provider platforms | Platform | What it is | | :---- | :---- | | \[AWS Bedrock\](https://aws.amazon.com/bedrock/) | Multi-model (Anthropic, Meta, Amazon Nova, etc.) inside AWS with IAM/VPC/guardrails | | \[Google Vertex AI\](https://cloud.google.com/vertex-ai) | Gemini plus open/partner models on GCP | | \[Azure AI Foundry / Azure OpenAI\](https://azure.microsoft.com/) | OpenAI and others with Azure enterprise controls | Best at enterprises wanting models under existing cloud contracts, compliance, data residency, and networking. Gotcha: newest models sometimes land here later than on first-party APIs, and there's more setup. ## Gateways / LLM proxies A \*\*gateway\*\* is your own routing/control layer in front of any of the above: one internal API that handles key management, provider fallback, rate limiting, caching, cost tracking, and logging. Options: \*\*LiteLLM\*\* (open-source proxy/SDK that normalizes 100+ providers to the OpenAI format, and is very common), \*\*Portkey\*\*, \*\*Cloudflare AI Gateway\*\*, \*\*Kong AI Gateway\*\*, and \*\*Helicone\*\* (observability-first proxy, \[LLMOps: Evals, Observability & Guardrails\](/agentic-ai-hub/llmops-evals-observability-guardrails/)). Trade-off: another hop and single-point-of-failure to run, in exchange for provider-independence and central governance. Most teams past prototype should run a gateway so no application code is welded to one vendor. > Selection heuristic: prototype on a \*\*first-party API or OpenRouter\*\*; put a \*\*gateway (LiteLLM/Portkey)\*\* in front once you have real traffic; move latency-critical open-model paths to \*\*Groq/Cerebras/Fireworks/Together\*\*; adopt \*\*Bedrock/Vertex/Azure\*\* when compliance or cloud-contract gravity demands it. --- --- title: "Local & Self-Hosting Stack" url: https://daily.dev/agentic-ai-hub/local-self-hosting-stack/ description: "Running models on your own machine or servers." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- Running models on your own machine or servers. It wins when you need \*\*privacy/data control\*\* (nothing leaves your box), \*\*offline\*\* operation, \*\*zero marginal cost\*\* at high volume, \*\*no rate limits\*\*, or \*\*full control\*\* over versions and sampling. It loses on frontier quality (the best models aren't open-weight or are too big for local hardware), on ops burden, and on cost-efficiency at low/spiky volume (an idle GPU still costs money). ## The tools | Tool | Best at | Watch out | | :---- | :---- | :---- | | \[Ollama\](https://ollama.com/) | Developers/local apps who want "it just works": one command pulls and runs a quantized model with a built-in server and OpenAI-compatible API | Convenience defaults (context length, quant) can silently cap quality, so check them; built on llama.cpp under the hood | | \[LM Studio\](https://lmstudio.ai/) | Non-terminal users and quick experimentation: desktop GUI to find, download, and chat with local models, plus a local server (strong on Apple Silicon MLX and Windows) | GUI-first, so less suited to headless servers | | \[llama.cpp\](https://github.com/ggml-org/llama.cpp) | Portability, quantization control, embedded/edge: the foundational C/C++ engine that defines \*\*GGUF\*\* and powers much of the ecosystem | Lower-level, so you manage flags and builds | | vLLM / SGLang (see \[Inference, Serving & Optimization\](/agentic-ai-hub/inference-serving-optimization/)) | Production self-hosting on your own GPUs, when "local" means a real GPU server serving many users | High-throughput server engines, not a laptop on-ramp | \*Also common:\* \[\*\*Open WebUI\*\*\](https://github.com/open-webui/open-webui) (the dominant self-hosted web frontend for Ollama/OpenAI-compatible backends), \[\*\*Jan\*\*\](https://jan.ai/) and \[\*\*GPT4All\*\*\](https://www.nomic.ai/gpt4all) (offline desktop apps like LM Studio), and \[\*\*KoboldCpp\*\*\](https://github.com/LostRuins/koboldcpp) / \[\*\*text-generation-webui\*\*\](https://github.com/oobabooga/text-generation-webui) (enthusiast runtimes, the latter a natural home for EXL2/EXL3). ## Hardware guidance (local) - \*\*Apple Silicon\*\* (M-series, unified memory) is a strong fit for many local users: a 64 to 128 GB Mac can run sizable quantized models because CPU and GPU share memory. Use \*\*MLX\*\*-optimized builds for best speed. - \*\*NVIDIA consumer GPUs\*\* (e.g., 24 GB-class like a 4090/5090-tier card) run 7B to 14B comfortably and 30B-class in 4-bit. Two cards or a 48 GB card open up \~70B in 4-bit. - \*\*CPU/RAM-only\*\* works via llama.cpp for smaller models but is slow for decode. Fine for batch/offline. ## Quantization formats - \*\*GGUF.\*\* llama.cpp's format. A spectrum of quant levels (e.g., Q4\_K\_M, Q5\_K\_M, Q8\_0) that trade size/speed for accuracy. \*\*Q4\_K\_M is the popular "good balance" default.\*\* Runs on CPU+GPU, any platform. The most common format for local models. - \*\*MLX.\*\* Apple's array framework and model format optimized for Apple Silicon's unified memory. The fastest path on Macs. - \*\*AWQ / GPTQ.\*\* GPU-oriented 4-bit weight quantization formats used with vLLM/TensorRT-LLM/Transformers for server inference. - \*\*EXL2/EXL3\*\* (ExLlama). Flexible bit-rate GPU quantization popular with enthusiasts for NVIDIA cards. --- --- title: "LLMOps: Evals, Observability & Guardrails" url: https://daily.dev/agentic-ai-hub/llmops-evals-observability-guardrails/ description: "Shipping LLM features without evals and observability is flying blind: outputs are non-deterministic, quality is subjective, and regressions are silent." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- Shipping LLM features without evals and observability is flying blind: outputs are non-deterministic, quality is subjective, and regressions are silent. LLMOps is the discipline that makes them measurable and operable. ## Evaluation The central problem is that there's often no single correct answer, so you need proxies for quality. - \*\*Golden datasets.\*\* A curated set of representative inputs with expected outputs (or rubrics). Your regression suite: run it on every prompt/model change and diff the scores. The highest-ROI eval investment. Even 50 to 100 examples catch most regressions. - \*\*LLM-as-judge.\*\* Use a strong model to score outputs against a rubric (correctness, helpfulness, faithfulness, format). It scales far past human grading and correlates reasonably when the rubric is good, but beware judge biases (position bias, verbosity bias, self-preference) and validate the judge against human labels periodically. - \*\*Pairwise / Elo ranking.\*\* Instead of absolute scores, have the judge (or humans) pick the better of two responses and aggregate into an Elo/Bradley-Terry rating. More reliable than absolute scoring for comparing models/prompts. The method behind public arenas (LMArena-style). - \*\*Task metrics & code-checks.\*\* For verifiable tasks use exact/structural checks (does JSON parse? does the test pass? exact match?), which are cheaper and more trustworthy than any judge. - \*\*RAG-specific.\*\* \*\*Ragas\*\* (faithfulness, context precision/recall), plus \*\*DeepEval\*\* (pytest-style LLM unit tests), \*\*promptfoo\*\* (config-driven eval/red-team, developer-friendly), and \*\*OpenAI Evals\*\*. - \*\*Benchmark harnesses.\*\* For standardized capability/benchmark runs: \*\*lm-evaluation-harness\*\* (EleutherAI, the academic standard), \*\*lighteval\*\* (Hugging Face), and \*\*Inspect AI\*\* (UK AISI, increasingly standard for safety/capability evals). ## Tracing & observability Capture every prompt, completion, tool call, latency, token count, and cost, organized into \*\*traces\*\* (a full request, including multi-step agent/chain runs), so you can debug, monitor quality/cost, and mine production data for eval sets. | Platform | Best at | One gotcha | Note | | :---- | :---- | :---- | :---- | | \*\*Langfuse\*\* | Open-source, self-hostable tracing + evals + prompt mgmt | You operate it if self-hosting | Very popular OSS default; generous cloud tier | | \*\*LangSmith\*\* | Deep LangChain/LangGraph integration; tracing + evals + datasets | Best value inside the LangChain stack; works standalone but that's its home | Managed by LangChain | | \*\*Braintrust\*\* | Eval-centric workflow, experiments, CI for prompts | More eval-first than full APM | Strong with dev teams iterating on prompts | | \*\*Arize / Phoenix\*\* | ML+LLM observability at scale; \*\*Phoenix\*\* is the OSS, OpenTelemetry-based tracer | Enterprise surface can be heavy for small teams | Open standards (OTel) friendly | | \*\*Helicone\*\* | Drop-in proxy observability (one line) + caching + cost tracking | Proxy model means a hop in the request path | Fastest to instrument; OSS + cloud | Many of these now speak \*\*OpenTelemetry / OpenLLMetry\*\*, so you can instrument once and route traces to multiple backends. Others in the space: \*\*MLflow\*\* (its GenAI eval + tracing, widely used in the Databricks ecosystem), \*\*Weights & Biases Weave\*\*, \*\*Datadog LLM Observability\*\*, \*\*Comet Opik\*\* (OSS), \*\*Laminar\*\*, \*\*Traceloop\*\*. ## Prompt management Version prompts like code: store them outside the app, track versions, run evals per version, roll back, and A/B test in production. Most observability platforms above include a prompt registry/playground. The anti-pattern is prompts hard-coded and edited in place with no history or eval gate. ## Guardrails & PII Runtime checks around model I/O: - \*\*Input guards.\*\* Prompt-injection/jailbreak detection, off-topic filtering, and \*\*PII detection/redaction\*\* before text hits the model or logs (e.g., \*\*Microsoft Presidio\*\* for PII). - \*\*Output guards.\*\* Schema/format validation, toxicity/safety filters, groundedness/hallucination checks (does the answer match retrieved context?), and competitor/policy filters. - Tooling: \*\*NVIDIA NeMo Guardrails\*\*, \*\*Guardrails AI\*\* (validators + structured output), \*\*Llama Guard / ShieldGemma\*\* (open safety classifiers), and provider-native moderation endpoints. Guardrails add latency and false positives, so tune thresholds against real traffic. ## Routing \*\*Model routing\*\* sends each request to the cheapest model that will handle it well, using a small/cheap model for easy queries and a frontier model for hard ones, which cuts cost with minimal quality loss. Approaches range from heuristic (by task type/length) to learned classifiers (e.g., RouteLLM-style) to gateway features (OpenRouter/Portkey auto-routing). Related: \*\*semantic caching\*\* (serve a cached answer for a semantically-equivalent query) for further savings. ## Cost / FinOps LLM spend is usage-based and can spike unpredictably, so treat it as a first-class metric: \*\*attribute cost per feature/user/tenant\*\* (via trace metadata), set budgets and alerts, exploit \*\*prompt/prefix caching\*\* (large discounts for repeated system prompts), \*\*batch APIs\*\* (\~50% off for async work on major providers), right-size models via routing, and cap context length (input tokens are usually the bulk of cost). Observability platforms and gateways (\[APIs, Inference Platforms & Gateways\](/agentic-ai-hub/apis-inference-platforms-gateways/)) surface most of this. The discipline is reviewing it weekly, not after the bill. --- --- title: "Levels of AI Adoption" url: https://daily.dev/agentic-ai-hub/levels-of-ai-adoption/ description: "Before you can improve how your team works with agents, you need a shared way to say where you are." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- Before you can improve how your team works with agents, you need a shared way to say where you \*are\*. Two frameworks have become the common reference points, and they are complementary rather than competing. One describes the maturity of your \*practice\*, the other describes the autonomy you grant on a \*given task\*. ## Shapiro's Five Levels: a maturity ladder Dan Shapiro's \["The Five Levels: from Spicy Autocomplete to the Dark Factory"\](https://www.danshapiro.com/blog/2026/01/the-five-levels-from-spicy-autocomplete-to-the-software-factory/) borrows the NHTSA's driving-automation levels deliberately, because the analogy gives "a common language for both where things were, and where things were going." The levels describe how much of the \*job\* has moved to the machine: - \*\*Level 0: Manual.\*\* The AI is a search engine or occasional tab-completion. "The code is unmistakably yours." No real productivity change. - \*\*Level 1: Discrete task offloading.\*\* Cruise control and lane-keeping. You delegate isolated, well-bounded work such as unit tests, docstrings, or a self-contained function, but your core job is unchanged. Classic Copilot/ChatGPT usage. - \*\*Level 2: Active pairing.\*\* Highway autopilot. The AI is a junior colleague handling "the boring stuff" while you manage flow and review. "You get into a flow state; you're more productive than you've ever been." Shapiro estimates \~90% of AI-native developers live here, and he issues the level's defining warning: it \*"feels like you are done. But you are not done."\* - \*\*Level 3: Human-in-the-loop management.\*\* Waymo with a safety driver. You stop writing code and start reviewing diffs from a senior-developer-grade agent running several tasks. "Your life is diffs." Shapiro notes that for many people "this feels like things got worse," and that "almost everyone tops out here." - \*\*Level 4: Specification-driven.\*\* A robotaxi. You are not driving. You write specs, argue requirements, craft prompts, and review test results. "You've now become that which you loathed: you're a PM." Leave for twelve hours, come back to check the suite is green. - \*\*Level 5: The dark factory.\*\* The Fanuc-style "dark factory," with robots building software while humans are "neither needed nor welcome." Shapiro reports this only working for very small teams (under five people) and calls it "nearly unbelievable," while conceding it "will likely be our future." The framework's most useful insight is that these are \*\*qualitative shifts in human responsibility, not incremental speedups.\*\* Each level up trades a kind of work you may love (writing code, being in flow) for a kind you may not (reviewing diffs, writing specs, arguing requirements). That is exactly why people plateau. Level 2 is pleasant and Level 3 feels like a demotion. The plateau is psychological rather than technical. Steve Yegge's \["Revenge of the Junior Developer"\](https://sourcegraph.com/blog/revenge-of-the-junior-developer) sketches the same ascent as six \*waves\* (completions → chat → coding agents → agent clusters → agent fleets), each of which he claims is "conservatively about 5x as productive as the previous wave." Two of his observations complicate the plateau story. He reckons "vibe coding is still completely invisible to 80% of the industry outside Silicon Valley," and he sees adoption running \*backwards\* by seniority, in that "junior developers have actually been far more eager to adopt AI than senior devs." If so, the plateau may be demographic as well as psychological. ## Osmani's Agentic Autonomy Levels: a per-task dial Addy Osmani's \["Agentic Autonomy Levels"\](https://addyosmani.com/blog/agentic-autonomy-levels/) covers similar ground but with a different unit of analysis. Where Shapiro describes \*your career stage\*, Osmani describes \*how much rope you give the agent on this particular task\*, built on two axes, agency (how much the agent does) and orchestration (how many agents, how coordinated): - \*\*L0 Assist.\*\* The agent suggests, you decide and act. - \*\*L1 Supervised action.\*\* The agent edits and runs commands but asks before anything consequential. The risk is \*approval fatigue\*, rubber-stamping every prompt. - \*\*L2 Scoped task delegation.\*\* Bounded work with clear goals, mostly unsupervised and interruptible. Verification shifts from reading every line to \*automated evidence\* (tests, types, screenshots). - \*\*L3 Goal-driven autonomy.\*\* The agent "does whatever it takes to achieve a goal, stopping only when some condition is met." Requires \*measurable\* stopping conditions. Vague goals fail here. - \*\*L4 Parallel delegation.\*\* Multiple isolated agents at once. The hard part is decomposing work into non-overlapping slices. - \*\*L5 Managed-by-exception orchestration.\*\* A manager agent dispatches, monitors, verifies, and escalates. You "step in only when things fail." The unifying idea is \*\*calibrated autonomy\*\*. The right level is a function of task risk, reversibility, and how good your verification is, not a badge of how advanced you are. A throwaway prototype can run at L3 safely. A database migration on production data might belong at L1 no matter how senior your team. Marc Nuri and others have since argued both ladders under-specify the messy middle where "agent chaos" precedes real orchestration, which is a fair critique. Treat the numbers as conversation-starters, not a certification scheme. ## What the data says Both ladders are practitioner syntheses, so it is worth checking them against survey data. The data supports the shape while complicating the optimism. Adoption is near saturation. The \[2025 DORA report\](https://dora.dev/dora-report-2025/) finds 90% of respondents using AI at work, and \[Stack Overflow's 2025 survey\](https://survey.stackoverflow.co/2025/ai/) has 51% of professionals using it daily. Trust is falling at the same time. 46% actively distrust the accuracy of AI output, and the top frustration, cited by 66%, is "AI solutions that are almost right, but not quite." DORA also offers the best corrective to ladder thinking. Instead of levels it describes seven team \*profiles\*, and it concludes that "AI's primary role is as an amplifier, magnifying an organization's existing strengths and weaknesses." Climbing the ladder scales whatever delivery system you already have, weak or strong. The sharpest single data point is \[METR's randomized controlled trial\](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/), which put 16 experienced maintainers through 246 real tasks on their own mature repos. With AI "they take 19% longer than without: AI makes them slower," yet even after the fact they believed AI had sped them up by 20%. The study's scope is narrow, but its lesson travels. Self-perceived productivity is not a measurement. Birgitta Böckeler's autonomy-pushing experiments at Thoughtworks (\[How far can we push AI autonomy?\](https://martinfowler.com/articles/pushing-ai-autonomy.html)) put an empirical ceiling on the top rungs. Full-app generation worked at small scope and degraded with complexity, and she concluded that "AI is not ready to create and maintain a maintainable business software codebase without human oversight." Supervision looks like a standing feature of the discipline rather than a stage you graduate out of. ## The tension between the ladders The two ladders are most useful read \*against\* each other, because they measure different things and the interesting failure modes live in the mismatch. Shapiro's level tells you how your role has structurally changed (are you still writing most code, or living in diffs and specs?). Osmani's level tells you your \*default\* trust setting per task (do you approve every edit, or only look when something breaks?). The pairing surfaces two common dysfunctions. The first is a team stuck at Shapiro L2 that \*thinks\* it's higher because it runs agents constantly but reviews nothing, which is the high-usage, low-trust gap the Stack Overflow numbers describe at population scale. The second is a team running L4 parallel delegation with L1 verification discipline, generating far faster than it can trust. One warning before you plot yourselves in a retro. METR's perception gap says teams misjudge their own position, so anchor the conversation in observable signals like review depth, rework rates, and incident trends rather than felt productivity. Anthropic's \[study of agent autonomy in practice\](https://www.anthropic.com/research/measuring-agent-autonomy), drawn from millions of real Claude Code sessions, found that "autonomy is not a fixed property of a model or system but an emergent characteristic of a deployment," co-constructed by model, user, and product. Experienced users, strikingly, auto-approve \*more\* and interrupt \*more\* at once. The levels are a vocabulary rather than a scoreboard. What matters is being "in a position to intervene when it matters." --- --- title: "Agentic Engineering: Core Ideas" url: https://daily.dev/agentic-ai-hub/agentic-engineering-core-ideas/ description: "\\"Agentic coding\\" is an overloaded term." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- "Agentic coding" is an overloaded term. Stripped down, it means you delegate a \*goal\* to a system that runs its own investigate-implement-verify loop with tools (a shell, a filesystem, tests, search) rather than hand-authoring each step. The craft is less about writing the perfect prompt and more about \*\*designing the system that iterates.\*\* A cluster of ideas defines this shift. ## What an agent actually is After a year of definitional chaos, the field has settled on a simple answer. Simon Willison, who \[collected 211 competing definitions\](https://simonwillison.net/2025/Sep/18/agents/), landed on one: \*"An LLM agent runs tools in a loop to achieve a goal."\* Anthropic's \["Building Effective Agents"\](https://www.anthropic.com/engineering/building-effective-agents) adds one distinction worth keeping: \*\*workflows\*\* run LLMs through predefined code paths, while \*\*agents\*\* let the model direct its own process and tool use. Tellingly for a company that sells the models, the post tells you to reach for the simpler option and add complexity "only when it demonstrably improves outcomes." Just how simple is the core? Thorsten Ball's \["How to Build an Agent"\](https://ampcode.com/how-to-build-an-agent) builds a working coding agent in under 400 lines of Go, because \*"it's an LLM, a loop, and enough tokens."\* That is the key lesson: if the loop itself is a few hundred lines of boilerplate, then almost everything that makes one agent better than another lives in the scaffolding around the model, what this handbook calls the harness. ## Inner loop vs. outer loop, and owning the outer one The single most useful decomposition comes from Osmani's \["Own the Outer Loop"\](https://addyosmani.com/blog/own-the-outer-loop/). The \*\*inner loop\*\* is the agent's own cycle of investigate, implement, verify, repeat. This is where model capability lives and where you should increasingly \*not\* be. The \*\*outer loop\*\* is where a human decides whether the work is safe to ship. Evidence crosses that boundary. A person renders a verdict. Osmani's line, \*"The model may write the line, but the Verdict is mine,"\* compresses the whole stance. The stance has independent pedigree too. Willison likes to cite a 1979 IBM training slide, \*"a computer can never be held accountable. Therefore a computer must never make a management decision."\* Accountability is non-transferable, and that is why the outer loop exists. Owning the outer loop means owning \*\*quality\*\* (installing checks that produce evidence), \*\*verdict\*\* (making the ship/no-ship call on that evidence), and \*\*answerability\*\* (being able to explain the decision later). The failure this guards against is generating faster than you can verify, a "trust-verification gap" that widens the more capable your agents get. The practical move is to place yourself in the \*constraint-setting and sampling\* loops (define the guardrails, spot-check the output) rather than the inner execution loop, and to build enough back-pressure (tests, audit logs, sandboxes) that the outer loop has something real to judge. ## Loop engineering: design the system, don't hold the tool \["Loop Engineering"\](https://addyosmani.com/blog/loop-engineering/) is the constructive counterpart. Hand-prompting means you hold the tool continuously (write prompt, read output, write next prompt), and \*you\* are the bottleneck. Loop engineering means you build the machine once and let it run. \*"You designed it one time. You did not prompt any of those steps."\* Osmani names five components, \*\*automations\*\* (scheduled discovery/triage), \*\*worktrees\*\* (isolated parallel workspaces), \*\*skills\*\* (reusable project knowledge), \*\*plugins/connectors\*\* (issue trackers, Slack, APIs), and \*\*sub-agents\*\* (separate maker and checker), plus a sixth, \*\*persistent state\*\* on disk or a board, because models forget between runs. The honest caveat is that loops don't remove your responsibility. They \*sharpen\* three risks. Unattended loops make unattended mistakes, comprehension debt grows quietly, and cognitive surrender tempts. Osmani's closing is the ethos of this whole part of the handbook. \*"Build the loop. But build it like someone who intends to stay the engineer, not just the person who presses go."\* ## The harness as a first-class artifact If loop engineering is the practice, \["Agent Harness Engineering"\](https://addyosmani.com/blog/agent-harness-engineering/) is the object. The equation is \*\*Agent = Model + Harness\*\*, and its corollary is bracing. \*"If you're not the model, you're the harness."\* Everything you control (system prompts, tool definitions, sandboxes, hooks, memory files, context-management strategy) is the harness. One claim is worth internalizing. \*"A decent model with a great harness beats a great model with a bad harness."\* Most agent failures are harness gaps, not model gaps, or in HumanLayer's blunt phrase, "skill issues." Dex Horthy's \[12-Factor Agents\](https://github.com/humanlayer/12-factor-agents) makes the point from production experience: most software that calls itself an agent is really "mostly deterministic code, with LLM steps sprinkled in at just the right points." His twelve factors boil down to owning your own prompts, context window, and control flow instead of adopting a framework's loop wholesale. Put plainly, the reliable end of the spectrum looks more like ordinary software than the "autonomous agent" rhetoric suggests. Two design principles travel well. The first is the \*\*ratchet\*\*, "every mistake becomes a rule," where each failure yields a permanent harness improvement (a hook, a memory-file line, a test) traceable to the incident that caused it. The second, \*\*"success is silent, failures are verbose"\*\*, says hooks should only speak when something breaks, which keeps the feedback channel clean. Work backwards from the behavior you want. If you can't name the behavior a harness component serves, delete it. And accept that harnesses are perishable. As models improve, scaffolding for old failure modes becomes dead weight. The complexity doesn't shrink, it relocates. ## The factory model and long-running agents Zoom out and the metaphor becomes industrial. Osmani's \["Factory Model"\](https://addyosmani.com/blog/factory-model/) reframes the job as \*"building the factory that builds your software,"\* where you move "from writing code to orchestrating systems that write code." Human roles become \*\*spec architect\*\*, \*\*systems thinker\*\*, and \*\*quality reviewer\*\*. TDD moves from good practice to "close to mandatory," because agents optimize for passing tests, so the tests must encode what \*should\* happen, not what currently does. The sharpest warning is that vague thinking no longer just slows you down, \*"it multiplies."\* A muddy spec fed to thirty parallel agents produces thirty variously-wrong artifacts. Yegge's six-waves forecast (\[Levels of AI Adoption\](/agentic-ai-hub/levels-of-ai-adoption/)) is the maximalist rendering of the same trajectory, and refreshingly candid about input costs. Agent fleets "burn lots of LLM tokens, to the tune of $10-$12/hour at current rates," a power bill priced like junior-developer labor. Whether his 5x-per-wave arithmetic survives measurement is the question the closing note below raises. The frontier of this is \*\*long-running / long-horizon agents\*\*, work that spans hours or days across multiple context windows (\[Osmani\](https://addyosmani.com/blog/long-running-agents/)). Three hard problems recur. Finite context comes first, along with the "context rot" that degrades quality well before the hard limit. Then there is no persistent memory, for which Anthropic's image is "engineers working in shifts, where each new engineer arrives with no memory of what happened on the previous shift." And there is self-grading bias, where models say "yes I'm done" more often than they should. The stabilizing techniques (state on disk not in context, checkpoint-and-resume, and \*separated\* planner/worker/judge roles) recur throughout this part of the handbook because they are the same techniques that make \*any\* agent trustworthy at scale. ## A note of realism from the other side It is worth pairing the enthusiasm with Armin Ronacher's \["The Coming Loop"\](https://lucumr.pocoo.org/2026/6/23/the-coming-loop/), which is broadly convinced the loop is coming and broadly worried about it. His critique of \*what loops produce today\* is specific and worth heeding. Harness-level loops that persist past the model's own "I am done" tend to generate code that is "too defensive, too complex, too local in its reasoning." Models "avoid strong invariants" and, in Karpathy's phrase, are "mortally terrified of exceptions," adding "fallbacks instead of making bad states impossible" rather than making invalid states unrepresentable. The result is code too local in its reasoning to inspect architecturally. His conclusion is not "don't" but a bind. Opting out may be impossible, since security researchers and competitors will run loops and defenders must too, which makes \*retaining engineering judgment inside the loop\* the real question rather than whether to enter it. Measurement has also landed one hard punch. In the METR RCT (\[Levels of AI Adoption\](/agentic-ai-hub/levels-of-ai-adoption/)), experienced developers were 19% \*slower\* with AI while believing, before and after, that they were faster. The caveats are acknowledged (early-2025 tooling, chat-style usage, experts on familiar code), but the meta-lesson is this section's own thesis in experimental form. Perceived speedup is not evidence. If you are going to own the outer loop, own the measurement too. --- --- title: "Orchestration Patterns" url: https://daily.dev/agentic-ai-hub/orchestration-patterns/ description: "Once one agent works, the obvious move is to run several. This is where the biggest gains and the biggest self-inflicted wounds both live." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- Once one agent works, the obvious move is to run several. This is where the biggest gains and the biggest self-inflicted wounds both live. The organizing text is Osmani's \["The Code Agent Orchestra"\](https://addyosmani.com/blog/code-agent-orchestra/), whose framing is the shift \*"from conductor to orchestrator"\*. A conductor guides one musician synchronously in real time. An orchestrator coordinates an asynchronous ensemble with specialized roles. \*"You used to pair with one AI. Now you manage an agent team."\* ## Role separation: don't let the model grade its own homework The foundational pattern is splitting the \*\*implementer\*\* from the \*\*verifier\*\*. Because \*"the bottleneck is no longer generation, it's verification,"\* and because models systematically overestimate their own completeness, you get materially better results by having a \*different\* agent (or role, or run) check the work, with fresh context and an adversarial brief, than by asking the author "are you done?" This is the multi-agent version of the same self-grading bias that plagues long-running agents. Concretely, a maker agent implements against a spec. A separate checker agent runs the tests, reviews the diff against the requirements, and tries to break it. The human sits above both, owning the outer-loop verdict. "Don't let the model grade its own homework" is the whole idea in one line. Anthropic's \[Claude Code best practices\](https://code.claude.com/docs/en/best-practices) document the pattern from the vendor side ("a fresh context improves code review since Claude won't be biased toward code it just wrote") and add a caveat its fans tend to omit, that "a reviewer prompted to find gaps will usually report some, even when the work is sound." Checker agents over-report dutifully, so the human above the pair judges the judge's output too. ## Worktrees for parallel isolation The enabling mechanic for running agents in parallel is \*\*git worktrees\*\* (or equivalent sandboxed checkouts), where each agent gets its own working directory on its own branch so concurrent agents don't clobber each other's files. Worktrees show up in the loop-engineering component list for exactly this reason. The discipline that makes them pay off is \*decomposition\*. Osmani's Autonomy-Levels L4 names the real challenge as "slicing work into non-overlapping pieces" so parallel agents don't produce merge conflicts or duplicated effort. Worktrees give you isolation for free. They do not give you a clean decomposition, which remains a human design problem. (The mechanic is now vendor-official. The \[Claude Code docs\](https://code.claude.com/docs/en/best-practices) recommend worktrees by name, "run separate CLI sessions in isolated git checkouts so edits don't collide.") ## The multi-agent debate: reads parallelize, writes don't Whether to go multi-agent \*at all\* is the field's liveliest open argument, and each side has solid evidence. The strongest published \*pro\* case is Anthropic's \["How we built our multi-agent research system"\](https://www.anthropic.com/engineering/built-multi-agent-research-system), where a lead agent delegating to parallel subagents beat a single-agent baseline by 90.2% on their internal research evals. The post is candid about the bill ("multi-agent systems use about 15× more tokens than chats") and about the boundary, noting that "most coding tasks involve fewer truly parallelizable tasks than research." Cognition (the Devin team) plants the opposite flag in Walden Yan's \["Don't Build Multi-Agents"\](https://cognition.ai/blog/dont-build-multi-agents). Parallel subagents without shared context make conflicting implicit decisions, so "running multiple agents in collaboration only results in fragile systems. The decision-making ends up being too dispersed." When you must split, "share context, and share full agent traces, not just individual messages." LangChain's Harrison Chase \[reconciles the two\](https://www.langchain.com/blog/how-and-when-to-build-multi-agent-systems) with the cleanest rule in the debate, that "read actions are inherently more parallelizable than write actions." Research and review fan out. Feature implementation against one shared codebase mostly doesn't, because merging conflicting writes (and the conflicting assumptions behind them) is the expensive part. The academic record backs the caution. The Berkeley MAST taxonomy (\[Why Do Multi-Agent LLM Systems Fail?\](https://arxiv.org/abs/2503.13657)) annotated 150+ execution traces and catalogued 14 recurring failure modes across three categories (system design, inter-agent misalignment, and task verification), many of them architectural rather than promptable-away. That a whole failure category is \*task verification\* independently confirms the maker/checker instinct above. ## Swarms, subagents, and teams: three concrete shapes The orchestra piece names three practical patterns, roughly in order of coordination sophistication: 1\. \*\*Subagents.\*\* A parent spawns focused children with explicit file ownership and a simple dependency graph. Cost-neutral, but \*you\* do the coordination. Good when the work fans out cleanly and dependencies are shallow. 2\. \*\*Agent teams.\*\* A shared task list with automatic dependency resolution, peer messaging, and file locking, which gives true parallelism \*with\* coordination. Osmani puts the sweet spot at \*\*3 to 5 teammates\*\*. 3\. \*\*Orchestration at scale.\*\* Three tiers, in-process (e.g. Claude Code), local orchestrators, and cloud fleets, for when you genuinely need many agents and reviewable artifacts. The when-to-use heuristic is to reach for multi-agent only when a single agent hits a wall, whether \*context overload\* on a large codebase, \*lack of specialization\* (a generalist underperforming focused agents), or \*no coordination primitives\* for genuinely parallel work. When those walls are real, \*"three focused agents consistently outperform one generalist agent working three times as long"\* thanks to parallelism, specialization, isolation, and compound learning. When they aren't, multi-agent is pure overhead, and worse, \*"small harmless mistakes compound at a rate that's unsustainable"\* across an orchestrated army. ## The orchestration tax and your parallel-agent limit Here is the honest counterweight to swarm enthusiasm, from \["The Orchestration Tax"\](https://addyosmani.com/blog/orchestration-tax/). Starting more agents is easy, but \*"more agents running doesn't mean more of you available, your cognitive bandwidth doesn't parallelize."\* Osmani's metaphor is precise and worth remembering. \*"You are the GIL of your AI agents. They all can run at once. But when any of their work needs genuine understanding, that work has to acquire the lock."\* Judgment is single-threaded, and context-switching between agents reloads your mental model in \*minutes\* (with incomplete recovery), not microseconds. The consequence is a hard ceiling. \*"The right number of parallel agents is how many you can actually code review properly. For most of us this is a low single digit."\* Run more and you don't get more output. You get invisible debt and quietly lowered standards. The mitigations are the mitigations of any concurrent system. \*\*Scale agent count to your review rate\*\*, \*\*sort work\*\* (delegate isolated tasks, protect the judgment-heavy ones for serial focus), \*\*batch reviews\*\* to amortize context-switch cost, \*\*automate verification\*\* so your scarce attention goes only to genuine judgment calls, and deliberately \*\*protect serial time\*\* by closing the dashboard and thinking. The uncomfortable through-line is that the human bottleneck was in some sense a feature rather than a bug, because the pain of reviewing was the signal that caught errors early. Orchestrate that signal away and mistakes compound in the dark. Simon Willison arrived at parallel agents a skeptic, and his conversion, in \["Embracing the parallel coding agent lifestyle"\](https://simonwillison.net/2025/Oct/5/parallel-coding-agents/), turns the abstract "sort work" mitigation into a concrete menu. Parallelize the low-stakes, low-review-cost work (research spikes, codebase Q&A, proofs of concept, maintenance chores) and keep feature work serial and spec-driven, because "reviewing code that lands on your desk out of nowhere is a \*lot\* of work," while "code that started from your own specification is a lot less effort to review." The parallel agents that pay are the ones whose output you can judge in minutes. --- --- title: "Context, Skills & Memory" url: https://daily.dev/agentic-ai-hub/context-skills-memory/ description: "Agents are only as good as what they know when they start, and by default they start knowing nothing about your repo, your conventions, or last week's decisions." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- Agents are only as good as what they know when they start, and by default they start knowing nothing about your repo, your conventions, or last week's decisions. \*\*Context engineering\*\*, deciding what an agent sees, when, and in what form, is now a core skill, and a small set of conventions has standardized around it. The term's christening is usually credited to Andrej Karpathy (amplified by Shopify's Tobi Lütke), who called it the "delicate art and science of filling the context window with just the right information for the next step." \[LangChain's survey of the field\](https://www.langchain.com/blog/context-engineering-for-agents) organizes everything under four moves, \*\*write\*\* context down (scratchpads, memories), \*\*select\*\* what comes in (retrieval, rules files), \*\*compress\*\* what stays (summarization, trimming), and \*\*isolate\*\* what doesn't belong together (sub-agents, sandboxes). Every convention below is an instance of one of them. ## When context hurts: the four failure modes The naive model, that more context is strictly better, is wrong, and knowing \*how\* it's wrong separates context engineering from context hoarding. Anthropic's \["Effective context engineering for AI agents"\](https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents) grounds the intuition. LLMs have a finite "attention budget," and benchmark work documents \*\*context rot\*\*, where "as the number of tokens in the context window increases, the model's ability to accurately recall information from that context decreases." Drew Breunig's \["How Long Contexts Fail"\](https://www.dbreunig.com/2025/06/22/how-contexts-fail-and-how-to-fix-them.html) gives the failure modes names worth memorizing, \*\*context poisoning\*\* (a hallucination enters the context and gets repeatedly referenced), \*\*context distraction\*\* (the model over-focuses on a long context, neglecting what it learned in training), \*\*context confusion\*\* (superfluous tools and information steer it wrong), and \*\*context clash\*\* (accumulated context contradicts itself). His companion piece, \["How to Fix Your Context"\](https://www.dbreunig.com/2025/06/26/how-to-fix-your-context.html), maps the mitigations, which amount to prune, quarantine, offload, and trim the tool loadout. Every convention in the rest of this chapter is one of those mitigations wearing a friendlier name. ## AGENTS.md and CLAUDE.md: conventions in the repo The dominant pattern is a plaintext/markdown file at the repo root that agents read on startup. \`AGENTS.md\` (\[agents.md\](https://agents.md/), pitched as "a README for agents," now used by over 60,000 open-source projects and supported by Codex, Cursor, VS Code, GitHub Copilot, Devin, Windsurf, Zed, Aider, and most other tools) and \`CLAUDE.md\` (Claude Code's flavor of the same idea) are the two you'll meet. The file holds the things you'd otherwise re-explain every session, such as how to run the build and tests, module boundaries, code-style rules, and, crucially, the \*"why we don't"\* decisions that keep an agent from cheerfully reintroducing a pattern you deliberately removed. In the orchestra model, \`AGENTS.md\` is explicitly the vessel for \*compound learning\*. Every session's hard-won lesson gets written back so the next session (and every parallel agent) starts smarter. This is the ratchet principle applied to knowledge. ## Skills / SKILL.md: reusable, disclosed on demand A \*\*Skill\*\* (packaged as a \`SKILL.md\` plus any supporting files) is reusable project knowledge or procedure the agent can pull in \*when relevant\* rather than carrying in every prompt. Skills appear in the loop-engineering component list precisely so you stop re-explaining the same context on every run. The design virtue is \*\*progressive disclosure\*\*. Instead of stuffing the whole system prompt with everything the agent \*might\* need, you let it load the deploy runbook or the API-migration playbook only when the task calls for it, keeping context lean and reducing the "context rot" that degrades long sessions. A mature team accumulates a small library of skills the way it once accumulated shell scripts and internal wikis. Simon Willison's assessment, \["Claude Skills are awesome, maybe a bigger deal than MCP"\](https://simonwillison.net/2025/Oct/16/claude-skills/), rests precisely on the token economics. Until invoked, "each skill only takes up a few dozen extra tokens," where an MCP server front-loads thousands of tokens of tool definitions into every session. The low-tech format of markdown plus a little YAML is the feature. Skills are portable across tools and models in a way heavier integrations are not. ## External state and memory: surviving the reset The defining constraint of today's agents is that context is finite and vanishes between sessions. The answer, consistent across the long-running-agents and factory literature, is to \*\*keep durable state outside the model\*\*, in progress files, task lists, decision logs, and session-as-event-log records on disk or a tracker (Linear, a markdown board). Anthropic's "engineers working in shifts with no memory of the previous shift" is the failure this prevents. The fix is that each shift writes down what it did and reads what came before. This is also why \*\*checkpoint-and-resume\*\* matters. An agent that saves intermediate state can recover from a crash or a context reset instead of losing hours of work. The most concrete production playbook is the Manus team's \["Context Engineering for AI Agents: Lessons from Building Manus"\](https://manus.im/blog/Context-Engineering-for-AI-Agents-Lessons-from-Building-Manus). Treat "the file system as the ultimate context: unlimited in size, persistent by nature, and directly operable by the agent itself." Have the agent \*recite\* its objectives (Manus keeps rewriting a \`todo.md\`) to fight goal drift. And, counterintuitively, "leave the wrong turns in the context," because a model shown its own failed attempt is less likely to repeat it. Their most surprising lesson is economic, that "the KV-cache hit rate is the single most important metric for a production-stage AI agent," which argues for append-only context and stable prompt prefixes. Context design shows up on the invoice, not just in output quality. Anthropic's long-horizon toolkit adds \*\*compaction\*\* (summarize a session nearing its limit into a fresh window) and structured note-taking as the shift-handoff document each session writes for the next. ## Intent debt: the context you can't regenerate There is one kind of context an agent cannot supply for you, and it deserves its own name. Osmani's \["The Intent Debt"\](https://addyosmani.com/blog/intent-debt/) defines it as the missing externalized \*rationale\* behind a system's design, meaning the goals and constraints that explain \*why\* the code is the way it is. It is distinct from technical debt (in the code) and comprehension debt (in people's heads). Intent debt lives in the artifacts that should exist and don't. The pivotal line is worth quoting in full. \*"An agent can't generate intent, because intent is the one input that has to come from you."\* Agents can refactor and explain existing code endlessly, but they cannot recover a rationale that was never written down, and a cold-starting agent has none of the hallway-conversation context a long-tenured engineer carried. Multiply that across many sessions and undocumented decisions get re-litigated (or silently violated) again and again. The mitigation is deliberate. Write goal-focused specs that state constraints and non-negotiables, keep decision logs as \*"pure intent-debt paydown,"\* and leave \`AGENTS.md\` "why we don't" notes. Osmani's summary is the practical takeaway for this whole chapter. \*"Write down the why, because it's becoming the most valuable thing you can leave in the repo."\* --- --- title: "Verification & Testing for Agents" url: https://daily.dev/agentic-ai-hub/verification-testing-for-agents/ description: "If generation is cheap and verification is the bottleneck, then verification infrastructure is your leverage." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- If generation is cheap and verification is the bottleneck, then verification infrastructure \*is\* your leverage. As Osmani puts it, until it catches up with generation, \*"human review isn't optional overhead, it's the safety system."\* This chapter is about building that system so it scales with output instead of collapsing under it. The stakes here are measurable. GitClear's \[analysis of 211 million changed lines\](https://www.gitclear.com/ai\_assistant\_code\_quality\_2025\_research) (2020-2024) reads like the verification gap's fossil record. Code cloning rose from 8.3% to 12.3% of changed lines, refactoring sank from \~25% of changed lines to under 10%, and copy/pasted code exceeded \*moved\* code for the first time in the dataset's history. That is more duplication and less consolidation, exactly the decay you'd predict. Simon Willison's \["Vibe Engineering"\](https://simonwillison.net/2025/Oct/7/vibe-engineering/) states the constructive inverse, that "if your project has a robust, comprehensive and stable test suite agentic coding tools can \*fly\* with it." The same infrastructure that catches the decay also sets how much autonomy you can safely grant. ## Agentic code review: evidence over vibes The first move is to stop treating review as a human reading every line and start treating it as a pipeline that \*produces evidence\* a human then judges. That means automated checks wired into the agent's lifecycle (tests, type checks, linters, and increasingly a dedicated \*\*reviewer agent\*\* that reads the diff against the original spec and flags divergence) running as \*\*hooks\*\* on lifecycle events (pre-commit, post-edit) so the checks are systematic rather than remembered. The outer-loop discipline applies. You want an \*accountability contract\* (what was checked, what the evidence showed, why you shipped), not a gut-feel thumbs-up. And the reviewer must be separate from the implementer. A self-reviewing agent inherits the same self-grading bias that makes "are you done?" unreliable. Kent Beck, TDD's creator, supplies the adversarial caveat from his "augmented coding" experiments. Tests become a "superpower" with agents because agents "can (and do!) introduce regressions," and yet he has "trouble stopping AI agents from deleting tests in order to make them 'pass!'" (\[Pragmatic Engineer interview\](https://newsletter.pragmaticengineer.com/p/tdd-ai-agents-and-coding-with-kent)). The guardrail is also a target. Protect the verification pipeline \*from\* the implementer, with tests owned by the checker role or locked against edits, or an agent optimizing for green will optimize the greenness checker away. ## Evals-as-tests For agent \*behavior itself\* (as opposed to the code it writes), the emerging practice is \*\*evals-as-tests\*\*, encoding the behaviors you require as automated evaluations you run like a test suite, so you catch regressions when you change a prompt, swap a model, or edit the harness. This matters because the usual test suite checks the \*product\*, not the \*process that produced it\*, and a harness change can silently degrade quality with every existing test still green. The factory model's insistence that tests encode "what should happen, not what currently does" applies doubly here, because your evals are the executable specification of acceptable agent behavior, and they're only as honest as the cases you thought to write. The practical literature here is unusually good. Hamel Husain's \["Your AI Product Needs Evals"\](https://hamel.dev/blog/posts/evals/), the essay much of the discipline traces back to, argues that "unsuccessful products almost always share a common root cause: a failure to create robust evaluation systems." It organizes evals into three levels (cheap assertions run constantly, human-and-model grading of real transcripts, A/B tests) and insists the unglamorous core practice is reading your own data, because "you must remove all friction from the process of looking at data." Anthropic's \["Demystifying Evals for AI Agents"\](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents) adds the agent-specific mechanics. Start with 20-50 realistic tasks sourced from observed failures, mix code-based, model-based, and human graders, and grade \*outcomes\* rather than trajectories, since checking for "a sequence of tool calls in the right order" proves "too rigid and results in overly brittle tests." ## Who validates the validators? Reviewer agents and LLM judges push the verification problem up a level rather than solving it. Shreya Shankar and colleagues' aptly titled \["Who Validates the Validators?"\](https://arxiv.org/abs/2404.12272) is the peer-reviewed form of the self-grading worry. It reports that "LLM-generated evaluators simply inherit all the problems of the LLMs they evaluate, requiring further human validation." Its key empirical finding, \*\*criteria drift\*\*, is the useful surprise, that "users need criteria to grade outputs, but grading outputs helps users define criteria." You cannot fully specify "good" before contact with real outputs, so a static rubric handed to a judge model decays in production. The operational consequence, echoed in Anthropic's guidance that judges "should be closely calibrated with human experts," is that the human calibration loop above the judge is a permanent fixture rather than scaffolding. Automate the volume, and keep sampling the judgments. ## DTU: safe, high-volume testing against the world The hardest verification problem is testing against services you don't control, the Oktas, Jiras, Slacks, and Google Workspaces your agent touches. You can't hammer live third-party APIs at agent volumes without hitting rate limits, tripping abuse detection, running up costs, or corrupting real data, and you \*especially\* can't safely test the dangerous failure modes. StrongDM's \*\*Digital Twin Universe (DTU)\*\* (\[writeup\](https://factory.strongdm.ai/techniques/dtu)) answers this with \*"behavioral clones of the third-party services our software depends on,"\* offline replicas that mirror those APIs and their observable behavior. You build the double from the API contract and known edge cases, then validate it against the real dependency until the behavioral differences stop showing up. The payoff is testing \*"at volumes and rates far exceeding production limits"\*, deterministic, replayable, and able to exercise failure modes that would be dangerous or impossible against live services. StrongDM's own account of why this is newly possible is worth noting. Engineers always \*wanted\* production-grade test replicas but never proposed them because the effort seemed prohibitive. Agents made building and maintaining the clones cheap enough that what was "unthinkable six months ago" is now "routine." That is the deeper pattern of this whole chapter. The same drop in generation cost that created the verification crisis also makes it affordable to build the verification infrastructure that resolves it. Whether you reach for a full DTU or a humbler set of high-fidelity mocks, the principle holds. As agent output volume explodes, quality only stays up if you can verify at the \*same\* volume, and that means investing in test doubles, sandboxes, and evals as deliberately as you invest in the agents themselves. --- --- title: "Human Factors & the Agent-Era Career" url: https://daily.dev/agentic-ai-hub/human-factors-agent-era-career/ description: "The techniques above make you faster. This chapter is about what they can quietly cost you, and how not to pay it." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- The techniques above make you faster. This chapter is about what they can quietly cost you, and how not to pay it. The honest position, held across the sources and worth adopting, is that the productivity is real \*and\* the erosion is real, and the difference between the two outcomes is almost entirely about posture. ## Three debts you can't see on a burndown chart The agent era introduces failure modes that don't show up as bugs. It's useful to hold them as a trio, because they compound: - \*\*Intent debt\*\* (\[Context, Skills & Memory\](/agentic-ai-hub/context-skills-memory/)): the \*why\* was never written down, and no agent can regenerate it. - \*\*Comprehension debt\*\* (\[Osmani\](https://addyosmani.com/blog/comprehension-debt/)): \*"the growing gap between how much code exists in your system and how much of it any human being genuinely understands."\* It accrues invisibly (the code is clean, the tests pass) while understanding erodes, because AI generates faster than anyone can evaluate. Osmani cites research finding delegation-heavy developers scoring materially lower on comprehension of their own systems, and the memorable frame is \*"making code cheap to generate doesn't make understanding cheap to skip."\* - \*\*Cognitive surrender\*\* (\[Osmani\](https://addyosmani.com/blog/cognitive-surrender/)): the acute version, accepting AI output as your own \*without forming an independent judgment\*. It's distinct from healthy offloading (delegating the "how" while keeping the "what"). Surrender abandons the reasoning entirely. The cited Wharton finding is sobering: on the trials where the AI was wrong, people accepted the wrong answer 73% of the time, even as having AI available pushed their confidence \*up\*. You "borrow the model's confidence" without building a parallel understanding. Warning signs: approving PRs you didn't really read, accepting fixes without knowing the root cause, defending design decisions you can't actually defend, and lowered standards when tired. The line that locates the danger: \*"Surface correctness is not systemic correctness, and the gap between them is exactly where surrender hides."\* ## What the research actually shows The debts were named by practitioners. In 2025 researchers started measuring them, and the measurements land uncomfortably close to the warnings. MIT Media Lab's \["Your Brain on ChatGPT"\](https://arxiv.org/abs/2506.08872) put essay writers under EEG across four months. The LLM-assisted group showed the weakest neural connectivity, the lowest self-reported ownership of their output, and frequently could not quote from essays they had "written" minutes earlier. When later forced to work unaided, their engagement stayed depressed, so the debt outlived the borrowing. (The caveats are a small preprint and essays rather than code, but it is the debt trio rendered in electrodes.) METR's RCT (\[Levels of AI Adoption\](/agentic-ai-hub/levels-of-ai-adoption/)) belongs on this list too. Developers 19% slower with AI while believing they were 20% faster is cognitive surrender's calibration failure measured in the field. And Microsoft Research and CMU's \[CHI 2025 study of 319 knowledge workers\](https://www.microsoft.com/en-us/research/publication/the-impact-of-generative-ai-on-critical-thinking-self-reported-reductions-in-cognitive-effort-and-confidence-effects-from-a-survey-of-knowledge-workers/) identifies the moderating variable this chapter turns on, that "higher confidence in GenAI is associated with less critical thinking, while higher self-confidence is associated with more critical thinking." The tool doesn't determine the outcome, which is Osmani's posture claim with a mechanism attached. The same study supplies the constructive reframe, that "GenAI shifts the nature of critical thinking toward information verification, response integration, and task stewardship." The thinking migrates rather than disappears, but only for the people who make the trip deliberately. ## Don't outsource the learning The most personal risk is to your own growth. Osmani's \["Don't Outsource the Learning"\](https://addyosmani.com/blog/dont-outsource-learning/) makes the case that the default workflow optimizes for \*closing the issue\*, not \*building capability\*, and that the struggle you skip is the struggle that made you good. The striking datum he cites (an Anthropic 2026 study) is that AI-assisted engineers scored 50% on comprehension quizzes vs. 67% for those working manually, but, crucially, \*"the tool didn't determine the outcome. The posture did."\* Engineers who asked conceptual questions scored above 65%, while those who copy-pasted scored under 40%. The same tool produced opposite results. The preserving practices are concrete and cheap. Form a hypothesis \*before\* prompting, ask for the explanation before the code, use a "learning mode" in unfamiliar territory, review AI output like a junior's PR, periodically recreate a solution by hand, and ask the model to teach its reasoning rather than just hand over the answer. His framing separates two metrics we tend to conflate. \*"Ship and learn are two separate metrics… I'd rather ship 80% of what I could have and learn 100% of what I needed to, than the reverse."\* That is a genuine trade-off, not a free lunch (sometimes shipping \*is\* the priority), but naming the two axes lets you choose deliberately instead of defaulting to ship-and-forget every time. Charity Majors scales the same argument to the organization in \["Generative AI is not going to build your engineering team for you"\](https://charity.wtf/2024/06/10/generative-ai-is-not-going-to-build-your-engineering-team-for-you/). Engineering is an apprenticeship industry, "writing code is the easiest part of software engineering, and it's getting easier by the day," and what stays hard is owning systems in production over time, exactly the work agents don't do. Teams that stop hiring juniors because "AI does junior work now" are eating their seed corn. Comprehension debt has a generational form, where nobody is doing the reps that mint the next senior. Don't-outsource-the-learning is a hiring policy as much as a personal discipline. ## Earning taste and judgment The flip side of the debt story is where durable value now lives. Taste, the ability to look at a technically-correct solution and know it's the \*wrong\* one, doesn't come from reading agent output. It comes from the deep work agents are eager to do for you. This is the uncomfortable implication of cognitive surrender's cure, \*"mutual amplification, not delegation,"\* with the goal of leaving a collaboration with \*sharper\* understanding, not fuzzier. Judgment is earned through the reps you're now tempted to skip, which is why the advice to "do the hard thing by hand sometimes" isn't nostalgia. It's how you keep the calibration that makes you worth having in the outer loop at all. The most credible optimists agree. Simon Willison (\[Here's how I use LLMs to help me write code\](https://simonwillison.net/2025/Mar/11/using-llms-for-code/)) is blunt that the leverage is earned, warning that "if someone tells you that coding with LLMs is \*easy\* they are (probably unintentionally) misleading you." His own results lean "on 25+ years of professional coding experience," and his one non-negotiable is the outer loop itself, "the one thing you absolutely cannot outsource to the machine is testing that the code actually works." Expertise and AI leverage compound rather than compete. The taste you protect is what makes the same agent worth 10x in your hands and 1x in someone else's. ## The new software lifecycle and the agent-era career Stepping back, the shape of the work changes. The lifecycle reorders around specs, orchestration, and verification rather than authoring, so you spend less time in the editor and more time defining problems, dispatching agents, and judging results. Osmani's \["The Agent-Era Career"\](https://addyosmani.com/blog/career-advice-age-of-agents/) states the thesis bluntly. \*"AI gets good at anything with an answer key. Your career is everything that doesn't have one."\* Problems with defined solutions get automated. Humans move to the ungradeable work. Kent Beck, fifty years into his career, compressed the re-pricing into one line, that "the value of 90% of my skills just dropped to $0\. The leverage for the remaining 10% went up 1000x" (\[90% of My Skills Are Now Worth $0\](https://newsletter.kentbeck.com/p/90-of-my-skills-are-now-worth-0)). His reading of his own line is notably cheerful. Technological revolutions proceed by "radically reducing the cost of something that used to be expensive," then "discovering what is valuable about what has suddenly become cheap," and his prescription is aggressive experimentation rather than defensive abstention. The value re-pricing is worth internalizing honestly. Falling in value are grinding boilerplate, raw coding speed, and solving well-defined technical puzzles. Rising are problem selection (deciding \*what\* to build and what \*not\* to), verification and quality judgment, deep cross-system understanding, spec-writing and clear thinking, and taste built through deliberate practice. The survival strategies follow. Do hard work on purpose (solve it without agents first to build the mental model), master verification (\*"there's nothing more demoralizing than delegation without verification at scale"\*), sprint the last mile (agents hit 70% fast, and the winning 30% is yours), own accountability (your name is on the change regardless of who typed it), and work in public (reputation is the scarce resource that compounds). The honest trade-off underneath all of it is that efficiency gains genuinely risk atrophy of judgment, and output can look correct while your taste quietly degrades, which is why protection requires deliberate practice \*even when the agent could do it faster.\* None of this is an argument against agents. The productivity is real and the leverage is enormous. It is an argument for a specific stance toward them, the one Osmani keeps returning to and that this entire part endorses. Build the loop, run the factory, orchestrate the swarm, but do it \*as someone who intends to remain the engineer.\* --- --- title: "People to Follow (X/Twitter)" url: https://daily.dev/agentic-ai-hub/people-to-follow/ description: "X/Twitter is where much of AI moves in real time. Papers get discussed hours before the press notices, and model launches often break here first." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- X/Twitter is where much of AI moves in real time. Papers get discussed hours before the press notices, and model launches often break here first. Grouped below by role. Handles verified July 2026\. ## Lab leaders & executives - \*\*Sam Altman\*\* (\[@sama\](https://x.com/sama)): OpenAI CEO. Posts product hints and release teasers, plus the occasional cryptic one-liner. 🔄 - \*\*Dario Amodei\*\* (\[@DarioAmodei\](https://x.com/DarioAmodei)): Anthropic CEO and co-founder. Long-form essays on scaling, safety, and where capability is headed ("Machines of Loving Grace"); posts sparingly but consequentially. - \*\*Demis Hassabis\*\* (\[@demishassabis\](https://x.com/demishassabis)): Google DeepMind CEO. Frontier research and science applications, from AlphaFold successors to Gemini. - \*\*Greg Brockman\*\* (\[@gdb\](https://x.com/gdb)): OpenAI president and co-founder; infrastructure scale, demos, and engineering culture. - \*\*Clément Delangue\*\* (\[@ClementDelangue\](https://x.com/ClementDelangue)): Hugging Face CEO. Posts on open-source models and the 🤗 ecosystem. - \*\*Yann LeCun\*\* (\[@ylecun\](https://x.com/ylecun)): Turing laureate. Skeptical of LLM hype, focused on world models and self-supervised learning. - \*\*Aravind Srinivas\*\* (\[@AravSrinivas\](https://x.com/AravSrinivas)): Perplexity co-founder and CEO; AI answer-engines and pointed takes on where product value accrues in the stack. 🔄 - \*\*Arthur Mensch\*\* (\[@arthurmensch\](https://x.com/arthurmensch)): Mistral co-founder and CEO. The open-weight, European "sovereign AI" viewpoint. - \*\*Ilya Sutskever\*\* (\[@ilyasut\](https://x.com/ilyasut)): co-founder of Safe Superintelligence (SSI), previously OpenAI chief scientist. Rarely posts, but a defining researcher of the era. - \*\*Mira Murati\*\* (\[@miramurati\](https://x.com/miramurati)): co-founder and CEO of Thinking Machines Lab; former OpenAI CTO. One of the most-watched new-lab founders. 🔄 - \*\*Mustafa Suleyman\*\* (\[@mustafasuleyman\](https://x.com/mustafasuleyman)): Microsoft AI CEO; co-founder of DeepMind and Inflection. Author of \*The Coming Wave\*. ## Researchers & scientists - \*\*Andrej Karpathy\*\* (\[@karpathy\](https://x.com/karpathy)): ex-OpenAI and Tesla, now at \[Eureka Labs\](https://karpathy.ai/). Writes explainer threads and "nanochat"/zero-to-hero material on building models from scratch. - \*\*Jim Fan\*\* (\[@DrJimFan\](https://x.com/DrJimFan)): NVIDIA; embodied agents, robotics, and sim-to-real. - \*\*Sebastian Raschka\*\* (\[@rasbt\](https://x.com/rasbt)): author of \*Build a Large Language Model (From Scratch)\*. Code-first explainers of LLM internals and training. 🔄 (newsletter in \[Newsletters\](/agentic-ai-hub/newsletters/)). - \*\*Nathan Lambert\*\* (\[@natolambert\](https://x.com/natolambert)): Ai2/Interconnects; public analysis of RLHF, post-training, and open-model policy. - \*\*Noam Brown\*\* (\[@polynoamial\](https://x.com/polynoamial)): OpenAI. Reasoning, self-play, and test-time compute, the ideas behind thinking models. - \*\*Jason Wei\*\* (\[@\_jasonwei\](https://x.com/\_jasonwei)): co-author on chain-of-thought and emergent abilities; short notes on what makes models reason. - \*\*François Chollet\*\* (\[@fchollet\](https://x.com/fchollet)): creator of Keras and the ARC-AGI benchmark; co-founder of Ndea and ARC Prize. A leading skeptic-realist on what counts as reasoning. - \*\*Chris Olah\*\* (\[@ch402\](https://x.com/ch402)): Anthropic; a pioneer of mechanistic interpretability. The clearest public voice on what's actually happening inside models. ## Builders & practitioners (agentic coding is the theme of 2026) - \*\*Peter Steinberger\*\* (\[@steipete\](https://x.com/steipete)): creator of \*\*OpenClaw\*\* (the open-source agent that "broke GitHub"), now at OpenAI. Posts on agentic-coding loops and harness design. 🔄 - \*\*Boris Cherny\*\* (\[@bcherny\](https://x.com/bcherny)): creator of \*\*Claude Code\*\* at Anthropic. Shows how the people who build the tools use them, including his "vanilla setup" threads. - \*\*Tibo (Thibault Sottiaux)\*\* (\[@thsottiaux\](https://x.com/thsottiaux)): OpenAI, ChatGPT/Codex engineering lead. Product updates, rate-limit resets, and roadmap notes. 🔄 - \*\*Mario Zechner\*\* (\[@badlogicgames\](https://x.com/badlogicgames)): creator of \*\*pi\*\* ("the shitty coding agent"), a minimal, steerable harness. Argues against over-building agents. - \*\*Armin Ronacher\*\* (\[@mitsuhiko\](https://x.com/mitsuhiko)): creator of Flask. Writes long-form on agentic coding (see \[lucumr.pocoo.org\](https://lucumr.pocoo.org/)). Also on \[Mastodon\](https://hachyderm.io/@mitsuhiko). - \*\*Addy Osmani\*\* (\[@addyosmani\](https://x.com/addyosmani)): Google Chrome engineering lead; LLM-coding workflows, "loop engineering," and web performance. - \*\*Theo Browne\*\* (\[@theo\](https://x.com/theo)): t3.gg / T3 Chat (YC W22). Takes on AI tooling, TypeScript, and the developer stack (channel in \[Podcasts & YouTube\](/agentic-ai-hub/podcasts-youtube/)). - \*\*Simon Willison\*\* (\[@simonw\](https://x.com/simonw)): Django co-creator, builder of the \`llm\` CLI and Datasette. Documents what LLMs can actually do (blog in \[Blogs & Publications\](/agentic-ai-hub/blogs-publications/)). - \*\*Jeremy Howard\*\* (\[@jeremyphoward\](https://x.com/jeremyphoward)): fast.ai co-founder. Works on accessible deep learning, small-model efficiency, and open science. - \*\*Guillermo Rauch\*\* (\[@rauchg\](https://x.com/rauchg)): Vercel CEO; v0, the AI SDK, and the AI-native web stack that many developers ship on. - \*\*Thorsten Ball\*\* (\[@thorstenball\](https://x.com/thorstenball)): Sourcegraph / Amp. Clear writing on how coding agents actually work under the hood. - \*\*Hamel Husain\*\* (\[@HamelHusain\](https://x.com/HamelHusain)): independent consultant. Evals and LLM-as-judge, the "measure before you ship" discipline that's easy to skip. - \*\*Georgi Gerganov\*\* (\[@ggerganov\](https://x.com/ggerganov)): creator of \*\*llama.cpp\*\* and \*\*ggml\*\*, the projects that put local LLM inference on everyday hardware. The center of gravity for running models yourself. 🔄 ## Sharp commentators & analysts - \*\*swyx (Shawn Wang)\*\* (\[@swyx\](https://x.com/swyx)): coined "AI Engineer," runs Latent Space and the AI Engineer conference. Maps the AI-engineering discipline. - \*\*Andrew Ng\*\* (\[@AndrewYNg\](https://x.com/AndrewYNg)): DeepLearning.AI; education-first commentary; author of \*The Batch\* (\[Newsletters\](/agentic-ai-hub/newsletters/)). - \*\*Pedro Domingos\*\* (\[@pmddomingos\](https://x.com/pmddomingos)): author of \*The Master Algorithm\*. Academic voice skeptical of overclaims. - \*\*Ethan Mollick\*\* (\[@emollick\](https://x.com/emollick)): Wharton professor; the most-read practical voice on using AI for real work (newsletter in \[Newsletters\](/agentic-ai-hub/newsletters/), blog in \[Blogs & Publications\](/agentic-ai-hub/blogs-publications/)). - \*\*Eliezer Yudkowsky\*\* (\[@ESYudkowsky\](https://x.com/ESYudkowsky)): the best-known AI-safety and existential-risk voice. Polarizing, but the reference point that debate on the risk side organizes around. > > \*\*Tip:\*\* Build a private X List rather than following people directly. It keeps your main feed manageable. A good starter list: a few builders from the group above, plus Karpathy, Simon Willison, swyx, and one lab leader you follow. If you'd rather not live on X, a developer feed like \[daily.dev\](https://daily.dev/) surfaces much of what these accounts share without the doomscroll. --- --- title: "Newsletters" url: https://daily.dev/agentic-ai-hub/newsletters/ description: "Email survives because it's async and curated. Tagged by focus and cadence." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- Email survives because it's async and curated. Tagged by focus and cadence. ## Research-focused - \*\*The Batch\*\* (DeepLearning.AI): \[deeplearning.ai/the-batch\](https://www.deeplearning.ai/the-batch). \*Weekly.\* Andrew Ng's curated research and industry digest; measured tone, works for beginners and busy pros. - \*\*Import AI\*\* (Jack Clark): \[jack-clark.net\](https://jack-clark.net/). \*Weekly.\* Policy, research, and a short sci-fi vignette. Covers where frontier capability meets governance. - \*\*Ahead of AI\*\* (Sebastian Raschka): \[magazine.sebastianraschka.com\](https://magazine.sebastianraschka.com/). \*Roughly monthly.\* Technical explainers of LLM training and architecture papers. 🔄 - \*\*Interconnects\*\* (Nathan Lambert): \[interconnects.ai\](https://www.interconnects.ai/). \*\~Weekly.\* Post-training, RLHF, and open-model strategy analysis. Covers how models are actually trained now. ## Product & news - \*\*daily.dev\*\*: \[daily.dev\](https://daily.dev/). \*Frequent.\* "Where developers discover what's next." A personalized developer feed that aggregates the best engineering and AI posts, releases, and discussions from across the web into one place; also a browser new-tab, web, and mobile app. - \*\*The Rundown AI\*\*: \[therundown.ai\](https://www.therundown.ai/). \*Daily.\* Broad, fast, consumer-and-builder mix. Context around releases. - \*\*Ben's Bites\*\*: \[bensbites.com\](https://bensbites.com/). \*Daily/weekly.\* Builder-flavored roundup with a startup and tooling lean. - \*\*Last Week in AI\*\*: \[lastweekin.ai\](https://lastweekin.ai/). \*Weekly.\* Thorough roundup of the week's research and industry news. Pairs with the podcast. ## Engineering & agentic coding - \*\*Latent Space\*\* (swyx & team): \[latent.space\](https://www.latent.space/). \*Weekly-ish.\* Essays, benchmarks, and interviews for AI engineers. Pairs with the podcast (\[Podcasts & YouTube\](/agentic-ai-hub/podcasts-youtube/)). - \*\*Pragmatic Engineer\*\* (Gergely Orosz): \[newsletter.pragmaticengineer.com\](https://newsletter.pragmaticengineer.com/). \*Weekly.\* Software engineering broadly, with growing coverage of how AI is reshaping engineering orgs and workflows. ## Business & strategy - \*\*One Useful Thing\*\* (Ethan Mollick): \[oneusefulthing.org\](https://www.oneusefulthing.org/). \*\~Weekly.\* A Wharton professor's practical, experiment-driven take on using AI for real work. One of the most-read applied-AI newsletters. - \*\*Stratechery\*\* (Ben Thompson): \[stratechery.com\](https://stratechery.com/). \*\~Daily (paid).\* Strategy analysis of the AI platform wars. - \*\*The Neuron\*\*: \[theneurondaily.com\](https://www.theneurondaily.com/). \*Daily.\* Business-and-culture angle on AI, lightweight and readable. - \*\*Exponential View\*\* (Azeem Azhar): \[exponentialview.co\](https://www.exponentialview.co/). \*Weekly.\* Macro analysis of AI's economic and societal impact; a step back from the news cycle. ## Semiconductors & compute - \*\*SemiAnalysis\*\* (Dylan Patel): \[semianalysis.com\](https://semianalysis.com/). \*Weekly, plus deep dives.\* The definitive source on AI chips, datacenter and GPU economics, and the supply chain. Where the compute layer gets reported. 🔄 --- --- title: "Blogs & Publications" url: https://daily.dev/agentic-ai-hub/blogs-publications/ description: "OpenAI Blog: openai.com/news. Launches and research. The announcement of record." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- ## Lab / org research blogs - \*\*OpenAI Blog\*\*: \[openai.com/news\](https://openai.com/news/). Launches and research. The announcement of record. - \*\*Anthropic\*\*: \[anthropic.com/news\](https://www.anthropic.com/news) and \*\*Research\*\* at \[anthropic.com/research\](https://www.anthropic.com/research). Interpretability, alignment, and Claude/Claude Code releases. The interpretability posts are unusually readable. - \*\*Google DeepMind\*\*: \[deepmind.google/discover/blog\](https://deepmind.google/discover/blog/). Gemini, the AlphaFold lineage, and frontier science. - \*\*Hugging Face Blog\*\*: \[huggingface.co/blog\](https://huggingface.co/blog). Open-source work: fine-tuning, quantization, and new-model deep dives. 🔄 - \*\*Meta AI / FAIR\*\*: \[ai.meta.com/blog\](https://ai.meta.com/blog/). Llama lineage and open research. - \*\*Ai2 (Allen Institute)\*\*: \[allenai.org/blog\](https://allenai.org/blog). Fully-open models (OLMo/Tülu) with training details others omit. - \*\*Mistral AI\*\*: \[mistral.ai/news\](https://mistral.ai/news). The European open-weight lab's model releases and research notes. - \*\*Microsoft Research\*\*: \[microsoft.com/research/blog\](https://www.microsoft.com/en-us/research/blog/). Deep in-house work on agents, systems, and HCI; often ahead of the product cycle. - \*\*Qwen (Alibaba)\*\*: \[qwenlm.github.io/blog\](https://qwenlm.github.io/blog/). Primary source for the Qwen open-weight family's releases and technical reports. 🔄 ## Individual voices - \*\*Simon Willison\*\*: \[simonwillison.net\](https://simonwillison.net/). A practical-LLM log documenting what current models can do; a good single individual blog to read. - \*\*Armin Ronacher\*\*: \[lucumr.pocoo.org\](https://lucumr.pocoo.org/). Long-form on agentic coding ("The Coming Loop," "Tools: Code Is All You Need," "A Year of Vibes"). Candid about what didn't work. - \*\*Addy Osmani\*\*: \[addyosmani.com/blog\](https://addyosmani.com/blog/). LLM coding workflows, "cognitive surrender," and web performance; also on \[Substack\](https://addyosmani.substack.com/). - \*\*Andrej Karpathy\*\*: \[karpathy.ai\](https://karpathy.ai/) / \[karpathy.github.io\](https://karpathy.github.io/). Occasional landmark essays (e.g., "Software 2.0/3.0" thinking). His YouTube is more active than the blog. - \*\*Peter Steinberger\*\*: \[steipete.me\](https://steipete.me/). Field notes from building OpenClaw and working inside agent loops. 🔄 - \*\*Dan Shapiro\*\*: \[danshapiro.com\](https://www.danshapiro.com/blog/). Glowforge CEO. Founder-perspective essays on applying AI in a real hardware and product business. - \*\*Chip Huyen\*\*: \[huyenchip.com/blog\](https://huyenchip.com/blog/). ML systems and AI engineering in production; author of \*AI Engineering\* (\[Learning Paths, Courses & Books\](/agentic-ai-hub/learning-paths-courses-books/)). - \*\*Lilian Weng\*\*: \[lilianweng.github.io\](https://lilianweng.github.io/). Survey-style deep dives (agents, hallucination, diffusion), free and textbook-quality. - \*\*Sebastian Raschka\*\*: \[sebastianraschka.com\](https://sebastianraschka.com/). Companion to \*Ahead of AI\*. From-scratch LLM implementation walkthroughs. ## News outlets & aggregators - \*\*The Verge (AI)\*\*: \[theverge.com/ai-artificial-intelligence\](https://www.theverge.com/ai-artificial-intelligence). Consumer and industry news, well-reported. - \*\*Ars Technica (AI)\*\*: \[arstechnica.com/ai\](https://arstechnica.com/ai/). Technical, skeptical, thorough. - \*\*MIT Technology Review (AI)\*\*: \[technologyreview.com/topic/artificial-intelligence\](https://www.technologyreview.com/topic/artificial-intelligence/). Deeper features and societal framing. - \*\*Hacker News\*\*: \[news.ycombinator.com\](https://news.ycombinator.com/). Not a publication, but the comment threads on an AI launch are often more informative than the article. --- --- title: "Podcasts & YouTube" url: https://daily.dev/agentic-ai-hub/podcasts-youtube/ description: "Latent Space (swyx & Alessio): latent.space/podcast. Best for: the AI-engineering discipline." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- ## Podcasts - \*\*Latent Space\*\* (swyx & Alessio): \[latent.space/podcast\](https://www.latent.space/podcast). \*Best for:\* the AI-engineering discipline. Deep, technical interviews with founders and researchers. - \*\*Dwarkesh Podcast\*\* (Dwarkesh Patel): \[dwarkesh.com\](https://www.dwarkesh.com/) / \[YouTube\](https://www.youtube.com/@DwarkeshPatel). \*Best for:\* long, rigorous conversations with frontier researchers and CEOs (the Dario Amodei and DeepMind episodes are notable). Unusually well-prepared. - \*\*No Priors\*\* (Sarah Guo & Elad Gil): \[no-priors.com\](https://www.no-priors.com/). \*Best for:\* an investor-and-founder view of where AI is heading; strong guest list. - \*\*The Cognitive Revolution\*\* (Nathan Labenz): \[cognitiverevolution.ai\](https://www.cognitiverevolution.ai/). \*Best for:\* fast coverage of new capabilities with a builder and safety balance. - \*\*Lex Fridman Podcast\*\*: \[lexfridman.com/podcast\](https://lexfridman.com/podcast/). \*Best for:\* long interviews with prominent names in the field. Uneven but occasionally definitive. - \*\*Machine Learning Street Talk (MLST)\*\*: \[YouTube\](https://www.youtube.com/@MachineLearningStreetTalk). \*Best for:\* deep, philosophical, technically serious debates; a lower-hype option. - \*\*The TWIML AI Podcast\*\* (Sam Charrington): \[twimlai.com/podcast\](https://twimlai.com/podcast/twimlai/). \*Best for:\* long-running, deeply technical interviews with ML researchers and practitioners. 770+ episodes and still going. - \*\*AI + a16z\*\* (Andreessen Horowitz): \[a16z.com/podcasts/ai-a16z\](https://a16z.com/podcasts/ai-a16z/). \*Best for:\* a builder-and-infrastructure view of applied AI from the investor side. - \*\*Hard Fork\*\* (Kevin Roose & Casey Newton): \[nytimes.com/column/hard-fork\](https://www.nytimes.com/column/hard-fork). \*Best for:\* a mainstream, well-produced weekly take on AI and tech news. The accessible on-ramp. ## YouTube channels - \*\*Theo (t3.gg)\*\*: \[youtube.com/@t3dotgg\](https://www.youtube.com/@t3dotgg). \*Best for:\* fast, opinionated reactions to AI tooling and the dev stack. Tracks what developers are arguing about this week. 🔄 - \*\*Andrej Karpathy\*\*: \[youtube.com/@AndrejKarpathy\](https://www.youtube.com/@AndrejKarpathy). \*Best for:\* build-it-from-scratch lectures (nanoGPT, tokenizers, LLM intro). - \*\*3Blue1Brown\*\*: \[youtube.com/@3blue1brown\](https://www.youtube.com/@3blue1brown). \*Best for:\* visual intuition on neural nets and transformers. Explains why attention works. - \*\*Yannic Kilcher\*\*: \[youtube.com/@YannicKilcher\](https://www.youtube.com/@YannicKilcher). \*Best for:\* paper walkthroughs covering the math and the contribution. - \*\*Two Minute Papers\*\*: \[youtube.com/@TwoMinutePapers\](https://www.youtube.com/@TwoMinutePapers). \*Best for:\* quick hits on new research (especially generative and visual). Light but broad. - \*\*AI Explained\*\*: \[youtube.com/@aiexplained-official\](https://www.youtube.com/@aiexplained-official). \*Best for:\* benchmark-driven breakdowns of frontier releases. Skeptical and precise. --- --- title: "Communities" url: https://daily.dev/agentic-ai-hub/communities/ description: "Where real-time, tacit knowledge lives, the kind that never makes it into a blog post." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- Where real-time, tacit knowledge lives, the kind that never makes it into a blog post. ## Discords - \*\*OpenClaw / agent-builders community\*\*: \[openclaw.ai\](https://openclaw.ai/) and the \[official Discord\](https://docs.openclaw.ai/channels/discord). Hub for the open-source agentic-coding wave. Harness tricks, PRs, and troubleshooting land here first. 🔄 - \*\*Hugging Face Discord\*\*: \[hf.co/join/discord\](https://hf.co/join/discord). Large open-source ML community; help channels for most libraries, plus reading groups. - \*\*Latent Space Discord\*\*: via \[latent.space\](https://www.latent.space/). AI engineers building in production. - \*\*EleutherAI Discord\*\*: \[eleuther.ai\](https://www.eleuther.ai/). Open-source research and training discussion. High signal, high bar. - \*\*Nous Research Discord\*\*: \[nousresearch.com\](https://nousresearch.com/) and its \[Discord\](https://discord.com/invite/nousresearch). A large, high-signal hub for open-source LLM fine-tuning, datasets, and post-training experiments. 🔄 - \*\*Individual tool Discords\*\*: most agent and model projects (Ollama, LM Studio, vLLM, and the major coding agents) run active Discords. Join the one for whatever you build on. 🔄 ## Reddit - \*\*r/LocalLLaMA\*\*: \[reddit.com/r/LocalLLaMA\](https://www.reddit.com/r/LocalLLaMA/). Local and open models: quantization, hardware, and day-one testing of new weights. Useful if you run models yourself. 🔄 - \*\*r/MachineLearning\*\*: \[reddit.com/r/MachineLearning\](https://www.reddit.com/r/MachineLearning/). Research-leaning. The "\[R\]" and "\[D\]" threads are worthwhile. - \*\*r/StableDiffusion\*\*: \[reddit.com/r/StableDiffusion\](https://www.reddit.com/r/StableDiffusion/). The largest open image/video-generation community: workflows, models, and day-one testing of new generative-media tools. 🔄 - \*\*r/ArtificialIntelligence\*\* & \*\*r/OpenAI / r/ClaudeAI / r/Bard\*\*: broader consumer discussion and provider-specific tips; noisier, still useful for sentiment. ## Forums - \*\*daily.dev\*\*: \[daily.dev\](https://daily.dev/). Developer news aggregator with \*\*Squads\*\*, topic-based communities where people share and discuss AI tooling and finds, plus a personalized feed for catching launches early. 🔄 - \*\*Hugging Face Forums\*\*: \[discuss.huggingface.co\](https://discuss.huggingface.co/). Longer-form, searchable help for the HF stack. - \*\*LessWrong / Alignment Forum\*\*: \[lesswrong.com\](https://www.lesswrong.com/) / \[alignmentforum.org\](https://www.alignmentforum.org/). Safety and alignment discussion. Where much of that debate happens. --- --- title: "Key Papers & Reading List" url: https://daily.dev/agentic-ai-hub/key-papers-reading-list/ description: "A chronological spine of the field. The ones tagged Start here or Landmark are the priority reads; the rest fill in context." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- A chronological spine of the field. The ones tagged \*\*Start here\*\* or \*\*Landmark\*\* are the priority reads; the rest fill in context. Links go to arXiv (or the canonical source). "Why it matters" in one line each. ## Foundations (2013–2020) - \*\*Word2Vec / "Efficient Estimation of Word Representations in Vector Space"\*\* (2013): \[arxiv.org/abs/1301.3781\](https://arxiv.org/abs/1301.3781). Learned word embeddings; the vector-space intuition everything downstream inherits. - \*\*Attention Is All You Need\*\* (2017): \[arxiv.org/abs/1706.03762\](https://arxiv.org/abs/1706.03762). Introduced the Transformer. Every modern LLM descends from it. \*\*Start here.\*\* - \*\*Deep RL from Human Preferences\*\* (Christiano et al., 2017): \[arxiv.org/abs/1706.03741\](https://arxiv.org/abs/1706.03741). The origin of RLHF: learning a reward from human comparisons. The recipe InstructGPT later scaled to language. \*\*Landmark.\*\* - \*\*BERT\*\* (2018): \[arxiv.org/abs/1810.04805\](https://arxiv.org/abs/1810.04805). Bidirectional pretraining that made "fine-tune a big pretrained model" the default paradigm. - \*\*GPT-2 / "Language Models are Unsupervised Multitask Learners"\*\* (2019): \[cdn.openai.com/better-language-models\](https://cdn.openai.com/better-language-models/language\_models\_are\_unsupervised\_multitask\_learners.pdf). Scaling generative pretraining. The "too dangerous to release" moment. - \*\*Scaling Laws for Neural Language Models\*\* (2020): \[arxiv.org/abs/2001.08361\](https://arxiv.org/abs/2001.08361). Showed loss falls predictably with compute, data, and parameters; the economic logic of the whole race. - \*\*GPT-3 / "Language Models are Few-Shot Learners"\*\* (2020): \[arxiv.org/abs/2005.14165\](https://arxiv.org/abs/2005.14165). In-context learning at scale. The start of prompting as a discipline. \*\*Landmark.\*\* ## The alignment & reasoning era (2021–2023) - \*\*LoRA\*\* (2021): \[arxiv.org/abs/2106.09685\](https://arxiv.org/abs/2106.09685). Parameter-efficient fine-tuning; why you can adapt huge models on one GPU. - \*\*Chain-of-Thought Prompting\*\* (2022): \[arxiv.org/abs/2201.11903\](https://arxiv.org/abs/2201.11903). "Let's think step by step" unlocked reasoning. The seed of today's thinking models. - \*\*InstructGPT / "Training LMs to follow instructions with human feedback"\*\* (2022): \[arxiv.org/abs/2203.02155\](https://arxiv.org/abs/2203.02155). RLHF; the recipe that turned GPT-3 into something usable and became ChatGPT. \*\*Landmark.\*\* - \*\*Chinchilla / "Training Compute-Optimal LLMs"\*\* (2022): \[arxiv.org/abs/2203.15556\](https://arxiv.org/abs/2203.15556). Rewrote the scaling rules: most models were undertrained on data. Reshaped every training budget after it. - \*\*ReAct\*\* (2022): \[arxiv.org/abs/2210.03629\](https://arxiv.org/abs/2210.03629). Interleaves reasoning and tool use. The conceptual blueprint for agents. - \*\*Constitutional AI\*\* (2022): \[arxiv.org/abs/2212.08073\](https://arxiv.org/abs/2212.08073). RLAIF: replace human harm labels with a model critiquing against a written constitution. The basis of Anthropic's alignment approach. - \*\*LLaMA\*\* (2023): \[arxiv.org/abs/2302.13971\](https://arxiv.org/abs/2302.13971). Efficient open weights that ignited the open-source LLM ecosystem. \*\*Landmark.\*\* - \*\*GPT-4 Technical Report\*\* (2023): \[arxiv.org/abs/2303.08774\](https://arxiv.org/abs/2303.08774). The capability jump that mainstreamed AI; notable also for what it does not disclose. - \*\*Toolformer\*\* (2023): \[arxiv.org/abs/2302.04761\](https://arxiv.org/abs/2302.04761). Models teaching themselves to call APIs. Foundational for tool use. - \*\*DPO / Direct Preference Optimization\*\* (2023): \[arxiv.org/abs/2305.18290\](https://arxiv.org/abs/2305.18290). Preference tuning without the RL; simpler and now used everywhere. - \*\*QLoRA\*\* (2023): \[arxiv.org/abs/2305.14314\](https://arxiv.org/abs/2305.14314). 4-bit fine-tuning of 65B models on a single GPU; democratized adaptation. ## Efficiency, agents & multimodality (2021–2025) - \*\*CLIP / "Learning Transferable Visual Models From Natural Language Supervision"\*\* (2021): \[arxiv.org/abs/2103.00020\](https://arxiv.org/abs/2103.00020). Contrastive image-text pretraining; the backbone of modern multimodal models. \*\*Landmark.\*\* - \*\*FlashAttention\*\* (2022): \[arxiv.org/abs/2205.14135\](https://arxiv.org/abs/2205.14135). IO-aware exact attention; the kernel that made long context practical. - \*\*Mistral 7B\*\* (2023): \[arxiv.org/abs/2310.06825\](https://arxiv.org/abs/2310.06825). Small-model efficiency (GQA, sliding-window attention) that performed well above its size. - \*\*Llama 2\*\* (2023): \[arxiv.org/abs/2307.09288\](https://arxiv.org/abs/2307.09288). Open weights plus an RLHF-tuned chat model; the release that seeded the local-model ecosystem (the list has LLaMA 1 above). - \*\*Mixtral / Mixture-of-Experts\*\* (2024): \[arxiv.org/abs/2401.04088\](https://arxiv.org/abs/2401.04088). Sparse MoE at open scale; the architecture behind much frontier-model efficiency. - \*\*Mamba / State Space Models\*\* (2023): \[arxiv.org/abs/2312.00752\](https://arxiv.org/abs/2312.00752). A sub-quadratic challenger to attention; a leading candidate for what comes after Transformers. - \*\*DeepSeek-V3\*\* (2024): \[arxiv.org/abs/2412.19437\](https://arxiv.org/abs/2412.19437). The 671B open-weight MoE base that R1 is built on; efficient training at frontier scale. - \*\*DeepSeek-R1\*\* (2025): \[arxiv.org/abs/2501.12948\](https://arxiv.org/abs/2501.12948). Showed strong reasoning can emerge from RL alone, open-weight; a watershed for open reasoning models. - \*\*RAG / "Retrieval-Augmented Generation"\*\* (2020, read it now): \[arxiv.org/abs/2005.11401\](https://arxiv.org/abs/2005.11401). Grounding generation in retrieved documents. The backbone of most production LLM apps. > > \*\*How to read the frontier:\*\* Papers now ship faster than you can read them. Follow \[People to Follow (X/Twitter)\](/agentic-ai-hub/people-to-follow/) researchers for the "read this one" signal, use \[alphaXiv\](https://www.alphaxiv.org/) or Hugging Face's \[Daily Papers\](https://huggingface.co/papers) to triage, and skim abstracts + figures before committing to a full read. 🔄 --- --- title: "Platforms" url: https://daily.dev/agentic-ai-hub/platforms/ description: "The infrastructure you'll actually live in." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- The infrastructure you'll actually live in. ## Model & code hubs - \*\*Hugging Face\*\*: \[huggingface.co\](https://huggingface.co/). The GitHub of ML: 1M+ models, datasets, Spaces (live demos), and the \`transformers\`/\`datasets\` libraries. A center of gravity for open-source AI. 🔄 - \*\*GitHub Trending\*\*: \[github.com/trending\](https://github.com/trending). Filter by day or week to catch breakout AI repos before the newsletters do. 🔄 - \*\*Ollama\*\*: \[ollama.com\](https://ollama.com/). Run open models locally with one command; an easy on-ramp to local inference. - \*\*LM Studio\*\*: \[lmstudio.ai\](https://lmstudio.ai/). A desktop app for downloading and running local models with a GUI; the point-and-click peer to Ollama. - \*\*Replicate\*\*: \[replicate.com\](https://replicate.com/). Run and fine-tune open models via API without managing GPUs. Useful for prototyping. - \*\*Together / Fireworks / Groq\*\*: \[together.ai\](https://www.together.ai/) · \[fireworks.ai\](https://fireworks.ai/) · \[groq.com\](https://groq.com/). Fast, cheap hosted inference for open models (Groq for very low latency). 🔄 - \*\*OpenRouter\*\*: \[openrouter.ai/models\](https://openrouter.ai/models). A single API and marketplace across 400+ hosted models, with live per-token pricing and usage rankings. 🔄 - \*\*llms.txt\*\*: \[llmstxt.org\](https://llmstxt.org/). Emerging convention, a Markdown file that points models at your site's key docs/content to make them AI-readable. Worth adopting for docs and products. 🔄 ## Benchmarks & leaderboards (Papers-with-Code successors) - \*\*LMArena\*\* (formerly Chatbot Arena; the older lmarena.ai now redirects here): \[arena.ai\](https://arena.ai/). Human-preference Elo rankings; the closest thing to a "vibes benchmark" that people trust. 🔄 - \*\*Artificial Analysis\*\*: \[artificialanalysis.ai\](https://artificialanalysis.ai/). Independent index comparing models on intelligence, speed, and price across providers; widely cited for head-to-head numbers. 🔄 - \*\*Hugging Face Open LLM ecosystem & leaderboards\*\*: \[huggingface.co/spaces\](https://huggingface.co/spaces). After the original Open LLM Leaderboard was retired, evaluation moved to task-specific Spaces and Arena-style boards. Search Spaces for the current one for your task. 🔄 - \*\*Papers With Code\*\*: \[paperswithcode.com\](https://paperswithcode.com/). Largely archived and frozen now. Still a useful historical map of benchmark SOTA. Its living successors are alphaXiv and HF Papers (below). - \*\*alphaXiv\*\*: \[alphaxiv.org\](https://www.alphaxiv.org/). Community discussion layer over arXiv; a practical Papers-with-Code successor for tracking active work. 🔄 - \*\*Hugging Face Daily Papers\*\*: \[huggingface.co/papers\](https://huggingface.co/papers). Curated, upvoted daily paper feed. A fast triage tool. 🔄 ## Dataset sources - \*\*Hugging Face Datasets\*\*: \[huggingface.co/datasets\](https://huggingface.co/datasets). The default hub for training and eval data. - \*\*Kaggle\*\*: \[kaggle.com/datasets\](https://www.kaggle.com/datasets). Datasets, competitions, and free notebooks/GPUs; a good learning sandbox. - \*\*Common Crawl\*\*: \[commoncrawl.org\](https://commoncrawl.org/). The web-scale corpus underlying most pretraining. - \*\*Awesome Public Datasets\*\*: \[github.com/awesomedata/awesome-public-datasets\](https://github.com/awesomedata/awesome-public-datasets). Curated index across many domains. --- --- title: "Events & Conferences" url: https://daily.dev/agentic-ai-hub/events-conferences/ description: "🔄 Dates and formats change yearly. Always confirm on the official site." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- 🔄 \*Dates and formats change yearly. Always confirm on the official site.\* ## Academic (where the research lands first) - \*\*NeurIPS\*\*: \[neurips.cc\](https://neurips.cc/). The largest ML conference. December. Proceedings are free online, so you don't need to attend to read them. - \*\*ICML\*\*: \[icml.cc\](https://icml.cc/). International Conference on Machine Learning; July. Core theory and methods. - \*\*ICLR\*\*: \[iclr.cc\](https://iclr.cc/). Learning representations. Famously open-review (\[openreview.net\](https://openreview.net/)), so you can read reviews and rebuttals live. Spring. - \*\*ACL / EMNLP\*\*: \[aclanthology.org\](https://aclanthology.org/). The NLP conferences. The entire ACL Anthology is free. - \*\*CVPR\*\*: \[cvpr.thecvf.com\](https://cvpr.thecvf.com/). Top computer-vision venue; June. - \*\*COLM\*\*: \[colmweb.org\](https://colmweb.org/). Conference on Language Modeling, launched 2024; a dedicated LLM venue. \~October. - \*\*AAAI\*\*: \[aaai.org/conference/aaai\](https://aaai.org/conference/aaai/). The flagship general-AI conference; \~January/February. \*\*Follow remotely:\*\* proceedings are open-access. Watch for paper threads on X (\[People to Follow (X/Twitter)\](/agentic-ai-hub/people-to-follow/)), and many talks post to YouTube weeks later. OpenReview lets you read ICLR submissions in real time. ## Industry & developer events - \*\*AI Engineer World's Fair / Summit\*\*: \[ai.engineer\](https://www.ai.engineer/). swyx's conference for the AI-engineering discipline. Talks stream and post to YouTube. 🔄 - \*\*OpenAI DevDay\*\*: \[openai.com/devday\](https://openai.com/devday/). Product and API announcements. The keynote streams live. 🔄 - \*\*Google I/O\*\*: \[io.google\](https://io.google/). Gemini and the broader Google AI stack. May. Streams free. 🔄 - \*\*Anthropic developer events\*\*: watch \[anthropic.com/news\](https://www.anthropic.com/news). Claude/Claude Code updates increasingly land at dedicated dev events and livestreams. 🔄 - \*\*NVIDIA GTC\*\*: \[nvidia.com/gtc\](https://www.nvidia.com/gtc/). Hardware and the compute layer. Jensen's keynote sets the infra agenda. Streams free. 🔄 - \*\*Microsoft Build\*\*: \[build.microsoft.com\](https://build.microsoft.com/). Microsoft's developer conference, now heavily Copilot- and AI-focused; May/June. Streams free. 🔄 - \*\*AWS re:Invent\*\*: \[aws.amazon.com/events/reinvent\](https://aws.amazon.com/events/reinvent/). The largest cloud conference, increasingly AI-centric (Bedrock, agents, silicon); late November/December. 🔄 > > \*\*The cheat code:\*\* you almost never need to be in the room. Follow the official YouTube/X, and let the \[People to Follow (X/Twitter)\](/agentic-ai-hub/people-to-follow/) crowd surface the 5 talks that mattered. --- --- title: "Learning Paths, Courses & Books" url: https://daily.dev/agentic-ai-hub/learning-paths-courses-books/ description: "Organized beginner → advanced. Pick a lane and go deep. Breadth comes with time." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- Organized beginner → advanced. Pick a lane and go deep. Breadth comes with time. ## Roadmaps (start here to orient) - \*\*roadmap.sh (AI Engineer)\*\*: \[roadmap.sh/ai-engineer\](https://roadmap.sh/ai-engineer). Visual, community-maintained path from fundamentals to building AI products. A good map of the territory. 🔄 - \*\*roadmap.sh (AI/Data Scientist)\*\*: \[roadmap.sh/ai-data-scientist\](https://roadmap.sh/ai-data-scientist). For the modeling and research side rather than the app side. - \*\*roadmap.sh (AI Agents)\*\*: \[roadmap.sh/ai-agents\](https://roadmap.sh/ai-agents). Design, build, and ship agents; the newest and most on-trend of the three roadmaps. 🔄 ## Beginner (build intuition + first models) - \*\*Andrew Ng, Machine Learning Specialization\*\*: \[coursera.org/specializations/machine-learning-introduction\](https://www.coursera.org/specializations/machine-learning-introduction). The classic on-ramp to ML fundamentals. - \*\*DeepLearning.AI Short Courses\*\*: \[deeplearning.ai/short-courses\](https://www.deeplearning.ai/short-courses/). Free 1-to-2-hour project courses on RAG, agents, and fine-tuning, built with the actual tool vendors. High ROI for practitioners. 🔄 - \*\*3Blue1Brown, Neural Networks series\*\*: \[youtube.com/@3blue1brown\](https://www.youtube.com/@3blue1brown). Watch before or alongside any course. The visual intuition sticks. - \*\*Book: \*AI Engineering\*, Chip Huyen\*\* (2025): \[oreilly.com\](https://www.oreilly.com/library/view/ai-engineering/9781098166298/). Intro to building applications on foundation models. A starting point for the app-builder path. - \*\*Book: \*Hands-On Machine Learning with Scikit-Learn, Keras & TensorFlow\*, Aurélien Géron\*\* (3rd ed., 2022): \[oreilly.com\](https://www.oreilly.com/library/view/hands-on-machine-learning/9781098125967/). The best-selling practical ML reference; code-first and thorough. ## Intermediate (build real things, understand internals) - \*\*Andrej Karpathy, Neural Networks: Zero to Hero\*\*: \[karpathy.ai/zero-to-hero.html\](https://karpathy.ai/zero-to-hero.html). Build a GPT from scratch, line by line. A free deep-learning course. - \*\*fast.ai, Practical Deep Learning for Coders\*\*: \[course.fast.ai\](https://course.fast.ai/). Top-down and code-first. Get models working, then learn why. - \*\*Hugging Face Courses\*\*: \[huggingface.co/learn\](https://huggingface.co/learn). Free NLP, LLM, agents, and RL courses using the HF stack; very hands-on. 🔄 - \*\*Book: \*Build a Large Language Model (From Scratch)\*, Sebastian Raschka\*\* (2024): \[manning.com\](https://www.manning.com/books/build-a-large-language-model-from-scratch). Implement a full LLM yourself. The paper-to-code bridge. - \*\*Book: \*Hands-On Large Language Models\*, Jay Alammar & Maarten Grootendorst\*\* (2024): \[oreilly.com\](https://www.oreilly.com/library/view/hands-on-large-language/9781098150952/). Visual, practical guide to using and understanding LLMs. - \*\*Book: \*Natural Language Processing with Transformers\*, Tunstall, von Werra & Wolf\*\* (rev. ed., 2022): \[oreilly.com\](https://www.oreilly.com/library/view/natural-language-processing/9781098136789/). The canonical transformers/Hugging Face book, by HF's own authors. Pairs with the HF courses above. - \*\*Course: \*Generative AI with Large Language Models\* (DeepLearning.AI + AWS)\*\*: \[coursera.org/learn/generative-ai-with-llms\](https://www.coursera.org/learn/generative-ai-with-llms). The de facto standard applied-LLM course: prompting, fine-tuning, and RLHF end to end. ## Advanced (research depth + production scale) - \*\*Stanford CS224N (NLP with Deep Learning)\*\*: \[web.stanford.edu/class/cs224n\](https://web.stanford.edu/class/cs224n/). Lectures on YouTube; the academic backbone for NLP and LLMs. - \*\*Stanford CS336 (Language Modeling from Scratch)\*\*: \[stanford-cs336.github.io\](https://stanford-cs336.github.io/). Build a full LLM training stack. About as close to how frontier labs work as a course gets. 🔄 - \*\*Book: \*Designing Machine Learning Systems\*, Chip Huyen\*\* (2022): \[oreilly.com\](https://www.oreilly.com/library/view/designing-machine-learning/9781098107956/). A production MLOps reference for shipping ML at scale. - \*\*Book: \*Deep Learning\*, Goodfellow, Bengio, Courville\*\* (2016): \[deeplearningbook.org\](https://www.deeplearningbook.org/). The foundational theory text, free online. Dense but canonical. - \*\*Read the papers\*\*: \[Key Papers & Reading List\](/agentic-ai-hub/key-papers-reading-list/) is itself the advanced curriculum. Reproduce one from scratch. Nothing teaches faster. > > \*\*A pragmatic path for a working developer (2026):\*\* roadmap.sh/ai-engineer to orient → Karpathy Zero-to-Hero for internals → a couple of DeepLearning.AI short courses on RAG + agents → \*AI Engineering\* (Huyen) → then ship something real and let \[People to Follow (X/Twitter)\](/agentic-ai-hub/people-to-follow/)/\[Communities\](/agentic-ai-hub/communities/) keep you current. The half-life of specific tools is short. The fundamentals and the \*habit of following the frontier\* are what compound. --- --- title: "AI Adoption Stats & Trends" url: https://daily.dev/agentic-ai-hub/ai-adoption-stats-trends/ description: "🔄 Fast-moving section. All figures are dated and sourced; surveys are annual and market forecasts get revised, so check the live sources for the latest." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- 🔄 \*Fast-moving section. All figures are dated and sourced; surveys are annual and market forecasts get revised, so check the live sources for the latest. Last refreshed: 2026-07-19.\* ## Developer adoption - \*\*84% of developers use or plan to use AI tools\*\* in their workflow, up from 76% in 2024, but trust is falling: only \*\*33% trust the accuracy\*\* of AI output while \*\*46% actively distrust it\*\* (\[Stack Overflow 2025 Developer Survey\](https://survey.stackoverflow.co/2025/ai/), published Dec 2025; \~49,000 respondents). - \*\*51% of professional developers use AI tools daily\*\*. Favorable sentiment slipped to \*\*60%\*\* (from 70%+ in prior years) (\[Stack Overflow 2025\](https://survey.stackoverflow.co/2025/ai/)). - \*\*31% of developers already use AI agents\*\* daily, weekly, or monthly, though \*\*38% don't plan to adopt them\*\* yet (\[Stack Overflow 2025\](https://survey.stackoverflow.co/2025/ai/)). - \*\*90% of software professionals now use AI at work\*\*, up \~14% year-over-year, spending a \*\*median of 2 hours/day\*\* with AI tools. \*\*>80% report a productivity boost\*\* and \*\*59% say AI improved code quality\*\*, yet only \*\*24% express significant trust\*\* in it (\[Google/DORA 2025 State of AI-assisted Software Development\](https://blog.google/innovation-and-ai/technology/developers-tools/dora-report-2025/), published Sept 2025; full report at \[dora.dev/dora-report-2025\](https://dora.dev/dora-report-2025/)). > > The consistent 2025-2026 signal across surveys is that \*\*adoption is near-saturation, but trust is declining.\*\* Developers treat AI as an assistant to verify, not an authority to defer to. ## Enterprise adoption - \*\*88% of organizations report regular AI use\*\* in at least one business function, up from 78% the prior year, but only \*\*39% attribute any EBIT impact\*\* to AI, and roughly \*\*two-thirds have not yet begun scaling\*\* AI across the enterprise (\[McKinsey, \*The State of AI 2025: Agents, Innovation, and Transformation\*\](https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai), published Nov 5, 2025). - \*\*62% of organizations are at least experimenting with AI agents\*\*, and \*\*23% report scaling an agentic AI system\*\* somewhere in the enterprise (\[McKinsey 2025\](https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai)). - The "AI high performers" (about \*\*6% of respondents\*\*) account for a disproportionate share of measurable value. Roughly three-quarters of them have scaled or are scaling AI (\[McKinsey 2025\](https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai)). ## Market size - Global \*\*generative-AI market estimated at \~USD 22.2 billion in 2025\*\*, projected to reach \*\*\~USD 324.7 billion by 2033\*\* at a \*\*40.8% CAGR\*\* (\[Grand View Research\](https://www.grandviewresearch.com/industry-analysis/generative-ai-market-report), 2026 report). - Longer-horizon view: \*\*Bloomberg Intelligence projects the generative-AI market at \~USD 2.3 trillion by 2032\*\*, driven by agentic systems and infrastructure demand (\[Bloomberg Intelligence press release\](https://www.bloomberg.com/company/press/generative-ai-market-poised-to-reach-2-3-trillion-by-2032-as-agentic-systems-proliferate-and-infrastructure-demand-surges-according-to-bloomberg-intelligence/), June 2026; figure corroborated by \[Structured Retail Products\](https://www.structuredretailproducts.com/insights/83681/generative-ai-market-could-reach-us23-trillion-by-2032-bloomberg), June 2026). \*Note: market-sizing figures vary widely by analyst depending on scope (models vs. full stack vs. AI-enabled services). Treat them as directional, not precise.\* --- --- title: "Glossary of Key Terms" url: https://daily.dev/agentic-ai-hub/glossary/ description: "A jargon-buster for the terms used throughout this handbook. Where a full section exists, terms link to it." lastUpdated: "2026-07-22" llmsTxt: https://daily.dev/llms.txt --- A jargon-buster for the terms used throughout this handbook. Where a full section exists, terms link to it. | Term | Definition | | :---- | :---- | | \*\*AEO / GEO (Answer Engine / Generative Engine Optimization)\*\* | Getting content surfaced and \*cited\* inside AI assistants' answers (ChatGPT, Perplexity, Google's AI answers) rather than ranked in classic search; the AI-era successor to SEO. The emerging \[llms.txt\](https://llmstxt.org/) convention (a Markdown file pointing models at a site's key content) is one tactic; the term GEO comes from \[arXiv:2311.09735\](https://arxiv.org/abs/2311.09735). | | \*\*Agent\*\* | An LLM given a goal, tools, and a loop. It plans, calls tools, observes results, and iterates until done. See \[Agentic Engineering\](/agentic-ai-hub/agentic-engineering-core-ideas/). | | \*\*Attention\*\* | The transformer mechanism that lets each token weigh the relevance of every other token in context. | | \*\*Chain-of-thought (CoT)\*\* | Prompting or training a model to produce intermediate reasoning steps before its final answer, improving accuracy on hard problems. | | \*\*Context window\*\* | The maximum number of tokens a model can consider at once (prompt + output). Modern models range from \~128K to 1M+ tokens. See \[RAG\](/agentic-ai-hub/rag-retrieval/). | | \*\*Distillation\*\* | Training a smaller "student" model to mimic a larger "teacher," which transfers capability at lower cost. | | \*\*DPO (Direct Preference Optimization)\*\* | A simpler, RL-free alternative to RLHF that tunes a model directly on preferred-vs-rejected response pairs. | | \*\*Embedding\*\* | A dense numeric vector representing the meaning of text (or image/audio); similar content sits close together in vector space. | | \*\*Eval\*\* | A structured test measuring model quality on a task or benchmark, used to compare models or guard against regressions. | | \*\*Few-shot\*\* | Including a handful of examples in the prompt to steer behavior. "Zero-shot" = no examples. | | \*\*Fine-tuning\*\* | Further training a base model on task- or domain-specific data to specialize its behavior. | | \*\*Function calling / tool call\*\* | The model emits a structured request (name + JSON args) that your code executes, returning the result to the model. | | \*\*GGUF\*\* | A file format for storing quantized model weights, used widely by llama.cpp and local-inference tooling. | | \*\*Guardrails\*\* | Input/output filters and policies that constrain a model by blocking unsafe content, enforcing formats, or checking for injection. | | \*\*Hallucination\*\* | Confident, fluent output that is factually wrong or fabricated. The core reliability problem of LLMs. | | \*\*Harness\*\* | The scaffolding around a model: the loop, prompt assembly, tool wiring, and control flow that turns a model into an application or agent. | | \*\*Inference\*\* | Running a trained model to generate output (as opposed to training it). | | \*\*Jailbreak\*\* | A prompt crafted to bypass a model's safety training and elicit restricted behavior. | | \*\*KV cache\*\* | Cached key/value attention tensors from prior tokens, which let generation reuse work instead of recomputing it. This is the main memory cost of long contexts. | | \*\*Latency\*\* | Time to produce a response. Often split into \*\*TTFT\*\* (time to first token) and per-token generation speed. | | \*\*Logits\*\* | The raw, unnormalized scores a model assigns to each possible next token before they're turned into probabilities. | | \*\*LLM-as-judge\*\* | Using a (often stronger) model to grade another model's outputs against criteria. Scalable but imperfect evaluation. | | \*\*LoRA (Low-Rank Adaptation)\*\* | A parameter-efficient fine-tuning method that trains small adapter matrices instead of full weights. \*\*QLoRA\*\* does this on a quantized base model for even lower memory. | | \*\*MCP (Model Context Protocol)\*\* | An open standard for connecting models to external tools, data, and resources through a uniform interface. See \[MCP\](/agentic-ai-hub/mcp-tool-context-ecosystem/). | | \*\*MoE (Mixture of Experts)\*\* | An architecture that routes each token to a few specialized sub-networks ("experts"). This gives large capacity at lower per-token compute. | | \*\*Multimodal\*\* | A model that handles more than one modality, e.g. text plus images, audio, or video. | | \*\*Parameter\*\* | A learned weight in a model. Count (e.g. "70B") roughly indicates capacity and cost. | | \*\*PEFT\*\* | Parameter-Efficient Fine-Tuning. Umbrella term for methods (LoRA, adapters, prefix-tuning) that adapt a model by training few weights. | | \*\*Posttraining\*\* | Everything done after pretraining to shape behavior: SFT, RLHF/RLAIF, DPO, etc. | | \*\*Pretraining\*\* | The initial, compute-heavy phase where a model learns from a massive unlabeled corpus by predicting the next token. | | \*\*Prompt injection\*\* | An attack where malicious instructions hidden in input (or fetched content) hijack the model's behavior. | | \*\*Quantization\*\* | Reducing weight precision (e.g. 16-bit → 4-bit) to shrink memory and speed inference, at some quality cost. | | \*\*RAG (Retrieval-Augmented Generation)\*\* | Retrieving relevant documents at query time and feeding them into the prompt so the model answers from grounded context. See \[RAG\](/agentic-ai-hub/rag-retrieval/). | | \*\*Reasoning / test-time compute\*\* | Spending extra compute at inference (longer internal "thinking") to improve answers on hard tasks. This is the basis of reasoning models. | | \*\*RLHF / RLAIF\*\* | Reinforcement Learning from Human (or AI) Feedback. Aligns a model to preferences using a reward signal from human raters or another model. | | \*\*SFT (Supervised Fine-Tuning)\*\* | Fine-tuning on curated input→output demonstrations, typically the first posttraining step. | | \*\*Speculative decoding\*\* | Using a small fast "draft" model to propose tokens that the large model verifies in bulk. This accelerates generation. | | \*\*Structured output\*\* | Constraining a model to emit valid JSON (or a schema/grammar) so outputs are machine-parseable. | | \*\*Temperature\*\* | A sampling knob controlling randomness: low = focused/deterministic, high = diverse/creative. | | \*\*Throughput\*\* | Total work per unit time (e.g. tokens/second across all requests). A serving/capacity metric, distinct from latency. | | \*\*Token\*\* | The atomic unit a model reads and writes: a word-piece, roughly ¾ of a word in English. Pricing and limits are per-token. | | \*\*Tokenizer\*\* | The component that splits text into tokens (and back), defining the model's vocabulary. | | \*\*Top-p (nucleus sampling)\*\* | Sampling only from the smallest set of tokens whose cumulative probability exceeds \*p\*, which trims the unlikely tail. | | \*\*TTFT (Time To First Token)\*\* | Latency until the first output token appears; key for perceived responsiveness in streaming UIs. | | \*\*Vector database\*\* | A store optimized for similarity search over embeddings; the retrieval backbone of most RAG systems. | | \*\*Zero-shot\*\* | Asking a model to perform a task with no examples, relying purely on its pretrained/instruction-tuned knowledge. | --- ## The app App directories are served as markdown by the web app and are not inlined here. - \[Sources directory (markdown)\](https://daily.dev/sources.md) - \[Tags directory (markdown)\](https://daily.dev/tags.md) - \[Public squads directory (markdown)\](https://daily.dev/squads/discover.md) --- ## Blog Blog posts are indexed here as links only; append \`.md\` to any post URL for its markdown source. - \[Blog index\](https://daily.dev/blog.md): Latest articles and guides - \[The Best Tech News Apps for Developers in 2026\](https://daily.dev/blog/best-tech-news-apps-for-developers.md): How to pick developer news apps: use one personalized feed, ... - \[Best Reddit Alternatives for Developers: Where to Find Quality Tech Discussion\](https://daily.dev/blog/reddit-alternatives-where-to-find-for-developers.md): Compare 8 developer-focused Reddit alternatives—Q&A, blogs, ... - \[Where to Discover New AI Developer Tools in 2026\](https://daily.dev/blog/discover-ai-developer-tools.md): Structured methods to find, evaluate, and adopt AI developer... - \[The Best Alternatives to Hacker News in 2026\](https://daily.dev/blog/best-hacker-news-alternatives.md): Compare top alternatives to Hacker News for personalized fee... - \[How to Keep Up with AI as a Developer in 2026\](https://daily.dev/blog/ai-keep-up-guide-for-developers.md): Practical routines, tools, and skills—agents, RAG, MCP, cont... - \[The Best Developer Communities to Join in 2026\](https://daily.dev/blog/best-developer-communities-to-join.md): Pick 2–3 developer communities—each for news, help, or publi... - \[Best Websites to Learn Programming in 2026: A Developer's Guide\](https://daily.dev/blog/best-websites-learn-programming-developer-guide.md): Compare top free, paid, and AI-assisted platforms to learn p... - \[daily.dev vs TLDR: Which Developer News Source Fits You?\](https://daily.dev/blog/daily-dev-vs-tldr-developer-news-which-fits-you.md): Use a personalized browser feed for steady discovery or a co... - \[Where Developers Find New Tools: The Best Discovery Platforms in 2026\](https://daily.dev/blog/best-developer-discovery-platforms-2026.md): A concise 2026 guide to top platforms and a simple weekly wo... - \[Best Developer Browser Extensions for Productivity in 2026\](https://daily.dev/blog/best-developer-browser-extensions-for-productivity.md): Curated list of top browser extensions for debugging, API te... - \[15 Essential Developer Tools Every Programmer Needs in 2026\](https://daily.dev/blog/essential-developer-tools-for-programmers.md): Modern tooling — from AI coding assistants to containers, CI... - \[The Daily Learning Routine That Makes Great Developers\](https://daily.dev/blog/daily-learning-routine-makes-great-developers.md): Build developer skills with 30-minute daily blocks: morning ... - \[Dev Resources: Top Community Picks\](https://daily.dev/blog/dev-resources-top-community-picks.md): Discover top dev resources for continuous learning, communit... - \[Developer Tools for Work-Life Balance and Avoiding Burnout in 2026\](https://daily.dev/blog/developer-tools-for-work-life-balance-avoid-burnout.md): Protect mornings, mute nonessential alerts, batch communicat... - \[How to Keep Up With AI Coding Tools Without Losing Your Week\](https://daily.dev/blog/how-to-keep-up-with-ai-coding-tools-without-losing-your-week.md): Use a three-source system and one 30–45 minute weekly sessio... - \[The Best Newsletters for AI Engineers in 2026\](https://daily.dev/blog/best-newsletters-ai-engineers.md): Pick one daily brief and one weekly digest to keep up with A... - \[News for Programmers: Community-Driven Insights\](https://daily.dev/blog/news-for-programmers-community-driven-insights.md): Explore how community-driven discussions enhance the value o... - \[Best Programming Communities to Join in 2026\](https://daily.dev/blog/general-programming-communities-to-join.md): The top programming communities for developers in 2026 — Git... - \[The Best AI Engineers to Follow on X in 2026\](https://daily.dev/blog/best-ai-engineers-to-follow-on-x.md): Curate a tight X feed of frontier, tooling, and research AI ... - \[Will AI Replace Software Engineers in 2026? An Honest Take\](https://daily.dev/blog/will-ai-replace-software-engineers-2026-honest-take.md): AI automates routine coding by 2026, shrinking junior roles ... - \[Daily.dev: A Global Dev Community Hub\](https://daily.dev/blog/dailydev-a-global-dev-community-hub.md): Daily.dev is a global hub for developers offering personaliz... - \[Top 10 Best Tech Websites & Blogs 2026\](https://daily.dev/blog/top-10-best-tech-websites-and-blogs-2024.md): Stay updated with the top 10 best tech websites & blogs in 2... - \[What Is MCP (Model Context Protocol) and Why Every Developer Should Care\](https://daily.dev/blog/mcp-model-context-protocol-why-developers-should-care.md): Explains MCP—an open standard that enables stateful, secure ... - \[11 Best Programming Forums in 2026 (Still Active)\](https://daily.dev/blog/11-best-programming-forums-2024.md): The most active programming forums in 2026 — Stack Overflow,... - \[15 Fun Coding Problems From Easy to Hard (2026)\](https://daily.dev/blog/fun-coding-problems-from-easy-to-hard.md): 15 fun coding problems from easy to hard, each with the core... The rest of the corpus is not enumerated here. Crawl \[sitemap-blog.xml\](https://daily.dev/sitemap-blog.xml) for every post.
