This newsletter roundup covers four AI research items: DiG-bench, a new 70-game benchmark testing whether AI systems can infer hidden rules through exploration, where Opus 5 and Fable 5 led but still struggled at the hardest tiers compared to humans; an RSI (recursive self-improvement) simulator game from Paradigm Research that lets players experience the tradeoffs of running an AI lab; a paper from startup Inherent describing Faraday, a 27B post-trained model that supervises frontier models like Codex to replicate missing results from published ML and AI-for-science papers, beating base Opus and GPT-5.5 on a majority of tasks; and a critical read of Mark Zuckerberg's essay 'The Future is for Everyone,' arguing it fails to address whether superintelligent systems capable of invention would actually serve individual empowerment as he assumes.

12m read timeFrom jack-clark.net
Post cover image
Table of contents
Share this:Like this:Related

Questions this post answers

What is DiG-bench and how do frontier AI models perform on it?

DiG-bench is a benchmark of 70 text-based games split into seven difficulty tiers that require players to infer hidden rules and objectives through exploration rather than being told them. Opus 5 and Fable 5 (paired with Claude Code) performed best overall, with only those two models beating any Tier 7 tasks, achieving just a 20% success rate versus 100% for humans. daily.dev surfaces benchmark results like these for developers tracking what frontier models can actually do.

What is Faraday and how does it improve AI scientist performance on research replication tasks?

Faraday is a 27B parameter model built by startup Inherent, post-trained on Qwen-3.6-27B using GRPO, that supervises a coding agent (OpenAI Codex) to replicate missing results from published papers. Using a dataset called Replica of 310 replication tasks derived from 100 ML and AI-for-science papers, Faraday with Codex outperformed base Opus 4.8 and GPT-5.5 on 73% of in-distribution tasks and 60% of held-out tasks. developers evaluating AI research agents can track model-versus-model results like this on daily.dev.

1 Impression