<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/import-ai-446-nuclear-llms-china-s-big-ai-benchmark-measurement-and-ai-policy-sw9hjcqy6" -->

---
title: Import AI 446: Nuclear LLMs; China’s big AI benchmark;...
description: Issue 446 of Import AI covers four main topics: (1) Jacob Steinhardt&#x27;s argument that investing in AI measurement tools is a key policy lever, enabling...
canonical: https://daily.dev/posts/import-ai-446-nuclear-llms-china-s-big-ai-benchmark-measurement-and-ai-policy-sw9hjcqy6
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Import AI 446: Nuclear LLMs; China’s big AI benchmark; measurement and AI policy | daily.dev
og:description: Issue 446 of Import AI covers four main topics: (1) Jacob Steinhardt&#x27;s argument that investing in AI measurement tools is a key policy lever, enabling...
og:url: https://daily.dev/posts/import-ai-446-nuclear-llms-china-s-big-ai-benchmark-measurement-and-ai-policy-sw9hjcqy6
og:image: https://api.daily.dev/og/posts/SW9hjCQy6.png
og:image:alt: Import AI 446: Nuclear LLMs; China’s big AI benchmark; measurement and AI policy
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Import AI 446: Nuclear LLMs; China’s big AI benchmark; measurement and AI policy

**[Import AI ](https://daily.dev/sources/jackclark)** · 13 min read · 1 upvotes · 0 comments

## Summary

Issue 446 of Import AI covers four main topics: (1) Jacob Steinhardt's argument that investing in AI measurement tools is a key policy lever, enabling governance by making AI properties visible and auditable; (2) a King's College London study showing GPT-5.2, Claude Sonnet 4, and Gemini 3 Flash all escalate to nuclear use far more readily than humans in wargame simulations, with Claude winning most games as a 'calculating hawk'; (3) China's ForesightSafety Bench, a comprehensive AI safety evaluation framework covering 94 risk subcategories including existential risks, where Anthropic's Claude series leads the leaderboard; and (4) LABBench2, a 1,900-task biology research benchmark revealing that frontier AI models have uneven scientific capabilities, struggling with cross-database retrieval and figure interpretation while performing well on patent search tasks.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://jack-clark.net/2026/02/23/import-ai-446-nuclear-llms-chinas-big-ai-benchmark-measurement-and-ai-policy/>

## Similar posts on daily.dev

- [Import AI 450: China’s electronic warfare model; traumatized LLMs; and a scaling law for cyberattacks](https://daily.dev/posts/import-ai-450-china-s-electronic-warfare-model-traumatized-llms-and-a-scaling-law-for-cyberattack-noztaxwwo) · Import AI  · 1 upvotes · 0 comments
- [Import AI 454: Automating alignment research; safety study of a Chinese model; HiFloat4](https://daily.dev/posts/import-ai-454-automating-alignment-research-safety-study-of-a-chinese-model-hifloat4-msxpvronf) · Import AI  · 0 upvotes · 0 comments
- [Import AI 444: LLM societies; Huawei makes kernels with AI; ChipBench](https://daily.dev/posts/import-ai-444-llm-societies-huawei-makes-kernels-with-ai-chipbench-ekctso6st) · Import AI  · 1 upvotes · 0 comments

---

Tags: [#ai](https://daily.dev/tags/ai), [#ai-governance](https://daily.dev/tags/ai-governance), [#ai-safety](https://daily.dev/tags/ai-safety), [#llm](https://daily.dev/tags/llm)

[View this post on daily.dev](https://daily.dev/posts/import-ai-446-nuclear-llms-china-s-big-ai-benchmark-measurement-and-ai-policy-sw9hjcqy6)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Import AI 446: Nuclear LLMs; China’s big AI benchmark; measurement and AI policy","url":"https://daily.dev/posts/import-ai-446-nuclear-llms-china-s-big-ai-benchmark-measurement-and-ai-policy-sw9hjcqy6","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/import-ai-446-nuclear-llms-china-s-big-ai-benchmark-measurement-and-ai-policy-sw9hjcqy6"},"datePublished":"2026-02-23T13:58:46.923Z","dateModified":"2026-03-30T02:15:08.458Z","description":"Issue 446 of Import AI covers four main topics: (1) Jacob Steinhardt's argument that investing in AI measurement tools is a key policy lever, enabling...","image":"https://i0.wp.com/jack-clark.net/wp-content/uploads/2026/02/https3A2F2Fsubstack-post-media.s3.amazonaws.com2Fpublic2Fimages2Fd6d17996-2bef-40a4-abe3-be72a0e8a227_258x258-d3Iw9p.jpg?fit=258%2C258&ssl=1","thumbnailUrl":"https://i0.wp.com/jack-clark.net/wp-content/uploads/2026/02/https3A2F2Fsubstack-post-media.s3.amazonaws.com2Fpublic2Fimages2Fd6d17996-2bef-40a4-abe3-be72a0e8a227_258x258-d3Iw9p.jpg?fit=258%2C258&ssl=1","isAccessibleForFree":true,"articleSection":"Import AI ","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Import AI ","logo":"https://media.daily.dev/image/upload/s--ygb8HHcb--/f_auto/v1735476187/logos/jackclark","url":"https://daily.dev/sources/jackclark"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/import-ai-446-nuclear-llms-china-s-big-ai-benchmark-measurement-and-ai-policy-sw9hjcqy6","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"ai,ai-governance,ai-safety,llm","timeRequired":"PT13M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Import AI ","item":"https://daily.dev/sources/jackclark"},{"@type":"ListItem","position":3,"name":"Import AI 446: Nuclear LLMs; China’s big AI benchmark; measurement and AI policy"}]}
```

