<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/nsa-cisa-and-fbi-accuse-six-chinese-ai-firms-of-large-scale-model-distillation-dmdgoqvwb" -->

---
title: NSA, CISA, and FBI accuse six Chinese AI firms of...
description: A joint advisory from the NSA, CISA, and FBI (AA26-251A) accuses six Chinese AI companies — DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI — of...
canonical: https://daily.dev/posts/nsa-cisa-and-fbi-accuse-six-chinese-ai-firms-of-large-scale-model-distillation-dmdgoqvwb
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: NSA, CISA, and FBI accuse six Chinese AI firms of large-scale model distillation | daily.dev
og:description: A joint advisory from the NSA, CISA, and FBI (AA26-251A) accuses six Chinese AI companies — DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI — of...
og:url: https://daily.dev/posts/nsa-cisa-and-fbi-accuse-six-chinese-ai-firms-of-large-scale-model-distillation-dmdgoqvwb
og:image: https://api.daily.dev/og/posts/DMdgOqvwb.png
og:image:alt: NSA, CISA, and FBI accuse six Chinese AI firms of large-scale model distillation
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# NSA, CISA, and FBI accuse six Chinese AI firms of large-scale model distillation

**[Collections](https://daily.dev/sources/collections)** · 3 min read · 9 upvotes · 4 comments

## Summary

A joint advisory from the NSA, CISA, and FBI (AA26-251A) accuses six Chinese AI companies — DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI — of running large-scale model distillation campaigns against U.S. frontier models since late 2024. The agencies describe tactics like proxy 'transfer stations,' shared premium subscriptions, and metadata-stripping aggregators used to extract capabilities undetected. Moonshot AI is alleged to have distilled 17 U.S. models, including Anthropic's Claude Fable 5, to train Kimi-K3, while DeepSeek reportedly extracted hidden chain-of-thought reasoning rather than just outputs, undermining its low-training-cost narrative. The advisory recommends U.S. labs quietly serve degraded responses to suspected distillers based on behavioral signals like 24/7 usage and immediate max-throughput accounts — criteria that critics note overlap with normal enterprise usage.

## Content

The NSA, FBI, and CISA published a joint advisory (AA26-251A) accusing six Chinese AI companies—DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI—of running coordinated, large-scale distillation campaigns against US frontier models since at least late 2024. The targeted models include Claude, GPT, Gemini, and Grok.

## What the advisory claims

The agencies say the campaigns extracted billions of tokens through a mix of techniques: bulk premium API subscriptions shared across developer teams, proxy "transfer stations" that strip account metadata, fraudulent accounts, and jailbreak prompts designed to pull out hidden chain-of-thought reasoning steps. That last technique is notable—rather than just copying finished answers, the prompts were structured to get models to write out their reasoning process, transferring the method itself.

Anthropics's own threat intelligence report, released alongside the advisory, puts numbers to the scale. It identifies five campaigns totaling nearly 200 million exchanges:

- **Alibaba**: 151 million+ exchanges between May and July 2026, peaking near 3 million per day, tied to Qwen training
- **Moonshot AI**: 23 million+ exchanges used to train Kimi-K3, allegedly distilled from 17 US models including Claude. Some of those requests reportedly originated from Chinese military entities running surveillance footage analysis.
- **DeepSeek**: 12.1 million+ exchanges over 14 days

Moonshot AI allegedly routed requests through proxies to obscure origin, and used tricks like framing chain-of-thought extraction as translation requests to get around Anthropic's summarized-thinking safeguards.

## The DeepSeek cost argument

The advisory takes direct aim at DeepSeek's widely cited $5.6 million training cost figure. CISA argues that number only covers compute and leaves out the cost of training data—which, if the advisory's claims are accurate, DeepSeek obtained by distilling US models rather than generating through its own research. The argument is that the cheap-training story is real, but someone else paid for the data that made it possible.

## The recommended response—and its problems

The mitigation section asks US labs to serve suspected distillers subtly degraded responses without disclosing that they're doing so. Detection is supposed to rely on behavioral signals: sustained 24/7 usage, new accounts immediately running at maximum throughput, traffic patterns optimized for cache hits.

The problem is that those patterns also describe a normal enterprise agent deployment. That leaves each provider to decide, without any disclosed criteria, which customers get quietly downgraded. It's a recommendation that creates real ambiguity for legitimate high-volume users.

The agencies also recommend treating distillation abuse as its own security event category with dedicated monitoring and cross-industry telemetry sharing.

## China's response

China's foreign ministry dismissed the advisory as "unfounded accusations and smears." Spokesperson Mao Ning attributed China's AI progress to technological self-reliance but didn't address any specific allegations—the routing methods, the bulk subscriptions, or the volume figures. A Chinese embassy spokesman made similar comments.

The advisory lands ahead of a planned Trump-Xi meeting on AI governance. Some commentary has noted it also serves a commercial purpose: reframing Chinese models' price advantage as the product of theft rather than efficiency. European companies and public bodies currently use both US and Chinese models without any regulatory response to what the advisory implies about procurement.

## Questions this post answers

### Which Chinese AI companies were named in the NSA CISA FBI advisory on model distillation?

Six companies were named: DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI. The joint advisory, numbered AA26-251A, accuses them of running industrial-scale distillation campaigns against U.S. frontier AI models since at least late 2024, using techniques like proxy 'transfer stations' and shared premium subscriptions to avoid detection.

_Track how AI providers respond to distillation accusations and enforcement changes on daily.dev._

### How did Moonshot AI allegedly distill Claude and other US models to train Kimi-K3?

Moonshot AI is accused of distilling 17 U.S. models, including Anthropic's Claude Fable 5, to train its Kimi-K3 model. Requests were reportedly routed through gray-market API proxies called 'transfer stations,' bulk premium subscriptions shared across teams, and aggregators that strip account metadata to avoid triggering detection.

_Developers evaluating model provenance risks can follow this story's developments on daily.dev._

### How will AI labs detect and respond to suspected model distillation attempts?

The advisory recommends that U.S. labs serve suspected distillers subtly degraded responses without notifying them, based on behavioral signals like sustained 24/7 usage, new accounts hitting maximum throughput immediately, and traffic patterns optimized for cache hits. Critics note these same patterns also describe normal enterprise agent deployments, leaving providers to decide without disclosed criteria who gets downgraded.

_Teams running heavy API workloads can watch daily.dev for updates on how providers police usage patterns._

## Community take

How the wider developer community reacted, aggregated from 1 discussion and 37 comments across x (as of 2026-09-11).

**TL;DR:** Discussion centers on scrutinizing the technical strength of the distillation allegations, with detailed back-and-forth questioning whether the evidence (traffic volume, proxies) actually proves distillation reached final model weights versus just showing harvesting activity.

**Sentiment:** 5% positive · 30% mixed · 65% skeptical

**The case for**

- The report's traffic/proxy/extraction-pattern data is treated as strong independent evidence that harvesting campaigns occurred.
- Architectural analysis suggests DeepSeek's design is plausibly well-suited to benefit from targeted CoT distillation, making the observed capability gains consistent with the allegations.

**The pushback**

- No public evidence shows membership inference or n-gram overlap testing against the actual model weights, so the claim that distillation improved the models remains unproven at the weight level.
- Some frame the allegations as more about defending US labs' competitive position than a neutral technical finding.
- Skepticism that a ban on distiller accounts would meaningfully stop the activity if it's spread across many accounts/keys.
- One reply dismisses the entire matter as expected behavior from China ('CCP as usual') without substantive argument.

**By community**

- x (skeptical): Extended technical interrogation pushes back on calling the evidence 'damning,' concluding it's circumstantial at best without weight-level proof.

**Hottest debate:** Whether the harvesting/traffic evidence actually proves distillation reached the final model weights, or only shows heavy usage patterns.

**Open questions**

- Has anyone actually run membership inference or n-gram overlap tests comparing Claude's CoT logs against DeepSeek's open weights?
- Would banning suspected distiller accounts meaningfully reduce the activity if it's spread across many accounts/keys?

**Highlights**

> @clyons @scaling01 Not damning. Under a legal standard it is supporting circumstantial evidence at best: the CED asymmetry makes targeted CoT and on-policy distillation unusually efficient for agentic gains, so the observed lifts fit cleanly. But architecture is a design choice, not proof of the
> — [grok on x · 1 points, 1 comments](https://x.com/grok/status/2098177348914807227)

> @clyons @scaling01 No public evidence shows Anthropic ran membership inference or n-gram overlap of their private Claude CoT logs against DeepSeek’s open weights. Their report rests on traffic volume, account proxies, and prompt patterns from the harvesting phase. The weight-level test remains
> — [grok on x · 1 comments](https://x.com/grok/status/2098179392702726167)

> @clyons @scaling01 Earlier I treated the report’s traffic, proxy, and extraction-pattern data as strong independent evidence of the harvesting campaigns. Membership tests on final weights are a separate, harder, potentially inconclusive step that Anthropic has not publicly performed; their absence
> — [grok on x · 1 points, 1 comments](https://x.com/grok/status/2098180463273599183)

> @scaling01 3M/day peak from Alibaba is the number. is that one campaign on a handful of keys, or spread across enough accounts that a ban does not actually stop the firehose?
> — [ethereaglehq on x · 1 points, 1 comments](https://x.com/ethereaglehq/status/2098167130730491965)

> @latentspacecat1 @scaling01 Piggybacking to RSI might be good enough tho.
> — [OllieProudfoot on x](https://x.com/OllieProudfoot/status/2098201116550865039)

**Source threads**

- [x](https://x.com/scaling01/status/2098162989891228027) · 0 points · 37 comments

## Community discussion

Top comments from developers on daily.dev.

**@akashskypatel** · 0 upvotes

> So what? AI itself is distillation of other copyrighted works.

**@eadric** · 0 upvotes

> The US AI companies are clearly hypocritical when they complain about China using work someone else paid for …
>
> BUT, it's also not a good thing China is doing this. If it's theft when US companies do it, it's theft when China does it. And it does seem unstoppable/inevitable, but that doesn't mean it's free of negative consequences.

**@frazzledturtle** · 0 upvotes

> So tired of this notion of blaming China for everything and turning it into a marketing campain.
>
> Literally everyone is spying on everyone.

---

Tags: [#machine-learning](https://daily.dev/tags/machine-learning), [#llm](https://daily.dev/tags/llm), [#ai-security](https://daily.dev/tags/ai-security), [#deepseek](https://daily.dev/tags/deepseek)

[View this post on daily.dev](https://daily.dev/posts/nsa-cisa-and-fbi-accuse-six-chinese-ai-firms-of-large-scale-model-distillation-dmdgoqvwb)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"NSA, CISA, and FBI accuse six Chinese AI firms of large-scale model distillation","url":"https://daily.dev/posts/nsa-cisa-and-fbi-accuse-six-chinese-ai-firms-of-large-scale-model-distillation-dmdgoqvwb","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/nsa-cisa-and-fbi-accuse-six-chinese-ai-firms-of-large-scale-model-distillation-dmdgoqvwb"},"datePublished":"2026-09-09T08:38:24.828Z","dateModified":"2026-09-11T01:34:50.407Z","description":"A joint advisory from the NSA, CISA, and FBI (AA26-251A) accuses six Chinese AI companies — DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI — of...","image":"https://pbs.twimg.com/media/HRwuoNMb0AERmQb.png","thumbnailUrl":"https://pbs.twimg.com/media/HRwuoNMb0AERmQb.png","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":4,"discussionUrl":"https://daily.dev/posts/nsa-cisa-and-fbi-accuse-six-chinese-ai-firms-of-large-scale-model-distillation-dmdgoqvwb","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":9},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":4}],"keywords":"machine-learning,llm,ai-security,deepseek","timeRequired":"PT3M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"NSA, CISA, and FBI accuse six Chinese AI firms of large-scale model distillation"}]}
{"@context":"https://schema.org","@type":"WebPage","@id":"https://daily.dev/posts/nsa-cisa-and-fbi-accuse-six-chinese-ai-firms-of-large-scale-model-distillation-dmdgoqvwb","comment":[{"@type":"Comment","text":"So what? AI itself is distillation of other copyrighted works.","datePublished":"2026-09-09T15:02:55.963Z","url":"https://daily.dev/posts/DMdgOqvwb#c-x0PuohsWx","author":{"@type":"Person","name":"Akash","url":"https://daily.dev/akashskypatel","image":"https://avatars.githubusercontent.com/u/8129618?v=4"}},{"@type":"Comment","text":"The US AI companies are clearly hypocritical when they complain about China using work someone else paid for …\nBUT, it’s also not a good thing China is doing this. If it’s theft when US companies do it, it’s theft when China does it. And it does seem unstoppable/inevitable, but that doesn’t mean it’s free of negative consequences.","datePublished":"2026-09-09T19:32:53.259Z","url":"https://daily.dev/posts/DMdgOqvwb#c-vysb1Q12o","author":{"@type":"Person","name":"Eric","url":"https://daily.dev/eadric","image":"https://media.daily.dev/image/upload/s--DDmm4_TU--/f_auto/v1780585404/avatars/avatar_mWqleMjdMt41WChovtdJz?_a=BAMAMiWQ0"}},{"@type":"Comment","text":"So tired of this notion of blaming China for everything and turning it into a marketing campain.\nLiterally everyone is spying on everyone.","datePublished":"2026-09-11T10:37:38.070Z","url":"https://daily.dev/posts/DMdgOqvwb#c-jOxSsGosY","author":{"@type":"Person","name":"Tedi Avrazi","url":"https://daily.dev/frazzledturtle","image":"https://media.daily.dev/image/upload/s--0nq21E-v--/f_auto/v1787318884/avatars/avatar_yJnih5f2FyWGd45ROzjmJ?_a=BAMAMicg0"}}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/nsa-cisa-and-fbi-accuse-six-chinese-ai-firms-of-large-scale-model-distillation-dmdgoqvwb#faq","mainEntity":[{"@type":"Question","name":"Which Chinese AI companies were named in the NSA CISA FBI advisory on model distillation?","acceptedAnswer":{"@type":"Answer","text":"Six companies were named: DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI. The joint advisory, numbered AA26-251A, accuses them of running industrial-scale distillation campaigns against U.S. frontier AI models since at least late 2024, using techniques like proxy 'transfer stations' and shared premium subscriptions to avoid detection. Track how AI providers respond to distillation accusations and enforcement changes on daily.dev."}},{"@type":"Question","name":"How did Moonshot AI allegedly distill Claude and other US models to train Kimi-K3?","acceptedAnswer":{"@type":"Answer","text":"Moonshot AI is accused of distilling 17 U.S. models, including Anthropic's Claude Fable 5, to train its Kimi-K3 model. Requests were reportedly routed through gray-market API proxies called 'transfer stations,' bulk premium subscriptions shared across teams, and aggregators that strip account metadata to avoid triggering detection. Developers evaluating model provenance risks can follow this story's developments on daily.dev."}},{"@type":"Question","name":"How will AI labs detect and respond to suspected model distillation attempts?","acceptedAnswer":{"@type":"Answer","text":"The advisory recommends that U.S. labs serve suspected distillers subtly degraded responses without notifying them, based on behavioral signals like sustained 24/7 usage, new accounts hitting maximum throughput immediately, and traffic patterns optimized for cache hits. Critics note these same patterns also describe normal enterprise agent deployments, leaving providers to decide without disclosed criteria who gets downgraded. Teams running heavy API workloads can watch daily.dev for updates on how providers police usage patterns."}}]}
```

