中國的前沿AI模型究竟用什麼硬體訓練?

發佈: 2026-07-28

Kimi K3是什麼,為何對硬體敘事如此重要?

On July 16, 2026, Beijing-based Moonshot AI released Kimi K3, a 2.8-trillion-parameter mixture-of-experts model that it describes as the world's first "open 3T-class" model — roughly 75% larger than DeepSeek's V4 Pro, the previous largest widely-used open-weight model. K3 ships with native multimodal input and a 1-million-token context window, and its weights were promised to be fully public by July 27, 2026, just over a week after the API launch. On benchmarks, K3 trailed Anthropic's Claude Fable 5 and OpenAI's GPT-5.6 Sol overall but beat Claude Opus 4.8 and GPT-5.5 on coding and general-agent tasks.

The hardware story starts with the model's sheer scale: no single Nvidia H100, H200, or B200 can hold the full 2.8-trillion-parameter model in memory. Moonshot recommends serving K3 on "supernodes" of 64 or more accelerators, keeping expert-parallel routing traffic inside a single high-bandwidth domain, with raw 4-bit quantized weights alone totaling roughly 1.4 terabytes. That scale is precisely why K3's release reopened a question compliance officers, competitors, and policymakers had been asking quietly for years: what compute did it actually take to train a model this large, under a compute-export regime explicitly designed to prevent exactly this?

月之暗面究竟用了什麼硬體?為何引發爭議?

Moonshot's own published kernel-optimization benchmarks reference Nvidia H200 GPUs alongside an unnamed "alternative vendor" accelerator — already a signal that the company's compute stack is not a simple, single-source story. The more consequential claim came five days after K3's release: on July 22, 2026, White House Office of Science and Technology Policy Director Michael Kratsios stated publicly that Moonshot AI "has also acquired GB300-equipped servers and has accessed GB300s in Thailand," likely to train its models, and separately alleged the company built "a sophisticated internal platform to conduct large scale distillation against U.S. models" — reportedly targeting Anthropic's models — engineered to "quickly switch between multiple methods of access to avoid detection."

It is important to be precise about what this is and isn't. This is a public political allegation from a senior US official, not a court finding, an Entity List action, or an admission by Moonshot. Treasury Secretary Scott Bessent said sanctions authority remains "on the table," but as of this writing no formal US enforcement action against Moonshot AI has been announced. Readers evaluating the story should hold the distinction between "alleged" and "adjudicated" clearly in mind — a distinction that matters for compliance teams as much as for casual observers, since headline allegations and enacted controls carry very different practical implications.

泰國轉運路徑為何重要——它暴露了怎樣的執法漏洞?

The Thailand allegation lands against a specific regulatory backdrop: a May 31, 2026 BIS memo affirmed that a headquarters-location licensing test applies to Nvidia Blackwell-class chips — including the GB300 — meaning that even servers nominally owned or operated through non-Chinese entities can trigger a license requirement if their ultimate beneficial control traces back to a Chinese headquarters. That memo was explicitly designed to close a shell-subsidiary loophole. AIChipMap has added a new regime entry, "US BIS Blackwell Southeast Asia Transshipment Enforcement (May 2026)," modeling this at the US→China level, reflecting that no formal entity-level action against Moonshot specifically has been taken as of this writing.

The underlying enforcement problem is structural, not unique to Moonshot: Southeast Asian jurisdictions — Thailand, Singapore, Malaysia — have historically had less-developed export-control enforcement infrastructure than the US, Netherlands, or Japan, making them attractive routing points for controlled chips destined for China. A December 2025 US enforcement action (nicknamed "Operation Gatekeeper" in reporting) against a roughly $160 million chip-smuggling network through the same region is the clearest precedent that this is a recognized, recurring pattern rather than an isolated incident. For compliance teams, the practical takeaway is that geography alone — a shipment routed through a non-China jurisdiction — no longer satisfies scrutiny; beneficial ownership and end-use tracing increasingly do the real work.

除了英偉達變通方案——中國國產算力究竟有哪些?

Whatever the truth of any individual transshipment allegation, China has spent the years since the 2022–2023 BIS advanced-computing rules building a genuine, if uneven, domestic compute stack. Huawei's Ascend 910B/910C remains the primary domestic training and inference alternative to Nvidia, fabricated by SMIC using DUV multi-patterning to approach 7nm-class density without EUV lithography access. Below Huawei sits a widening tier of fabless GPU/accelerator startups — Cambricon and Biren (both under US Entity List restrictions since 2020 and 2022 respectively), and a newer cohort of Moore Threads, MetaX, and Enflame — all fabbed at SMIC on mature nodes rather than TSMC's leading edge, and all facing the same structural constraint: Nvidia's roughly two-decade CUDA software head start, which no Chinese accelerator vendor has closed regardless of how competitive their raw specifications look on paper.

That design-and-fab layer rests on a broader domestic equipment and tooling base that is also expanding. NAURA and AMEC — China's two largest domestic toolmakers, forming a "Big Three" with ACM Research — now supply etch and deposition tools to SMIC, YMTC, CXMT, and Hua Hong, with NAURA climbing to roughly fifth place globally in equipment sales by 2025–2026. Lithography remains the hardest layer to replace: SMEE, China's state-owned lithography champion, has spun off a Huawei-linked subsidiary, Shanghai Yuliangsheng, now testing what's reported as China's first domestic 28nm-class immersion DUV scanner at SMIC, with mass production targeted "as early as 2027" and sub-10nm capability not expected before 2030. On the software side, Empyrean Technology — China's largest domestic EDA vendor, though still holding only about 6% of the total Chinese EDA market against Cadence and Synopsys — has grown specifically because the 2022 US EDA export restrictions created policy-driven demand for a local alternative.

DeepSeek的效率路線如何改變算力取得的計算邏輯?

DeepSeek's January 2025 "moment" — when its V3 and R1 models matched frontier US-lab performance at a reported fraction of the training cost — demonstrated that software efficiency can partially substitute for restricted hardware access, and that lesson has propagated across the Chinese AI industry since. Multi-head latent attention compresses the KV cache to cut memory-bandwidth pressure during inference; sparse mixture-of-experts routing activates only a small fraction of total parameters per token; and FP8 mixed-precision training reduces memory footprint and improves throughput on export-compliant, interconnect-limited chips like the H800.

The practical reframing this forces is important: the right question is no longer simply "does China have enough advanced chips," but "how much model capability can be extracted per unit of available compute, and how fast is that ratio improving." Kimi K3's own architecture — a 2.8-trillion-parameter MoE model designed from the outset around distributed expert-parallel serving across accelerator "supernodes" rather than assuming abundant single-chip capability — reflects the same efficiency-first design logic DeepSeek popularized, regardless of how its specific training hardware is ultimately characterized.

2026年中國AI資本市場熱潮釋放了怎樣的訊號?

2026 has been a watershed year for Chinese AI-hardware capital markets. Zhipu AI — added to the US Entity List in January 2025, which forced a pivot toward training on domestic Ascend hardware — listed on the Hong Kong Stock Exchange in January 2026 as "Knowledge Atlas Technology," becoming the first Chinese frontier AI foundation-model company to go public; its market cap reached roughly HK$1 trillion (about $128 billion) by June 2026 after its GLM-5.2 release, up roughly 2,400% from its IPO price. Moore Threads followed with a Shanghai STAR Market IPO that was oversubscribed more than 4,000-fold by retail investors, and MetaX posted a roughly 700% single-day debut pop on its own STAR Market listing months earlier — drawing even more retail interest than Moore Threads. Enflame, backed heavily by Tencent, cleared STAR Market IPO registration in July 2026.

Whatever the resolution of any individual export-control dispute, this wave of oversubscribed IPOs signals sustained Chinese capital-market and state appetite for domestic AI-hardware self-sufficiency that is unlikely to reverse regardless of how enforcement actions against any single company play out.

合規、採購和投資者讀者接下來應關注什麼?

Several concrete threads are worth tracking over the coming months: whether Moonshot AI faces a formal US enforcement action (Entity List addition or specific sanction) rather than remaining at the level of public allegation; whether Kimi K3's weights actually ship in full on the promised July 27, 2026 timeline; whether BIS extends or tightens headquarters-location licensing enforcement to other Southeast Asian jurisdictions beyond the current Thailand-specific allegation; and whether Yuliangsheng's domestic 28nm immersion DUV tool reaches SMIC mass production on its roughly 2027 target.

The broader frame worth holding onto is bilateral, not one-sided: the United States restricts advanced chips, EDA tools, and now headquarters-traced server access to Chinese entities, while China restricts its own exports of gallium, germanium, graphite, and antimony — materials that feed directly back into the same semiconductor supply chain the US is trying to protect. Neither side's controls operate in isolation, and reading China's frontier-AI hardware story — Kimi K3 included — only in terms of what it can't get, rather than also what it is building domestically and what leverage it holds in return, misses half the picture.