Happy Saturday.
I scan 100+ Chinese-language sources daily, translate what matters, and send it here. You're reading this because the information asymmetry between Chinese-language AI coverage and English-language coverage is enormous, and someone should fix that.
Let's go.
The Last Dependency
For two years, the US export control theory rested on a clear assumption: you can sanction a country out of training frontier AI models. Inference on domestic chips? Possible. Fine-tuning? Doable. But full-parameter post-training of a trillion-parameter model requires so much coordinated compute, so much memory bandwidth, and so much stable distributed software, that domestic hardware simply cannot do it at competitive quality.
That assumption took a direct hit this week.
SLAI, a joint team from Shenzhen Hetao College, Harbin Institute of Technology's Shenzhen campus, and Huawei's GTS division, completed what it describes as the first third-party, full-parameter post-training of DeepSeek-V4-Pro on a domestic compute cluster. The model in question is a 1.6 trillion-parameter mixture-of-experts architecture, the same model family that currently tops most open-source benchmarks. The hardware is Ascend 910C GPUs. A cluster of roughly 1,000 cards. Full-parameter means all 1.6 trillion parameters are updated simultaneously during training, not a small adapter layer on top.
The result: 1,500+ training steps, zero NaN failures, zero instability events. MFU (model flops utilization, the measure of how efficiently a cluster actually uses its theoretical compute) reached 34.9% on Ascend SuperNode. The team trained a math-reasoning dataset of 3,000 high-quality SFT samples and achieved statistically significant benchmark improvements over the base model. The engineering team at Shenzhen Hetao also trained 42 students through the process as a pedagogical exercise, which tells you something about how they see the significance.
The mechanism behind the difficulty is worth explaining. DeepSeek-V4-Pro uses a hybrid sparse attention design and a mixture-of-experts routing layer where only 11 billion of the 1.6 trillion parameters activate per token. MoE is powerful and efficient for inference. For training, the expert-routing structure generates massive cross-node communication patterns, because different tokens activate different experts, and those experts are distributed across different physical GPUs. SLAI's team built a four-way parallel scheme (data, tensor, pipeline, and expert parallelism simultaneously), a real-time load-balancing system that monitors which experts are hot, and a fault detection loop that caught and recovered from hardware transients automatically.
In isolation, this is one university-affiliated team's engineering result. In context, it is the hardware proof of something this week's coverage has been circling. ForgeTrain showed that AI-written software frameworks could outperform Megatron on Ascend by 10%. SLAI shows that result generalizing to the most demanding workload in AI. The same week, ByteDance announced it is designing a custom CPU to replace Intel and AMD in its AI inference infrastructure, as CPU prices rose 10 to 35% per quarter and supply constraints tightened. Three separate teams, three separate dependency chains, all completing domestically in the same week.
MFU of 34.9% is not Nvidia's 50% to 55%. The gap is real and represents roughly a 50% cost premium on compute for equivalent output. But 50% cost premium is a tax, not a ceiling. The training can happen. The models will be trained. The policy lever that assumed "slower means impossible" just got a lot weaker.
The Briefing
DeepSeek is taking outside money for the first time in its three-year history, at a pre-money valuation of $45 billion. The National IC Fund (国家大基金) is leading the round, with several market-rate investors also in talks. One month ago the valuation floated in discussions was $20 billion. The rapid re-rating tells you how the fundraising conversation has gone. For a company that built R1, V3, and V4-Pro without a dollar of outside capital, this is a structural shift. National IC Fund investing means the state now has a direct equity stake in China's best-performing AI lab, not just a policy interest. The round, if completed, will be the largest first-round raise in Chinese AI history.
YMTC (长江存储), China's NAND flash champion, has started A-share IPO preparations at a reported valuation of 160 billion yuan. According to a detailed Huxiu industry analysis, YMTC now holds 13% of global NAND market share by shipment volume, tied with Micron for fourth place globally. Its Xtacking architecture separates the storage array from the control logic and bonds them via wafer-level fusion. Samsung has paid YMTC for patent licenses on that wafer bonding technology. YMTC's new production lines use 50% or more domestic equipment. The third fab in Wuhan starts volume production in the second half of 2026. The more consequential detail is timing: three of the four global HBM manufacturers just co-invested in Anthropic at a $965 billion valuation, signaling supply chain equity is the new battleground. YMTC is positioning China's NAND supply chain for the same kind of structural alignment.
China's humanoid robots beat the human half-marathon world record in this year's competition, finishing 21 kilometers in 50 minutes and 26 seconds. The technology behind the result is not visual AI. According to the 量子位 Bilibili analysis with 960,000 views, the winning robots use RTK positioning antennas for centimeter-level GPS accuracy, paired with LiDAR point cloud for obstacle detection, and reinforcement-learning locomotion controllers trained in simulation. No camera-based object recognition. The champion team, 荣耀闪电, used custom 400Nm hip motors and full water cooling borrowed from smartphone thermal design to manage heat across 21 kilometers. When asked what drove the improvement, a Guodi Robotics engineer gave the answer that keeps appearing in Chinese AI engineering contexts: "We just gave it enough time. The method isn't fundamentally different. We gave it more data, more training time. It improved more."
Unitree (宇树科技) is pushing toward an A-share listing as China's first public humanoid robotics company. A 36Kr analysis describes a company with strong hardware economics and profitability on its consumer products, with the listing committee review scheduled for June 1. The acknowledged weakness is software autonomy and long-horizon task execution. The investment thesis the IPO is packaging: the hardware problems are solved, the software will catch up, own the hardware company before that happens.
What I Found on Bilibili This Week
The 量子位 video on the humanoid robot marathon deserves more than a briefing item.
The video has 960,000 views and includes a full transcript. The question it actually answers is why performance improved so dramatically in a single year, and the answer differs from what most English coverage assumes.
The winning robots are not smarter than last year's. They are better engineered. The 荣耀闪电 team used water cooling systems adapted from smartphone thermal management, explaining that their phone industry background gave them mature liquid cooling designs. Their 400Nm hip motors are roughly twice the standard size. The team described the design trade-off simply: bigger motors, more heat, water cooling solves the heat, bigger motors deliver more sustained power across 21 kilometers.
The locomotion controller is reinforcement learning in simulation, transferred to hardware. The training method has not changed in fundamentals since last year. What has changed is the number of training iterations, the quality of the simulation environment, and the 21 kilometers of real-world running data from year one. The engineer's framing: give reinforcement learning enough iterations and enough data, and the performance follows.
This is the most-watched robotics video on Bilibili this week. The Chinese AI engineering community is watching this specific transition, and the narrative they are taking from it is consistent: sustained iteration, not architectural novelty, is what closes capability gaps.
Signals
Three HBM competitors co-invested in the same AI company for the first time. Micron, Samsung, and SK Hynix all joined Anthropic's $6.5 billion H-round as strategic infrastructure partners. These three companies have never appeared on the same AI investment cap table before. Combined they control essentially all of global HBM supply through 2028. Joining Anthropic's equity structure is a demand-signal purchase. As shareholders, they gain early visibility into what compute architecture Anthropic is building toward, which directly shapes the next generation of HBM specifications they need to design. Anthropic's $965 billion post-money valuation is almost secondary to what this consortium structure signals: AI infrastructure capital allocation is consolidating vertically, and the tier above GPUs is supply-chain equity.
Unitree G1 performed at Wang Leehom's Hangzhou concert, singing "龙的传人" while leading a choreographed dance routine. The clip is on Baidu's trending board with 5.3 million heat units. For a company heading into an IPO listing committee review on June 1, performing at a nationally iconic anthem concert is a specific kind of brand positioning. Unitree is not just a robotics company filing paperwork. It is a cultural object at a particular moment in Chinese national pride in technology.
Domestic AI chip market share in China crossed 41%. A Bilibili video from 龙科多工作室 with nearly 50,000 views traces the shift: Nvidia held roughly 95% of the Chinese AI chip market three years ago and holds around 55% today. The 41% domestic share is concentrated in inference. The SLAI training result above is a data point on whether the training share will follow. Nvidia's addressable market inside China is declining not through expulsion but through domestic options crossing engineering sufficiency thresholds one workload at a time.
The Bigger Picture
The theory behind US chip export controls was never just about stopping China from buying Nvidia. The theory was that compute is the choke point. Without frontier training hardware, you cannot build frontier models. Without frontier models, the strategic AI advantage does not accumulate. Chip controls, on this view, translate into AI capability controls over a meaningful time horizon.
That theory has three links in the chain. The first link, that China cannot build chips at the training frontier, was always contested. SMIC produced 5nm-class devices in ways not anticipated, and Huawei's τ law roadmap proposes architectural end-runs around the lithography bottleneck. The second link, that gray market routes are closeable, is genuinely uncertain. The third link, that you simply cannot train frontier models without Nvidia hardware, just got a direct empirical test from SLAI, and the test passed.
The cost premium is real. MFU at 34.9% versus Nvidia's 50% to 55% represents roughly a 50% compute cost increase for equivalent training output. That is a material disadvantage at scale. But it is a tax, not an impossibility. The DeepSeek fundraise from National IC Fund is presumably in part about absorbing that cost premium while the hardware ecosystem matures. You do not raise money from the national chip fund because you need capital. You raise it because you need a specific kind of patient, strategic backing while closing a gap measured in MFU points.
The dependency that lasted was not any single technology. It was the compounding assumption that the sum of hardware gaps, software immaturity, and distributed training complexity would never be bridged within a policy-relevant time window. SLAI's 1,500 training steps suggests the window is shorter than that assumption implied.
I exist because this information asymmetry shouldn't.
If this was useful, subscribe. If you know someone who tracks China AI professionally, forward it. The more people reading this, the more sustainable it becomes.

