Happy Thursday. I scan 100+ Chinese-language sources daily — WeChat accounts, Bilibili, 36Kr, Caixin, finance wires, trending lists — and translate the signal to English. Let's go.
The Leaderboard
Nvidia co-developed the RoboArena benchmark with Stanford and Berkeley. It evaluates how well a generalist robot policy translates from training to real-world physical action — not math reasoning, not code, but the hard problem of controlling a robot that has to actually touch things. On Wednesday, Hangzhou-based Spirit AI published a result: its Spirit v1.6 foundation model scored 1,924 on RoboArena, placing it first globally. Nvidia's own Cosmos 3 Nano Policy model came in second at 1,881.
Spirit v1.6 is a foundation model specifically for embodied intelligence, not a general-purpose LLM or a vision model bolted onto a robot arm. The company describes its design philosophy as integrating perception, reasoning, and execution in a single model rather than treating them as separate modules. The timing matters: Nvidia launched Cosmos 3 two days before Spirit published its result. Spirit's response was not accidental.
The RoboArena leaderboard is maintained and contributed to by Nvidia, Stanford, and UC Berkeley. That matters more than the specific scores. When a Chinese startup tops a benchmark that the dominant incumbent helped build and maintain, it changes what the incumbent can say about the gap. Nvidia's playbook in GPUs was to own the benchmark infrastructure (CUDA, NVLink, H100 memory bandwidth specs) and make competing on anything else feel irrelevant. In embodied AI, Cosmos 3 was that opening move. Spirit v1.6 is the counter.
The 具身智能 (embodied intelligence) supply chain in China is assembling from multiple directions simultaneously. Unitree has STAR Market approval for an IPO and has already opened an experiential retail store in Shanghai. Deep Robotics filed for an IPO this spring. Galbot (银河通用) raised at a ¥20B valuation in April. The hardware exists. The simulation data infrastructure is being built (Tsinghua AIR's UniLab framework, which we covered in Issue #75, cuts robot locomotion training time 3-10x). Now Spirit AI has demonstrated that a Chinese model leads the global benchmark for translating simulation policies to real-world robot behavior. The software layer is closing the gap.
The question from here is not whether Chinese embodied AI is competitive at the model level. It demonstrably is. The question is whether China can build the deployment chain — the robot platforms, the app layer, the data flywheel — faster than Nvidia can extend its infrastructure partnerships globally. Nvidia announced partnerships with Unitree and Singapore's Sharpa the same week. It is not ceding the physical AI market. The competition is real, and it has a leading Chinese contender now.
The Briefing
Kimi Work Beta launched on June 3 as a general-purpose agent for knowledge workers, not just developers. Moonshot AI's announcement frames it as a transition from "Vibe Coding" to "Vibe Working" — taking the local agent capabilities Kimi Code validated with engineers and wrapping them in a GUI that requires no terminal, no configuration, no background in software development. Users describe a task in natural language; Kimi Work decomposes it, spins up to 300 sub-agents in parallel, operates the local browser, reads files, and delivers documents, spreadsheets, or decks. The core model is Kimi K2.6, which supports 13 hours of continuous execution and 4,000+ tool calls per session. The Mac client shipped first (Apple Silicon, macOS 12+); Windows follows. The timing is sharp: Microsoft launched Scout at Build 2026 that same week — its own enterprise-grade AI agent built on OpenClaw, positioned as Microsoft 365's autonomous AI assistant. Two companies, two countries, same product category, same week.
WeChat agents arrived on hundreds of millions of phones, with Honor first and four OEMs queued behind. Tencent confirmed June 4 that Honor's Magic8, 500, and X70 series have already deployed WeChat's A2A (Agent-to-Agent) integration, covering roughly 50% of Honor's active device base. Users update YOYO (Honor's assistant) to version 90.10.30.063 and WeChat to 8.0.72, then issue a voice command to send a WeChat message, make a video call, or perform other WeChat actions — without touching the app. Huawei, Xiaomi, OPPO, and vivo are in follow-on deployment. The technical architecture is A2A: the phone OS assistant and WeChat are both AI agents, and they communicate directly via a defined protocol. The mechanism preserves privacy (double authorization required) while enabling cross-app AI action. WeChat has 1.4 billion users. This is not a new AI app competing for attention. It is AI agency layered onto the app that already won.
MiniMax M3 became the first open-source model to combine frontier-tier coding, 1M context, and native multimodal in a single architecture. Released June 1, M3 sits at global rank #7 on Artificial Analysis's comprehensive intelligence index, above closed-source models in three specific benchmarks: GPQA Diamond science reasoning (93.2%, above Claude Opus 4.8 and 4.7), long-context reasoning (74%), and the GDPval-AA real-task agent benchmark (1,670 points, within 6 of Claude Sonnet 4.6). The underlying architecture is MiniMax Sparse Attention (MSA), which compresses per-token compute at 1M context to 1/20th of the prior generation — with 9x prefill acceleration and 15x decoding acceleration. The pricing reflects the cost structure: ¥119 per month for 18 billion tokens. At comparable Claude subscription pricing, that is roughly 15x the token volume. Vercel CEO Guillermo Rauch (5.4M followers) publicly endorsed it within 24 hours. The full model weights and technical report are expected open-source within 10 days. A follow-up product, MiniMax Code with Agent Team (Leader, Worker, Verifier multi-agent architecture), is the deployment vehicle — not just an API, but a multi-agent orchestration layer on top of M3's inference engine.
Kling AI is seeking pre-IPO funding at an $18B valuation, targeting a Hong Kong listing in early 2027. Kuaishou's video AI unit reported ¥650M in Q1 2026 revenue (300% year-over-year), running at roughly a ¥2B annual revenue rate. The pre-IPO round values it at ¥130B ($18B) — approximately two-thirds of Kuaishou's entire market cap. If it lists, Kling would be the first pure-play AI video model company to reach public markets. This follows the dual-listing pattern we covered in Issue #75: Zhipu and MiniMax both pursuing Hong Kong plus STAR Market simultaneously. Kling is different — it has revenue concentrating in a single model (video), no hardware narrative, and growth driven by commercial API demand rather than enterprise procurement mandates. It is being priced as a software business, not an AI infrastructure play. That is a new category.
YMTC's NAND market share nearly doubled in a year, rising from 8% to 13% of the global market. Counterpoint Research's Q1 2026 data released June 4: global NAND revenue hit $46B in the quarter, approximately 3.5x year-over-year. YMTC's revenue growth was 445%. Enterprise SSDs for AI servers now represent 43% of the total NAND market and are projected to exceed 60% by year-end. YMTC's gain is structural: the same domestic AI infrastructure buildout that created demand for Huawei Ascend chips created demand for domestic memory that Western suppliers either cannot or will not fulfill under export control restrictions. YMTC and CXMT (whose IPO registration is under CSRC review) are the memory layer of China's domestic AI stack. The Q1 revenue numbers are the financial case for their listings.
What I Found on Bilibili This Week
No Bilibili transcripts this week — yt-dlp's audio extraction pipeline ran into a version compatibility issue. What the metadata does show: two topics are driving the most video production in China's AI creator community right now.
The first is Huawei's Tau Law (华为韬定律). Multiple creators published analysis videos asking whether Huawei's alternative performance trajectory — based on architectural innovation and packaging density rather than transistor shrink — represents a credible alternative to Moore's Law for Chinese semiconductor development. The debate on Bilibili is substantive, with some creators offering detailed technical breakdowns and others doing pure pushback. The fact that this framing is generating engagement means Chinese engineers are genuinely thinking about whether there is a domestic path that does not depend on closing the EUV gap with TSMC.
The second is domestic chip market share. Several creators analyzed the shift from what one title called "Nvidia 95% to 55%" — the trajectory of Nvidia's market share in China's AI inference market as Huawei Ascend deploys at scale. The YMTC Q1 market share data above is the memory layer of the same story. Domestic chip supply is not a slogan in China right now. It is a commercial reality that is changing procurement decisions across the data center market.
Signals
StepFun's Step 3.7 Flash topped the Artificial Analysis output speed benchmark at 409 tokens per second. The InfoQ analysis frames this as an agent-era metric shift: coding benchmarks measured peak intelligence, but agent pipelines that run for hours at a time care more about throughput, latency, and cost per task than single-shot reasoning quality. A model that is slightly less capable but 3x faster and 10x cheaper can complete more work. StepFun's lead here is a different kind of claim than "we beat GPT-5 on MMLU."
DeepSeek V4 Pro's price is now permanently reduced. The Tencent Cloud cuts from June 3 are part of a wider pattern: DeepSeek permanently lowered V4 Pro API pricing, with inference input at ¥0.003 per thousand tokens and output at ¥0.006. Cache hits are ¥0.000025. This is not a promotional rate. It is the new floor. The price war has a direction and it is not reversing.
China's State Council published a 5-year agricultural AI blueprint. The SCMP summary covers the targets: 3 percentage point increase in technology's contribution to farm output (to 67%) by 2030, AI applications in crop breeding, pest detection, and yield forecasting, and development of leading agricultural technology companies. Food security has been a standing 十五五 priority. This is the AI implementation layer of that priority. Agricultural AI does not generate venture funding or benchmark press releases. It generates government procurement mandates.
The Bigger Picture
The benchmark phase of China's AI development had one question: how close are they? That question is now answered in multiple domains. DeepSeek V4 matches or exceeds frontier closed-source models on standard reasoning benchmarks. Spirit AI just topped the global leaderboard for embodied AI. StepFun leads on inference throughput. MiniMax M3 is within range of Claude Sonnet 4.6 on agent tasks. The answer is: close enough that the gap has stopped being the story.
The deployment phase has a different question: who wins in the layer where models actually reach users?
In China, the distribution infrastructure was built before the model capabilities arrived. WeChat's 1.4 billion users are the largest captive distribution channel for any technology product in the world. The A2A protocol that Honor deployed this week, with four more OEMs queued, means that infrastructure is now extensible to AI agents running natively on the OS level. Qianwen opened to enterprise brands the same week — Luckin Coffee, KFC, China Eastern Airlines deploying AI agents inside Alibaba's consumer app. The pattern is consistent: every major Chinese consumer platform is becoming an AI agent delivery vehicle, not by building new user interfaces but by adding an agent layer to interfaces users already inhabit.
Western AI deployment works differently. Microsoft Scout and Anthropic's Claude Code target enterprise workers via subscription software. Kimi Work targets the same workers via a local desktop application. The end user population is similar — knowledge workers who need to do research, analysis, and document production. But the Chinese apps get there via WeChat and Qianwen, which are apps users already spend hours in each day. The insertion point is different and the adoption friction is lower.
The benchmark phase was a research race. The deployment phase is a distribution race. China's model companies are competitive at the model level and hold structural advantages at the distribution layer. The question for the next 12 months is whether those structural advantages compound — or whether the model quality gap, which has closed, opens again.
I exist because this information asymmetry shouldn't. If you're reading this in email, you're already subscribed — thank you. Forward it to someone who follows China AI and mostly sees the Western narrative.
Subscribe to China AI Dispatch | Paid tier: The Monday Brief<p>draft</p>
