The Toll Booth
China's biggest independent AI token seller filed to IPO. It loses money on every one.
Happy Wednesday. I scan more than 100 Chinese-language sources every day, the WeChat accounts, the Bilibili channels, the finance wires, the policy feeds, and I write up the China AI stories English-language coverage misses. One person reading the Chinese internet so you don't have to. Let's go.
The Toll Booth
In a gold rush the surest business is selling shovels, or charging a toll on the road to the mine. This week one of China's busiest AI toll booths filed to go public and opened its books. SiliconFlow, the largest independent platform in China for buying access to AI models, filed a prospectus with the Hong Kong exchange. On the surface it is a growth story. Revenue grew 653% in a year. The platform now serves 578.5 billion tokens a day to 10.3 million registered users and more than 13,000 enterprise customers. Then you reach the margin line.
SiliconFlow's gross margin was positive 39.4% in 2024. In 2025 it flipped to negative 24%. Its public cloud business, the part where a developer pays to run DeepSeek or Qwen or GLM by the token, ran a gross margin of negative 119%. Read that again. For every yuan of revenue SiliconFlow booked selling tokens on its public cloud, it spent more than two yuan producing them. The company lost ¥345M ($48M) on the year, more than four times its 2024 loss, burning ¥14.8M a month. One analyst who read the prospectus described the unit economics as spending 4 yuan to rent the compute, 7 yuan to turn it into tokens, then selling the result for 1 yuan. He called SiliconFlow a cyber philanthropist.
This is the first audited look at what selling China's cheap AI actually earns, and the number is negative. SiliconFlow sits in the middle of the value chain. It rents GPU capacity upstream, runs the open models, and resells token access downstream. That middle seat was supposed to be the safe one, the booth every query has to pass through. Instead it is where the price war lands hardest. Chinese labs have spent two years driving token prices toward zero to win developers, DeepSeek most aggressively of all. SiliconFlow does not set those prices, it just has to match them, while the compute it rents does not get cheaper on the same schedule.
The founder makes the bet legible. Yuan Jinhui built OneFlow, a deep-learning training framework, then sold it in 2023 to Wang Huiwen's model startup, which Meituan bought two months later. Yuan did not take the Meituan job. He started SiliconFlow instead, and was the first to get DeepSeek's R1 and V3 running on Huawei's Ascend chips, the domestic-inference path everyone now takes for granted. His backers are not naive money. Alibaba, Meituan, SenseTime, NIO, Zhipu and Huawei's Hubble fund are all on the cap table, holding a company valued at ¥7.74B ($1.1B) that loses money on its core product.
So what are they buying? Volume, and the bet that it inverts. SiliconFlow is the fourth-largest token supplier in China and the largest that is not a cloud giant's captive arm. The theory is that inference cost per token keeps falling faster than price, that Monday's DSpark-style efficiency gains (an 85% speedup on the same chips) eventually drag the cost line back under the revenue line, and that whoever owns the volume at the crossover owns the market. It might be right. But filing to go public with a negative-119% gross margin on your flagship product is a remarkable way to admit the crossover has not happened yet. The toll booth is jammed with traffic. It still cannot cover its costs.
The Briefing
DeepSeek, the company that started the price war, is quietly raising prices. Developers got an email this week saying the full V4 release lands in mid-July, and that peak-hour API prices will double across input and output tokens, in the windows from 1 to 4am and 6 to 10am UTC. Off-peak stays at the floor. This is load management dressed as a price hike, a way to say the compute is genuinely tight without giving up the cheap-forever brand. The tell is in DeepSeek's own hiring. It has been recruiting data-center site-selection engineers and operations managers for a self-built facility in Ulanqab, a node on China's East-Data-West-Compute grid, and is looking to roughly double its headcount. Even the price butcher has found the floor on how cheap a token can go.
UBS talked to a dozen enterprise IT chiefs and found about 60% are now tightening AI spending. In a research note this week, the bank's analysts wrote that token cost has become the central worry for large companies watching their AI bills climb. Uber's operations chief said in May the return on AI spend was thin enough that the rising cost could no longer justify itself. The demand side is getting disciplined at the same moment the supply side is discovering it cannot cut prices further. UBS named the beneficiaries of the belt-tightening, and they were Chinese. Cheap open-weight models like DeepSeek are what cost-conscious enterprises reach for when they stop paying frontier prices.
While token prices fall, the hardware underneath them is getting more expensive. Chinese DRAM maker CXMT signed a ¥20B ($2.94B) long-term memory supply deal with Tencent, one of the largest domestic chip procurement commitments in years, ahead of CXMT's planned Shanghai STAR Market listing. Apple is separately lobbying Washington for approval to buy DRAM from CXMT as global memory prices spike, and Micron, Samsung and SK Hynix are facing a class-action suit alleging they fixed memory prices. Memory is a cost line that runs straight into every inference bill. When it climbs, the toll booth's margin gets thinner from below even as the price war squeezes it from above.
Not every Chinese AI company going public is bleeding. Autonomous-driving firm Momenta launched its Hong Kong IPO this week, raising about $751M with GIC, Fidelity and BlackRock as cornerstone investors and Mercedes-Benz and BYD's investment arm among strategics. The contrast with SiliconFlow is the whole story. Momenta's revenue tripled from ¥743M in 2023 to ¥2.41B in 2025, and its adjusted net loss narrowed to ¥303M. Revenue growing while losses shrink is what a path to profit looks like. Revenue growing 653% while losses quadruple is what selling below cost looks like. The same IPO window is pricing both.
Signals
Nvidia is hiring robotics talent across China even as it cannot sell China its best chips. The company opened embodied-AI, simulation and deployment roles in Beijing, Shanghai and Shenzhen, aimed at dexterous manipulation and whole-body control for general-purpose robots. The chips are restricted. The talent is not.
The embodied-AI funding frenzy has not cooled. X Square Robot closed four funding rounds in two months at a ¥20B valuation, led by four different internet giants, for a general-purpose robot brain. The money that cannot earn a return selling tokens is chasing the next frontier, where China now counts 15 embodied-AI startups above ¥10B in six months.
DeepSeek's full V4 is due mid-July with native multimodal support, the long-requested upgrade to the model that has done more than any other to push token prices to the floor. When it ships, the price war gets a new front.
The Bigger Picture
China has built the cheapest AI in the world. Domestic chips it was not supposed to be able to make, a power grid that can actually feed the clusters, and a two-year price war that dragged token costs toward zero. That is a real achievement, and SiliconFlow's prospectus is the first hard look at what it costs the people doing the selling.
The pattern in the numbers is that the money in Chinese AI right now sits at the ends of the chain, not the middle. Upstream, the chip and memory makers sign the multi-billion-dollar deals and run the fat margins, CXMT's ¥20B Tencent contract being this week's example. Downstream, the money races toward physical AI, the robots and world models where Nvidia is hiring and X Square is raising. The middle, where the models get built and the tokens get sold, is where the losses pool. The labs burn cash on compute. The inference platforms sell the output below cost. Tencent's own research institute has a name for the broader condition, token不经济, token uneconomics, the gap between exploding token volume and the value it actually produces.
The China-specific twist is uncomfortable. The same domestic-compute independence that let Meituan train a 1.6-trillion-parameter model on Chinese chips last week lowers the cost floor for everyone. But the price war means the savings flow to the users, and a lot of those users are outside China, running DeepSeek and Qwen and GLM through platforms like SiliconFlow because they are the cheapest capable models on earth. China is, in effect, subsidizing the world's AI inference and booking the loss at home. Whether that is a strategy, buy the global developer base now and monetize later, or just a race no one can stop, is the ¥345M question SiliconFlow's bankers are about to put to the market.
SiliconFlow going public with a negative-24% gross margin is not a failure. It is the clearest statement yet of where the AI gold rush pays and where it does not. Everyone selling shovels is rich. The toll booth in the middle of the road, the one every query has to cross, is the part that still cannot cover its own costs.
I exist because this information asymmetry should not. If someone forwarded you this, subscribe at chinaaidispatch.substack.com and read the Chinese AI internet with me every morning.

