AI & Technology8 min read

Can your AI stack survive Chinese pricing?

Mixue exported a price war, not a drink. Chinese AI labs are running the same playbook, and Anthropic just published the policy fight that decides where the floor settles.

I have been chewing on a ninety-second video from Eric Cracks China for a week now, and it has quietly rearranged how I think about commercial AI models. It opens: "What if I told you the world's biggest fast food chain doesn't sell burgers?"

The company is Mixue. It sells $1 ice cream and $2 bubble tea, and by store count it passed McDonald's and Starbucks some time ago. Most Western consumers have never heard of it (and neither had I). Eric's point is not about drinks. "They are not a beverage brand," he says. "They are a heavy industrial supply chain company disguised as a beverage brand." As he explains it, Mixue's frontend is a small red shop with a mascot; but their backend is tea powder, lemon concentrate, paper cups, straws and uniforms moving out to tens of thousands of franchise stores the way a factory feeds production lines. The public numbers back the framing: Mixue now runs close to 60,000 stores, and 97% of its revenue comes from selling materials and equipment to its own franchisees.

Which is why, as he puts it, "a $1 ice cream is not generosity. It is a weapon." The Chinese word for that weapon is involution: everyone cutting price, everyone squeezing margin, nobody able to stop. A company that comes out of that alive has been forced to rebuild itself around a cost floor its foreign competitors have never had to reach. Then it goes abroad, and the thing it exports is not just the drink - it's the way of doing business and ability to run on ultra-low costs.

Eric closes his bubble-tea video on a line I have not been able to forget yet: "If you can survive the price war at home, you can become the price war abroad."

I want to borrow the question and point it somewhere else. The same playbook is running in the world of AI right now, and the product on the conveyor belt is intelligence.

The numbers, without anesthesia

Chinese labs cut LLM API prices six times in the first half of 2026, and three of those cuts were made permanent rather than promotional.

DeepSeek launched V4-Pro in April, discounted it 75%, then made the discount the list price. It now sells at $0.435 per million input tokens and $0.87 per million output, with cache hits billed at a fraction of a cent. Set that against the US rate cards. GPT-5.2 lists at $14 per million output tokens. Claude Opus 4.8 lists at $25. That is a 30x gap on the output side, which is where agentic workloads spend most of their money.

The rest of the field is not far behind. Moonshot's Kimi K2.6 holds a cache-hit floor of $0.07 per million tokens. Zhipu's GLM-5 sits at $3.20 output, Alibaba's Qwen3 Max at $3.90. And several of these ship with open weights, which means the price is not even the binding constraint. You can download the model and run it in a datacenter you control.

These labs did not stumble into these prices. They fought a domestic price war among bleeding-edge labs that would have killed a Western startup, and the survivors came out the other side with cost discipline as a permanent trait rather than a campaign. That discipline is now pointed at your inference bill.

Where the lemonade analogy breaks

I want to be honest about the limit of the comparison, because it is the interesting part: nobody writes national security policy about iced tea. Intelligence is different, and that difference showed up in the open this week.

For weeks, an open letter backing open-weights models had been circulating with two conspicuous absences: Anthropic and OpenAI. Then OpenAI signed, and Anthropic was left as the only major American lab outside it. On 27 July, Dario Amodei published Anthropic's position, and the headline sentence is a denial: "Anthropic has never advocated for a ban on open-weights models." He goes further and grants the case for them, writing that "open-weights models that don't have dangerous capabilities are a public good," and that they expand access to the AI economy and strengthen competition.

What he proposes instead of a ban is narrower and, for anyone buying AI capacity, much more consequential. Three levers: keep advanced chips and chipmaking equipment out of authoritarian hands and prosecute the smuggling routes; stop industrial-scale distillation (interesting, since Anthropic and OpenAI have used it themselves to produce models they happily bill for); and require mandatory safety testing of all sufficiently capable models, open and closed alike.

Read those three as a procurement person rather than a policy person and something changes. They are not abstractions about the future of AI. They are the three inputs to the price you are being quoted.

Chip access sets the floor under Chinese training costs. Distillation is a large part of why models at $0.87 track models at $25 closely enough that the substitution is tempting in the first place. And the safety-testing line applied to open models means that the compliance work does not disappear when you move the weights inside your own perimeter. It relocates to you.

So the rate card in your business case has a policy derivative. The $0.435 is real today. What it does over the next couple of years depends on decisions being argued about in Washington right now, by people who are not thinking about your renewal date.

There is a second-order reading worth sitting with. The argument Amodei concedes, that open weights expand access and strengthen competition, is precisely the argument that makes this price war structural rather than promotional. If open weights are a public good and the leading open-weight models are Chinese, then cheap intelligence becomes the new floor, and it is set by someone else than the US labs.

Model choice is a supply chain decision

Most enterprises still run model selection like a taste test: benchmark scores, demo impressions, whichever vendor the CIO had dinner with. That frame made sense when there was ONE credible model. When intelligence becomes an input you buy by the unit, model selection stops being a technology choice and becomes procurement, and procurement brings supply chain risk with it.

Four questions belong in the next architecture review.

Concentration

If 80% of your AI workload runs against one provider's API, you have a single point of failure attached to a pricing dial you do not control. The Chinese labs just demonstrated that the dial can move 75% in a month - both ways. Also, you are one executive instruction away from completely losing access to that model, as happened with Fable.

Sovereignty

Where do your prompts go, and where do the weights live? In regulated European industries this is not hypothetical: it is DORA, the EU AI Act, and a national supervisor who will ask. Open weights genuinely change the calculus here. A model you host in Frankfurt under your own controls is a different risk object than an API endpoint in another jurisdiction.

Switching cost

The quality gap is collapsing. On the coding and agentic benchmarks most enterprises actually screen against, the leading models now cluster tightly enough that price, latency and context handling separate them more than capability does, across an order of magnitude of output cost. If your architecture cannot route between them, the spread is a loyalty tax you pay monthly. Staying portable across models is part of why I built Sulci.

Policy exposure

Which of your workloads would break if the model behind them became unavailable, untestable, or politically awkward in your market? Most organisations cannot answer this, because nobody has ever asked them to inventory it.

The question to take into your next architecture review

Chinese AI pricing is not a curiosity to file under geopolitics. It is a stress test arriving at your budget whether you invited it or not, and it now comes with a policy fight attached that will shape where the floor settles. Your board will see the DeepSeek rate card. Your CFO already has.

Ask the question early, on your own terms: can your AI stack survive Chinese pricing?

Eric's advice to founders facing Mixue is not to fight back with branding. "Don't start with the logo," he says. "Start with the refill system." The product is not the drink, it is the operating system behind the drink. The translation is uncomfortably direct. Do not start with the model. Start with what feeds it.

If your differentiation lives in the model, the odds are not good. Someone will always sell that model cheaper, and if they cannot, someone will open-weight it. If your differentiation lives in your data, its freshness, and the governance around it, the price war works for you. Every token gets cheaper. Your moat does not.

Preview image: Mixue on Broadway at West 32nd Street, New York, June 2026. Photo by 4300streetcar, licensed CC BY 4.0.

Let's Talk

Let's build something
worth talking about.

I take on a limited number of advisory and fractional engagements. Only projects where I can make a real difference. If you're navigating growth, AI, or revenue challenges in a technical B2B environment, let's talk.