This website uses cookies

Read our Privacy policy and Terms of use for more information.

Reading time: 4 minutes

Greetings from above,

Two rival labs, one week, one very pointed benchmark chart. Alibaba just published numbers that put its newest model right next to Anthropic's most capable one — on purpose.

I went through the release notes and the coverage line by line this weekend, because "comparable to Anthropic" is a big claim, and big claims deserve a close read before you believe them.

Today we're covering:

  • What Qwen3.8-Max actually is and what it claims to match

  • Why the pricing matters as much as the benchmarks

  • What this means if you're choosing between open and closed models

Let's get into it.

What Alibaba Actually Released

Alibaba released its biggest-ever AI model, claiming performance on par with Anthropic in the latest Chinese breakthrough aimed at challenging U.S. rivals. The new Qwen3.8-Max is built on 2.4 trillion parameters and ranks higher on several benchmarks than Moonshot's recently unveiled Kimi K3.

The comparison target is specific. Alibaba shared results showing Qwen3.8-Max delivering comparable, and in some cases better, scores than Anthropic's Fable 5 — on Arena.AI's leaderboard, it placed second worldwide for analyzing images and video, behind only Claude Fable 5.

Worth noting on the technical side: the model uses a "mixture-of-experts" approach, activating only 95 billion parameters at a time to keep costs and response delays down, despite the 2.4 trillion total parameter count. Alibaba also highlighted the model's autonomous coding and "long-horizon execution" — its ability to complete tasks requiring many sequential steps — as a core focus, and said it was able to independently perform a software engineering project over 16 days in internal testing.

The Pricing Angle Matters More Than the Benchmarks

Benchmark charts are easy to publish and hard to verify from the outside. Pricing is harder to fake.

Qwen is priced aggressively at $2 per million input tokens and $6 per million output tokens, which makes it look attractive compared to the best U.S.-built models, even accounting for the fact that different models use different numbers of tokens per task.

There's also a structural difference worth understanding if you're new to this space: open models like Qwen3.8-Max or Kimi K3 let developers download, run, and modify the AI directly, unlike proprietary models like ChatGPT and Claude, whose underlying systems stay private. Neither OpenAI nor Anthropic has disclosed the total parameter counts of their models, which is part of why direct comparisons are messy — Alibaba is publishing numbers its competitors don't.

Why This Is Happening Now

The release comes as Chinese AI firms work to close the gap with U.S. competitors, despite export controls limiting access to the most advanced AI chips, including Nvidia's H200. Qwen3.8-Max specifically positions itself against Moonshot's Kimi K3, launched July 16, which had already captured global attention on a scale not seen since DeepSeek's R1.

It's not a settled race, though. Prediction markets still put Anthropic's odds of holding the best-model title through the end of August 2026 at roughly 90.5%, against a marginal 0.4% for Alibaba, which tells you the market isn't fully buying the "on par" framing yet — benchmarks and real-world reliability aren't the same thing.

What This Means For You

  • If you're picking a model for autonomous coding or long, multi-step tasks, Qwen3.8-Max is now a real budget-friendly option worth benchmarking against whatever you're currently using

  • Open-weight models are catching up fast on published numbers, but remember: you can verify what you can run yourself, and you generally can't verify a closed model's internals

  • Don't switch on a headline. Run your own task against both models before moving anything production-critical

Wrap Up

What you learned today:

  • Alibaba's Qwen3.8-Max claims performance close to Anthropic's Fable 5, at a fraction of the price, with full parameter transparency Anthropic and OpenAI don't offer

  • The U.S.-China AI race is now genuinely close on paper, even if trust in benchmarks still favors the incumbents

  • Pricing and openness, not just raw capability, are becoming the real competitive lines to watch

The gap everyone assumed was permanent is looking smaller by the month. Worth keeping an eye on regardless of which model you use daily.

Thanks for being part of this community,

Keep learning,

🔑 Robert from God of Prompt

Reply

Avatar

or to participate