A two-year-old Brooklyn startup just launched an open-weight AI model — and opened with an admission no other American lab has made on a launch day: China’s best open models are still better, and Beam isn’t trying to beat them on power. According to Reuters, Nvidia-backed Reflection AI unveiled Beam on Monday, October 5, positioning it not as a king-slayer but as a cost-cutter — a model that claims to match China’s GLM-5.2 on advanced reasoning while using 3–4x less inference compute.
The pitch is refreshingly concrete. Beam is a text-only mixture-of-experts model with 501 billion total parameters, of which just 23 billion activate per task. Compare that with Z.ai’s GLM-5.2 at roughly 744 billion total and 40 billion active: Beam is the smaller engine, and it’s proud of it. It was pretrained on 23.8 trillion tokens and carries a 1-million-token context window.
The Honest Benchmark Table
Here’s where Reflection’s candor gets interesting. In the company’s own scorecard, Beam scores 80.1 on Terminal Bench 2.1 against GLM-5.2’s 81.0 — a genuine near-match. But the same scorecard concedes that newer Chinese models sit ahead, with Moonshot’s Kimi K3 at 88.3 and the newer GLM-5.3 also beating Beam. The wire coverage ran the parameter counts and TechCrunch noted the efficiency pitch — but nobody published the full side-by-side picture of Beam against GLM-5.2, GLM-5.3, and Kimi K3. Here it is:
| Model | Total / active parameters | Terminal Bench 2.1 |
|---|---|---|
| Reflection Beam | 501B / 23B | 80.1 |
| Z.ai GLM-5.2 | ~744B / 40B | 81.0 |
| Moonshot Kimi K3 | Not disclosed | 88.3 |
| Z.ai GLM-5.3 | Not disclosed | Ahead of Beam (per Reflection) |
One critical asterisk: every figure in that table is company-reported and unverified. TechCrunch flagged this explicitly, and it matters more than the scorecard itself — until someone else runs these benchmarks, the numbers are a marketing claim, not a result.
The Margin Argument
So why launch at all when your rivals are ahead? Because every point of reasoning quality at one-quarter the token cost is money for whoever runs it. For an API provider or an enterprise running inference at scale, matching GLM-5.2’s reasoning with 3–4x less compute isn’t a consolation prize — it’s a margin business. That framing, plainly stated, is more interesting than yet another “we beat the state of the art” press release.
The company itself is a familiar kind of founding story: Reflection was started in 2024 by ex-DeepMind researchers Misha Laskin and Ioannis Antonoglou, it has Nvidia’s backing, and it signed a compute deal with SpaceX earlier this year for capacity at the Colossus 2 data center.
The Catch: You Can’t Check Their Homework Yet
And here is the part launch-day coverage glossed over: the open weights aren’t public yet. Reflection says they’ll drop under Apache 2.0 later this month, with an early version currently gated behind a waitlist. Until independent researchers can download Beam and reproduce those numbers, every claim in this article — and every row in that table — is a promise, not a fact. Judge Beam when the weights land.
Why This Matters
The open-model race has a candor problem. Every launch claims supremacy; most quietly retreat when independent benchmarks land. Reflection’s bet is the opposite: tell the truth about where you rank, compete on the cost of every token, and let the margin argument do the talking. If the weights prove out later this month, Beam becomes the model American companies reach for not because it’s the smartest — but because it’s the smartest per dollar. In a market where inference bills are the new cloud bills, that’s the more honest race to win.
Meanwhile in AI Tech: Google just made its budget model tier free — and Tesla is trimming memory from the Optimus program to cut costs. The compute-cost squeeze is this week’s real story.


