AINews

Reflection AI Beam: specs, release timing and benchmarks

Reflection AI Beam is a 501B open-weight model due this month under Apache 2.0. Its own table trails GLM-5.2 on reasoning, and nobody has tested it yet.

Source-based. Written from the documents, reporting and reviews linked in the text. Nothing here was tested hands-on by The Ruling Desk. How we work

An Nvidia GB200 NVL72 server rack on show at Computex 2024, lit from the edges, beside a liquid cooling cabinet
Photo: Geekerwan / Wikimedia Commons, CC BY 3.0

Reflection AI Beam is the first open-weight model from Reflection, the Nvidia-backed lab in Brooklyn, announced on October 5, 2026. It's a 501-billion-parameter model that Reflection says matches leading Chinese open models on reasoning while using far less computing power. You can't download it yet: the weights, a technical report and a model card are due later this month under the Apache 2.0 license, and until then every performance number is Reflection's own.

Key takeaways

  • Beam is big but sparse. It has 501 billion parameters in total, only 23 billion of which work on each token, and Reflection says it was pretrained on 23.8 trillion tokens with a context window of up to 1 million tokens. It reads and writes text only.
  • It isn't downloadable yet. Reflection opened an early-access waitlist and says the weights, technical report and model card arrive later in October 2026.
  • The parity claim is softer than it sounds. Reflection says Beam scores comparably to Z.ai's GLM-5.2 on reasoning with 3 to 4 times less inference compute. In Reflection's own table, Beam trails GLM-5.2 on all four reasoning tests where both have scores, by 0.7 to 4.6 points, while leading on most agentic and tool-use tests.
  • Nobody has checked it independently. As of October 6, 2026, we found no outside evaluation of Beam; the comparisons are Reflection's, using other labs' scores from third-party trackers.

What Reflection AI announced with Beam

Reflection laid out Beam in a launch post on its blog. It's a mixture-of-experts model, a design that splits the network into many specialist blocks and runs only a few of them for each word, which is why 501 billion parameters cost roughly what a 23-billion-parameter model does per token. Reflection says it built Beam for coding, reasoning and agentic work, meaning tasks where the model runs tools and takes several steps on its own.

The main numbers, all as Reflection states them:

SpecBeam
ArchitectureSparse mixture-of-experts, text only
Total parameters501 billion
Active per token23 billion
Pretraining data23.8 trillion tokens
Context windowUp to 1 million tokens
LicenseApache 2.0, when the weights ship
Training hardware6,144 Nvidia GB300 GPUs for pretraining, 10,500 for reinforcement learning

Reflection says it ran its reinforcement learning, the trial-and-error training stage that rewards correct answers, for four weeks and over 100 million attempts, with a maximum context of 256,000 tokens. A separate midtraining stage, it says, stretched the effective context to 1 million tokens. Beam also has a reasoning effort setting: lower values give shorter answers, higher ones let it think longer on hard tasks.

When you can download Beam

Not yet. Reflection says Beam is "undergoing final red-teaming and evaluations" and is going first to a select group through a waitlist. The company promises the weights under Apache 2.0, a permissive license that allows commercial use, plus documentation and tools for running, evaluating and fine-tuning the model, all "later this month." TechCrunch reports that the weights will reach developers through big cloud providers, smaller GPU clouds and open-source libraries.

Reflection hasn't published hardware requirements. As rough arithmetic of our own, 501 billion parameters take about 1 TB of memory at 16-bit precision and about half that at 8-bit, so this is a model for multi-GPU servers, not a laptop. For a smaller option that's already out, see our explainer on Aleph Alpha's 78B Kolibri, another Western open-weight release from the same week.

How Beam compares with GLM-5.2 on Reflection's own table

Reflection's headline claim is that on advanced reasoning tests Beam "achieves scores comparable to GLM-5.2" with 3 to 4 times less inference compute. GLM-5.2 is the open model from China's Z.ai, which TechCrunch puts at 744 billion parameters with 40 billion active, and which Z.ai ships under the MIT license on Hugging Face.

Here's a selection from the benchmark table in Reflection's post. The Beam scores are Reflection's; the others come from the trackers Artificial Analysis and DataCurve, Reflection says.

Test (Reflection's table)BeamGLM-5.2GLM-5.3
AIME 2026 (math)97.899.2not reported
Humanity's Last Exam, no tools36.240.542.3
GPQA Diamond (science)90.591.291.7
CriPT AA (reasoning)16.320.919.1
Terminal Bench 2.1 (coding)80.181.088.2
SWE Bench Pro v1 (coding)65.562.1not reported
MCP Atlas (tool use)78.777.884.2
AutomationBench public (agents)37.026.248.2
IFBench (instruction following)79.773.3not reported

Three things stand out. On reasoning, "comparable" means slightly behind: Beam trails GLM-5.2 on every reasoning test where both have a score. On agentic and tool-use tests, Beam is ahead of GLM-5.2 on most rows. And GLM-5.3, Z.ai's newer model, which Anthropic recently tested for exploit writing, beats Beam on every test where both are reported. Reflection itself says "frontier open models like Kimi K3 remain ahead on raw capability"; Kimi comes from the Chinese lab Moonshot AI, whose Hong Kong IPO plans we cover separately.

The compute claim is an estimate, not a measurement. Reflection says it counted roughly 2 times the active parameters times the average number of tokens each model generated, leaving out prompt processing and serving overhead, and calls the result "an approximate compute comparison rather than measured inference cost." By our arithmetic, Beam's smaller active count (23 billion against GLM-5.2's 40 billion, per TechCrunch) accounts for a factor of about 1.7; the rest of the claimed 3 to 4 times would have to come from Beam writing fewer tokens per answer.

Why a US lab is pitching against Chinese open models

Reflection was founded in 2024 by former Google DeepMind researchers Misha Laskin and Ioannis Antonoglou. When it raised US$2 billion at a US$8 billion valuation in October 2025, Laskin told TechCrunch that if the US didn't answer DeepSeek and Qwen, the global standard of intelligence "will be built by someone else", and that many companies and governments avoid Chinese models for legal reasons. TechCrunch's launch story says Reflection has now raised about US$4.7 billion from backers including Nvidia and Sequoia, and was last valued at US$25 billion pre-money, meaning before that round's new cash.

That's the pitch behind Beam: an American model with open weights that businesses can run and customize themselves. Reflection's post says Beam "advances the Western open-weight frontier." Early commentary collected by Latent Space's AINews placed Beam around GLM-5.2's level and called it a small win for US open models, but that reading rests on the same Reflection numbers.

What happens next

  • Weights, report and model card, later in October 2026, Reflection says. The technical report is also where it promises its safety evaluation results.
  • Independent benchmarks, once the weights are public. That's the first real test of the 3-to-4-times efficiency claim.
  • More models. Reflection calls Beam "the first model in a series" and says it's already training the next one.

Bottom line

Reflection AI Beam looks like a serious US entry in open-weight AI, with a license businesses can use and a strong pitch on efficiency. But it isn't downloadable yet, every number comes from Reflection, and its own table shows it slightly behind GLM-5.2 on reasoning and further behind GLM-5.3 and Kimi K3. If you build on open models, join the waitlist or wait for the weights and the first independent tests later this month before you plan around it.

FAQ

When will Beam be available to download?

Reflection says the weights, technical report and model card come "later this month," meaning October 2026, under the Apache 2.0 license. It hasn't given a specific date. For now, Beam is limited to a select group of early-access users from a waitlist.

Is Beam open source?

Beam is open-weight: Reflection says it will publish the trained model under Apache 2.0, which allows commercial use and modification. In 2025, Laskin told TechCrunch that Reflection would release model weights while keeping its datasets and full training pipeline proprietary, so you won't be able to rebuild Beam from scratch.

Is Reflection AI the company behind Reflection 70B?

No. Reflection 70B was a 2024 model from the startup HyperWrite whose claimed benchmark scores could not be reproduced. Reflection AI, the company behind Beam, is a separate lab founded by former Google DeepMind researchers.

Filed under AI

Newsletter

Console and phone guides, by email.

Fixes, settings and buying decisions for the console and phone you own, from the guides we publish. Free. Unsubscribe in one click. Your email is kept by beehiiv, our newsletter service, and used only for this newsletter.