Live
Abstract illustration of a lattice of dim purple nodes with a single glowing coral path routing through five lit nodes
AI & ML

Reflection AI’s Beam, Explained: The 501B Open-Weight Model Aimed at China

The most revealing number in Reflection AI’s launch post for Beam is one where Beam loses. On Terminal Bench v2.1, a test of how well a model gets real work done in a command line, Reflection’s own table gives Beam 80.1. It gives DeepSeek V4.1 Flash 90.6.

Reflection AI's announcement graphic for Beam
Reflection AI announced Beam on Oct. 5. Image: Reflection AI

Reflection published that comparison anyway, on Oct. 5, alongside the claim that Beam “advances the Western open-weight frontier.” Both things can be true. By Reflection’s numbers, Beam beats every other Western open model it tested, and it still sits behind the best Chinese open models on several of the tests the company chose. Here is what Beam is, what the jargon in its spec sheet means, and what you can actually do with it.

What is Reflection AI’s Beam?

Beam is Reflection’s first model. The company describes it as “a sparse Mixture-of-Experts model with 501 billion total parameters, 23 billion active, built for coding, reasoning, and agentic workloads.” It is text-only. Reflection says it pretrained Beam on 23.8 trillion tokens of web and licensed data using 6,144 Nvidia GB300 GPUs in under four weeks, then ran four more weeks of reinforcement learning on about 10,500 GB300s, with more than 100 million rollouts and roughly 1.3 billion software sandboxes used to train and grade the model. Midtraining extended its effective context to one million tokens, the company says.

Those GPU counts matter less as bragging rights than as a signal of who is paying. New York-based Reflection was founded in 2024 by former Google DeepMind researchers, Semafor reported, and in June it closed a funding round at a US$25 billion (about CA$35 billion) pre-money valuation, with investors including Nvidia, Sequoia and Citigroup. CEO Misha Laskin led reward modelling for Gemini at DeepMind, and CTO Ioannis Antonoglou co-created AlphaGo, according to Alex Heath’s Sources. TechCrunch reported that Reflection signed compute deals worth more than US$7 billion (about CA$9.7 billion) with SpaceX and Nebius this summer for access to GB300 chips through 2029.

Mixture of experts and ‘active parameters,’ explained

A parameter is one of the adjustable numbers a model learns during training. In a conventional “dense” model, every parameter does work on every word it reads or writes. A mixture-of-experts model splits much of the network into many smaller sub-networks, the experts, and a router picks a handful of them for each token.

That is why Beam has two sizes. It holds 501 billion parameters of knowledge, but only about 23 billion of them fire for any given token. The compute cost of generating text tracks the active number. The memory cost of storing the model tracks the total.

Think of a hospital with 500 specialists on staff. You see three. The building still has to hold all 500.

This design is not new. DeepSeek, Alibaba’s Qwen team and Z.ai all ship large MoE models, and MarkTechPost’s breakdown of the architecture notes that Beam borrows DeepSeek-V3’s auxiliary-loss-free method for keeping its experts evenly used. Reflection’s pitch is that it squeezes more out of each active parameter than rivals do.

Beam benchmarks: where it wins and where it trails

All of the following numbers come from Reflection. None has been independently verified, as TechCrunch noted.

Horizontal bar chart of Terminal Bench v2.1 scores published by Reflection AI: DeepSeek V4.1 Flash 90.6, Kimi K3 88.3, GLM 5.3 88.2, Qwen 3.8 Max 86.6, GLM 5.2 81.0, Beam 80.1, Inkling 63.8, Nemotron 3 Ultra 56.4
Reflection’s own Terminal Bench v2.1 numbers put Beam ahead of Western open models and behind the top Chinese ones. Scores are company-published and not independently verified. Graphic: prompt/power
  • Against Western open models, Beam leads. On SWE-bench Verified it scores 80.9, ahead of Thinking Machines Lab’s Inkling at 77.6 and Nvidia’s Nemotron 3 Ultra at 70.7. On Terminal Bench v2.1 it scores 80.1, against 63.8 for Inkling and 56.4 for Nemotron 3 Ultra.
  • Against Z.ai’s GLM 5.2, it is roughly even. Beam edges GLM 5.2 on SWE Bench Pro v1 (65.5 to 62.1) and DeepSWE v1.1 (44.4 to 44.0) and trails it slightly on Terminal Bench (80.1 to 81.0).
  • Against the newest Chinese models, it trails. On Terminal Bench, Qwen 3.8 Max scores 86.6, GLM 5.3 88.2, Kimi K3 88.3 and DeepSeek V4.1 Flash 90.6. On DeepSWE, DeepSeek V4.1 Flash scores 74.2 to Beam’s 44.4.

Reflection’s real argument is efficiency. “On advanced reasoning benchmarks, it achieves scores comparable to GLM-5.2 while using 3–4× less inference compute,” the company writes. That figure needs a footnote of its own. Help Net Security pointed out on Oct. 6 that the calculation multiplies active parameters by tokens generated and leaves out prompt processing, attention costs and serving overhead. It is a model-compute estimate, not a measured bill.

Beam is the best American open model on Reflection’s own scorecard, and that scorecard still has a Chinese model in first place.

What ‘open weights’ under Apache 2.0 means for you

Open weights means Reflection will publish the trained parameters themselves, the files that make up the model, so anyone can download and run Beam on their own hardware rather than renting it through an API. It is not the same as fully open source in the strict sense: Reflection has not said it will release its 23.8-trillion-token training data or the training code.

The licence is the part that matters for businesses. Reflection says “We will release the weights under an Apache 2.0 license,” one of the most permissive licences in software. It allows commercial use, modification and redistribution, and it carries none of the user caps or acceptable-use clauses that some earlier “open” model licences added. A Canadian bank, a hospital network or a provincial ministry could, in principle, fine-tune Beam on its own data and run it entirely on its own servers.

“In principle” is doing work there. Until the weights ship, the licence is a promise in a blog post.

How to try Reflection Beam, and the hardware it needs

Right now there is one route: an early-access waitlist on Reflection’s platform, linked from the launch post. Reflection says it will release “the weights, technical report, model card, and developer artifacts later this month.” As of Oct. 6 the weights had not been posted, and Reflection has published no API pricing in its launch materials.

When the weights do arrive, Beam will not run on a laptop. Our rough math: at 8-bit precision, 501 billion parameters take about 501 GB just to store; squeezed to 4-bit, about 250 GB, before any working memory for long conversations. That is multi-GPU server territory, or a cloud rental. The 23 billion active parameters make each answer cheap to compute once the model is loaded. They don’t shrink what has to be loaded.

If you want an open model on your own machine today, the realistic options are far smaller; we rounded up open-weight models worth running on a laptop. For most developers, Beam’s practical arrival will come when cloud providers host it. TechCrunch reported that Reflection plans to distribute it through hyperscalers and neoclouds.

Why a US open model matters in the race with DeepSeek

Lately the strongest open-weight models have come from Chinese labs. We covered how DeepSeek V4.1 Flash narrowed the US lead on LiveBench, and Reflection’s own table now shows the same pattern on agentic coding.

Reflection is selling to the buyers that gap leaves stranded. The company sees a market in “businesses, governments, and other entities building sovereign AI systems that don’t or can’t use Chinese models,” Laskin told Semafor. “They don’t really have very good options today,” he said. That pitch is aimed at places like Ottawa. In May, AI Minister Evan Solomon said advancing a B.C. data centre project with Telus meant “taking concrete action to build sovereign AI capacity here in Canada,” according to the federal announcement. Sovereign compute still needs a model to run, and Reflection is pitching Beam as that model.

Laskin also wants a seat in Washington for companies like his. “You want multiple voices around the table, both open and closed, and they should be working collaboratively on helping the government regulate AI,” he told Semafor.

Reflection is already training its successor, which Laskin told Semafor would be “much more” powerful than Beam. It will need to be. The number it has to beat is 90.6, and Reflection printed it.

// Contributor, AI Explained
Catherine Crowe

Catherine Crowe covers AI explained for prompt/power: the plain-English guides that break down how the technology works, what the jargon means and what it changes for everyday people. Originally from Canada, she writes from New Zealand.

Latest from prompt/power

  1. Thomson Reuters Won the First AI Training Appeal. Footnote 7 Is the CatchOct 7
  2. The Family Safe Word: How to Beat AI Voice-Clone Emergency ScamsOct 7
  3. Your SSN or SIN Leaked in a Breach? Do These 8 ThingsOct 7
  4. Apple’s Oct. 13 Event Rumour, Plus iPhone Duo Pre-Order Dates for CanadaOct 7
  5. Why Grindr Is Paying US$250M for Calgary PrEP Clinic FreddieOct 7

Leave a Reply

Your email address will not be published. Required fields are marked *