AI Tools Review
Reflection Beam 501B: Specs, Benchmarks & Verdict

Insights

Reflection Beam 501B: Specs, Benchmarks & Verdict

AI Tools Review Editorial Team10 October 2026

    Quick Answer:

    Beam is Reflection AI's first open-weight model, announced on 05/10/2026: a sparse Mixture-of-Experts model with 501 billion total and 23 billion active parameters, built for coding, reasoning and agentic work. Reflection reports 80.9 on SWE-bench Verified and 80.1 on Terminal-Bench 2.1, says it is competitive with GLM 5.2 at three to four times less inference compute, and plans to release the weights, a technical report and a model card later in October 2026. It is not yet downloadable, and every number below is Reflection's own.

    For most of 2026 the best open-weight models have come from China. DeepSeek, Zhipu, Moonshot and Alibaba have set the pace, while American labs have mostly kept their strongest models behind APIs. Reflection AI, a Nvidia-backed start-up, wants to change that with Beam, and the YouTube commentary has already started calling it "American DeepSeek".

    This guide is based on Reflection's own launch post, which is unusually detailed about pretraining and reinforcement learning, plus trade-press coverage of the announcement. We separate what Reflection has published from what remains a promise, we reproduce the company's figures and name them as the company's, and we have not run the model ourselves because the weights are not out yet.

    AI Revolution X covers the Beam announcement and frames it as the Western answer to DeepSeek.

    Executive Summary

    What it is: Beam is a text-only, sparse Mixture-of-Experts model with 501B total parameters and 23B active per token. It was pretrained on 23.8 trillion tokens of web, public and proprietary licensed data, then pushed through what Reflection calls one of the largest reinforcement learning runs any open lab has attempted. The API model name is Beam-501B-A23B, and trade coverage reports a 1 million-token context window.

    • Headline scores (Reflection): SWE-bench Verified 80.9, Terminal-Bench 2.1 80.1, SWE-bench Pro v1 65.5, GPQA Diamond 90.5, AIME 2026 97.8, MCP Atlas 78.7, Humanity's Last Exam (no tools) 36.2.
    • Positioning: competitive with GLM 5.2 and approaching Qwen 3.8 Max on coding and agentic tasks, with frontier open models such as Kimi K3 still ahead on raw capability. The pitch is intelligence per token, not the top of the leaderboard.
    • Training scale: 6,144 NVIDIA GB300 GPUs for pretraining in under four weeks; 10.5K GB300 GPUs for four weeks of reinforcement learning with more than 100 million rollouts.
    • Status: announced, not shipped. Reflection says it is in final red-teaming and evaluation and that the weights, technical report and model card follow later this month. An Apache 2.0 licence has been signalled, and a beta API waitlist exists with no published price.
    • Biggest caveat: self-reported benchmarks, a table with many "not reported" cells, and a baseline on Terminal-Bench 2.1 that trails every open model it is compared with except Inkling and Nemotron 3 Ultra.

    Our view: Beam is a credible and unusually transparent debut. If the weights land under a genuinely permissive licence, it gives Western enterprises a frontier-adjacent open model they can self-host without the political baggage some buyers attach to Chinese labs. Whether it earns the "American DeepSeek" label depends on independent tests that do not exist yet.

    Why Beam Matters: The Western Open-Weight Gap

    Open-weight models matter because they can be run on your own hardware, fine-tuned on private data and inspected in ways closed APIs never allow. In 2026 the frontier of that category has been dominated by Chinese labs. We have reviewed several of them: GLM-5.3 and its predecessor GLM-5.2 from Zhipu, Kimi K3 from Moonshot, Qwen 3.8 Max from Alibaba, and DeepSeek V4.1 Flash. Those four names are exactly the comparison set in Reflection's benchmark table, which tells you which models Reflection thinks it has to beat.

    The American entries have been fewer. Our Thinking Machines Inkling coverage describes a debut open model from a well-funded US lab, and Beam's own table includes Inkling as a baseline: on SWE-bench Verified Beam scores 80.9 against Inkling's 77.6, and on Terminal-Bench 2.1 it scores 80.1 against 63.8. NVIDIA's Nemotron 3 Ultra is the other Western reference point, and Beam scores higher on most rows where both report a score (IFBench is the exception: 79.7 against 81.7).

    The policy context is just as important. The debate over whether the United States should encourage or restrict open weights has run through our coverage of the US-China AI race and lobbying and the open-weights letter from NVIDIA and Anthropic's response. A strong, permissively licensed American model strengthens the case that openness and national competitiveness can go together. That is why a launch post from a start-up gets this much attention.

    Architecture and Pretraining

    Reflection's post spends as much time on how Beam was built as on what it scores, which is welcome. The architecture is a 52-layer sparse Mixture-of-Experts transformer that combines interleaved local and global attention, fine-grained routed experts, a controlled residual stream and several forms of load balancing. Only 23B of the 501B parameters fire for any token, which is the basis of the efficiency claim: roughly 4.6% of the weights are active per token, so serving cost tracks the active count rather than the headline size, although the full weight set still has to be stored and distributed across hardware.

    Data and compute

    Beam was pretrained on 23.8 trillion tokens from the web, public sources and proprietary licensed datasets, end to end in under four weeks on a cluster of 6,144 NVIDIA GB300 NVL72 GPUs. Reflection says it trained its own quality classifiers for web, code and STEM content, sorted data into fine-grained quality tiers and weighted training towards the best material. Two figures stand out: about 95% of raw internet tokens are eliminated through parsing, deduplication and curation, and Reflection says conventional filters would have missed roughly 1.8 trillion tokens it kept, including 87% of its curated web-code tokens.

    The company says it trains on almost all publicly accessible, unrestrictively licensed code and documentation on the web, and built a pipeline that uses a vision-language OCR model to process petabytes of PDF material so the base model has STEM knowledge. Code and technical content was repeated several times over the training horizon, which Reflection describes as requiring careful fuzzy deduplication and attention to the "science of overtraining".

    Two scatter charts of validation loss against training FLOPs for code and web data, showing Beam Base on a straight scaling line and below DeepSeek v4 Flash Base and Nemotron 3 Ultra Base
    Beam's pretraining recipe scales predictably across four orders of magnitude in compute, and Beam Base sits at or below DeepSeek v4 Flash Base and Nemotron 3 Ultra Base on decontaminated code and web validation loss (lower is better). Source: Reflection AI.

    The chart above is the evidence for Reflection's claim that the final Beam Base matches its predicted performance and matches or outperforms comparable open base models. Note what it does and does not show: the metric is validation loss in bits per byte on Reflection's own decontaminated validation sets, which is a useful scaling-law check but not a downstream benchmark. On the code panel Beam Base is lowest of the three plotted models; on the web panel it is close to Nemotron 3 Ultra Base and clearly below DeepSeek v4 Flash Base.

    Line chart of Beam pretraining loss in bits per byte falling smoothly over the full run of 23.8 trillion tokens on 6,144 NVIDIA GB300 GPUs
    Beam's training loss across the pretraining run (23.8 trillion tokens, 6,144 NVIDIA GB300 GPUs). Reflection reports no instabilities or large irrecoverable spikes. Source: Reflection AI.

    Reflection reports nine semi-automatic rewinds during the run, attributed to non-deterministic gradient-norm spikes or suspected silent data corruption, and a goodput of 92.3% towards the end, meaning the share of wall-clock time spent on training steps retained in the final model. It built a topology-aware Kubernetes-based scheduler, a node lifecycle system and a silent-data-corruption detector in-house.

    Stable, balanced MoE

    Mixture-of-Experts models can waste capacity if the router sends most tokens to a few favourite experts. Reflection built on the auxiliary-loss-free load balancing introduced by DeepSeek, added a cosine decay of expert-bias updates to reduce routing perturbations late in training, and layered on sequence-level balancing so utilisation stays even on data outside the pretraining distribution, which matters because reinforcement learning shifts the distribution.

    Line chart of Beam expert load balance during pretraining, falling from about 2x to 1.04x the uniform load by the end of training
    The busiest expert's load relative to uniform routing, averaged across MoE layers, falls to 1.04x by the end of pretraining (1x is perfectly balanced; lower is better). Source: Reflection AI.

    The result, according to Reflection, is almost-perfect uniform utilisation, so every expert can contribute to learned reasoning. For the residual stream the team used depth-based scaling, SandwichNorm, elementwise attention gating and FP32 residual accumulation to keep activations bounded across all 52 layers. These are engineering details, but they explain the thesis of the whole post: a numerically healthy base is what lets reinforcement learning run for weeks without collapsing.

    The Reinforcement Learning Run

    The reinforcement learning stage is where Beam's capabilities were made. Reflection used 10.5K NVIDIA GB300 GPUs for four weeks, generating more than 100 million rollouts with a maximum context length of 256K tokens. Training and grading used roughly 1.3 billion sandboxes, and the team assembled a pool of nearly one million environments covering software engineering, terminal use, competitive coding, STEM, web search, tool use and general knowledge work.

    Reflection compares its rollout count with other open efforts: its text says Inkling was trained on 30 million rollouts and MiMo on 753 thousand, and it calls Beam's run one of the largest by any open lab to date. That is a claim about process, not about final quality, and it shows how much of Beam's capability Reflection attributes to reinforcement learning rather than pretraining alone.

    The most useful lesson in the post is about data quality. Reflection says compromises in task quality led to capability plateaus and other training problems, and that it filtered tasks for difficulty (neither always solvable nor impossible) and quality (not underspecified, misleading, guessable or hackable). Independent judges re-screened passing solutions for verifier exploits, and replayable records made rewards inspectable. Anyone who has watched models learn to game weak graders will appreciate why that matters.

    Infrastructure numbers

    Reflection published a long list of platform statistics. We reproduce the ones that give a sense of scale:

    • An average of 110K concurrent rollouts, and up to 170K concurrent sandboxes.
    • More than one billion sandbox creation requests across over 20 clusters, two clouds and four regions, with 90% of new sandboxes ready in under 10 seconds.
    • Inference-to-training GPU ratios from 3.9:1 to 5.4:1, adjusted as the workload evolved, with the trainer resized across five GPU mesh configurations without losing state.
    • New weights reaching the inference fleet in a median of about 12 seconds, with hierarchical distribution cutting cross-rack traffic by 75% and making fleet-wide adoption 2.2 times faster.
    • 71 inference incidents handled without terminating training, with median recovery of eight minutes and 0.02% of serving GPU-minutes lost.
    • Training batches 99.99% full on average through dynamic packing, with per-GPU throughput held within 1.5% as mean rollout length grew almost 70%.

    The asynchronous design deserves a mention. Rollouts are generated by older checkpoints while the trainer learns, so tokens in long rollouts can come from several policy versions. Reflection says it developed algorithms to keep learning stable under that policy staleness and training-inference mismatch, and that training remained stable even when learning from interactions generated more than a day earlier, which it puts at 107 weight versions behind the current policy.

    Learning to reason efficiently

    Reflection trained Beam with a controllable length penalty that rewards successful solutions while discouraging unnecessary tokens. Early in training, performance improved even as completions got shorter; later, as agentic ability grew, completions lengthened again but with real gains. Users get a reasoning-effort parameter to trade response length against performance on hard tasks.

    This matters because verbose reasoning is the hidden cost of modern models. Reflection estimates inference compute as roughly two times the active parameter count times the mean generated tokens per attempt, using data from Artificial Analysis and DataCurve, and excludes prompt prefill, attention over context and serving overhead. By that estimate Beam matches GLM-5.2 on advanced reasoning with three to four times less compute, and the gap widens against 2T-plus models such as Qwen 3.8 Max. Treat it as a back-of-envelope comparison, as Reflection itself does: it is an approximation, not measured cost.

    Benchmarks: What Reflection Reports

    Reflection's table compares Beam with Inkling, Nemotron 3 Ultra, GLM 5.2, GLM 5.3, Kimi K3, Qwen 3.8 Max and DeepSeek V4.1 Flash. The company updated Beam's results on 08/10/2026 and says the current tables reflect the latest numbers. Cells marked "NR" are scores that have not been reported, so a missing number is not a zero. We reproduce selected rows below. Every figure is Reflection's.

    Coding and terminal work

    BenchmarkBeamInklingGLM 5.2GLM 5.3Kimi K3Qwen 3.8 MaxDeepSeek V4.1 Flash
    SWE-bench Verified80.977.6NRNRNRNRNR
    SWE-bench Pro v165.554.362.1NRNR67.7NR
    SWE-bench Pro v2-Hard77.256.9NR84.388.2NRNR
    Terminal-Bench 2.180.163.881.088.288.386.690.6
    DeepSWE v1.144.4NR44.061.068.051.074.2
    SWE-bench Multilingual78.0NRNRNRNRNRNR

    The pattern is consistent. Beam clearly beats the other Western open models where both report a score, roughly ties GLM 5.2 on Terminal-Bench 2.1 (80.1 against 81.0) and is ahead of it on DeepSWE (44.4 against 44.0), but sits well behind the latest Chinese models on harder suites: 61.0 and 68.0 for GLM 5.3 and Kimi K3 on DeepSWE, and 74.2 for DeepSeek V4.1 Flash, against Beam's 44.4. That DeepSeek V4.1 Flash result is notable because it is a "flash" model, and it is a reminder that the gap Reflection describes as "raw capability" is real.

    Reasoning and general capability

    BenchmarkBeamInklingGLM 5.2GLM 5.3Kimi K3Qwen 3.8 MaxDeepSeek V4.1 Flash
    AIME 202697.897.199.2NRNRNRNR
    GPQA Diamond90.587.291.291.793.592.690.9
    HLE (no tools)36.229.740.542.346.943.639.1
    SciCode49.746.1NR59.058.752.152.0
    AA-LCR79.377.378.379.788.780.384.0
    IFBench79.779.873.3NRNR82.8NR

    On reasoning Beam is solid but not leading. GPQA Diamond at 90.5 is within a point of GLM 5.2 and DeepSeek V4.1 Flash, while Humanity's Last Exam without tools at 36.2 is below every Chinese comparison model that reports a score. Reflection's efficiency claim is therefore the right one to focus on: these are mid-pack scores achieved with far fewer generated tokens, by its estimate.

    Tool use and search

    On MCP Atlas, a tool-calling benchmark, Beam scores 78.7 against Inkling's 76.0 and GLM 5.2's 77.8, but behind GLM 5.3 (84.2), Kimi K3 (82.3) and Qwen 3.8 Max (84.5). On BrowseComp with context management it posts 77.4 against Kimi K3's 91.2, and on DeepSearchQA 80.1 against 95.0. On AutomationBench Beam scores 37.0, ahead of GLM 5.2 (26.2) and Inkling (not reported) but well behind DeepSeek V4.1 Flash (54.8). On tau3 banking it posts 38.0 against Qwen 3.8 Max's 55.2. If you want an agent that browses and researches, the current open leaders are clearly ahead; if you want a coding workhorse that is cheap to run, Beam's case is stronger. For readers choosing an agent harness, our best AI coding agents guide covers the tooling side.

    Reflection's Demonstrations

    The launch post includes four worked examples. They are vendor-selected, so read them as capability illustrations rather than tests:

    • Astronaut free-fall game: a 3D p5.js game in which an astronaut dodges or destroys asteroids, built by a text-only model reasoning about how visuals should look.
    • Land and water puzzle: given a fixed 180 by 90 grid of 16,200 longitude and latitude points, Beam reasoned in text and got 95.5% coverage right. Reflection places this between Opus 5 (92.5%) and Fable 5 (97.8%).
    • NYC subway live map: Beam found documentation, checked authentication needs, located map geometries, built frontend and backend and kept the server running for live updates.
    • Fine-tuning notebook: plugged into OpenCode, it researched Unsloth and the latest smallest Gemma-4 model and wrote a Text2SQL fine-tuning notebook, which Reflection says raised Gemma's held-out accuracy by 66.5%.

    Reflection also notes an interesting generalisation result: during training on reasoning, software engineering and terminal tasks it saw consistent gains in browsing even though no browsing tasks were in the reinforcement learning mix. When given web access, it organically learned to search for and query other large language models and to use OCR APIs to read documents. It is the company's own observation rather than an independent finding.

    Julian Goldie walks through the Beam announcement and what a 501B open-weight model could mean for builders.

    Safety, Licence and Release Status

    This is the section where Beam differs most from the closed-model launches we cover, such as Claude Opus 5.5, which arrived with a full system card. Reflection says Beam is undergoing final red-teaming and evaluations and that the model card, technical report, weights and developer artifacts will be released later in October 2026. The blog post has a short safety and alignment section, but there is no model card yet, so we cannot report dangerous-capability evaluations, refusal behaviour or alignment findings. Anyone planning to deploy Beam in a regulated setting should wait for that document.

    On licensing, trade coverage reports that Reflection plans an Apache 2.0 licence together with documentation and the full stack for running, evaluating and fine-tuning the model. Apache 2.0 would be about as permissive as open-weight licences get, but nothing is verifiable until the weights ship. Treat the licence as a stated intention.

    Pricing and Access

    There is no published API price. Early access is by sign-up and the OpenAI-compatible API (Beam-501B-A23B) is in a beta waitlist. Reflection's argument is that more intelligence per token means lower cost, but it has not put a number on that. Compare that with StepFun's Step 5 Preview, which is already priced at $1.00 input and $2.70 output per million tokens, and with our reviews of DeepSeek V4.1 Flash for the cheap end of the market.

    Self-hosting is the other route. At 501B total parameters the weights will need a multi-GPU node even though only 23B are active per token, so smaller teams will probably use hosted endpoints from inference providers once the weights are public. The sparse design lowers compute per token, not memory footprint.

    How It Compares

    • Versus Inkling and Nemotron 3 Ultra: ahead on most rows where both report a score (with small exceptions such as IFBench and AA Omniscience against Inkling), which supports the claim that Beam advances the Western open-weight frontier. See our Inkling article.
    • Versus GLM 5.2: roughly level on coding and reasoning, behind on tool use and search on some rows, ahead on efficiency by Reflection's estimate.
    • Versus GLM 5.3, Kimi K3 and Qwen 3.8 Max: behind on raw capability across most published rows. See GLM-5.3, Kimi K3 and Qwen 3.8 Max.
    • Versus StepFun Step 5 Preview: a different bet on the same idea, a large sparse model with a cheap active footprint. Step 5 is 600B total and 27B active and is API-first; Beam is 501B and 23B active and open-weight-first.
    • Versus closed frontier models: not in the same tier on difficult agentic work, and Reflection does not claim otherwise.

    Limitations and Open Questions

    • Self-reported results. Reflection's figures are not independently reproduced, and the table was already revised on 08/10/2026. Independent scores, from groups such as Artificial Analysis, will matter.
    • Selective comparison. Many cells are "NR", so head-to-head comparisons are patchy. Some third-party coverage disagrees about which Kimi model is the right comparison point.
    • Text only. No native image or audio input, which matters for computer-use and design tasks.
    • Efficiency is an estimate. The 3 to 4 times compute claim excludes prefill, attention cost and serving overhead.
    • Not yet released. No weights, model card, technical report or confirmed licence at the time of writing, and no price.
    • Safety evidence pending. We cannot assess misuse risk, refusals or alignment until the model card appears.

    Who Should Care

    Watch closely if you are an enterprise or public-sector buyer that wants a self-hostable model from a Western lab, a developer building coding agents who cares about cost per solved task, or a researcher who wants to study a model trained with a documented reinforcement learning recipe. Wait for independent tests if you need the strongest open model today (the Chinese leaders still score higher), or if you need browsing and research performance. Do not deploy yet if you need a published safety evaluation, because it does not exist.

    The Bottom Line

    Beam is the most serious Western open-weight debut we have seen this year, and the launch post is a rare piece of transparency about frontier training. The numbers show a model that leads its Western peers, matches GLM 5.2 on several coding and reasoning tests and trails the best Chinese open models, with an efficiency argument that is plausible but unproven. The headline risks are straightforward: self-reported scores, no model card and a licence that is still a promise. If Reflection ships Apache 2.0 weights later this month as stated, expect a wave of independent benchmarks and fine-tunes. We will update this page when they arrive.

    Sources

    Images and charts: Reflection AI, Introducing Beam (hero image, pretraining scaling, expert load balance and pretraining loss figures), reproduced for review and commentary with credit.

    Last updated: 10/10/2026. Sourced from Reflection AI's launch post and trade coverage. We have not tested Beam ourselves; all benchmark figures are vendor-reported, and the weights, model card and licence were not yet public at the time of writing.

    Free Guide

    Get the free guide: Claude vs ChatGPT, Gemini & Grok

    A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.

    Pop your email in to get it free
    Preview of the free guide: Claude vs ChatGPT, Gemini and Grok, 2026 features, pricing and what-you-can-do comparison.

    Frequently Asked Questions

    What is Reflection AI's Beam?
    Beam is Reflection AI's first open-weight model, announced on 05/10/2026. It is a sparse Mixture-of-Experts model with 501 billion total parameters and 23 billion active per token, built for coding, reasoning and agentic workloads. It is text-only and was pretrained on 23.8 trillion tokens.
    Can I download Beam's weights yet?
    Not at the time of writing. Reflection says Beam is in final red-teaming and evaluation and that it will release the weights, technical report, model card and developer artifacts later in October 2026. Early access is by sign-up. Reflection has indicated an Apache 2.0 licence, but you should confirm the terms when the weights actually ship.
    How good is Beam on coding benchmarks?
    Reflection reports 80.9 on SWE-bench Verified, 80.1 on Terminal-Bench 2.1, 65.5 on SWE-bench Pro v1 and 78.0 on SWE-bench Multilingual. On Terminal-Bench 2.1 that is behind GLM 5.3 (88.2), Kimi K3 (88.3), Qwen 3.8 Max (86.6) and DeepSeek V4.1 Flash (90.6). All figures are Reflection's own and have not been independently reproduced.
    Is Beam the American answer to DeepSeek?
    It is positioned that way by commentators. Reflection says Beam advances the Western open-weight frontier and is competitive with GLM 5.2, but that frontier open models such as Kimi K3 remain ahead on raw capability. Its pitch is efficiency: scores comparable to GLM-5.2 on advanced reasoning at three to four times less inference compute, by Reflection's own estimate.
    What hardware was Beam trained on?
    Pretraining ran in under four weeks on 6,144 NVIDIA GB300 NVL72 GPUs. The reinforcement learning stage used 10.5K NVIDIA GB300 GPUs for four weeks, generating more than 100 million rollouts with a maximum context length of 256K tokens.

    Explore more AI tool comparisons

    In-depth reviews, benchmarks and guides to help you choose the right AI tools.

    Browse all reviews
    AI Tools Review Editorial Team

    AI Tools Review Editorial Team Expert verified

    Our editorial team consists of veteran AI researchers, software engineers, and industry analysts. We spend hundreds of hours benchmarking frontier models natively to provide you with objective, actionable intelligence on agentic AI capabilities and cybersecurity landscapes.