Reflection launches Beam the first open-weight AI model with 501 billion parameters designed for coding and reasoning tasks

    Reflection launches Beam the first open-weight AI model with 501 billion parameters designed for coding and reasoning tasks

    Reflection has unveiled Beam, its inaugural open-weight model, boasting an impressive architecture of 501 billion parameters with 23 billion active units. This Mixture-of-Experts model is meticulously engineered for coding, reasoning, and dynamic tasks, exemplifying the cutting-edge progress in AI development.

    Beam’s performance is driven by significant efforts in both pretraining and reinforcement learning. The model was pretrained on a staggering 23.8 trillion high-quality tokens sourced from diverse web and proprietary datasets, achieving performance that rivals or surpasses similar models of comparable size. Simultaneously, the development of robust algorithms and infrastructure facilitated extensive and efficient reinforcement learning, using 10.5K NVIDIA GB300 GPUs to complete over 100 million rollouts in just four weeks.

    The results of these sophisticated training methodologies have established Beam as a formidable player in the context of open-weight models, delivering efficient inference computations. The model is currently in the final stages of red-teaming and evaluations. Early access sign-ups are available here, with plans to release the full model weights and supporting documentation later this month.

    Specifically tuned for coding and practical applications, Beam stands at the forefront of performance among open-weight models, competing effectively against larger counterparts like GLM 5.2 and nearing the capabilities of Qwen 3.8-Max in similar tasks. While some frontier models may excel in raw capabilities, Beam’s strength lies in its enhanced efficiency during inference.

    Beam’s architecture achieves impressive results in coding and logical reasoning tasks, demonstrating comparable performance to GLM-5.2 while using substantially less computational power, achieving a 3-4 times improvement in efficiency compared to models in the 2 trillion parameter category such as Qwen 3.8-Max.

    These advancements translate into a powerful tool for enterprises focused on coding and agentic workloads, offering a potent combination of high intelligence per token at a reduced operational cost. Beam’s development used high-compute reinforcement learning as an important scaling method, facilitating extensive exploration of problem-solving strategies, extended rollouts, and adaptability to feedback from various environments.

    With 10.5K GPUs engaged in generating over 100 million rollouts within a wide context length of 256K tokens, Beam’s training represents one of the largest reinforcement learning efforts conducted within an open lab. As the training progressed, an increasing richness in capabilities was noted, indicating continuous improvements without any signs of reaching a plateau.

    Beam uses asynchronous policy gradients to maintain learning stability across longer rollouts, overcoming challenges related to token freshness. Newly developed algorithms ensure the model’s performance remains consistent despite the use of older tokens during training. Users have the ability to adjust the reasoning parameter within Beam, allowing them to configure performance in line with their task needs and computational resources.

    In the pursuit of developing versatile reasoning abilities, Beam has demonstrated generalization across various tasks, even excelling in unfamiliar domains such as web searching and querying other models. Early demonstrations showcased Beam’s ability to perform complex tasks, including developing live subway dashboards using public data and creating intricate machine learning notebooks.

    To cater to the demands of frontier-scale reinforcement learning, the training used a repository of nearly one million challenging environments, curated through a systematic process designed to ensure both top notch and appropriate difficulty levels. This rigorous approach counteracted previous shortcomings observed during training runs that had insufficient data quality.

    The architecture for Beam’s training underscores a robust foundation, essential for integrating complex reasoning and dynamic task management capabilities. Continuous efforts to enhance data quality and infrastructure have resulted in a stable and efficient pretraining framework that allows Beam to maintain its learning trajectory and operational efficiency.

    Looking ahead, Reflection is committed to providing Beam as a widely accessible resource, with plans to release its model under an Apache 2.0 license along with essential documentation and tools for developers. As the first model in an ongoing series, Beam marks a significant milestone in the company’s ambition to push the boundaries of intelligence in the open AI landscape.

    Leave a Reply