Google’s Medical AI Chatbot Got Its First Real-Clinic Test. Here’s How It Did

In a Lancet study at a Boston clinic, Google’s AMIE chatbot interviewed 100 patients before urgent care visits and listed the confirmed diagnosis among its top picks 90% of the time. A doctor watched every chat.
HomeTech & AIReflection’s Beam Is a 501-Billion-Parameter Open AI Model Built to Take On...

Reflection’s Beam Is a 501-Billion-Parameter Open AI Model Built to Take On China’s Best

An American AI lab is making a bid to win back the open-model crown. Reflection, a startup founded by Misha Laskin and Ioannis Antonoglou, on Monday unveiled Beam, a 501-billion-parameter model whose weights the company says it will release for anyone to download later this month.

That matters because, for the past couple of years, the best “open-weight” models (the kind developers can run on their own hardware and customize) have mostly come from Chinese labs, with model families such as Qwen, GLM and Kimi. SiliconANGLE described Beam as the first U.S. open model to post performance comparable to those Chinese alternatives.

What Reflection announced

According to Reflection’s announcement, Beam is:

  • A “mixture-of-experts” model with 501 billion total parameters, of which only about 23 billion are active for any given word it processes. Think of it as a big team of specialists where only the relevant few are called in for each task, which keeps running costs down.
  • Built for coding, reasoning and “agentic” work, meaning multi-step tasks where the AI uses tools, runs code or searches the web on its own.
  • Able to handle very long inputs, with a context window extended to 1 million tokens (roughly several thick books’ worth of text).
  • Text-only. It can work with other kinds of content only when they are turned into text first, for example through OCR.
  • Licensed under Apache 2.0, a permissive license that allows commercial use.

How it was trained

Reflection says Beam’s base model was trained on 23.8 trillion tokens of web data and licensed datasets, using 6,144 Nvidia GB300 GPUs for under four weeks. A second phase of reinforcement learning, where the model practices tasks and is rewarded for getting them right, ran on about 10,500 GB300 GPUs for another four weeks. The company says that phase produced more than 100 million practice attempts across roughly a million coding, agent and science environments.

The computing power came in part through a $6.3 billion deal with SpaceX to rent Nvidia GB300 systems, SiliconANGLE reported. TechCrunch first reported the SpaceX compute deal in June.

How good is it?

Reflection published a set of coding benchmarks. Among the highlights it reported:

  • SWE-Bench Verified (fixing real bugs in open-source projects): 80.9
  • SWE-Bench Pro v1: 65.5, ahead of GLM 5.2 (62.1) and just behind Qwen 3.8 Max (67.7)
  • Terminal-Bench v2.1 (working in a command line): 80.1, roughly level with GLM 5.2 (81.0) but behind Kimi K3 (88.3)

The company also claims Beam matches GLM-5.2’s reasoning performance while using three to four times less computing power to generate answers. That is the core of its pitch: not the smartest open model on every test, but a strong one that is cheaper to run.

Why “open weights” matter

Most of the AI people use day to day, like ChatGPT or Gemini, runs on closed models: you send your question to the company’s servers and get an answer back. An open-weight model is different. Companies, universities and hobbyists can download it, run it on their own computers or cloud accounts, and fine-tune it for their own needs.

That can mean lower costs, more control over sensitive data and no dependence on a single vendor. It is also why many U.S. businesses have been quietly building on Chinese open models, which have led on quality. A competitive American option gives them another choice.

The catch

  • You can’t download it yet. Reflection says the weights, a technical report and a model card will arrive “later this month,” after final red-teaming and evaluations. Until then, access is limited to an early-access signup.
  • The numbers are the company’s own. Independent testers haven’t yet had a chance to verify the benchmarks, and Reflection notes its efficiency figures leave out some real-world serving costs, such as processing long prompts.
  • It isn’t the top performer. Reflection itself acknowledges that larger rivals such as Kimi K3 and Qwen 3.8-Max remain ahead on raw capability, and Beam trails closed frontier models.
  • Running it isn’t trivial. Even with only 23 billion active parameters, a 501-billion-parameter model needs serious server hardware. This isn’t something you’ll run on a laptop.
  • Safety details are pending. Reflection says its safety evaluations will be published in the technical report and that it plans to open-source its safety benchmarks so others can test the model.

What’s next

Reflection says Beam is the first in a series and that each release should bring open models closer to the cutting edge. The real test comes when the weights are public and developers can see how Beam holds up outside the company’s own charts.

More on Contoh

Sources