Two labs announced "open-weight" models at the trillion parameter scale this week, and as of this morning you cannot download either one. Reflection AI announced Beam on October 5: 501 billion parameters, 23 billion active, Apache 2.0, weights "later this month." Mistral followed on October 6 with Mistral Large 4: 1.05 trillion parameters, 49 billion active, weights "end of this month," license not published. Both ship an API today. Both ship the thing they're calling open at some later date.

I've watched this pattern harden over the past year. The announcement runs on the day the API goes live. The weights come later, if they come at all, and the license shows up with them. "Open weight" now describes a press release. The download link comes later.

What was announced, by the numbers

Mistral's model is a sparse mixture of experts with a 1.6 billion parameter vision encoder bolted on, a 1 million token context window, and coverage of more than 160 languages. Mistral trained it on 3,800 Nvidia Grace Blackwell GPUs in its own European datacenters. The preview API costs $1.36 per million input tokens and $4.18 per million output tokens at list, discounted by half during the preview. Mistral says the weights will ship in FP8 and FP4 formats and that the model fits on four to eight B200 or B300 GPUs at FP4.

Reflection's Beam is smaller and sparser. Pretraining ran on 23.8 trillion tokens across 6,144 GB300 GPUs in under four weeks. The reinforcement learning phase took another four weeks on 10,500 GB300s, generating more than 100 million rollouts inside roughly 1.3 billion sandboxes. Reflection reports 80.9 on SWE-Bench Verified, 80.1 on Terminal Bench 2.1, and 97.8 on AIME 2026. The company's claim is that Beam matches GLM 5.2, a model with about 250 billion more parameters, while using a third to a quarter of the inference hardware.

Reflection was founded in March 2024 by Misha Laskin and Ioannis Antonoglou, both out of Google DeepMind. It raised $2 billion at an $8 billion valuation in October 2025 in a round led by Nvidia. SiliconANGLE reports the company is now valued at $25 billion and has a $6.3 billion agreement with SpaceX to rent GB300 NVL72 racks. Beam is its first open-weight model. Or will be.

On quality, Simon Willison ran Mistral Large 4 through the preview API and put it at 38 on the Artificial Analysis index, just behind DeepSeek 4.1 Flash. Mistral Large 3 scored 9 on the same index in December 2025. His read is that Mistral is back to about six months behind the frontier. Mistral's own human evaluation, run by Surge AI, scored Large 4 at 3.74 out of 5 on coding against 4.22 for Claude Opus 5.

Total parameters are the hosting bill

Active parameter counts tell you about latency and cost per token. Total parameter counts tell you how much HBM you have to buy before the first token comes out. Sparse models decouple those two numbers, and the marketing leans on the small one.

Run the arithmetic on Large 4. At FP8, 1.05 trillion parameters is about 1.05 terabytes of weights before you allocate a byte for KV cache. A DGX B200 node has 1,440 GB of GPU memory across eight cards. So FP8 is one full node, and the KV cache for a 1 million token context eats much of what's left. At FP4 you're down to roughly 525 GB, which is where Mistral's "four to eight GPUs" claim comes from. Four B300s at 288 GB each gives you 1,152 GB. It fits. Whether the FP4 checkpoint holds up on your workload is a question nobody outside Mistral can answer until October 27.

Beam is a different hosting story. At FP8, 501 billion parameters is about 500 GB, which fits in an eight H100 node with 640 GB total. At BF16 you need an H200 node or two H100 nodes. That's hardware a lot of teams already own or can rent by the hour without a procurement cycle. With 23 billion active parameters per token, the compute per request looks like a 23 billion parameter dense model from 2024.

That difference matters more to me than the benchmark gap. A 1 trillion parameter model that fits on one Blackwell node is a model for people who have a Blackwell node. A 500 billion parameter model that fits on Hopper hardware is a model for people who have a datacenter from two years ago, which is most of us.

Nobody has read the license yet

Reflection says Apache 2.0, in writing, in the announcement. That is the right answer and I'll give them credit for it on the day the weights land.

Mistral has said nothing about the license. Its model page lists it as "Open" with no terms attached. VentureBeat reported the weights will ship October 27 under a custom Mistral license, which is the same shape as the Mistral Research License that covered Large 2 in 2024 and barred commercial use without a separate agreement. Large 3 shipped under Apache 2.0 in December 2025, so the company has gone both ways in the past two years.

For anyone planning a deployment, the license is the whole question. A research license on a 1 trillion parameter model means you can run it in a lab and nowhere else. An Apache license means you can put it in a product. Those are two different products, and Mistral has announced which one you're getting only by implication.

My prediction, which I'd be happy to get wrong: Large 4's weights ship under a license that forbids commercial use, with a sales contact for anyone who wants more. The 49 billion active parameter count and the European datacenter line are aimed at sovereign buyers and regulated industries, and those customers pay for contracts, not for Apache headers.

Shipping the API first is the tell

Both companies could have held the announcement until the weights were on Hugging Face. They didn't. The API went up on announcement day and the weights went into a queue three weeks long.

Part of that is practical. Quantized checkpoints need validation, model cards take time, and the inference stack integrations (vLLM, SGLang, llama.cpp) don't write themselves. Reflection says a technical report and fine-tuning tooling will accompany the weights, and that work is real.

But part of it is that the API is the business and the weights are the marketing. Three weeks of exclusive hosted access is three weeks of benchmark posts, leaderboard entries, and usage data, all of which point back at the API. By the time the weights arrive, the model has a reputation, and the open release reads as a bonus on top of a product that already exists.

I don't think either lab is lying. Reflection has a funding round and a GPU lease that depend on shipping something, and Mistral has a track record of following through on weights. But "open-weight model released" and "open-weight model announced, API available, weights in three weeks, license pending" are two different sentences, and most of the coverage this week used the first one.

If you run inference infrastructure, do nothing yet. Put October 27 on the calendar, check the Mistral license the hour it posts, and watch for Beam's Hugging Face repository to appear. Then size a node, pull the FP8 checkpoint, and run your own evals before you believe any of the numbers above, including mine.

If Beam ships under Apache 2.0 on Hopper hardware with the scores Reflection posted, it will be the most useful open model release of the year for self-hosters. If Large 4 ships under a research license, it will be a good API with a trillion parameters of press.