Bolmo Reads Bytes. The Tokenizer Stayed in Stage 1

Nature published Ai2’s byteification paper on Oct 7. Bolmo 7B beat BLT 7B by 16.5 points on STEM. Bwen 8B is the stronger retrofit. Character tasks are the reason to care.

Workbench with a printed tokenizer vocabulary sheet beside a hex dump of the same sentence in bytes, cool lab light, no logos

Nature put the Bolmo paper out on October 7: Retrofitting language models to operate over bytes. Ai2’s same-day blog is the lab note. Byteification is a two-stage conversion. You start with a subword model you already trained. You do not throw the weights in the trash and pretrain on raw bytes from scratch. That is the claim that matters. “Bytes instead of tokens” as a slogan has been around for years. The paper is about a retrofit that, in their tests, finally sits near the source model.

We already treated omission in generated reports as a different evaluation failure. This week is the input side. If your benchmark is character-level, the tokenizer was part of the score. If your benchmark is MMLU-style STEM, compare Bolmo to BLT, not to a closed chat model you did not run.

What they converted

Ai2 introduced Bolmo last December as a fully open byte-level family built from Olmo. The Nature paper is that method in a journal. The blog adds checkpoints: the same conversion applied to Qwen 3 8B and Llama 3 8B, producing Bwen 8B and Blama 8B. Both, Ai2 says, come close to the models they were derived from. Bwen 8B is their strongest byteified model so far and beats Bolmo 7B on the aggregate suite.

They also released Stage 1 checkpoints. In that stage the original global model stays frozen while the new byte-level components train. Researchers who want to poke the architecture without paying for a full second stage get a shorter on-ramp. That is a packaging decision as much as a scientific one. Frozen trunks are how you keep the source ecosystem: the same downstream heads, the same evaluation harness, a different front-end.

The Nature abstract’s procedure is two-stage conversion with “minimal extra training.” Stage 2, in the FLOPs note we pulled, is about three times the global model’s compute (forward plus backward). An ablation that skipped stage 1 added 6.5 billion tokens to stage 2, lengthening it about 17 percent. Do not round that into “cheap.” Do round it into “cheaper than pretraining BLT from random initialization,” which is the baseline they actually beat.

Bolmo 1B and Bolmo 7B, from Olmo, are what Ai2 calls the first fully open byte-level language models competitive with strong subword models across a broad task range. “Fully open” is their label. “Competitive” is on their suite. Read both as claims with a leaderboard attached, not as a funeral for BPE.

The 16.5 points are against BLT

Nature: Bolmo 7B achieved a +16.5 percent absolute improvement in STEM tasks over BLT 7B, which was trained from random initialization. That sentence has two objects. Bolmo is a retrofit of Olmo 3. BLT is a from-scratch byte model. A 16.5-point STEM gap is evidence that starting from a trained subword trunk helped. It is not evidence that Bolmo 7B beat every 7B subword model on STEM. The paper also says Bolmo 7B greatly outperformed source Olmo 3 on character understanding, and was better in certain coding settings.

Those are the two axes you should keep separate. Character understanding is where subword tokenization is structurally awkward: spaces, typos, overlapping morphemes, code identifiers, scientific strings. STEM multiple choice is where a strong trunk already knew a lot. If you only quote the STEM number, you are advertising the inheritance. If you only quote character tasks, you are advertising the point of bytes.

Bwen 8B outperformed Bolmo 7B and, Nature says, reached performance close to and sometimes surpassing the source Qwen model. That “sometimes” is doing work. A retrofit that matches the parent on average and wins on character tasks is a different product than a retrofit that loses 4 points of MMLU to save a tokenizer. Ai2’s blog is more promotional on Bwen (“strongest byteified model yet”). Nature’s wording is the one to keep for the parent comparison.

Eval hygiene in the figure caption we saw: 95 percent confidence interval on a core-task macro average from binomial sampling error, n = 53,879 evaluation instances. Each plotted point is a single 150,000-step training run. Single runs are still single runs. The n is large for the tasks, not for the training seeds. If you reproduce, budget for that.

Why bytes were supposed to lose

Subword tokenizers compress English well and other scripts less well. They also freeze a vocabulary before you see the user’s mess. Byte-level models skip that freeze. They also, historically, lagged in quality and were expensive at the softmax because the sequence got longer. Nature’s efficiency figure argues the subword LLM becomes Pareto-dominated as softmax cost grows with vocabulary size, and byte-level models become Pareto-optimal there. SuperBPE tokenizer transfer is in that figure as a subword comparison point. Compression targets versus attained compression are labeled. I am not going to pretend I reran the Pareto plot. The paper’s claim is that bytes can win the efficiency curve once vocabularies get large, if you patch bytes instead of attending every one.

They say you can further speed byteified models by training with higher ratios of bytes per patch, something subword models can only do in a limited way. Patching is the trick that keeps inference from being a per-byte transformer tax. A patch is a group of bytes the global model sees as one step. Raise the ratio and you shorten the transformer sequence. Raise it too far and you are back to a clumsy tokenizer you invented in the data loader. “Practical inference speeds” in the abstract is that knob, not magic hardware. If a checkpoint card omits the ratio, you cannot compare latency to Olmo 3.

Nature’s wrap-up lists compute and energy as potential advantages, along with less English-centric bias and fine-grained scientific text. Potential is doing a lot of work. Sequence length, patch ratio, and softmax size fight each other. If you have a 128k-vocab commercial model, their Pareto story is aimed at you. If you have a 32k-vocab 7B, measure. SuperBPE appears in the same efficiency figure as a tokenizer-transfer baseline on the subword side. It is there so they are not only fighting 2010s BPE. I did not rerun the plot. The paper’s claim is that bytes can win that curve once vocabularies get large.

A tokenizer trained on English-heavy web text carves other languages into awkward pieces. Bytes do not play favorites among Unicode. That does not make Bolmo a multilingual champion by default. It removes one source of bias. The Russian-to-English poetry fine-tune Ai2 mentions is an existence proof that someone already took Bolmo 7B into a character-sensitive job. It is not a BLEU table. Lyrics and verse care about rhyme, stress, and line breaks. Subword merges smear those. If you were going to pick a hobby fine-tune to show why bytes exist, verse is a fair one. It is still a hobby fine-tune. Do not put it in a product slide.

Scientific and technical strings are the third promised dividend: identifiers, formulas, codes, the stuff BPE splits into junk. Character-level reasoning tasks are where the paper says the models excel. If your product never sees those strings, you are buying a paper. If your product is code or OCR-adjacent text, you are in the test distribution. OCR output in particular is full of one-character errors. A subword model sees a new token. A byte model sees a flipped byte. That difference is the whole pitch, and it is also why people kept training byte models that lost on everything else.

What you should not do with this

Do not delete your tokenizer because Nature published. The method reuses the source model. The tokenizer is still in stage 1 as a teacher or as the frozen trunk, depending how you read the staging. The output is byte-level. The inheritance is subword.

Do not compare Bolmo 7B to GPT-5.5 on a blog chart. Compare it to Olmo 3, BLT 7B, Bwen 8B, and Blama 8B, which is the family the authors actually ran. We already warned about small open models getting over-read. Same discipline.

Do not treat Stage 1 checkpoints as production weights. They are, in Ai2’s description, a faster starting point with the global model frozen. Stage 2 is the expensive alignment of that trunk with the byte front-end. If you only train stage 1 and ship, you shipped a research scaffold.

Do not collapse “better in certain coding settings” into “Bolmo is the new coding model.” Certain is a hedge the paper earned. Run your repo. Character-level wins on code often show up in rename-the-variable and syntax-repair tasks. They may not show up in long-horizon agent evals. Those evals have other failures, including the reporting omission problem.

If you actually train

Start from the published Stage 1 checkpoint for the family you care about: Olmo for Bolmo, Qwen 3 8B for Bwen, Llama 3 8B for Blama. Do not mix trunks. The point of the retrofit is the source ecosystem. Mixing is a new paper.

Keep the character-level tasks in the eval. If those go up and STEM goes down, you learned something. If both go up, you matched the blog. If only STEM goes up, you may have just continued pretraining and called it byteification.

Log bytes per patch. That knob is the inference story. Hide it and you will not know why latency moved.

If you need a from-scratch byte baseline, BLT 7B is the one they used. Beating it by 16.5 STEM points is the inheritance result. Replicate that before you claim a new method.

Hugging Face will have more clones by next week. Check whether a card says Stage 1 or the full conversion. Check whether it is Bolmo, Bwen, or Blama. The names are ugly on purpose. Use them. “The Nature byte model” is how you download the wrong 8B. If the card cites last December’s Bolmo announcement and not the October 7 Nature version, you may be missing Bwen, Blama, and the Stage 1 dumps.

The authors write that the results remove a long-standing performance barrier to end-to-end byte-level language modelling. The engineering version is shorter. You can convert Olmo, Qwen 3 8B, and Llama 3 8B, and the character tasks are where the source models were leaving points on the table.