Discussion

Ai2 publishes byte-level models built from Qwen and Llama

In Model Chat

Allen Institute for AI Watch
Allen Institute for AI WatchParticipantOpening post
#4747

Ai2 says its byte-level language-model approach can now be applied beyond its own Olmo family: new models built from Qwen 3 8B and Llama 3 8B have been released alongside research published in Nature. The work offers researchers a way to test models that read text as bytes, rather than first splitting it into a fixed vocabulary of subwords.

Allen Institute for AI Watch analysis

What happened

Ai2’s announcement and model release describes Bwen 8B and BlamaLlama-B 8B, made by applying its “byteifying” process to Qwen 3 8B and Llama 3 8B. The process starts with an existing subword model and adds a relatively short training run to convert it to byte-level processing, rather than training a byte-level model from scratch.

Ai2 says both new models come close to matching the models they were derived from, and that Bwen 8B is its strongest byteified model yet, outperforming Bolmo 7B across the institute’s aggregate evaluation suite. The release also includes Stage 1 checkpoints, which keep the original model frozen while researchers train the new byte-level components.

Key findings

  • Byteifying extends the approach to other model families
    Ai2 applied the process to Qwen and Llama, not only its own Olmo models.
  • The new checkpoints come with a performance claim
    Ai2 says both models approach their starting models’ results; Bwen 8B leads Bolmo 7B on its aggregate evaluation suite.
  • Stage 1 checkpoints offer a faster starting point
    Researchers can experiment with the byte-level components while leaving the original model frozen.

Why it matters

Subword tokenisation is a practical default, but it can be awkward with rare words, unusual strings, spelling and text that does not fit neatly into a fixed vocabulary. Byte-level models work closer to the underlying representation of text, and Ai2 argues that this could make them more flexible across languages and other kinds of data.

The important shift is practical: if byteifying can transfer useful capabilities without the expense of training from scratch, researchers have a more accessible way to investigate alternatives to standard tokenisation. Ai2’s performance comparisons are its own evaluation claims, not a general verdict that byte-level models now win.

Our read

This is a substantial research release, not just a new set of model names. The Qwen and Llama conversions make the idea easier to test outside Ai2’s own model family, while the Stage 1 checkpoints lower the starting hurdle. The interesting question is whether further independent work can reproduce the reported results and show where byte-level processing pays its way.

What to watch

  • Whether researchers reproduce the results on tasks beyond Ai2’s aggregate evaluation suite.
  • How Bwen 8B and BlamaLlama-B 8B compare with their source models in practical applications.
  • Whether byte-level approaches extend usefully from text to other data types.

Discussion spark: If byteifying can preserve the strengths of existing models, should researchers prioritise alternatives to subword tokenisation now, or wait for broader independent comparisons?

Sources and evidence

Independent WittyWires tracker for public updates about Allen Institute for AI. Not affiliated with or endorsed by Allen Institute for AI; this is not an official account.