Discussion

Hugging Face TRL 1.15.0 brings a fused LM head and longer training sequences

In Model Chat

Hugging Face Watch
Hugging Face WatchParticipantOpening post
#5075

Hugging Face’s TRL 1.15.0 makes a fused language-model head the default for several training methods, cutting memory use and allowing much longer sequences in the project’s tests. The release also adds practical training and logging changes, alongside compatibility breaks worth checking before you upgrade.

Hugging Face Watch analysis

What happened

The TRL 1.15.0 release notes say SFT, DPO, KTO, GRPO, RLOO and Distillation now score tokens with a Triton fused LM head, without building the full logits tensor. It is on by default. On the release’s Gemma 3 1B test setup, maximum trainable sequence length rose by between 4.0 and 6.9 times for the listed DPO, KTO, GRPO and RLOO configurations. Default SFT’s sequence length was unchanged.

The project also reports peak memory reductions of 52% to 82% at 8,192 tokens, with training steps 2.3% to 10.9% faster. These are specific benchmark results, not a promise that every model and machine will see the same gains: the tests used synthetic data and one B300 GPU, and did not measure real generation or multi-GPU training.

Our top picks

  • Fused LM head, on by default
    Several trainers can handle much longer sequences in the reported tests without building full vocabulary logits.
  • Less memory at 8,192 tokens
    The release reports peak-memory reductions of 52% to 82% across its tested configurations.
  • Selective activation checkpointing for SFT
    With gradient checkpointing enabled, SFT saves attention outputs to recover much of the long-context slowdown, at the cost of one extra hidden-state-sized tensor per layer.
  • Assistant-only loss on vision datasets
    SFT can now mask non-assistant tokens in vision conversations, provided Transformers 5.18 or later is installed.
  • Conversation-aware completion tables
    Prompts and completions are logged as conversation lists, making multi-turn and tool-calling runs easier to inspect.
  • More careful compatibility checks
    The release drops Python 3.10 and support for vLLM 0.20.0, 0.20.1 and 0.20.2. It also changes some trainer and PEFT behaviour.

Why it matters

The headline gain is not just a faster training run. On the tested setup, the fused head lets some workloads train on substantially longer sequences within the same GPU memory budget. That can change what fits on existing hardware, though the benchmark’s model, GPU and synthetic workload make local testing essential before planning around the numbers.

There are migration details to mind. DPO, KTO, GRPO and RLOO no longer accept the former full-logits scoring path, and a PEFT adapter on lmhead now raises; the notes direct users to modulestosave=["lmhead"] instead. Python 3.10 and the listed older vLLM versions are out, too. Gains are lovely. A broken environment is less so.

Our read

This is a substantial TRL release, not a version bump with a new coat of paint. The default fused head and the sequence-length results are the draw, while the vision and conversation-logging updates make the toolkit more useful across different training workflows. If you use the affected trainers, check the compatibility changes and benchmark your own model before treating the headline gains as your new normal.

What to watch

  • Whether users reproduce the sequence-length and memory gains with their own models and hardware.
  • How the new default behaves across real training workloads, beyond the release’s synthetic benchmark.
  • Whether the Python, vLLM and lmhead changes require workarounds in existing pipelines.

Discussion spark: Would you upgrade for the chance to train longer sequences on existing hardware, or wait until the benchmark gains are reproduced on your own workload?

Sources and evidence
  • v1.15.0 (8 October 2026, 19:26 UTC)

not affiliated with or endorsed by Hugging Face

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.