MiniMax released M2.7 and M2.7-highspeed on 18 March 2026, describing M2.7 as the first of its models to participate deeply in its own development. The interesting part is narrower than the slogan: MiniMax documented an internal M2.7 agent changing a programming scaffold, evaluating each attempt and choosing whether to keep it.
MiniMax Watch analysis
What happened
According to MiniMax, its research teams used M2.7 to maintain memory, build skills for reinforcement-learning experiments and support data pipelines, training environments, debugging, code changes and smoke tests. Researchers still supplied guidance and made critical decisions; the company estimated that the agent handled 30 to 50 percent of that workflow.
In a more tightly bounded test, an internal version ran more than 100 autonomous rounds on a programming scaffold. It analysed failures, adjusted sampling settings and workflow instructions, added loop detection, reran evaluations, then kept or reverted changes. MiniMax reported a 30 percent improvement on internal evaluation sets. Separate official release notes confirm the 18 March launch and the two model variants, but do not validate that result.
Why it matters
That distinction matters. This was not a model independently redesigning and retraining its own weights from first principles. The reported system improved the harness around a model and parts of the research workflow, inside objectives and evaluations chosen by people. Even so, agents that can propose, test and reject changes could compress the dull but expensive cycle between an experiment and a useful engineering decision.
Our read
‘Self-evolution’ is the sort of phrase that can arrive wearing a lab coat and leave with the silverware. MiniMax deserves some credit for describing a concrete loop rather than offering only incense. The proof is still first-party: no public internal evaluation set, intervention log or reproducible run accompanies the 30 percent figure, so it should be read as a documented company result, not independent confirmation.
What to watch
- Whether MiniMax publishes reproducible scaffold experiments or evaluation traces.
- How often humans redirected, approved or rescued the claimed autonomous runs.
- Whether improvements transfer beyond one internal scaffold and its chosen tests.
- What portion of later model training, rather than harness tuning, agents can safely control.
Discussion spark: When a lab calls model-assisted training self-evolution, what minimum level of autonomous hypothesis, experimentation and validation should that term require?
Sources and evidence
- MiniMax M2.7: Early Echoes of Self-Evolution (18 March 2026, 00:00 UTC)
- MiniMax API model release notes (18 March 2026, 00:00 UTC)
not affiliated with or endorsed by MiniMax