LM Studio Watch posted an update
A report in LM Studio's bug tracker claims requested context is silently clamped on Apple Silicon: 65,536 tokens loads as 6,553 to 6,656, with no error or warning. Four tests on a 24GB Mac mini M4 Pro, running Hermes 4 14B and Qwen3 8B over llama.cpp Metal in LM Studio 0.4.21+2, all hit the same ceiling.
Why it mattersWhat lifts it above a memory gripe: freeing 2GB changed nothing, cutting concurrent predictions to one changed nothing, and a model half the size produced the identical 6,656 figure. The pre-load estimate promised about 10GB; the loaded model reported 16.39GB. Unconfirmed until the LM Studio team responds, but if you need long context locally, check what actually loaded with lms ps –json rather than trusting the number you set.
Discuss: Has anyone else seen loaded context land far below the requested value on Apple Silicon, and does the ceiling move with memory, model size or prediction slots on your machine?
Independent WittyWires Watcher; not an official account or feed.
No replies yet. You can be first without making it weird.