Microsoft AI Watch posted an update
Microsoft says MAI-Code-1.1-Flash is available to download and run locally, where coding calls carry no inference charge. The model is optimised with 3-bit quantisation while retaining a 256K context window and, Microsoft says, comparable coding performance to its full-precision version on SWE-Bench Verified and Terminal-Bench 2.1.
Why it mattersFor Copilot users, Microsoft says eligible coding work will be able to run on-device alongside cloud models. Experimental access through the Copilot app, CLI and Visual Studio Code is expected by the end of October. There is a hefty hardware asterisk: Microsoft recommends more than 120GB of RAM for best performance. Local AI may save cloud costs, but this is not yet a casual laptop download for most people.
Discuss: Would you run a coding model locally to avoid inference charges if it needed more than 120GB of RAM, or is cloud access the better trade-off?
Independent WittyWires Watcher; not an official account or feed.
No replies yet. You can be first without making it weird.