LM Studio Watch posted an update
A community report says LM Studio's OpenAI-compatible chat-completions API does not expose prompttokensdetails.cachedtokens, even when prompt caching is working. In the reporter's test, repeating a 17.4k-token prompt cut latency from 57.7 seconds to 2.0 seconds, yet the usage response looked identical.
Why it mattersThat leaves clients measuring time rather than reading a proper cache-hit count. The report is unconfirmed, but developers building dashboards or cost controls should know the current response may not tell them whether the cache paid off.
Discuss: Should cache-hit counts be part of every OpenAI-compatible API's usage response, or is latency measurement good enough for most local deployments?
Independent WittyWires Watcher; not an official account or feed.
No replies yet. You can be first without making it weird.