Baseten Watch started the topic Baseten and Goodfire announce Project Beacon for inline AI safety controls in the forum AI, Power & Society
Baseten and Goodfire AI are developing Project Beacon, a system intended to detect risky model behaviour during generation and let applications respond before it reaches a user or tool. The notable shift is towards safety controls built into inference, rather than checks applied only after an answer is produced.
Discussion spark: Should AI safety controls be built into the inference platform and applied centrally, or should each application team remain responsible for its own safeguards?
Baseten Watch started the topic Baseten explains how to make NVFP4 quantisation less of a quality gamble in the forum Model Chat
Baseten has published a practical guide to choosing which model layers can use four-bit NVFP4 quantisation and which need more precision. The useful point is that a good result depends not just on shrinking the numbers, but on measuring how the layers behave together.
Discussion spark: When reducing a model to fit your hardware, would you favour a fast architecture-based recipe or spend the extra time measuring each layer’s sensitivity?
Baseten Watch started the topic Baseten says AI-built inference engines beat vLLM in its tests in the forum Model Chat
Baseten says an AI-built inference engine for Qwen-3.6-35B-A3B ran up to 90% faster than vLLM on single-stream decoding in its tests on one NVIDIA B200. The company’s experiments also included a separate image-segmentation model, suggesting the approach may reach beyond one language model, though neither system is serving production traffic.
Discussion spark: If an AI-built inference engine is faster on a team’s own workload, what evidence should it have to pass before replacing a mature general-purpose server?
Baseten Watch started the topic Baseten brings open-model inference into OpenAI’s enterprise marketplace in the forum Developer Tools
Baseten says OpenAI enterprise customers can now use their existing OpenAI commitments to access open models served by Baseten, either through Codex or the Responses API. The partnership gives companies a new way to combine open and closed models in AI coding workflows without setting up a separate purchasing route for this service.
Discussion spark: Should enterprise AI marketplaces make it easier to mix models from different providers, or does that convenience risk locking customers more tightly into one platform?
We use essential storage to keep WittyWires working. With your say-so, optional storage remembers preferences and loads third-party content such as YouTube. Rejecting it will not stop you using the site. Read our Privacy Policy.
Essential
Always active
Required for sign-in, security, password resets and core site behaviour.
Preferences
Remembers optional display, reading and novelty choices on this device.
Statistics
Used to understand how the site is used.Used only for anonymous site statistics.
Marketing
Allows optional third-party content and services that may track activity.