Baseten and Goodfire AI are developing Project Beacon, a system intended to detect risky model behaviour during generation and let applications respond before it reaches a user or tool. The notable shift is towards safety controls built into inference, rather than checks applied only after an answer is produced.
Baseten Watch analysis
What happened
Baseten says Goodfire is developing monitors for specific behaviours on supported models, using a model’s internal activations alongside text checks. Baseten will connect those signals to systems that can apply customer-defined policies. The company says an architecture diagram shows a frontier model weighing in when the activation and text monitors disagree; monitoring runs alongside generation and does not block it.
Depending on the policy, an application could request approval, refuse, fall back to another response or log an event for administrators. Baseten identifies prompt injection, unauthorised actions, sensitive-data exposure and cyber misuse as risks it wants to address. The companies plan to bring capabilities to market with a small number of early partners, beginning with selected models and monitored behaviours. Baseten says it expects to release further capabilities over the coming months. Read Baseten’s Project Beacon announcement.
Why it matters
As AI agents read documents, call tools and take multi-step actions, a check at the end of a workflow may arrive too late to prevent a consequential step. Project Beacon’s proposed approach is to detect selected behaviours during inference and give applications a policy-based way to respond. That could give teams a more central way to manage controls across models and workloads, rather than building separate filters for each application.
The practical scope is still limited: Baseten describes a planned system, not a generally available product, and its initial monitors will cover selected behaviours and supported models. The announcement provides no performance results for detecting unsafe behaviour.
Our read
The useful idea is the combination of behaviour monitoring with controls that can change what an application does next. Keeping monitoring parallel to generation could also avoid making every response wait for another review step. But a safety control is only as useful as the behaviours it can detect and the policy attached to them. This is worth watching as a product direction, not yet treating as a ready-made safety net.
What to watch
- Which models and behaviours are included in the first release.
- How customers can test detection quality, including missed alerts and false alarms.
- Whether Baseten publishes details on latency and how its policies act on monitor signals.
Discussion spark: Should AI safety controls be built into the inference platform and applied centrally, or should each application team remain responsible for its own safeguards?
Sources and evidence
- Source update (9 October 2026, 17:01 UTC)
not affiliated with or endorsed by Baseten