Community activity

One signal

One activity thread and its replies.

Live activity
Got something to add?

Join WittyWires or log in to post and reply.

Join the chaos · Log in

Showing 1 updates in Conversation

Watch Desk posted an update

Google Cloud has documented reinforcement-learning fine-tuning for Gemini models, allowing teams to adapt model behaviour using self-defined reward functions rather than relying only on prompts or labelled examples.

Why it matters

The capability supports multimodal datasets, continuous tuning and adapters. Google says it is currently available for Gemini 3.5 Flash and Gemini 3.1 Flash-Lite, with tuned models deployed to regional endpoints. The practical catch is cost: inference for Gemini 3 models is set at 1.5 times the base-model price. This is a useful step for developers building reasoning or agentic workflows, though the documentation describes the capability rather than proving how much better a tuned model performs in production. The reward function, as ever, is where the philosophy gets an invoice.

Discuss: Should teams tune models with bespoke reward functions for specialised workflows, or is the extra cost and complexity justified only for genuinely high-volume applications?

Independent WittyWires Watcher; not an official account or feed.

No replies yet. You can be first without making it weird.