Taiwanese platform uniopen adapted Amazon Nova 2 Lite to classify customer interactions against its own retail moderation rules, AWS says. The case study gives a practical example of fine-tuning improving both behaviour and subject classification, with a prompt change then pushing both measures above the company’s production targets.
AWS AI Watch analysis
What happened
Uniopen’s moderation system sorts interactions across nine behaviour categories and three subject types: brand, other and forbidden. AWS says the base model struggled with this business-specific taxonomy, so the team fine-tuned Nova 2 Lite using 3,391 labelled conversation windows and tested it on 737 held-out windows.
On that test set, the behaviour score rose from 0.5852 to 0.8364, while the subject-type score climbed from 0.4162 to 0.8302. A subsequent prompt change, which simplified the output format and clarified how to return multiple behaviours, raised the scores to 0.8550 and 0.8491 respectively. Uniopen’s production targets were 0.8500 and 0.8200. The AWS case study describes the results.
Key findings
- Fine-tuning addressed a specific gap
The largest gains came from training on examples of uniopen’s own moderation policy. - A prompt change added gains without more training
The output-format adjustment lifted both scores above their stated production targets. - Two measures acted as release gates
Behaviour and subject classification were checked separately, so strength in one could not hide weakness in the other. - People remain in the correction loop
Human reviewers verify candidate corrections before they enter the training set, and review ambiguous cases.
Why it matters
A general-purpose model does not automatically understand the categories a particular business needs it to apply. This example shows a route from broad capability to a narrower operational task: measure the model against the actual policy, improve the weak spots, and require both dimensions to clear release gates.
For teams considering a similar system, the useful detail is not simply that fine-tuning helped. Uniopen reports that a prompt-level adjustment also made a measurable difference, while its workflow keeps human checks for uncertain cases and proposed training corrections.
Our read
The strongest lesson is the disciplined evaluation, not a promise that this recipe will work unchanged elsewhere. Uniopen tested against its own policy and thresholds, and AWS reports the resulting scores; other businesses would need their own representative data and release gates. Still, it is a refreshingly concrete account of how a model can be adapted without treating generated corrections as ground truth.
What to watch
- Whether later tests show the scores hold up on new boundary cases from real traffic.
- How uniopen’s taxonomy and thresholds change as its moderation policy evolves.
- Whether the same evaluation pattern transfers to other retail moderation tasks.
Discussion spark: For a business-specific moderation system, should a model be allowed to handle routine cases automatically once it clears measured release gates, or should human review remain the default?
Sources and evidence
- How uniopen customized Amazon Nova to their retail moderation policies for production deployment (1 October 2026, 15:33 UTC)
not affiliated with or endorsed by Amazon Web Services (AWS)