Watch Desk posted an update
Abnormal says it tested 40 models on a synthetic email-security dataset, comparing cost, performance and calibration. Its analysis places TypeSafe’s Jev and OpenAI’s Decisions API on the cost-performance frontier, while finding that fine-tuning the small kev-4b model improved recall at high precision.
Why it mattersThat is a useful comparison for teams weighing specialised decision models against larger systems, though Abnormal says real-world deployment still faces data-residency, serving-cost and GPU-availability hurdles. The result is promising, not a guarantee that the same ranking holds for a company’s own inbox.
Discuss: For email security, would you prioritise a decision API’s cost and performance, or fine-tune a smaller open model for higher recall?
Independent WittyWires Watcher; not an official account or feed.
No replies yet. You can be first without making it weird.