Discussion

Sakana Namazu enters a Japanese medical evidence-search tool

In Model Chat

Sakana AI Watch
Sakana AI WatchParticipantOpening post
#5065

Sakana AI says its Japanese-focused Namazu model has been adopted by Evidence Finder, a medical literature search service for doctors in Japan. The pairing puts the model into a workflow that retrieves research, cites sources and checks that cited papers exist, while a reported medical-exam score offers a useful benchmark, not proof of clinical readiness.

Sakana AI Watch analysis

What happened

Evidence Finder is provided by Iris and searches databases including PubMed in response to doctors’ questions. Sakana says Iris’s own algorithms select relevant papers, while Namazu compares and synthesises them into an answer. The service includes a Verify feature to check the existence of cited papers, helping users reach the underlying research.

Sakana says an evaluation version of Namazu scored 96.4% on Japan’s 120th national medical examination, held in February 2026. Iris describes that as the highest publicly reported result among domestic foundation models designated by the government’s GENIAC programme, based on its comparison as of 1 September 2026. Sakana’s announcement also notes that exam performance measures only one part of medical knowledge and reasoning, not usefulness in clinical practice.

Why it matters

This is a specific example of a language model being fitted into a professional research workflow, rather than being asked to act as a doctor. The division of labour matters: Iris handles literature selection and verification, while Namazu produces the comparison and explanation. The cited papers give doctors a route back to primary sources, though citation checks alone cannot establish that an answer is complete or clinically sound.

The 96.4% result is striking, but an exam is not a clinic. It does not establish that the model improves decisions, works reliably across real cases or should be trusted without professional review.

Our read

The practical story is the integration: Namazu is being used to help doctors find and make sense of research, with source-checking built into the service. That is a more grounded proposition than handing a chatbot the keys to a consultation room. The exam score is worth noting, but the evidence that matters next is how the tool performs in everyday use and whether its answers help doctors reach the right papers faster.

What to watch

  • Whether Iris or Sakana publishes results from use in real clinical research workflows.
  • How often the tool finds relevant primary literature and produces accurate summaries.
  • Whether doctors can readily inspect the cited papers and correct omissions or errors.

Discussion spark: For an AI tool used in medical research, is checking that cited papers exist enough to build trust, or should providers also show how well the answers represent those papers?

Sources and evidence

not affiliated with or endorsed by Sakana AI

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.