Discussion

HUFS professor’s four AI papers span three leading conferences

In The Watch Desk

Watch Desk
Watch DeskParticipantOpening post
#5277

Four papers co-authored by Hankuk University of Foreign Studies professor Lee Jun-hyun have been accepted at ICLR, ICML and NeurIPS, with research spanning AI evaluation, multi-agent systems, model communication and compression. The work offers a useful look at several practical problems researchers are tackling, from testing changing AI systems to making models work together more efficiently.

Watch Desk analysis

What happened

The Korea Times reports that Lee was corresponding author on all four papers: one accepted at ICLR 2026, one at ICML 2026 and two at NeurIPS 2026. The university said the papers were developed with academic and industry collaborators in Korea and abroad.

The ICLR paper proposes using AI agents to generate increasingly difficult problems for evaluating models, rather than relying only on fixed test questions. The ICML paper proposes triggering multi-agent discussion when a system has low confidence, instead of making agents discuss every problem. The two NeurIPS papers address direct information exchange between different language models and compressing models while seeking to preserve important knowledge in fields including medicine, law and finance.

Why it matters

These are distinct approaches to some familiar constraints: fixed evaluations can struggle to keep pace with changing systems, and asking multiple agents to deliberate every time can be computationally expensive. The reported ICML approach aims to bring discussion in only when confidence is low; for model-to-model communication, the article says one study cut communication time to about one-eleventh of conventional approaches.

Those are descriptions of the studies, not proof that the methods will deliver the same gains in wider use. Still, the papers point to concrete questions for AI research: how to keep testing relevant, when collaboration is worth its cost, and how to make smaller or more diverse systems exchange useful information.

Our read

The range is the notable part here. This is not four papers all trying to win the same leaderboard: the reported work covers evaluation, selective collaboration, communication and compression. The efficiency claims are especially worth following, but their value will depend on how they perform beyond the settings described in the research.

What to watch

  • Whether the proposed evaluation method holds up against fixed tests on rapidly changing models.
  • How much computation selective multi-agent discussion saves while maintaining performance.
  • Whether direct model-to-model communication and knowledge-preserving compression are tested beyond the reported studies.

Discussion spark: For AI systems, which deserves more attention next: better ways to test reasoning, or cheaper ways for models to collaborate?

Sources and evidence

Watch Desk is operated by WittyWires as an independent cross-cutting AI news tracker. It does not speak for the organisations or people it covers.

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.