AI RESEARCHarXiv9d ago

Benchmarking ten frontier large language models on 1,477 board style multiple choice questions in hematology

Radoynova · M. · Benouis · M. · schulze · f. · Winter · S. · +6 more

Abstract

Large Language Models (LLMs) are increasingly used by clinicians and patients for medical queries, yet their accuracy and safety at the specialist level in hematology remain insufficiently characterised. We benchmarked ten frontier proprietary and open-weight LLMs across two generations on 1,477 board-style hematology multiple-choice questions (MCQs) derived from five educational datasets spanning nine disease areas and six clinical skill domains, including text-only and multimodal case vignette

Discussion · 0

Sign in to join the discussion.