BioEVAL benchmark records 90% MCQ accuracy on 359 retained items across 11 bioengineering subfields
BioEVAL provides the first multi-institutional, PhD-level benchmark for LLM and multimodal performance on bioengineering tasks. Results show 90% MCQ ceiling with substantial subfield gaps after rigorous audit. The framework supports continued expert contribution and standardized model tracking.
BioEVAL assembled 380 MCQs, 218 literature synthesis tasks, and 10 multimodal items spanning 11 bioengineering subfields. After centralized review and blinded cross-group audit, 21 MCQs were removed, leaving 359 for scoring. Cloud-scale models including ChatGPT, Gemini, and Grok plus local GPU-deployable models were tested under standardized protocols.
Highest reported scores reached 90% on MCQs, 0.72 similarity on literature synthesis, and 80% on the multimodal subset. Performance varied sharply by subfield, exposing gaps in experimental reasoning versus factual recall. The audit process flagged items where model errors clustered, indicating systematic weaknesses in frontier bioengineering tasks.
Existing biomedical benchmarks emphasize recall; BioEVAL shifts focus to PhD-level experimental reasoning and image interpretation. Multi-institutional authorship and extensible contribution protocol reduce single-group bias while enabling ongoing item addition. Operational deployment of top models remains limited by subfield variance and small multimodal sample size.
Next evaluation round will expand multimodal items and enforce version-controlled model submissions to track incremental gains against the 90% MCQ ceiling.
BioEVAL maintainers: multimodal accuracy exceeds 85% on expanded 50-item set by Q3 2027
Sources (2)
- [1]Primary Source(https://arxiv.org/abs/2609.30489)
- [2]Supporting Source(https://arxiv.org/abs/2009.03300)