Searcharxiv⌕ Search

arXiv · 2609.30489

BioEVAL: A global, multi-institutional benchmark of large language and multimodal models for bioengineering

Abstract

Large Language Models (LLMs) have demonstrated historic breakthroughs in general reasoning with early successes in biomedical science. However, existing LLM benchmarking emphasizes factual recall, offering limited insight into model performance on frontier and multimodal tasks. We assembled BioEVAL (BioEngineering Validation of AI and LLMs), a global, multi-institutional initiative designed to assess experimental reasoning capability across bioengineering (BE) subfields. BioEVAL spans 11 major BE subfields plus a set of uncategorized items, bringing together 22 research groups to create a PhD-level benchmark comprising 608 evaluation items: 1) 380 multiple-choice questions (MCQs, 359 retained after audit), 2) 218 literature synthesis tasks, and 3) 10 multimodal problems with experimental image interpretation. Benchmark items underwent authoring-group expert review and centralized quality control before evaluation. Following evaluation, a blinded cross-group consensus audit of the highest- and lowest-accuracy MCQ items flagged 21 questions for revision or removal; these were withheld, and all reported MCQ results are computed on the 359 retained items. We evaluated diverse cloud-scale foundation/multimodal models (e.g., ChatGPT, Gemini, and Grok) and locally deployable models suitable for inference on consumer-grade GPUs. Models achieved the highest accuracy of up to 90% on MCQs, similarity score of 0.72 on literature synthesis, and accuracy of 80% on a small sample of multimodal reasoning questions, with substantial performance variation across subfields. Leaderboard rankings characterize current capabilities, limitations, and development priorities across the evaluated BE task categories. BioEVAL is maintained as an extensible benchmark with standardized protocols for continuing expert item contribution and model evaluation.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Shun Ye, Vinny Chandran Suja, Chenlong Li, Chongming Jiang, Reza Zamani, Xiang Li, Christopher Bain, Yuqi Zhou, Walker Peterson, Huidong Wang, Chenglang Hu, Jongchan Park, Xiao Cheng, Benjamin Swedlund, Sandra Murillo, Anjali Sivanandan, Shiyu Sun, Liang Lanfeng, Mohammad Tariqul Islam, Baju C. Joy, Ishaq N. Khan, Sreedhar S. Kumar, Gabriel Mercado-Vásquez, James V. Vizzard, Jonathan M. Matthews, Helen Huang, Xiaolu Guo, Ethan Nicklow, Guorui Chen, Ryan A. Neff, Surjendu Maity, Hyeonjin Park, Han-ho Joo, Katherine Dong, Yuyan Cai, Weihang Huang, Yichen Zou, Rui Yan, Raphael Figueroa, Artem Goncharov, Bella Rose Schremmer, Lian Elsa Linton, Keisuke Goda, Liang Gao, Ke Cheng, Leonardo Morsut, Jennifer L. Wilson, Jianping Fu, Lim Chwee Teck, Deblina Sarkar, Andreas Hierlemann, Savaş Tay, Alexander Hoffmann, Donald Richieri Griffin, Jun Chen, Shana O. Kelley, Shyni Varghese, Jinwoo Cheon, Wilbur A. Lam, James J. Moon, Wilson W. Wong, Samir Mitragotri, Dino Di Carlo. 2026-09-24. BioEVAL: A global, multi-institutional benchmark of large language and multimodal models for bioengineering. https://arxiv.org/abs/2609.30489

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Hierarchical Reasoning Model

Reasoning, the process of devising and executing complex goal-oriented action sequences, remains a critical challenge in AI. Current large language models (LLMs) primarily employ Chain-of-Thought (CoT) techniques, which suffer from brittle task decomposition, extensive data requirements, and high latency. Inspired by the hierarchical and multi-timescale processing in the human brain, we propose the Hierarchical Reasoning Model (HRM), a novel recurrent architecture that attains significant computational depth while maintaining both training stability and efficiency. HRM executes sequential reasoning tasks in a single forward pass without explicit supervision of the intermediate process, through two interdependent recurrent modules: a high-level module responsible for slow, abstract planning, and a low-level module handling rapid, detailed computations. With only 27 million parameters, HRM achieves exceptional performance on complex reasoning tasks using only 1000 training samples. The model operates without pre-training or CoT data, yet achieves nearly perfect performance on challenging tasks including complex Sudoku puzzles and optimal path finding in large mazes. Furthermore, HRM outperforms much larger models with significantly longer context windows on the Abstraction and Reasoning Corpus (ARC), a key benchmark for measuring artificial general intelligence capabilities. These results underscore HRM's potential as a transformative advancement toward universal computation and general-purpose reasoning systems.

cs.AI↗

A memory-based active inference model of DishBrain-like adaptive behaviour

Recent and rapid advances in artificial intelligence (AI) make it increasingly important to understand the foundations of adaptive behaviour in autonomous agents, especially for building safe and efficient systems. While artificial neural networks have dominated the development of AI, recent work has begun to explore living biological neuronal networks as an alternative substrate for computation. These systems promise remarkable data and sample efficiency and rich dynamics, and may also inspire explainable and biologically plausible models. Here, we develop an experiment-informed active inference framework to model decision-making in closed-loop agents that mirror experimental setups using biological neurons. Using a generative model whose dimensions are matched to an experiment protocol, we systematically compare three decision-making schemes within this common generative model. Under matched episode counts (i.e. total data available for learning) to the in-vitro experiment, our simulations show that agents with short memory horizons reach a level of performance close to that of mouse and human cortical cultures (DishBrain platform), whereas longer memory horizons depart from it substantially. Increasing the planning horizon, by contrast, confers no comparable benefit. Because all model parameters are explicit, we can also track the quantities in our generative model that accompany this improvement, such as the risk term and the entropy of the transition and state-action mappings. Together, these results illustrate how active inference offers a formal language for comparing decision-making schemes in similar closed-loop control environments.

cs.AI↗

A Quantitative Study of Sustained Focus in Large Language Models via Repetitive Deterministic Prediction Tasks

We investigate the performance of large language models (LLMs) on repetitive deterministic prediction tasks and study how the sequence accuracy rate (SAR) scales with output length. Each such task involves the repetition of the same operation $N$ times. Examples of such tasks include letter replacement in letter strings following a given rule, integer addition, and multiplication of string operators in many-body quantum mechanics. If the LLM performs the task by a simple repetition algorithm, the success rate would follow an exponential decay with sequence length. In contrast, our experiments on leading LLMs reveal a crossover that is sharper than exponential: $-\log\mathrm{SAR}$ grows super-linearly with $N$, and accuracy collapses around a characteristic length $N_*$, the accuracy cliff that separates reliable from unreliable generation. The hypothesis of independent per-step errors is rejected for every model and task we studied. The crossover is well described by a double-exponential accumulation law, $\mathrm{SAR}=\exp(-β_0 Nα^{N-1})$, whose crossover scale $N_*$ does not depend on the functional form chosen to fit it. To interpret this behaviour we introduce a minimal effective model in which step-correctness variables interact through dense random couplings and compete with an external field set by the prompt. Solved by direct enumeration, the model reproduces the super-linear error accumulation and the accuracy cliff qualitatively, and it assigns to each model--task pair two interpretable parameters, an intrinsic error rate and an error-accumulation factor.

cs.AI↗