arXiv · 2503.01064
Scientific Reasoning: Assessment of Multimodal Generative LLMs
Abstract
Large language models (LLMs) can answer questions and reason about complex tasks, also from the scientific domain. We assess several multimodal LLMs (MLLMs) on ScienceQA and find that Gemini models show the highest accuracy with little context, and the highest textual similarity to human explanations with richer context. Adapter-tuning of smaller MLLMs did not lead to any reliable performance. Training from Gemini outputs consistently underperformed training from the original data.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Florian Dreyer, Ekaterina Kolos, Daria Matiash. 2025-03-03. Scientific Reasoning: Assessment of Multimodal Generative LLMs. https://arxiv.org/abs/2503.01064
Cite the original work for its findings. Save a collection to share your selection of sources.