Multimodal Foundation Models for Scientific Literature Understanding: A Comprehensive Evaluation
Santos, R., Oliveira, L., Ferreira, M., Costa, A.. Multimodal Foundation Models for Scientific Literature Understanding: A Comprehensive Evaluation. Loshu Comput. Intell..
Vol.3, No.2. Apr 2026. https://doi.org/10.58921/ljci.2026.0202
Article
Recommended articles
Cited by 15
Metrics
Highlights
- Scientific literature increasingly contains complex multimodal content—figures, tables, equations, molecular diagrams—that unimodal language models cannot fully process.
- We evaluate 12 multimodal foundation models (including GPT-4V, Gemini Ultra, LLaVA-1.6, and SciPhi) on a newly constructed benchmark, SciMMU, comprising 8,400 question-answer pairs spanning figure interpretation, table reasoning, equation understanding, and cross-modal synthesis across chemistry, biology, and physics.
- GPT-4V achieves highest overall accuracy (71.3%), but all models show substantial weaknesses in equation-grounded reasoning (average 42.7%) and cross-paper synthesis tasks (average 38.1%).
Abstract
Scientific literature increasingly contains complex multimodal content—figures, tables, equations, molecular diagrams—that unimodal language models cannot fully process. We evaluate 12 multimodal foundation models (including GPT-4V, Gemini Ultra, LLaVA-1.6, and SciPhi) on a newly constructed benchmark, SciMMU, comprising 8,400 question-answer pairs spanning figure interpretation, table reasoning, equation understanding, and cross-modal synthesis across chemistry, biology, and physics. GPT-4V achieves highest overall accuracy (71.3%), but all models show substantial weaknesses in equation-grounded reasoning (average 42.7%) and cross-paper synthesis tasks (average 38.1%). We release SciMMU as a community resource for evaluating AI scientific comprehension.