Home Issues Editorial Board News
Submit manuscript
Losharu Journal of Computational Intelligence Volume 3, Issue 2 Research Article
Loshu Comput. Intell.
Losharu Journal of Computational In...
Losharu Journal of Computation...
Date: April 2026
Article: ljci.2026.0202
Published by
Research Article Full text access Get rights and content ↗

Multimodal Foundation Models for Scientific Literature Understanding: A Comprehensive Evaluation

Article
Recommended articles
Cited by 15
Metrics
Highlights
  • Scientific literature increasingly contains complex multimodal content—figures, tables, equations, molecular diagrams—that unimodal language models cannot fully process.
  • We evaluate 12 multimodal foundation models (including GPT-4V, Gemini Ultra, LLaVA-1.6, and SciPhi) on a newly constructed benchmark, SciMMU, comprising 8,400 question-answer pairs spanning figure interpretation, table reasoning, equation understanding, and cross-modal synthesis across chemistry, biology, and physics.
  • GPT-4V achieves highest overall accuracy (71.3%), but all models show substantial weaknesses in equation-grounded reasoning (average 42.7%) and cross-paper synthesis tasks (average 38.1%).
Abstract
Scientific literature increasingly contains complex multimodal content—figures, tables, equations, molecular diagrams—that unimodal language models cannot fully process. We evaluate 12 multimodal foundation models (including GPT-4V, Gemini Ultra, LLaVA-1.6, and SciPhi) on a newly constructed benchmark, SciMMU, comprising 8,400 question-answer pairs spanning figure interpretation, table reasoning, equation understanding, and cross-modal synthesis across chemistry, biology, and physics. GPT-4V achieves highest overall accuracy (71.3%), but all models show substantial weaknesses in equation-grounded reasoning (average 42.7%) and cross-paper synthesis tasks (average 38.1%). We release SciMMU as a community resource for evaluating AI scientific comprehension.
Keywords
From Same Issue
Neural Architecture Search for Edge Computing: Efficient Models for Re...
Wei, L., Zhang, F., Liu, Y., Sun, Q., Ha...