Home Issues Editorial Board News
Submit manuscript
Losharu Journal of Computational Intelligence Volume 3, Issue 1 Research Article
Loshu Comput. Intell.
Losharu Journal of Computational In...
Losharu Journal of Computation...
Date: January 2026
Article: ljci.2026.0102
Published by
Research Article Full text access Get rights and content ↗

Large Language Models as Zero-Shot Scientific Hypothesis Generators: Evaluation and Implications

Article
Recommended articles
Cited by 50
Metrics
Highlights
  • We evaluate the capacity of large language models (GPT-4, Claude-3, Gemini-1.5) to generate novel scientific hypotheses across physics, chemistry, and biology using a structured prompting protocol.
  • Human expert evaluation of 500 generated hypotheses reveals a 31% novelty rate and 67% scientific validity, with significant variation across disciplines and models.
  • Biology-domain hypotheses show highest novelty (38%) while physics hypotheses demonstrate superior validity (79%).
Abstract
We evaluate the capacity of large language models (GPT-4, Claude-3, Gemini-1.5) to generate novel scientific hypotheses across physics, chemistry, and biology using a structured prompting protocol. Human expert evaluation of 500 generated hypotheses reveals a 31% novelty rate and 67% scientific validity, with significant variation across disciplines and models. Biology-domain hypotheses show highest novelty (38%) while physics hypotheses demonstrate superior validity (79%). We propose a reproducible evaluation framework comprising five dimensions: novelty, plausibility, testability, specificity, and interdisciplinary breadth, and identify critical limitations in experimental design reasoning that must be addressed before autonomous research deployment.
Keywords
From Same Issue
Quantum-Classical Hybrid Algorithms for Combinatorial Optimization: Be...
Nakamura, H., Tanaka, Y., Fujiwara, K.,...
Diffusion Models for Protein Structure Prediction: Benchmarking Agains...
Chen, Y., Liu, X., Wang, Z., Zhou, J., L...