Large Language Models as Zero-Shot Scientific Hypothesis Generators: Evaluation and Implications
Al-Rashidi, A., Hassan, M., Ibrahim, K., Wang, L.. Large Language Models as Zero-Shot Scientific Hypothesis Generators: Evaluation and Implications. Loshu Comput. Intell..
Vol.3, No.1. Jan 2026. https://doi.org/10.58921/ljci.2026.0102
Article
Recommended articles
Cited by 50
Metrics
Highlights
- We evaluate the capacity of large language models (GPT-4, Claude-3, Gemini-1.5) to generate novel scientific hypotheses across physics, chemistry, and biology using a structured prompting protocol.
- Human expert evaluation of 500 generated hypotheses reveals a 31% novelty rate and 67% scientific validity, with significant variation across disciplines and models.
- Biology-domain hypotheses show highest novelty (38%) while physics hypotheses demonstrate superior validity (79%).
Abstract
We evaluate the capacity of large language models (GPT-4, Claude-3, Gemini-1.5) to generate novel scientific hypotheses across physics, chemistry, and biology using a structured prompting protocol. Human expert evaluation of 500 generated hypotheses reveals a 31% novelty rate and 67% scientific validity, with significant variation across disciplines and models. Biology-domain hypotheses show highest novelty (38%) while physics hypotheses demonstrate superior validity (79%). We propose a reproducible evaluation framework comprising five dimensions: novelty, plausibility, testability, specificity, and interdisciplinary breadth, and identify critical limitations in experimental design reasoning that must be addressed before autonomous research deployment.