Galactica: A Large Language Model for Science

Abstract

Information overload is a major obstacle to scientific progress. Theexplosive growth in scientific literature and data has made it ever harder todiscover useful insights in a large mass of information. Today scientificknowledge is accessed through search engines, but they are unable to organizescientific knowledge alone. In this paper we introduce Galactica: a largelanguage model that can store, combine and reason about scientific knowledge.We train on a large scientific corpus of papers, reference material, knowledgebases and many other sources. We outperform existing models on a range ofscientific tasks. On technical knowledge probes such as LaTeX equations,Galactica outperforms the latest GPT-3 by 68.2% versus 49.0%. Galactica alsoperforms well on reasoning, outperforming Chinchilla on mathematical MMLU by41.3% to 35.7%, and PaLM 540B on MATH with a score of 20.4% versus 8.8%. Italso sets a new state-of-the-art on downstream tasks such as PubMedQA andMedMCQA dev of 77.6% and 52.9%. And despite not being trained on a generalcorpus, Galactica outperforms BLOOM and OPT-175B on BIG-bench. We believe theseresults demonstrate the potential for language models as a new interface forscience. We open source the model for the benefit of the scientific community.

Quick Read (beta)

loading the full paper ...