Can Hallucinations Help? Boosting LLMs for Drug Discovery

  • 2025-08-22 12:12:09
  • Shuzhou Yuan, Zhan Qu, Ashish Yashwanth Kangen, Michael Färber
  • 0

Abstract

Hallucinations in large language models (LLMs), plausible but factuallyinaccurate text, are often viewed as undesirable. However, recent work suggeststhat such outputs may hold creative potential. In this paper, we investigatewhether hallucinations can improve LLMs on molecule property prediction, a keytask in early-stage drug discovery. We prompt LLMs to generate natural languagedescriptions from molecular SMILES strings and incorporate these oftenhallucinated descriptions into downstream classification tasks. Evaluatingseven instruction-tuned LLMs across five datasets, we find that hallucinationssignificantly improve predictive accuracy for some models. Notably,Falcon3-Mamba-7B outperforms all baselines when hallucinated text is included,while hallucinations generated by GPT-4o consistently yield the greatest gainsbetween models. We further identify and categorize over 18,000 beneficialhallucinations, with structural misdescriptions emerging as the most impactfultype, suggesting that hallucinated statements about molecular structure mayincrease model confidence. Ablation studies show that larger models benefitmore from hallucinations, while temperature has a limited effect. Our findingschallenge conventional views of hallucination as purely problematic and suggestnew directions for leveraging hallucinations as a useful signal in scientificmodeling tasks like drug discovery.

 

Quick Read (beta)

loading the full paper ...