Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description Framework

  • 2025-09-19 11:06:48
  • Laura Kopf, Nils Feldhus, Kirill Bykov, Philine Lou Bommer, Anna Hedström, Marina M. -C. Höhne, Oliver Eberle
  • 0

Abstract

Automated interpretability research aims to identify concepts encoded inneural network features to enhance human understanding of model behavior.Within the context of large language models (LLMs) for natural languageprocessing (NLP), current automated neuron-level feature description methodsface two key challenges: limited robustness and the assumption that each neuronencodes a single concept (monosemanticity), despite increasing evidence ofpolysemanticity. This assumption restricts the expressiveness of featuredescriptions and limits their ability to capture the full range of behaviorsencoded in model internals. To address this, we introduce Polysemantic FeatuReIdentification and Scoring Method (PRISM), a novel framework specificallydesigned to capture the complexity of features in LLMs. Unlike approaches thatassign a single description per neuron, common in many automatedinterpretability methods in NLP, PRISM produces more nuanced descriptions thataccount for both monosemantic and polysemantic behavior. We apply PRISM to LLMsand, through extensive benchmarking against existing methods, demonstrate thatour approach produces more accurate and faithful feature descriptions,improving both overall description quality (via a description score) and theability to capture distinct concepts when polysemanticity is present (via apolysemanticity score).

 

Quick Read (beta)

loading the full paper ...