Uncovering the Semantics of Wikipedia Categories

Abstract

The Wikipedia category graph serves as the taxonomic backbone for large-scaleknowledge graphs like YAGO or Probase, and has been used extensively for taskslike entity disambiguation or semantic similarity estimation. Wikipedia'scategories are a rich source of taxonomic as well as non-taxonomic information.The category 'German science fiction writers', for example, encodes the type ofits resources (Writer), as well as their nationality (German) and genre(Science Fiction). Several approaches in the literature make use of fractionsof this encoded information without exploiting its full potential. In thispaper, we introduce an approach for the discovery of category axioms that usesinformation from the category network, category instances, and theirlexicalisations. With DBpedia as background knowledge, we discover 703k axiomscovering 502k of Wikipedia's categories and populate the DBpedia knowledgegraph with additional 4.4M relation assertions and 3.3M type assertions at morethan 87% and 90% precision, respectively.

Quick Read (beta)

loading the full paper ...