A Data-free Universal Prior over Syntactic Structures

  • 2026-09-26 17:07:02
  • Fermín Moscoso del Prado Martín
  • 0

Abstract

Probability is fundamental to theories of language comprehension, production, acquisition, and evolution, as well as to large language models. Existing theories assume that the probabilities of syntactic structures emerge from language-specific experience. An unexplored possibility is that these probabilities can have a universal component that is independent of any language-specific experience. Here, I show that a universal prior over syntactic structures emerges from a cognitively motivated model of human language production, in which words are progressively integrated into syntactic structure through network growth. The resulting prior assigns probabilities to syntactic structures --represented as dependency trees-- without fitting parameters to linguistic data. It assigns higher probabilities to attested than to random trees in all 138 typologically diverse languages examined. These prior probabilities correlate positively with those estimated from corpora in 33 of 34 languages. The results indicate that part of the probability structure of syntax can arise independently of language-specific learning. Linguistic experience may therefore refine probabilities that are already structured by the process of language production, rather than estimating them from scratch. This identifies a possible cognitive origin for part of the probability distribution over syntactic structures, linking language production and statistical learning while providing a data-independent structural bias for probabilistic models of language.

 

Quick Read (beta)

loading the full paper ...