Disentangling the Factors of Convergence between Brains and Computer Vision Models

  • 2025-08-25 17:23:27
  • Joséphine Raugel, Marc Szafraniec, Huy V. Vo, Camille Couprie, Patrick Labatut, Piotr Bojanowski, Valentin Wyart, Jean-Rémi King
  • 0

Abstract

Many AI models trained on natural images develop representations thatresemble those of the human brain. However, the factors that drive thisbrain-model similarity remain poorly understood. To disentangle how the model,training and data independently lead a neural network to develop brain-likerepresentations, we trained a family of self-supervised vision transformers(DINOv3) that systematically varied these different factors. We compare theirrepresentations of images to those of the human brain recorded with both fMRIand MEG, providing high resolution in spatial and temporal analyses. We assessthe brain-model similarity with three complementary metrics focusing on overallrepresentational similarity, topographical organization, and temporal dynamics.We show that all three factors - model size, training amount, and image type -independently and interactively impact each of these brain similarity metrics.In particular, the largest DINOv3 models trained with the most human-centricimages reach the highest brain-similarity. This emergence of brain-likerepresentations in AI models follows a specific chronology during training:models first align with the early representations of the sensory cortices, andonly align with the late and prefrontal representations of the brain withconsiderably more training. Finally, this developmental trajectory is indexedby both structural and functional properties of the human cortex: therepresentations that are acquired last by the models specifically align withthe cortical areas with the largest developmental expansion, thickness, leastmyelination, and slowest timescales. Overall, these findings disentangle theinterplay between architecture and experience in shaping how artificial neuralnetworks come to see the world as humans do, thus offering a promisingframework to understand how the human brain comes to represent its visualworld.

 

Quick Read (beta)

loading the full paper ...