Authors
Maithe van Noort
L. Korthals
M.L.S. de Heer Kloots
G. Aldegheri
M. Heilbron
Date (dd-mm-yyyy)
2025
Title
Compositional Meaning in Vision-Language Models and the Brain
Publication Year
2025
Number of pages
4
Document type
Abstract
Abstract
What is the role of compositional structure in the alignment of visual and linguistic brain areas to computational semantic embeddings? Vision-language models (VLMs) have shown meaningful alignment to the brain in their representations of semantic structure, for both images and text. However, the extent to which these representations capture compositional structure -- i.e. changes in meaning based on changes to the combinatorial structure of parts -- remains uncertain. Here we leverage Winoground, a dataset designed to test compositionality in multimodal representations, to compare the compositional structure captured by different model embeddings, as well as fMRI responses collected as part of a larger study on multi-modal meaning (with 2760 image and 2760 semantically equivalent language trials). In contrast to VLM embeddings, neural representations in the brain show a striking absence of compositional processing (chance level performance) when evaluated on the Winoground benchmark -- despite robust semantic encoding of individual concepts as measured by voxel activity predictions. This is intriguing as distinctions between stimuli in Winoground are trivial to any English-speaking human, highlighting the challenge of identifying the substrates of compositional processing in the brain. Our targeted dataset and evaluation pipeline lay the foundation for systematic, cross-modal evaluations of compositionality in both artificial and biological neural representations.
Permalink
https://hdl.handle.net/11245.1/4bbe81c9-d43f-414f-98c4-5092f12ede22