Skip to content

Does visual narrative comprehension involve a grammar?

From cave paintings to murals, visual storytelling has long existed as a communicative tool. When comprehending visual narratives, do we parse them using a grammar, similar to that of language? Adapting tools from computational psycholinguistics, we demonstrate that a model with grammatical properties can segment visual narratives similar to humans. Read more below:


Top row: model surprisal as a function of distance from a narrative boundary, for the Earley parser, HMM, trigram and bigram models under different information conditions. Bottom row: boundary agreement rising with each model's normalized surprisal
Model comparison. Panel titles give the information each model had access to — semantics (changes in space, location, causality) and narrative grammar (the E, I, P, R categories and their pairing rules). Top: surprisal by distance from the boundary; only models with grammatical structure spike at the boundary itself. Bottom: human boundary agreement rises with each model's normalized surprisal. Error bars show standard error, shaded regions 95% CI.

Upadhyayula & Cohn (2025), Cognitive Science. Paper · Data & code · Talk


Back to How is information organized?