Can eye movements tell us what a scene means?¶
Where people look is a readout of what they find meaningful — so gaze should be predictable from semantics alone. Ongoing work with Sophie Su uses CLIP embeddings together with eye fixation maps to characterize the semantic content of a film as it unfolds, and to ask how much of viewers' gaze it explains.
Su, Upadhyayula & Zacks (in preparation).