Lila Gleitman Prize lecture: cognitive tools

Speaker: Judy Fan, this year's Lila Gleitman Prize recipient

Before the lecture

The session opened with a tribute to Lila Gleitman, remembered as a brilliant, joyful presence who trained a long line of cognitive scientists, many in the room that day, and who passed away in August 2021 as Professor Emerita of Psychology and Linguistics at Penn. The prize, established in 2022 by the Cognitive Science Society and the Society for Language Development (which Gleitman founded), is funded by donations from her family, friends, students, and colleagues.

Judy Fan was introduced as an assistant professor of psychology, with courtesy appointments in education and computer science, at Stanford. PhD from Princeton in 2016, undergraduate work in neurobiology and statistics at Harvard. Her research was described as bridging psychology, neuroscience, vision science, and education around one central question: how do people use drawings, sketches, and other physical representations of thought to learn, communicate, and solve problems.

The talk: what makes a thinking tool

Fan opened with the number line and the Cartesian coordinate system, clean, familiar, and completely invented. There are no straight lines waiting in nature to be discovered. Humanity built these tools, and they became so useful that essentially every math curriculum on Earth now teaches them. Her question was where the capacity to invent and use tools like this comes from.

She organized the talk around a simple map: individual cognition versus social cognition, crossed with cognitive tools (physical or informational objects that shape thought) versus engineering (using understanding to go out and build something new). A full account, she argued, needs all four quadrants, not just the traditional individual-cognition focus most of the field defaults to.

Part one: visual abstraction

She distinguished three building blocks: visual perception (raw input becoming a meaningful percept), visual production (leaving a lasting, meaningful mark), and visual communication (arranging marks to have a specific effect on someone else). On why a drawing is meaningful at all, she contrasted a resemblance account, it just looks like the thing, with a convention account, meaning is purely socially agreed. She described a 2014 finding, with Dan Yamins and Nikhil Bhatia, that general-purpose vision networks trained only on labeled photographs generalized surprisingly well to simple line drawings, supporting a modernized version of the resemblance account. She mentioned, with a smile, that she first presented an early version of this result at her very first CogSci meeting, in 2015.

A follow-up study with Robert Hawkins, Bao Wu, and Noah Goodman used a two-player drawing game. A sketcher had to help a partner pick out a target image from either same-category or different-category distractors. People drew more detail when distractors were close, similar objects, and got away with far more abstract, category-level sketches when distractors were obviously different. A clean demonstration that sketching is a pragmatic act, calibrated to exactly how much detail the moment actually requires, not a fixed visual skill.

Work with Holly Huey, Charles Lu, and Karen Wang looked at how people convey causal, mechanistic knowledge through drawing, how a circuit turns on a light, say, finding that people spontaneously reach for devices like motion arrows, even at real cost to visual fidelity. That shows the underlying representation is flexible enough to support very different communicative goals from the same raw experience.

Part two: reasoning with data

She framed data visualization as one of the most powerful cognitive technologies ever built. Unlike a telescope or microscope, a plot lets you see patterns too large, slow, or noisy to perceive directly, and she illustrated this with William Playfair's 1786 trade-balance plot, among the very first time-series visualizations ever drawn. She announced a new research network, launched in 2024 with NSF support, bringing together academic groups to study how people learn to reason about data, measurement, inference, and uncertainty. She named several student collaborators and asked, among other things, why today's AI systems can be so capable with multimodal data and yet so brittle, in ways that might reveal something about what genuinely robust multimodal reasoning requires.

Part three: building things together

Work with her former student Will McCarthy, now at Autodesk Research, used a digital block-building environment where people reconstruct target tower shapes from a fixed inventory of blocks. Despite many valid solution sequences, people converge, even on a first attempt and more so on repeat attempts, on a small set of common intermediate moves. A related project with Felix Binder, now at Meta, modeled which visual sub-goals people pick and in what order. Work led by Leo Wang modeled people's object descriptions as inferring an implicit mental "graphics program," explaining why people settle on one level of description over another.

No cognitive tool, however powerful, has ever changed the world on its own. Writing, computers, whatever comes next. Real, lasting change only happens once we build environments that let everyone actually learn to use the tool.

She closed by naming two directions ahead. Education, framed as designing environments (including AI tutoring systems) that support rather than replace the work teachers and learners do together. And design, understood as the broader question of how people build experiences, games, music, stories, that are meaningful in themselves, not just efficient. She tied this back to Gleitman directly: the greatest gift at this point in a career, she said, is helping the next generation of researchers find their way, exactly as Gleitman did for so many people in that room.

Questions afterward

One questioner pushed on whether basic visual primitives like the straight line really don't exist in the world, given that horizon lines and shortest paths do occur naturally. Fan took a conciliatory both-and position, suggesting the deeper source of structure is the set of problems nature poses to a species occupying a particular ecological niche, and argued for returning to closer naturalistic observation to help settle ongoing debates about whether current AI models are getting visual abstraction right. Another asked whether black-box neural networks can really tell us which visual inputs are meaningful, or whether more symbolic approaches are needed. She said she wants all the tools available, noting that brains themselves are complex recurrent systems we're still learning to interpret, and that symbolic, intermediate-level descriptions remain scientifically useful for prediction and explanation regardless of how the deeper, harder question of subjective experience eventually gets resolved.