Dubova opened with something personal. Seven years ago she'd been denied a visa and couldn't attend her first Cognitive Science Society meeting, posted about it on Twitter to her one follower, and the resulting attention led the Society to introduce a policy rotating conference locations every few years to make the meeting more globally accessible. This year was the first year that policy actually took effect, and she was visibly moved to be standing there because of it. She thanked her advisor, Rob Goldstone, and collaborators including Sabina Sloman.
Her framing for the whole talk: around 2019 everyone was worried about the replication crisis in psychology, results that don't hold up. She'd become increasingly bothered by something upstream of that though, a conceptual crisis, where we're not even sure we know what we're supposed to be studying, or whether the standard rules of the scientific method we take for granted are actually justified by any evidence about how learning works, rather than just assumed.
With Goldstone, she built a model where simulated agents inhabit a world structured as a high-dimensional multivariate Gaussian, a "ground truth," and collect data by picking a variable to fix and observing a conditional sample. Agents build theories, which are simplified compressions of the data they've collected, and share theories or data with each other.
She used this to compare experimentation strategies: confirmation-driven (design experiments to confirm your own theory), disagreement-driven, and more exploratory strategies like random or novelty-seeking search. The trick was measuring two different kinds of performance. Perceived performance, meaning how well an agent's theory fits the data it has actually collected. And actual performance, meaning how well it fits the true underlying distribution, which you can only check because it's a simulation.
The result was that strategies that look best by perceived performance, confirmation-driven experimentation especially, are often the worst generalizers. They look impressive because the agent has boxed itself into a narrow, self-consistent slice of the space that it never leaves. More exploratory strategies, like random or novelty-driven search, generalize far better to the true underlying structure, even though they look messier along the way. The ranking almost completely reverses depending on which kind of performance you're looking at.
She and Goldstone followed this up with real neuroscientists, giving them different conceptual framings of an unfamiliar system and watching what experiments they chose to run next. Conceptual expectations clearly shaped experimental choices, in ways that mapped onto the disagreement-versus-confirmation distinction from the simulations. Those expectations helped when they happened to be accurate, and hurt when they were wrong.
The second half of the talk, done with Sabina Sloman, pushed on something more fundamental: does learning a good theory always require compressing your experience down into something simpler? The standard picture, which she called the constrained-capacity view, says yes. A theory is a lossy compression of the data, like fitting a straight line through a cubic function's output, or an autoencoder squeezing information through a narrow bottleneck.
She laid out three regimes. Constrained capacity: not enough room to fit everything, so you have to simplify, the classic view. Sufficient capacity: just enough room to memorize everything exactly, which risks catastrophic overfitting, like a twentieth-degree polynomial through twenty data points. And excess capacity: far more room than you'd need to memorize everything, a thousand-degree polynomial say, which sounds like it should be an overfitting disaster.
She and Sloman have written this up as a target article, now out in Behavioral and Brain Sciences, with a set of diagnostic criteria for figuring out whether a given learner in a given task is operating with constrained, sufficient, or excess capacity, and what that implies about how it will generalize. She was careful to say the dissertation doesn't answer the big questions so much as open them up for other people to run with.
One question pushed on how you'd tell meaningful experimental variation apart from irrelevant noise when designing a sampling strategy in a real, structured world rather than a simulation. She agreed this was a real limitation of the simulations and an open question she wants to study further. Another asked whether this connects to Bayesian model averaging, since it also avoids collapsing onto one possibly-wrong guess. She thought the analogy was a good intuition, closer to how Bayesian inference preserves many hypotheses rather than committing to one, versus frequentist point estimates, but wasn't sure of the precise technical relationship.
I was scribbling fast during the personal parts of this one, so some of the biographical details might not be word for word.