Speakers and topics · Wednesdays 4:30–5:30 pm · CoDA E160
The seminar meets Wednesdays, 4:30–5:30 pm in CoDA E160, in person only, through all three quarters of the academic year. Each talk below states its room, and any talk held elsewhere is flagged. Speakers for Winter and Spring are added here as they confirm.
Any changes to the schedule will be reflected here, so check this page often. Or subscribe and we will send you the coming week’s talk every Sunday.
Fall 2026
WarningTwo talks meet in a different room
The September 23 and October 7 talks are in Packard 101, not CoDA E160. All other talks are in CoDA E160 as usual.
September 23
Different room this week:Packard 101, not CoDA E160.
Mechanistically interpretable models of computation in the early visual system
The early visual system has long provided a tractable setting for studying the mechanisms of neural computation. Yet even the retina and primary visual cortex transform signals through multiple excitatory and inhibitory pathways that interact nonlinearly, a complexity that until recently made the neural code for natural stimuli inaccessible. Over the past decade we have developed approaches using small convolutional neural networks fit to neural data to recover both computational and mechanistic insight into this code. In the retina, three-layer CNNs with fewer cell types than the retina itself accurately capture responses to natural scenes and generalize to reproduce a diverse set of ethological phenomena related to motion encoding, adaptation, and predictive coding. Moreover, these models are interpretable in that interneurons recorded separately, and never fit directly, correlate strongly with model interneurons. We have further developed mechanistic interpretability methods that automatically generate hypotheses for how interneurons with different spatiotemporal responses combine to produce visual computations. We have used these models to track visual computations from retina to cortex in the mouse, and to discover new phenomena. These results demonstrate a general approach for studying the circuit mechanisms of visual computation in biological and artificial networks.
Generative video models have shown remarkable prowess, yet their evolution into true ‘world models’—stable, controllable, and long-horizon simulators—remains a grand challenge in Spatial AI. In this talk, we detail recent advances in video world models, specifically focusing on enhancing temporal consistency and reasoning across long contexts as well as mitigating auto-regressive drift. We explore the integration of multi-agent dynamics and the development of effective human-centric conditioning and control strategies. Importantly, we demonstrate how world models enable unique strategies for training robust policies to drive autonomous agents by embracing scaling laws.
October 7
Different room this week:Packard 101, not CoDA E160.
Missing data presents a persistent challenge in biomedical research. Data imputation techniques have evolved from single-modality approaches to multi-modal approaches, which show great promise for imputing one modality based on the availability of another. Recent advancements in large, pre-trained artificial intelligence (AI) models, known as foundation models, offer even more powerful solutions for data imputation. We introduce the concept of cross-modal data modeling, a methodology harnessing foundation models to impute missing data and also generate realistic synthetic samples. Multi-modal modeling empowers researchers to model complex interactions among diverse biomedical data types, including omics and imaging. This approach can illuminate how one modality influences another, facilitating in-silico exploration of disease mechanisms without the need for extensive and costly real-world data collection. We highlight ongoing efforts in multi-modal modeling in spatial omics, digital pathology and radiology, and anticipate its substantial contributions to understanding disease biology and enhancing healthcare practices.
Training Methods for a Medical Imaging Foundation Model
For the past several years, AI has been in an era of scale, during which the developers of frontier models have increased training dataset size and computing capacity, leading to reliable advances in performance. Healthcare has lagged this trend, in part because of the need to protect patient privacy and the difficulties of curating and aggregating AI-ready healthcare datasets. As we overcome these challenges, we have a unique opportunity to develop self-supervision methods for massive high-quality healthcare datasets. We will describe plans to train a medical imaging foundation model on Stanford’s entire digital radiology archive, approximately 2 petabytes of diagnostic medical images and associated reports. We will show how new methods can dramatically reduce the cost of training, improving the efficiency and accuracy of machine learning applied to massive imaging datasets.
The human genome encodes functional DNA words, syntax and grammar that regulate gene activity in cell-type specific manner. Deciphering this regulatory code is essential for understanding how genes are controlled across cell types, individuals, and species, and for explaining how genetic variation shapes traits and disease. Large-scale molecular profiling efforts now provide rich, cell-type–specific regulatory maps, offering an unprecedented opportunity to apply machine learning to decode sequence-function relationships. In this talk, I will present lightweight deep learning models trained on diverse molecular profiles coupled with interpretation frameworks that can be used to (1) uncover causal sequence syntax and its context-specificity across molecular and cellular contexts, (2) detect and correct experimental biases, (3) infer fundamental biophysical parameters, (4) prioritize functional regulatory variants underlying common and rare diseases, and (5) edit and design regulatory DNA for potential therapeutic applications. Our models achieve or surpass the performance of large multi-task supervised foundation models and self-supervised DNA language models across diverse benchmarks. By systematically interpreting ~5,000 models trained on bulk and single-cell datasets spanning diverse fetal and adult contexts, we expose the remarkable complexity and context-specificity of regulatory sequence lexicons, syntax and variation encoded in the human genome. Our models and derived interpretation products serve as a foundational resource for understanding the human genome.
Large-Scale Computational Models at the Neuroscience-AI Intersection
I’ll first give a primer on the recent history of the recently-developed field of Cognitive NeuroAI, which uses neural networks to explain the brain and human behavior. Then I’ll discuss two new large-scale projects at the cutting edge of Cognitive NeuroAI: Probabilistic Structure Integration (PSI), a new class of cognitively-inspired world models of use in computer vision, robotics, and neural modeling applications; and the Stanford Digital Brain Project, a recently-launched experimental and computational effort to create a massive digital twin of the human brain.
Multi-Modal Foundation Models for Precision Medicine
Clinical decision-making requires integrating diverse data types — medical images, clinical text, and molecular omics profiles. AI methods that effectively leverage multi-modal data hold transformative potential for biomedical discovery and clinical care. This talk explores how multi-modal foundation models can generate biologically relevant insights and clinically meaningful predictions, such as identifying which patients respond to immunotherapy. I will present our recent work developing multi-modal foundation models that integrate pathology images with text, as well as single-cell spatial transcriptomics and spatial proteomics, showing how these advances can improve diagnostic accuracy and guide personalized cancer treatment.
Detecting coordinated behavior across multiple fully encrypted domains
Coordinated inauthentic behavior threatens societal stability, markets, and security. Advances in generative AI amplify these threats, enabling effortless content creation, amplifying actors’ influence. Detection is hindered by cross-domain activity, where pseudonymous profiles operate across encrypted platforms, and by privacy constraints limiting content analysis. We develop a robust and scalable cross-domain identity matching framework, based on bursty dynamics, independent of content or interaction data. It outperforms state-of-the-art temporal and structural approaches, remains resilient to incomplete data, and is fairly stable over time. By framing identity matching within the “network of networks” perspective, we demonstrate how coordinated behavior propagates across domains. This dual methodological and theoretical contribution paves the way for innovative strategies to combat digital threats in an increasingly complex, encrypted, and adversarial landscape. Joint work with Shahar Somin, Jeremy Kepner, and Tom Cohen.
You … your memories and ambitions, your sense of personal identity and free will, are in fact no more than the behavior of a vast assembly of nerve cells …” Crick’s words capture the profound challenge of decrypting the neural code. This challenge has long been hindered by two limitations: our ability to record activity from large neuronal populations under the complex, variable conditions in which brains evolved, and our capacity to model the intricate relationships between stimuli, behaviors, and neural activity. Recent breakthroughs are beginning to overcome these barriers. Cutting-edge technologies now enable large-scale recordings, while AI can construct predictive brain models that link stimuli, neural activity, and behavior. These digital twins open the door to virtually limitless in silico experiments, testing theories that would otherwise be impossible to probe at scale in living brains. I will discuss our work building brain foundation models and digital twins to uncover the mechanisms of neural representation, with predictions validated through closed-loop experiments.