Statistics Seminar Series

Choose which semester to display:

Schedule for Fall 2026

Seminars are on Mondays
Time: 4:10 pm - 5:00 pm

Location: Room 903 SSW, 1255 Amsterdam Avenue

 

14-Sept.-26

Speaker: Roshni Sahoo (Columbia University)

Title: Data Fusion for High-Resolution Estimation

Abstract:  High-resolution estimates of population health indicators are critical for precision public health. We propose a method for high-resolution estimation that fuses distinct data sources: an unbiased, low-resolution data source (e.g. aggregated administrative data) and a potentially biased, high-resolution data source (e.g. individual-level online survey responses). We assume that the potentially biased, high-resolution data source is generated from the population under a model of sampling bias where observables can have arbitrary impact on the probability of response but the difference in the log probabilities of response between units with the same observables is linear in the difference between sufficient statistics of their observables and outcomes. Our data fusion method learns a distribution that is closest (in the sense of KL divergence) to the online survey distribution and consistent with the aggregated administrative data and our model of sampling bias. This approach significantly reduces bias in high-resolution estimates compared to baselines that rely on a single data source alone on a testbed that includes repeated measurements of three indicators measured by both the (online) Household Pulse Survey and ground-truth data sources at two geographic resolutions over the same time period.

Bio: Roshni Sahoo is assistant professor in the Decision, Risk, and Operations Division at Columbia Business School. She works on data-driven decision making for social impact. Her research develops statistical methodology to address challenges in public health and economic development, combining tools from machine learning, causal inference, and optimization. Roshni completed her PhD in Computer Science at Stanford University in 2026. She also received Bachelor of Science degrees in Computer Science and Mathematics at MIT in 2020.

21-Sept.-26

Speaker: Rajarshi Mukherjee (Harvard)

Title:  Statlib: A Library for Formal Statistical Theory in Lean 4

Abstract:  In this talk, we introduce and discuss Statlib, a verifiable library for mathematical statistics in Lean 4. As large language models now produce an increasing volume of plausible statistical arguments, machine-checked rigor is being considered a critical step. We will discuss how we can build Statlib on the most trusted library, Mathlib, to expand foundational mathematical and statistical results in measure theory, probability, and functional analysis, and to feed universal statistical abstractions back into Mathlib. Beyond programming, we have two things in mind: a deep understanding of the statistical language we're using, and a community of statisticians who are ready for this language to develop new research and education.With the aim of building an open, transparent, and trustworthy environment, we will discuss coordinating targeted formalization projects across classical and modern methods, aspects of autoformalization and its subtleties, and promises, developing comprehensive tutorials to onboard future contributors, and establishing a collaborative forum to address shared architectural themes and implementation challenges.

Bio: Rajarshi Mukherjee is an Associate Professor in the Department of Biostatistics at Harvard T.H. Chan School of Public Health. Previously, he was an Assistant Professor in the Division of Biostatistics at UC Berkeley following his time as a Stein Fellow in the Department of Statistics at Stanford University. He obtained his PhD in Biostatistics from Harvard University, advised by Prof. Xihong Lin. He is generally interested in understanding broad aspects of causal inference in observational studies in modern data settings, with a focus on learning about fundamental challenges in the statistical analysis of environmental mixtures and their effects on the cognitive development of children and cognitive decline in aging populations.

28-Sept.-26

Speaker: Linjun Zhang (Rutgers)

Title: Statistics for AI: Auditing What Models Learn and Measuring What They Can Do

Abstract:  As AI systems become increasingly capable yet opaque, two fundamental questions have taken on growing importance: What data did a model learn from, and how reliably can we measure what it can do? This talk develops a “statistics for AI” perspective on these questions through two case studies.
The first concerns data misappropriation in large language models. The training of LLMs has raised significant privacy and legal concerns, particularly when copyrighted or otherwise protected materials are used without proper attribution or authorization. To determine whether a target LLM has incorporated data generated by another model, we embed carefully designed watermarks in the protected data and formulate detection as a hypothesis-testing problem. This leads to a general statistical framework for constructing test statistics and rejection thresholds with explicit control of Type I and Type II errors. The second concerns the evaluation of LLMs when benchmarks are small and model outputs are stochastic. On difficult problems, an LLM may fail to generate a correct solution while still reliably distinguishing the better of two candidate solutions. We leverage this generation–verification gap to develop an evaluation framework that combines standard labeled outcomes with pairwise comparison signals elicited from the model. Treating these signals as control variates, we construct a semiparametric estimator based on the efficient influence function. The resulting estimator improves statistical efficiency and enables principled uncertainty quantification.

Bio: Linjun Zhang is an Associate Professor in the Department of Statistics and an affiliated faculty member in Computer Science at Rutgers University. He received his Ph.D. in Statistics from the Wharton School at the University of Pennsylvania in 2019, where he received the J. Parker Bursk Memorial Prize and the Donald S. Murray Prize for excellence in research and teaching, respectively. He is a recipient of the NSF CAREER Award, the Rutgers Presidential Teaching Award in 2024, and the Warren I. Susman Award for Excellence in Teaching in 2025. His current research interests include the statistical foundations of large language models, algorithmic fairness, privacy-preserving data analysis, and deep learning theory.

5-Oct-26

Speaker: Eric Kolaczyk ( McGill University)

Title:  Minority representation and fairness in network ranking, with an application to school contact diary data

Abstract:  Considerations of bias, fairness and representation are a prerequisite of responsible modern statistics, including statistical network analysis. Systematic bias can arise in network analysis when observation errors depend on node attributes, potentially distorting node rankings and group representation. In this talk, we are motivated by a high school contact network constructed from self-reported contact diaries, where nodes represent students, directed edges represent reported contacts, and groups are defined by sex. For networks with two known groups of nodes, we define minority representation as the proportion of nodes from the smaller group among the top-ranked nodes. We characterize its asymptotic behavior in degree-based rankings under a two-group stochastic block model. We model observation error through group-dependent missing-edge mechanisms and develop statistical methods to estimate the error rates and test for systematic bias. When correction is necessary, we apply a ranking correction procedure that targets the estimated group representation expected in the absence of observation error. Application to the contact diary data provides evidence that male students are more likely to omit contacts with female students than vice versa. These findings highlight the importance of accounting for group-dependent observation errors when assessing and correcting group representation in network rankings. Joint work with Hui Shen and Peter MacDonald.

Bio:  Eric Kolaczyk is a professor at McGill University in the Department of Mathematics and Statistics, and the inaugural director of the McGill Computational and Data Systems Institute (CDSI). He holds a Canada Research Chair in Statistical Network Analysis. He is also an associate academic member at Mila, The Quebec AI Institute.

12-Oct-26

Speaker: Siva Balakrishnan (Carnegie Mellon University)

Title:

Abstract:

Bio:

19-Oct-26

Speaker: Chen Zhou ( Erasmus University Rotterdam)

Title: Generalized Linear Models for Extremes: Estimation and Inference in High Dimensions

Abstract: We propose a regression model for the extreme tail of a response variable where a single covariate-dependent function characterizes the entire conditional tail. The tail shape itself is left unrestricted: heavy-, light- and short-tailed responses are covered by the same framework. We specify the function through a link function and a linear combination of the covariates, in the spirit of a generalized linear model. In estimation, we match the parametric specification to the underlying tail function under a Bregman divergence, over a region localized at the largest observations. We allow the number of covariates to exceed the effective sample size. We derive the convergence rate of the penalized estimator and propose a debiased estimator that is asymptotically normal, yielding confidence intervals for individual coefficients. Its asymptotic variance is determined by the covariance of the score, which under tail localization differs from the Hessian and must be estimated separately. We apply the method to automobile insurance claims data.

Bio: Chen Zhou is an Professor of Mathematical Statistics and Risk Management in the Erasmus School of Economics at Erasmus University Rotterdam
as well as a Research Fellow at the Tinbergen Institute. His research focuses on extreme value statistics and its applications in quantitative risk management, financial stability and financial regulation. He is an Area Editor of Extremes and Academic Director of the Master program in Quantitative Finance at Erasmus School of Economics.

26-Oct-26

Speaker:  TBD

Title:

Abstract:

Bio:

2-Nov-26

HOLIDAY

9-Nov-26

Speaker: TBD

Title:

Abstract:

Bio:

16-Nov-26

Speaker:TBD

Title:

Abstract:

Bio:

23-Nov-26

Speaker:TBD

Title:

Abstract:

Bio:

30-Nov-26

Speaker: Ilya Shpitser (Johns Hopkins)

Title:

Abstract:

Bio:

7-Dec-26

Speaker: Xiaowu Dai (UCLA)

Title:

Abstract:

Bio:

14-Dec-26

Speaker: Tracey Ke (Harvard)

Title:

Abstract:

Bio: