Working papers, etc.
Inferential Models. My primary research focus is on foundations of statistics and probability, in particular, developing a framework of valid and efficient probabilistic inference that is stronger than the Bayesian framework in various ways. One key advantage over the Bayesian solution is that this framework's probability-like output is provably valid in the following sense: the method assigning high "probability" to false hypotheses is a provably rare event. This is (intentionally) similar to frequentist-style error rate control, but stronger in certain respects. It turns out that achieving this validity property requires considerations beyond ordinary probability, and our inferential model framework relies heavily on random sets; more recently, I've been focusing on possibility measures and other kinds of more general imprecise probabilities. My current efforts aim to extend these developments to modern problems that involve structured-but-high-dimensional unknowns. The key to these developments is the use of partial prior information that drives the regularization needed to achieve efficiency in high-dimensions.
—Induction and the rule of succession through a possibilistic inferential model lens (with S.-N. Prim, M. Raner, and J. Williams). [arXiv]
—Valid and efficient possibilistic structure learning in Gaussian linear regression (with N. Singer and J. Williams). [arXiv]
—Divide-and-conquer with finite sample sizes: valid and efficient possibilistic inference (with L. Cella and E. Hector). [arXiv]
—Fisher's underworld and the behavioral--statistical reliability balance in scientific inference. [researchers.one] [arXiv]
—Valid and efficient imprecise-probabilistic inference with partial priors, III. Marginalization. [researchers.one] [arXiv]
—Valid and efficient imprecise-probabilistic inference with partial priors, II. General framework. [researchers.one] [arXiv]
—Valid and efficient imprecise-probabilistic inference with partial priors, I. First results. [researchers.one] [arXiv]
—An imprecise-probabilistic characterization of frequentist statistical inference. [researchers.one] [arXiv]
Anytime-valid inference, e-values, etc. A major challenge to classical statistics is the mathematical fact that error rate control properties break down when, e.g., the sample size is determined in "real time" as opposed to being predetermined and fixed. As an example, researchers might reach their planned/budgeted sample size, note that they're close to a significant result, and then decide to collect more data in hopes of reaching the significance threshold. While these strategies are practically very natural, it's a serious issue if following these procedures ruins the reliability of the statistical methods employed. Anytime-valid inference is about, among other things, the design of methods that are provably reliable even under dynamic data collection schemes. The workhorses behind these methods are called e-values or e-processes, which have connections to classical ideas dating back at least to Wald and Robbins.
—Universal inference for model selection on networks (with J. Williams and E. Yanchenko) [arXiv]
Generalized Bayesian inference & prediction. Generalized Bayes can mean various things, and here I mean the construction of a "posterior distribution" when the quantity of interest is not directly defined as the parameter of a statistical model. One situation where this would be appropriate is when the quantity of interest is defined as a minimizer of an expected loss, as is common in machine learning applications. For such cases, a Gibbs posterior is a data-dependent probability distribution constructed directly from that loss—no likelihood required; see M. & Syring's review chapter. Since one can't count on having a likelihood (due to computational intractability or simply not being willing to specify one and risk model misspecification biases), if there's a more direct way to carry out a Bayesian(-like) posterior construction, then the risk of model misspecification biases can be reduced/eliminated without the need for overly-complex nonparametric models. But there's no free lunch: by mis- or under-specifying the model, the calibration enjoyed by well-specified Bayesian posteriors (e.g., via the Bernstein–von Mises theorem) no longer holds, so the calibration needs to be handled manually; for this, I personally like the generalized posterior calibration algorithm in Syring & M. Biometrika 2019.
—Generalized Bayes inference on a linear personalized MCID (with P.-S. Wu). [researchers.one] [arXiv]
—Gibbs posterior inference on a Levy density under discrete sampling (with Z. Wang). [researchers.one] [arXiv]
—Calibrating generalized predictive distributions (with P.-S. Wu). [researchers.one] [arXiv]
Miscellaneous. I have ongoing work on, e.g., empirical Bayes inference in high-dimensional problems, conformal prediction, and nonparametric estimation of mixing distributions, data fusion, among other things.
— Variational empirical Bayes variable selection in high-dimensional logistic regression (with Y. Tang). [arXiv]
—Permutation-based uncertainty quantification about a mixing distribution (with V. Dixit). [arXiv]
