Source-linked AI summary
Bayesian Mode Regression
Keming Yu, Katerina Aristodemou
TL;DR
Mode regression summarizes how regressors affect the conditional mode, addressing a gap in Bayesian inference for this target. The paper develops parametric, nonparametric, and empirical-likelihood Bayesian approaches, with estimates that are consistent and asymptotically normal under standard conditions, even under likelihood misspecification. The approaches provide proper inference tools and are reported as strong competitors of classical mode-regression estimates.
Problem
Existing mode-regression inference methods are not Bayesian, while conditional-density approaches are not direct mode estimators and face dimensionality and interpretability challenges.
Method
The paper develops parametric, nonparametric, and empirical-likelihood Bayesian mode-regression methods, including a mode-uniform likelihood and Dirichlet process mixtures.
Results
The proposed estimates are consistent and asymptotically normal under standard conditions, even when the likelihood is misspecified, and the approaches offer proper inference tools.
Takeaways & Limitations
The proposed Bayesian approaches are reported as strong competitors of classical mode-regression estimates.
Takeaways & Limitations
Existing nonparametric conditional-density methods face the curse of dimensionality and make estimated conditional modes difficult to interpret through predictors or covariates.
Abstract
from arXiv · showhide
Like mean, quantile and variance, mode is also an important measure of central tendency and data summary. Many practical questions often focus on "Which element (gene or file or signal) occurs most often or is the most typical among all elements in a network?". In such cases mode regression provides a convenient summary of how the regressors affect the conditional mode and is totally different from other regression models based on conditional mean or conditional quantile or conditional variance. Some inference methods have been used for mode regression but none of them from the Bayesian perspective. This paper introduces Bayesian mode regression by exploring three different approaches. We start from a parametric Bayesian model by employing a likelihood function that is based on a mode uniform distribution. It is shown that irrespective of the original distribution of the data, the use of this special uniform distribution is a very natural and effective way for Bayesian mode regression. Posterior estimates based on this parametric likelihood, even under misspecification, are consistent and asymptotically normal. We then develop a nonparametric Bayesian model by using Dirichlet process (DP) mixtures of mode uniform distributions and finally we explore Bayesian empirical likelihood mode regression by taking empirical likelihood into a Bayesian framework. The paper also demonstrates that a variety of improper priors for the unknown model parameters yield a proper joint posterior. The proposed approach is illustrated using simulated datasets and a real data set.
1 Introduction
Mode captures the most likely or most typical value and can preserve distributional features that mean and median may obscure. Existing conditional-density approaches face dimensionality and interpretability problems, motivating this paper’s fully Bayesian framework for direct mode-regression inference using three approaches.
- Mode identifies the most likely or most typical value and is used across fields including biology, astronomy, economics, and finance.
- Unlike the mean or median, mode can preserve distributional features such as wiggles that averaging may remove.
- Conditional mode is typically estimated through conditional density estimation using nonparametric methods.
- These approaches are not direct conditional-mode estimators, and conditional-density estimation may suffer from the curse of dimensionality.
- Estimated conditional modes are difficult to describe and interpret in terms of predictors or covariates.
- Existing direct methods have limited application because of inadequate inference tools, slow coefficient convergence, bandwidth selection, and approximate normal confidence intervals.
- The paper introduces a fully Bayesian framework for direct mode-regression inference using parametric, nonparametric, and empirical-likelihood approaches.
2 Bayesian mode regression
The paper develops Bayesian mode regression through a mode-uniform working likelihood, with parametric, nonparametric, and empirical-likelihood formulations. The parametric approach supports consistent and asymptotically normal posterior inference, while improper priors can still yield proper posteriors.
- Mode regression formulation: Mode regression models the conditional mode as mode(y|x) = x′β, using a step-loss formulation that becomes focused around the mode as σ approaches 0.The conditional mode is reformulated as regression with a zero-mode error term.
- Parametric Bayesian method: The parametric Bayesian method treats the mode-regression objective as a likelihood based on a uniform distribution over a window of width 2σ.Posterior estimates are obtained by combining this likelihood with a prior for β, using MCMC when no standard conjugate prior is available.
- Consistency and asymptotic normality: Even when the working likelihood is misspecified, posterior estimators remain consistent under regularity conditions, in the sense of minimizing Kullback–Leibler distance within the parametric family.The paper also establishes asymptotic normality of the posterior as sample size increases.
- Covariance estimation: MCMC posterior draws provide a natural and efficient way to estimate the covariance matrix and other asymptotic quantities of the classical mode-regression estimator.The Bayesian approach yields a full posterior density rather than only a point estimate.
- Prior selection: The method uses priors for the window parameter σ, including a Uniform(w1, w2) prior guided by empirical, Chebyshev, and Silverman-based rules.Choosing a suitable prior for σ can be difficult in practice; a Dirichlet process prior is introduced to provide a more flexible nonparametric mixture model.
- Nonparametric and empirical-likelihood extensions: The paper extends Bayesian mode regression with Dirichlet process mixtures and an empirical-likelihood formulation, while showing that improper uniform priors for β can produce proper joint posteriors.The nonparametric model relaxes dependence on the mode-uniform distribution assumption.
3 Numerical experiments
The numerical experiments evaluate parametric, empirical-likelihood, and nonparametric Bayesian mode regression on simulated and WECO data. Bayesian estimators recover parameters well, generally outperform classical alternatives, and yield narrower intervals in several settings.
- Experimental design: The experiments use two simulated examples and one real WECO productivity dataset to assess Bayesian mode regression.The WECO analysis examines productivity as a function of gender, physical-dexterity exam score, and education.
- Experimental design: PBMR is fitted across simulation cases, while ELBMR and NBMR provide additional comparisons in selected cases.PBMR and ELBMR use independent improper uniform priors; NBMR uses a truncated Dirichlet process mixture model.
- Simulation example 1: Absolute parameter-estimation biases under PBMR range from 0.01 to 0.26, while ELBMR and NBMR successfully recover the true β0 and β1 values.These results cover simulated settings with symmetric, skewed, and contaminated error distributions.
- Simulation example 1: ELBMR and NBMR produce smaller standard deviations and shorter confidence intervals than PBMR in the first simulation example.The reported comparison concerns both model parameters.
- Productivity of Western Electric Workers - WECO: For WECO productivity, the baseline female worker’s estimated mode productivity is 4.93 units, with higher mode productivity associated with exam score and education and lower levels for male workers.The analysis also finds that the NBMR results are similar to PBMR results but have much smaller confidence intervals.
4 Conclusions
The paper introduces a Bayesian mode regression framework with three approaches and reports theoretically justified, practically usable inference that performs competitively with classical methods.
- The framework includes parametric, nonparametric, and empirical likelihood-based Bayesian approaches.These approaches establish a fully Bayesian framework for direct mode regression inference.
- Bayesian mode regression fills a gap because mode regression had no Bayesian-perspective literature.
- The estimates are consistent and asymptotically normal under standard conditions, even when the likelihood is misspecified.
- The methods provide implementable inference tools and credible intervals irrespective of sample size.
- The numerical studies suggest that Bayesian mode regression estimates are strong competitors of classical mode regression estimates.
Appendix A
The appendix establishes conditions under which posterior moments and Bayesian mode regression estimates can be obtained for a full-rank design.
- For σ > 0 and m > p, the appendix gives the moments of the posterior distribution.
- The coefficient matrix X is assumed to have full rank p.
- A subset of p constraints can ensure bounded, nonzero coefficient magnitudes under the stated conditions.