Source-linked AI summary
On efficiency gains via augmenting a tiny sample with a massive auxiliary sample
Yen-Chi Chen
TL;DR
The paper asks how a massive but biased auxiliary sample can improve inference from a tiny target sample. It develops IPW and FL analyses using Tukey’s factorization, showing that FL can deliver full efficiency gain for some parameters while IPW can attain a parametric target-sample rate nonparametrically. The paper also characterizes key efficiency conditions and scope limitations.
Problem
The paper studies how to improve target-population inference when the target sample is tiny, the auxiliary sample is massive and biased, and the efficiency conditions are theoretically unclear.
Method
Using Tukey’s factorization, the paper studies IPW and FL for exponential families, mixtures, and nonparametric density estimation, with additional efficiency analyses for Ising models.
Results
FL can consistently estimate some target-model parameters as n1 →∞ with fixed n0, while IPW lacks consistency in that regime; nonparametric IPW density estimation reaches a parametric rate in n0.
Takeaways & Limitations
Full efficiency gain is driven by parameter orthogonality, whereas IPW remains useful for robustness to target-model misspecification and can achieve parametric target-sample rates nonparametrically.
Takeaways & Limitations
Full efficiency gain is not general for location-scale models, and the nonparametric IPW improvement remains limited by the target sample size rather than the full sample size.
Abstract
from arXiv · showhide
In this paper, we study the problem of augmenting a tiny target sample with a massive auxiliary sample. Utilizing Tukey's factorization, there are two popular approaches: the inverse probability weight (IPW) and the full-likelihood (FL) methods. We show that the IPW approach suffers from the limited target sample problem while the FL method may estimate some model parameters at the rate of the massive auxiliary sample size, a phenomenon we call full efficiency gain. We study the theory behind the full efficiency gain for exponential families and mixtures of exponential families. We also study the efficiency gain for the IPW method under a nonparametric procedure and show how it can achieve a parametric rate of the target sample size. As a side note, we also discuss how one may use FL to train neural network models simultaneously for both the target distribution and the odds model.
1. Introduction
The paper asks whether a massive but biased auxiliary sample can improve inference for a tiny target population. It contrasts IPW and FL, then studies when FL achieves full efficiency gain and when IPW improves nonparametrically.
- Motivation: The central problem is extracting information from a massive, biased auxiliary sample to improve inference on a target population represented by a small sample.Examples include rare astronomical objects and special epidemiological subpopulations.
- Approaches: IPW estimates an odds model before using weighted likelihood, whereas FL jointly optimizes the odds model and target distribution in one full likelihood.Both approaches use Tukey’s factorization to bridge the target and auxiliary populations.
- Motivating example: In a Gaussian-logistic example with n0 = 50 and n1 = 5000, FL has uniformly smaller MSE, with σ2 nearly 100 times more accurate than under IPW.The comparison uses 1,000 iterations with both models correctly specified.
- Theory: The paper explains full efficiency gain through parameter orthogonality, allowing some FL parameters to exploit the massive auxiliary sample while others improve only mildly.The variance parameter σ2 is orthogonal to the unidentifiable curve in the motivating example.
- Scope: The theory covers exponential families, mixtures of exponential families, and Ising network structure, while the IPW analysis shows nonparametric estimators can reach a parametric target-sample rate.The work focuses on efficiency conditions that remain theoretically obscure in related data-integration and covariate-shift settings.
2. Background and problem setup
The setup combines target and auxiliary observations distinguished by a population indicator, with the auxiliary sample potentially much larger. Tukey’s factorization expresses the target distribution through odds and auxiliary-distribution information, motivating weighting.
- Data structure: The variable X is observed with an indicator A, where A = 0 denotes the target population and A = 1 denotes the auxiliary population.The analysis begins in a univariate setting and can be generalized to multivariate data.
- Sample sizes: The target and auxiliary sample sizes are denoted n0 and n1, respectively, with total sample size n = n0 + n1.The regime of interest has the target fraction tending toward zero as the auxiliary sample dominates.
- Tukey’s factorization: Tukey’s factorization relates the target and auxiliary distributions through the odds of population membership given X.This density-ratio relationship permits information from auxiliary observations to be used for target-distribution learning.
3. Inverse probability weighting (IPW) approach
IPW estimates the odds from pooled data and plugs the estimated odds into a weighted target-likelihood procedure. Its effective precision remains constrained by the smaller target sample.
- Procedure: IPW first estimates the odds model, then substitutes the estimated odds into Tukey’s factorization to estimate the target distribution.For the Gaussian case, the resulting estimator uses a weighted likelihood based on target and auxiliary observations.
- Gaussian implementation: Under the Gaussian model, IPW estimates the target mean and variance by maximizing a weighted likelihood, equivalent to weighted sample mean and variance calculations.The same construction extends to other parametric likelihoods.
- Inefficiency: When n0 ≪ n1, odds-model uncertainty has effective order O_P(1/√n0), because the smaller dataset controls the effective sample size.Summing the weighted likelihood over all observations therefore does not remove the target-sample bottleneck.
- Inefficiency: IPW is not statistically consistent when n0 is fixed and n1 →∞, so no target-distribution parameter can be consistently estimated in that regime.Consistency is recovered only as the target sample size also tends to infinity.
4. Full-likelihood (FL) model
FL combines the target distribution and odds model into a joint likelihood, retaining information from both populations. In the Gaussian case, orthogonality lets the variance estimator exploit the massive auxiliary sample.
- Full-likelihood construction: FL jointly estimates the target distribution and odds model by maximizing their full likelihood rather than freezing estimated odds in a second stage.This modification can improve the effective sample size for some parameters.
- Likelihood decomposition: Tukey’s factorization yields a full-data likelihood that decomposes into target-distribution, odds-model, and marginal-population components.The marginal probability is data-independent for Fisher-information calculations and can be omitted from the score analysis.
- Gaussian-logistic model: In the Gaussian-logistic model, FL estimates all four parameters simultaneously under the correctly specified joint model.The resulting estimator is analyzed through its asymptotic covariance matrix.
- Efficiency gain: The full-likelihood analysis identifies settings in which some parameters improve at rates associated with the auxiliary sample size rather than only the target sample size.A related asymptotic result also applies when the target-to-auxiliary sample ratio tends to zero.
- Efficiency gain: The variance parameter can achieve full efficiency gain because it is orthogonal to the unidentifiable curve, allowing FL to exploit the massive auxiliary sample.This explains why variance precision improves much more than precision for µ, α, and β.
Because n0n
Under a fixed target-to-auxiliary ratio, the variance parameter’s convergence depends on the total sample size and achieves full efficiency. With fixed target size and growing auxiliary size, FL can consistently estimate this parameter.
- The variance parameter’s convergence is improved drastically and depends on the total sample size.
- With fixed n0 and n1 →∞, FL consistently estimates the variance parameter in the Gaussian-logistic model.
5. The geometry of full efficiency gain
Full efficiency arises when target-distribution parameters are orthogonal to the auxiliary sample’s unidentifiable directions. The resulting fast rate applies to selected parameters in exponential families and mixtures, but not generally to location-scale parameters such as the Uniform scale.
- 5. The geometry of full efficiency gain: Under n0/n1 →0, λ1 is asymptotically normal at the fast total-sample rate O(1/√n).
- 5. The geometry of full efficiency gain: The auxiliary likelihood identifies λ1 while λ2 and β remain constrained along an unidentifiable curve.
- 5. The geometry of full efficiency gain: Reparameterization isolates a rank-k variance bottleneck and rescues orthogonal directions from the target-sample rate.
- 5. The geometry of full efficiency gain: Including an orthogonal feature in the logistic odds model creates Fisher-information confounding and forces its parameter back to the target-sample rate.
- 5. The geometry of full efficiency gain: In mixtures, orthogonal secondary parameters and λk,1 achieve the total-sample rate, while parallel parameters and mixture weights retain the target-sample bottleneck.
- 5. The geometry of full efficiency gain: Full efficiency for Gaussian variance is special; for the Uniform distribution, selection intertwines scale and odds, preventing auxiliary identification of scale.
6. Robustness and efficiency gain of the IPW approach
IPW is more robust to target-model misspecification, while FL can be more efficient when its parametric assumptions are correct. Nonparametric IPW renormalization reaches a target-sample parametric rate, but not the full-sample rate.
- 6. Robustness and efficiency gain of the IPW approach: When the target model is misspecified but the odds model is correct, IPW remains consistent, whereas FL fails through model feedback.
- 6. Robustness and efficiency gain of the IPW approach: FL achieves the Cramér-Rao lower bound and can sharply reduce variance, but is less robust than IPW.
- 6. Robustness and efficiency gain of the IPW approach: The target-sample augmented IPW KDE cannot improve on the target-only KDE rate because half of it uses the target-only estimator.
- 6. Robustness and efficiency gain of the IPW approach: ASR-KDE renormalizes the auxiliary KDE using the estimated odds model and normalizing constant.
- 6. Robustness and efficiency gain of the IPW approach: The renormalized nonparametric estimator can achieve a parametric convergence rate in n0 when n0 ≪ n1 ≍ n.
- 6. Robustness and efficiency gain of the IPW approach: ASR-KDE and IPW-KDE are asymptotically equivalent under the stated local expansion argument.
- 6. Robustness and efficiency gain of the IPW approach: This improvement remains limited by n0 and depends critically on a correctly specified odds model.
7. Training a full-likelihood model with neural nets
The paper proposes using FL to train neural networks for the target distribution and odds model jointly. The procedure relies on sampling from the target model and evaluating score-function gradients, but its practical efficiency remains unestablished.
- FL can train neural networks simultaneously for the target distribution and odds model.
- The method assumes that samples can be generated from the modeled target distribution for Monte Carlo approximation.
- Training requires gradients with respect to both parameter sets, while direct evaluation of the full-likelihood is unnecessary.
- REINFORCE can provide the gradient approximation but is known for high Monte Carlo variance; reparameterization is suggested as a remedy.
- The practical efficiency of this neural-network training procedure remains unclear and is left for future work.
8. Simulations
The simulations examine when FL can exploit massive auxiliary samples, focusing on Gaussian, Ising, mixture, and nonparametric density-estimation settings. Across parametric experiments, correctly restricted FL accelerates estimation for parameters orthogonal to the odds model, while target-sample limitations remain for entangled parameters and misspecified methods.
- 5D Gaussian and multiple contrasts: Correct Restricted FL achieves the fast convergence rate for Gaussian variance and three orthogonal contrasts, while individual means remain bottlenecked at O(1/n0).The contrasts are µ1 − 2µ2, µ4 − µ5, and µ1 − µ2 − µ4.
- 5D Gaussian and multiple contrasts: Efficiency gains require the odds model to be correctly specified and strictly restricted along orthogonal directions.Unrestricted FL destroys the fast rate through Fisher information confounding, while incorrectly restricted FL produces severe bias and massive MSE.
- Efficiency gain in the Ising model: In the Ising experiment, FL reduces interaction-parameter MSE by up to 27.6× versus IPW as n1 increases to 100,000 with n0 = 200.IPW reaches a stable limit once n1 = 20,000 for N = 2, whereas FL continues improving.
- Scaling analysis for Gaussian mixture models: For the Gaussian mixture, FL’s orthogonal variance estimate improves with n1, reaching up to 493.2× efficiency gain over IPW, while entangled parameters plateau.The mixture weight and centers are entangled with the odds tilt; µ1 and w1 do not improve much with n1.
- Nonparametric density estimation via IPW-KDE: IPW-KDE and ASR-KDE achieve the parametric O(n0^-1) rate when n0 increases, but all KDE methods show little improvement when only n1 increases.TAIPW-KDE, IPW-KDE, and ASR-KDE receive only minor gains from improved auxiliary KDE accuracy because target-sample uncertainty dominates.
9. Discussion
The paper contrasts FL’s full efficiency gain with IPW’s target-sample bottleneck and recommends using the methods complementarily according to model suitability. It also reports robustness and scope considerations for both approaches.
- Discussion: When n0 is fixed and n1 →∞, IPW is inconsistent, whereas FL can consistently estimate some target-model parameters.IPW remains robust to target-model misspecification when n0 →∞, and nonparametric IPW density estimation can attain a target-sample parametric rate.
- Discussion: FL achieves full efficiency when target-distribution parameters remain untouched by the odds model, allowing both samples to inform those parameters.This mechanism is developed for exponential families and mixtures, with orthogonality to the unidentifiable curve identified as critical.
- Discussion: IPW achieves a parametric efficiency gain with IPW-KDE or ASR-KDE because the massive auxiliary sample permits accurate auxiliary-density estimation.The remaining bottleneck is uncertainty in the odds estimator rather than nonparametric density estimation itself.
- Discussion: The paper recommends choosing among target-only, FL, and IPW estimators based on odds-model feasibility and the appropriateness of exponential-family target models.If no reasonable odds model exists, target-only parametric or nonparametric inference is recommended.
- Discussion: In Gamma and Beta examples, the parameter structurally independent of the odds mechanism receives the fast rate, while the confounded parameter remains target-sample limited.Linear odds preserve Gamma shape, logarithmic odds preserve Gamma scale, and statistic-specific Beta odds preserve the complementary parameter.
- Discussion: The Gaussian simulation uses n0 = 50, n1 = 5000, and 1,000 Monte Carlo datasets to compare IPW and FL under a standard-Gaussian target.The misspecification figure instead compares odds-parameter MSE when the target is Uniform but FL assumes Gaussian.
B.2. Simulation: model mis-specification
The simulations examine model misspecification and Ising interactions under increasing auxiliary sample sizes. They show that FL exploits auxiliary data for structurally invariant parameters, while IPW remains target-sample limited.
- B.2. Simulation: model mis-specification: Under Uniform target data with FL incorrectly assuming Gaussian, IPW remains consistent for the correctly specified odds model, whereas FL loses consistency.The FL estimate of β has substantially higher MSE than IPW in this misspecified setting.
- B.2. Simulation: model mis-specification: In the Ising study, interactions are excluded from the odds model and therefore achieve the fast O(1/n) rate predicted by Theorem 2.The study varies dimensions N = 2, 3, 5 and uses sparse interactions for N = 5.
- B.2. Simulation: model mis-specification: The scaling experiments fix n0 at 100 or 200 while increasing n1 from 5,000 to 100,000 over 500 Monte Carlo iterations.Figures 6 and 7 report interaction-parameter MSE across dimensions for the two target-sample sizes.
- B.2. Simulation: model mis-specification: With n0 fixed, FL interaction-parameter MSE continues decreasing at the fast O(1/n) rate as auxiliary data increase.This validates Theorem 2 for structurally invariant parameters.
- B.2. Simulation: model mis-specification: IPW MSE reaches a variance floor as n1 approaches 100,000, and doubling n0 from 100 to 200 halves that floor.The result confirms that IPW remains bounded by target-sample information.
D.2. Results
Results across Gaussian mixtures and the supporting Fisher-information analysis show that FL can achieve fast rates for invariant parameters, whereas confounded parameters retain the target-sample bottleneck.
- D.2. Results: For invariant GMM variance σ1^2, IPW MSE stagnates near 0.16 while FL improves nearly 535× at n1 = 100,000.Confounded means and mixture weights retain an O(1/n0) floor for both methods.
- D.2. Results: Exponential tilting changes both Gaussian component parameters and component weights in a mixture model.The mass factor favors components with higher means when β > 0, creating cross-component effects beyond mean shifts.
- D.2. Results: If β or mean separation is large, the minority component can undergo exponential starvation; with β = 2, its auxiliary weight falls below 0.03%.Even 100,000 auxiliary observations may then provide only a few minority-component points, causing finite-sample instability.
- D.2. Results: The Fisher-information derivation uses score covariance, conditioning on A, and Schur complements to isolate fast and target-limited parameter blocks.The intercept’s rank-one contribution is canceled during inversion, leaving the non-intercept covariance determined by conditional score variance.
- D.2. Results: The orthogonal statistic S2,⊥ is absent from the odds mechanism, so its corresponding λ2,⊥ parameter achieves the fast total-sample rate.The scaled variance in these directions is zero under the target-sample bottleneck.
E.4. Proof of Theorem 4
The proof establishes fast convergence for mixture-component parameters orthogonal to the odds mechanism by analyzing the auxiliary likelihood’s null space and the perturbed Fisher information matrix.
- E.4. Proof of Theorem 4: The mixture parameter vector is partitioned into component parameters, mixture weights, intercept, and odds parameters, with localized component scores weighted by posterior responsibilities.The full score is organized into component blocks before analyzing the Fisher information.
- E.4. Proof of Theorem 4: The auxiliary likelihood has a k-dimensional unidentifiable subspace generated by coordinated shifts of β, component parameters, and mixture weights.The auxiliary density remains unchanged along these perturbations, producing a corresponding Fisher-information null space.
- E.4. Proof of Theorem 4: λ1 is orthogonal to the auxiliary null space because its sufficient statistic is structurally independent of the odds statistic.Its null-matrix block is exactly zero, so the missingness penalty does not affect λ1.
- E.4. Proof of Theorem 4: Rotating into null-space and orthogonal-complement coordinates shows that only parameters with nonzero null-space components receive the O(1/ε) variance penalty.The remaining inverse-information blocks stay bounded at O(1).
- E.4. Proof of Theorem 4: Because λ1’s inverse Fisher information is O(1), standard M-estimation theory gives asymptotic normality at the fast total-sample rate.This completes the proof for the fast component-parameter block.
- E.4. Proof of Theorem 4: For the nonparametric estimator, the proof separately bounds the unnormalized density ratio and normalization constant before applying a Delta-method expansion.Compact support, bounded odds, and smooth auxiliary density assumptions control the denominator and integrated errors.