Source-linked AI summary

Multiple Source Adaptation and the Renyi Divergence

Yishay Mansour, Mehryar Mohri, Afshin Rostamizadeh

arXiv:1205.2628v1cs.LGstat.ML

TL;DR

Multiple source adaptation addresses learning when training and test distributions differ. This paper develops Renyi-divergence-based guarantees for distribution-weighted source combinations across broader target, distribution-knowledge, source-distribution, and labeling-function settings, with theoretical and empirical support.

  • Problem

    Multiple source adaptation concerns learning when training and test distributions differ, extending beyond the shared-distribution assumption underlying standard generalization analyses.

  • Method

    The paper analyzes distribution-weighted combinations using Renyi divergence for arbitrary targets, known or unknown target distributions, approximate source distributions, and differing source labeling functions.

  • Results

    Theoretical and empirical results show that distribution-weighted combinations can provide effective multiple-source adaptation solutions, including when target distributions are not mixtures of source distributions.

  • Takeaways & Limitations

    Renyi divergences provide the paper’s natural distribution-distance framework for analyzing practical multiple-source adaptation problems.

  • Takeaways & Limitations

    The analysis assumes bounded losses and, with multiple labeling functions, requires each source function to be close to the target function on the target distribution.

Abstract

from arXiv · show

This paper presents a novel theoretical study of the general problem of multiple source adaptation using the notion of Renyi divergence. Our results build on our previous work [12], but significantly broaden the scope of that work in several directions. We extend previous multiple source loss guarantees based on distribution weighted combinations to arbitrary target distributions P, not necessarily mixtures of the source distributions, analyze both known and unknown target distribution cases, and prove a lower bound. We further extend our bounds to deal with the case where the learner receives an approximate distribution for each source instead of the exact one, and show that similar loss guarantees can be achieved depending on the divergence between the approximate and true distributions. We also analyze the case where the labeling functions of the source domains are somewhat different. Finally, we report the results of experiments with both an artificial data set and a sentiment analysis task, showing the performance benefits of the distribution weighted combinations and the quality of our bounds based on the Renyi divergence.

1 Introduction

Multiple source adaptation addresses prediction on a target domain when labeled data are scarce and several related source domains are available. This paper uses Rényi divergence to broaden distribution-weighted guarantees beyond source mixtures and to cover approximate source distributions and differing labeling functions.

  • Problem: Multiple source adaptation combines hypotheses learned from several labeled domains to predict on a target domain with little or no labeled data.The learner receives source distributions and hypotheses, then constructs a target-oriented hypothesis; the target distribution may be known or unknown.
  • Prior work: Simple convex combinations can incur classification error of half even when every source hypothesis is error-free on its own domain.Distribution-weighted combinations instead weight source hypotheses according to source distributions and achieve an ε loss guarantee for mixtures of source distributions.
  • Contributions: The paper bounds distribution-weighted combination loss for arbitrary target distributions using the maximum source loss and Rényi divergence to mixtures of source distributions.This extends earlier guarantees that applied only when the target was a mixture of the source distributions.
  • Contributions: The analysis also covers approximate source distributions, with guarantees depending on divergence from the true distributions.The paper studies how replacing each true source distribution with an approximation affects the resulting loss bound.
  • Contributions: When source labeling functions differ, the guarantees extend under the assumption that those functions remain close to the target function on the target distribution.The closeness is required on the target distribution rather than on the individual source distributions.
  • Experiments: Experiments on artificial and sentiment-analysis data show performance benefits for distribution-weighted combinations and support the quality of Rényi-divergence-based bounds.The experiments include a sentiment-analysis task and a target that is not a mixture of the base domains, where the distribution-weighted combination outperforms either base hypothesis.

2 Preliminaries

The preliminaries formalize loss, the multiple-source adaptation setting, combining rules, and bounded convex losses. They then introduce Rényi entropy and divergence, including divergence to the class of mixtures of source distributions.

  • Multiple Source Adaptation Problem: The adaptation setting supplies a target distribution, k source distributions, and source hypotheses with source loss at most ε, then seeks a low-loss target hypothesis.A combining rule maps the source hypotheses to one hypothesis evaluated under the target distribution.
  • Multiple Source Adaptation Problem: Distribution-weighted combining uses a parameter z in the simplex to construct a hypothesis from the source hypotheses.The resulting family is denoted H in the paper.
  • Assumptions: The loss function is assumed non-negative, convex in its first argument, and bounded, with absolute and 0-1 loss given as examples.These assumptions support the analysis of the combining rules and their target-domain loss.
  • Rényi Entropy and Divergence: Rényi entropy is parameterized by α, with limiting cases including support-size entropy at α=0, Shannon entropy at α=1, and collision-based entropy at α=2.Rényi entropy is described as non-negative and decreasing as α increases.
  • Rényi Entropy and Divergence: Rényi divergence measures distribution discrepancy, coincides with KL divergence at α=1, and is zero exactly when the two distributions are identical.The paper also uses its base-2 exponential form for some analyses.
  • Rényi Entropy and Divergence: For a class of distributions, the paper defines divergence as the infimum over that class and applies it to all mixtures of the k source distributions.The mixture class consists of distributions Qλ formed with simplex weights λ.

3 Multiple Source Adaptation Guarantees

The paper derives Renyi-divergence guarantees for combining multiple source hypotheses under known and unknown target distributions, and establishes a nearly tight lower bound. It also analyzes target-independent r-norm combinations and recovers prior guarantees for target mixtures.

  • 3.1 Known Target Distribution: For any target distribution P, a distribution-weighted hypothesis based on the source mixture minimizing Dα(P∥Qλ) has a Renyi-divergence-dependent loss guarantee.The construction selects λ by minimizing divergence between P and Qλ, then combines source hypotheses using those weights.
  • 3.1 Known Target Distribution: When P is a mixture of the source distributions, the divergence equals one and a distribution-weighted combination achieves loss at most ǫ.This recovers the earlier multiple-source result as a special case.
  • 3.2 Unknown Target Distribution: With an unknown target distribution, the paper constructs a hypothesis without target knowledge whose loss bound for any P is similar to the known-target guarantee.The construction depends on source distributions and matching hypotheses, while the bound uses the divergence between P and an optimized source mixture.
  • 3.3 Lower Bound: The lower bound shows that for a suitable target P, any fixed source hypothesis with source loss ǫ can incur target loss determined by its Renyi divergence from Q.The bound is nearly tight relative to the upper bound from Lemma 1.
  • 3.4 Simple Combining Rules: Target-independent r-norm combinations achieve LP(hr-norm, f) ≤ ρkǫ when P is (ρ, r)-norm-bounded by the source distributions.This family includes natural rules such as uniform and maximum combinations.

4 Approximate Distributions

The paper extends multiple-source adaptation guarantees to approximate source distributions, using divergences between approximate and true distributions to control the resulting loss bounds.

  • 4 Approximate Distributions: The learner replaces each true source distribution Q_i with an approximation bQ_i and retains source hypotheses whose losses are measured under Q_i.The analysis treats known and unknown target-distribution cases separately.
  • 4 Approximate Distributions: The approximate method selects bλ by minimizing D_α(P∥bQ_µ) and constructs the combination using the approximate source distributions.This modifies the corresponding procedure based on the true source distributions.
  • 4 Approximate Distributions: The analysis establishes supporting divergence and norm-boundedness lemmas, including a mixture relation and a triangle inequality-like property with an increased parameter.These lemmas supply the ingredients for the main approximate-distribution bounds.
  • 4 Approximate Distributions: Theorem 13 bounds performance using both the divergence between P and mixtures of true distributions and divergences between approximate and true source distributions.The theorem compares λ, minimizing divergence to a true-source mixture, with bλ, minimizing divergence to an approximate-source mixture.

5 Multiple Target Functions

The paper extends multiple-source adaptation to settings with distinct source labeling functions, assuming they remain close to the target labeling function on the target distribution.

  • 5 Multiple Target Functions: Source labeling functions f_i may differ, but the analysis assumes each is within δ loss of the target function f under P.Closeness is required on the target distribution rather than on each source distribution.
  • 5 Multiple Target Functions: Under a convex loss satisfying the triangle inequality, Theorem 16 gives a bound for combining source hypotheses across these distinct labeling functions.The bound applies for any mixture parameter λ.
  • 5 Multiple Target Functions: A relaxed β-inequality yields a corresponding extension, with the bound adjusted by the factor β.This broadens the result beyond losses obeying the exact triangle inequality.

6 Experiments

Experiments on artificial and sentiment-analysis data evaluate distribution weighted combinations, showing agreement with the Rényi-divergence analysis and improved performance over base hypotheses.

  • 6 Experiments: The artificial dataset uses Gaussian mixtures with a quadrant-based labeling function to test source combinations against a target containing all four components.Distribution weighting selects the appropriate base hypothesis according to the input quadrant.
  • 6 Experiments: The artificial-data error curve follows the same shape as the Rényi-divergence curve across mixture parameter λ, as predicted by the bounds.The endpoint values λ=0 and λ=1 recover the two basic hypotheses.
  • 6 Experiments: Distribution weighted combinations significantly improve sentiment-analysis MSE even when each base domain is relatively powerful.The experiment uses four product-review domains and evaluates mixtures on held-out data.
  • 6 Experiments: When the target is not a mixture of the base domains, distribution weighting performs significantly better than either base hypothesis.The comparison uses two single-domain base hypotheses and a target formed from the other two domains.

7 Conclusion

The paper concludes that distribution weighted combinations are effective for multiple-source adaptation, while Rényi divergences provide the natural distance measure for its guarantees.

  • 7 Conclusion: The theoretical and empirical results indicate that distribution weighted combinations can solve multiple-source adaptation problems, including real-world applications.The conclusion also covers approximate distributions and multiple labeling functions.
  • 7 Conclusion: The analysis identifies the family of Rényi divergences as the appropriate distribution distance for these adaptation results.This conclusion summarizes the role of Rényi divergences throughout the theoretical analysis.
Loading 1205.2628v1…