Source-linked AI summary
Models for Paired Comparison Data: A Review with Emphasis on Dependent Data
Manuela Cattelan
TL;DR
Paired comparison analysis requires models that accommodate covariates, ordinal outcomes, and dependence, but dependent-data models create numerical and identification difficulties. This paper reviews these developments, explores pairwise likelihood for dependent comparisons, and compares estimation methods through simulation, using university preferences for illustration.
Problem
Dependent paired-comparison models are more realistic but present numerical and identification difficulties, while model assessment can also face sparse contingency tables and nonstandard test distributions.
Method
The paper reviews model extensions, proposes pairwise likelihood estimation for dependent comparisons, compares it with maximum likelihood and limited information estimation in simulation, and illustrates the methods with university data.
Results
Pairwise likelihood performs quite well relative to limited information estimation in the reported simulation, although covariance-parameter estimation is more problematic than worth-parameter estimation.
Takeaways & Limitations
The review emphasizes pairwise likelihood as a practical estimation approach for dependent paired-comparison models and discusses implementation issues alongside applications.
Takeaways & Limitations
Dependent Thurstonian models remain sensitive to identification restrictions, which can change the estimated covariance-matrix class, standard errors, and significance of other parameters.
Abstract
from arXiv · showhide
Thurstonian and Bradley-Terry models are the most commonly applied models in the analysis of paired comparison data. Since their introduction, numerous developments have been proposed in different areas. This paper provides an updated overview of these extensions, including how to account for object- and subject-specific covariates and how to deal with ordinal paired comparison data. Special emphasis is given to models for dependent comparisons. Although these models are more realistic, their use is complicated by numerical difficulties. We therefore concentrate on implementation issues. In particular, a pairwise likelihood approach is explored for models for dependent paired comparison data, and a simulation study is carried out to compare the performance of maximum pairwise likelihood with other limited information estimation methods. The methodology is illustrated throughout using a real data set about university paired comparisons performed by students.
1. INTRODUCTION
Paired comparison data are widely used because people can compare objects in pairs more easily than rank lists. The paper reviews recent extensions of the Thurstone and Bradley–Terry models, emphasizing dependent data and implementation.
- Paired comparison data record choices between couples of objects and arise especially when human judgments are involved.
- The compared entities may include beverages, lotteries, players, physical stimuli, and other objects judged by subjects or agents.
- The paper updates earlier reviews by focusing on extensions of the Thurstone and Bradley–Terry models.
- The review covers ordinal outcomes, explanatory variables, dependent observations, pairwise likelihood estimation, simulation comparisons, software, and a university-preference illustration.
2. INDEPENDENT DATA
Independent paired-comparison models represent object worth through pairwise preference probabilities and can be extended to ordinal outcomes and covariates. The university example illustrates estimation, ranking, uncertainty assessment, and subject–object interactions.
- 2.1 Traditional Models: Traditional models assign each object a worth parameter, with pairwise preference probability determined by the difference in worths.
- 2.1 Traditional Models: The Thurstone model uses a normal cumulative distribution, whereas the Bradley–Terry model uses a logistic cumulative distribution.
- 2.1 Traditional Models: Reference-object constraints identify worth parameters as differences relative to the selected reference object.
- 2.1 Traditional Models: Quasi-standard errors provide an alternative way to approximate inference on contrasts without reporting the full covariance matrix.
- 2.1 Traditional Models: The university study compares six universities using 303 students, with aggregated results reported for 15 paired comparisons.
- 2.1 Traditional Models: Under the fitted Thurstone model, Stockholm is least preferred and London most preferred; London is estimated preferred to Paris with probability 0.66.
- 2.4 Explanatory Variables: Subject–object interaction effects allow covariates such as language knowledge to modify preferences for particular universities, whereas non-interacting subject covariates cannot be included.
- 2.4 Explanatory Variables: For a student knowing English and French, the estimated probabilities for London, no preference, and Paris are 0.46, 0.13, and 0.41, respectively.
3.1 Intransitive Preferences
Independent paired-comparison models impose strong stochastic transitivity, but real choices can be systematically intransitive. Modeling dependence can accommodate such comparisons while still producing an overall ranking.
- Transitivity: Independent Bradley–Terry and Thurstone models satisfy strong stochastic transitivity, requiring consistent probabilistic preferences across objects.Strong stochastic transitivity requires πik ≥ max(πij,πjk).
- Intransitive Preferences: Systematic intransitivity can arise when different aspects of the same objects dominate different comparisons.Such choices form circular triads, where i beats j, j beats k, but k beats i.
- Alternative Extensions: Multidimensional Bradley–Terry extensions represent object worths on multiple dimensions but do not provide a final ranking of all objects.The cited two-dimensional model represents worth parameters on a plane.
- Dependent Comparisons: Modeling dependence among comparisons allows systematic intransitivity while yielding a ranking of all objects.This approach became feasible through developments in inferential techniques for dependent data.
3.2 Multiple Judgment Sampling
Multiple judgments by the same person create dependence among paired comparisons, motivating correlated Thurstonian and choice-model extensions. These models improve realism but introduce identification and estimation difficulties.
- Motivation: When people make all paired comparisons, comparisons by the same person are more realistically treated as dependent rather than independent.The setting is multiple judgment sampling, where S people make all N comparisons.
- Thurstonian Models: Thurstonian models represent compared stimuli as normally distributed latent variables, with covariance structures capturing dependence and judge-level variation.The original Thurstone model includes correlation among observations; within-judge variability is modeled separately.
- Dependent Structures: Extensions add pair-specific errors and random effects to represent variation associated with particular comparisons or judging mechanisms.Takane’s formulation uses a design matrix for paired comparisons and pair-specific errors independent across subjects.
- Identification: Reduced-parameter dependent Thurstonian models still require identification restrictions, including constraints on covariance and worth parameters.The unrestricted model is over-parametrized, and the original covariance specifications cannot be recovered uniquely from paired-comparison data.
- Choice Models: Choice models formulate decisions through utility maximization and can express paired comparisons as choices between a preferred alternative and the remaining options.In this representation, paired comparisons do not literally occur, so inconsistent choices cannot be observed.
- Choice Models: Logit and nested-logit models are easy to estimate but cannot account for random taste variation or temporally correlated unobserved factors.Multivariate probit models allow those features, but their estimation is not straightforward.
- Choice Models: Combining stated and revealed preferences can create identification problems and require high-dimensional integral approximations.These difficulties arise when additional disturbances and unobserved covariates are included.
3.3 Object-Related Dependencies
Object-related dependence arises when repeated comparisons share an object, such as in animal contests. Random object effects and covariates provide a multivariate probit formulation for these data.
- Motivation: Comparisons can be correlated when the same object appears in multiple pairings, including contests involving the same animal.This dependence does not require a human judge.
- Random Effects: Object-specific random effects capture unobserved quantities affecting comparison results alongside observed animal characteristics.The random effect Ui is assumed to have mean zero.
- Model Structure: The object-related model combines covariates, object effects, independent errors, and a paired-comparison design matrix in a multivariate probit model.The design matrix records which comparisons are observed, which need not include every possible pair.
- Data Structure: Compared with multiple-judgment sampling, object-based applications may involve many objects and lack independent replications of every comparison, increasing inferential difficulty.This distinction is especially relevant in sports tournaments and animal-behavior data.
3.4 Inference
Inference for dependent paired-comparison models involves a trade-off between full-information accuracy and computational feasibility. The paper compares multivariate-probit, limited-information, and pairwise-likelihood approaches, finding that pairwise likelihood can perform well while avoiding some high-dimensional integration difficulties.
- Full-information methods: Multivariate-probit inference requires approximating S integrals of dimension N = n(n−1)/2, which grows quadratically with the number of objects.Several proposed approximations involve randomness, slow computation, or high computational cost at large dimensions.
- Limited-information estimation: Limited-information estimation uses univariate win proportions, bivariate win proportions, and a final optimization over the model parameters.The procedure estimates thresholds first, tetrachoric correlations second, and structural parameters third.
- Pairwise likelihood: Pairwise likelihood estimates dependent-comparison models by multiplying likelihood contributions for pairs of observations relative to individual judges.It is a special case of composite likelihood.
- Scope and limitations: Object-related dependencies were excluded from the simulation study, although pairwise likelihood can still be employed when the number of objects is large and the number of subjects is small.The asymptotic behavior of the maximum pairwise-likelihood estimator remains more problematic for long sequences of dependent observations, and theoretical results are lacking for all paired comparisons.
- Simulation studies: In one simulation setting, all methods performed comparably well, while empirical coverage for covariance parameters was systematically below nominal levels.The reported coverage results suggest that bootstrap procedures may be needed for more accurate coverage probabilities.
- Simulation studies: Maximum likelihood performed best in the second simulation setting, but optimization was not always straightforward because the Hessian sometimes was not negative definite.Pairwise likelihood performed quite well relative to limited-information estimation at S = 100.
- Simulation studies: Pairwise likelihood had lower bias and shorter confidence intervals than limited-information estimation in the second simulation setting.Maximum percentage bias fell from 44.1% to 15.4% for limited-information medians, versus 16.1% to 4% for pairwise-likelihood medians.
4. SOFTWARE
R packages support classical paired-comparison models and several extensions, but no single package combines the desired covariates, link functions, categories, and dependent-data models.
- Several R packages facilitate fitting classical paired-comparison models and, in some cases, more complicated models.The reviewed packages include tools for elimination-by-aspects, Bradley–Terry, ordinal, covariate, and subject-partitioning models.
- The eba package fits elimination-by-aspects models, which reduce to Bradley–Terry models when each object has one relevant aspect.
- The prefmod package fits log-linear Bradley–Terry models and allows ordinal comparisons, but reduces outcomes to two or three categories.
- BradleyTerry2 supports logit, probit, and cauchit links, comparison-specific covariates, and several estimation methods.
- No available package fits the Section 3.2.1 dependent-data models, although pairwise likelihood implementation is straightforward through bivariate normal probabilities.
5. CONCLUSIONS
The review finds established methods for independent paired-comparison data but continuing methodological and computational challenges for dependent, incomplete, and socially influenced comparisons.
- The review covers many Bradley–Terry and Thurstone extensions but leaves multidimensional object evaluations and other aspects untreated.
- Dependent-data models face identification problems because different restrictions can imply different covariance-matrix classes.
- Missing comparisons complicate limited-information estimation and goodness-of-fit testing, whose quadratic statistics assume complete comparison designs.
- Large designs may create missing data, while prolonged judgments may require accounting for subject fatigue or elapsed time.
- Future models may need to represent influence among subjects, leaders, and social or cultural context, increasing model complexity.
- Object-related dependencies remain difficult because comparisons are jointly dependent and often unbalanced, complicating standard-error computation and model assessment.