Source-linked AI summary

Offline Evaluation Measures of Fairness in Recommender Systems

Theresia Veronika Rampisela

arXiv:2604.25032v1cs.IR

TL;DR

Fairness measures for recommender systems lack robust analysis of their interpretability, empirical behavior, and computability. This thesis theoretically and empirically investigates these measures, develops corrections, and finds substantial agreement in fairness rankings despite differing and sometimes problematic score ranges.

  • Problem

    Existing recommender-system fairness measures lack sufficient analysis of their robustness, score interpretation, empirical distributions, and computability.

  • Method

    The thesis combines theoretical and empirical investigations with corrected measures that bound scores, handle edge cases, and address identified limitations.

  • Results

    Most fairness measures strongly agree when ranking recommenders, although their score ranges differ and some measures are incomparable, incomputable, or constant-valued.

  • Takeaways & Limitations

    Fairness evaluation requires interpreting measure ranges and considering multiple fairness types because existing measures can be redundant or conflicting.

  • Takeaways & Limitations

    Non-realisability remains unresolved when best and worst fairness scores require impractical constrained optimization on large recommender-system datasets.

Abstract

from arXiv · show

The evaluation of recommender system fairness has become increasingly important, especially with recent legislation that emphasises the development of fair and responsible artificial intelligence. This has led to the emergence of various fairness evaluation measures, which quantify fairness based on different definitions. However, many of such measures are simply proposed and used without further analysis on their robustness. As a result, there is insufficient understanding and awareness of the measures' limitations. Among other issues, it is not known what kind of model outputs produce the (un)fairest score, how the measure scores are empirically distributed, and whether there are cases where the measures cannot be computed (e.g., due to division by zero). These issues cause difficulty in interpreting the measure scores and confusion on which measure(s) should be used for a specific case. This thesis presents a series of papers that assess and overcome various theoretical, empirical, and conceptual limitations of existing recommender system fairness evaluation measures. We investigate a wide range of offline evaluation measures for different fairness notions, divided based on the evaluation subjects (users and items) and for different evaluation granularities (groups of subjects and individual subjects). Firstly, we perform theoretical and empirical analysis on the measures, exposing flaws that limit their interpretability, expressiveness, or applicability. Secondly, we contribute novel evaluation approaches and measures that overcome these limitations. Finally, considering the measures' limitations, we recommend guidelines for the appropriate measure usage, thereby allowing for more precise selection of fairness evaluation measures in practical scenarios. Overall, this thesis contributes to advancing the state-of-the-art offline evaluation of fairness in recommender systems.

Resum´e

Afhandlingen undersøger robustheden og begrænsningerne ved offline fairness-evalueringsmetoder i anbefalingssystemer. Den udvikler forbedrede evalueringsmetoder og retningslinjer for deres anvendelse på tværs af brugere, genstande, grupper og individer.

  • Problem: Eksisterende fairness-evalueringsmetoder er ofte foreslået og anvendt uden tilstrækkelig robusthedsanalyse, hvilket begrænser forståelsen af deres resultater og anvendelighed.Problemerne omfatter blandt andet fortolkning, empiriske scorefordelinger og tilfælde, hvor mål ikke kan beregnes.
  • Metode og dækningsområde: Afhandlingen analyserer teoretiske, empiriske og konceptuelle begrænsninger ved et bredt udvalg af offline fairness-mål for brugere og genstande på gruppe- og individniveau.Undersøgelsen dækker forskellige fairness-definitioner og evalueringssubjekter.
  • Bidrag: På baggrund af analyserne udvikler afhandlingen nye evalueringsmetoder og mål, der afhjælper identificerede begrænsninger.Bidragene sigter mod mere præcis og anvendelig offline evaluering af fairness i anbefalingssystemer.
  • Anvendelsesretningslinjer: Afhandlingen foreslår retningslinjer for valg og anvendelse af fairness-mål, så de kan udvælges mere præcist i praktiske scenarier.Retningslinjerne tager højde for målenes identificerede begrænsninger.
  • Samlet bidrag: Samlet bidrager afhandlingen til at fremme offline evaluering af fairness i anbefalingssystemer.Bidraget omfatter både analyse af eksisterende mål og udvikling af tilgange, der håndterer deres begrænsninger.

Executive Summary

This thesis examines theoretical, empirical, and conceptual limitations in offline fairness measures for recommender systems, develops corrective approaches, and recommends appropriate measure usage. It addresses ambiguity in score interpretation, poorly understood empirical behavior, and measures that omit intended fairness aspects.

  • Theoretical limitations: Unknown score ranges and unattained fairness endpoints make it difficult to determine how close recommendations are to fairest or most unfair scenarios.Ambiguous equations and omitted range information can leave maximum or minimum achievable scores unknown, making measure scores difficult to interpret.
  • Empirical limitations: Fairness measures can have poorly understood score distributions, conflicting rankings, stability issues, efficiency costs, and limited expressiveness or redundancy.These empirical limitations are examined across the thesis, with stability studied in Paper 3 and efficiency in Papers 3–5.
  • Conceptual limitations: Measures may fail to quantify intended fairness concepts when they assess disparity across individuals without accounting for their similarity.Individual fairness is commonly defined as equal treatment of similar individuals [63].
  • Objectives: The thesis assesses existing fairness measures’ theoretical, empirical, and conceptual limitations, designs approaches or measures to address them, and recommends appropriate usage.Its overall aim is to improve offline evaluation of fairness in recommender systems.
  • Thesis contributions: Papers 1 and 3 analyse theoretical limitations and formulate corrections or explain why they cannot be resolved, while Papers 4–6 address conceptual limitations.The conceptual contributions include joint evaluation of individual item fairness and effectiveness, a similarity-aware individual user fairness measure, and a unified individual- and group-fairness evaluation for users.

Evaluation Measures of Individual Item Fairness for Recommender Systems: A Critical Study

This study critically evaluates exposure-based, relevance-independent measures of individual item fairness in recommender systems. It identifies five theoretical limitations, proposes corrections or explains unresolvable issues, and derives guidance for appropriate measure use.

  • Corrections and guidance: The study proposes theoretical corrections where possible, explains why some limitations cannot be resolved, and provides guidance for selecting and correctly using individual item fairness measures.None of the examined measures is without limitations.
  • Theoretical limitations: The study identifies five theoretical limitations affecting exposure-based individual-item fairness measures, including non-realisability, quantity-insensitivity, and undefinedness, with impacts ranging from rare edge cases to practical scenarios.The limitations are independent of the recommender algorithm for top-k recommendation settings.
  • Theoretical limitations: Non-realisability prevents some measures from reaching their theoretical fairest or unfaires​t values, causing dataset-dependent score ranges and possible fairness underestimation.For k = 1, m = 2, n = 3, the lowest II-D and AI-D scores are 2/9 and 1/18; for k = m = W = 2 and n = 3, their minima are 0.02 and 0.005.
  • Theoretical limitations: Some measures have narrower achievable ranges because exposure weighting, item similarity, or exponential-like exposure prevents theoretical extremes, as illustrated by ↓Gini-w values from 0.0373 to 0.156.For k = n = 3 and m = 2, ↓Gini-w cannot reach 0 or 1 under non-uniform exposure.
  • Theoretical limitations: Quantity-insensitivity makes QF ignore repeated recommendations across users, so it can assign QFori = 0.6 to scenarios with different exposure distributions and miss popularity-bias unfairness.The limitation reflects the measure’s design choice, despite sensitivity to transfer being a basic inequality criterion.

2.9 Appendix

The appendix derives attainable bounds for several fairness measures and reports recommender performance, measure correlations, and controlled exposure-insertion experiments on Amazon-* and Book-x datasets.

  • Fairness bounds: The appendix derives minimum and maximum attainable values for Jain, QF, Ent, Gini, FSat, and VoCD from the most unfair and fair recommendation scenarios.For Jain, QF, Ent, Gini, and FSat, the corresponding extrema are explicitly assigned to the most unfair or fairest cases; weighted Gini is also treated separately.
  • Fairness bounds: For VoCD, the maximum with one similar-item pair occurs when the items are recommended 1 and m times, and adding pairs does not increase that maximum.These claims are established as Theorems 2.1 and 2.2.
  • Empirical evaluation: BPR generally performs best in relevance, except on Amazon-lb where NeuMF is best, while ItemKNN gives the best fairness scores.The original Ent scores are undefined because of zero-division errors, and II-D has constant scores.
  • Empirical evaluation: The appendix also reports Kendall’s Tau correlations, max/min fairness experiments, sliding-window experiments, and recommender hyperparameter search spaces and optima.The cited passages identify these analyses and their corresponding figures or tables without supplying their numerical values.
  • Empirical evaluation: Artificially inserting least-exposed and relevant items produces less stable score changes for m = {100, 500} than for m = 1000, but preserves the general trends.Inserting most-exposed copies or irrelevant items produces similar but opposite trends, with only k unique items ultimately receiving exposure in the former setup.

Can We Trust Recommender System Fairness Evaluation? The Role of Fairness and Relevance

The study finds that joint relevance-and-fairness measures for individual item fairness have serious limitations in alignment, interpretability, sensitivity, and practical usage. Across nine measures, none reliably captures both relevance and fairness in a balanced way.

  • Agreement: Correlations with single-aspect measures are inconsistent: IAA/HD/II-F do not align with fairness, IFD/MME/AI-F highly disagree with relevance, and IBO/IWO vary inconsistently.Within measure clusters, score ranges can still differ by up to ∆≈0.7, complicating interpretation.
  • Sensitivity: Joint measures are much less sensitive to rank changes than relevance-only and fairness-only measures, with single-aspect score changes up to two magnitudes greater.The study attributes this insensitivity to relevance effects being masked by fairness effects and vice versa.
  • Interpretability: Most joint measures consistently score almost perfect fairness even when single-aspect measures indicate highly irrelevant and unfair recommendations.Some measures have scores on the order of 10^-3 or less, with ranges such as (0, 0.0015) versus [0,1] for relevance, fairness, and IBO/IWO.
  • Practical implications: Practical use requires caution because IFD÷ ignores different cut-offs, MME and ↓IFD× are costly to compute, and some measures worsen when recommendations become more relevant and fair.The authors recommend improving joint measures so one score reflects relevance and fairness more accurately and balancedly.
  • Overall findings: Of 9 joint measures, 3 align with relevance-only measures, 4 align more with fairness-only measures, and the rest behave inconsistently.IBO/IWO behave inconsistently; IAA/HD/II-F align more with relevance, while IFD/MME/AI-F align more with fairness.

Relevance-aware Individual Item Fairness Measures in Recommender Systems: Limitations and Usage Guidelines … 4.2 Evaluation Measures of Fairness for Individual Items

This work examines relevance-aware individual item fairness measures, which jointly evaluate item exposure and relevance, and identifies limitations that hinder their interpretation and applicability. It amends the measures, empirically validates the corrections on real-world and synthetic data, and provides usage guidelines for practical scenarios.

  • Abstract: Relevance-aware measures evaluate individual item fairness by considering whether item exposure is comparable for items with similar relevance, unlike exposure-only measures that ignore relevance.The work focuses on individual fairness for items rather than group fairness, and addresses a class of measures that had received less analysis than exposure-only measures.
  • 4.1 Introduction: The study identifies five theoretical limitations of relevance-aware individual item fairness measures and analyses their prevalence across evaluation scenarios.It extends prior empirical work by systematically examining limitations that had not previously been analysed for these measures.
  • 4.1 Introduction: The amended measures resolve identified limitations where possible, while the study justifies why some limitations cannot be resolved.The corrected measures are evaluated on both real-world and synthetic data, where they overcome the limitations of the original formulations.
  • 4.1 Introduction: The paper recommends detailed measure-selection guidelines for practical evaluation scenarios, informed by the measures’ identified limitations and their corrected formulations.Table 4.3 summarises the theoretical limitations of relevance-aware individual item fairness measures, including two causes of non-realisability.
  • 4.2.1 Notation and Examination Functions: The framework represents item exposure through examination functions that model viewing probability as a function of rank position, using linear, DCG-, RBP-, or inverse-rank-based discounts.The notation covers users, items, full and top-k recommendation lists, relevance, rank positions, and repeated recommendation rounds.
  • 4.2.2 Exposure-based Fairness Measures: Exposure-based measures assess fairness from the aggregated item-exposure distribution, with greater uniformity generally corresponding to fairer recommendations.Examples include measures adapted from other domains, such as the Gini and Jain indices, alongside recommendation-specific measures such as QF.
  • 4.2.3 Relevance-aware Fairness Measures: The reviewed relevance-aware measures jointly combine exposure with relevance, and all except Hellinger Distance are defined for multiple recommendation rounds or stochastic rankings.The set comprises seven Joint measures published for recommender systems up to 1 July 2024, with score direction indicated separately for fairness and unfairness measures.

4.3 Measure Limitations

This section identifies five theoretical limitations affecting the fairness measures in §4.2.3, showing that some impair interpretation or applicability while others can make measures unusable. [177]

  • Overview: Five limitations affect the measures, with some restricting interpretation or usage and others rendering fairness evaluation impossible.The limitations are summarised in Tab. 4.3; their effects range from requiring cautious use to complete unusability.
  • Non-realisability: Non-realisability affects all Joint measures because their theoretical maximum or minimum fairness score may be unattainable across possible rankings [177].The attainable best and worst scores can depend on dataset distributions, recommendation cut-off k, relevance labels, and exposure weights; a new cause is also identified beyond the four previously reported causes [177].
  • Non-realisability: For k = m = 2 and n = 3, ↓HD reaches at most 0.707 rather than its theoretical maximum of 1, while ↓II-F ranges only from 0.007 to 0.88.These restricted ranges can make fair or unfair recommendations appear misinterpreted when scores are compared with the theoretical range.
  • Unobserved items: Treating unobserved items as irrelevant versus relevant changes IAA from 0.8 to 0.2 and ↓IFD÷(u) from 0.092 to 0.114, with one extra relevant item raising ↓IFD÷(u) to 0.136.Because relevance may be unknown when items were not shown or not rated/interacted with, sparse datasets can produce substantial fairness-score variation and possible fairness overestimation.
  • Non-localisation: Non-localisation makes IBO, IWO, II-F, AI-F, and HD depend on relevance information beyond the predicted ranking’s local top-k region.IBO/IWO and II-F/AI-F require relevance information for all user-item pairs, while HD requires obtaining the top-k ground-truth items by partially sorting all items.
  • Undefinedness: Undefinedness prevents scoring when division by zero occurs: IAA fails at k = 1, IBO/IWO when an item is irrelevant for every user, and IFD× when n < 2.IFD÷ was incorrectly reported as undefined for irrelevant items [161] [242], because the original measure computes fairness only over relevant items [205]; it is instead undefined when a user has no relevant item.

4.4 Resolving Limitations

This section resolves several fairness-measure limitations through modified formulas, normalization, and corrected exposure functions, while showing that other limitations remain unresolved because their bounds or required item-level information are unavailable.

  • Resolving limitations: The proposed modifications resolve top-k-insensitivity for IFD÷ and partially resolve Cause 2 non-realisability for IAA, IFD, and II-F through top-k indicators and per-user min-max normalization.These resolutions apply to single-round or deterministic multi-round settings with binary relevance, which are common recommender-system evaluation scenarios [228].
  • Resolving limitations: IFD÷-our zeros exposure-relevance scores beyond the cutoff, enabling different k values to quantify disparity between items inside and outside the top-k.Users with only one relevant item require exclusion or a fixed fairest score because the original pairwise-difference measure is always zero for them.
  • Resolving limitations: Undefinedness is resolved for IAA, IBO, and IWO, while zero-exposure is resolved for IAA by assigning nonzero exposure through rank k and zero exposure thereafter.The corrected examination function must also be used when computing the corrected IAA measure.
  • Unresolvable limitations: Non-realisability from Cause 1 remains unresolved because the fairest and unfairest recommendation lists are unknown, explicit bounds are unavailable, and dataset-scale optimization is impractical.If theoretical bounds were known, HD, MME, IBO/IWO, and AI-F could be normalized similarly to the resolved measures.
  • Unresolvable limitations: Non-localisation generally remains unresolved because the measures require item relevance information, although IFD÷, II-F, and AI-F can technically be modified to reduce this limitation.The proposed changes include aggregating exposure-relevance across all item pairs or replacing the target exposure with another target.

4.5 Experimental Setup

The experiments compare original Joint fairness measures with corrected versions across four real-world recommendation datasets, using standard preprocessing, recommender models, reranking, and effectiveness and fairness metrics. They also assess the robustness and limitations of the original measures.

  • The study compares original Joint measures with corrected versions while analysing how extensively the original measures exhibit their known limitations.The corrected measures address some limitations identified for the original measures.
  • Experiments use Lastfm, Amazon-lb, QK-video, and ML-10M, covering music, e-commerce, videos, and movies.QK-video uses only ‘sharing’ interactions; the other datasets follow the specified sources.
  • Datasets undergo 5-core filtering, duplicate removal, rating binarisation where applicable, and 6:2:2 global temporal or random train/validation/test splits.Users with fewer than five training interactions are removed after splitting.
  • Items are ranked with ItemKNN, BPR, MultiVAE, and NCL, then the top 25 items are reranked to improve fairness while retaining a cutoff of 10.The base recommenders are not fairness-optimised; BPR, MultiVAE, and NCL use 300-epoch training with early stopping, selecting the best validation NDCG@10 model.
  • Evaluation includes Joint and corrected relevance-aware fairness measures, recommendation effectiveness metrics, and exposure-only fairness measures.Effectiveness includes HR, MRR, P, R, MAP, and NDCG; exposure-only measures include Jain, QF, Entropy, and FSat, with specified parameter settings and edge-case handling.

4.6 Empirical Analysis

The empirical analysis shows that Joint fairness measures often disagree with single-aspect measures and can be difficult to interpret because of scale, undefinedness, and correction-related issues. Overall, no Joint measure reliably captures both recommendation effectiveness and exposure-based fairness.

  • Measure alignment: The best model selected by Eff measures generally differs from the fairest model, while Joint measures disagree more often about the best model.Eff measures consistently agree with one another across datasets, as do Fair measures, but these two groups select different models except for QF in Amazon-lb.
  • Correlation analysis: Joint-measure correlations follow three broad groups, but original and corrected measures can rank models differently despite generally strong agreement.IAA/HD/II-F correlate positively with Eff measures, IFD/MME/AI-F negatively, and IBO/IWO inconsistently; original-versus-corrected correlations range from τ ∈[0.57, 1].
  • Score properties: Original Joint scores suffer from near-zero ranges, scale mismatches, and undefined values, limiting model discrimination and score interpretation.Original ↓Joint scores can be ≤10−3, IBOori/IWOori produce ‘nan’ through division by zero, and lower-is-better measures can span radically different ranges despite comparable labels.
  • Score properties: Corrected IFD× and II-F scores are easier to distinguish across models because they normalize based on achievable max/min scores.Original and corrected versions can nevertheless produce very different values, such as ↓IAAori ≈0.010 versus ↓IAAour ≈0.9 and ↓II-Fori ≈0.000 versus ↓II-Four ≈0.9 on ML-10M.
  • Measure alignment: No Joint measure reliably accounts for both recommendation effectiveness and fairness quantified purely through exposure.IBO/IWO has inconsistent relationships with single-aspect and Joint measures; IAA/HD/II-F strongly disagrees with Fair measures, while IFD/MME/AI-F strongly disagrees with Eff measures.

4.7 Guidelines for Measure Usage

The guidelines select relevance-aware item fairness measures by considering alignment, computability, interpretability, expressiveness, stability, and efficiency. Overall, corrected IFD×-our is recommended as a strong choice, with IBOour/IWOour as alternatives and corrected measures preferred whenever possible.

  • Alignment: Prioritise non-Eff-aligned measures, avoid redundant measures within alignment groups, and compute one representative per group of highly similar measures.Eff-aligned measures are IAA/HD/II-F, while Fair-aligned measures are IFD/MME/AI-F; other measures have inconsistent alignment.
  • Computability: IFD×/MME are the most flexible choices because they are computable in all five conditions, whereas original IAA is restricted and should be replaced by corrected IAA.II-F/AI-F are computable in four of five cases, and IFD×/MME are especially useful when the total number of relevant items per user is unknown.
  • Interpretability: For binary relevance and single-round recommendations, prioritise IAAour, IFDour, or II-Four, and always use IAAour rather than faulty IAAori exposure quantification.The corrected IAA version avoids the interpretability issue caused by treating the item at position k as unexposed.
  • Expressiveness: Use HDori/IBO/IWO or corrected IAA/IFD/II-F measures for expressiveness, avoid original IFD×/MME/AI-F, and use IFD÷-ori only for full rankings.Most original measures produce scores close to the fairest value and respond weakly to exposure or relevance changes; IFD÷-ori is insensitive to cut-off k.
  • Stability and efficiency: Compute IFD× and MME for stability, monitor users with a single relevant item when using IFD÷, and prefer IBOour/IWOour as efficient alternatives without expressiveness issues.MME and IFD× are pairwise and computationally expensive, taking up to 30 minutes and approximately 12 minutes respectively on larger datasets, but can be parallelised.
  • Overall recommendations: IFD×-our is one of the better overall choices because it has desirable alignment, computability, interpretability, expressiveness, and stability, despite taking over 10 minutes on larger datasets.Its computation time can potentially be improved through parallelisation.

4.8 Related Work

Prior work surveys fairness in ranking and recommender systems across stakeholders and granularity, but generally lacks in-depth analysis of fairness measures. Related approaches distinguish group from individual fairness and jointly assess relevance with exposure or fairness.

  • Related surveys: Most fairness surveys provide broad overviews across stakeholders and evaluation granularities, but none offers an in-depth analysis of existing fairness measures [133] [173] [228].
  • Group and individual fairness: Group fairness seeks similar treatment for dominant and protected groups, whereas individual fairness seeks similar treatment for similar individuals; most existing work focuses on group fairness [63] [65].
  • Joint relevance and exposure evaluation: Alternative evaluations combine relevance and exposure through harmonic-mean ranking or compare model scores against the Pareto Frontier to represent relevance–fairness trade-offs [40] [181].

4.9 Conclusion

The work theoretically and empirically investigates relevance-aware individual item-fairness measures, identifies their limitations, and addresses them through redefinitions or explanations of unresolvable issues. It primarily studies deterministic single rankings with binary relevance, leaving broader ranking distributions and multistakeholder fairness for future research.

  • 4.9 Conclusion: The study investigates existing relevance-aware individual item-fairness measures that account for both item exposure and item relevance, examining their theoretical limitations under common and extreme evaluation settings.Where possible, the work redefines measures; otherwise, it explains why certain limitations cannot be resolved.
  • 4.9 Conclusion: The analysis focuses primarily on deterministic single rankings with binary relevance, while treating non-binary relevance and multi-round scenarios only conceptually.Future work should investigate fairness evaluation more deeply for distributions of rankings, including stochastic rankings and multi-round recommendations.
  • 4.9 Conclusion: Future research could examine the relationship between individual and group fairness measures and jointly evaluate user and item fairness in multistakeholder settings [229].

4.10 Appendix

The appendix derives fairest and unfairest rankings for IAA(u), IFD÷(u), IFD×(u), and II-F(u). Across these measures, fairness depends on placing relevant items at the top or bottom according to each measure’s exposure formulation.

  • Appendix: The appendix supplies formal pairwise theorems and derivations characterizing optimal or worst rankings for the four fairness measures.The derivations use item exposure at rank positions and, where relevant, the top-k cutoff and target exposure.
  • IAA(u): For ↓IAA(u), ranking items by non-increasing relevance produces the fairest ranking, while non-decreasing relevance produces the unfairest ranking.The pairwise results show that placing higher-relevance items closer to the top cannot worsen fairness, whereas placing lower-relevance items closer to the top cannot improve it.
  • IFD÷(u): For ↓IFD÷(u), ranking relevant items beyond the top-k is fairest, compared with ranking one or both relevant items within the top-k.The measure is lower when both relevant items are beyond the cutoff than when one or both receive top-k exposure.
  • IFD×(u): For ↓IFD×(u), placing a lower-relevance item closer to the top than a higher-relevance item produces a fairer or equal ranking.The proof considers both items beyond k, both within the top-k, and one on each side of the cutoff.
  • II-F(u): For ↓II-F(u), ranking higher-relevance items closer to the top produces a fairer or equal ranking than reversing their order.The result follows because higher relevance implies a higher target exposure, while exposure is non-increasing with rank position.

Joint Evaluation of Fairness and Relevance in Recommender Systems with Pareto Frontier

The section introduces DPFR, a Pareto-frontier-based method for jointly evaluating recommender-system relevance and fairness. Across datasets and measure pairs, DPFR identifies models differently from relevance, fairness, and most existing joint measures, producing a more balanced evaluation.

  • Motivation: Fairness and relevance often trade off, making separate evaluation ambiguous because different models may be best for each objective, while existing joint measures are difficult to interpret.Relevance and fairness measures may aggregate scores over different subjects, further complicating their combination.
  • DPFR approach: DPFR jointly evaluates relevance and fairness by measuring distance to a Pareto frontier of feasible recommendation scores, with a controllable fairness weight α.This makes DPFR modular, interpretable, and compatible with existing relevance and fairness measures.
  • DPFR approach: The Pareto frontier contains recommendations that are Pareto-optimal across fairness toward individual items and average user relevance, retaining the best fairness score for duplicate relevance values.The empirical frontier is a close approximation rather than a verifiable match to the theoretical frontier.
  • Measure suitability: DPFR is unsuitable for relevance measures based on a single relevant item, such as HR and MRR, because generating the frontier requires ranking items.The fairness measures QF and FSat can also behave inconsistently depending on dataset properties.
  • Empirical evaluation: Across all datasets and measure pairs, DPFR’s best model differs from the best relevance model, while existing Fair+Rel measures select a relevance- or fairness-best model 73.3% of the time.The best DPFR model is also different from the best fairness model half the time, and is less skewed toward either aspect.
  • Empirical evaluation: DPFR generally orders models differently from existing Fair+Rel measures, with weak or negative Kendall correlations for IAA and II-F on datasets including ML-10M and QK-video.AI-F correlates most strongly with DPFR, but its ranking remains non-equivalent on five of six datasets.

5.7 Appendix

The appendix documents dataset statistics, algorithm pseudocodes and complexity bounds for Oracle and Oracle2Fair, plus implementation assumptions, modifications, and edge cases. It also reports that replacing frequently recommended items does not significantly affect overall recommendation performance.

  • Oracle complexity: Oracle has time complexity O(R2m + Rm log m + Rmk log k + k2m2), with the k2m2 block typically dominated by O(k2m) when m >> H.The bound is dominated by the costs of the main iterative blocks.
  • Oracle2Fair complexity: Oracle2Fair has complexity O(R2m + Rm log m + Rmk log k + k2m2 + km2 log m + kmn) when k ≥ H, or O(R2m + Rm log m + Rmk log k + Hkm2 + km2 log m + Hmn) otherwise.Further assumptions about one or more variables are needed to simplify these bounds.
  • Edge cases: Oracle2Fair may fail to halt when users’ train/validation sets contain almost all dataset items and re-recommendation is prohibited, although such datasets are rare.Some item counts can remain above ⌈km/n⌉ in this edge case.
  • Modified GS: The modified GS treats all items as one similarity cluster and improves efficiency by caching replacement-loss information, sorting candidate pairs, and attempting only the first P pairs, where P is 25% of recommendation slots.Replacements are skipped when the item to be replaced no longer appears in a user’s recommendation list.
  • Replacement effects: Replacing frequently recommended items in Oracle2Fair and fair rerankers does not significantly affect overall recommendation performance.These replacements are used during Pareto Frontier generation and as part of the fair rerankers; the Pareto Frontier scores come from Oracle2Fair-generated lists rather than model-generated lists.

Measuring Individual User Fairness with User Similarity and Effectiveness Disparity

Existing individual-user fairness measures often fail to jointly capture user similarity and recommendation effectiveness. The proposed Pairwise User unFairness (PUF) measure addresses this by weighting pairwise relevance disparities by user similarity and is more responsive to effectiveness and similarity changes.

  • Sensitivity to effectiveness and similarity: PUF is more sensitive than existing measures to changes in recommendation effectiveness and user-similarity distributions, while most alternatives ignore effectiveness disparity or user similarity.Only SD consistently follows expected effectiveness-disparity trends, whereas PUF and UF respond to similarity-distribution changes; non-similarity-based measures do not.
  • Pairwise User unFairness (PUF): PUF measures individual user fairness through recommendation-relevance disparities between user pairs, weighted by their similarity, thereby accounting for both user similarity and utility.PUF ranges in [0, 1] and is modular, allowing any similarity and effectiveness measures that satisfy the range requirement.
  • Agreement among measures: Existing fairness measures frequently disagree with PUF on model rankings; only SD aligns substantially with PUF, with τ ≥0.62, while UF and PUF never agree on the fairest model.Other measures, including UF, show weak or negative correlations with PUF, with τ ∈ [−1, 0.24].
  • Sensitivity to user similarity: PUF distinguishes fairness levels across similarity distributions, whereas non-similarity-based measures remain unable to reflect these changes and can therefore misinterpret fairness.As similarity skewness increases, PUF and UF become fairer; similarity-independent SD remains constant, and PUF is more sensitive to negatively skewed distributions than UF.
  • Effectiveness-based fairness: Under MostFair assignments, PUF scores remain close to the fairest value 0, while UF stays around ∼0.8 and fails to distinguish MostFair from MostUnfair cases.PUF therefore captures maximal and minimal fairness more accurately, whereas UF overestimates effectiveness-based unfairness and is nearly constant across recommendation effectiveness.

Stairway to Fairness: Connecting Group and Individual Fairness

The study empirically connects group and individual fairness in recommender systems using comparable measures across multiple user-grouping schemes. It finds that group-level fairness can conceal substantial individual and within-group unfairness, motivating evaluation at all three levels.

  • Findings: No individual Fair measure consistently matches group Fair rankings, although CV always correlates strongly with an individual Fair measure.Group Fair measures show weak-to-strong agreement with Gini_ind/Atk_ind and weak-to-strong disagreement with SD_ind; CV has τ ∈[.71, .79].
  • Findings: Fairness generally worsens as more sensitive attributes define intersectional groups, while within-group fairness remains comparatively stable as the number of groups increases.Between-group fairness generally worsens with more groups, whereas within-group fairness is usually stable, except for SD on JobRec.
  • Findings: Within-group unfairness is almost as high as individual unfairness and always higher than between-group unfairness, showing why fairness evaluation should extend beyond between-group comparisons.Within-group unfairness can even exceed individual unfairness, such as for SD on JobRec.
  • Findings: Recommender systems that are fair for groups can remain very unfair for individual users, providing empirical evidence that these fairness concepts are disjoint.Individual Fair scores were always worse than group Fair scores across all measures and datasets, indicating that group scores can mask within-group disparities.
Loading 2604.25032v1…