Source-linked AI summary
Best-Worst Scaling More Reliable than Rating Scales: A Case Study on Sentiment Intensity Annotation
Svetlana Kiritchenko, Saif M. Mohammad
TL;DR
Rating scales are widely used but can be inconsistent, while BWS is claimed to provide efficient, high-quality annotation without systematic comparative evidence. This paper directly compares both methods under controlled annotation budgets and finds that BWS is significantly more reliable, particularly for complex phrases and smaller annotation totals.
Problem
The paper asks whether BWS really provides high-quality annotations with an annotation count comparable to rating scales, since this claim had not been systematically established.
Method
The authors directly compare rating-scale and BWS sentiment-intensity annotations while controlling the total number of annotations and evaluating reproducibility.
Results
ρ = 0.98 for BWS versus ρ = 0.95 for rating scales in average split-half reliability, a statistically significant difference (p < .001).
Takeaways & Limitations
BWS produces more reliable sentiment rankings than rating scales, especially for linguistically complex phrases and when about 5N or fewer annotations are available.
Takeaways & Limitations
Both methods are imperfect proxies for true word–sentiment intensities, which are not known directly.
Abstract
from arXiv · showhide
Rating scales are a widely used method for data annotation; however, they present several challenges, such as difficulty in maintaining inter- and intra-annotator consistency. Best-worst scaling (BWS) is an alternative method of annotation that is claimed to produce high-quality annotations while keeping the required number of annotations similar to that of rating scales. However, the veracity of this claim has never been systematically established. Here for the first time, we set up an experiment that directly compares the rating scale method with BWS. We show that with the same total number of annotations, BWS produces significantly more reliable results than the rating scale.
1 Introduction
Rating scales are widely used but can produce inconsistent and biased annotations. BWS offers a comparative alternative intended to retain annotation efficiency, yet this paper directly tests whether it is more reliable.
- Rating scales let annotators assign categorical or numerical values representing a measurable characteristic, with multiple responses usually averaged into an item score.
- Rating-scale annotation can suffer from disagreements across annotators, temporal inconsistency within annotators, scale-region bias, and fixed granularity.
- Paired comparisons avoid these rating-scale problems but require order N^2 annotations for N items.
- BWS asks annotators to identify the best and worst items in each n-tuple, enabling efficient inference of real-valued scores and rankings.For 4-tuples, two selections reveal five of six pairwise comparisons; typically 1.5N–2N tuples are annotated.
- The paper quantitatively compares RS and BWS on 3,207 English terms and reports that BWS produces significantly more reliable rankings, especially with about 5N or fewer annotations.
- Prior BWS annotation studies created datasets but did not systematically compare BWS with rating scales.
2 Complexities of Comparative Evaluation
Because neither method directly reveals true sentiment intensity, the study evaluates quality through reproducibility under equal annotation budgets. It also acknowledges that the methods impose different cognitive demands.
- Neither rating scales nor BWS perfectly captures native speakers’ true word–sentiment intensities, which are not directly known.
- The study uses reproducibility as a quality measure: similar repeated independent annotations should yield similar sentiment scores.
- The comparison controls total annotations because RS evaluates items individually while BWS evaluates groups of four items repeatedly.
- BWS and rating-scale questions involve different cognitive tasks, and quantifying their relative cognitive load is beyond the paper’s scope.
3 Annotating for Sentiment
The experiment collected sentiment-intensity annotations for 3,207 English terms using both rating scales and BWS. Rating-scale annotations showed notable temporal inconsistency, while BWS used repeated four-term comparisons to derive scores.
- 3,207 English terms were annotated for sentiment intensity by native English-speaking U.S. crowdsourcing workers using both methods.
- The term set contained 1,621 single words and 1,586 short phrases combining words with negators, modals, or degree adverbs.
- Rating-scale workers assigned each term a value from −4 to 4, with 0 marking no positive or negative sentiment.
- 20N total rating-scale annotations were collected, with quality checks and mean aggregation used to produce each term’s final score.
- 37% of repeated same-worker annotations differed, with an average difference of 1.27 points; inconsistency increased across longer intervals.
- BWS presented four terms at a time, collected 20N total annotations, and varied the number of tuples from 1N to 2N.
- BWS scores were computed as the percentage chosen most positive minus the percentage chosen most negative, ranging from −1 to 1.
4 How different are the results obtained by rating scale and BWS?
With equal annotation budgets, BWS and rating scale produce noticeably different sentiment scores and rankings, especially for linguistically complex phrases. The divergence is strongest for phrases containing negations and modal verbs.
- Differences between BWS and rating-scale outcomes are markedly larger with 3N or 5N annotations, but remain notable even with 20N.The comparison uses transformed scores, rank differences, Spearman correlation, and Pearson correlation.
- BWS and rating scale agree more on single terms than on phrases, with the lowest correlations for phrases containing negations and modal verbs.Positive phrases with negators show dramatically reduced correlation between the methods.
- Rating-scale annotations are more inconsistent for complex phrases, including positive phrases with negators, than for the full term set.The standard deviation is σ = 1.17 for these phrases versus σ = 0.81 for the full set.
5 Annotation Reliability
Split-half reliability is higher for BWS than for rating scales when methods are compared using the same total number of annotations. The advantage is largest with smaller annotation budgets and persists for linguistically complex phrases.
- BWS annotations are more reliable than rating-scale annotations at the same total annotation count, especially when 5N or fewer annotations are available.Reliability is measured using average split-half Spearman rank correlation over repeated random half-splits.
- BWS reliability is similar across 1N, 1.5N, and 2N sets of annotated 4-tuples when the total number of annotations is held constant.Reliability can therefore be improved by increasing either unique 4-tuples or independent annotations per tuple.
- BWS reaches ρ = 0.95 with 3N annotations per half-set, which is 30% of the amount needed for rating scales.
- BWS retains a reliability advantage for linguistically complex phrases, including positive phrases with negators and other phrase classes.The rating-scale reliability drop is especially severe for positive phrases with negators, whereas the BWS drop is much smaller.
6 Conclusions
The experiment finds that BWS produces significantly more reliable sentiment annotations than rating scales when total annotation counts are controlled. Its advantage is clearest with smaller budgets and linguistically complex phrases.
- BWS produced significantly more reliable results than rating scales when the total number of annotations was controlled.
- The reliability difference was more marked when an N-item set received about 5N or fewer total annotations.
- BWS was more reliable for linguistically complex items, including phrases with negations and modals.