Source-linked AI summary
Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation
Victor Le Pochat, Tom Van Goethem, Samaneh Tajalizadehkhoob, Maciej Korczyński, Wouter Joosen
TL;DR
Researchers use popularity rankings to study representative web samples, but the paper finds that four major rankings have biases, instability, and manipulation vulnerabilities. It evaluates these properties, develops manipulation techniques, and introduces TRANCO, which combines rankings and filters domains for more stable and reproducible research. The study reports that Alexa can be manipulated with a single HTTP request, while TRANCO changes by at most 0.6% daily and is more resistant to manipulation.
Problem
Commercial popularity rankings have undisclosed data methods and potentially skewed userbases, limiting researchers’ ability to assess their validity and representativeness.
Method
The paper evaluates four rankings across similarity, stability, representativeness, responsiveness, and benignness, develops manipulation techniques, and constructs TRANCO by combining lists and filtering undesirable domains.
Results
A single HTTP request can place a domain in Alexa’s top million, while TRANCO changes by at most 0.6% per day and is more resilient against manipulation.
Takeaways & Limitations
TRANCO provides an online, archived ranking intended to support reliable, verifiable, and reproducible studies of popular domains.
Abstract
from arXiv · showhide
In order to evaluate the prevalence of security and privacy practices on a representative sample of the Web, researchers rely on website popularity rankings such as the Alexa list. While the validity and representativeness of these rankings are rarely questioned, our findings show the contrary: we show for four main rankings how their inherent properties (similarity, stability, representativeness, responsiveness and benignness) affect their composition and therefore potentially skew the conclusions made in studies. Moreover, we find that it is trivial for an adversary to manipulate the composition of these lists. We are the first to empirically validate that the ranks of domains in each of the lists are easily altered, in the case of Alexa through as little as a single HTTP request. This allows adversaries to manipulate rankings on a large scale and insert malicious domains into whitelists or bend the outcome of research studies to their will. To overcome the limitations of such rankings, we propose improvements to reduce the fluctuations in list composition and guarantee better defenses against manipulation. To allow the research community to work with reliable and reproducible rankings, we provide Tranco, an improved ranking that we offer through an online service available at https://tranco-list.eu.
I. INTRODUCTION
The paper finds that widely used popularity rankings have undisclosed biases, instability, poor representativeness, and manipulation weaknesses that can affect security research. It proposes TRANCO as a more stable, reproducible, and manipulation-resistant alternative.
- Motivation: 133 top-tier studies relied on commercial rankings whose undisclosed methods and limited, potentially skewed userbases undermine assessment of their validity.Researchers rarely reported retrieval dates, and historical list composition was often unavailable, limiting reproducibility.
- Findings: The four rankings disagree substantially, change over time, include non-representative or malicious sites, and can skew measurements of vulnerabilities or security practices.The paper identifies these problems across Alexa, Cisco Umbrella, Majestic, and Quantcast.
- TRANCO: TRANCO combines existing rankings, filters undesirable domains, changes by at most 0.6% daily, and requires at least four times the manipulation effort to reach the same rank.It is provided through an online service and archive to support reproducible research.
- Manipulation: Adversaries can manipulate every examined ranking at scale; for Alexa, a single HTTP request can enter the top million, and a rank as good as 28 798 was easily achieved.The paper reports that manipulation can affect thousands of domains and can exploit rankings used as whitelists or research inputs.
- Findings: Alexa’s list shifted to a one-day averaging period without announcement, while sites ranked below 100 000 can experience large rank changes from small traffic differences.The shorter averaging period contributes to instability, and Alexa itself warns that lower ranks are not statistically meaningful.
B. Cisco Umbrella
The paper compares popularity lists using their data sources, ranking metrics, and research-relevant properties. Cisco Umbrella uses DNS traffic and unique querying IPs, while the comparison evaluates similarity, stability, representativeness, responsiveness, and benignness.
- Cisco Umbrella: Cisco Umbrella publishes a daily list of one million entries ranked using aggregated traffic from domains and their subdomains.Any domain name may be included in the list.
- Cisco Umbrella: Umbrella derives rankings from DNS traffic to OpenDNS resolvers, using the number of unique IPs issuing queries for each domain.The provider claims more than 100 billion daily requests from 65 million users and applies sampling and normalization.
- Comparison criteria: The comparison evaluates similarity, stability, representativeness, responsiveness, and benignness as properties relevant to security research.These properties concern agreement, temporal rank changes, web-wide popularity, site availability, and absence of malicious domains.
- Evaluation: The study uses rankings downloaded between January 1 and November 30, 2018, and crawls sites from the four lists using a distributed ten-machine setup.The crawl was performed on May 11, 2018 at 13:00 UTC with headless Chromium.
A. Similarity
The four rankings share relatively few domains, and their similarity remains low even when higher ranks receive greater weight. This limited overlap means that choosing a different list can substantially change the domains included in a study.
- The four lists contain around 2.82 million sites but agree on only around 70 000 sites.
- Even when heavily weighting the top 100, RBO remains low at 24%–33% among Alexa, Majestic and Quantcast.Umbrella’s full-list RBO with the others is lower, between 4.5% and 15.5%, partly because it includes subdomains.
- Restricting Umbrella to pay-level domains raises its RBO with the other lists to around 30%.
- Quantcast’s removal of non-quantified sites caused RBO to fall below 5.5%, with no overlap among the top 10.
- The small overlaps indicate disagreement about which sites are most popular, so switching lists can change measured tracker prevalence.
- Majestic and Quantcast usually change by at most 1% per day, whereas Umbrella changes by an average of 10%.Alexa’s later shift to a one-day average made around half of its top million change daily.
E. Benignness
The rankings contain malicious or otherwise unsuitable domains despite their widespread use in security research and whitelists. Their composition and changing availability therefore create concrete risks for study validity and security tools.
- All four rankings contain domains flagged as potentially harmful, with Majestic reaching 0.22% of its list.In Alexa’s top 10 000, four sites were flagged for social engineering, while one Majestic site was flagged in its top 10 000.
- Whitelisting popular domains is particularly dangerous because security tools and Quad9 whitelist domains from Alexa or Majestic.This can expose users to ranked malicious domains while increasing the impression that those sites are safe to browse.
- The rankings’ validity and representativeness directly affect security-study results, and forged domains could let adversaries influence research findings.
- Popular-domain rankings are used to measure issue prevalence, evaluate domains, rank or bin sites, and whitelist them.
- Alexa is used by 133 recent security studies, mainly for measuring prevalence or evaluating popular domains.
- Most studies omit download timing, visit timing, and the reachable proportion of listed sites, hampering reproducibility amid daily list changes.
B. Influence on security studies
Popularity rankings are widely used to evaluate security and privacy practices, yet their limited and potentially biased data sources can distort studies and enable adversarial influence. The paper quantifies how ranking manipulation could hide fingerprinting providers and proposes more robust alternatives.
- Incentives: Security research frequently uses popularity rankings to represent real-world site usage and evaluate security practices.
- Case study: 7 032, 1 652, 74 and 24 manipulated domains could remove at least one fingerprinting provider from Alexa’s top 1M, 100K, 10K and 1K, respectively.
- Case study: 15 fingerprinting providers could disappear from Alexa’s top 1M with fewer than 100 000 manipulated domains.
- Feasibility: Manipulation can boost domains already present in a ranking, avoiding the need to insert new domains and reducing cost and detectability.
A. Alexa
The Alexa ranking collects traffic through browser-extension and analytics channels, but experiments show that forged traffic can obtain strong ranks with very few requests. Alexa’s processing and geographic coverage also introduce important qualifications.
- Extension: Alexa collects traffic through its browser extension and Certify analytics service, then ranks domains using proprietary unique-visitor and page-view measures.
- Extension: Alexa accepted page visits from generated profiles without imposing profile-creation limits, enabling experiments across 1 152 profile configurations.
- Caveat: Alexa appears to block reporting from EU and EEA countries, potentially biasing the ranking toward traffic from other regions.
- Extension: One request yielded a rank within Alexa’s top million, while 12 requests achieved rank 370 461.
- Extension: Alexa’s estimated request-to-rank model requires only 1 000 page views for rank 10 000, using an upper-bound calculation.
- Extension: Alexa accepted a nonexistent domain and did not check test domains during forged page-visit submission, further lowering manipulation costs.
2) Certify:
Alexa Certify derives rankings from a website’s tracking script, and forged requests produced sustained high ranks. However, the technique is slower and more expensive to scale than extension-based manipulation.
- Certify: Alexa Certify measures website traffic through an installed tracking script and requires a subscription starting at USD 19.99 per month per website.
- Certify: The experiment forged tracking requests that appeared to come from new users and diversified source IP addresses through Tor.
- Certify: For 48 days, the test domain reached Alexa’s top 100 000 and achieved rank 28 798.
- Certify: Alexa’s metrics reported 100.0% real traffic and no excluded traffic, suggesting the automated requests were not detected.
- Certify: Certified rankings require 21 days after subscription, creating a substantial delay before manipulation affects the ranking.
- Certify: Scaling Certify manipulation is quickly prohibitive because each inserted site requires a separate subscription, registration and real website.
- Certify: A domain received a Certified rank without simulated extension traffic, suggesting Alexa does not verify consistency between its two data sources.
B. Cisco Umbrella
Cisco Umbrella ranks domains by unique client IPs issuing DNS requests, making manipulation possible through large pools of addresses and crafted queries. Experiments achieved rank 200 000 with 1 000 unique IPs, while fake domains and subdomain aggregation further increase scalability.
- Umbrella ranks websites by the number of unique client IPs issuing DNS requests, requiring access to many addresses and requests to its resolvers.
- Cloud providers: 10 000 different IPs can cost less than USD 1 through repeatedly launching and stopping the cheapest AWS instances, but obtaining them takes time.Allocating Elastic IP addresses is faster but costs USD 0.10 per remap, making 10 000 IPs cost USD 1 000.
- Rank 200 000 was achieved with only 1 000 unique IP addresses, and all manipulation attempts succeeded.Umbrella appeared to count one day of DNS traffic for two days, reducing requests needed per day.
- Fake domains can be inserted because DNS names are freely chosen and Umbrella applies no filtering to invalid entries.
- 12 subdomains were ranked simultaneously with one request set, while aggregation across subdomains could also improve the pay-level domain’s rank.
- Umbrella manipulation can scale very broadly because fake domains, subdomain inclusion, and absent filtering or manipulation detection operate together.
1) Backlinks:
Majestic ranks domains using backlinks from distinct subnets, so manipulation requires creating or obtaining links that its crawler discovers. Paid backlinks successfully inserted a test domain, while reflected URLs offered a cost-free but effort-intensive alternative with limited achievable scale.
- Backlinks: USD 500 and two and a half months of backlink curation successfully inserted the test domain into Majestic’s ranking.
- Backlinks: Backlink manipulation trades cost against speed because more diverse or expensive links reduce the time needed to accumulate qualifying subnets.Majestic considers links for at least 120 days, affecting the cost of maintaining long-term manipulation.
- Reflected URLs: 1 041 pages reflected the test domain’s URL through GET parameters, and submitting those URLs to Majestic successfully ranked the domain.One backlink to a nonexistent domain was also crawled and counted as a referring subnet.
- Reflected URLs: Reflected-URL manipulation costs no money but requires substantial effort to find suitable pages and may not cover enough subnets for very high ranks.Deeper crawling could discover more pages, while permanent page changes or XSS would be more aggressive alternatives.
- Reflected URLs: A reflected URL continues counting indefinitely unless the site is reconfigured or taken offline, and one susceptible site can promote multiple attacker-chosen domains.
1) Quantified:
Quantcast derives ranking data mainly from tracking-script traffic and requires domain verification before processing it. Forged traffic was acknowledged, but the test domain was not yet ranked, likely because Quantcast updates slowly and the domain was young.
- Quantified: Quantcast traffic manipulation used requests from 479 US VPN servers, generating 400 users per day because its ranking counts US traffic.
- Quantified: The forged traffic was acknowledged and reported as 6,696 US users, but the test domain still did not appear in the ranking.The authors attribute this likely to the domain’s short age and Quantcast’s slow update frequency.
- Quantified: Above 5 000 visits, additional traffic can produce large rank improvements across blocks of estimated domains; the test domain’s theoretical rank was around 367 000.
- Quantified: Quantcast verifies that a tracking pixel is present before processing traffic, so manipulating multiple domains requires registering domains and setting up real websites.
- Quantified: Over 2 000 ranked domains reported 0 visits, suggesting that merely registering for tracking may sometimes suffice for inclusion.More than half of these domains were registered by DirectEmployers.
- Alternatives: Quantcast also states that it uses traffic data from ISPs and toolbar providers, but the authors could not determine which providers were involved.
VI. AN IMPROVED TOP WEBSITES RANKING
The paper proposes improving research rankings because existing generation methods can distort their properties and enable manipulation. TRANCO combines available ranking data, supports configurable lists, and preserves permanent records for reproducible studies.
- Existing ranking methods produce undesirable properties that can potentially sway security-study results and conclusions.
- Providers can improve resilience by detecting singular fraud and making large-scale manipulation prohibitively costly, though small-scale attacks may remain possible.
- Requiring accounts, filtering source IPs, and restricting traffic to known ISP ranges can make automated multi-address manipulation more difficult.
- Link-based rankings can be hardened through detecting reflected-URL attacks and using reputation signals such as page age or flow metrics.
- Availability checks requiring ranked domains to resolve and host real content can reduce unavailable and potentially fake entries.
- TRANCO combines existing ranking data to cancel respective deficiencies while allowing researchers to configure traffic sources and stability.
- Permanent records of lists, configurations, and construction methods make historical rankings easier to retrieve and studies easier to replicate.
1) Combination options and filters:
TRANCO combines existing rankings across providers and time, then applies configurable filters to improve similarity, stability, representativeness, responsiveness, and benignness. Its archived list records support reproducible studies.
- Combination options: TRANCO averages ranks across selected providers using Borda count or the Zipf-inspired Dowdall rule.The standard list uses the Dowdall rule.
- Combination options: Averaging rankings over the past 30 days reduces short-term fluctuations and filters can exclude briefly popular or manipulated domains.One-day lists remain available when short-term effects are the research target.
- Filters: TLD, pay-level-domain, subdomain, responsiveness, status-code, and content-length filters tailor the sampled domains to research needs.Additional options use Chrome User Experience Report membership and Google Safe Browsing to refine popularity and benignness.
- Evaluation: The evaluation combines lists from March 1, 2018 to November 14, 2018 and truncates them to one million domains.The period avoids distortions from Alexa’s and Quantcast’s method changes.
- Evaluation: RBO with Alexa and Majestic is 46.5–53.5% and 46.5–52%, while RBO with Quantcast and Umbrella is 31.5–40% and 33.5–40.5%.No provider has a disproportionately high influence on the combined list.
- Evaluation: Less than 0.6% of the all-provider combined list changes daily after 30-day averaging, compared with 1.8% for Alexa and 0.65% for Umbrella.The authors report that this supports longitudinal use because domain membership changes little.
- Reproducibility: Permanent links, citation templates, exact domain downloads, and configuration records make generated lists easier to retrieve and reference later.These features address the difficulty of recovering historical list composition.
4) Manipulation:
The paper finds that commercial popularity rankings can be manipulated at low cost and that their weaknesses affect both research and whitelisting. TRANCO raises the effort required for manipulation but remains dependent on the underlying lists.
- 4) Manipulation: TRANCO still inherits susceptibility from its four source lists, although its combinations and filters make successful insertion harder.Domains appearing on all source lists simultaneously receive favored treatment.
- 4) Manipulation: Combining lists and applying filters increases manipulation effort because an attacker must sustain changes longer and satisfy multiple list or domain-quality conditions.Removing unavailable domains can thwart some manipulated entries.
- 4) Manipulation: Because providers use separate traffic sources, manipulating one list is isolated; reaching a comparable combined rank requires manipulating all four lists.For the October 31, 2018 combined list, boosting one list to at least rank 11 091 was required to reach the top million.
- 4) Manipulation: All four existing rankings can be manipulated at scale, including Alexa through a single HTTP request that reached rank 28 798.The authors developed techniques that forge the data underlying domain rankings.
- 4) Manipulation: Alexa’s list changed by half every day, Umbrella contained only 49% HTTP-200 domains, and Majestic included more than 2 000 Google Safe Browsing-marked malicious domains.These findings illustrate instability, limited responsiveness, and benignness problems across the rankings.
- 4) Manipulation: Manipulation can insert malicious domains into whitelists and influence security research because rankings are widely trusted and rarely scrutinized.The paper reports that only two studies questioned their ranking methods.
- 4) Manipulation: The combined lists change by at most 0.6% per day, and manipulating one source list into its top 1 000 yields rank 100 000 in the combined list.The service provides reproducible access to these more stable and manipulation-resistant rankings.