Source-linked AI summary
A Long Way to the Top: Significance, Structure, and Stability of Internet Top Lists
Quirin Scheitle, Oliver Hohlfeld, Julien Gamba, Jonas Jelten, Torsten Zimmermann, Stephen D. Strowes, Narseo Vallina-Rodriguez
TL;DR
Researchers widely use Internet top lists, yet their construction, representativity, bias, stability, and overlap are poorly understood. The paper surveys their use, analyzes three popular lists and their ranking mechanisms, and reproduces measurements to assess list-dependent results. It finds substantial structural and temporal variation, including distorted comparisons with the general population and instability for some lists.
Problem
Researchers rely on Internet top lists as domain samples, but their creation, representativity, biases, stability, and overlap are poorly understood.
Method
The study surveys networking research, analyzes Alexa, Cisco Umbrella, and Majestic list structure and stability, examines ranking mechanisms, and evaluates measurement impacts against the general population.
Results
Top lists can significantly distort research results, with list choice and acquisition day affecting measurements; some lists also exhibit abrupt changes, weekly patterns, and substantial churn.
Takeaways & Limitations
Measurements based on top lists should be interpreted cautiously because list-specific sampling biases can limit generalisation to the Internet.
Takeaways & Limitations
The study excludes other top lists that are little used, inconsistently available, or fluctuate in size.
Abstract
from arXiv · showhide
A broad range of research areas including Internet measurement, privacy, and network security rely on lists of target domains to be analysed; researchers make use of target lists for reasons of necessity or efficiency. The popular Alexa list of one million domains is a widely used example. Despite their prevalence in research papers, the soundness of top lists has seldom been questioned by the community: little is known about the lists' creation, representativity, potential biases, stability, or overlap between lists. In this study we survey the extent, nature, and evolution of top lists used by research communities. We assess the structure and stability of these lists, and show that rank manipulation is possible for some lists. We also reproduce the results of several scientific studies to assess the impact of using a top list at all, which list specifically, and the date of list creation. We find that (i) top lists generally overestimate results compared to the general population by a significant margin, often even an order of magnitude, and (ii) some top lists have surprising change characteristics, causing high day-to-day fluctuation and leading to result instability. We conclude our paper with specific recommendations on the use of top lists, and how to interpret results based on top lists with caution.
1 INTRODUCTION
Internet top lists are widely used as supposedly representative domain samples, but their opaque construction leaves their bias, stability, and representativity poorly understood. This study examines their significance, structure, stability, ranking mechanisms, and impact on research results.
- Significance: 69 papers (10.0%) out of 687 surveyed networking publications used at least one Internet top list.The survey covered networking-related papers published in 2017.
- Structure: Top lists show distinctive structural properties, including invalid TLDs, less than 30% overlap, and disjoint domain classifications.These properties are investigated across popular lists.
- Stability: Daily churn reaches 50% for some top lists, demonstrating substantial instability.The study conducts longitudinal analyses of list stability.
- Ranking Mechanisms: Controlled experiments and Alexa-toolbar reverse engineering examine ranking mechanisms, including placing an unused test domain at rank 22k in Umbrella.The experiment demonstrates that ranking mechanisms can be studied and influenced under controlled conditions.
- Research Result Impact: Top lists significantly exaggerate measured characteristics relative to the general domain population, and results depend on the day of week when a list was obtained.The study evaluates this impact through experiments using top lists and all com/net/org domains.
2 DOMAIN TOP LISTS
The study focuses on Alexa, Cisco Umbrella, and Majestic, whose lists are updated daily but derive from different data sources and ranking methods. Several other lists are excluded because they are little used, inconsistently available, or variable in size.
- Alexa: Alexa ranks domains using web activity from its browser plugin, direct sources, and more than 25,000 browser extensions over three months.The user base is undisclosed, leaving possible geographic or age-related biases unresolved.
- Cisco Umbrella: Cisco Umbrella ranks FQDNs observed through Cisco OpenDNS, covering Internet services beyond websites.Its DNS-based collection differs fundamentally from measuring website visits or links.
- Majestic: Majestic ranks sites by the number of /24 IPv4 subnets linking to them using a web crawler.The list is included because its mechanism is orthogonal, openly licensed, and available over several years.
- Other Top Lists: Other top lists are not examined in detail because they are little used, inconsistently available, or fluctuate in size.Examples include Quantcast, Statvoo, Chrome UX Report, and SimilarWeb.
3 SIGNIFICANCE OF TOP LISTS
A survey of 687 networking papers shows substantial reliance on Internet top lists, especially Alexa, for measuring security, privacy, performance, and network characteristics. Many studies depend on list contents, while precise retrieval and measurement dates are rarely reported.
- Survey Scope and List Use: Internet measurement relied on top lists in 22.2% of surveyed papers, compared with 8.5% in security, 6.4% in systems, and 7.9% in web technology.The survey covered ten network-related venues in 2017.
- Top Lists Used: Alexa Global Top 1M was the most common choice, appearing in 29 studies, alongside numerous Alexa subsets such as Top 10k.One paper used Umbrella Top 100, while no surveyed paper used Majestic.
- Characterisation of Studies: Top-list studies covered security, privacy and censorship, performance, web infrastructure, DNS, IP, and TLS/HTTPS.Security was the largest purpose category, with 38 papers; web content was the most common layer, with 22 papers.
- Dependence on Lists: 45 studies had results that could depend on the selected list, while 17 used lists for verification and eight used them without necessary result dependence.These categories distinguish list-dependent outcomes from verification and list-independent uses.
- Reproducibility: Only 7 of 69 papers reported list retrieval dates, 9 reported measurement dates, and 2 reported both.Precise dates are described as important first steps for reproducibility, though their absence does not prove irreproducibility.
4 TOP LISTS DATASET
The dataset combines long-running daily snapshots of Alexa, Cisco Umbrella, and Majestic with a joint overlapping period for comparative analyses. Additional archived and community snapshots are used only where continuous daily data are available.
- Dataset Construction: Daily snapshots for the three lists were collected as far back as possible from the authors’ archives, community contributions, and the Internet Archive.Table 2 summarizes the resulting datasets and metrics.
- Dataset Construction: Alexa data include daily snapshots from January 2009–March 2012 and April 2013–April 2018.These periods are stored as datasets AL0912 and AL1318.
- Comparative Dataset: The JOINT dataset spans the overlapping period from June 2017 through April 2018 for comparative analyses across lists.Other individual snapshots were used only for periods with continuous daily data.
5 STRUCTURE OF TOP LISTS
The three top lists differ substantially in coverage, domain structure, and membership, reflecting distinct collection methods and producing limited overlap. Umbrella is especially DNS-oriented, while Alexa and Majestic provide more web-specific views.
- Coverage: ≈700 valid TLDs cover only about 50% of active TLDs in the JOINT period, so top-list measurements may miss up to 50% of TLDs.
- TLD validity: 2.3% of Umbrella’s Top 1M consists of 23k names under 1,347 invalid TLDs, compared with none in Alexa and only 35 names in Majestic.The authors associate this with invalid names queried by misconfigured hosts or outdated software.
- Domain depth: Umbrella contains 28% base domains and subdomains reaching level 33, whereas Alexa and Majestic contain almost exclusively base domains.Umbrella’s DNS-lookup origin allows deeply labeled names to enter the list.
- Aliases: Domain aliases account for about 5% of most lists but only 1.5% of Majestic, with the most common base name appearing about 200 times.
- Intersection: Only 99k domains are shared by all three Top 1M lists, while Alexa and Majestic overlap by an average of 29%.Pairwise Top 1k overlap is also low: 295 Alexa–Majestic, 56 Alexa–Umbrella, and 65 Umbrella–Majestic domains on average.
- Collection differences: Umbrella has more mobile-traffic and third-party tracking domains and the lowest overlap with the other lists, whereas Alexa and Majestic provide web-specific views.The comparison uses disjunct Top 1k domains and their presence in mobile, blacklist, and other-list datasets.
6 STABILITY OF TOP LISTS
Top-list stability varies sharply by provider, rank, and time: Majestic is relatively stable, while Umbrella and post-January-2018 Alexa show substantial churn and weekly variation. Lower-ranked domains fluctuate more, and rank order diverges over longer windows.
- Daily changes: 6k domains change daily in Majestic, 118k in Umbrella, and 483k in Alexa after its January 2018 change.Alexa changed from 21k daily changes before January 2018 to becoming the most unstable list afterward.
- Rank dependence: Instability increases at higher ranks for Alexa and Umbrella but not Majestic, showing that lower-ranked domains fluctuate more.
- Churn composition: 20% to 33% of daily changing domains are new, while 66% to 80% repeatedly leave and re-enter a list.
- Weekly patterns: Alexa and Umbrella exhibit weekly, non-monotonic membership changes, with domains leaving and rejoining at weekly intervals.The authors associate these patterns with different Internet usage on weekends, including leisure-oriented domains.
- Membership duration: About 90% of Alexa Top 1M domains remain for 50 or fewer days, while 40% of Majestic Top 1M domains remain for the full year.
- Weekly patterns: ≈35% of Alexa Top 1M domains and over 15% of Umbrella domains have completely non-overlapping weekday and weekend rank distributions.The effect is less pronounced in the Top 1k lists.
- Rank order: For day-to-day Top 1k comparisons, Kendall’s τ exceeds 0.95 for 99% of Majestic, 72% of Alexa, and 40% of Umbrella ranks.Against a fixed reference day, this very strong correlation falls below 5% for all lists.
7 UNDERSTANDING AND INFLUENCING TOP LISTS RANKING MECHANISMS
The paper examines how Alexa, Umbrella, and Majestic construct rankings and whether their mechanisms permit manipulation. Experiments show that Umbrella is driven more by distinct sources than query volume, while test-domain rankings disappear quickly after measurement stops.
- Alexa: Alexa ranks sites using visitor and page-view data gathered through browser extensions and other sources, but its data-collection mechanism is only partly documented.The authors reverse engineer the Alexa toolbar to investigate what data it gathers.
- Umbrella: Umbrella mainly reflects domains frequently resolved through OpenDNS, including automated DNS activity that may not represent human visits.The paper notes that scanning machines can appear in the ranking because their DNS lookups contribute to the list.
- Umbrella: 10k probes issuing 1 query per day achieved rank 38k, showing that probe count has a stronger influence than query volume per probe.The experiment varied 100, 1k, 5k, and 10k probes and frequencies of 1, 10, 50, and 100 queries per probe per day.
- Umbrella: Umbrella’s reliance on unique sources reduces susceptibility to individual heavy hitters, but test domains disappeared within 1–2 days after measurements stopped.The observed disappearance indicates that injected ranking effects were not persistent after traffic ceased.
- Majestic: Majestic ranks domains using referring /24 IPv4-subnets over 90 days, and the authors identify purchased backlink services as a way to influence rankings.Normalising by referring subnets limits the influence of single IP addresses, but referral services can still increase a domain’s popularity.
8 IMPACT ON RESEARCH RESULTS
The study evaluates how top-list choice and list timing affect Internet measurement results across DNS, hosting infrastructure, and protocol adoption. Compared with the general population, top lists generally produce significantly distorted and more extreme measurements, sometimes varying by weekday.
- Record Type Perspective: 11.5% of Umbrella and 2.7% of Majestic domains returned NXDOMAIN, versus 0.8% in the general com/net/org population.NXDOMAIN is used as a proxy for list-entry quality because allegedly popular domains are expected to exist.
- Overall impact: Top lists significantly distort measurements of Internet characteristics compared with the general population across nearly all evaluated metrics.Table 5 compares top lists with com/net/org domains and marks significant deviations from baseline values.
- Record Type Perspective: 11–13% of top-list domains supported IPv6, compared with 4% in the general population.IPv6 support was measured through AAAA records or CNAME chains of up to 10 names.
- Record Type Perspective: Top-list CAA adoption reached 1–2%, compared with 0.1% in the general population, while Top 1k lists reached up to 28%.The authors describe this as a distortion of the general-population value by two orders of magnitude.
- Hosting Infrastructure Perspective: All Top 1M lists had at least twice the general-population CDN prevalence, while all Top 1k lists had at least 20 times the prevalence.Weekend-versus-weekday effects on CDN ratios were minor, but the relative share of the top five CDNs generally exceeded 80% and differed by list and rank.
- TLS and Protocol Adoption: HTTP/2 adoption averaged 7.84% in the general population, versus up to 26.6% in Top 1M lists and around 35% or more in Top 1k lists.HTTP/2 adoption also differed by list and, for lists with weekday patterns, by weekday.
9 DISCUSSION
Top lists offer convenient, compact samples but can misrepresent the general Internet, fluctuate substantially, and produce results that depend on list choice and retrieval timing. The paper therefore recommends matching lists to study purposes, measuring repeatedly, documenting list details, and improving provider transparency and stability.
- Disadvantages: Top 1M lists introduce significant bias, while Top 1k lists can produce excessive magnitude bias, limiting generalisation to the broader Internet.Domains in top lists behave significantly differently from the general population.
- Disadvantages: Lists can change by up to 50% per day, and weekday-versus-weekend differences can make one-off measurements unstable.Repeated longitudinal measurements help reduce both temporal instability and day-of-week bias.
- Disadvantages: Different top lists can significantly alter measurements because their sampling biases differ, although those differences may help locate domains adopting specific technologies.The paper specifically notes effects on CDN or AS structure.
- Recommendations: Researchers should choose lists according to what their domains represent, record the exact list and dates, and share the list when possible.The recommendations distinguish, for example, Umbrella’s DNS-traffic perspective from web-specific lists such as Alexa and Majestic.
- Recommendations: Top-list providers should improve consistency and transparency, while offering both long-term and short-term versions to balance stability against recency.The proposed long-term version could use a 90-day sliding window, while the short-term version would use only the most recent data.
- Ethical considerations: The study follows stated harm-minimisation practices for active scans and used query volumes intended not to burden OpenDNS or RIPE Atlas.The authors describe blacklists, dedicated scanning infrastructure, and distributed probe scheduling as safeguards.
10 RELATED WORK
Prior work offers general measurement guidance and isolated warnings about popularity rankings, but this paper addresses the broader structure, stability, and research limitations of popular top lists. Existing studies also examine narrower issues such as subdomains, routing failures, and list manipulation.
- Sound Internet Measurements: General Internet-measurement guidelines do not specifically address the use of top lists.The paper positions top-list analysis as a distinct issue within sound measurement practice.
- Measuring Web Popularity: Earlier popularity research and SEO discussions note instrumentation bias and anecdotal list problems but lack systematic analyses.The cited warnings particularly concern Alexa ranks for low-traffic sites.
- Limitations of Using Top Lists in Research: Prior research mentions top-list limitations in specific studies, including variation from including www subdomains and causes such as routing failures.Other work focuses on challenges in web measurements rather than comprehensive list analysis.
- Limitations of Using Top Lists in Research: Earlier work also investigates manipulation of top lists, whereas this paper combines that issue with systematic analysis of list content and stability.The distinction is between focused prior studies and the paper’s broader scope.
11 CONCLUSION
The paper presents a comprehensive study of popular Internet top lists, covering their use, structure, stability, ranking mechanisms, and effects on measurements. It finds substantial distortion and temporal dependence, then derives recommendations for cautious use and interpretation.
- Conclusion: The study reports up to 50% daily churn for some lists and distinctive structural characteristics across lists.It combines structural analysis with stability measurements.
- Conclusion: Controlled experiments show that a test domain’s Umbrella rank could be manipulated, while reproduced measurements show distortion relative to the general population.The conclusion links ranking-mechanism analysis with research-result impact.
- Conclusion: Measurement results can depend on the day of week, so top-list-based findings require careful interpretation.The conclusion identifies temporal timing as a source of result variation.
- Conclusion: The paper closes with desirable properties for top lists and recommendations for their use in science, supported by shared code and data.An archival mirror is provided for long-term access.