Source-linked AI summary

PhiShark2026: A Multi-Layer Active-Web Raw-Evidence Dataset for Phishing Website Research

Furkan Çolhak, Ferhat Demirkıran, Hasan Dağ, Alexander Iliev

arXiv:2608.23199v1cs.CR

TL;DR

Phishing datasets often discard the raw, multi-layer evidence needed to study rapidly changing websites and revisit analytical assumptions. This paper constructs a hosting-aware active-web corpus that preserves artifacts and acquisition metadata across web, transport, domain, and infrastructure layers. The resulting 67,502-scan resource supports inspectable reanalysis while characterizing systematic phishing–benign differences across multiple evidence layers.

  • Problem

    Phishing datasets often reduce observations to fixed tables or narrow URL, HTML, and SSL features, limiting auditability and cross-layer reanalysis.

  • Method

    The study builds a scan-centered active-web corpus that preserves raw multi-layer artifacts and separates tenant evidence from provider-owned infrastructure signals.

  • Results

    67,502 scans are preserved across web, registration, TLS, security, and infrastructure layers, with characterization showing systematic differences between phishing and benign observations.

  • Takeaways & Limitations

    The corpus provides an inspectable foundation for alternative feature extraction, cross-layer analysis, and future reanalysis of phishing observations.

  • Takeaways & Limitations

    The phishing sources are operational feeds rather than a complete census, and access to the raw corpus is controlled for approved research users.

Abstract

from arXiv · show

Phishing websites are short-lived and rapidly changing, yet many phishing datasets reduce observations to URLs or precomputed features, constraining researchers to predefined representations and discarding the underlying evidence needed to derive alternative features, apply new extraction methods, examine cross-layer relationships, and reanalyze observations as phishing techniques evolve. This study addresses this limitation with a multi-layer active-web dataset comprising 67,502 scans, including 33,387 phishing observations from operational feeds and 34,115 screened benign reference observations. The corpus preserves raw evidence across HTML content and screenshots, URL and redirect behavior, HTTP and security headers, compliance files, TLS certificates, DNS and domain registration, open ports, geolocation and accessibility measurements, and network infrastructure, while explicitly recording unavailable evidence rather than treating it as negative observations. To avoid misleading infrastructure attribution on shared platforms, the study applies a hosting-aware evidence model that masks provider-owned infrastructure signals for free-hosted tenant pages while retaining meaningful page- and transport-level evidence. Characterization reveals systematic differences between phishing and benign websites across web-resource usage, domain maturity, mail and policy configuration, security headers, and infrastructure context. By preserving raw artifacts together with acquisition metadata and explicit evidence availability, the corpus provides an inspectable and reproducible foundation for future phishing measurement and dataset research.

I. INTRODUCTION

Phishing research is constrained by static, pre-aggregated datasets that inadequately represent dynamic campaigns and discard source evidence. PhiShark2026 addresses this gap with a multi-layer raw-evidence corpus spanning web, domain, certificate, and infrastructure observations.

  • Static datasets often inadequately represent dynamic phishing tactics, class imbalance, and structural variability.
  • Pre-aggregated tables restrict reproducibility and make derived fields difficult to revisit as phishing tactics evolve.
  • Multi-layer evidence links page-, domain-, and infrastructure-level properties without reducing analysis to a single field.
  • PhiShark2026 preserves DNS, WHOIS, IP, TLS, ports, HTTP headers, compliance files, screenshots, HTML, favicons, redirects, cookies, and page statistics.
  • The study contributes a 67,502-scan raw-evidence corpus, hosting-aware tenant/provider separation, cross-layer measurements, and robustness analyses.
  • Existing Phishing Datasets: Prior datasets progressed from URL-centric tables toward screenshots and richer HTML, but usually retained only partial surrounding evidence.

B. Measurement Gaps

The literature provides incomplete evidence for phishing measurement: many datasets are tabular, modality coverage is partial, infrastructure analysis is fragmented, and active-web realism is limited.

  • Many datasets remain pre-processed tables with limited auditability, coarse SSL proxies, and screenshots that are rarely paired with broader evidence.
  • URL, basic HTML, and coarse SSL signals describe only a narrow slice of rendered behavior, infrastructure, certificates, and domain context.
  • Infrastructure-aware studies often append DNS, WHOIS, TLS, IP, or hosting fields to conventional tables instead of releasing raw cross-layer artifacts.
  • Current datasets rarely model phishing as an active-web evidence ecology combining lexical, rendered, certificate, registration, and infrastructure layers.

C. Design Requirements

The paper requires an inspection-ready active-web benchmark that preserves raw artifacts across evidence layers and supports ecological interpretation rather than fixed feature tables. Its scan-centered design canonicalizes targets, dispatches specialized fetchers, and joins outputs for controlled reanalysis.

  • The design requirement is a well-documented, inspection-ready evidence ecosystem for studying layer availability, interactions, and robustness to shortcuts.
  • The corpus preserves URL, HTML, screenshots, favicons, DNS, WHOIS, IP WHOIS, certificates, redirects, headers, ports, and compliance evidence.
  • Raw evidence preservation, ecological interpretation, and free-host control are studied together rather than as isolated additions to earlier datasets.
  • Each scan records a target URL at a specific time together with successfully observed page, transport, registration, and infrastructure evidence.
  • Targets come from OpenPhish and the Malicious Links List, so the phishing sources are operational feeds rather than a complete census.
  • Acquisition Architecture: Canonicalized URL, host, and registrable-domain forms are routed to browser, network, DNS, TLS, WHOIS/RDAP, geolocation, compliance, and port workers, then joined by scan identifier.

C. Benign Reference Construction

Benign references were randomly sampled from OpenPageRank, with free-hosted candidates selected within provider strata to approximate phishing-provider composition. Candidates were screened by two services, while the resulting observations and evidence retained explicit missingness and hosting-aware scope.

  • Sampling: Benign candidates were randomly sampled from the OpenPageRank Top 10 Million Domains list.Independently registered candidates were sampled without restrictions; free-hosted candidates were sampled within provider strata.
  • Sampling: Free-hosted candidates were provider-stratified to reduce provider-composition imbalance between phishing and benign groups.This design preserved ordinary tenant pages on shared platforms while approximating the phishing provider distribution.
  • Screening: Each benign candidate was screened independently by Google Web Risk and the PhiShark URL-analysis platform before retention.Only candidates receiving no malicious or phishing indication from either service were retained.
  • Screening: The author-operated PhiShark screening service contributed no released scores or characterization features.Its outputs were used only for candidate screening.
  • Scope: Benign observations were targeted snapshots rather than a continuously sampled feed, so calendar time was not used for a two-class temporal comparison.This differs from the repeatedly collected phishing feeds.
  • Evidence scope: Missing evidence remained unavailable rather than being recoded as negative, while provider-owned infrastructure was masked for free-hosted tenant pages.Page- and transport-level observations remained meaningful across hosting strata.

F. Dataset Composition and Access

The controlled corpus contains 67,502 scans with linked metadata and per-scan archives, combining operational phishing feeds with benign references. Its characterization uses explicit layer-specific denominators and hosting-aware analytical boundaries.

  • Composition: 67,502 scans comprise 33,387 phishing and 34,115 benign observations.OpenPhish contributes 30,878 scans and the Malicious Links List contributes 2,509; 30,445 scans use free hosting and 37,057 use non-free hosting.
  • Release organization: Each retained record is linked by a unique scan_id to metadata and a corresponding <scan_id>.zip archive.The metadata records provenance, hosting stratum, identifiers, timestamps, status, and artifact-presence flags.
  • Release organization: Researchers filter the release CSV, open the matching archive, and join fetcher-specific observations through the enclosing scan identifier.Archives contain metadata.json, data/*.json observations, and captured artifacts.
  • Access: Raw corpus access is controlled to approved research users, while documentation and access instructions are publicly available.The documentation describes the metadata schema, archive layout, evidence modules, and access process.
  • Analytical scope: Analyses compare page-level observations within hosting strata and restrict provider-owned infrastructure analysis to independently registered non-free sites.Results use layer-specific available-case denominators, and missing fetches are not treated as observed negatives.

A. Web-Resource Ecology

Web-resource characterization finds thinner and lighter captured page ecosystems for non-free phishing than benign sites, while hosting and weighting choices materially affect technology comparisons. The reported thresholds are descriptive rather than deployment rules.

  • Page footprint and retrieval: 43 versus 30 and 37 versus 27 characters are the median phishing-versus-benign target URL lengths in free and non-free hosting.Non-free phishing pages also have median HTML payloads of 30.4 versus 78.8 KB, screenshots of 130.2 versus 270.4 KB, and extracted URLs of 18 versus 71.
  • Resource graph: 16 versus 43 HTTP transactions, 3 versus 11 scripts, 1 versus 3 cookies, 3 versus 5 resource domains, and 181.3 versus 691.8 KB are the median phishing-versus-benign non-free resource values.HTTP error transactions are the exception, with medians of one for phishing and zero for benign pages.
  • Interaction stack: 33.3% versus 12.9% of non-free phishing versus benign scans exhibit the descriptive minimal interaction stack.The complementary rich stack occurs in 4.3% and 31.2%, respectively; the thresholds are not optimized cutoffs or deployment rules.
  • Technology ecology: Median detected technologies are 1 and 3, with means of 2.03 and 3.04, for non-free phishing and benign scans respectively.Provider-standardized prevalence is benign-high for tag managers, analytics, JavaScript libraries, and content-management systems.
  • Technology ecology: Angular, Ruby on Rails, and React are phishing-high by 8.76, 8.66, and 7.23 percentage points at scan level but reverse after registration-domain deduplication.The result demonstrates why repeated-domain composition must be controlled before interpreting framework associations.

B. DNS and Registration

DNS, registration, TLS, security-header, port, and geolocation measurements show systematic class differences, but these signals describe observed infrastructure and maturity patterns rather than deterministic labels or actor identity.

  • DNS ecology: MX occurs in 29.7% of phishing and 86.7% of benign scans, while SPF occurs in 28.4% and 50.7%.The mean number of populated DNS record families is 3.82 versus 4.68, indicating a thinner mail and policy profile for phishing domains.
  • DNS ecology: 2,435 phishing and 849 benign DNS-covered records have no A answer, and CNAME-only or IPv6-only service does not explain these cases.Transient resolution state, incomplete resolver output, proxy behavior, and collection timing remain plausible explanations.
  • Registration maturity: 21.0% of phishing domains are 0–7 days old, whereas 91.5% of benign domains are older than five years.These are maturity patterns rather than deterministic labels because young benign domains and long-lived compromised domains remain possible.
  • Infrastructure: Median open-port counts are 4 for phishing and 3 for benign scans, while ports 80 and 443 dominate both groups.Port differences describe host-exposure context rather than a monotonic risk measure.
  • TLS: Let’s Encrypt accounts for 68.0% of covered non-free phishing scans and 38.3% of benign scans.At most 91-day validity windows cover 92.8% of non-free phishing certificates, while 38.2% of benign certificates exceed 180 days; valid TLS is not proof of legitimacy.
  • Security headers: Zero supported security-header families occur in 44.3% of non-free phishing scans and 22.9% of benign scans.Six or more occur in 8.1% and 28.5%, respectively, while free-hosting defaults can reverse individual relationships.
  • Policy and compliance: CSP presence does not establish a restrictive policy, and robots.txt results cannot exclude catch-all routes or soft-404 responses.Among CSP-covered non-free scans, unsafe-inline appears in 77.2% of phishing and 62.2% of benign policies.
  • Geolocation: Observed server country and geolocation describe the hosting environment rather than actor origin.Full six-region accessibility is common in both groups, while partial accessibility is more frequent among phishing scans.

D. Cross-Layer Artifact Reuse

The analysis uses exact artifact hashes and cross-domain recurrence to identify reuse candidates, while showing that multi-layer agreement reduces benign collisions more effectively than single artifacts.

  • Higher-confidence reuse candidates require the same combined signature to recur under at least two registration domains.The analysis reports these intersections as candidates rather than confirmed campaigns.
  • 89.7% of phishing favicons and 50.6% of phishing screenshots belong to repeated exact-hash clusters, versus 20.6% and 8.9% for benign artifacts.Exact compressed-HTML reuse is much lower: 3.9% for phishing and less than 0.1% for benign.
  • Security-header completeness is measured across twelve header families observed by the raw web fetcher.The supplied figure caption identifies the comparison dimensions but does not state an outcome.
  • Screenshot-plus-favicon signatures cover 32.67% of eligible non-free phishing scans and 2.07% of benign scans.Screenshot-plus-hosting-IP covers 13.52% and 0.93%, while favicon-plus-TLS-serial covers 9.49% and 0.57%.
  • The strict screenshot-plus-favicon-plus-HTML signature identifies 46 phishing scans and no benign scans.Multi-layer intersections provide stronger candidates for manual validation than any single hash, but generic artifacts and shared infrastructure can still collide.

V. ROBUSTNESS ANALYSIS

Robustness analyses test whether observed phishing–benign differences persist after accounting for provider mix, repeated domains, page complexity, and threshold choices.

  • Provider standardization: 33.37% of phishing and 11.41% of benign scans meet the minimal interaction-stack definition after provider standardization.The Mantel-Haenszel odds ratio is 3.80, and 14 of 15 eligible providers retain the phishing-high direction.
  • Provider standardization: Provider-stratified results preserve the benign-high direction for HTTP transactions, script transactions, and transferred-resource volume, with HTTP error transactions as the principal exception.The thinner resource ecology is therefore not explained only by different provider mixtures.
  • Registration-domain weighting: All 22 core observations retain their direction under registration-domain equal weighting, although several magnitudes change.This weighting prevents repeatedly scanned domains from dominating the comparison.
  • Registration-domain weighting: Angular, Ruby on Rails, and React reverse from phishing-high at scan level to the opposite direction after registration-domain equal weighting.These scan-level associations reflect repeated-domain composition rather than intrinsic framework risk.
  • Complexity control: Under complexity control, domain age, MX presence, and transferred-resource volume remain benign-high, while target URL length and HTML size are phishing-high.Several pooled differences attenuate toward chance, separating persistent properties from signals tracking overall page richness.

D. Threshold Sensitivity

The threshold analysis finds a persistent phishing-high interaction-stack gap across tested definitions, while the dataset’s interpretation remains bounded by hosting, temporal, availability, and attribution constraints.

  • Threshold sensitivity: Every tested interaction-stack definition retains a phishing-high prevalence gap ranging from +12.81 to +33.37 percentage points across 81 threshold combinations.The result is not an artifact of the single descriptive threshold used earlier.
  • Scope boundaries: Temporal robustness is omitted because phishing was collected over approximately 114 days while benign observations came from short source-specific bursts.A two-class temporal split would confound calendar time with source and acquisition design.
  • Hosting-aware evidence: The 67,502-scan dataset supports page-level and transport comparisons across free-hosting and non-free-hosting contexts, while masking provider-owned infrastructure for free-hosted tenants.This defines the analytical boundary for infrastructure interpretation.
  • Cross-layer characterization: Non-free phishing pages are consistently thinner across HTTP transactions, scripts, cookies, contacted domains, IPs, ASNs, and transferred bytes.Provider adjustment, domain weighting, threshold sensitivity, and complexity control distinguish stable differences from composition-sensitive ones.
  • Attribution limits: Infrastructure measurements describe observed hosting context rather than identifying an attacker, operator, or country of origin.This caution also applies to exact artifact reuse, whose intersections remain candidates until generic and shared-provider artifacts are excluded.
  • Evidence availability: Explicit availability denominators prevent missing fetches from being treated as observed negatives.Uneven WHOIS/RDAP, geolocation, and port availability requires available-case denominators and does not justify assuming random missingness.

VII. CONCLUSION

The paper releases a hosting-aware active-web corpus that preserves raw evidence across layers and supports cross-layer, robustness, and reuse analyses. Its conclusions remain bounded by operational coverage, web volatility, incomplete evidence, cadence differences, and dual-use risks.

  • Corpus contribution: 67,502 active-web scans include 33,387 phishing and 34,115 benign observations organized into free-hosting and non-free-hosting strata.The strata prevent shared provider infrastructure from being attributed to tenant pages.
  • Main findings: The dataset’s value lies in relationships among layers, with thinner web-resource ecology, younger registration profiles, and less complete mail, policy, and security-header posture for non-free phishing pages.Provider adjustment and registration-domain weighting preserve these directions, while complexity control identifies which persist among equally sparse pages.
  • Artifact reuse: Cross-domain screenshot-plus-favicon and screenshot-plus-hosting-IP intersections reduce benign collisions and provide stronger candidates for manual campaign validation.Generic error pages and blank renders show why single hashes are not campaign labels.
  • Limitations and future work: The corpus remains bounded by operational-feed coverage, active-web volatility, incomplete layers, and incompatible phishing–benign acquisition cadence.Future releases can extend it with longitudinal captures, validated public-file content, and domain- and artifact-cluster annotations.
  • Ethics and access: Raw URLs, HTML, screenshots, and related artifacts improve auditability but create dual-use and residual risks that motivate controlled research access.Sensitive content may include victim-specific parameters, embedded forms, attacker-controlled scripts, and operational kit structure.
Loading 2608.23199v1…