Methodology
How top1m.org builds and maintains an independent website popularity ranking.
Overview
top1m.org continuously crawls its domain corpus and publishes a new ranking after the daily calculation completes. Completed daily editions are retained as immutable archives for historical comparison and downloads.
Coverage and recrawling
Crawlers work through the full registrable-domain corpus in repeated rounds. A successful crawl updates that domain's latest page and outbound-link evidence. Failures use increasing backoff so repeatedly unreachable domains do not consume the same capacity as healthy sites. Links can add newly discovered registrable domains to future crawl rounds; subdomains are recorded as hosts rather than separate registrable domains.
Domain ranking
A domain's score is the number of distinct source domains currently observed linking to it. Multiple links from one source, including links through different source hosts, count as one source-domain vote. This makes repeated links and single-owner link amplification less effective.
- Canonical evidence: one current source-domain to target-domain relationship.
- Equal source votes: a source contributes at most one vote to a target.
- Burst protection: suspicious coordinated discovery bursts can be held out of publication.
- Complete ranks: every registrable domain in the published corpus receives a rank; zero-vote ties use a stable ordering.
Rank is a relative link-popularity measure, not an estimate of visitors, traffic, revenue, or bandwidth.
Hosts, IPv4, and IPv6
Hosts receive a separate ranking based on distinct source domains linking to each observed hostname. IPv4 and IPv6 are separate products ranked from current hosting observations, including hosted-domain and hosted-host counts. These datasets are not interchangeable with the registrable-domain ranking.
Profiles and research signals
DNS, WHOIS, technology, contact, and registration-lifecycle observations enrich public profiles and research pages. They do not add votes to the core domain popularity score. Observations can be incomplete or stale until the relevant domain is crawled or enriched again.
Publication
Crawl data: updated continuously as rounds progress
Rankings: published after the daily calculation completes
Daily archives: Domains, Hosts, IPv4, and IPv6
Edition: dated YYYY.MM.DD
Limitations
- Coverage depends on what the crawler can reach and can under-represent some regions or networks.
- Sites that block automated access may retain older evidence or have less complete profiles.
- Parked domains, shared infrastructure, and CDN front domains can behave differently from visitor-based popularity measures.
- Technology, contact, DNS, and WHOIS values are observations, not guarantees of ownership or accuracy.
Data integrity
Published daily editions are immutable. Historical ranks are not rewritten when crawling or ranking code changes.