Crawler Knowledge Atlas
The Crawler Knowledge Atlas is a public reference and a live research corpus. Thirty-six focused articles explain how discovery, rendering, directives, HTTP, structure, and measurement work in practice.
Every article is published as canonical semantic HTML and linked alternate Markdown and JSON representations. That makes format choice observable without sacrificing a useful human-readable site.
- Discovery and crawl frontiersHow crawlers learn that URLs exist, prioritize them, and avoid revisiting the same resource.
- Rendering and client executionWhat changes when a crawler reads server HTML, executes scripts, or inspects a rendered DOM.
- Crawler and indexing directivesThe standards and hints sites use to request crawling, indexing, following, and canonicalization behavior.
- HTTP behavior and efficiencyHow status codes, redirects, validators, caching, compression, and negotiation shape a crawl.
- Content structure for machinesMarkup patterns that preserve meaning across browsers, parsers, search crawlers, and AI retrieval systems.
- Quality, provenance, and observabilityHow to keep a crawl corpus current, attributable, non-duplicative, safe, and measurable.
Research contract
The atlas records anonymous request metadata, never raw IP addresses. A request proves that a URL was fetched; it does not prove indexing, ingestion, citation, training use, or a human visit.