Reference Library · 26 Documents · MIT + CC BY 4.0

Every signal a web crawler reads.

Living reference covering the HTML meta tags, HTTP status codes, and HTTP header categories that determine whether a search engine, AI engine, or social platform understands your site. Each document is a canonical reference — not a tutorial, not a marketing post.

Cite as: Anady, J. W. (2026). Crawler Signal Reference Library. Zenodo. https://doi.org/10.5281/zenodo.20405342
Companion to the open-source toolchain at /tools/ and the methodology paper at /research/.

High-demand crawler references

These reference pages are already receiving search impressions and should be treated as priority crawler-signal resources.

HTML meta tags 14 documents

HTML meta author tag HTML meta charset HTML meta color-scheme HTML meta content-language HTML meta copyright HTML meta generator HTML meta keywords HTML meta referrer HTML meta refresh HTML meta robots HTML meta theme-color HTML meta viewport HTML Open Graph (og:) tags HTML Twitter Cards (twitter:)

HTTP status codes 4 documents

HTTP 2xx success HTTP 3xx redirection HTTP 4xx client error HTTP 5xx server error

HTTP headers 8 categories

Caching headers Content headers CORS headers Performance headers Rate-control headers Request headers Security headers SEO headers (X-Robots-Tag, Vary)

Related work

The reference library is the empirical foundation for two MIT-licensed tools and a published methodology paper.

Author

Maintained by Joseph W. Anady, founder of ThatDeveloperGuy. Wikidata: Q139901957. ORCID: 0009-0008-8625-949X. Reference content is CC BY 4.0; companion software is MIT licensed.

Your preferences

A little more control.

Essential functions keep this site and its security checks working. Optional Google Analytics and Google Tag Manager load only after you allow measurement. They may use analytics cookies. You can decline and still use the entire site.

Motion preferences are stored only in your browser. Read the privacy notice.