DK logodavorkarafiloski
AboutCase StudiesSpeakingGlossary Work With Me
Glossary/Robots.txt
Core / Technical SEO

Robots.txt

FoundationsPractitionerSenior lens
Quick definition

A root-level file that tells crawlers which parts of a site they may or may not request.

01
Foundations
New to SEO? Start here.

Robots.txt is a plain text file at the root of your domain that gives crawlers instructions about which paths they may fetch. It is a crawl-control tool, not a security or privacy tool — anything you truly need private must be behind authentication, because the file itself is public and only well-behaved bots obey it.

02
Practitioner
Doing the work day to day.

Use Disallow to steer crawlers away from low-value areas — internal search results, cart and account pages, endless parameter combinations — so crawl demand flows to pages that matter. Keep it minimal and test every change in Search Console before shipping; a stray Disallow: / has taken entire sites out of Google. Reference your XML sitemap at the bottom of the file so crawlers can find it.

03
Senior lens
Strategy, trade-offs, judgement.

The subtle, career-defining detail: robots.txt controls crawling, not indexing. A URL you block can still appear in results as a bare link if other pages point to it, because Google indexed the reference without seeing the page. To actually remove something from the index you must allow the crawl and serve a noindex — blocking it in robots.txt guarantees Google never sees that noindex. Get this backwards and you cement unwanted pages in the index permanently.

DKDavor’s take

Robots.txt is the most dangerous file on your site relative to how simple it looks. One wrong line, shipped on a Friday, can deindex everything. Treat edits to it like a database migration.

Common mistakes
Using robots.txt to try to remove a page from Google — that requires noindex on a crawlable page instead.
Blocking CSS or JS that Google needs to render and understand the page.
Leaving a staging-site Disallow: / in place after launch, quietly blocking the whole production site.
In practice
Disallow: /cart/ keeps crawlers out of session URLs, but to remove an already-indexed thank-you page you must allow the crawl and return noindex.
Related terms
Crawl Budget
The number of URLs a search engine will crawl on your site within a given window.
Indexation
The process by which a search engine stores and makes a crawled page eligible to appear in results.
XML Sitemap
A machine-readable list of a site’s important URLs, submitted to help search engines discover and prioritise them.
Keep learning · Core / Technical SEO
Canonical TagCore Web VitalsStructured Data (Schema Markup)Internal Linking
← Previous
Core Web Vitals
Next →
XML Sitemap
← Back to the full glossary

Want this applied to your pipeline?

I turn concepts like these into quarterly roadmaps and measurable organic revenue for SaaS teams.

Work with me →
davorkarafiloski

Proven SEO systems for SaaS teams that refuse to fall behind in AI-era search.

© 2026 Davor Karafiloski. All rights reserved.
Skopje, North Macedonia · SEO Director at SmartClick