Robots.txt & Crawl Control
Robots.txt controls which crawlers may access which parts of a site. Correct configuration preserves crawl budget for pages worth indexing and makes an explicit decision about AI crawlers: whether being cited in AI answers is worth more than restricting access to your content.
Prices shown are starting points. The final cost depends on the size and scope of your project, and we confirm it in writing before any work begins.
Robots.txt is small, powerful and easy to get badly wrong. A single misplaced disallow can remove an entire site from search, and because nothing errors, it can run for weeks before anyone notices traffic has gone.
The other half is AI crawlers. GPTBot, PerplexityBot, ClaudeBot and Google-Extended can each be allowed or blocked, and the choice has real consequences: block them and you are absent from AI answers; allow them and your content informs responses you do not control. For most service businesses trying to be found, being citable wins.
What's included
- Robots.txt audit against what is actually indexed
- Crawl budget analysis on large sites
- Disallow rules for filters, search and admin paths
- AI crawler policy set deliberately per bot
- Sitemap declaration
- Crawl-delay where a server needs protection
- Verification in Search Console's robots tester
Questions people actually ask
Does robots.txt stop a page being indexed?
No, it stops crawling, not indexing. A blocked page can still appear in results if other sites link to it, showing without a description. To keep a page out of the index, allow crawling and use a noindex meta tag. Blocking it prevents Google from ever seeing that tag.
Should I block AI crawlers?
It is a business decision, not a technical one. Blocking protects content from training use and removes you from AI answers entirely. For most businesses trying to be discovered, appearing in those answers is worth more. Publishers with licensable archives often decide differently.
What is the most common robots.txt mistake?
Shipping the staging file to production: a blanket Disallow: / that blocks the whole site. It is silent, it looks fine to visitors, and it is usually found only after traffic has collapsed.
Should I block AI crawlers like GPTBot?
It depends on what you sell. If customers find you by asking ChatGPT or Perplexity a question, blocking those crawlers removes you from the answer. If your content is the product and being reproduced costs you money, block them. We make the call per crawler and write down the reasoning.
What is crawl budget and does my site have a problem with it?
Crawl budget is how much of your site Google will fetch in a given period. Small sites rarely hit the limit. It becomes real on large stores with faceted filters, where thousands of near-identical filter URLs absorb the crawl and the pages that matter get visited less often.

