Skip to content

Crawl Budget: What It Is and How to Improve It for Large Sites

Last Updated: September 29, 2026

Crawl budget gets blamed for indexing problems it does not cause. It is not a number Google publishes, not a field in Search Console, and not a penalty you can recover from. Formally, it is the interaction between two limits: how many requests Googlebot is willing to make against your host in a given window (the crawl rate limit), and how much your site currently looks worth fetching (crawl demand).

That distinction matters because the two halves fail differently, and the fix for one does nothing for the other. A site that answers slowly fixes its problem with infrastructure. A site nobody wants to crawl fixes it with content and demand. Conflate them and you spend months tuning the wrong thing.

It matters for a second reason too: for most sites the honest answer is that crawl budget is not your problem. More on that below. If you already run a large or heavily parameterized site, our crawl budget waste finder is a fast way to see where requests are going.

The two limits behind one number

Google describes crawl budget as the interplay of crawl rate limit and crawl demand. Treat it as a multiplication rather than a balance:

  • Crawl rate limit: how many requests per hour Googlebot will send you without degrading real users. Server speed, uptime, error rates, and timeouts all move this. Consistently slow or erroring responses make Google back off, sometimes sharply, and sometimes for a while after you have recovered.
  • Crawl demand: how much Google wants your pages. Fresh, linked-to, well-performing URLs raise it. Duplicate URL spaces, soft 404s, and thin boilerplate lower it.

A high rate limit with low demand means Google politely declines to crawl most of your site. High demand against a crawling host means requests queue, time out, and push the rate limit down. Both produce the same symptom: important pages sit unrecrawled for weeks while Google fetches the same few thousand filter URLs over and over.

Who this affects, and who should stop reading

Plainly: if your site has 40 pages, crawl budget is irrelevant to you. Google will discover and recrawl 40 pages many times over on any normal day, and nothing you do to logs, robots.txt, or sitemaps will change your rankings. The same goes for a small local business site, a portfolio, or a blog with a few hundred posts. You have bigger levers available.

Crawl budget only becomes a constraint when the number of crawlable URLs is large relative to what Google wants, or when the site changes faster than Google can keep up. That means at least one of these is true:

  • Catalogs, marketplaces, and directories in the tens of thousands of listings or more.
  • Inventory or listings that turn over daily, where a slow recrawl costs real money.
  • Programmatic sites where templates multiply into parameter combinations: a few thousand pages times a few dozen filter and sort variants is already an unbounded crawl space.
  • High-volume publishing sections where recrawl speed decides how quickly a new page appears in results.

Below that threshold you are optimizing the wrong lever. Content quality, internal linking, and backlinks will outperform every crawl optimization available to you.

Where to see what Googlebot is actually doing

Server log files

Logs are the ground truth: every request, its user agent, response code, and timestamp. Extract the Googlebot entries, verify them by resolving the IP address (crawler IPs reverse-resolve to googlebot.com), and bucket by response code, by URL pattern, and by whether the request was served from cache. The pattern to look for is simple: if a large share of Googlebot hits go to URLs you would never put in a sitemap, the budget is being spent elsewhere.

A rough pass needs little more than filtering on the user agent and grouping by path prefix. If you want that structure pre-built, our log file analyzer pulls out the Googlebot share, the response-code distribution, and the URL patterns consuming the most hits.

Search Console crawl stats

Search Console, under Settings and then Crawl stats, gives Google's own view: total requests per day, response distribution, crawl purpose (discovery, refreshing, rendering), and the top crawled pages. It is aggregated and delayed, but it comes straight from Google and needs no server access. Read it alongside your Search Console workflow so crawl data sits next to your index coverage and performance numbers instead of in a tab nobody opens.

Use logs for diagnosis and crawl stats for trends. Logs tell you which URL class is eating requests; crawl stats tells you whether last month's fix moved anything.

What actually wastes crawl budget

  • Faceted navigation. Filter and sort combinations on e-commerce and directory sites generate millions of crawlable URLs with almost no unique content. This is the largest single source of waste on most big sites.
  • Infinite parameter URLs. Session IDs, tracking parameters, sort orders, and calendar pagination that never terminate. A parameter audit usually reveals dozens of variants per canonical page.
  • Soft 404s. Pages that return 200 with an empty state or a "no results" message. Google keeps fetching them as crawlable near-duplicates. They are also a content signal problem, so pair the fix with our thin content guide.
  • Redirect chains. Every hop is an extra request that returns nothing indexable. Collapse chains to a single hop with our redirect chain guide and sweep them in bulk with the bulk chain checker.
  • Orphan and discovery gaps. The opposite waste: pages with no internal links that Google only finds through the sitemap, if at all. Run the orphan page finder to catch URLs Google sees but your own site never points to.
  • Slow or erroring responses. 5xx spikes and multi-second server times lower the rate limit directly, which suppresses crawling of your good pages too.

How to prioritize the fixes

Order matters more than speed here. Work in this sequence:

  1. Stop the infinite spaces first. Anything that generates unbounded URLs, such as filters, sorts, or calendars, gets blocked in robots.txt or canonicalized to the clean variant. This changes the size of the problem instead of shaving a few percent off it. Validate every rule with the robots.txt checker before deploying, because a bad pattern blocks more than you intended.
  2. Collapse redirect chains and fix soft 404s. Each one wastes a request on every hit and both are mechanical to repair.
  3. Repair discovery. Add internal links to priority orphan pages and keep the sitemap to canonical, indexable URLs only, following the same rules the XML sitemap guide walks through.
  4. Improve speed and stability. Only after the URL space is bounded does tuning the rate limit pay off, because otherwise you are simply fetching waste faster.

The rule of thumb: fix anything that multiplies URL count before fixing anything that merely costs one request at a time.

How to verify the effect afterwards

Do not declare victory on a single good day. Compare like for like across at least a two- to four-week window:

  • In crawl stats, watch total requests per day and the share landing on parameter or filtered URLs.
  • In logs, compare the percentage of Googlebot hits hitting canonical indexable pages before and after the change.
  • In the Sitemaps report and index coverage, confirm priority URLs are discovered and crawled, not merely submitted.
  • Track time-to-index for a handful of newly published pages, because that is the metric users actually feel.

If nothing moves after a clean fix, the constraint was demand, not rate. Google does not consider those pages worth fetching, and the conversation shifts to content, internal links, and site architecture. That is usually the more productive conversation to have.

Related Articles

Continue learning with these practical SEO guides:

Related Tools

TheLinksMaster editorial team

Reviewed by TheLinksMaster

SEO Editorial Team

Our editorial team tests every tool and guide on this site against real websites. We update content when search algorithms change so you always get current, practical SEO advice.