Blog · 5 min read

What Is Crawl Budget and Why It Matters

Crawl budget explained for site owners. What it is, when it actually limits your site, and the fixes that stop search engines wasting crawls.

Part of the The Complete Technical SEO Guide guide.

By Umit Caybas, Founder and SEO Consultant. Updated .

Crawl budget is the number of URLs a search engine bot will fetch from your site within a given period. It matters for large or frequently changing sites because wasted crawls on duplicate, thin, or blocked pages delay the discovery and indexing of the pages that actually earn traffic.

Key takeaways

  • 01Crawl budget only constrains large or slow sites with many URLs.
  • 02Duplicate and thin URLs waste crawls that should reach revenue pages.
  • 03Server speed and clean internal linking are the biggest levers.

What crawl budget is

Search engines do not have unlimited time for any single site. Crawl budget is the practical ceiling on how many URLs a bot will request from your site over a period, set by how fast your server responds and how much the search engine wants your content. Google describes it as the combination of two limits. Crawl capacity is the rate the crawler believes your server can sustain without degrading the experience for real visitors; if response times climb or errors appear, the crawler backs off. Crawl demand is how much the engine wants to fetch, driven by how popular your URLs are, how often they change, and how much of the site it has already seen. The glossary entry for crawl budget has the formal definition; this article is about what it means in practice.

The important consequence is that the budget is spent on whatever URLs the crawler finds, not on the URLs you care about. If your site exposes a million addresses and only four thousand of them are real pages, the crawler will spend most of its visits on the other 996,000, and the real pages will be fetched rarely and re-checked slowly.

When it becomes a problem

Small brochure sites rarely notice a limit. A site with two hundred pages and a responsive server will be crawled in full every few days, and nothing about crawl budget will ever show up in its traffic. The problem arrives with scale or with structure. Large ecommerce catalogues, publishers with deep archives, and sites with faceted navigation can generate millions of URL variants, and bots spend their time on those instead of on new or updated pages.

The symptoms are recognisable. New products or articles take weeks to appear in results. Search Console's page indexing report shows a large "discovered, currently not indexed" bucket, which means the engine knows the URLs exist but has not got around to fetching them. Updated pages keep showing their old titles in results long after the change. And in the logs, the crawler's visits cluster on URLs nobody would ever search for.

Common sources of waste

  • Faceted navigation and sort parameters, which multiply every category page by every combination of filter
  • Session identifiers and tracking parameters in URLs, which make one page look like thousands
  • Redirect chains and soft 404s, where the crawler fetches a page only to be sent somewhere else or shown an empty result with a 200 status
  • Infinite calendar or pagination loops, where "next month" or "next page" never ends
  • Internal search result pages that are linked from the site and generate a new URL for every query
  • Staging or duplicate hostnames that are reachable and not canonicalised

Each of these is a template problem rather than a page problem, which is the good news: one fix repairs every affected URL at once.

How to measure it

Start with server logs to see what bots actually fetch; a crawl tool can only tell you what could be fetched. Four weeks of logs, filtered to verified search engine user agents, will show the share of requests reaching each template, the share wasted on redirects and errors, and how long since each important page was last visited. Search Console's crawl stats report gives a summarised view for Google alone, including response time trends and the split between HTML, images and other resources.

Both platforms in our Semrush and Ahrefs comparison can crawl the site to confirm what the logs show, and to produce the list of URL variants that need consolidating. Then put a value on the problem: the SEO ROI calculator turns the traffic a fix would free up into leads and revenue, which is the number that gets engineering time scheduled.

How to fix it

The fixes, in the order they usually pay off:

  1. Consolidate URL variants. Give every filtered, sorted and parameterised variant a canonical tag pointing at the clean URL, and promote only the handful of variants with real search demand to indexable landing pages of their own.
  2. Stop generating the waste. Change filter links so the crawlable version of each page carries no parameters, block session and sort parameters in robots.txt, and remove internal links to search result pages and infinite archives.
  3. Collapse redirect chains. Every internal link should point at the final URL. A chain of three redirects costs three fetches and passes less signal each hop.
  4. Speed up the server. Crawl capacity rises when responses get faster and errors disappear. Caching, a CDN and fixing slow database queries all raise the ceiling.
  5. Point internal links at what matters. Link the pages that earn revenue from the home page and from category pages, regenerate the XML sitemap to include only canonical, in-stock, indexable URLs, and let the crawler find its way to the right places.

The technical SEO guide covers each of these steps in context, including what the logs look like before and after a fix.

What to expect afterwards

Crawler behaviour changes first, usually within days: the share of visits reaching real pages rises and the wasted share falls. Index coverage moves next, over two to six weeks, as pages that were "discovered, currently not indexed" are finally fetched and stored. Rankings and traffic follow. None of this requires new content. The same pages simply become reachable, which is why crawl budget work is often the highest-return technical fix on a large site.

Frequently asked questions

Does crawl budget matter for a small site?

Most small sites never hit a crawl limit. Crawl budget becomes a real constraint once a site has tens of thousands of URLs, heavy faceted navigation, or slow server responses that throttle how quickly bots can fetch pages.

How can I see how my crawl budget is being spent?

Server log files are the most reliable source. They show every request a bot made, which URLs it fetched, and how the server responded, so you can spot wasted crawls on parameters, redirects and error pages that never should have been requested.

Will blocking pages in robots.txt free up crawl budget?

It can, but blocked pages can still be indexed from links alone and cannot pass signals. Fixing the source of the wasted URLs, such as faceted navigation or session parameters, is usually a better long-term answer than blocking them.

  • Glossary

    Crawl Budget

    Definition of crawl budget, the number of URLs a search engine will fetch from a site in a period, and why wasting it on duplicates costs indexation.

  • Comparisons

    Semrush vs Ahrefs for Australian SMBs

    Semrush and Ahrefs compared on keyword data, site auditing, backlinks, rank tracking, local features and price for Australian small and mid-sized businesses.

  • Tools

    SEO ROI Calculator

    A free calculator that turns your organic traffic, expected uplift, conversion rate and lead value into extra leads, revenue and return on SEO spend.

  • FAQs

    Technical SEO Frequently Asked Questions

    Straight answers to the technical SEO questions Australian businesses ask most, from crawl budget and canonicals to structured data and speed.

Next step

Book the Technical SEO.

Technical SEO fixes for Australian websites, from $1,800 + GST. Crawl, indexation, rendering and structured data problems found and fixed in priority order.

Book the Technical SEO