Czym jest Crawl?

How search engine bots discover pages by following links and fetching them for processing. A page must be crawled before it can rank.

Crawling is how search engines discover your content: automated bots (Googlebot and peers) follow links from page to page, fetch what they find, and hand it off for processing. It is step one of the search pipeline, ahead of indexing and ranking; a page that is never crawled cannot appear in results at all.

You steer crawling with a few mechanisms. An XML sitemap lists the URLs you want found; robots.txt tells bots which paths not to fetch; internal links carry crawlers through your site, and pages many clicks deep or linked from nowhere get crawled rarely or never. Two classic mistakes: leaving a site-wide crawl block in robots.txt after launch, and assuming blocking a crawl hides a page (a blocked URL can still be indexed from external links, just without its content). Very large sites also manage crawl budget, but under a few thousand pages it is rarely the constraint.

Crawlers announce themselves in the user agent and are filtered out by analytics tools, which is why your pageview counts exclude bot hits.

Powiązane pojęcia

Zobacz te metryki na własnej stronie

Analyse śledzi każdą metrykę z tego słownika bez cookies i banera zgody. Analityka, lejki i silnik SEO w jednej karcie. Za darmo przez 14 dni.

Rozpocznij bezpłatny okres próbny