Built-in SEO crawling

SEO crawler for technical audits and migration QA

Discover URLs through sitemaps, robots.txt and internal links. Inspect technical, content and structural issues at URL level, then compare crawl snapshots across releases or migrations.

7-day Pro trial. No credit card required.

URLMatcher SEO crawler dashboard
Crawl coverage

What the URLMatcher SEO crawler discovers and checks

Use the shared crawl dataset for a technical audit, release comparison or migration workflow without reducing every page to a few spreadsheet columns.

Discovery and scope

Control where the crawl starts and how URL inventory is discovered.

  • Start URL and page limit
  • XML sitemaps and sitemap indexes
  • Sitemap directives in robots.txt
  • Internal-link discovery
  • Crawl depth and subdomain controls
  • Query-parameter and trailing-slash variants

Technical and indexability

Keep the signals that determine whether a URL can be crawled and indexed.

  • HTTP status and final URL
  • Redirect chains, hops and loops
  • Canonical, noindex and nofollow
  • robots.txt directives and crawl access
  • Content type, HTML size and mixed content
  • Hreflang, structured data and sitemap coverage

Content and metadata

Inspect the page elements that can disappear or drift between releases.

  • Page title and title length
  • Meta description and description length
  • H1, H2, H3 and H4 headings
  • Word count and content fingerprint
  • Image count and missing alt text
  • Duplicate and near-duplicate signals

Links and site structure

Preserve the relationships that reveal crawl paths and internal authority.

  • Internal and external links
  • Inlinks, outlinks and source pages
  • Crawl depth and internal PageRank
  • Orphan, deep and poorly connected pages
  • Broken internal links and assets
  • Wrong host, protocol and domain variants
Automatic issue detection

Turn crawl data into issues your team can investigate

URLMatcher groups crawl findings by severity and keeps the affected URLs, context and underlying fields available for review.

Crawlability and status

Surface 4xx and 5xx responses, redirect chains and loops, blocked or noindex pages, mixed content and oversized HTML.

Metadata and content

Find missing, duplicate or unusually long titles and descriptions, heading gaps, thin pages and images without alt text.

Links and site structure

Identify broken links and assets, orphan or deep pages, weak inlink coverage, internal redirects and structural PageRank signals.

Canonical, hreflang and sitemaps

Flag canonical conflicts, hreflang problems, sitemap coverage gaps and unexpected protocol, host or domain variants.

crawl-input.csv
Common crawl-data gaps
Parameter variants were removed
Staging and production hosts were combined
No previous snapshot is available
Internal links were not included
The problem

A cleaned export can hide the URLs you needed to inspect.

Query parameters, trailing slashes and staging hosts can represent separate technical states.

Without the link graph, orphan pages and low-inlink sections are harder to diagnose.

Without snapshots, a release comparison becomes another manual spreadsheet exercise.

How it works

From domain to crawl data in 4 steps

01

Add the site

Enter a domain, reuse a project or start from a migration workflow.

02

Set the scope

Choose the page limit, crawl depth, subdomain behavior and robots.txt handling.

03

Run the crawl

Follow live progress while pages are fetched and parsed in batches.

04

Inspect the data

Open URLs, issues, links and crawl differences in the project.

Capabilities

Control the crawl and inspect the URL-level output

Keep scope, extracted page data, link relationships and crawl changes connected to migration and project workflows.

Crawl setup and live progress

Start a one-shot or project crawl and follow discovery and processing from the crawl-status view.

Sample data
URLMatcher crawl status using sample data

Scope controls

Set the start URL, page limit, depth, subdomain behavior and robots.txt handling before the crawl starts.

Crawl configurationIllustrative flow
Start URL
Page limit
Depth limit
Follow subdomainsOptional
Respect robots.txtOn

Sitemap and host discovery

Discover sitemap locations and keep production, staging and subdomain hosts visible as separate crawl decisions.

Discovery and host checksIllustrative flow
Sitemaps
Crawl queue
www.example.comIncluded
staging.example.comSeparate

URL-level extraction

Collect status, canonical, indexability, metadata, headings, body text, images, links and redirects.

Fields stored per URLIllustrative flow
Status
Canonical
Indexability
Title
Headings
Body
Links
Response

Internal link graph

Store incoming, outgoing and internal relationships for issue analysis, PageRank and link opportunities.

Internal link graphIllustrative flow
Home
Category
Guide
Target

Incoming and outgoing relationships remain available at URL level.

Snapshot comparison

Compare two crawls and inspect URL, metadata, status and structural changes between them.

Snapshot comparisonIllustrative flow
Previous crawl
Before release
Current crawl
After release
Status changedMetadata changedURL removed
Use cases

Where the crawler fits into the workflow

01

Migration crawl preparation

Crawl the old and new site directly before URL matching starts.

Domain, locale and staging checks keep the two URL sets usable for the matching workflow.

Old and new site - crawl both sources
Sitemap discovery - common locations checked
Parameter variants - kept visible
Matching input - reuse crawl results
02

Technical release checks

Compare crawl snapshots after a release, redesign or template change.

Status, canonical, metadata, headings and internal links stay available at URL level for comparison.

Snapshot comparison - before and after
Status changes - redirects and errors
Metadata changes - titles and headings
Structure changes - depth and inlinks
03

Ongoing project monitoring

Keep crawl data current inside the same project as issues and GSC analysis.

Scheduled crawls feed URL Explorer, issue reports, PageRank, duplicate detection and internal link analysis.

Scheduled recrawls - plan-based frequency
URL Explorer - raw page data
Project Issues - severity and explanation
Link graph - PageRank and opportunities
FAQ

SEO crawler questions

Start your next audit with a complete crawl.

Set the scope, crawl the site and inspect technical issues and URL-level data from the same project workspace.

7-day Pro trial. No credit card required.