Technical SEO

Technical SEO Audit: A Step-by-Step Guide for 2026

A technical SEO audit checks whether search engines can crawl, render and index your site. Work through crawl, indexing, robots.txt, sitemaps, canonicals, redirects, rendering and speed, then fix by impact.

By SEORecheck Research Team, SEO auditorsPublished 10 min read

A technical SEO audit is a structured check of whether search engines can discover, crawl, render and index the pages you want to rank — and whether anything in your site's infrastructure holds them back. It covers crawling, indexing, robots.txt, sitemaps, canonicals, redirects, status codes, JavaScript rendering, mobile, Core Web Vitals, structured data and server logs.

Content and links only pay off once this foundation works. A brilliant article that returns noindex, sits behind a redirect chain or renders blank without JavaScript simply does not compete.

If you are new to audits in general, start with what an SEO audit is; for a broader tick-list that also covers content and links, see the SEO audit checklist.

What does a technical SEO audit include?

A complete audit answers five questions, in this order:

  1. Can search engines find the pages? (internal links, sitemaps, robots.txt)
  2. Can they crawl them? (status codes, redirects, server errors, crawl traps)
  3. Can they render them? (JavaScript, blocked resources)
  4. Will they index them? (noindex, canonicals, duplicates, quality)
  5. Do the indexed pages deliver a good experience? (mobile, Core Web Vitals, HTTPS, structured data)

Each layer depends on the one before it. There is no point tuning LCP on a template Google never indexes.

Area What to check
Crawl Status codes, depth, orphan pages, crawl traps
Indexing Page indexing report, excluded URLs
robots.txt Blocked sections and resources
Sitemaps Only indexable, canonical 200 URLs
Canonicals Self-referencing, consistent signals
Redirects Chains, loops, 302 vs 301
Rendering Content and links present in rendered HTML
Mobile Parity with desktop, usable layout
Core Web Vitals LCP, INP, CLS field data
Structured data Valid, matches visible content
Logs What Googlebot actually requests

Step 1: Crawl the site like a search engine

Start with a full crawl using a desktop or cloud crawler such as Screaming Frog SEO Spider, Sitebulb or a similar tool. Set the user agent to Googlebot Smartphone, respect robots.txt on the first pass, and connect Search Console and analytics if the tool allows it.

From the crawl, pull out:

  • Status codes: every 4xx and 5xx URL that is linked internally.
  • Click depth: important pages more than three or four clicks from the homepage.
  • Orphan pages: URLs in your sitemap or analytics that no internal link points to.
  • Crawl traps: faceted filters, calendars or session parameters that generate endless URLs like /shoes?color=red&size=42&sort=price&page=37.

Compare the number of crawlable URLs with the number of pages you actually want indexed. A 400-page business site that exposes 12,000 crawlable URLs has a parameter or pagination problem.

Step 2: Check indexing in Google Search Console

The Page indexing report in Google Search Console shows how many of your URLs are indexed and why the rest are not. It is the single most useful screen in a technical audit.

How do I read the Page indexing report?

Open Indexing → Pages and focus on the "Why pages aren't indexed" table. Not every reason is a problem: "Alternate page with proper canonical tag" and "Page with redirect" are usually expected. The reasons that deserve investigation are "Excluded by 'noindex' tag" on pages you want ranked, "Duplicate without user-selected canonical", "Duplicate, Google chose different canonical than user", "Soft 404", "Blocked by robots.txt" for important URLs, and the two discovery statuses: "Discovered – currently not indexed" and "Crawled – currently not indexed". The first often points to crawl budget or weak internal linking on large sites; the second usually means Google fetched the page but did not consider it valuable or distinct enough to index. Click a reason, export sample URLs, then run the URL Inspection tool on a few of them to see the Google-selected canonical, the last crawl date and the rendered HTML.

Step 3: Review robots.txt

robots.txt controls crawling, not indexing. A URL blocked in robots.txt can still appear in results (without a snippet) if other pages link to it, and Google cannot see a noindex tag on a page it is not allowed to fetch.

A typical, safe setup looks like this:

User-agent: *
Disallow: /cart/
Disallow: /checkout/
Disallow: /*?sessionid=

Sitemap: https://example.com/sitemap.xml

Check for these common mistakes:

  • Disallow: / left over from a staging environment.
  • Blocked CSS or JavaScript folders that Google needs to render the page.
  • Trying to remove pages from the index with robots.txt instead of noindex or a 404/410.
  • Rules for AI crawlers (GPTBot, ClaudeBot, PerplexityBot, Google-Extended) that accidentally block Googlebot too. Blocking Google-Extended does not affect Google Search; it only controls use of your content for Gemini. More on this in our guide to SEO for AI search.

Step 4: Validate XML sitemaps

A sitemap should list only URLs you want indexed: status 200, indexable, and self-canonical. Anything else sends mixed signals.

Check that:

  • The sitemap is referenced in robots.txt and submitted in Search Console.
  • It contains no redirected, 404, noindex or non-canonical URLs.
  • <lastmod> reflects real content changes, not the date the file was generated.
  • Each file stays under 50,000 URLs and 50 MB uncompressed; larger sites use a sitemap index.
  • On multilingual sites, all language versions are included (optionally with hreflang annotations — see our hreflang guide).

Step 5: Audit canonical tags and duplicates

Every indexable page should carry a self-referencing canonical with an absolute URL:

<link rel="canonical" href="https://example.com/services/seo-audit/">

Duplicates usually come from protocol and host variants (http://, www), trailing slashes, uppercase letters, tracking parameters, sorting and filtering, and printer or AMP versions. Your canonical, internal links, sitemap and redirects should all point to the same preferred URL.

Remember that a canonical is a hint, not a directive. If the URL Inspection tool shows "Google-selected canonical" different from yours, look for conflicting signals: internal links pointing to the other version, near-identical content, or a canonical that points to a redirected or noindex page.

Step 6: Fix redirects, chains and status codes

What is a redirect chain and why does it matter?

A redirect chain happens when one URL redirects to another that redirects again, for example http://example.com/page → https://example.com/page → https://www.example.com/page → https://www.example.com/page/. Each hop adds latency for users and an extra fetch for crawlers, and Googlebot follows only a limited number of hops (Google Search Central documents up to 10) before giving up. Chains also dilute clarity: the internal link, the sitemap entry and the canonical may each point to a different step. The fix is simple: redirect every old URL directly to its final destination in a single 301 (or 308), then update internal links and sitemaps to use the final URL so that the redirect is only a safety net for external links and bookmarks. Loops — A → B → A — are worse, because the page never resolves at all, and they usually appear after overlapping rules in the CMS and server configuration.

Also check:

  • 302 redirects used for permanent moves (use 301/308).
  • Soft 404s: "not found" pages returning 200.
  • 5xx errors in the Crawl stats report, which make Google slow down crawling.
  • Internal links to 3xx and 4xx URLs — fix the link, not just the redirect.

Step 7: Test JavaScript rendering

Google renders JavaScript, but rendering happens after crawling and can fail or be delayed. Other search engines and many AI crawlers do not execute JavaScript at all.

Compare the raw HTML (view-source) with the rendered HTML (URL Inspection → "View crawled page"). Check that the title, meta robots, canonical, main content, internal links and structured data are present in the server response, or at least in the rendered version. Common failures:

  • Links implemented as <div onclick> instead of <a href>.
  • Content loaded only after user interaction (clicks, scroll, tabs that fetch on open).
  • A noindex in the initial HTML that JavaScript later removes — Google may never render the page.
  • Blocked API or script files in robots.txt.

For content-heavy sites, server-side rendering or static generation remains the safest option.

Step 8: Check mobile experience

Google uses mobile-first indexing for all sites, so the mobile version is the version that gets indexed. Verify that mobile pages have the same main content, headings, internal links, structured data and meta tags as desktop. Hidden navigation is fine; missing content is not. Also check tap targets, font sizes, intrusive interstitials and a correct viewport tag:

<meta name="viewport" content="width=device-width, initial-scale=1">

Step 9: Measure Core Web Vitals

Core Web Vitals measure real-user experience: LCP (loading, good at ≤ 2.5 s), INP (responsiveness, good at ≤ 200 ms; it replaced FID in March 2024) and CLS (visual stability, good at ≤ 0.1). Google assesses them at the 75th percentile of field data from the Chrome UX Report (CrUX).

Use the Core Web Vitals report in Search Console to find failing URL groups (templates), then PageSpeed Insights to diagnose individual pages. Treat lab scores from Lighthouse as debugging aids, not the verdict. Core Web Vitals are part of page experience signals and work more like a tie-breaker than a primary ranking factor — but slow templates also hurt conversions. Our Core Web Vitals guide covers causes and fixes per metric.

Step 10: Validate structured data

Check JSON-LD with Google's Rich Results Test and the Schema Markup Validator. Markup must describe content that is actually visible on the page; mismatches can lead to manual actions. Prioritize types that still produce search features — Product, Review snippets, Article, Breadcrumb, Organization, LocalBusiness, Event, Video. Do not add FAQPage or HowTo markup expecting rich results: Google dropped HowTo rich results in 2023 and retired FAQ rich results for all sites in May 2026.

Step 11: Analyze server log files

Logs show what search engine bots actually request, not what a crawler simulates. Filter for verified Googlebot (reverse DNS lookup, since the user agent can be spoofed) and look for:

  • Share of bot hits spent on parameter URLs, redirects and errors.
  • Important sections that Googlebot rarely visits.
  • Spikes in 5xx responses or slow response times.
  • Activity from AI crawlers such as GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot, if you want to know how they use your site.

No log access? The Crawl stats report in Search Console (Settings → Crawl stats) gives a useful summary by response code, file type and Googlebot type.

Which tools do you need for a technical SEO audit?

Tool Cost Best for
Google Search Console Free Indexing, sitemaps, CWV field data, crawl stats
PageSpeed Insights Free Field + lab performance per URL
Rich Results Test Free Structured data eligibility
Bing Webmaster Tools Free Second index view, IndexNow
Screaming Frog / Sitebulb Freemium / paid Full crawls, JS rendering, redirects
Log analyzer (e.g. Screaming Frog Log File Analyser) Paid Real bot behavior

How do you prioritize technical SEO issues?

A technical audit on a real site can easily surface hundreds of warnings. Most of them do not matter. Prioritize each issue by three factors: impact (does it stop pages from being crawled or indexed, or merely slightly weaken a signal?), scope (one URL or an entire template?), and effort (a one-line config change or a replatforming project?). Blocking issues on revenue templates come first: an accidental noindex, a Disallow on a key folder, broken canonicals, server errors, content that does not render. Next come template-wide inefficiencies such as redirect chains in navigation, duplicate URL variants and failing Core Web Vitals on high-traffic groups. Cosmetic warnings — missing alt text on decorative images, long titles on archived posts — go last. Group fixes into a 30/60/90-day plan, assign an owner to each, and schedule a recheck so you can confirm that every fix was actually deployed and indexed.

Key takeaways

  • Audit in dependency order: discover → crawl → render → index → experience.
  • The Page indexing report and URL Inspection tool are your ground truth for what Google indexes.
  • robots.txt controls crawling, not indexing; use noindex or 404/410 to remove pages.
  • Canonicals, internal links, sitemaps and redirects must all agree on one preferred URL.
  • Don't implement FAQ or HowTo markup for rich results; they no longer appear.
  • Prioritize by impact × scope ÷ effort, and recheck after fixes ship.

Get a technical audit without doing it all yourself

If you would rather see where your site stands before investing hours in crawls and exports, you can request a free SEO audit preview: an SEO score plus several real issues with examples from your own pages. Curious what a full report looks like first? Browse the sample report.

All articles

Technical SEO

Core Web Vitals Explained: LCP, INP and CLS (2026 Guide)

Core Web Vitals are three Google metrics for real-user experience: LCP (loading, ≤ 2.5 s), INP (responsiveness, ≤ 200 ms) and CLS (visual stability, ≤ 0.1), measured at the 75th percentile of field data.

10 min read

International SEO

Hreflang Guide: How to Set Up Multilingual SEO Right

Hreflang is a signal that tells Google which language or regional version of a page to show each user. Add it with HTML link tags, HTTP headers or an XML sitemap, with self-references and return links.

9 min read

SEO audit

How Much Does an SEO Audit Cost? 2026 Price Guide

An SEO audit can cost nothing (automated tools) or run from a few hundred to several thousand dollars for a one-off professional audit, and more for large enterprise sites. Scope, site size and depth drive the price.

7 min read