The Website SEO Audit Checklist, in the Order a Site Actually Fails

Perry Lam · FounderPublished

A website SEO audit is a sequence of checks establishing, in order, whether a page can be fetched, indexed, understood and cited. Dependency order matters more than completeness: a page missing from the raw HTML fails every later check, and no title fix will save it.

The eight checks below run in dependency order, each with a free way to run it. Platform facts are stamped as of September 2026 and attributed to the document stating them, because three items most 2026 checklists carry are ones Google's documentation says do nothing.

Run the checks in the order a site fails, not in category order

Audit order should follow dependency, not category. A page absent from the raw HTML cannot be indexed. An unindexed page cannot be cited. A page competing against three of its own siblings cannot rank whatever its title says. Category sorting hides those dependencies.

The eight checks in failure order, with the free way to run each, as of September 2026
The check, in failure orderFree way to run itWhat failing it costs
1. Is the page text in the HTML the server returns?View source, or reload with JavaScript offAI crawlers read raw HTML only
2. Is the URL on Google, and snippet-eligible?Search Console URL InspectionNothing below reaches an unindexed page
3. Do two of your pages chase this query?A site: search on your own domainTwo pages split one signal
4. Do the first 100 words answer the query?Read the top and look for the answer38% of AI Overview citations sit there
5. Is the title unique, with the intent phrase first?Search Console Performance, by pageDuplicate titles are a cannibalization tell
6. Does anything on the site link here?A site: search, then find-in-pageAn orphan is crawled late
7. Does the markup match the visible text?Rich Results Test, read beside the pageMarkup claiming what nobody can see
8. Does the page hold still on a phone?PageSpeed Insights, field dataA slow page loses the click

Step one: read the page the way an AI crawler reads it

Open the page source, or reload with JavaScript disabled, and read what is left. That is what the AI crawlers get: Vercel's analysis of its own network traffic, run with MERJ, reports that none of the major AI crawlers render JavaScript.

In that sample, ChatGPT's crawler fetched JavaScript in about 11.5% of requests and Claude's in about 23.84%, and neither ran it. Googlebot renders; Gemini inherits that from Google's infrastructure. The test is not whether the page looks right in a browser, but whether the sentence you want quoted is there before any script runs.

The second half of this check is robots.txt. OpenAI documents OAI-SearchBot for ChatGPT search and GPTBot for training, says each setting is independent, and states that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers. Anthropic documents ClaudeBot, Claude-User and Claude-SearchBot the same way. A blanket disallow copied from a template usually blocks the wrong one.

Step two: confirm the page is indexed and eligible for a snippet

Paste the URL into Search Console's URL Inspection tool and read two things: whether the status says the URL is on Google, and what View crawled page returns as the HTML Google actually fetched. Google's Search Console documentation describes both, and together they answer step one and step two at once.

Eligibility is stated plainly. Google's "AI features and your website" documentation says: "To be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet, fulfilling the Search technical requirements." A stray noindex, a leftover nosnippet and a wrong canonical are the cheapest findings an audit returns.

Step three: one page per intent, and the search that finds the duplicates

Search your own domain with a site: query plus the phrase you think a page owns. If two or three of your pages come back, they are splitting one signal, and the fix is consolidation rather than more work on each. An answer engine choosing between near-duplicates may cite neither.

Google's AI-optimization guide, last updated July 2026, says creating separate content for every possible variation of how people might search, done primarily to manipulate rankings or generative AI responses, violates its scaled content abuse spam policy. One page per intent, and a redirect for the loser.

Step four: the first 100 words, where most AI citations come from

Read the first 100 words and ask whether a reader who saw only that would have their question answered. Surfer's analysis of more than 100,000 AI citation placements across over 10,000 AI Overviews responses, collected on 2 June 2026, found 38% of citations came from there.

In the same analysis, pages that confirmed and answered the query early were cited 45% of the time against 23%. Surfer says the data is correlational and it sells a product built on the finding, so read it as a strong pattern, not a proven mechanism.

Three checks share one failure mode: telling a search engine something the page does not support. A title duplicated across two pages, a page nothing internal links to, and structured data describing content the reader cannot see are the same error. Search Console and the Rich Results Test answer all three free.

Sort Search Console's Performance report by page and look for two of your pages matching one query: that is step three, in data rather than in a site: search. For orphans, run the site: search and check whether anything else links to the page. For markup, Google's AI-features documentation lists making sure structured data matches the visible text among the things worth doing.

The three checks to take off your list, in Google's own words

Three items appear on most 2026 audit checklists that Google's documentation says do nothing for Search: publishing an llms.txt file, chunking content into small pieces for AI, and adding special schema for AI features. Google's AI-optimization guide, last updated July 2026, addresses all three.

On the first, that guide says: "You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn't use them." On the second, no requirement to break content into tiny pieces. On the third, structured data is not required for generative AI search and there is no special markup to add.

Ahrefs measured the same file from the other side. Across the 137,210 domains in its Web Analytics data with traffic in May 2026, 28% published an llms.txt file and 97% of those were fetched by nothing that month; no AI bot requested one that did not exist. None of this makes structured data pointless. It makes "add AI markup" a line item with no documented effect.

What the fixes look like when they run as a system

An audit produces a list; a system stops the list refilling. Mirastart builds pages as server-rendered HTML so the text is there before any script runs, keeps one page per query by design, and puts indexing, calls and bookings on one screen, so a regression shows up the week it happens.

The booking system behind this site's own discovery-call page calculates real availability and confirms automatically; the online booking and repair-status software Quick Auto NC runs its bays on is ours, and so is the Carolina Sky Painting website. Audit findings get built into those, not a document that ages.

Sources

  1. Google Search Central, "AI features and your website" - States that to be eligible as a supporting link in AI Overviews or AI Mode a page must be indexed and eligible to be shown in Google Search with a snippet, and lists making sure structured data matches the visible text among the things worth doing. Page footer reads Last updated 2025-12-10 UTC.
  2. Google Search Central, "Optimizing your website for generative AI features on Google Search" - Google's statement that no machine readable files, AI text files, markup or Markdown are needed because Search does not use them; that there is no requirement to chunk content; that structured data is not required for generative AI search; and that spinning out content per query variation to manipulate rankings violates the scaled content abuse policy. Page footer reads Last updated 2026-07-10 UTC.
  3. Vercel, "The rise of the AI crawler" - Vercel's network analysis with MERJ: none of the major AI crawlers render JavaScript; ChatGPT's crawler fetched JavaScript files in about 11.50% of requests and Claude's in about 23.84% without executing them; Gemini inherits Google's rendering. Read via search result summary on September 29, 2026; the domain is blocked to direct fetches from this environment.
  4. Surfer, "Why Almost 40% of AI Citations Come from Your First 100 Words" - Over 100,000 AI citation placements across more than 10,000 AI Overviews responses, collected June 2, 2026: 38% of citations come from a page's first 100 words, and pages that confirm and answer the query early were cited 45% of the time against 23%. Surfer states the data is correlational and sells a product built on the finding.
  5. Ahrefs, "We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read" - Published June 15, 2026. Of 137,210 domains in Ahrefs Web Analytics with traffic in May 2026, 28% published an llms.txt file and 97% of those received zero requests; no AI bot requested an llms.txt file that did not exist. The authors note their customer base skews technical, so the adoption figure is an upper bound.
  6. OpenAI, "Overview of OpenAI Crawlers" - OAI-SearchBot surfaces sites in ChatGPT's search features and GPTBot crawls for model training; each robots.txt setting is independent of the others, and sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers. The page does not say whether the agents execute JavaScript.
  7. Anthropic Help Center, "Does Anthropic crawl data from the web?" - Documents ClaudeBot (training), Claude-User (user-directed fetches) and Claude-SearchBot (search quality) as separate agents controlled separately in robots.txt, and states the bots honor robots.txt directives.
  8. Google Search Console Help, "URL Inspection tool" - Documents the index status values, including that URL is on Google means the page has been indexed, and the View crawled page control that shows the raw HTML returned along with HTTP headers and page resources. Read via search result summary on September 29, 2026.
Questions

Website SEO audit checklist, answered.

Do I need a paid crawler to audit a small site?

Not for a site of a few dozen pages. Every check above has a free way to run it: view source or JavaScript off for raw HTML, Search Console URL Inspection for indexing and the fetched HTML, a site: search for duplicates and orphans, the Rich Results Test for markup, PageSpeed Insights for speed. A paid crawler earns its cost when the site is large enough that you cannot hold the page list in your head, or when you need the same audit repeated on a schedule without a person doing it. Below that, the tool is not the bottleneck; running the checks in order is.

How often should a site be audited?

Steps one and two are worth re-running after every deploy that touches templates, routing or the robots file, because those are the changes that silently remove a page from the index. The rest is a quarterly job for a stable site. The exception is a redesign that changes URLs, which resets everything and deserves a full pass before and after the switch, with the old-to-new URL mapping checked line by line rather than sampled.

Does any of this help with ChatGPT and Perplexity, or only with Google?

Steps one, three and four apply to every engine, and step one applies most of all to the non-Google ones. Vercel's analysis reports that the AI crawlers read raw HTML without executing JavaScript, so a page whose text only exists after a script runs is invisible to ChatGPT and Perplexity while remaining perfectly visible to Googlebot, which does render. Step two is Google-specific in its tooling, though the underlying idea is not: an engine cannot cite a page it was never allowed to fetch. OpenAI, Anthropic and Perplexity each document their own user agents, and robots.txt is where you answer them.

Keep reading
Charlotte

We Tested 20 Charlotte Marketing Agency Websites for AI Search Readiness (2026)

An original measurement of 20 Charlotte-area agency homepages: which AI crawlers get in, what they see in raw HTML, how the first 100 words read, and what the schema says. Mirastart measured itself the same way and reports it separately.

Read
Comparison

SEO vs. GEO: How Charlotte Businesses Get Found in 2026

Ranking on Google and being recommended by AI are now two different games. Here is what each one takes, and where they overlap.

Read
Guide

How Long Does SEO Take? An Honest Timeline, and the Three Clocks Behind It

Most timelines answer this with a range and no mechanism. This one ties each window to the Google-documented process that governs it, separates the three clocks that run at different speeds, and names what to check at 30, 90 and 180 days instead of what to hope for.

Read
Guide

What Is Generative Engine Optimization (GEO)? A Definition

A plain definition of generative engine optimization, built from the three content changes the paper that coined the term actually measured and the one sentence Google's own documentation uses to draw the boundary.

Read
Guide

How to Read a Marketing Agency Report: Six Numbers and What They Prove

Most guides to reading an agency report define the metrics. This one names the columns that cannot mean what their labels imply, using the platforms' own documentation, and gives the question that turns each one into a number about your business.

Read
Charlotte

The Best SEO Companies in Charlotte (2026)

Twelve SEO firms with a Charlotte-area address, the national agencies holding page one with a Charlotte city page, and the three things that changed in local search this year. Written by one of the firms on the list, and disclosed.

Read
Charlotte, NC

Charlotte SEO, from inside Charlotte

SEO for Charlotte businesses from a Charlotte company: map pack visibility, local service pages, reviews, and the AI answers customers now read.

In Charlotte

SEO company in Charlotte, NC

Local search, technical SEO, content strategy, and visibility in the AI answers your customers now read.

The practice

Web design and development in Charlotte, NC

Fast, accessible business sites, eCommerce experiences, landing pages, redesigns and booking journeys, built to load quickly and point at one obvious next step.

The practice