A website SEO audit is a sequence of checks establishing, in order, whether a page can be fetched, indexed, understood and cited. Dependency order matters more than completeness: a page missing from the raw HTML fails every later check, and no title fix will save it.
The eight checks below run in dependency order, each with a free way to run it. Platform facts are stamped as of September 2026 and attributed to the document stating them, because three items most 2026 checklists carry are ones Google's documentation says do nothing.
Run the checks in the order a site fails, not in category order
Audit order should follow dependency, not category. A page absent from the raw HTML cannot be indexed. An unindexed page cannot be cited. A page competing against three of its own siblings cannot rank whatever its title says. Category sorting hides those dependencies.
| The check, in failure order | Free way to run it | What failing it costs |
|---|---|---|
| 1. Is the page text in the HTML the server returns? | View source, or reload with JavaScript off | AI crawlers read raw HTML only |
| 2. Is the URL on Google, and snippet-eligible? | Search Console URL Inspection | Nothing below reaches an unindexed page |
| 3. Do two of your pages chase this query? | A site: search on your own domain | Two pages split one signal |
| 4. Do the first 100 words answer the query? | Read the top and look for the answer | 38% of AI Overview citations sit there |
| 5. Is the title unique, with the intent phrase first? | Search Console Performance, by page | Duplicate titles are a cannibalization tell |
| 6. Does anything on the site link here? | A site: search, then find-in-page | An orphan is crawled late |
| 7. Does the markup match the visible text? | Rich Results Test, read beside the page | Markup claiming what nobody can see |
| 8. Does the page hold still on a phone? | PageSpeed Insights, field data | A slow page loses the click |
Step one: read the page the way an AI crawler reads it
Open the page source, or reload with JavaScript disabled, and read what is left. That is what the AI crawlers get: Vercel's analysis of its own network traffic, run with MERJ, reports that none of the major AI crawlers render JavaScript.
In that sample, ChatGPT's crawler fetched JavaScript in about 11.5% of requests and Claude's in about 23.84%, and neither ran it. Googlebot renders; Gemini inherits that from Google's infrastructure. The test is not whether the page looks right in a browser, but whether the sentence you want quoted is there before any script runs.
The second half of this check is robots.txt. OpenAI documents OAI-SearchBot for ChatGPT search and GPTBot for training, says each setting is independent, and states that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers. Anthropic documents ClaudeBot, Claude-User and Claude-SearchBot the same way. A blanket disallow copied from a template usually blocks the wrong one.
Step two: confirm the page is indexed and eligible for a snippet
Paste the URL into Search Console's URL Inspection tool and read two things: whether the status says the URL is on Google, and what View crawled page returns as the HTML Google actually fetched. Google's Search Console documentation describes both, and together they answer step one and step two at once.
Eligibility is stated plainly. Google's "AI features and your website" documentation says: "To be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet, fulfilling the Search technical requirements." A stray noindex, a leftover nosnippet and a wrong canonical are the cheapest findings an audit returns.
Step three: one page per intent, and the search that finds the duplicates
Search your own domain with a site: query plus the phrase you think a page owns. If two or three of your pages come back, they are splitting one signal, and the fix is consolidation rather than more work on each. An answer engine choosing between near-duplicates may cite neither.
Google's AI-optimization guide, last updated July 2026, says creating separate content for every possible variation of how people might search, done primarily to manipulate rankings or generative AI responses, violates its scaled content abuse spam policy. One page per intent, and a redirect for the loser.
Step four: the first 100 words, where most AI citations come from
Read the first 100 words and ask whether a reader who saw only that would have their question answered. Surfer's analysis of more than 100,000 AI citation placements across over 10,000 AI Overviews responses, collected on 2 June 2026, found 38% of citations came from there.
In the same analysis, pages that confirmed and answered the query early were cited 45% of the time against 23%. Surfer says the data is correlational and it sells a product built on the finding, so read it as a strong pattern, not a proven mechanism.
Step five: titles, internal links and markup that matches the page
Three checks share one failure mode: telling a search engine something the page does not support. A title duplicated across two pages, a page nothing internal links to, and structured data describing content the reader cannot see are the same error. Search Console and the Rich Results Test answer all three free.
Sort Search Console's Performance report by page and look for two of your pages matching one query: that is step three, in data rather than in a site: search. For orphans, run the site: search and check whether anything else links to the page. For markup, Google's AI-features documentation lists making sure structured data matches the visible text among the things worth doing.
The three checks to take off your list, in Google's own words
Three items appear on most 2026 audit checklists that Google's documentation says do nothing for Search: publishing an llms.txt file, chunking content into small pieces for AI, and adding special schema for AI features. Google's AI-optimization guide, last updated July 2026, addresses all three.
On the first, that guide says: "You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn't use them." On the second, no requirement to break content into tiny pieces. On the third, structured data is not required for generative AI search and there is no special markup to add.
Ahrefs measured the same file from the other side. Across the 137,210 domains in its Web Analytics data with traffic in May 2026, 28% published an llms.txt file and 97% of those were fetched by nothing that month; no AI bot requested one that did not exist. None of this makes structured data pointless. It makes "add AI markup" a line item with no documented effect.
What the fixes look like when they run as a system
An audit produces a list; a system stops the list refilling. Mirastart builds pages as server-rendered HTML so the text is there before any script runs, keeps one page per query by design, and puts indexing, calls and bookings on one screen, so a regression shows up the week it happens.
The booking system behind this site's own discovery-call page calculates real availability and confirms automatically; the online booking and repair-status software Quick Auto NC runs its bays on is ours, and so is the Carolina Sky Painting website. Audit findings get built into those, not a document that ages.
Sources
- Google Search Central, "AI features and your website" - States that to be eligible as a supporting link in AI Overviews or AI Mode a page must be indexed and eligible to be shown in Google Search with a snippet, and lists making sure structured data matches the visible text among the things worth doing. Page footer reads Last updated 2025-12-10 UTC.
- Google Search Central, "Optimizing your website for generative AI features on Google Search" - Google's statement that no machine readable files, AI text files, markup or Markdown are needed because Search does not use them; that there is no requirement to chunk content; that structured data is not required for generative AI search; and that spinning out content per query variation to manipulate rankings violates the scaled content abuse policy. Page footer reads Last updated 2026-07-10 UTC.
- Vercel, "The rise of the AI crawler" - Vercel's network analysis with MERJ: none of the major AI crawlers render JavaScript; ChatGPT's crawler fetched JavaScript files in about 11.50% of requests and Claude's in about 23.84% without executing them; Gemini inherits Google's rendering. Read via search result summary on September 29, 2026; the domain is blocked to direct fetches from this environment.
- Surfer, "Why Almost 40% of AI Citations Come from Your First 100 Words" - Over 100,000 AI citation placements across more than 10,000 AI Overviews responses, collected June 2, 2026: 38% of citations come from a page's first 100 words, and pages that confirm and answer the query early were cited 45% of the time against 23%. Surfer states the data is correlational and sells a product built on the finding.
- Ahrefs, "We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read" - Published June 15, 2026. Of 137,210 domains in Ahrefs Web Analytics with traffic in May 2026, 28% published an llms.txt file and 97% of those received zero requests; no AI bot requested an llms.txt file that did not exist. The authors note their customer base skews technical, so the adoption figure is an upper bound.
- OpenAI, "Overview of OpenAI Crawlers" - OAI-SearchBot surfaces sites in ChatGPT's search features and GPTBot crawls for model training; each robots.txt setting is independent of the others, and sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers. The page does not say whether the agents execute JavaScript.
- Anthropic Help Center, "Does Anthropic crawl data from the web?" - Documents ClaudeBot (training), Claude-User (user-directed fetches) and Claude-SearchBot (search quality) as separate agents controlled separately in robots.txt, and states the bots honor robots.txt directives.
- Google Search Console Help, "URL Inspection tool" - Documents the index status values, including that URL is on Google means the page has been indexed, and the View crawled page control that shows the raw HTML returned along with HTTP headers and page resources. Read via search result summary on September 29, 2026.