Charlotte marketing agency websites are mostly open to AI search crawlers, but not entirely: as of September 2026, 15 of the 20 agency homepages we measured let every major AI crawler in, 3 of 20 return an error to OAI-SearchBot, the crawler that decides whether a site appears in ChatGPT search, and none block anything in robots.txt.
The headline findings
Twenty Charlotte-area agency homepages, measured on September 3, 2026, produced these numbers: 0 of 20 restrict an AI crawler in robots.txt, 5 of 20 refuse at least one AI crawler at the server, 16 of 16 fetchable homepages deliver real text in raw HTML, 19 of 20 carry JSON-LD, and 4 of 16 pass the first-100-words test.
| Measure | Count | What it means |
|---|---|---|
| robots.txt allows all six AI crawlers | 20 of 20 | Nobody opts out on paper |
| Serves HTTP 200 to all six crawler user agents | 15 of 20 | Access is lost at the server |
| Returns an error to OAI-SearchBot | 3 of 20 | Ineligible for ChatGPT search answers |
| Raw HTML under 500 words as GPTBot | 0 of 16 | No empty JavaScript shells |
| First-100-words test: pass | 4 of 16 | Most open with a tagline, not an answer |
| JSON-LD present | 19 of 20 | Near-universal; not required for AI features |
| Named founder or Person node, homepage or /about | 10 of 20 | 3 had neither; 7 had no /about to check |
Who blocks which crawler
Nobody in the sample blocks an AI crawler in robots.txt: all 20 files allow OAI-SearchBot, GPTBot, ClaudeBot, Claude-SearchBot, PerplexityBot and Applebot to fetch the homepage. Access is lost one layer down. Five of 20 homepages return HTTP 400 or 403 to at least one crawler's documented user agent while serving 200 to a browser.
The five split three ways. Two refuse all six user agents. One refuses OAI-SearchBot, ClaudeBot and Claude-SearchBot but serves GPTBot. Two refuse only GPTBot, the training crawler, and serve every search crawler, which is the exact separation OpenAI's crawler documentation describes: "Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links."
Three agencies wrote deliberate rules. Overtop Media, Knowmad Digital Marketing and Idea Forge Studios each name the AI crawlers in robots.txt and allow them; the other 17 files allow everything by default. Two of the three also allow Applebot-Extended, Apple's training opt-out, which Apple's documentation says does not crawl pages at all.
What an AI crawler actually sees
Every homepage we could fetch as GPTBot delivers its text in raw HTML. Across the 16 fetchable homepages the median is 1,147 words, the range is 603 to 5,432, and none falls under 500 words, the region where a client-rendered page starts to look empty to a crawler that does not run JavaScript.
That matters because, per Lantern's analysis updated in June 2026, no major AI crawler renders JavaScript. Med Rank Interactive and Overtop Media exceed 5,000 words of raw text on the homepage alone.
| Agency | Raw-HTML words as GPTBot | AI crawlers allowed | JSON-LD | Title has "Charlotte" |
|---|---|---|---|---|
| Crimson Park Digital | 603 | 6 of 6 | yes | no |
| Premier Marketing | 2,058 | 6 of 6 | yes | yes |
| The Social Rook | 1,191 | 6 of 6 | yes | yes |
| Epic Notion | 669 | 3 of 6 | yes | yes |
| Pinckney Harmon | n/a (HTTP 400) | 5 of 6 | yes | yes |
| Yellow Duck Marketing | 1,009 | 6 of 6 | yes | yes |
| WiT Group | 1,052 | 6 of 6 | yes | no |
| Southline Digital | n/a (HTTP 403) | 0 of 6 | yes | no |
| Med Rank Interactive | 5,432 | 6 of 6 | yes | yes |
| Overtop Media | 5,122 | 6 of 6 | yes | yes |
| LAIRE | 910 | 6 of 6 | no | no |
| Knowmad Digital Marketing | 1,375 | 6 of 6 | yes | no |
| CGR Creative | n/a (HTTP 400) | 5 of 6 | yes | no |
| The Branding Agency | 1,591 | 6 of 6 | yes | yes |
| Kashmer Interactive | 884 | 6 of 6 | yes | no |
| Idea Forge Studios | 1,204 | 6 of 6 | yes | yes |
| GoBeyond SEO | n/a (HTTP 403) | 0 of 6 | yes | yes |
| IMCG Creative | 875 | 6 of 6 | yes | yes |
| 3Bug Media | 1,103 | 6 of 6 | yes | no |
| Organic Clicks | 2,289 | 6 of 6 | yes | yes |
| Mirastart (author, not ranked) | 1,398 | 6 of 6 | yes | yes |
Structure and the first 100 words
Structure is where the sample is weakest. On the extractability script, which grades a page the way AI retrieval reads it, the 16 scored homepages have a median of 51 out of 100 and a range of 29 to 88. Only 4 of 16 pass the first-100-words check; 7 are partial and 5 fail.
The check asks whether the first 100 words of body text contain a definition or a sentence on the page's own topic, plus a number, without a pronoun opener. Surfer's June 2026 analysis of 100,000 citation placements across 10,000 AI Overviews found 38% of citations come from those words. Med Rank Interactive scored 88; Kashmer Interactive, The Social Rook and Organic Clicks also pass.
Schema, titles and canonicals
Structured data is nearly universal and mostly generic: 19 of 20 homepages carry JSON-LD, 16 declare an Organization or Corporation, 7 a LocalBusiness or ProfessionalService, 2 a FAQPage, and 4 a Person node anywhere in the graph. Google's guide is explicit: "Structured data isn't required for generative AI search, and there's no special schema.org markup you need to add."
Titles run long: the median tag is 64.5 characters, 13 of 20 exceed 60, and 12 of 20 contain the word Charlotte. Canonicals are clean on 19 of 20; one points at a different hostname.
What this means if you are choosing a Charlotte agency
Ask any agency to fetch its own homepage with the OAI-SearchBot user agent before you sign. An agency whose server turns that crawler away cannot be expected to know the rule for your site, and as of September 2026, 3 of 20 in this sample are in that position. Structure and named people are where the field separates.
Read the numbers with care: a 403 to a user-agent string is often a hosting firewall's default rather than a decision, fixed in minutes once noticed. What the table cannot tell you is whether the agency will notice on your site; ask that in the sales call.
Mirastart's own numbers
Mirastart's homepage, measured the same way on the same day: HTTP 200 to all six crawlers, a robots.txt that names all six and allows them, 1,398 words of raw HTML as GPTBot, an extractability score of 65 with a first-100-words pass, JSON-LD for Organization, LocalBusiness, ProfessionalService, WebSite and FAQPage, a 56-character title containing Charlotte, and a same-host canonical.
What it does badly: neither the homepage nor the About page carries a Person node; the founder is named in text only, and the Person schema sits on a team page this study did not look at. The score of 65 is mid-pack, well behind Med Rank Interactive's 88.
Method
The study measured 21 homepages once each on September 3, 2026, with stdlib Python scripts at one request per second, raw HTML only, no JavaScript rendering. The sample is 20 agencies with a Charlotte-area office or headquarters, drawn from our two Charlotte agency lists and Expertise.com's Charlotte SEO and digital marketing pages, plus Mirastart as a separately reported 21st row.
- Word count: homepage fetched with the GPTBot user agent used by extractability.py; script, style, noscript, template and SVG blocks removed; remaining visible words counted.
- Crawler access: robots.txt parsed with Python's urllib.robotparser for six tokens (no file counts as allowed), then the homepage fetched once under each vendor's documented user-agent string; non-200 responses re-checked once. Real crawlers arrive from published IP ranges, which a firewall may treat differently from a user-agent string; only the string was tested.
- Structure: scripts/extractability.py scored the raw HTML and returned the first-100-words verdict; JSON-LD, title and canonical were read from the same HTML.
- People: a Person node in JSON-LD or a visibly named founder on the homepage or the page at /about (then /about-us); a /about that redirected to the homepage or a blog post was recorded as unclear.
- Limits: homepages only, one measurement per site. Clutch's Charlotte lists returned HTTP 403 and were not used; Expertise.com's pages were last updated September 1, 2026. Raw data: docs/data/charlotte-agency-ai-search-readiness-2026.csv in the Mirastart repository, 21 rows.
Sources
- Raw data: the study's per-site measurements (CSV, CC BY 4.0) - every number in the tables above, one row per site
- OpenAI, Overview of OpenAI Crawlers - Documents OAI-SearchBot (search), GPTBot (training), ChatGPT-User and OAI-AdsBot as separate robots.txt settings; source of the quoted sentence.
- Anthropic, Does Anthropic crawl data from the web, and how can site owners block the crawler? - Names ClaudeBot, Claude-User and Claude-SearchBot and states all three respect robots.txt.
- Apple, About Applebot - Published June 8, 2026. Applebot powers Spotlight, Siri and Safari search; Applebot-Extended is a training opt-out that does not crawl pages.
- Google Search Central, Optimizing your website for generative AI features on Google Search - Last updated July 10, 2026. Source of the quoted sentence on structured data.
- Surfer, Why Almost 40% of AI Citations Come from Your First 100 Words - 100,000 citation placements across 10,000 AI Overviews, data collected June 2, 2026: 38% of citations come from the first 100 words.
- Lantern, AI Crawlers Do Not Render JavaScript - Updated June 11, 2026. As of June 2026 none of the major AI crawlers execute JavaScript; GPTBot downloads JavaScript files about 11.5% of the time without running them.
- Expertise.com, Best SEO Agencies in Charlotte, NC - Last updated September 1, 2026; 17 agencies with addresses and websites, used to assemble the sample.
- Expertise.com, Best Digital Marketing Agencies in Charlotte, NC - Last updated September 1, 2026; 18 agencies, used to assemble the sample.