AI Visibility
We Audited 45 Tampa Real Estate Websites. 1 in 4 Blocks ChatGPT.
Short answer: We ran every Tampa real estate agency listed on Google Maps through an AI-visibility audit. 24% actively block GPTBot (the crawler that feeds ChatGPT) with a hard 403. 42% publish no structured data at all, so AI can’t tell what the business even is. Put together, 44% are effectively invisible to AI answer engines. When a buyer moving to Tampa asks ChatGPT “who’s a good agent here,” nearly half the market can’t be recommended, no matter how good they are.
Original research by Specularis, July 2026. Raw data and methodology are published below.
What did the study find?
We audited 45 Tampa real estate websites (neutral market sample, see methodology).
| Finding | Result |
|---|---|
| Actively block GPTBot (hard 403) | 24% (11 of 45) |
| …of those, blocked by Cloudflare | 5 |
| Publish no structured data at all | 42% (19 of 45) |
| Have any Organization/LocalBusiness schema | 33% (15 of 45) |
| Have proper RealEstateAgent schema | 11% (5 of 45) |
Have a real llms.txt | 13% (6 of 45) |
| Effectively invisible to AI | 44% (20 of 45) |
Two separate failures are stacking on top of each other.
How many Tampa real estate websites block ChatGPT?
Eleven of forty-five, or 24%, return a hard 403 Forbidden when the request comes from OpenAI’s GPTBot user-agent. Five of those blocks come from Cloudflare, which serves a JavaScript challenge page (“Just a moment…”) that a crawler can never solve.
Almost none of these firms chose this. A 403 to GPTBot is usually a default setting in a security product, switched on by a web host or a well-meaning IT contractor. The business never made a decision. It just quietly disappeared from AI answers.
Does blocking AI crawlers actually protect your website?
No. This is the most expensive misconception in the data.
robots.txt is a polite request, not a wall. The bots that actually threaten your site, like scrapers, credential stuffers, and DDoS traffic, ignore robots.txt entirely. So disallowing GPTBot buys you exactly zero protection from attackers. What it does buy you is removal from the answer when a buyer asks an AI for a recommendation.
Real protection lives at the firewall/WAF layer, and it stays fully intact whether or not you let the well-behaved AI crawlers through. Allowing GPTBot and PerplexityBot does not make you less secure.
Why does structured data matter so much here?
Getting crawled isn’t enough. AI also has to understand what you are. 42% of the sites we audited publish no structured data whatsoever, and only 11% use RealEstateAgent schema, the machine-readable label that says “this is a real estate professional, in this city.”
Without it, an AI engine sees text and photos of houses. It can’t confidently say “this is a real estate agent serving Tampa.” So when it’s asked to name three agents, it reaches for the ones it can verify. You don’t lose to a better agent. You lose to a more legible one.
We ran the identical audit in Miami and found the same story, only worse: 63% of Miami real estate sites are invisible to AI. Two markets, one pattern.
How we ran the study (methodology)
- Sample. Pulled every “real estate agency” listing in Tampa, FL from Google Maps that has a website (July 2026), then deduplicated to unique root domains. No pre-filtering for AI visibility.
- Exclusions. Removed 3 sites unreachable or returning 404 to every user agent, including a normal browser. Broken sites aren’t “blocking AI.”
- Re-testing. One site returned a 429 (rate limit); on re-test it returned 200, so we reclassified it as allowed. Transient errors are not blocks.
- Crawler test. Requested each homepage using OpenAI’s GPTBot user-agent. A hard 403, or a Cloudflare challenge, counts as blocked.
- Identity test. Counted
application/ld+jsonblocks and checked forRealEstateAgent,LocalBusiness, andOrganizationtypes. - Emerging signals. Checked
robots.txtfor AI directives and/llms.txt.
We report aggregates only. We don’t name individual businesses.
How can I check my own website?
- Crawler access. Open
yoursite.com/robots.txtand look forDisallowunder GPTBot, ClaudeBot, PerplexityBot. Then check whether your host or Cloudflare has an “AI scrapers” block on. - Structured data. View your homepage source and search for
application/ld+json. Nothing there means no machine-readable identity. - The real test. Ask ChatGPT and Perplexity “who’s a good real estate agent in Tampa?” and see if you appear.
Or run the free AI visibility audit. It scores your site 0 to 100 across five pillars and emails the exact fixes. Free, no credit card.
FAQ
How many Tampa real estate websites block ChatGPT?
24%. Eleven of the 45 sites we audited return a hard 403 to GPTBot, the crawler behind ChatGPT. Five of those blocks come from Cloudflare.
Does blocking GPTBot protect my site from attacks?
No. Malicious bots ignore robots.txt and crawler rules entirely. Blocking AI crawlers provides no security benefit. It only removes you from AI-generated answers. Real protection comes from your firewall/WAF, which keeps working whether or not AI crawlers are allowed.
What percentage of real estate sites have proper schema?
Only 11% of the sites we audited use RealEstateAgent schema, and 42% publish no structured data at all, so AI cannot reliably determine what the business is.
What is llms.txt and do agents need it?
llms.txt is a plain-text file at your site root that tells AI what your site is about. It's an emerging standard, and just 13% of the sites we audited had one. It won't fix a blocked crawler, but it helps AI understand you faster.
Is this study repeatable?
Yes. The method is simple and reproducible: sample from Google Maps, request each homepage as GPTBot, and check for structured data. Our raw data is published alongside this article.
Want to know your number?
Run the free AI visibility audit. It scores your site 0–100 across five pillars and emails the exact fixes in minutes.
Run the free audit →