← Back to Blog
Uncategorized

How to Test Your AI Visibility: A 5-Platform Prompt Framework

Most brands assume they show up in AI-generated answers because they rank on Google. Those are two different things. This framework tests your real visibility across ChatGPT, Gemini, Claude, Perplexity, and Google AI Mode.

July 24, 2026Alex Rodriguezai visibility testinganswer engine optimizationchatgpt gemini claude perplexity
FIG. 01Uncategorized — Visual Reference
How to Test Your AI Visibility: A 5-Platform Prompt Framework

Most people don't know whether they show up in AI-generated answers. They assume they do because they rank on Google. Those are two completely different things, and confusing them is costing real business.

This is the third article in a three-part series. Part one showed why different AI engines read your pages differently — Gemini parses JSON-LD directly, ChatGPT and Claude only see visible prose. Part two gave you the dual-channel checklist to fix that gap. This article gives you the testing framework to verify whether your fixes are actually working — and to establish a baseline before you touch anything.

The framework covers five platforms: ChatGPT, Gemini, Claude, Perplexity, and Google AI Mode. Each one retrieves and surfaces information differently. A brand can dominate one and be completely invisible in another. You need to test all five to know where you actually stand.


Why Your Google Rankings Tell You Nothing About AI Visibility

There is only 8–12% overlap between URLs cited by ChatGPT and top-10 Google rankings for the same commercial queries.[^1] For comparison queries, the correlation is actually negative. Strong Google rankings and zero AI visibility coexist all the time.

The reason is architectural. Traditional search ranks pages by authority and relevance signals. AI answer engines select sources based on how clearly a page answers a specific question — entity definition, structured answer blocks, prose clarity, and third-party validation. These are different selection criteria, and optimizing for one does not automatically improve the other.

There is also a measurement gap. Google Search Console tells you impressions, clicks, and position. It tells you nothing about whether ChatGPT recommended you to someone who never clicked a link. AI-sourced traffic converts at 4.4 times the rate of organic traffic,[^2] which means the channel you can't see in your analytics is often the most valuable one.

The only way to know your AI visibility is to test it directly.


The 5-Platform Testing Framework

Each platform has a different retrieval architecture, citation behavior, and content preference. You need a consistent prompt set run across all five to get a complete picture. Here is how each one works and what to look for.

Platform 1: ChatGPT (Web Browsing Mode)

How it retrieves: ChatGPT pulls from Bing's index when web browsing is enabled. It does not parse JSON-LD schema — it reads rendered DOM text only. Free and Plus versions behave differently; Plus users see fresher results and more citations.

What to test: Always use a fresh incognito session. Prior conversation history shapes responses. Run each prompt three times across separate sessions to account for variability — only 30% of brands stay visible from one AI answer to the next.[^3]

Prompt set for ChatGPT:

Prompt TypeExample Prompt
Brand recognition"What does [Your Brand] do?"
Problem-solution"What are the best options for [your service category] in [your city/region]?"
Comparison"How does [Your Brand] compare to [Competitor]?"
Expertise"Who are the leading experts in [your specialty]?"
Specific service"What is [specific service you offer] and who provides it?"

What to record: Direct citation with link, brand mention without link, no mention. Note the exact language used to describe you — that is the AI's current understanding of your positioning.


Platform 2: Gemini (Google AI Mode)

How it retrieves: Gemini sits on Google's native indexing infrastructure and directly parses `@graph` JSON-LD. It is the most schema-aware of the five platforms. If your structured data is correct, Gemini will extract it accurately — as demonstrated in the live test documented in part one of this series.

What to test: Gemini is your ground truth for schema accuracy. If Gemini gets your entity data wrong, your JSON-LD has an error. If Gemini gets it right but ChatGPT gets it wrong, your prose reinforcement is insufficient.

Prompt set for Gemini:

Prompt TypeExample Prompt
Entity extraction"What type of business is [Your Brand] and what do they specialize in?"
Contact/location"Where is [Your Brand] located and how do you contact them?"
Service specifics"What services does [Your Brand] offer?"
Relationship mapping"What is the relationship between [Your Brand] and [Parent Brand/Org]?"
Authority signal"What is [Your Brand] known for in [your industry]?"

What to record: Whether Gemini correctly identifies your entity type, location, services, and organizational relationships. Any error here is a schema error, not a prose error.


Platform 3: Claude (with Web Search)

How it retrieves: Claude uses a hard-wall approach — it only fetches URLs that actively appear in live search result snippets. New subdomains, recently launched pages, or pages with no index footprint will trigger a refusal rather than a hallucinated answer. This is actually a useful signal: if Claude refuses to answer, your page has no search footprint yet.

What to test: Claude's refusal behavior is diagnostic. A refusal means the page isn't indexed or visible enough to appear in search snippets. A successful answer that gets your details wrong means your visible prose is unclear.

Prompt set for Claude:

Prompt TypeExample Prompt
Existence check"Can you find information about [Your Brand] at [your URL]?"
Positioning"What does [Your Brand] specialize in?"
Credibility"What evidence exists that [Your Brand] is an authority in [your field]?"
Specific claim"Does [Your Brand] work with [specific client type or industry]?"
Direct comparison"Is [Your Brand] a good fit for [specific use case]?"

What to record: Refusal (no index footprint), successful answer with correct details, successful answer with incorrect details. The last category is the most actionable — it tells you exactly what your visible prose is communicating incorrectly.


Platform 4: Perplexity

How it retrieves: Perplexity is the most transparent of the five. It always shows a numbered sources panel, crawls the web in near real time, and blends multiple search indexes. The visible sources panel makes it the easiest platform to audit — you can see exactly which pages are influencing the answer.

What to test: Perplexity is your best tool for competitive share-of-voice analysis. Run the same category queries you use for ChatGPT and compare which brands appear in the sources panel. If your competitors are in the panel and you are not, that is a direct content gap signal.

Prompt set for Perplexity:

Prompt TypeExample Prompt
Category query"What are the best [your service type] providers in [your market]?"
Problem query"I need help with [specific problem you solve] — who should I talk to?"
Comparison"Compare [Your Brand] vs [Competitor] for [use case]"
Recent activity"What has [Your Brand] published recently about [your topic]?"
Local/geo"[Your service] near [your city] — who are the top options?"

What to record: Whether your domain appears in the sources panel, your position in the panel (first-cited sources shape the answer more than later ones), and which competitors appear alongside or instead of you.


Platform 5: Google AI Overviews

How it retrieves: Google AI Overviews pull from Google's index using a combination of traditional ranking signals and structured data. They appear at the top of search results for informational queries and are increasingly appearing for local and service queries.

What to test: AI Overviews are the highest-traffic AI surface for most service businesses. They are also the most directly connected to your traditional SEO work — pages that rank well and have strong structured data are more likely to be cited.

Prompt set for Google AI Overviews:

Prompt TypeExample Prompt
Definitional"What is [your specialty or service type]?"
How-to"How do I [specific task you help with]?"
Local intent"[Your service] in [your city]"
Comparison"Best [your service type] for [specific use case]"
FAQ"[Common question your clients ask]"

What to record: Whether you appear in the AI Overview box (not just in organic results below it), the exact language used, and whether your structured data is being pulled into the answer.


Building Your Baseline: The Scoring Sheet

Run all five platforms against the same core prompt set. Score each result on a simple three-point scale:

ScoreMeaning
2Cited with correct details and positive/neutral framing
1Mentioned without citation, or cited with incorrect details
0Not mentioned

A perfect score across all five platforms and all five prompt types is 50 points. Most sites start between 10 and 25. The score is not the goal — the gap analysis is. Where you score 0 tells you where to focus first.

Priority order for fixing gaps:

  1. 1.Gemini scores 0 on entity prompts — your JSON-LD has an error. Fix the schema first.
  2. 2.Gemini scores 2, ChatGPT/Claude score 0 — your prose reinforcement is missing. Apply the dual-channel checklist.
  3. 3.Perplexity sources panel shows competitors but not you — you have a content gap. They have published content on that topic; you have not.
  4. 4.Claude refuses — your page has no search footprint. You need indexation, backlinks, or both.
  5. 5.Google AI Overviews show competitors — your structured data or content format is not matching the query intent. Review the page structure against the AI Overview that is showing.

The Repeat Testing Protocol

A single test is a snapshot. AI visibility is probabilistic — only 30% of brands stay visible from one answer to the next, and only 20% remain visible across five consecutive runs.[^3] You need a cadence.

Minimum viable testing schedule:

  • Baseline audit: Run the full 5-platform framework before making any changes. Document every result.
  • Post-implementation check: Run the same prompts 2–3 weeks after making schema or prose changes. AI systems need time to re-crawl and re-index.
  • Monthly monitoring: Run the brand recognition and category prompts monthly across all five platforms. Track movement, not just snapshots.
  • Quarterly competitive sweep: Add your top three competitors to the comparison prompts. Share of voice matters as much as absolute citation rate.

The metric that matters most is trajectory. Moving from 12 to 22 points on your baseline score over a quarter means your fixes are working. Staying flat means something in the implementation is wrong.


What Good Looks Like: Benchmark Targets

Based on the research and real-world testing, here are working targets for service businesses:

MetricStarting BenchmarkStrong PerformanceMarket Leader
Branded query citation rate40–60%80%+90%+
Category query citation rate5–10%15–25%30%+
Gemini entity accuracyCorrect on 3/5 promptsCorrect on 5/5 promptsCorrect + rich detail
Perplexity sources panelAppears in 1–2 category queriesAppears in 3–4Appears in 4–5
Google AI Overview inclusion0–1 queries2–3 queries3–5 queries

If your branded citation rate is below 60%, that is the first thing to fix. AI systems should be able to find clear, consistent information about your brand when someone searches for you by name. If they can't, your entity definition is weak or your information is inconsistent across sources.


The Three-Article Loop

This framework closes the loop on the three-article series:

Article 1Why ChatGPT Can't Read Your Schema (and Gemini Can): Diagnosed the problem. Different AI engines have different retrieval architectures. Your JSON-LD works for index-based systems; your prose works for fetch-based systems. Most sites are only doing one.

Article 2How to Reinforce Schema in Prose: The Dual-Channel Checklist: Gave the fix. Twenty specific prose reinforcement patterns mapped to schema data categories. Covers entity identity, contact information, services, authority signals, and profile links.

Article 3 — This article: Gives the measurement method. Run the 5-platform prompt framework, build your baseline score, prioritize gaps in the right order, and test on a cadence that accounts for AI answer variability.

The sequence is: diagnose → fix → measure → repeat. If your scores are not moving after implementing the dual-channel checklist, come back to the testing framework and look at where the gaps are. The framework will tell you whether the problem is schema, prose, content depth, or indexation.


FAQ

How long does the 5-platform test take?

A full baseline audit takes 2–3 hours the first time. Once you have your prompt set saved, monthly monitoring runs take about 30 minutes across all five platforms.

Do I need paid accounts for any of these platforms?

No. All five platforms have free tiers that are sufficient for testing. ChatGPT Plus gives fresher results and more citations, so it is worth testing on both free and Plus if you have access. Perplexity Pro tests different models, which is useful for a quarterly competitive sweep.

What if Claude refuses to answer all my prompts?

That is a useful signal, not a failure. It means your pages have no search footprint visible to Claude's retrieval mechanism. The fix is indexation and backlinks, not schema or prose. Focus on getting your pages cited in external sources first.

How do I track this over time without a paid tool?

A simple spreadsheet works. Columns: platform, prompt type, date, score (0/1/2), notes. Run the same prompts each month and track the score movement. You do not need a monitoring platform to establish a baseline and track improvement.

Should I test my competitors too?

Yes, especially on the category and comparison prompts. Seeing which competitors appear in Perplexity's sources panel for your category queries tells you exactly what content you need to create or improve.


References

[^1]: Only 8-12% overlap between ChatGPT citations and Google top-10 rankings — Maximus Labs AEO research

[^2]: AI-sourced traffic converts at 4.4x the rate of organic traffic — Okoone marketing research

[^3]: Only 30% of brands stay visible from one AI answer to the next; 20% remain visible across five consecutive runs — AirOps 2026 State of AI Search Report


New to AEO? Start with the foundational overview: What Is Answer Engine Optimization? A Practical Definition for Service Businesses

About the Author

Alex Rodriguez is an AI-first SEO operator based in Cedar Park, TX. 15+ years building content systems that drive AI visibility and organic growth.

About Alex →

Want This for Your Site?

I build content systems optimized for AI answer selection. Start with an audit.

Request an Audit