AI Visibility
The AI Visibility Partner Scorecard
21 questions to ask before you sign.
August 14, 2026 · Vikram Jayanand
Why this scorecard exists
Every SEO agency now sells AI visibility. Most of them sell monitoring. Monitoring tells you that you are absent from AI answers. It does not tell you why, and it does not fix it. A dashboard that scans ten engines daily is worth very little if the reason your brand never appears is that your CDN has been returning 403 to OAI-SearchBot since March.
This scorecard is built around a different order of operations. Access first, because a site that cannot be fetched cannot be cited. Then extractability, because retrieval works on passages rather than pages. Then corroboration, because models weight claims that independent sources repeat. Measurement comes fourth, not first. Commercial outcome comes last but functions as a gate: an engagement that cannot connect to pipeline is a reporting subscription, not a growth programme.
Score each question 0, 1 or 2. 0 for no capability, or a generic answer that would apply to any client. 1 for partial capability, a manual process, or capability without evidence. 2 for demonstrated capability with a worked example, tooling or client data. Maximum score is 42.
Section A · Access and Extractability
10 points · Gate section
If an agency cannot diagnose this layer, nothing else in the engagement can work. This is where most invisible brands are actually losing, and it is the section most agencies skip because it requires engineering literacy rather than content skills.
1. Do they audit crawler access at the agent level?
Ask them to name the agents they check and explain what each one does. A credible answer distinguishes training crawlers from live retrieval fetchers: GPTBot and Google-Extended govern training, while OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, Claude-SearchBot and Applebot-Extended govern whether you can be retrieved and cited right now. Blocking the first group has no effect on your presence in AI answers. Blocking the second removes you entirely.
2: They name specific agents, explain the training versus retrieval split, and show a robots.txt audit from a live client. 1: They check robots.txt but treat all AI bots as one category. 0: They talk about "AI crawlers" generically, or say robots.txt is a technical SEO issue outside their scope.
2. Do they check the edge, not just the origin?
Your robots.txt can be perfectly configured while Cloudflare, Akamai, Fastly or Vercel silently blocks the same requests. Bot management rules, AI scraping toggles and aggressive rate limiting are frequently enabled by default or switched on by an IT team with no marketing consultation. This is the single most common cause of total AI invisibility in businesses that otherwise rank well.
2: They test actual fetch responses from AI user agents and inspect the WAF or CDN configuration. 1: They mention CDN blocking as a possibility but do not test for it. 0: No awareness of the edge layer.
3. Do they verify what the fetcher actually receives?
Most AI retrieval fetchers do not execute JavaScript. They take the raw HTML response. If your site is client-side rendered, or if service detail sits inside tabs and accordions that hydrate after load, the fetcher sees an empty shell. Ranking well on Google proves nothing here, because Googlebot renders and these fetchers largely do not.
2: They fetch the raw HTML without JavaScript and show you the difference between what a browser sees and what a fetcher sees. 1: They know JavaScript rendering is a risk but assess it by inspecting the framework rather than testing. 0: They assume Google indexing implies AI retrievability.
4. Do they check fetch performance, redirect hygiene and Bing indexation?
AI fetchers time out faster than Googlebot and handle redirect chains poorly. Separately, ChatGPT's search layer leans heavily on Bing's index, so Bing indexation is a prerequisite that almost nobody checks. IndexNow submission is a cheap and underused lever.
2: They measure time to first byte for fetcher requests, audit redirect chains, and confirm Bing indexation status. 1: They cover general site speed but not fetcher-specific behaviour or Bing. 0: Neither is mentioned.
5. Who implements the fixes?
An access diagnosis is worthless as a PDF. Ask directly whether they ship the change, brief your developer with specifications, or hand over recommendations and hope.
2: They implement, or they produce developer-ready specifications and verify the deployment afterwards. 1: They produce recommendations and re-scan when you tell them it is done. 0: Recommendations only, with no verification loop.
Gate rule: below 6 out of 10 in Section A, stop scoring. The rest of the engagement rests on a layer they cannot see.
Section B · Content and Entity
8 points
Once the machine can read you, the question is whether what it reads is usable. This is where most "AI content optimisation" offers are just on-page SEO with a new label.
6. Do they build for passage-level retrieval?
Retrieval operates on chunks, not documents. Each section needs to be semantically self-contained, resolving its own pronouns and context so it survives being lifted out of the page. Answer-first construction, headings that function as retrieval anchors, one claim per passage. A beautifully written page that only makes sense read top to bottom will lose to a plainer page built in extractable blocks.
2: They can explain chunking, show a rewritten passage, and articulate why self-containment matters. 1: They talk about FAQ sections and clear headings without the underlying mechanic. 0: Standard readability and keyword advice.
7. Are your facts machine-readable and internally consistent?
Pricing, locations, credentials, service scope, team bios and comparison tables need to exist as HTML text, not as images, PDFs or interactive widgets. Separately, contradictory claims across your own site cause models to drop or hedge the claim entirely: a different founding year on the About page than in the press kit, a service listed in the nav that no longer appears anywhere else.
2: They audit fact placement and run a consistency check across the estate. 1: They cover schema and structured data but not extractability or consistency. 0: Neither addressed.
8. Do they establish you as a resolvable entity?
Before a model can recommend you, it has to resolve you as one entity rather than three ambiguous strings. That means Wikidata, Crunchbase, LinkedIn, trade registries and Google Business Profile, connected through sameAs, with consistent legal and trading names. Name collisions are common in Gulf B2B and are frequently the whole problem.
2: They audit entity resolution across public knowledge sources and have a plan to fix gaps. 1: They implement Organization schema and consider the job done. 0: No concept of entity establishment.
9. Is the prompt set built from real customer language?
A generic GEO keyword export is not a prompt set. The questions people put to an assistant are longer, more conversational, more comparative and more situational than search queries, and they differ by market. Ask to see how prompts are grouped by intent, journey stage and location, and ask where the language came from: sales call recordings, support tickets, win/loss interviews, review text, or a spreadsheet.
2: Prompts are derived from your customers' actual language, with a stated source, and built for your market rather than translated into it. 1: Prompts are researched but derived from keyword tools and competitor analysis. 0: A standard prompt list that would apply to any client in your category.
Section C · Corroboration
8 points
A claim that exists only on your own website is a claim. The same claim appearing across independent sources becomes a fact the model will repeat unprompted. This is the mechanic that off-site work is actually for, and describing it as backlink building gives away that an agency has not understood it.
10. Do they measure the gap between what you say and what others confirm?
Most audits measure where you are cited. The more useful diagnostic is which of your core claims are corroborated anywhere other than your own domain. Market leadership, client count, sector specialism, founding date, credentials: if these appear nowhere independent, the model has no basis to assert them on your behalf.
2: They produce a claim-by-claim corroboration map showing which assertions are independently supported and which are unsupported. 1: They map external citations without connecting them to specific claims. 0: Off-site work is described purely as coverage or links.
11. Is there a specific plan for independent third-party presence?
The sources that shape AI answers are largely outside your control: industry media, directories, review platforms, comparison pages, Reddit, YouTube, LinkedIn, association listings, research citations. The plan needs to be platform-appropriate and genuinely useful, not the same promotional paragraph distributed everywhere.
2: A prioritised, named source plan tied to the prompts where you are absent, with realistic assessment of which sources are influenceable. 1: A generic list of directory and PR activity. 0: "We will build authority."
12. Is there a correction path for false or outdated claims?
When a model states something wrong about you, someone has to trace it back to the source that caused it and get that source corrected. Ask what they do when this happens, and whether they have done it before.
2: A defined remediation process, with an example of a correction they have traced and resolved. 1: They would investigate case by case. 0: No process.
13. Will they tell you when the problem is not GEO?
Sometimes you are not recommended because your positioning is unclear, or because you are not genuinely a leading option in the category you are targeting. No amount of optimisation fixes a category relevance problem. An agency willing to say this in the pitch is worth more than one with a better dashboard.
2: They raised a positioning or category concern unprompted, or can describe a client they declined or redirected. 1: They acknowledge the limit when you raise it. 0: Every problem has a GEO solution.
Section D · Measurement and Method
8 points
Measurement matters. It just does not come first, and the quality bar is higher than most reporting suggests.
14. Can they tell retrieval from memory?
Some answers are grounded in live retrieval and are addressable in weeks. Others come from parametric memory built during training and move on a scale of quarters, if at all. An agency that cannot distinguish the two will promise timelines it has no mechanism to hit.
2: They diagnose which answers are retrieval-grounded, show citation evidence, and set different expectations for each. 1: They understand the distinction conceptually but do not diagnose per answer. 0: All AI answers treated as one thing.
15. Is the testing protocol clean and variance-aware?
Logged-in sessions carry memory and personalisation, which contaminates readings. One response is not a measurement. Ask for the protocol: fresh sessions, logged out, geography controlled, a stated number of runs per prompt, and a defined threshold before a result is called stable.
2: A written protocol with run counts, session hygiene and a stability threshold. 1: They run multiple checks but without a formalised method. 0: Screenshots of single responses.
16. How do they separate their work from model drift?
A provider ships a new model and every number moves at once. Without a control, any agency can claim credit for a version bump. Ask for a held-out cluster of prompts they deliberately do not work on, and a stated procedure for what happens when a major model update lands mid-engagement.
2: A control cluster plus a documented drift procedure. 1: They note model changes in reporting but run no control. 0: Improvements attributed to their work without qualification.
17. Do you get the underlying data, and do you keep it?
Prompts, raw responses, providers, citations, source URLs, competitor data, scan dates and history should be inspectable rather than summarised into a single percentage. Then ask the question buyers forget: who owns the prompt set and the response history if the relationship ends.
2: Full data access, exportable, with a clear statement that the data and prompt set are yours on exit. 1: Full data access within their platform, ownership on exit unclear. 0: Reports and screenshots only.
Section E · Commercial
8 points · Gate section
The typical agency checklist measures visibility a dozen different ways and never once measures money. This section is where most AI visibility offers quietly fail.
18. Did they size the opportunity before proposing the work?
In many Gulf B2B categories, AI referral is still low single-digit percentages of sessions. That may well justify moving early, but it is a different conversation from the one the pitch deck implies. An agency that skips this step is selling on hype.
2: They estimated current and projected AI-referred volume for your category and were candid about the size of the prize. 1: They cite general market growth statistics. 0: No sizing, urgency framing only.
19. How does AI visibility connect to pipeline?
Much AI-referred traffic arrives without a referrer and lands in direct. Ask how they identify it: referrer patterns where available, landing page and session signatures, self-reported attribution added to forms, CRM tagging so AI-sourced leads can be traced to revenue. Ask what conversion rate they expect, because AI-referred sessions typically convert at a higher rate on lower volume, which makes good traffic look bad against standard engagement benchmarks.
2: A concrete attribution plan spanning analytics, form capture and CRM, with a stated view on expected conversion behaviour. 1: They track referral traffic where the referrer is visible. 0: Success is defined entirely by visibility metrics.
20. What is the plan when visibility rises and traffic falls?
If the assistant answers the question completely, the user never clicks. Your mention rate can improve while sessions decline. This is a foreseeable outcome, not a failure, and it needs a stated response: which metrics replace clicks, and how the brand captures value from a zero-click mention.
2: They raised zero-click as a strategic consequence and proposed what to measure instead. 1: They report zero-click mentions as a metric without addressing the implication. 0: Not considered.
21. Do they guarantee outcomes, or process?
Nobody can guarantee that a model will recommend a specific company. Models change, sources change, responses vary by wording, date, location and provider. What can be guaranteed is method: research quality, implementation, monitoring cadence, evidence-based prioritisation and reporting. Treat "top position in ChatGPT" as a disqualifier.
2: Explicit process commitments, explicit refusal to guarantee placement. 1: Vague about what is being promised. 0: Outcome guarantees.
Gate rule: below 4 out of 8 in Section E, you are buying a reporting subscription rather than a growth programme, whatever the total says.
Scoring
Add the sections: Access and Extractability out of 10, Content and Entity out of 8, Corroboration out of 8, Measurement and Method out of 8, Commercial out of 8. Maximum total is 42.
- 35 to 42. Strong candidate. Verify the Section A evidence is from live client work rather than a demo.
- 27 to 34. Promising. Identify which section is dragging and ask for specific evidence there before committing.
- Below 27. High execution risk.
Gates override the total. Section A below 6, or Section E below 4, means the total is not meaningful. An agency scoring 30 with a 3 in Section A is very good at measuring a problem it cannot fix.
The question to ask
Do not ask an agency whether they can get you into ChatGPT. Ask this instead:
Can an AI assistant fetch and read our site today, and how do you know? Which of our claims are confirmed by anyone other than us? Which prompts matter to our buyers, where are we absent, and is that because we are not retrievable, not corroborated, or not genuinely relevant to the category? What will you change, who will implement it, and how will we see it in pipeline?
Every part of that should get a specific answer.
ViMi Digital works on AI visibility as an engineering problem before a content problem. If you want your current position established, we run an AI Visibility Health Check covering crawler access, extractability, entity resolution, corroboration and answer-level presence across the major assistants.
Book a call to scope your Health Check →About the author. Vikram Jayanand is the Co-Founder of ViMi Digital, where he works with B2B teams across Asia and the Gulf on AI visibility, signal engineering and demand generation.