Fix · vendor accountability
Verify an agency's AI visibility audit before you pay the retainer
An agency sent you an 'AI visibility audit' with a score and no evidence trail. Before you pay for the retainer, replay their claims against the actual assistants. A score isn't evidence, and a deck isn't verifiable — the questions below apply to any AEO vendor, including us.
1. A score isn't evidence
An "AI visibility score" — especially one presented as a single number or letter grade — collapses a multi-dimensional question (which assistants, which buyer queries, what kind of citation, on what date) into an aggregate that can't be verified. If the report doesn't show you which specific queries were tested, against which specific assistant products, on which date, with the raw assistant responses, you can't replay the test yourself or check if it was representative of the queries your actual buyers ask.
A proper visibility audit should include: the exact buyer-phrased queries tested, which assistant lanes were probed (ChatGPT web search, Perplexity Sonar, Gemini grounded search, etc.), the date of the probe, and whether your domain was cited (with a link), mentioned (named but not linked), or absent — per query, per lane. Anything less than that is an opinion dressed up as data.
2. Replay the claims with live probes
The only way to verify an agency's audit is to re-run the same queries against the same assistants and see if you get the same answer. If their report says you're "invisible to ChatGPT" but doesn't list which queries they tested, you can't confirm or refute it. If it does list queries, paste them into ChatGPT web search, Perplexity, and Gemini yourself and check whether your domain appears in the citations. A legitimate finding should be reproducible, at least directionally, within a few days of the original test.
Expect some variation — assistant answers are generated fresh per query and vary with phrasing, model version, and retrieval-index state — but if an agency claims zero visibility and you find your domain cited on 3 out of 5 queries when you replay them, that's a red flag. Either the audit was stale, the queries were unrepresentative, or the finding was fabricated.
3. Our evidence standard (and yours)
We run the same AEO engine we sell on our own site, in public, with every probe dated and every change committed to a version-controlled ledger — see the Agent Run Ledger. The ledger shows the queries we test, the lanes we probe, the citation counts per run, and the interventions we ship, with no edits after the fact. If an intervention fails to move the needle across two consecutive weekly probes, it gets marked as a "no-change" failure on the same page. That's the accountability standard we hold ourselves to.
You should expect something similar from any AEO vendor: dated probe runs, listed queries, per-lane results, and a commitment to report failures, not just wins. If a vendor only shows you the citations that worked and never mentions the ones that didn't, you're seeing marketing, not science.
4. Questions to ask any AEO vendor (including us)
Before signing a retainer or paying for optimization work, ask:
- Which specific buyer queries will you test against, and how did you choose them? — Generic brand searches ("your company name") aren't useful; you want queries your actual buyers type when they don't yet know your brand exists.
- Which assistant lanes are you probing? — ChatGPT web search, Perplexity Sonar, Gemini grounded, Claude with search, Bing Copilot? If the answer is vague ("all of them"), press for specifics.
- Will you show me the raw assistant responses, or just a summary score? — A proper audit includes screenshots or text exports of the actual citations, not just "you scored 42/100."
- What does a citation mean in your audit? — Does "cited" mean your domain was linked as a source, or just mentioned in passing text? These are different outcomes, and conflating them inflates the score.
- How often will you re-probe, and will you report no-change results? — A one-time audit can't tell you if a change worked. Weekly or biweekly re-probes, with a commitment to report failures, are the minimum.
- What happens if the interventions don't move your visibility? — If the answer is "that never happens" or "we'll keep trying until it works," you're being sold certainty no one can deliver. A vendor with an honest calibration will tell you what they've tried that failed, and under what conditions they'd recommend stopping.
If any of these questions get deflected with "proprietary methodology" or "trust us, we've done this before," that's not rigor — it's opacity. Methodology can be proprietary; evidence can't be.
Get a dated, verifiable scan before you hire anyone
A free sampled visibility scan across three live assistant-search lanes on up to 5 buyer intents, plus deterministic checks on crawl access, llms.txt, and schema. You keep the raw data — no account, no card, and nothing withheld behind a "schedule a call" gate.
Run a free scanUse this as your baseline before any vendor (us included) starts work. If you pay for our $79 audit after, you get the same data structure — dated, per-query, per-lane — so you can compare runs yourself, not take our word for it.
Frequently asked
What if I already paid for an audit and it looks like this?
Run the queries they listed (if they listed any) against the same assistants yourself. If the findings don't reproduce, that's your leverage to ask for a refund or a re-do with proper evidence. If they didn't list queries, you can't verify it at all — which is itself the problem.
Doesn't every vendor have some proprietary scoring model?
A scoring model is fine if it's disclosed and the underlying data is verifiable. "We weight Perplexity 2× more than ChatGPT because our buyers use it more" — that's a methodology. "We give you a 42/100 and won't show you which queries we tested" — that's not.
Is AttractOS claiming to be the only honest vendor in this space?
No — we're claiming that any vendor (us included) should be able to answer the questions above with specifics, not deflection. If we ever can't, hold us to this same page.
Related
For what a verifiable, append-only audit ledger actually looks like: the Agent Run Ledger — our own site, tested weekly, with every intervention and failure logged. For the fundamental capability gaps any audit should check: Why doesn't ChatGPT recommend my business?
Last updated 2026-07-20. This is vendor-accountability guidance, not a probe report. It applies to any AEO vendor's claims, including our own.