Guide

How Accurate Is Your AI Visibility Tool, Really? A Buyer's Guide

8 min readBy Katy BarnsSeptember 16, 2026

Key Takeaways

  • Most AI visibility scores use "union" math: if any one engine names you, you get credited across the board. That's not accuracy, that's rounding up.
  • Per-model scoring counts each engine's answer separately, so being named by 2 of 3 engines is 66%, not 100%.
  • Domain matching matters more than people think. Tools that don't canonicalize URLs will credit you for a competitor with a similar name.
  • Sample size and question realism decide whether a score means anything. Ten generic prompts won't tell you what real buyers see.
  • A tool that only measures and never fixes anything is half a product. You want diagnosis plus a weekly to-do list, not just a chart.

Why "accuracy" is the wrong first question

Everyone shopping for an AI visibility tool asks the same thing: how accurate is it? Reasonable question, wrong framing. Accuracy isn't a single number a vendor can hand you in a sales call. It's a function of three things: how the score is calculated, how questions are sourced, and how your brand gets matched against the answer text. Get any one of those wrong and the tool can be technically measuring AI visibility while still handing you a number that means nothing.

I've looked at enough of these dashboards to notice the pattern. The demo always shows a clean, high score. The methodology page, if there is one, is three sentences long. That gap is where the real question lives.

The union-scoring trick that inflates every dashboard

Here's the trick, and it's not even subtle once you know to look for it. Say you're tracked against three AI engines: ChatGPT, Gemini, and Perplexity. You ask 50 buyer-style questions. ChatGPT names you in 40 of them. Gemini names you in 10. Perplexity never mentions you.

A union-scored tool says: were you named by at least one engine in this question? If yes, count it as visible. Add those up and you get a score that looks like 40-something out of 50, roughly 80 percent. Sounds great. Except two of your three engines are barely mentioning you, and the number hides that completely.

Per-model scoring treats each engine as its own test. ChatGPT: 80 percent. Gemini: 20 percent. Perplexity: 0 percent. Average that honestly and you're sitting around 33 percent, not 80. Same underlying data, wildly different story, and only one of those stories is true.

This isn't a hypothetical. It's the default math most tools shipped with because it produces a bigger, more sellable number. We ripped that math out of our own product for exactly this reason. A number that flatters you is worse than useless. It tells you to relax when you shouldn't.

What per-model scoring actually looks like in practice

Per-model scoring means every engine gets graded on its own curve, then the scores get shown side by side instead of blended into one misleading average. If you're strong on ChatGPT and invisible on Gemini, you should see that as two separate bars, not one averaged-away number.

The value here isn't just honesty for its own sake, though that matters. It's that different engines pull from different sources and weight them differently. ChatGPT leans on a mix of web content and its training data. Perplexity leans hard on live citations. Gemini pulls from Google's index in ways that overlap with, but don't match, traditional SEO signals. If your score is blended, you can't tell which engine is failing you or why. If it's per-model, the fix is obvious: your Perplexity numbers are bad because you have no recent citeable sources, so go get some.

Domain matching is the other place tools cheat

Union scoring gets talked about because it's the more dramatic inflation. Domain matching is quieter but just as damaging. Say your business is Greenhouse Dental and there's an unrelated company somewhere called Greenhouse that gets cited constantly for something totally different. A tool doing sloppy string matching on brand names will credit your dashboard for every one of those mentions. Your score goes up. Your actual visibility hasn't moved an inch.

The fix is registered-domain and canonical-URL matching: the tool checks the actual link behind a citation, not just whether a name shows up in the text. It's more work to build. It's also the only version that tells you the truth. Before you buy anything, ask the vendor directly: do you match by domain, or by brand name in the text? If they hesitate, you have your answer.

Sample size and question quality decide whether a score means anything

A tool that runs your brand against 10 questions once a month is not measuring your AI visibility. It's taking a single snapshot and calling it a trend line. AI answers are not static. Ask the same question twice in the same week and you can get different results depending on model updates, current events, and plain randomness in generation.

What you want is a large enough panel of realistic, unbranded buyer questions, asked repeatedly, across engines, so the noise averages out and what's left is signal. Best HVAC company near me asked once tells you almost nothing. Asked 30 times across a quarter, against real competitor names, it starts to tell you something you can act on.

Question realism matters just as much as volume. A tool that only tests branded queries like is your company good is testing something buyers rarely type. You want the unbranded version: what does someone ask when they don't know you exist yet. That's the actual battlefield.

Red flags to check before you buy any AI visibility tool

A short list, because this decision doesn't need to be complicated.

Ask for the methodology in writing, not a slide. If a vendor can't explain in plain language how they score a mention, that's a red flag on its own.

Ask specifically whether scoring is per-model or union. If they don't know the difference, or dodge the question, move on.

Ask how domain matching works. We check the brand name is not the same as we check the canonical URL.

Ask what happens after the score. A dashboard that stops at diagnosis is half a product. You want to walk away with a list of things to actually go fix, ideally with someone willing to help fix them.

Ask about the price relative to what's included. Enterprise AI visibility tools can run into the thousands per month for measurement alone, no action included. That math only makes sense if you have a team ready to act on the findings. Most small businesses don't, which is a different problem entirely.

What we do differently at OMG

We built OMG around the assumption that most buyers of this category get burned by inflated numbers, so accuracy is the whole pitch, not a footnote. Per-model scoring is the default, not an upgrade. Every plan runs a panel of unbranded buyer questions against multiple engines and shows you the breakdown by model, not a blended average that hides your weak spots.

We also didn't want to build a tool that just tells you you're losing and leaves you there. OMG generates the actual fixes, weekly: blog posts, FAQ pages, schema and JSON-LD, llms.txt and robots.txt files. With the Do It For Me add-on, the technical files publish automatically and the content drafts queue up for your approval before they go live. You're not paying for a report. You're paying for the report plus the work that closes the gap it finds.

Pricing and how to get started

Plans start at $99 a month for a Starter tier with a 2-3 engine panel, scale to $249 a month for Pro with a wider panel and more frequent runs, and go up to $599 a month for Done For You service where we're actively managing the fixes, not just generating them. Enterprise runs $2,499 a month for larger multi-location or multi-brand accounts. The Do It For Me add-on is $399 a month on any tier.

Agencies get a white-label version with multi-client management built in, so you can run this under your own brand for every client on your roster.

If you want to see where you actually stand before committing to anything, run the free audit. It'll show you a real per-model breakdown, not a blended vanity number, so you can decide for yourself whether the gap is worth closing.

FAQs

What does per-model scoring mean in AI visibility tools?+

It means each AI engine, ChatGPT, Gemini, Perplexity, and so on, is scored separately instead of being blended into one number. Being named by 2 of 3 engines shows as roughly 66 percent, not 100 percent, because the tool isn't crediting you for the engine that ignored you.

Why do most AI visibility scores look inflated?+

Because most tools use union scoring: if any single engine mentions your brand, the question counts as a win across the board. That math produces bigger, more sellable numbers, but it hides which engines are actually ignoring your business.

How many questions should an AI visibility tool test to be reliable?+

More than a handful, run repeatedly over time. A one-time check against 10 questions is a snapshot, not a trend. AI answers shift with model updates and random variation, so a reliable score needs volume and repetition to average out the noise.

What is domain matching and why does it matter?+

It's checking the actual canonical URL behind a citation instead of just matching your brand name in the text. Without it, a tool can credit your score for mentions of an unrelated company that happens to share part of your name.

Does OMG just measure AI visibility, or does it also fix it?+

Both. OMG runs the per-model measurement and then generates the actual deliverables, blog posts, FAQ pages, schema markup, llms.txt and robots.txt files, needed to close the gap. The Do It For Me add-on auto-publishes the technical files and queues content drafts for approval.

Related reading

See how AI sees your business

Ready to see where you stand? Get started in minutes.