How to Measure AI Search Visibility Without Buying an Expensive Tool

Measure AI search visibility with a frozen prompt set you run yourself, a three-state score, and the free platform reports. Record four things per run: were you named in the answer, were you linked in the sources, was what the answer said about you true, and did anyone arrive afterwards. Score each run green, amber, or red, and read presence rate on a four-week rolling window instead of week to week. Tie it to leads with GA4 annotations, the AI Assistant channel, and a self-reported source field on your booking form. This costs no software and about an hour a week, which is not the same thing as free.
Most agency owners land on this question after a client mentions that ChatGPT recommended somebody else. The instinct is to buy a dashboard. Buy the instrument last.
This is the loop we run on blog.magicteams.ai every week, so the mechanics below are what we actually do rather than a suggested workflow. What follows is the real price of the entry-level tools, the scoring rule, how to connect any of it to leads, and the point where paying for it starts making sense.
If the underlying idea is new, start with what answer engine optimization means for B2B.
What do entry-level AI visibility tools actually cost?
Less than most people assume. They also buy less than most people assume.
Profound’s Starter plan is $99 a month billed yearly for 50 prompts on ChatGPT only. Growth is $399 a month for 100 prompts across ChatGPT, Perplexity, and Google AI Overviews. Otterly’s Lite plan is $29 a month for 15 prompts across four engines, with Standard at $189 for 100.
Read the prompt counts rather than the prices. Fifteen prompts is a smaller set than you should be running by hand. A hundred prompts is more than most $1M to $10M agencies can act on in a quarter.
So at the low end you’re paying for something too small to be stable, and at the high end you’re paying for volume you have no capacity to respond to. The middle of that range is where a tool earns its money, and you can’t tell where you sit in it until you’ve run the loop once.
Is doing it yourself actually free?
No. It’s unpriced, which is a different thing, and pretending otherwise is how these projects die in week four.
Here’s the arithmetic on our loop. Twenty prompts across two engines is forty runs. Reading and logging each one takes well under a minute once you’re practiced, so call the whole thing an hour a week including the review conversation at the end.
That’s about 4.3 hours a month. Put your own internal cost on it. At $75 an hour you’re spending roughly $325 a month of somebody’s time, which is more than every entry plan on the market except Profound’s Growth tier.
So price isn’t what you’re saving by starting with a spreadsheet. What you get for that hour is calibration.
A vendor score compresses presence, position, sentiment, and share into one figure you can’t take apart. When it drops four points, you can’t tell whether one of your pages slipped, a competitor published something, or the model simply rolled differently that morning.
After four weeks of reading raw answers with your own eyes, you can tell, and the same vendor score becomes worth paying for.
What are you actually measuring?
Four separate things that get mashed into one word.
| What you record | The question it answers | Why it stands alone |
|---|---|---|
| Presence | Were we in the answer at all? | The base rate. Meaningless in one run, useful across dozens |
| Citation accuracy | Was what the answer said about us true? | Being described wrong is worse than being absent |
| Source position | Cited in the body, in the top links, or buried? | Separates supporting evidence from footnote |
| Downstream action | Did anyone arrive, read, or book? | The only column your P&L recognizes |
Presence without accuracy is a liability. Presence without action is a vanity number. Rolling all four into a single score is why so many AI visibility reports feel busy and change nothing.
Accuracy deserves more weight than it usually gets. Researchers at Columbia’s Tow Center tested eight AI search tools with 1,600 queries in March 2025, giving each one a direct excerpt and asking it to identify the article behind it.
That study looked at news attribution, not at how engines describe a vendor, so don’t read the number as your error rate. Read it as a reason to check what the answer says about you, not only whether your URL showed up.
How do you score a run without arguing about it?
Make the rule mechanical, or two people will score the same answer differently and the trend line becomes an opinion.
Log the run itself in the citation log covered in tracking AI search citations alongside Search Console. This is the scoring layer that sits on top of those columns, and it exists because a yes/no “cited” field throws away the most common state of all, which is being named without being linked.
| Score | Rule |
|---|---|
| Green | Brand named in the answer body and one of your URLs appears in the cited sources |
| Amber | One but not both. Named without a link, or linked without being named or described |
| Red | Neither. You don’t appear in the answer or in its sources |
| Flag | Any run, green ones included, where the answer states something false about you |
Two numbers come off that, and both are read as a four-week rolling figure rather than a weekly one:
- Presence rate = (green + amber) ÷ total runs
- Green rate = green ÷ total runs
A flag isn’t a score. It’s an interrupt. If an engine says you serve an industry you don’t, or quotes a price you don’t charge, that gets fixed this week regardless of how the rest of the sheet reads.
Two mechanics matter more than they sound. Run the prompts in a signed-out or temporary session, or personalization and chat memory will hand you a flattering answer assembled partly from your own browsing.
And freeze the prompt wording for a full quarter, because a prompt list you keep tuning measures your edits rather than the market. The selection method sits in building an AI search prompt set.
Magic Teams keeps this in a spreadsheet on purpose. The column that earns the hour is the one where somebody writes down the single change the row triggered, and that column only gets filled in when a human has just read the answer. Automating collection before you’ve run the loop by hand for a month tends to produce a tidy chart nobody argues with.
How much of this do the free reports cover?
The platform side is a monthly pull, not a weekly one, because it moves slowly and there’s less in it than the headlines suggest.
Google announced Search generative AI performance reports in Search Console on 3 June 2026, then rolled them out in stages, so check your Performance menu before assuming your property is missing data.
There’s exactly one metric in them, impressions, split by pages, countries, devices, and dates. No queries, no clicks. Bing Webmaster Tools gets closest to real citation counts for free, though Microsoft says its figures are a sample of overall citation activity.
Neither one tells you what an answer said about you. That gap is the entire reason the manual prompt set exists. The full source-by-source comparison, including what each one can’t show, is in the Search Console tracking piece.
Worth keeping in view while you read any of it: Google describes AI features as part of the existing Search experience, with those impressions already inside your overall Search Console traffic. You’re measuring a surface, not a separate ranking system.
How do you connect this to leads without calling it attribution?
Three free mechanisms, plus one question on your booking form. None of them is attribution and you shouldn’t present them as such.
Annotations. GA4 lets you drop dated notes on the timeline. You need Analyst access or above to create them, titles cap at 60 characters and descriptions at 150, and each property holds 1,000.
Write one every time you publish, rewrite, or fix a flagged claim. Six months later that timeline is the only honest record of what you actually did and when.
The AI Assistant channel. GA4 groups arrivals from ChatGPT, Gemini, Deepseek, Copilot, and Grok into an AI Assistant channel, and the same page states it excludes Google’s AI Overviews and AI Mode. Google’s list doesn’t name Perplexity, so sort your Referral rows by source before you treat this channel as a total. It’s a floor on assistant traffic.
Assisted paths. In GA4’s Advertising section, the key event attribution paths report shows which channels initiate, assist, and close. Assistant sessions often sit mid-path rather than last, so a last-click view will understate them.
Then the tie-breaker: a “how did you hear about us” field on your booking form, free text, no dropdown. It’s messy and self-reported, and it’s still the only place a buyer can tell you that ChatGPT named you and they searched your brand afterwards.
One figure worth seeing before you set expectations with anyone. Seer Interactive published a single-client case study covering just under 11,000 AI sessions from October 2024 to April 2025.
Small volume, high intent is the pattern worth testing on your own data. It isn’t a multiplier you can forecast with, and Seer says as much.
Google’s own position points the same direction without the drama. In August 2025 it said total organic click volume had been relatively stable year over year while average click quality had increased, where a quality click is one the user doesn’t quickly click back from.
Plainly, then: none of this establishes that an AI answer caused a booking. It’s directional. Say that out loud in any report that goes to a client or a board.
A worked example: reading one prompt over four weeks
Take prompt P-07, a vendor-compare prompt: “who can help a 20-person marketing agency stop the founder approving every client email”. Run it weekly on two engines for four weeks. That’s eight runs.
Say it comes back 2 green, 3 amber, 3 red. Presence rate is 5 ÷ 8, so 63%. Green rate is 2 ÷ 8, so 25%. Those figures are illustrative, not ours.
The read is specific. You’re being surfaced in the conversation in nearly two runs out of three, but carrying a link in only one out of four. That usually means the engines know the topic exists on your site and are quoting somebody else’s page for the specifics.
Now look at which amber rows they are. If two of the three cite a competitor’s comparison page, the gap is a page that doesn’t exist yet. If they cite your page but paraphrase it without a link, the gap is inside the page, and what makes a B2B page citation-worthy is the next thing to read.
One row, one diagnosis, one change. Eight numbers did that. No dashboard was required, and no dashboard would have produced the diagnosis on its own.
When is a paid tool worth buying?
When collection is your bottleneck rather than interpretation.
The third branch is the common one. Owners buy the platform hoping it will create the discipline, and the discipline is a recurring hour with a named owner. That’s the same shape as an AI daily brief: a fixed input, a fixed time, and a human decision at the end.
Frequently asked questions
How many prompts, and how often?
Twenty prompts weekly is enough for a single-market service business, split across problem-aware, solution-aware, vendor-compare, and objection stages. Fewer than fifteen and the presence rate swings too hard to read. More than thirty and nobody runs it past week three.
Will Search Console tell me which prompt cited me?
No. The generative AI report has no query dimension and no click data. That’s exactly the gap the prompt set fills.
Does a signed-in ChatGPT account change the answers?
Enough to matter. Memory and chat history bias results towards what you’ve already looked at, which is why the runs go in a signed-out or temporary session every time.
Can I automate the collection?
Yes, and eventually you should. Run it by hand for at least four weeks first. Reading the raw answers is what teaches you which fields are worth logging, and a pipeline built before that usually logs the wrong ones.
Can I put this in a client report?
Only with the caveat attached. Presence rate is a sampling result from your own runs, not a market share figure, and no free method here proves an AI answer caused a sale. If your work touches regulated advice, in law, finance, or health, have a qualified professional review the claims before an engine starts quoting them.
Does any of this predict rankings?
No. There’s no published ranking system for AI answers, and a tool that implies otherwise is selling confidence. You’re measuring a base rate over time, which is good for deciding what to fix and useless for forecasting.
Where this fits
AI search visibility is one measurement loop among several, and it earns its hour the same way the others do, by ending in a decision somebody writes down. If you’re assembling the wider picture of what your automation and content work returns, how to measure ROI on AI automation applies the same discipline to operations.
Run it yourself for a month before you spend anything. You’ll know your own presence rate, you’ll have a list of flagged inaccuracies to correct, and you’ll know whether a $99 plan would tell you something you don’t already have.
If at that point collection is the part eating the hour, that’s the piece worth automating. Installing loops like this one, with a named owner and a human decision at the end, is what a one-week Magic Teams AIOS intensive does. Bring your log and we’ll read it with you.