August 15, 2026

How to Build an AI Search Prompt Set (The Prompt Ledger)

Prompt map tracing each buyer stage from problem-aware to vendor-compare to the source type an AI answer tends to cite for it
Photo: Magic Teams AI / generated in the build

Build an AI search prompt set from real buyer language rather than from keywords. Twenty to forty prompts is enough for most service businesses. Spread them across four intent stages, trace every prompt back to something a human actually said or typed, freeze the list for a quarter, then run each prompt several times a week and score presence rate across runs instead of position in any single answer. One run tells you nothing. Three hundred runs tell you where you stand.

Most advice on this stops at “track your brand prompts” and points you at a dashboard. That skips the only hard part, which is deciding which questions belong in the set and what counts as a result.

What follows is the ledger format, the sampling rule, the decision table, and a full worked set for one topic cluster on this blog that you can check against the published pages. It’s written for owners of agencies and professional practices who are doing this themselves rather than buying a visibility platform. If the underlying idea is new, read what answer engine optimization means for B2B first.

Why can’t you just read this out of Search Console?

Because the query data isn’t there.

Google launched Search generative AI performance reports in Search Console on 3 June 2026. They show impressions, pages, countries, devices, and dates for AI Overviews and AI Mode. Search Engine Journal’s read of them lists what’s absent: queries, clicks, click-through rate, average position, citation placement, and the passage used to support the answer.

Two more limits matter. The data starts on 18 May 2026 with no backfill, and the first rollout went to UK sites only, a detail Google’s own announcement left out.

So you can see that a page appeared inside an AI feature. You can’t see what anyone asked to make it appear. ChatGPT, Claude, and Perplexity report nothing to you at all.

A prompt set closes that gap. It’s a fixed, written instrument you point at the engines yourself, on a schedule, so “are we showing up” stops being a vibe someone formed after asking ChatGPT once on a Tuesday.

Why does one run of a prompt prove nothing?

Because the same prompt rarely returns the same answer twice.

Rand Fishkin at SparkToro and Patrick O’Donnell at Gumshoe.ai ran 2,961 prompts through ChatGPT, Claude, and Google’s AI Overviews in late 2025, testing 12 brand-recommendation prompts 60 to 100 times each per platform.

Here’s the part that makes prompt sets workable anyway. Across 994 responses about headphones, Bose, Sony, Sennheiser, and Apple still appeared in 55% to 77% of them. The list shuffles every time. The underlying leaders hold.

Some of that variance is built into the product. Google documents that AI Mode and AI Overviews may use a “query fan-out” technique, issuing multiple related searches across subtopics and data sources to build a response. Two people asking the same thing can be answered from different sub-searches.

That’s the design principle for the whole method. Individual answers are noise. Presence rate across many runs is signal. Any tool or agency reporting your “rank in ChatGPT” from a single pull is reporting a dice roll.

What actually goes in the set?

Four intent stages, with each prompt written the way a buyer would write it.

Semrush tracked 1,094 US categories in ChatGPT monthly from January to June 2026, using five representative prompts per category covering definition, comparison, alternatives, use case, and buying question. Their ownership finding is the useful bit. Only 15.2% of categories had a clear owner, defined as a brand appearing in at least four of the five prompts with at least a five percentage point lead over the runner-up. Another 31.2% had an emerging leader, and 53.7% were unsettled, with no brand appearing in even three of five.

Winning one prompt is not owning a topic. That’s why the set is a grid rather than a list.

Buyer stage What the prompt sounds like What the answer tends to cite What a win looks like for you
Problem-aware Long and situational. “I run a 14-person agency and I’m still approving every client email myself.” Explainer posts, definitions, forum and community threads You’re named in how the problem gets framed, before any vendor list appears
Solution-aware “What are the options for X” or “how do businesses usually handle X” Category explainers, comparison pages, methodology posts Your approach appears as one of the named options
Vendor-compare “Best X for Y”, “alternatives to Z”, “who should I hire for X” Listicles, review sites, Reddit threads, directories Your brand appears in the list at all
Objection and risk “Is X safe”, “is X legal in the EU”, “is X worth the money” Regulator pages, vendor policy docs, expert posts with sources Your page is cited as the evidence behind the answer

Notice that the expected source type changes by stage. That tells you which of your pages could plausibly get pulled in, and whether such a page exists yet.

Phrase the prompts the way people phrase them. Semrush analyzed more than a billion lines of US clickstream data from October 2024 to February 2026 and found the average search-enabled ChatGPT prompt grew from 4.7 to 8.7 words year over year, while 65% to 85% of prompts matched no keyword in a 27-billion-keyword database. Keyword-shaped language is rising too, from 18.9% of prompts in October 2025 to 34.9% by February 2026.

So keep both registers in the set. Short keyword-shaped prompts and long situational ones. They pull different sources.

Where do the prompts come from?

From artifacts, not brainstorms. This rule separates a useful set from a wish list.

That last rule is the one people break. If you add and drop prompts every month, a rising presence rate might only mean you swapped in easier questions.

On volume, SE Ranking suggests 20 to 40 prompts to start, weighted toward consideration-stage questions, run across two or three engines for at least 30 days. That’s vendor guidance with no published dataset behind it, so treat it as a starting point rather than a finding.

The arithmetic points the same way. Size the set from your run budget, not the other way round. Twenty-four prompts at five runs each across three engines is 360 pulls a week, which is a focused morning or a scripted job. Forty prompts run once each looks bigger and tells you less.

What does the weekly sample look like?

Same day, same conditions, logged the same way every time.

The fourth field is the one most teams skip and the one that pays. Which URLs got cited tells you what kind of evidence the engine trusts for that question.

If four of five answers cite a regulator page and a review site, and your blog post is the shape of neither, that’s your instruction for the next piece of work.

A worked example: the AI disclosure cluster

Here’s the grid filled in for one topic we publish on at Magic Teams AI, so you can see the derivation rather than a template.

The topic is whether and how businesses should disclose AI use in client email. Each prompt below is a question we chose to answer on this blog, written in the register a buyer would use, and each links to the page we’d want cited.

# Stage Prompt Page we’d want cited
1 Problem-aware “we’ve started using AI to draft client emails and I’m not sure if we have to tell anyone” Do AI email disclosure laws apply to my business
2 Solution-aware “how do agencies usually handle AI disclosure with clients” AI disclosure requirements by industry
3 Objection and risk “does GDPR require disclosing AI-written emails” GDPR and AI email disclosure rules
4 Objection and risk “does telling people an email was AI-written hurt reply rates” Does disclosing AI emails hurt response rates
5 Vendor-compare “who helps agencies set up reviewed AI email workflows” Service page

Five prompts, four stages, one topic. Prompt 5 is the only one where a vendor list is even likely, which is the point. The other four are won by having a checkable answer published, not by being on a list.

Personal insight

Before a prompt earns a row in our ledger, we open the page behind it and ask whether it looks like something a machine would quote: a direct answer near the top, a named source with a date, and a heading that matches the question. You can run that same check on the five pages above and judge our work for yourself. When a prompt has no page behind it, it goes on the writing list instead of the tracking sheet.

Then you log runs. The row below is illustrative, showing the format rather than published results.

W33 | P3 | ChatGPT | run 3/5 | appeared: no | competitors: 2 law firms, 1 SaaS blog | cited: ico.org.uk, 2 vendor blogs

Read that row and the work assigns itself. The engine wants a regulator source alongside a practitioner explanation. If your page doesn’t link the regulator, it isn’t the shape of thing that gets pulled in. That’s a page-level fix, and it’s the same fix an AEO content audit surfaces.

What do you do with the numbers?

Convert presence rate into one of four decisions.

Those thresholds are our working rule of thumb, not a published benchmark. Pick your own bands if your category is thinner or more crowded, then keep them fixed so the readings stay comparable.

Two cautions on reading these. A presence rate measures a moving system you don’t control, so treat month-over-month direction as the signal and any single figure as approximate.

And presence is not revenue. Google’s AI optimization guide says plainly that “optimizing for generative AI search is optimizing for the search experience, and thus still SEO.” Keep leads and pipeline as the outcome metric and use the prompt set as a diagnostic. The same discipline applies to automation spend, which we cover in how to measure ROI on AI automation.

What do people get wrong?

  • Tracking brand prompts and calling it visibility. Asking “what is Magic Teams AI” tells you what a model already knows about you. It says nothing about whether you reach someone who’s never heard of you.
  • Running each prompt once. With under a 1% chance of the same list twice, a single run is a coin flip with a spreadsheet around it.
  • Inventing prompts that sound like keywords. With 65% to 85% of ChatGPT prompts matching no keyword in a 27-billion-keyword database, a set built purely from your keyword tool is testing questions nobody asks.
  • Editing the set mid-quarter. Your trend line then measures your editing.
  • Chasing one prompt. Semrush’s data puts ownership at appearing in at least four of five prompts in a category, so a topic-level set beats a trophy prompt.
  • Expecting technical tricks to move it. Google states there are “no additional technical requirements” to appear in AI Overviews or AI Mode beyond being indexed and snippet-eligible, and no special file or schema is needed.

Frequently asked questions

How many prompts should I track?

Twenty to forty for a single-service business, weighted toward solution-aware and vendor-compare questions. Add a block per service line if you sell more than one thing, rather than stretching one set to cover everything. More prompts with fewer runs each is the wrong trade.

How often should I run the set?

Weekly for the full set, reported monthly. More frequent produces noise you’ll be tempted to react to. Less frequent and you’ll miss a competitor’s push until it’s a quarter old.

Do I need separate prompt sets for ChatGPT, Gemini, and Perplexity?

No. Keep one set of questions and run it against each engine, because the buyer’s question doesn’t change by platform. What changes is what gets cited, so log sources per engine.

Should I use a tracking tool or a spreadsheet?

A spreadsheet until the manual runs become the bottleneck, usually somewhere past 30 prompts and 3 engines. Build the set by hand first either way. A tool that picks your prompts for you will pick prompts that flatter the tool.

Can I set a target citation rate for my team?

Don’t. You don’t control the system generating the number, and a target on an uncontrolled metric turns into gaming. Set targets on inputs you do control, like pages published against unanswered prompts and sources added to existing pages.

One boundary worth stating plainly. A prompt set measures visibility, not accuracy or compliance. If an AI answer misstates something legal, financial, medical, or security-related about your business, that’s a correction and possibly a legal matter, and it belongs with a qualified professional rather than your content calendar.

Building the set takes an afternoon. Running it every week for a year, keeping the ledger honest when three people touch it, and turning the cited-URL column into published work is the operating problem underneath, and it’s the part that usually collapses back onto the founder. That’s the kind of reviewable, recurring workflow we install during an AIOS week at Magic Teams AI, and it’s why I keep this ledger for our own blog before recommending it to anyone. If that’s the bottleneck you’re staring at, a fit call is a reasonable next step.