How to track brand mentions in AI search: a 4-step playbook across five motors
Most guides on this topic push you toward paying for a tool before you have measurement discipline. This one walks the four-step methodology first — prompt library, per-motor firing, scoring, aggregation — with a free path that scales to roughly fifty prompts before automation matters. Includes the honest limits every vendor guide skips, most importantly that ChatGPT web_search does not fire on many tool and how-to queries at all.
To track brand mentions in AI search: build a prompt library of 50-150 buyer-intent queries, fire the library across ChatGPT, Perplexity, Google AI Overview, Gemini, and Claude weekly, score each response on three distinct mention types (citation, named-in-text, recommended), and aggregate results into share-of-model and citation-share metrics. The methodology can start free with manual firing on 10-15 prompts and scale to automated tools when volume justifies. Honest limit: ChatGPT web_search does not fire on many tool and how-to queries at all — some motors are structurally harder to measure than others.
I have commercial interests in AI content-generation categories broadly but do not sell any AI-search-visibility tool. Every claim in this piece derives from public methodology any reader can execute against the same endpoints, or from primary academic sources cited inline. The 5-motor citation data referenced in the honest-limits section was collected on 2026-08-19 and the raw JSON is attached in the methodology appendix.
Three distinct mention types you need to separate
Most guides on this topic conflate three different things. They are measured differently, they have different strategic value, and they require different firing methodology. Before you build a prompt library, get this taxonomy straight. Confusing them is the number-one reason teams end up with dashboards that look busy but do not tell them anything actionable.
The mistake most teams make is measuring only citations because they are the easiest surface to check. This under-counts your real visibility on Claude by a wide margin (Claude names brands in text without citation on a majority of responses in our testing) and misses the recommended-shortlist signal entirely on Perplexity. A serious measurement setup captures all three types per response.
The four-step measurement workflow
The rest of this piece drills into each step. This is the skeleton to keep in mind — every measurement problem in AI search reduces to failing one of these four steps.
- 01
Build a prompt library
50-150 queries mirroring real buyer intent. Composition matters more than absolute volume. Sourced from real Google People Also Ask data, sales transcripts, community discussion — not invented from vendor bias.
- 02
Fire the library across motors
Weekly minimum. Per-motor methodology varies materially. Perplexity via Sonar API for scale. ChatGPT and Google AI Overview manual for reliability. Claude and Gemini manual with text-scan for named-in-text detection.
- 03
Score each response
Per-response fields: mention type (citation, named, recommended), position, sentiment, context, motor, date. No composite score until you have four weeks of baseline data.
- 04
Aggregate over time
Weekly cadence. Track share-of-model per motor, citation share per motor, sentiment trend, position shift. Twenty-percent week-over-week drift on any metric triggers investigation.
Building the prompt library
This is where most teams under-invest. Prompt library quality is the ceiling on measurement quality — no amount of tooling or automation fixes a badly constructed library. Get this right first.
- 01
Volume: 50-150 prompts
Under 50, sample noise dominates and weekly variance obscures real signal. Over 150, refresh cadence and firing cost become unsustainable for a team without dedicated measurement headcount. Start at 50, expand toward 150 as you identify gaps.
- 02
Composition: four query classes
Split across category head queries (broadest reach), competitor comparison queries (highest strategic signal), use-case-specific queries (buyer-persona coverage), and long-tail buyer-intent queries (extraction target for AI motors). Aim for rough parity across the four classes, adjusted for your category's search-demand distribution.
- 03
Sourcing: real user language only
Pull from Google People Also Ask data, buyer-interview quotes, sales-call transcripts, subreddit and LinkedIn discussion. Do not invent prompts. Invented prompts encode vendor bias and produce measurements that reflect what you think users ask rather than what they actually ask.
- 04
Refresh cadence: quarterly
Every quarter, add new prompts as PAA and buyer-intent language shifts, and retire prompts that no longer reflect real queries. Flag every prompt-library change in reporting so historical trend comparisons stay honest.
- 05
Discipline: neutral phrasing, no leading questions
Write 'What are the best X tools?' not 'Why is Brand X the best Y tool?' Leading questions produce answers that flatter your brand and generate measurements that overstate your position. Neutrality is a measurement-integrity concern, not a stylistic one.
A prompt library built to these rules will produce measurement data that any independent analyst could reproduce against the same queries. Reproducibility is the difference between a monitoring dashboard and a serious measurement program.
Firing the library across five motors
Each of the five major motors requires different firing methodology. Treating them as interchangeable is one of the most common mistakes and produces measurement gaps that skew the whole aggregate view. Below is the honest per-motor methodology, ordered from most-automatable to most-manual.
Firing methodology as of 2026-08-24. Motor behavior shifts as models update; re-verify per-motor methodology quarterly.
For teams scaling beyond ~50 prompts across all five motors weekly, dedicated tools are worth evaluating. Otterly, Peec, Gauge, AthenaHQ, and Profound all automate multi-motor firing with different strengths and price bands. See the buyer's guide at /writing/complete-guide-ai-search-visibility-tools-2026 for the full comparison across pricing tiers, feature depth, and methodology transparency.
Scoring each response
Raw responses become measurable data through a scoring rubric applied uniformly across motors and weeks. The rubric below is the minimum viable set. Add derived metrics later; do not start with a proprietary composite score before you have baseline data to calibrate it against.
Do not invent a proprietary composite score before you have baseline data. Start with counts (mentions per week per motor), then add derived metrics (share-of-model, citation share, sentiment trend) once you have four or more weeks of history. Composite scores calibrated on thin data optimize for the wrong thing.
Aggregating over time
Turning per-response scores into a monitorable trend. This is where measurement becomes actionable — one-time snapshots are useful for orientation but do not tell you whether your work is moving the needle.
- 01
Weekly cadence at minimum
Daily is overkill for most teams and creates alert fatigue. Monthly misses model-update-driven shifts. Weekly is the honest floor. If your library is 150 prompts and firing takes hours, cut library size before you cut cadence.
- 02
Four primary metrics
Share of model (percentage of prompts where you appear, per motor). Citation share (percentage of your appearances that are linked citations versus named-in-text). Sentiment trend (four-week rolling average). Position shift (first-mentioned percentage vs later-mentioned percentage). Every dashboard should present all four side-by-side.
- 03
Drift-detection threshold: 20 percent
Any week-over-week change greater than 20 percent on any metric triggers investigation. Model updates, prompt library refreshes, and index re-crawls all cause real drift. The threshold separates noise from signal — smaller shifts are usually noise; larger shifts usually have a cause worth understanding.
- 04
Reporting rhythm
Monthly cross-motor summary for stakeholders (share-of-model + citation share aggregated). Quarterly deep-dive with methodology notes, any prompt-library changes flagged, and honest limits section restated. Do not report aggregate numbers without disclosing prompt-library changes — a library expansion can look like a visibility gain.
If you do not have budget yet (start-free workflow)
The full methodology above scales cleanly with automation. But you do not need automation to start. Below is a genuinely free path that produces meaningful signal — reduce prompt library, fire manually, spreadsheet the scoring, complement with two free tools. Scale up to paid monitoring only when the manual workflow becomes the bottleneck, not before.
- 01
Reduce library to 10-15 highest-priority queries
Category head, top three competitor comparisons, top three use-case queries. Fifteen prompts is the honest sizing for a workflow that takes thirty minutes weekly.
- 02
Fire manually across five motors weekly
Perplexity, ChatGPT, Google AI Overview, Claude, Gemini — all in incognito with target market's location. Roughly thirty minutes total per week for fifteen prompts across five motors, once you have the rhythm.
- 03
Score in a spreadsheet using the Step 3 rubric
One row per response. Six columns for the six fields. Weekly rollup on a second tab. Nothing fancier is required to produce genuinely actionable measurement data at this scale.
- 04
Complement with two free tools
HubSpot AEO Grader for a one-time snapshot signal on your domain. Ahrefs AI Visibility Checker for a crawler-access check on the AI-search crawlers (GPTBot, OAI-SearchBot, PerplexityBot, ClaudeBot, Google-Extended). Both require no signup or credit card. Together they answer whether you have a content problem, a crawler-access problem, or both.
- 05
Scale up only when manual becomes the bottleneck
The free path scales cleanly to about fifty prompts across five motors weekly before the workflow overhead becomes unsustainable for most teams. When you cross that threshold — not before — evaluate paid monitoring tools with a specific bottleneck to solve, not as a general capability upgrade.
This path does not measure everything a $500-per-month tool would measure. It measures what matters. That is the difference between a monitoring program that produces decisions and a dashboard that produces slides.
The honest limits of AI-search measurement
Sections most competitor guides skip. Reader trust depends on naming these limits, because measurement work that pretends to be more precise than it can be produces confident-sounding reports that lead to wrong decisions.
- 01
The ChatGPT web_search paradox
On tool and how-to queries, ChatGPT web_search does not fire on six to ten of ten samples across our workstreams. When web_search does not fire, there is no citation panel and no directly measurable surface — the response is generated from training data alone. Named-in-text mentions still happen (ChatGPT can name your brand from training data) but cannot be tracked at scale without qualitative review of every response. For a full breakdown of this pattern and what to do about it, see /writing/how-to-get-cited-by-chatgpt — the companion citation-side playbook.
- 02
Claude's named-without-cited pattern
Claude frequently names brands in the answer text with no citation panel entry. Measurement requires scanning answer text on every response, not just checking citation panels. Teams that measure only citation panels systematically under-count Claude visibility by a wide margin.
- 03
Retrieval without attribution
All five motors retrieve more sources than they cite. Your content may be actively contributing to the answer being generated without appearing in any user-visible surface. Retrieved-but-not-cited is unmeasurable from outside the model. Do not build strategy on the assumption that citation share equals actual retrieval share — the two are meaningfully different.
- 04
Personalization drift
Authenticated users see different answers than logged-out users, and users on ChatGPT Plus see different model versions than users on Free. Almost no measurement tool captures the authenticated surface. The methodology in this piece measures the logged-out surface only, and this limitation should be flagged in every report to stakeholders.
- 05
Model-update discontinuities
When a motor updates its underlying model, citation patterns can shift by twenty to fifty percent in a single week. Distinguishing model-update shifts from genuine visibility change requires historical baseline data plus awareness of announced model releases. Perplexity Sonar model updates, OpenAI GPT version releases, Google AI Overview refreshes, and Anthropic Claude releases all cause measurable drift when they land.
- 06
Referrer traffic is not visibility
AI-motor referrer traffic to your site is a downstream signal, not a measurement of citation share. A page can earn high Perplexity citation share and generate zero referral clicks because the answer is complete without a click-through. The absence of referrer traffic does not mean the absence of AI-search visibility.
Six common mistakes teams make
Even with the tools working, these six mistakes cost measurement quality. Each is a common pattern; all six can be avoided without additional budget.
- 01
Vanity metric conflation
Mention count is not citation. Citation is not referral traffic. Referral traffic is not revenue. Treating any of these as a proxy for another produces reports that celebrate the wrong thing. Track each as its own signal and resist the urge to combine them into a single composite before you have baseline calibration.
- 02
Single-motor bias
Measuring only Perplexity because it is the easiest to measure (via API, with a clear citation panel) produces an incomplete picture. Perplexity behavior does not predict ChatGPT behavior does not predict AIO behavior. If you measure one motor, you are measuring one channel of a five-channel discipline.
- 03
Snapshot-only measurement
One-time firings produce orientation, not measurement. Weekly cadence is the honest floor. Teams that fire quarterly cannot distinguish real trend from single-firing noise.
- 04
Scraping-vs-API assumption
Tools using scraped AI-motor responses versus API responses produce meaningfully different measurements — Surfer's 1,000-prompt study found only 24 percent brand overlap between the two on the same prompts. Neither is 'correct.' Ask your tool vendor which they use before comparing their numbers against another tool's.
- 05
Prompt library staleness
Queries retire and buyer-intent language shifts, but many teams never refresh their prompt library after initial setup. Quarterly refresh is the floor. Report any library changes explicitly so historical comparisons stay honest.
- 06
Ignoring the retrieval-without-attribution ceiling
Assuming citation share equals actual visibility misses the retrieval-without-attribution layer. This ceiling means measurable citation share is always a lower bound on real visibility, never a full accounting. Report accordingly.
What this piece does not cover
Deliberately narrow scope. Adjacent topics that matter for a complete AI-search visibility program but are covered elsewhere on this site or deferred to future pieces.
- 01
How to actually get cited in the first place
The causal complement to measurement is optimization. If measurement shows you are invisible, the fix is on the citation-earning side. See /writing/how-to-get-cited-by-chatgpt for the full playbook including the ChatGPT paradox and the 79-percent independent-voice Perplexity citation pattern.
- 02
Full tool selection
The tools that automate the methodology in this piece are covered in depth in the buyer's guide at /writing/complete-guide-ai-search-visibility-tools-2026 — sixty-plus tools across five categories with dated pricing and honest verdicts.
- 03
Motor-specific playbooks
For deeper per-motor tactics on Perplexity and ChatGPT specifically, see /writing/perplexity-seo and /writing/chatgpt-seo. This piece covers the cross-motor measurement discipline; the motor-specific pieces cover the tactics within each channel.
- 04
Revenue attribution from AI-search referrals
Attributing revenue to specific AI-search referrals is a harder problem than measuring mention share. It requires session-level tracking of referrer sources ChatGPT and Perplexity do not always populate, plus multi-touch attribution modeling. This is a possible follow-up piece; the honest current answer is that referrer-attribution methodology is not yet mature across the five motors.
- 05
Sentiment analysis at scale
The scoring rubric in this piece captures sentiment as a per-response field. Automating sentiment classification at scale requires NLP tooling that is out of scope here. Manual sentiment tagging on a fifteen-prompt library takes minutes per week; automation matters only at library sizes above one hundred.
Frequently asked questions
How to track brand mentions in ChatGPT specifically?
Fire your prompt library manually at chatgpt.com in incognito with your target market's location. Check the citation panel when web_search fires and scan the answer text for named-in-text mentions when it does not. Be aware that ChatGPT web_search does not fire on many tool and how-to queries at all — on six to ten of ten samples in our testing — meaning ChatGPT is structurally the hardest of the five motors to measure at scale.
Is it possible to track brand mentions in AI search?
Yes, with three important caveats. Perplexity, Google AI Overview, and Claude produce measurable citation panels or named-in-text mentions on most queries. ChatGPT is measurable only when web_search fires, which is inconsistent on tool queries. Retrieval-without-attribution — cases where your content contributes to an answer without appearing in any user-visible surface — is not measurable from outside the model. Plan around these ceilings.
How much does AI brand mention tracking cost?
The free path in this piece scales to about fifty prompts across five motors weekly and takes roughly thirty to sixty minutes per week. Paid monitoring tools start around twenty-nine dollars per month (Otterly) for entry-tier single-user monitoring and scale to two thousand plus per month for enterprise (Profound). Team-scale mid-market monitoring typically lands at two-hundred-and-ninety-five to five-hundred per month for full-loop platforms.
What is a good AI visibility score?
There is no absolute benchmark that applies across categories. Interpretation requires comparison against category competitors — your share-of-model versus theirs on the same prompt library. As a rough sizing anchor: on a fifty-prompt library, appearing on ten to twenty percent of prompts across all five motors is common for an established brand, forty percent plus is category-leading, and under five percent is typical for smaller or newer brands. Category variance is wide; use competitor comparison rather than absolute thresholds.
How often should I measure AI brand mentions?
Weekly is the honest floor. Daily is overkill for most teams and creates alert fatigue. Monthly misses model-update-driven shifts that can move citation share by twenty to fifty percent in a single week. If your library is large enough that weekly firing is a burden, cut library size before you cut cadence.
Can I track brand mentions in Perplexity for free?
Yes. Fire manually at perplexity.ai in incognito with your target market's location. Perplexity's citation panel is directly visible and copyable. For scale beyond about fifty prompts weekly, the Perplexity Sonar API costs roughly five-thousandths of a dollar per query, which is genuinely negligible for measurement-scale volumes but requires basic API integration.
How to track AI search performance?
Track four metrics weekly per motor: share of model (percentage of prompts where you appear), citation share (percentage of your appearances that are linked citations versus named-in-text), sentiment trend (four-week rolling average), and position shift (first-mentioned percentage versus later-mentioned). Report all four side-by-side and set a twenty-percent week-over-week drift threshold to distinguish signal from noise.
What is the difference between AI ranking and AI citation?
Citation means appearing as a linked source in the motor's citation panel. Ranking means being named or recommended in the answer text with prominence — first mentioned, quoted directly, or recommended by name. A brand can be ranked without being cited (Claude and ChatGPT frequently name brands without citation panels) or cited without being prominently ranked (Perplexity often cites fifteen to twenty sources per answer with only three or four named in the response text).
Do I need a tool to track AI brand mentions?
No, for libraries under about fifty prompts weekly. The free path in this piece — manual firing across five motors, spreadsheet scoring, plus HubSpot AEO Grader and Ahrefs AI Visibility Checker for complementary signals — produces measurement data of the same quality as paid tools at this scale. Paid tools become worth their cost when library size or motor coverage crosses the manual workflow bottleneck.
How do I check my AI citations?
Fire the head queries most relevant to your brand across Perplexity, Google AI Overview, ChatGPT, Claude, and Gemini in incognito. For each response, check the citation panel for your domain and scan the answer text for named-mention of your brand. This is the same methodology paid tools automate; running it manually on ten to fifteen high-priority queries produces a directional signal in under an hour.
- Aggarwal et al. — GEO: Generative Engine Optimization (KDD 2024)arxiv.org →
- Chen et al. — Earned-media bias in AI search (arXiv 2509.08919, 2025)arxiv.org →
- Kevin Indig — ChatGPT citations: 44% from first third of content (1.2M responses)almcorp.com →
- Surfer — Scraped AI answers vs API results: 1,000-prompt study (Dec 2025)surferseo.com →
- Vismore — Best Ways to Track Brand Mentions in AI Search (750 Response Study)vismore.ai →
- LLM Pulse — How to Track Brand Mentions in AI Searchllmpulse.ai →
- Perplexity Sonar API documentationperplexity.ai →
- DataForSEO — Perplexity Sonar, Google organic advanced, ChatGPT endpoints (methodology reference)dataforseo.com →
Methodology appendix
The methodology in this piece is derived from a combination of primary academic sources cited inline (Aggarwal Princeton KDD 2024, Chen et al. arXiv 2509.08919, Kevin Indig's 1.2M ChatGPT response study, Surfer's 1,000-prompt scraping-vs-API study) and our own 5-motor citation baseline fired on 2026-08-19.
The ChatGPT web_search paradox referenced in the honest-limits section is drawn from six workstreams testing ChatGPT web_search behavior across tool and how-to query classes. Across those six workstreams, web_search fired zero times on the vast majority of samples — a pattern documented in the companion citation-earning piece with the raw data.
Search behavior and motor behavior shift over time. The methodology in this piece is dated August 2026. Re-verify per-motor firing methodology quarterly, especially for motors with fast-moving APIs (Perplexity Sonar, DataForSEO endpoints). If ChatGPT web_search behavior changes materially — for example if web_search begins firing reliably on tool queries — the honest-limits section here will need updating.
- 5-motor citation baseline
- 10 head queries × 5 motors (Perplexity Sonar, ChatGPT, Google AIO, Gemini, Claude), full annotation capture, fired 2026-08-19
- Head SERP audit for this piece
- Top-6 audit for `how to track brand mentions in ai search` fired 2026-08-24 via WebSearch
- Cross-workstream ChatGPT web_search pattern
- 0/N web_search fires across 6 prior workstreams on tool and how-to query classes
- Reproducibility
- The four-step workflow in this piece is deterministic — same prompt library against same motors on the same date will produce reproducible measurement data. Motor behavior shifts week to week and month to month; the methodology is reproducible, the specific measurements are dated.