How to track brand mentions in AI search: a 4-step playbook across five motors

Most guides on this topic push you toward paying for a tool before you have measurement discipline. This one walks the four-step methodology first — prompt library, per-motor firing, scoring, aggregation — with a free path that scales to roughly fifty prompts before automation matters. Includes the honest limits every vendor guide skips, most importantly that ChatGPT web_search does not fire on many tool and how-to queries at all.

To track brand mentions in AI search: build a prompt library of 50-150 buyer-intent queries, fire the library across ChatGPT, Perplexity, Google AI Overview, Gemini, and Claude weekly, score each response on three distinct mention types (citation, named-in-text, recommended), and aggregate results into share-of-model and citation-share metrics. The methodology can start free with manual firing on 10-15 prompts and scale to automated tools when volume justifies. Honest limit: ChatGPT web_search does not fire on many tool and how-to queries at all — some motors are structurally harder to measure than others.

I have commercial interests in AI content-generation categories broadly but do not sell any AI-search-visibility tool. Every claim in this piece derives from public methodology any reader can execute against the same endpoints, or from primary academic sources cited inline. The 5-motor citation data referenced in the honest-limits section was collected on 2026-08-19 and the raw JSON is attached in the methodology appendix.

Three distinct mention types you need to separate

Most guides on this topic conflate three different things. They are measured differently, they have different strategic value, and they require different firing methodology. Before you build a prompt library, get this taxonomy straight. Confusing them is the number-one reason teams end up with dashboards that look busy but do not tell them anything actionable.

Type
Definition
Example
Measurable?
Citation
Linked source in a citation panel. Numbered footnote in Perplexity. References block in Google AI Overview. Source panel in ChatGPT or Claude when web_search fires. Directly measurable — check the panel.
Perplexity cites your domain as [3] in the response, with the URL visible in the source panel.
Yes, per motor, at scale
Named-in-text
Brand name appears in the answer paragraph without a linked citation. Claude does this frequently. ChatGPT does this in training-data-only responses when web_search does not fire. Measurable only by scanning answer text.
Claude names your brand inside the answer paragraph but provides no citation panel entry.
Yes, but requires text-scan per response
Recommended
Brand appears in a shortlist or comparison — 'options include Brand A, Brand B, Brand C' — a distinct form that combines named-in-text with implicit endorsement. High strategic value because it signals the motor treats you as a category peer.
Perplexity answers 'best X tools include Brand A, Brand B, Brand C' with your brand in the shortlist.
Yes, requires structured extraction from answer text

The mistake most teams make is measuring only citations because they are the easiest surface to check. This under-counts your real visibility on Claude by a wide margin (Claude names brands in text without citation on a majority of responses in our testing) and misses the recommended-shortlist signal entirely on Perplexity. A serious measurement setup captures all three types per response.

The four-step measurement workflow

The rest of this piece drills into each step. This is the skeleton to keep in mind — every measurement problem in AI search reduces to failing one of these four steps.

  1. 01

    Build a prompt library

    50-150 queries mirroring real buyer intent. Composition matters more than absolute volume. Sourced from real Google People Also Ask data, sales transcripts, community discussion — not invented from vendor bias.

  2. 02

    Fire the library across motors

    Weekly minimum. Per-motor methodology varies materially. Perplexity via Sonar API for scale. ChatGPT and Google AI Overview manual for reliability. Claude and Gemini manual with text-scan for named-in-text detection.

  3. 03

    Score each response

    Per-response fields: mention type (citation, named, recommended), position, sentiment, context, motor, date. No composite score until you have four weeks of baseline data.

  4. 04

    Aggregate over time

    Weekly cadence. Track share-of-model per motor, citation share per motor, sentiment trend, position shift. Twenty-percent week-over-week drift on any metric triggers investigation.

Building the prompt library

This is where most teams under-invest. Prompt library quality is the ceiling on measurement quality — no amount of tooling or automation fixes a badly constructed library. Get this right first.

  1. 01

    Volume: 50-150 prompts

    Under 50, sample noise dominates and weekly variance obscures real signal. Over 150, refresh cadence and firing cost become unsustainable for a team without dedicated measurement headcount. Start at 50, expand toward 150 as you identify gaps.

  2. 02

    Composition: four query classes

    Split across category head queries (broadest reach), competitor comparison queries (highest strategic signal), use-case-specific queries (buyer-persona coverage), and long-tail buyer-intent queries (extraction target for AI motors). Aim for rough parity across the four classes, adjusted for your category's search-demand distribution.

  3. 03

    Sourcing: real user language only

    Pull from Google People Also Ask data, buyer-interview quotes, sales-call transcripts, subreddit and LinkedIn discussion. Do not invent prompts. Invented prompts encode vendor bias and produce measurements that reflect what you think users ask rather than what they actually ask.

  4. 04

    Refresh cadence: quarterly

    Every quarter, add new prompts as PAA and buyer-intent language shifts, and retire prompts that no longer reflect real queries. Flag every prompt-library change in reporting so historical trend comparisons stay honest.

  5. 05

    Discipline: neutral phrasing, no leading questions

    Write 'What are the best X tools?' not 'Why is Brand X the best Y tool?' Leading questions produce answers that flatter your brand and generate measurements that overstate your position. Neutrality is a measurement-integrity concern, not a stylistic one.

A prompt library built to these rules will produce measurement data that any independent analyst could reproduce against the same queries. Reproducibility is the difference between a monitoring dashboard and a serious measurement program.

Firing the library across five motors

Each of the five major motors requires different firing methodology. Treating them as interchangeable is one of the most common mistakes and produces measurement gaps that skew the whole aggregate view. Below is the honest per-motor methodology, ordered from most-automatable to most-manual.

Motor
Method
Cadence
Cost
Captures
Perplexity
Sonar API for scale, or manually at perplexity.ai for smaller libraries.
Weekly minimum
~$0.005 per query via Sonar API
Full citation panel URLs, answer text, source rankings. Best automation target of the five.
Google AI Overview
Manual at google.com incognito with target market's location. DataForSEO Google organic advanced endpoint captures AIO blocks at scale.
Weekly minimum
$0 manual, ~$0.005 via DFS
AIO reference panel URLs, AIO paragraph text (which names tools by name — capture separately from citation panel).
Claude
Manual at claude.ai. Scan both citation panel (when web_search fires) and answer text for named-mention.
Weekly minimum
$0 manual
Citation panel URLs when populated, plus full answer text for named-in-text detection.
Gemini
Manual at gemini.google.com. YouTube dominance skew: for tool queries, Gemini's citation surface is heavily YouTube-weighted.
Weekly minimum
$0 manual
Citation panel URLs (many will be YouTube), answer text for named-in-text.
ChatGPT
Manual at chatgpt.com incognito. DataForSEO ChatGPT endpoint exists but returned inconsistent web_search fires in our testing (0 fires on 8/10 tool queries).
Weekly minimum, with acknowledgment that many queries produce no measurable surface
$0 manual, ~$0.005 via DFS (unreliable)
Citation panel URLs when web_search fires. Text-only answers when it does not — capture answer text for named-in-text analysis but flag that these are training-data-only responses.

Firing methodology as of 2026-08-24. Motor behavior shifts as models update; re-verify per-motor methodology quarterly.

For teams scaling beyond ~50 prompts across all five motors weekly, dedicated tools are worth evaluating. Otterly, Peec, Gauge, AthenaHQ, and Profound all automate multi-motor firing with different strengths and price bands. See the buyer's guide at /writing/complete-guide-ai-search-visibility-tools-2026 for the full comparison across pricing tiers, feature depth, and methodology transparency.

Scoring each response

Raw responses become measurable data through a scoring rubric applied uniformly across motors and weeks. The rubric below is the minimum viable set. Add derived metrics later; do not start with a proprietary composite score before you have baseline data to calibrate it against.

Field
Values
Purpose
Mention type
citation / named / recommended / none
Distinguishes the three measurement surfaces per the definitions section. Determines which strategic play the mention represents.
Position
1st mentioned / 2nd / 3rd / later / not mentioned
Position bias is real in AI-generated answers. Being named third in a five-tool shortlist is materially different from being named first.
Sentiment
positive / neutral / negative
Motor answers can name your brand critically. Sentiment tracks whether mentions help or hurt.
Context
free text — one sentence
What the answer said about the brand. Comparison, recommendation, criticism, feature highlight. Enables qualitative pattern detection at aggregation time.
Motor
perplexity / chatgpt / aio / claude / gemini
Aggregate share-of-model per motor. Motor variance is real and material — one motor's citation graph does not predict another's.
Date
YYYY-MM-DD
Single-day snapshot. Model updates and index refreshes cause real shifts; date-stamping every capture enables drift detection.

Do not invent a proprietary composite score before you have baseline data. Start with counts (mentions per week per motor), then add derived metrics (share-of-model, citation share, sentiment trend) once you have four or more weeks of history. Composite scores calibrated on thin data optimize for the wrong thing.

Aggregating over time

Turning per-response scores into a monitorable trend. This is where measurement becomes actionable — one-time snapshots are useful for orientation but do not tell you whether your work is moving the needle.

  1. 01

    Weekly cadence at minimum

    Daily is overkill for most teams and creates alert fatigue. Monthly misses model-update-driven shifts. Weekly is the honest floor. If your library is 150 prompts and firing takes hours, cut library size before you cut cadence.

  2. 02

    Four primary metrics

    Share of model (percentage of prompts where you appear, per motor). Citation share (percentage of your appearances that are linked citations versus named-in-text). Sentiment trend (four-week rolling average). Position shift (first-mentioned percentage vs later-mentioned percentage). Every dashboard should present all four side-by-side.

  3. 03

    Drift-detection threshold: 20 percent

    Any week-over-week change greater than 20 percent on any metric triggers investigation. Model updates, prompt library refreshes, and index re-crawls all cause real drift. The threshold separates noise from signal — smaller shifts are usually noise; larger shifts usually have a cause worth understanding.

  4. 04

    Reporting rhythm

    Monthly cross-motor summary for stakeholders (share-of-model + citation share aggregated). Quarterly deep-dive with methodology notes, any prompt-library changes flagged, and honest limits section restated. Do not report aggregate numbers without disclosing prompt-library changes — a library expansion can look like a visibility gain.

If you do not have budget yet (start-free workflow)

The full methodology above scales cleanly with automation. But you do not need automation to start. Below is a genuinely free path that produces meaningful signal — reduce prompt library, fire manually, spreadsheet the scoring, complement with two free tools. Scale up to paid monitoring only when the manual workflow becomes the bottleneck, not before.

  1. 01

    Reduce library to 10-15 highest-priority queries

    Category head, top three competitor comparisons, top three use-case queries. Fifteen prompts is the honest sizing for a workflow that takes thirty minutes weekly.

  2. 02

    Fire manually across five motors weekly

    Perplexity, ChatGPT, Google AI Overview, Claude, Gemini — all in incognito with target market's location. Roughly thirty minutes total per week for fifteen prompts across five motors, once you have the rhythm.

  3. 03

    Score in a spreadsheet using the Step 3 rubric

    One row per response. Six columns for the six fields. Weekly rollup on a second tab. Nothing fancier is required to produce genuinely actionable measurement data at this scale.

  4. 04

    Complement with two free tools

    HubSpot AEO Grader for a one-time snapshot signal on your domain. Ahrefs AI Visibility Checker for a crawler-access check on the AI-search crawlers (GPTBot, OAI-SearchBot, PerplexityBot, ClaudeBot, Google-Extended). Both require no signup or credit card. Together they answer whether you have a content problem, a crawler-access problem, or both.

  5. 05

    Scale up only when manual becomes the bottleneck

    The free path scales cleanly to about fifty prompts across five motors weekly before the workflow overhead becomes unsustainable for most teams. When you cross that threshold — not before — evaluate paid monitoring tools with a specific bottleneck to solve, not as a general capability upgrade.

This path does not measure everything a $500-per-month tool would measure. It measures what matters. That is the difference between a monitoring program that produces decisions and a dashboard that produces slides.

The honest limits of AI-search measurement

Sections most competitor guides skip. Reader trust depends on naming these limits, because measurement work that pretends to be more precise than it can be produces confident-sounding reports that lead to wrong decisions.

  1. 01

    The ChatGPT web_search paradox

    On tool and how-to queries, ChatGPT web_search does not fire on six to ten of ten samples across our workstreams. When web_search does not fire, there is no citation panel and no directly measurable surface — the response is generated from training data alone. Named-in-text mentions still happen (ChatGPT can name your brand from training data) but cannot be tracked at scale without qualitative review of every response. For a full breakdown of this pattern and what to do about it, see /writing/how-to-get-cited-by-chatgpt — the companion citation-side playbook.

  2. 02

    Claude's named-without-cited pattern

    Claude frequently names brands in the answer text with no citation panel entry. Measurement requires scanning answer text on every response, not just checking citation panels. Teams that measure only citation panels systematically under-count Claude visibility by a wide margin.

  3. 03

    Retrieval without attribution

    All five motors retrieve more sources than they cite. Your content may be actively contributing to the answer being generated without appearing in any user-visible surface. Retrieved-but-not-cited is unmeasurable from outside the model. Do not build strategy on the assumption that citation share equals actual retrieval share — the two are meaningfully different.

  4. 04

    Personalization drift

    Authenticated users see different answers than logged-out users, and users on ChatGPT Plus see different model versions than users on Free. Almost no measurement tool captures the authenticated surface. The methodology in this piece measures the logged-out surface only, and this limitation should be flagged in every report to stakeholders.

  5. 05

    Model-update discontinuities

    When a motor updates its underlying model, citation patterns can shift by twenty to fifty percent in a single week. Distinguishing model-update shifts from genuine visibility change requires historical baseline data plus awareness of announced model releases. Perplexity Sonar model updates, OpenAI GPT version releases, Google AI Overview refreshes, and Anthropic Claude releases all cause measurable drift when they land.

  6. 06

    Referrer traffic is not visibility

    AI-motor referrer traffic to your site is a downstream signal, not a measurement of citation share. A page can earn high Perplexity citation share and generate zero referral clicks because the answer is complete without a click-through. The absence of referrer traffic does not mean the absence of AI-search visibility.

Six common mistakes teams make

Even with the tools working, these six mistakes cost measurement quality. Each is a common pattern; all six can be avoided without additional budget.

  1. 01

    Vanity metric conflation

    Mention count is not citation. Citation is not referral traffic. Referral traffic is not revenue. Treating any of these as a proxy for another produces reports that celebrate the wrong thing. Track each as its own signal and resist the urge to combine them into a single composite before you have baseline calibration.

  2. 02

    Single-motor bias

    Measuring only Perplexity because it is the easiest to measure (via API, with a clear citation panel) produces an incomplete picture. Perplexity behavior does not predict ChatGPT behavior does not predict AIO behavior. If you measure one motor, you are measuring one channel of a five-channel discipline.

  3. 03

    Snapshot-only measurement

    One-time firings produce orientation, not measurement. Weekly cadence is the honest floor. Teams that fire quarterly cannot distinguish real trend from single-firing noise.

  4. 04

    Scraping-vs-API assumption

    Tools using scraped AI-motor responses versus API responses produce meaningfully different measurements — Surfer's 1,000-prompt study found only 24 percent brand overlap between the two on the same prompts. Neither is 'correct.' Ask your tool vendor which they use before comparing their numbers against another tool's.

  5. 05

    Prompt library staleness

    Queries retire and buyer-intent language shifts, but many teams never refresh their prompt library after initial setup. Quarterly refresh is the floor. Report any library changes explicitly so historical comparisons stay honest.

  6. 06

    Ignoring the retrieval-without-attribution ceiling

    Assuming citation share equals actual visibility misses the retrieval-without-attribution layer. This ceiling means measurable citation share is always a lower bound on real visibility, never a full accounting. Report accordingly.

What this piece does not cover

Deliberately narrow scope. Adjacent topics that matter for a complete AI-search visibility program but are covered elsewhere on this site or deferred to future pieces.

  1. 01

    How to actually get cited in the first place

    The causal complement to measurement is optimization. If measurement shows you are invisible, the fix is on the citation-earning side. See /writing/how-to-get-cited-by-chatgpt for the full playbook including the ChatGPT paradox and the 79-percent independent-voice Perplexity citation pattern.

  2. 02

    Full tool selection

    The tools that automate the methodology in this piece are covered in depth in the buyer's guide at /writing/complete-guide-ai-search-visibility-tools-2026 — sixty-plus tools across five categories with dated pricing and honest verdicts.

  3. 03

    Motor-specific playbooks

    For deeper per-motor tactics on Perplexity and ChatGPT specifically, see /writing/perplexity-seo and /writing/chatgpt-seo. This piece covers the cross-motor measurement discipline; the motor-specific pieces cover the tactics within each channel.

  4. 04

    Revenue attribution from AI-search referrals

    Attributing revenue to specific AI-search referrals is a harder problem than measuring mention share. It requires session-level tracking of referrer sources ChatGPT and Perplexity do not always populate, plus multi-touch attribution modeling. This is a possible follow-up piece; the honest current answer is that referrer-attribution methodology is not yet mature across the five motors.

  5. 05

    Sentiment analysis at scale

    The scoring rubric in this piece captures sentiment as a per-response field. Automating sentiment classification at scale requires NLP tooling that is out of scope here. Manual sentiment tagging on a fifteen-prompt library takes minutes per week; automation matters only at library sizes above one hundred.

Frequently asked questions

Q · 01

How to track brand mentions in ChatGPT specifically?

Fire your prompt library manually at chatgpt.com in incognito with your target market's location. Check the citation panel when web_search fires and scan the answer text for named-in-text mentions when it does not. Be aware that ChatGPT web_search does not fire on many tool and how-to queries at all — on six to ten of ten samples in our testing — meaning ChatGPT is structurally the hardest of the five motors to measure at scale.

Q · 02

Is it possible to track brand mentions in AI search?

Yes, with three important caveats. Perplexity, Google AI Overview, and Claude produce measurable citation panels or named-in-text mentions on most queries. ChatGPT is measurable only when web_search fires, which is inconsistent on tool queries. Retrieval-without-attribution — cases where your content contributes to an answer without appearing in any user-visible surface — is not measurable from outside the model. Plan around these ceilings.

Q · 03

How much does AI brand mention tracking cost?

The free path in this piece scales to about fifty prompts across five motors weekly and takes roughly thirty to sixty minutes per week. Paid monitoring tools start around twenty-nine dollars per month (Otterly) for entry-tier single-user monitoring and scale to two thousand plus per month for enterprise (Profound). Team-scale mid-market monitoring typically lands at two-hundred-and-ninety-five to five-hundred per month for full-loop platforms.

Q · 04

What is a good AI visibility score?

There is no absolute benchmark that applies across categories. Interpretation requires comparison against category competitors — your share-of-model versus theirs on the same prompt library. As a rough sizing anchor: on a fifty-prompt library, appearing on ten to twenty percent of prompts across all five motors is common for an established brand, forty percent plus is category-leading, and under five percent is typical for smaller or newer brands. Category variance is wide; use competitor comparison rather than absolute thresholds.

Q · 05

How often should I measure AI brand mentions?

Weekly is the honest floor. Daily is overkill for most teams and creates alert fatigue. Monthly misses model-update-driven shifts that can move citation share by twenty to fifty percent in a single week. If your library is large enough that weekly firing is a burden, cut library size before you cut cadence.

Q · 06

Can I track brand mentions in Perplexity for free?

Yes. Fire manually at perplexity.ai in incognito with your target market's location. Perplexity's citation panel is directly visible and copyable. For scale beyond about fifty prompts weekly, the Perplexity Sonar API costs roughly five-thousandths of a dollar per query, which is genuinely negligible for measurement-scale volumes but requires basic API integration.

Q · 07

How to track AI search performance?

Track four metrics weekly per motor: share of model (percentage of prompts where you appear), citation share (percentage of your appearances that are linked citations versus named-in-text), sentiment trend (four-week rolling average), and position shift (first-mentioned percentage versus later-mentioned). Report all four side-by-side and set a twenty-percent week-over-week drift threshold to distinguish signal from noise.

Q · 08

What is the difference between AI ranking and AI citation?

Citation means appearing as a linked source in the motor's citation panel. Ranking means being named or recommended in the answer text with prominence — first mentioned, quoted directly, or recommended by name. A brand can be ranked without being cited (Claude and ChatGPT frequently name brands without citation panels) or cited without being prominently ranked (Perplexity often cites fifteen to twenty sources per answer with only three or four named in the response text).

Q · 09

Do I need a tool to track AI brand mentions?

No, for libraries under about fifty prompts weekly. The free path in this piece — manual firing across five motors, spreadsheet scoring, plus HubSpot AEO Grader and Ahrefs AI Visibility Checker for complementary signals — produces measurement data of the same quality as paid tools at this scale. Paid tools become worth their cost when library size or motor coverage crosses the manual workflow bottleneck.

Q · 10

How do I check my AI citations?

Fire the head queries most relevant to your brand across Perplexity, Google AI Overview, ChatGPT, Claude, and Gemini in incognito. For each response, check the citation panel for your domain and scan the answer text for named-mention of your brand. This is the same methodology paid tools automate; running it manually on ten to fifteen high-priority queries produces a directional signal in under an hour.

Related reading

  • How to Get Cited by ChatGPT (5-motor playbook)/writing/how-to-get-cited-by-chatgpt

    Causal complement to this piece. If measurement shows you are invisible, the citation-earning side is where the fix lives. Documents the ChatGPT web_search paradox referenced in the honest-limits section here.

  • Best AEO Tools and AI Search Visibility Tools 2026/writing/complete-guide-ai-search-visibility-tools-2026

    Tool selection for teams scaling beyond the manual free path. 60+ tools ranked across five categories, dated pricing, non-vendor voice.

  • Perplexity SEO/writing/perplexity-seo

    Motor-specific measurement depth for Perplexity. This piece covers cross-motor discipline; the motor-specific pieces cover tactics within each channel.

  • ChatGPT SEO/writing/chatgpt-seo

    Motor-specific measurement depth for ChatGPT. Broader ChatGPT SEO discipline than the tracking angle covered here.

  • GEO vs SEO vs AEO/writing/geo-vs-seo

    Structural comparison. Establishes the measurement discipline this piece operationalizes.

  • E-E-A-T SEO in 2026/writing/eeat-seo

    Trust axis rationale for citation-source weighting. Explains why some cited sources carry more authority signal than others.

  • Aggarwal et al. — GEO: Generative Engine Optimization (KDD 2024)arxiv.org
  • Chen et al. — Earned-media bias in AI search (arXiv 2509.08919, 2025)arxiv.org
  • Kevin Indig — ChatGPT citations: 44% from first third of content (1.2M responses)almcorp.com
  • Surfer — Scraped AI answers vs API results: 1,000-prompt study (Dec 2025)surferseo.com
  • Vismore — Best Ways to Track Brand Mentions in AI Search (750 Response Study)vismore.ai
  • LLM Pulse — How to Track Brand Mentions in AI Searchllmpulse.ai
  • Perplexity Sonar API documentationperplexity.ai
  • DataForSEO — Perplexity Sonar, Google organic advanced, ChatGPT endpoints (methodology reference)dataforseo.com

Methodology appendix

The methodology in this piece is derived from a combination of primary academic sources cited inline (Aggarwal Princeton KDD 2024, Chen et al. arXiv 2509.08919, Kevin Indig's 1.2M ChatGPT response study, Surfer's 1,000-prompt scraping-vs-API study) and our own 5-motor citation baseline fired on 2026-08-19.

The ChatGPT web_search paradox referenced in the honest-limits section is drawn from six workstreams testing ChatGPT web_search behavior across tool and how-to query classes. Across those six workstreams, web_search fired zero times on the vast majority of samples — a pattern documented in the companion citation-earning piece with the raw data.

Search behavior and motor behavior shift over time. The methodology in this piece is dated August 2026. Re-verify per-motor firing methodology quarterly, especially for motors with fast-moving APIs (Perplexity Sonar, DataForSEO endpoints). If ChatGPT web_search behavior changes materially — for example if web_search begins firing reliably on tool queries — the honest-limits section here will need updating.

5-motor citation baseline
10 head queries × 5 motors (Perplexity Sonar, ChatGPT, Google AIO, Gemini, Claude), full annotation capture, fired 2026-08-19
Head SERP audit for this piece
Top-6 audit for `how to track brand mentions in ai search` fired 2026-08-24 via WebSearch
Cross-workstream ChatGPT web_search pattern
0/N web_search fires across 6 prior workstreams on tool and how-to query classes
Reproducibility
The four-step workflow in this piece is deterministic — same prompt library against same motors on the same date will produce reproducible measurement data. Motor behavior shifts week to week and month to month; the methodology is reproducible, the specific measurements are dated.

New GEO research, as it ships.

Occasional essays on AI search visibility — nothing else, no sponsored content. Unsubscribe anytime.