Blog · · By HeardOf

How to track your brand in ChatGPT yourself: a spreadsheet method, and where it breaks

The short answer

Write 20 to 50 buyer prompts, ask each engine once, score every answer yes or no per brand, named and cited. It breaks at one run per prompt: standard error 0.370, per Schulte, Bleeker and Kaufmann, April 2026, historical.

Does AI name your brand? Ask Gemini now — free, no account →

What do you need to track brand mentions in ChatGPT by hand?

A list of prompts written the way a buyer types them, one answer per prompt from each engine, and a sheet with one row per brand per answer and two yes-or-no columns, named and cited. That is the whole apparatus needed to track brand mentions in ChatGPT by hand; the rules for the two columns, and the reason one pass through the sheet is not enough, are the rest of this post. Which kinds of buying question belong on that list is in how to choose the prompts you track.

The worked example is our run of 4 September 2026, as how to get your SaaS cited by ChatGPT and Perplexity describes it: five buyer prompts in each of five B2B software categories, put to ChatGPT and Perplexity for ten answers per category, six tracked vendors per category, 300 brand-answer pairs; for every answer, which vendors were named, which were cited by URL, and every source the engine returned. All 25 prompts are printed there verbatim, five per category, with the caveat the post states itself: the third prompt in every category names three of the six vendors outright, so those three get a mention rate partly handed to them by the question. Its advice to a reader is 20 to 50 prompts, run against each engine on the same day of the month, every named vendor and every cited URL logged; the first run is the baseline.

How do you build the sheet?

One row per tracked brand per answer, so 25 prompts on two engines with six brands per category is 300 rows. The columns:

ColumnWhat goes in itThe post that states the rule
DateThe day the answers were collectedGet cited: run them against each engine on the same day of the month
EngineChatGPT or PerplexityGet cited: mentions and citations separately, per engine
PromptThe exact wording from your list, the same on every runGet cited: re-running the same prompts; all 25 printed verbatim
Answer textEverything the engine wrote, inline link text and web addresses includedNamed but not cited: the rule as applied on 23 September 2026
SourcesEvery URL the engine returned as a source for that answerGet cited: every source the engine returned
BrandOne tracked brand; each answer gets one row per brandNamed but not cited: scored once per brand
Alternative namesThe other names that count for this brandNamed but not cited: Motive also as KeepTruckin
Registered domainThe brand's own domain, with www. removedNamed but not cited: kvCORE's is insiderealestate.com
NamedYes or no, by the rule in the next sectionNamed but not cited
CitedYes or no, by the rule in the next sectionNamed but not cited
The columns of a one-run tracking sheet, in the order they are filled. Every rule in the third column is stated in one of two posts of ours, read live on 24 September 2026; the table adds none they do not state.

Write the prompts first and keep the wording. Ask each one once in each engine on one day. Paste the whole answer and every source into the rows before scoring anything. Score one brand at a time, then sum per engine: a brand's named rate on ChatGPT is its yes rows over the number of ChatGPT answers, the same for cited, and the engines stay apart because, in the earlier post's words, "an average across engines hides both". In our run of 4 September 2026 the sums were 108 of 240 rows named and 72 cited, 45.0% and 30.0%, over the 24 brands in the four categories whose stored answers could be re-counted on 23 September 2026, and the engines differed, ChatGPT 60 named and 49 cited of 120, Perplexity 48 and 23, as named but not cited reports.

What counts as named, and what counts as cited?

Named: the brand's name, or one of the alternative names you listed for it, appears in the answer text as a whole word, in any case, link text and addresses the engine printed inline included. Cited: at least one URL among the answer's sources has the brand's registered domain as its host, www. removed, or a subdomain of it, so help.fleetio.com counts for Fleetio. Both are the rules named but not cited states as applied on 23 September 2026, each yes or no per brand per answer: a brand named three times in one answer counts once, and so does a brand cited three times.

The same post shows where the lines fall. ChatGPT once cited a page at propertybase.lwolf.com, on 4 September 2026, in an answer that named Propertybase; that host is not propertybase.com, so the row reads named yes, cited no. In one staffing answer of 4 September 2026 the one occurrence of JobAdder was inside a web address the engine printed in the text; the rule counts it as named, and a reader would not. Decide such cases before you score and keep the decision between runs; a rule that moves is a second source of change on top of the one below.

Where does the method break?

At one run per prompt, in the words of the study that put a number on it: "A single run (SE = 0.370) is essentially uninformative". Schulte, Bleeker and Kaufmann (arXiv 2604.07585, 8 April 2026, historical by our 90-day rule) asked eight prompts in each of four Swiss-German campaign verticals up to ten times in succession on ChatGPT, Gemini, Google AI Mode and Perplexity, between 21 and 25 March 2026, and computed for 1,216 per-brand series, from the three verticals whose brand detection cleared its quality threshold, how far an estimate from a given number of runs sits from the ten-run mean. Its coverage table says the daily data of 24 January to 20 March 2026 was collected via web-interface scraping; the paper does not say how the 21 to 25 March runs behind the table below were collected.

Runs per promptStandard error95% interval, plus or minus
10.3700.724
20.2460.483
30.1880.369
40.1510.296
50.1230.241
60.1010.197
70.0810.158
80.0620.121
90.0410.081
Table 16 of Schulte, Bleeker and Kaufmann, arXiv 2604.07585, 8 April 2026, read on 24 September 2026: subsampling standard error of the estimated per-brand detection rate by number of runs, averaged across 1,216 per-brand series, the ten-run mean standing in for the true rate. The paper's note: the nine-run row is subject to finite population correction and underestimates the error of truly independent runs.

Their recommendation is "at least 7 runs per prompt per day for brand visibility monitoring, and at least 8 runs when source-level coverage matters"; at seven, the interval of plus or minus 0.158 is "adequate for detecting large differences (e.g., a brand detected in 80% vs. 20% of runs)" and not for finer ones. For the sheet, a one-run row records whether a brand was in that answer, and the rate over 25 prompts is a rate over prompts; it is not a finding that you are absent for any one prompt, since in the paper's description "brands or sources may appear in one response and disappear entirely in another". One more of its lines bites, for its own runs of March 2026: "ChatGPT activates web search only for specific queries, leaving 57.8% of its runs with zero citations", so a Cited column that reads no all the way down on ChatGPT may be an engine that did not search, which an empty Sources column shows and the rate does not.

In one study, the citation decision moved on its own even with the sources held fixed. Selvam and Ghosh (CITECHOICE, arXiv 2609.15164, 14 September 2026) froze the full transcripts of a GPT-5.4 search agent, retrieved text included, and generated the final answer a second time: "In 120 independently regenerated cells, binary target citation agrees 85.0% of the time, meaning one in seven decisions flips." The unit is a cell, one answer-target family under one of four replay conditions, 30 families; what was compared is whether the target document was cited at all between the first generation and an independent decoding pass: 102 cells unchanged, 6 gained a citation, 12 lost one. One retrieval provider and one model, as its limitations say, and frozen transcripts rather than a live engine: the flip is the generation step alone, measured on that setup and not on ChatGPT.

Does our own audit have the same limit?

Yes, and it says so. The four engines we ask, as updated 23 September 2026, states that each prompt is asked once per engine per UTC day and that the report gives rates per engine with a headline that is their mean, both our own rules as read on 19 September 2026; it cites the same Schulte study, and it does not tell a subscriber "you are absent for this prompt", because at one answer that claim is close to a coin flip. Forty prompts in a fixed mix and four brands scored per run, from the same 19 September reading; four engines as of 23 September 2026. The sheet above is the same one-answer design at a different size; nothing in this post measures anything new.

How was this read?

Every page was fetched on 24 September 2026 and read in full: our three posts and the two arXiv abstract and HTML pages, with curl and a browser user-agent between 19:29 and 19:30 UTC, every span quoted above found in the served text. The papers are dated from their abstract pages; Schulte's, at 169 days, is labelled historical. Every figure is the study's or the post's own, none was recomputed, and nothing was run against any engine. What this post says about our own audit comes from our post on which engines we check and cannot be checked from here; the two posts the method is taken from and the two papers it breaks on carry a link and the day they were read.

Common questions

How many prompts do you need to track brand mentions in ChatGPT?

Our post of 4 September 2026 on getting cited says 20 to 50, written the way a buyer types them; the run it describes used 25, five per category, all printed verbatim. Schulte, Bleeker and Kaufmann (arXiv, 8 April 2026, historical) write that monitoring based on one or two prompts reflects the idiosyncrasies of those prompts rather than campaign-level visibility.

How many times should you run each prompt?

At least 7 per prompt per day for whether a brand appears, and 8 when the cited sources matter, per Schulte, Bleeker and Kaufmann (arXiv, 8 April 2026, historical), whose single-run standard error of 0.370 falls to 0.081 at seven. Our run of 4 September 2026 asked each prompt once per engine, and so does our audit, our own rule as read on 19 September 2026, per our post on which engines we check, as updated 23 September 2026; a one-run sheet gives a rate over prompts, not a verdict on one of them.

What counts as a brand mention in a ChatGPT answer?

In the rules our post named but not cited states as applied on 23 September 2026: the brand's name, or an alternative name listed for it, as a whole word in the answer text, any case, inline link text and addresses included, yes or no per brand per answer. A citation is a separate count: a source URL whose host is the brand's registered domain or a subdomain of it. In that re-count of our run of 4 September 2026, ChatGPT and Perplexity pooled, 37 of 240 pairs were named without a citation and 1 was cited without a name.

Is the cited column less stable than the named column?

Schulte, Bleeker and Kaufmann (arXiv, 8 April 2026, historical) find the convergence slower for sources, 8 runs against 7 for brand detection. Selvam and Ghosh (arXiv, 14 September 2026) measured citations only: with the retrieved sources frozen, whether a target document was cited at all agreed 85.0% of the time across two generations of 120 cells, one decision in seven flipping with nothing but the generation step varying.

These are our numbers. Yours are one audit away.

One sheet, two columns per brand, and the limit printed at the step where it bites: one answer per prompt gives a rate over your prompts, not a verdict on any one of them. Everything this post says about our own audit comes from our post on which engines we check, as updated 23 September 2026, and cannot be checked from here; the two posts the method is taken from and the two papers it breaks on carry a link and the day they were read. If you would rather the counting were done for you, with that limit stated the same way — your buyers' questions, you against up to three competitors you name, our own rule as read on 19 September 2026, who got named and who got cited — that is what HeardOf does.

How to choose the prompts you track: buying questions, not keywords