Table of content

AI Recommendations Are Random by Nature — And That’s Intentional

The best evidence for this comes from a study by Rand Fishkin (SparkToro) and Patrick O’Donnell (Gumshoe.ai). 600 volunteers ran identical prompts through ChatGPT, Claude, and Google AI nearly 3,000 times, and the results fundamentally shift the way many companies measure AI visibility:

  • Probability of getting the same list twice: less than 1%
  • Probability of getting the same list in the same order: about 1 in 1,000
  • Even the list length varied greatly — from 2 to over 10 recommendations

With those numbers in mind, it doesn’t make sense to say “we rank number 3 on ChatGPT.” The next query could put you in first place — or leave you out entirely. (Source: Trustmary, based on Search Engine Land’s reporting of the study)

What’s interesting is what the study found instead: even though the order swung wildly, certain brands consistently appeared across the queries — in 60% to 90% of cases within certain categories. And even when users phrased their questions completely differently, the AI models managed to recognize the intent and surfaced the same set of brands. Searching for headphones, Bose, Sony, Apple, and Sennheiser kept showing up again and again — regardless of whether the question was phrased in a hundred different ways.

That’s the core of the problem with tracking too few prompts, and it’s also the key to what you should do instead.

Why Fewer Than 25 Prompts Per Topic Is Statistically Worthless

Imagine testing 5 prompts about your category once a week. With a variation where the same list almost never repeats, and the order changes in 999 out of 1,000 cases, you don’t get a picture of reality — you get a snapshot of noise.

A few concrete consequences of testing too few prompts:

  1. You mistake randomness for a trend. If your brand shows up in 2 out of 5 prompts one week and 4 out of 5 the next, it looks like progress on paper. In reality, it can just be normal statistical noise around your actual visibility level.
  2. You miss intent variation. Customers don’t ask just one question about your category — they ask it in hundreds of ways (“best X for Y,” “X vs Z,” “cheapest X that lasts”). With few prompts, you’re typically only testing one or two phrasings and missing how you perform across the full spectrum of purchase intents.
  3. You can’t separate signal from single occurrences. It takes 25+ prompts before patterns start to stabilize, because that’s the scale at which random fluctuations statistically begin to cancel out and a real visibility percentage — that is, how often your brand actually appears — becomes visible.
  4. The competitive picture gets distorted. With few samples, a competitor who happened to show up three times in a row can look like they’re “winning” the category, even though the real picture over 25-50 prompts would show something entirely different.

In short: with a small prompt sample, you’re not measuring your position. You’re measuring the game of chance.

Why “Rank Tracking” Is the Wrong Goal Altogether

Traditional SEO has trained us to think in positions: number 1, number 3, top 10. But as the numbers above show, a position in an AI answer isn’t a fixed quantity — it’s a snapshot of a probabilistic process. Chasing “position 3 on ChatGPT” is therefore a fight against an opponent that’s deliberately designed to vary its answers.

The right question isn’t “where do we rank?” It’s: “How often do we appear at all — across every way people actually ask about our category?”

That’s Why Topic Authority Is the Right Way Forward

If AI models consistently surface the same brands across hundreds of different phrasings — as with Bose, Sony, Apple, and Sennheiser in the headphones example — it’s clear that this isn’t about optimizing for one specific prompt or one specific keyword. It’s about building authority across the entire topic.

Topic authority means your presence covers the category broadly enough that you show up no matter which angle the user asks from — comparisons, recommendations, use cases, pricing questions, “best for X” questions, and so on. Instead of optimizing a single page for a single keyword, you build an ecosystem of content, mentions, and third-party sources that together signal to AI models that you’re a natural part of the answer to the topic — no matter how the question is phrased.

This also fits with how AI models actually work: they identify the user’s intent behind a question and match it against a broader pattern of sources and signals, not against a single optimized document. The broader and more consistent your presence is across the topic — on your own website, in third-party mentions, in comparison articles, in reviews — the greater the probability that you get caught by the broad brush the AI model is painting with, regardless of which of the thousand phrasings the user chooses.

What Does This Mean in Practice?

  1. Track with volume, not samples. Use at least 25-50 varied prompts per topic, run repeatedly over time, to get a statistically sound visibility percentage instead of a random snapshot.
  2. Measure frequency, not position. Ask “in what share of relevant prompts do we appear?” — not “what rank do we hold?”
  3. Cover the full intent spectrum. Build prompt sets that reflect the different ways customers actually phrase things: comparisons, recommendations, prices, use cases — not just your preferred keyword.
  4. Build breadth over depth on a single piece of content. Invest in owning the topic across your own pages, third-party mentions, comparison content, and reviews, so you appear consistently no matter which angle the AI model takes.
  5. Make it an ongoing process. Because AI model answers change month to month, topic authority isn’t something you achieve once — it’s a continuous investment in becoming a fixed part of how the models “understand” your category.

Conclusion

The randomness in AI answers isn’t a bug you can optimize around with a few prompts a week — it’s a built-in property of how the models generate answers. The only way to see through the noise is to test broadly enough for the patterns to emerge, and the only way to win the game isn’t to chase a single position, but to build a presence across the entire topic so broad and consistent that, statistically speaking, you’re almost impossible to leave out — no matter how the question is asked.