Gå til hovedindhold

From AI answers to action

The challenge is that AI answers are opaque. You can’t “view source” on a language model’s reasoning the way you can inspect a search engine’s ranking factors. So how do you actually improve your presence in AI answers, rather than guessing? This is where our action plan methodology comes in: a systematic, evidence-based process for turning raw AI-response data into concrete, prioritized tasks — and then repeating the cycle as the landscape shifts.

The Core Principle: Plans Are Built From Evidence, Not Assumptions

The foundation of our approach is simple but easy to violate in practice: every recommendation in an action plan must be traceable back to real data — specifically, real prompts, the real answers AI models gave to those prompts, and the real content those models drew on or cited when constructing their answers.

This matters because it’s tempting to build SEO-style playbooks based on generic best practices (“write more content,” “get more backlinks,” “improve your meta descriptions”) and simply relabel them as “AI visibility” advice. That approach is fast, but it’s fundamentally disconnected from how large language models actually decide what to mention and what to omit. Our methodology instead starts from the observed behavior of the models themselves, and works backward to figure out why they behaved that way — then forward to what should change.

Step 1: Capture the Real Prompt Landscape

The process begins by defining the set of prompts that matter for a given brand — the actual questions real customers or prospects are likely to type into an AI assistant when they’re in a buying or research mindset. These aren’t keywords in the traditional SEO sense; they’re natural-language questions, comparisons, and recommendation requests.

For each prompt, we capture the AI model’s full response, not just a summary. This includes:

  • Which brands, products, or services are mentioned by name
  • The order and prominence in which they appear
  • The specific claims or attributes associated with each brand
  • Any sources, citations, or content types the model appears to be drawing from

This step produces a dataset that isn’t hypothetical — it’s a factual record of what AI models are currently telling people about a market, a category, and the brands within it.

Step 2: Identify Visibility Gaps

With that dataset in hand, the next step is comparative analysis. The central question we ask is straightforward:

Is our customer being mentioned in these answers? If not, who is being mentioned instead, and why?

This is where the methodology becomes genuinely diagnostic rather than descriptive. A visibility gap on its own — “the customer wasn’t mentioned” — isn’t actionable. It’s just a symptom. The valuable work happens in explaining the gap.

When a competitor appears in an AI answer and our customer doesn’t, we don’t stop at noting the absence. We dig into the content layer: what specific piece of content, page, article, comparison, dataset, or type of resource does the competitor have that appears to be feeding the model’s answer? Is it:

  • A detailed comparison page that directly addresses the prompt’s framing?
  • A specific product spec sheet or pricing table?
  • Third-party reviews, testimonials, or press coverage?
  • A well-structured FAQ or “best for X” style article?
  • Original research, statistics, or data that gets cited repeatedly?

The goal is to identify the content pattern that correlates with being mentioned — not just the fact that a competitor was mentioned.

Step 3: Translate Gaps Into Content Recommendations

Once we know what kind of content is driving the competitor’s visibility, we check whether our customer has an equivalent asset. In most cases, the gap is precisely this: the customer simply doesn’t have that type of content on their site, or what they have is thinner, older, or less directly aligned with how the prompt is phrased.

The recommendation that follows is deliberately concrete. Rather than a vague instruction like “improve your content,” the action plan specifies:

  • The content type to create or upgrade — e.g., a direct comparison page, a buyer’s guide, an updated spec table
  • The angle or framing — that matches how the AI model, and by extension real users, are asking about the topic
  • The competitive reference point — i.e., what the competitor’s equivalent content does that seems to be working

This turns an abstract visibility problem into a concrete production task that a content or marketing team can actually execute.

Step 4: Prioritize and Assign

Not every gap carries equal weight. A prompt that represents a high-intent, high-volume buying question deserves more urgency than a niche or rarely-asked variant. As part of building the action plan, gaps are prioritized based on factors such as:

  • How frequently the underlying prompt pattern is asked
  • How commercially relevant the prompt is (early research vs. late-stage comparison/purchase intent)
  • How consistently the competitor advantage shows up across multiple prompts and models
  • The effort required to close the gap

The result is a prioritized list of tasks — not a data dump, but a plan.

Step 5: Execute, Then Re-Measure

This is the step that separates a one-off audit from an actual system. Once the identified tasks are completed — new content published, existing pages restructured, missing information added — the cycle doesn’t end. We return to Step 1 and re-run the same prompts against the AI models.

This re-measurement step is essential for two reasons:

  • Validation — it confirms whether the content changes actually moved the needle. Did the customer start appearing in answers where they were previously absent? Did their position or framing improve?
  • Discovery of new gaps — the competitive landscape and the models themselves are not static. Once the “low-hanging” gaps are closed, a new layer of gaps often becomes visible: perhaps the customer now appears, but a different competitor has since sharpened their own content, or the models have shifted which sources they favor.

Based on this new data, a new action plan is generated — not from scratch, but as an evolution of the previous one. This makes the whole process iterative and self-correcting rather than a static report that goes stale the moment it’s delivered.

Why This Approach Works

There are a few reasons this evidence-first, content-linked methodology holds up better than generic AI-visibility advice:

  • It’s causally grounded. Instead of assuming what should influence an AI model’s answer, we observe what does — the actual content associated with mentioned brands — and use that as the basis for recommendations.
  • It’s specific enough to act on. “You’re not mentioned” is a diagnosis. “You’re not mentioned because you lack a direct comparison page addressing this exact question, while your competitor has one” is a brief for a content team.
  • It compounds over time. Because the process loops — measure, act, re-measure — the action plan gets sharper with every cycle. Early plans tend to catch large, obvious gaps; later plans catch increasingly nuanced positioning and content-quality issues.
  • It adapts to a moving target. AI models are updated frequently, and the content that “wins” citations can shift. A one-time audit can’t keep up with that; a repeating measurement cycle can.

The Bigger Picture

What this methodology ultimately reflects is a shift in how visibility itself needs to be understood. Where classic SEO action plans were often built around technical and on-page signals (keywords, backlinks, page speed), AI-answer visibility is much more directly tied to the substance of the content: does it actually contain the specific information, comparison, or answer that the AI model needs to construct a good response?

By anchoring every action plan to real prompts, real AI answers, and the real content behind those answers — and by closing the loop with re-measurement — the process avoids the trap of generic advice and instead produces a living, evidence-based roadmap that keeps pace with how AI-driven discovery actually works.

Introduction to GEOtracking.ai

The problem: AI visibility isn’t keyword tracking with a new name

Tools like Morningscore and Accuranker have done well within classic SEO, where keyword tracking makes a lot of sense – there’s search volume, there’s historical data, and a keyword like “best running shoes” is searched more or less the same way by everyone.

That logic has been carried over directly into AI visibility: the user is asked to type in a set of prompts, and the tool then tracks whether the company gets mentioned in the answer to those exact prompts.

The problem is that prompts don’t behave like keywords. There’s no search volume on prompts, and in all likelihood there never will be. A prompt isn’t a fixed search string – it’s a natural-language expression, and two users who are genuinely looking for the same answer rarely phrase it the same way. One user writes “best CRM for a small B2B sales team,” another writes “which CRM should I choose as a small business,” a third asks “what’s the difference between HubSpot and Pipedrive for sales.” Three completely different prompts that cover essentially the same information need – and none of them has any search volume to measure against.

When AI visibility is tracked based on a narrow set of manually entered prompts, you’re actually measuring very little. You get a snapshot of whether the model mentions the company in that exact phrasing – not an accurate picture of the company’s real visibility within a topic.

The future is topic-based – not prompt-based

Everything points to both the AI models themselves and the tools that measure them moving toward a topic-based approach rather than a prompt-based one. The clearest sign of this is already visible at Bing, which has started exposing so-called grounding queries in its Webmaster Tools.

In its simplest form, a grounding query is the specific information gap a language model needed to fill, and that a given website helped close. Instead of showing a user-phrased prompt, it shows the underlying topic or question the model actually searched for information about. If a user asks the AI for a comprehensive travel guide, the grounding query behind it might be something like “tax rules for digital nomads in Portugal.” If your page gets used as a source here, it’s not because you happened to match one specific prompt – it’s because you own the authority on the underlying topic.

This is exactly what will, in the future, be referred to as Topical Authority in AI search: the ability to prove to a language model that your page is the most reliable source for answering a specific subtopic – not the ability to match a single phrasing.

In other words: the broader and more systematically you cover a topic, the higher the probability that you show up, regardless of how any individual user phrases their question. That’s the principle our entire tool is built around.

We’ve turned tracking upside down

Most AI visibility trackers start with the prompt. The user is asked to type in the prompts they want to track themselves – often because that’s the easiest way to build a product, not because it produces the best data foundation.

We do the opposite. We don’t start with prompts – we start with topics.

Instead of the user having to guess which phrasings are relevant, our model starts from the topics that are relevant to the client’s business, and then pulls data across multiple keyword databases to uncover how those topics are actually being searched and asked about. Based on that data, we systematically generate a broad set of prompts that together cover the topic from many angles – instead of betting everything on a handful of manually selected phrasings.

That’s also why our cheapest package starts at 200 prompts. By structure, this ensures a minimum of four distinct topics—each containing between 25 and 50 carefully crafted prompts. It’s not an arbitrary number – it’s the volume required to say anything statistically meaningful about a company’s visibility on a topic, rather than simply measuring a few random samples.

The broader you track, the higher the statistical probability that you actually capture the cases where a company does – or doesn’t – get mentioned. Narrow tracking on a handful of manually chosen prompts gives a distorted picture, because you’re really only seeing a small, random fraction of the overall conversation happening between users and AI models about a given topic.

From data to action plan

Measuring visibility is only half the job. The other half – and it’s at least as important – is turning that data into action.

That’s why we don’t stop at showing how often a company gets mentioned. We build an action plan grounded in data: how are the client’s competitors being mentioned on the same topics? Which subtopics do competitors own that the client isn’t yet present on? And which specific information gaps – the same kind of gaps that Bing’s grounding queries reveal – does the client need to fill to increase the probability of being cited by AI models going forward?

The result isn’t just a dashboard full of numbers, but a prioritized plan for where effort needs to go in order to actually move the needle on visibility.

What sets us apart

In summary, three things set our approach apart from most other AI visibility trackers on the market:

  1. We build on data, not guesswork. Many competing tools ask the user to manually type in the prompts to be tracked. We generate prompts from topics and real keyword data instead, so the selection doesn’t depend on what one person happens to think of.
  2. We track broadly enough to be statistically useful. Where many competitors’ packages start as low as 20 prompts – which doesn’t provide any real statistical basis for assessing topical authority – our cheapest package starts at 200 prompts, precisely because breadth is a prerequisite for reliable results.
  3. We don’t stop at measurement – we deliver an action plan. Many tools only track and report. We turn data on competitors’ visibility and a client’s own gaps into a concrete plan for how visibility can actually be increased.

AI search is rapidly moving away from individual prompts and toward topics and authority. The companies that understand this early – and measure accordingly – will be in a significantly stronger position as AI models increasingly become the first point of contact between customers and businesses.

Why Tracking Fewer Than 25 Prompts Per Topic Never Gives You an Accurate Picture of Your AI Visibility

AI Recommendations Are Random by Nature — And That’s Intentional

The best evidence for this comes from a study by Rand Fishkin (SparkToro) and Patrick O’Donnell (Gumshoe.ai). 600 volunteers ran identical prompts through ChatGPT, Claude, and Google AI nearly 3,000 times, and the results fundamentally shift the way many companies measure AI visibility:

  • Probability of getting the same list twice: less than 1%
  • Probability of getting the same list in the same order: about 1 in 1,000
  • Even the list length varied greatly — from 2 to over 10 recommendations

With those numbers in mind, it doesn’t make sense to say “we rank number 3 on ChatGPT.” The next query could put you in first place — or leave you out entirely. (Source: Trustmary, based on Search Engine Land’s reporting of the study)

What’s interesting is what the study found instead: even though the order swung wildly, certain brands consistently appeared across the queries — in 60% to 90% of cases within certain categories. And even when users phrased their questions completely differently, the AI models managed to recognize the intent and surfaced the same set of brands. Searching for headphones, Bose, Sony, Apple, and Sennheiser kept showing up again and again — regardless of whether the question was phrased in a hundred different ways.

That’s the core of the problem with tracking too few prompts, and it’s also the key to what you should do instead.

Why Fewer Than 25 Prompts Per Topic Is Statistically Worthless

Imagine testing 5 prompts about your category once a week. With a variation where the same list almost never repeats, and the order changes in 999 out of 1,000 cases, you don’t get a picture of reality — you get a snapshot of noise.

A few concrete consequences of testing too few prompts:

  1. You mistake randomness for a trend. If your brand shows up in 2 out of 5 prompts one week and 4 out of 5 the next, it looks like progress on paper. In reality, it can just be normal statistical noise around your actual visibility level.
  2. You miss intent variation. Customers don’t ask just one question about your category — they ask it in hundreds of ways (“best X for Y,” “X vs Z,” “cheapest X that lasts”). With few prompts, you’re typically only testing one or two phrasings and missing how you perform across the full spectrum of purchase intents.
  3. You can’t separate signal from single occurrences. It takes 25+ prompts before patterns start to stabilize, because that’s the scale at which random fluctuations statistically begin to cancel out and a real visibility percentage — that is, how often your brand actually appears — becomes visible.
  4. The competitive picture gets distorted. With few samples, a competitor who happened to show up three times in a row can look like they’re “winning” the category, even though the real picture over 25-50 prompts would show something entirely different.

In short: with a small prompt sample, you’re not measuring your position. You’re measuring the game of chance.

Why “Rank Tracking” Is the Wrong Goal Altogether

Traditional SEO has trained us to think in positions: number 1, number 3, top 10. But as the numbers above show, a position in an AI answer isn’t a fixed quantity — it’s a snapshot of a probabilistic process. Chasing “position 3 on ChatGPT” is therefore a fight against an opponent that’s deliberately designed to vary its answers.

The right question isn’t “where do we rank?” It’s: “How often do we appear at all — across every way people actually ask about our category?”

That’s Why Topic Authority Is the Right Way Forward

If AI models consistently surface the same brands across hundreds of different phrasings — as with Bose, Sony, Apple, and Sennheiser in the headphones example — it’s clear that this isn’t about optimizing for one specific prompt or one specific keyword. It’s about building authority across the entire topic.

Topic authority means your presence covers the category broadly enough that you show up no matter which angle the user asks from — comparisons, recommendations, use cases, pricing questions, “best for X” questions, and so on. Instead of optimizing a single page for a single keyword, you build an ecosystem of content, mentions, and third-party sources that together signal to AI models that you’re a natural part of the answer to the topic — no matter how the question is phrased.

This also fits with how AI models actually work: they identify the user’s intent behind a question and match it against a broader pattern of sources and signals, not against a single optimized document. The broader and more consistent your presence is across the topic — on your own website, in third-party mentions, in comparison articles, in reviews — the greater the probability that you get caught by the broad brush the AI model is painting with, regardless of which of the thousand phrasings the user chooses.

What Does This Mean in Practice?

  1. Track with volume, not samples. Use at least 25-50 varied prompts per topic, run repeatedly over time, to get a statistically sound visibility percentage instead of a random snapshot.
  2. Measure frequency, not position. Ask “in what share of relevant prompts do we appear?” — not “what rank do we hold?”
  3. Cover the full intent spectrum. Build prompt sets that reflect the different ways customers actually phrase things: comparisons, recommendations, prices, use cases — not just your preferred keyword.
  4. Build breadth over depth on a single piece of content. Invest in owning the topic across your own pages, third-party mentions, comparison content, and reviews, so you appear consistently no matter which angle the AI model takes.
  5. Make it an ongoing process. Because AI model answers change month to month, topic authority isn’t something you achieve once — it’s a continuous investment in becoming a fixed part of how the models “understand” your category.

Conclusion

The randomness in AI answers isn’t a bug you can optimize around with a few prompts a week — it’s a built-in property of how the models generate answers. The only way to see through the noise is to test broadly enough for the patterns to emerge, and the only way to win the game isn’t to chase a single position, but to build a presence across the entire topic so broad and consistent that, statistically speaking, you’re almost impossible to leave out — no matter how the question is asked.