Last updated September 2026.
Say a tracker tells you your brand shows up in about a third of ChatGPT answers for a topic you watch. Where did that number come from? Not from one universal source. It came from a browser that opened ChatGPT and read the screen, a developer key that called a model endpoint, or a data feed the vendor bought from someone else who did one of the first two things at scale.
Those three pipelines can return different numbers for the identical prompt, and almost no vendor spells out which one feeds which engine. Here is how each pipeline actually works, why two of the five major engines leave no other choice, and what six vendors do and do not disclose about their own plumbing.
Three pipelines, one dashboard
Every visibility platform in this space builds on one of three collection methods, or a mix of them. Here is what each one touches, why a vendor reaches for it, and where it breaks.
| Stage | UI automation | Official model API | Data reseller layer |
|---|---|---|---|
| What it touches | The real consumer product: ChatGPT’s web app, a live Google search, the Gemini or Copilot interface | A developer endpoint the model maker publishes, called directly with a key | A third party’s own scrape or API pipeline, licensed and resold |
| Why a vendor uses it | ChatGPT’s web citations and Google AI Overviews have no public API at any price | Faster and cheaper once a real endpoint exists, and it sidesteps anti-bot defenses entirely | Skips building a collection pipeline in the first place |
| What it returns | What a real user would see on screen right now, browsing tools included | A structured response tied to whatever model version and settings the vendor requests | Whatever the reseller’s pipeline captured, on the reseller’s own refresh schedule |
| Where it breaks | Rate limits, layout changes, and CAPTCHA walls that demand constant upkeep | The API can run a different model checkpoint or grounding setting than the live consumer app | Every dashboard built on the same feed inherits the same blind spots, on the same delay |
Why ChatGPT web and AI Overviews force the UI lane
OpenAI publishes a developer API, and it is genuinely useful for a lot of things. It is not the same product as ChatGPT.com. The consumer app runs live web browsing and drops source links inside the answer; a standard API call does not reproduce that behavior by default, and nothing OpenAI publishes lets a third party pull the exact citation list a ChatGPT.com user just saw. A vendor that wants those citations has to open the same web app a person would use and read what loads.
Google AI Overviews sit in the same spot for a different reason. There is no AI Overviews API, public or paid, from Google. The search-data infrastructure providers that build scraping tools for exactly this problem say so directly: SerpApi confirms the only way to capture an AI Overview and its cited sources is to render the actual Google results page and parse what comes back. Every tool that tracks AI Overviews does that rendering itself or buys it from someone who does.
That covers two of the five engines in this comparison where UI automation, or a reseller feed built on UI automation underneath, is not a design choice. It is the only door in.
Where the official API lane works, and where it quietly diverges
Gemini and Perplexity both publish developer APIs a vendor can call directly, no browser required. That lane runs faster at scale and skips the anti-bot defenses a consumer app throws at automated traffic. It also comes with a catch few dashboards mention: an API response is not guaranteed to match what the consumer app shows for the same prompt.
The gap has a few usual causes. The API call can hit a different model checkpoint than the one live in the app that week. It can skip a browsing or grounding tool the consumer product switches on by default. It arrives with no session history, no account context, and a plainer system prompt than the branded product wraps around it. None of that makes the API wrong. It makes it a measurement of a different, related surface.
Evertune treats that gap as the point rather than an inconvenience. Its published methodology describes a dual-layer approach: querying foundation models directly through their APIs, and separately capturing what the consumer apps answer to the same prompts. Reading both layers side by side shows whether a visibility problem sits inside the model’s own training or inside the retrieval layer a chat product bolts on top of it, a distinction a single-lane tool cannot draw at all.
The reseller layer: one feed, several dashboards
Not every vendor in this space builds its own collection pipeline. Some buy one.
DataForSEO’s LLM Mentions API is the clearest public example. It aggregates mention and citation data from Google’s AI Overviews and ChatGPT at scale, and DataForSEO says plainly that the same feed already powers a new generation of AI visibility and generative-search tracking tools built by other companies. That is the business model in one sentence: one company runs the collection pipeline once, and multiple software vendors put a different dashboard on top of the exact same rows.
None of that is hidden or improper. Building and maintaining browser automation at scale, against engines that actively try to block it, is genuinely hard, and buying that layer from a specialist is a rational call for a small team. The catch sits with the buyer. Two tools that look independent can share a data source, a refresh cadence, and a sampling method you never see, and a vendor’s own marketing rarely says so.
What six vendors disclose about their own pipeline
Public documentation on this specific question is thin across the board. Most vendors publish which engines they cover and how many prompts your plan allows. Far fewer say whether the numbers behind those prompts come from a browser, an API, or a licensed feed.
| Tool | What it discloses about collection | Customer-facing export API |
|---|---|---|
| Ahrefs Brand Radar | Runs a pre-built index of more than 470 million real, search-backed prompts through each AI platform it tracks (Ahrefs, 2026) | Yes, API and MCP access on a paid Ahrefs plan |
| Profound | Not publicly disclosed | Enterprise tier only, terms not published |
| Peec AI | Not publicly disclosed | Not publicly confirmed |
| Otterly.AI | Not publicly disclosed | Yes, launched June 2026 |
| Temso | States directly that it reads real user interfaces, not API calls | Not publicly confirmed |
| Evertune | Dual-layer: direct foundation-model API calls plus consumer-app answers | Not publicly confirmed |
Two things stand out here. First, a customer-facing export API and a vendor’s own internal collection method are two different questions, and a vendor that answers one rarely answers both. Otterly.AI shipped a public API for pulling your own account data out in June 2026, and that says nothing about how Otterly.AI’s own pipeline reads ChatGPT in the first place.
Second, Ahrefs Brand Radar takes a mechanically different approach from the other five. Instead of a hand-built prompt list an analyst wrote, it starts from real search queries in Ahrefs’ own keyword database, expands them into natural-language questions, and runs the result, well over 470 million prompts a month, through each platform it tracks. Every other tool in this table samples a smaller panel someone configured on purpose.
Why the same prompt can score two different ways
Put the three lanes together and the practical takeaway is simple: a visibility percentage is only comparable to another visibility percentage collected the same way. A UI-automation reading of ChatGPT and an API reading of ChatGPT measure two related but genuinely different surfaces. Browsing tools, model checkpoints, and session context can all differ between them. A reseller-fed dashboard adds a third variable on top: someone else’s refresh schedule and someone else’s sampling choices, both usually invisible to you.
None of that makes any single lane wrong. It means a number on a dashboard is only as trustworthy as the pipeline behind it, and that pipeline is the one detail most vendor pricing pages skip entirely.
Before you trust a citation percentage, ask the vendor which of these three lanes produced it, and for which engine specifically. Then check that answer against the full nine-tool field on the LLM visibility benchmark, where every tracker on this site is graded against the same testing protocol.