Ask an AI assistant for the best expense software, running shoe or family SUV and you will not receive a familiar results page. You get a paragraph, perhaps a short list, and sometimes a handful of citations. A few brands are named. The rest cease to exist for the length of that answer. For marketers trained to read rankings, traffic and click-through rates, this new surface is maddeningly opaque. Profound and Evertune have arrived to make it observable.
The companies belong to a young category variously called answer engine optimization, generative engine optimization, AI search optimization or, when everyone tires of inventing acronyms, AI visibility. Their basic wager is easy to understand: if consumers increasingly ask ChatGPT, Perplexity, Gemini, Claude and Google's AI experiences what to buy, brands will pay to know what those systems say about them.
The harder question is what counts as knowing. An AI answer is not a fixed search result. It can vary with phrasing, geography, model version, retrieval mode and one run to the next. A screenshot can start a meeting, but it cannot support a strategy. The useful work begins when thousands of slippery answers become a dataset that can be compared over time.
Two lenses on the same black box
Profound's Answer Engine Insights is organized around prompts. A team defines a brand, competitors and topics, then tracks questions across consumer-facing answer engines. The platform reports visibility, share of voice, sentiment, position and citations. Its documentation says prompts are run daily and that the company captures the front-end experiences people see, rather than relying only on model APIs. That distinction matters because the answer in a browser can include search, citations and product features that a bare model response does not.
The citation view is the useful center of gravity. It shows which domains and pages are being pulled into answers, then lets a marketer segment them as owned sites, competitors, earned media, social pages, institutions and other source types. A prompt where your company is mentioned but never cited tells one story. A prompt where an independent review is repeatedly cited while your own documentation is ignored tells another.
Follow the prompt
Best understood as a diagnostic microscope: prompt-level answers, visibility, source citations, sentiment and comparisons across engines, topics and time.
Survey the market
Best understood as a competitive panel: repeated sampling, unaided brand awareness, position, share of answer and source influence across a category.
Evertune approaches the same uncertainty with the instincts of advertising measurement. Its founders, Brian Stempeck, Ed Chater and Poul Costinsky, were early members of The Trade Desk. The platform emphasizes running prompts repeatedly, benchmarking a brand against the competitive category and attaching confidence to changes. Its Visibility Score measures the percentage of sampled AI responses that mention a brand. Its Share of Answer measures the brand's portion of mentions across all answers. Its AI Brand Score combines the probability of appearing with average position.
That design exposes an uncomfortable flaw in vanity benchmarking. A brand can look dominant if the comparison set quietly excludes the company AI systems actually recommend. Evertune says its category reports discover the broader competitive field, while custom prompt trackers can follow named competitors and keywords. The principle is worth stealing even if you never buy the software: define the denominator before celebrating the percentage.
The new search question is not where you rank. It is whether the answer remembers you at all.YesPress analysis
The metric is not the outcome
Both products can produce a clean score, but each score compresses several editorial decisions. Which prompts count? Were they suggested by a marketer, observed in real use or generated from a category? How often were they repeated? Which model experience was queried? Does a brand count when it appears in a citation but not the prose? What happens when the answer names a product without the parent company?
Profound's documentation offers a concrete example. Visibility is calculated from responses that include at least one brand, then asks how many of those responses include yours. Evertune separately defines visibility, average position and share of answer. Neither definition is inherently universal. A buyer comparing vendors needs to ask for the formula, sampling method, platform coverage and export options before comparing one vendor's score with another's.
There is a second trap. Visibility sounds like influence, but a mention can be negative, incidental or buried. A citation signals retrieval, not persuasion. A favorable answer can still send no traffic. And a rise in one dashboard cannot prove that a particular page edit caused the change. Model updates, changing sources and competitor activity remain confounding forces.
The proper use is diagnostic. If a competitor appears across high-intent prompts and your brand does not, investigate the sources that support the answer. If your documentation is cited for technical questions but ignored for comparison questions, the gap may be format, authority or missing evidence. If the shift appears on one model and not others, resist the urge to declare a universal win.
What a small team can steal
The category's most transferable idea is not a proprietary score. It is a disciplined operating loop. Start with a compact set of questions that represent actual buying, comparison and troubleshooting intent. Include brand-free prompts such as “Which payroll tools work for a distributed team?” because those reveal unaided visibility. Add branded prompts to test accuracy, positioning and stale facts.
Then preserve the prompt set long enough to see change. Run questions more than once, across the AI products your customers use, and save the full answers with dates and citations. Separate three observations: whether you were named, where you appeared and which sources supported the response. A spreadsheet can do this at small scale. Platforms become valuable when the manual sampling burden, regional needs or competitor set grows.
- Choose prompts from customer language. Sales calls, support logs and onsite search are better starting points than a brainstorm of flattering questions.
- Record the denominator. Keep models, regions, run frequency and competitor definitions stable enough for comparison.
- Read the citations. Find the pages repeatedly shaping answers, then classify them as owned, earned, social, institutional or competitive.
- Make one legible change. Improve a product fact page, publish missing evidence or pitch a relevant independent source. Avoid changing everything at once.
- Review trends, not trophies. A durable pattern across repeated runs matters more than a single flattering answer.
This loop also keeps content production honest. The temptation is to flood the web with machine-written pages engineered around every prompt gap. That may satisfy a dashboard for a moment while degrading the source environment on which answer engines depend. Better work is slower and more specific: clear product documentation, original research with methods, comparison pages that acknowledge tradeoffs, expert commentary, and corrections to facts AI systems repeatedly get wrong.
The scoreboard cannot play the game
Profound and Evertune increasingly describe workflows that move from observation toward recommendations, content and distribution. Yet the enduring value of their measurement products is the boundary they reveal. Software can identify a source gap. It cannot make weak evidence credible. It can surface an inaccurate claim. It cannot decide which correction deserves legal, product or editorial approval. It can rank publishers by citation frequency. It cannot manufacture the trust required to appear there.
That is why the best owner of AI-search measurement may not be an SEO specialist working alone. The dashboard crosses product marketing, communications, documentation, analytics and engineering. A false pricing answer belongs with the web and product teams. A missing category association belongs with positioning. A competitor dominating third-party sources belongs with communications. A bot unable to read a page belongs with engineering.
The companies are selling visibility into a channel that is still changing under them. That is not a reason to ignore the category. It is a reason to demand methodological humility. AI discovery is already consequential enough to measure and still too unstable to reduce to one number.
The winning team will not be the one with the prettiest weekly score. It will be the one that turns a repeated pattern into a better page, a clearer fact, a stronger source or a more useful product, then watches long enough to learn whether the answer changed.
Keep exploring
Questions people ask
What is answer engine optimization?
Answer engine optimization is the practice of improving how accurately and often a brand or source appears in AI-generated answers. It overlaps with generative engine optimization, or GEO.
What does Profound measure?
Profound tracks brand visibility, share of voice, citations, sentiment and competitor performance across prompts, platforms, regions and time.
What does Evertune measure?
Evertune measures brand mention probability, position, sentiment, share of answer, source influence and performance against competitors across AI systems.
Why do AI visibility scores change?
AI systems are probabilistic, prompts vary and model or retrieval updates can change answers. Meaningful analysis therefore requires repeated observations and a stable methodology.
Do these tools guarantee better visibility?
No. They can diagnose patterns and opportunities, but outcomes also depend on product facts, accessible owned content, credible third-party sources and changes made by the AI platforms themselves.