The five core metrics
Mention rate: the percentage of relevant buyer prompts where your brand appears at all, measured across ChatGPT, Gemini, Claude and Perplexity. Answer position: whether you are the first recommendation or a trailing mention. Sentiment: how the answer frames you when you appear. Share of voice: your mentions as a share of all brand mentions across your tracked prompts. Citation coverage: how many of the sources assistants cite for your category mention you.
Each metric answers a different question. Mention rate tells you if you exist. Position tells you if you win. Sentiment tells you if the win is worth having. Share of voice tells you who is beating you. Citation coverage tells you why.
KPIs worth putting on a dashboard
For a quarterly executive view, three KPIs carry the story: share of voice against your top three competitors, mention rate on your ten highest-revenue buying moments, and first-recommendation rate on the same set. Everything else is diagnostic detail for the team doing the work.
Resist the urge to track hundreds of prompts. Twenty well-chosen buying moments, tested in several phrasings each, produce more decision-grade signal than a thousand generic keywords.
Benchmarks: what good looks like
From our research across 1,200 buyer moments: being mentioned in over half of relevant buying-moment prompts puts you in the top tier for most categories. Leading the answer in a quarter of them usually indicates category dominance. New entrants commonly start under ten percent mention rate, which is a baseline, not a verdict.
Local and niche categories run higher: a strong regional firm can reach seventy percent mention rate for local moments, while crowded SaaS categories fight for every point.
Why measurement wobbles, and how to handle it
Assistants retrain, sample sources and personalize, so the same prompt can return different brands on different days. A five-point weekly swing is noise. The honest methodology tests multiple phrasings per moment, measures across all four major assistants, and judges monthly trends instead of daily snapshots.
Any vendor showing you a daily AI visibility score without confidence ranges is selling precision that does not exist.
From metrics to movement
Metrics only matter if they change what you do. The working loop: baseline the five metrics, trace weak moments to their causes in the four layers of AI presence, content, structure, authority and entity, ship fixes, and re-measure monthly. Aethon automates the measurement across the four assistants and pairs it with the execution, and the free audit produces your starting scorecard with real answer screenshots.