What to track, exactly
Four layers make tracking decision-grade rather than decorative. Presence: is your brand named in the answer to each tracked question, per assistant. Position and framing: first mention or afterthought, recommended or merely listed, described accurately or from stale facts. Sources: which pages and platforms the assistant cites or leans on, because sources are where fixes happen. And competitors: who else is named in your moments, since share is relative by definition. Track those four across a frozen basket of buyer questions, the design rules are in how to start doing AEO and GEO, and you have a scoreboard that survives scrutiny.
Cadence and variance, the two traps
Two technical traps ruin most homegrown tracking. Cadence: daily checks feel rigorous but mostly re-measure sampling noise, while quarterly checks miss model transitions entirely; monthly matches how the underlying systems actually change, with an extra pass after major model releases. Variance: assistants answer probabilistically, so a single ask per question is a coin flip pretending to be data, credible tracking samples each question multiple times and reports frequency, you appeared in seven of ten runs, rather than binary presence. Any tool or process that shows you one answer per question per month is showing you anecdotes with a dashboard on top, part of the scorecard critique in AEO tools.
From tracking to action
Tracking earns its cost only when readings become fixes. A moment where you never appear routes to coverage work: does any page of yours answer that question directly? A moment where you appear inconsistently routes to evidence work: thicker citations and consistency, per how to get recommended by AI. A moment where you appeared and stopped routes to source diagnosis: what changed in the cited documents. And framing problems, named but described wrong, route to fact alignment. This routing is the difference between monitoring and managing, and automating the whole loop, sampling, diagnosis, fixes, re-measurement, is exactly what Aethon does across ChatGPT, Gemini, Claude, and Perplexity, with the free as the entry baseline.
A minimal tracking sheet that actually works
If you start manual, structure beats effort. One row per question per assistant per month; columns for named yes-no across three samples, first-brand mentioned, framing note, and cited sources. Three samples per question is the floor that makes frequency meaningful, ask, regenerate, ask again in a fresh session, and the source column is the one teams skip and regret, because it is the diagnosis. Twenty questions, four assistants, three samples is 240 asks: batched with copy-paste discipline, a long afternoon monthly, which is precisely the tedium that either becomes a protected ritual or becomes the reason to hand the loop to software. Either outcome beats the middle path of tracking one sample of five questions and calling it measurement.