
How to Test Whether Your AEO Changes Are Actually Working
Making changes to improve AI visibility is one thing. Knowing whether those changes worked is harder. Here's a practical framework for attributing AEO improvements to specific actions.
The attribution problem in AEO
You update your homepage description. You get 15 new G2 reviews. You publish a comparison page. A few weeks later, you check your AI visibility and it seems better.
But which change did it?
This is the core challenge of AEO experimentation. Unlike paid advertising, where you can turn a campaign on and off and see immediate results, AEO changes work through signal accumulation. AI engines don't tell you why they changed their answers. And multiple changes are usually happening at the same time.
Without a testing framework, AEO becomes a cargo cult: you do things that seem like they should work, and you hope the numbers move.
Start with a baseline
Before making any changes, you need to know where you are. A baseline is a snapshot of your current AI visibility across a specific set of queries, measured consistently across the engines you care about.
Your baseline should capture:
- Mention rate: the percentage of queries where your brand appears in the answer
- Query coverage: which query types surface you (category queries, comparison queries, use-case queries)
- Source citations: which external sources AI engines reference when they mention you
- Description accuracy: how AI engines characterize your product when they do mention you
Capture this before you make any changes. If you're already mid-experiment with multiple changes underway, wait for the current changes to stabilize before establishing a fresh baseline.
A baseline measured once is a data point. A baseline measured weekly is a trend. You need at least three consistent measurements before treating any reading as reliable.
Design for isolation
The single most important principle in AEO testing is one variable at a time. This is harder than it sounds because AEO changes often compound: publishing a new FAQ page also adds content to your site, which affects how crawlers index you, which may also affect whether new reviews reference that content.
The changes easiest to isolate are:
- On-site content changes (rewriting specific pages)
- Review volume increases on a specific platform
- A new comparison or alternatives page
- A press mention or editorial placement
- Adding or updating structured data markup
For each change, document the specific action, the date you made it, and the state of other variables. If you make multiple changes, stagger them by at least four weeks so you can observe signal lag before introducing the next variable.
Signal lag by change type
Different AEO changes have different absorption times. Measuring before the signal has propagated produces misleading results.
| Change type | Typical signal lag | Notes |
|---|---|---|
| Homepage or about page rewrite | 2 to 4 weeks | Fastest; live retrieval engines pick it up quickly |
| New FAQ or comparison page | 2 to 6 weeks | Depends on crawl frequency and page authority |
| Review platform additions | 4 to 8 weeks | Platform aggregation plus engine indexing adds time |
| Press or editorial coverage | 6 to 12 weeks | Training data updates are less frequent than retrieval |
| Analyst report or research citation | 8 to 16 weeks | High trust, slow propagation |
| Wikipedia or Wikidata edits | 1 to 3 weeks | Checked frequently by engines |
Most brands measure too early and conclude their changes didn't work. Wait the full signal lag period before drawing any conclusions.
What counts as a meaningful result
Not every change in your mention rate is signal. AI engine outputs have inherent variance: the same query can produce different answers on different days.
A result is meaningful when:
- It persists across multiple measurement sessions, not a one-time reading
- It appears on more than one AI engine, not a quirk of one model's current state
- The direction of change matches the type of change you made (if you published a use-case page, you'd expect improvement in use-case queries specifically)
If your mention rate improved but only on one engine and only for one week, treat it as noise. If it improved across all three engines over multiple measurement sessions, treat it as signal.
Building a testing log
The practical version of all of this is a simple log. For each test, record:
- The change made (specific, not vague: "rewrote the homepage first paragraph to include 'automated contract management for solo attorneys'")
- The date the change went live
- The expected signal lag
- The queries you're monitoring for change
- Baseline measurements before the change
- Follow-up measurements at the end of the lag period
- Your conclusion
This log becomes your institutional knowledge about what works for your specific brand. How to track AEO performance covers the monitoring infrastructure. The log is what turns that monitoring data into learning you can act on.
Confounding variables to watch
A few external factors can shift your AI visibility with no action on your part.
Competitor activity. If a competitor starts appearing more often in answers, your relative share may decrease even if your absolute signal hasn't changed. Auditing competitor AI visibility helps separate your absolute visibility from your relative standing.
AI engine updates. The engines themselves change how they retrieve and weight sources. A sudden change in your mention rate that coincides with no change on your part is often an engine-side update, not something you caused or can fix.
Organic press or community coverage. A mention you didn't actively pursue, or a viral forum thread about your product, can move your visibility significantly. If this happens during a test period, note it in your log as a confound.
The minimum viable AEO test
If you want to start testing without a sophisticated infrastructure, here's the simplest version:
- Pick one query where you currently don't appear.
- Identify the most likely reason (missing FAQ, no comparison page, thin review volume).
- Make exactly one change to address that reason.
- Wait the appropriate signal lag period.
- Run the query five times across each engine and record whether you appear.
That's a test. It's not statistically rigorous, but it's better than making five changes at once and guessing which one worked.
QuickAEO runs your keywords across ChatGPT, Perplexity, and Gemini with multiple trials and tracks your mention rate over time. That measurement layer is the foundation for running clean AEO experiments rather than making changes and hoping the numbers move.