// open source
answer-engine-benchmark
- Install
pip install answer-engine-benchmark- Runs on
- Python 3.10 or later
- Licence
- MIT, free to use
answer-engine-benchmark asks four AI engines your buyers’ questions, three times each, and counts how often the answers cite your site.
pip install answer-engine-benchmark
aeb run questions/template.yaml --out runs/first \
--set brand="Acme Analytics" --set domain=acme.example \
--set category="invoice software" --set audience="small accounting firms"
What we found with it
We ran it on our own site the day after we moved it to a new build. 30 questions, four engines, three runs each: 360 answers for $1.39 in API fees, with the Claude answers on a subscription.
Asked about us by name, the engines cited our site in 105 of 156 answers. Asked for a provider or for advice, they cited it once in 204. The full write-up is AEO vs SEO: what we measured on our own site.
How it works
The question set has four groups. Discovery and comparison questions never name you. They show whether engines find you when nobody asked for you. Brand and trust questions do, and show what the engines say about you. You fill in your name, domain, category and audience, and a run won’t start until you have.
An answer cites you when one of its URLs is on your domain, in the text or in the sources the engine attached. The tool also records where your first citation sits among the sites cited, and flags phrases near your name that describe you wrongly. In our run, 41 of 156 brand answers mentioned work from our old site.
Every question is asked three times per engine by default, because one answer
is one sample. aeb noise re-asks a few questions later and checks whether the
new results sit within one citation of the old. Ours did, for all 15 pairs we
re-asked.
Every answer is kept in a JSONL file with its full text, sources, model, tokens and cost. Any number in the report can be traced to the answers behind it.
The engines
| Engine | Model by default | Key |
|---|---|---|
| OpenAI | gpt-5-mini |
OPENAI_API_KEY |
| Gemini | gemini-3.5-flash |
GEMINI_API_KEY |
| Perplexity | sonar |
PERPLEXITY_API_KEY |
| Claude | claude-sonnet-5 |
ANTHROPIC_API_KEY |
Keys are read from the environment and never written to the results. Any model
can be changed with --model.
Sharing a report
--anonymise replaces every domain that isn’t yours with “Source A”, “Source B”
and so on. Use it before a report leaves your company.
Related
- Answer engine optimization, where we run this panel for clients and report the sample size with every reading
- answer-crawler-check, which checks that AI answer crawlers can fetch your pages at all
- Open source, everything else we have published or had merged