
Foundry Field Reports
Market Research for Machine Answers
AI visibility tools are not scarce, and the best of them are growing quickly. What the category inherited is SEO's mental model: a list of queries, a position, one number that goes up. But an assistant does not rank you. It describes you — differently every time somebody asks. Describing is what research measures, and research is designed, not tracked.
The category is measuring it like SEO. It behaves like a customer study.
Contents (8)
- 01An AI answer is an observation, not a ranking.
- 02The questions you pick decide the number you get.
- 03Five things change the answer while your site sits still.
- 04Citations are not links to collect. They are quotes from whoever the machine asked.
- 05What gets cited tells you what to make next.
- 06“41% of what?” is the question that kills most visibility scores.
- 07Attribution is the hard part. Almost nothing arrives labeled.
- 08Start with your brand, then widen the questions.
01
An AI answer is an observation, not a ranking.
Ask ChatGPT which vendors a buyer in your category should evaluate. Then ask it again. You will get a different answer.
That is not a glitch in your monitoring tool. That is the whole problem with how this category gets sold to you.
Rank tracking trained all of us to think visibility is a position. You ranked fourth. Someone passed you. Now you rank fifth. The number described a slot, and the slot sat still long enough to matter.
An assistant holds no slots. It writes an answer, once, and it may write a different one an hour later with nothing changed on your site or your competitor's. Ask four assistants who a mid-market ops team should evaluate and you get four different shortlists. Ask one assistant twice and the order moves.
A ranking gives you one slot you either hold or lose. The same question asked eight times gives you a distribution — named in three, absent from five.
That movement is the thing you are measuring. A model does not store a ranking of your market. It writes an account of your market on demand — from what it retrieves and what it already absorbed — and every answer is one sample of that account.
Treat that sample like a rank and you make two mistakes. You panic at normal variance. And you celebrate one good answer as a position you now hold. You hold nothing.
A rank is a position. An answer is an observation.
02
The questions you pick decide the number you get.
Two agencies measure the same brand on the same afternoon. One reports 20%. One reports 60%. Both are telling the truth.
One asked forty questions about the category. The other asked forty questions about the brand by name. Neither number is wrong. Neither number compares to the other — and that kills the idea of a universal visibility score, permanently.
A percentage with no question list attached is decoration. Pollsters publish their questions and their wording, because a century of survey research proved the wording is the finding. Nobody in this category publishes anything.
| Survey research | Measuring machine answers |
|---|---|
| Target population | The questions your buyers ask |
| Sampling frame | The prompt corpus available to you |
| Questionnaire | Your tracked prompt portfolio |
| Question-wording effects | Phrasing, framing and context |
| Respondent segments | Personas, regions, decision stages |
| Survey mode | ChatGPT, Claude, Gemini, Perplexity |
| Repeat waves | Scheduled runs over time |
| Raw responses | Preserved answers and citations |
Two things that comparison does not license.
- A persona is a context you typed, not a demographic you sampled.
- Synthetic questions are not demand. They are your best guess at what buyers ask until somebody shows you real query data. Take the discipline. Leave the claim.
03
Five things change the answer while your site sits still.
Wording. Model. Persona and place. Whether it searched the web this time. When you asked.
| What moved | What it does to the answer |
|---|---|
| Wording | “Best CRM for a small team” and “affordable CRM for a five-person agency” are the same question. They return different answers. |
| Model | One assistant favors your competitor for reasons you cannot see and cannot argue with. |
| Persona and place | Berlin gets a different list than Boston. One study across seven models and fourteen countries found source selection shifts by platform, language and market. (Published by Profound.) |
| Retrieval | Some runs fetch live pages. Some answer from memory. The same question, researched or recalled. |
| Time | The index moved on Tuesday. Nobody told you. |
So your visibility dropped six points. Did a competitor publish something? Did an assistant change how it retrieves? Or did you just sample a different corner of the same distribution?
You cannot tell. And you are about to spend a quarter fixing whichever cause you guessed first. Hold four of the five still and you get an answer instead of a guess.
04
Citations are not links to collect. They are quotes from whoever the machine asked.
Everyone arriving from SEO reads citations as backlinks with extra steps. Get cited more, rank higher, win. That instinct will cost you a year.
A backlink is authority you accumulate. Volume matters, anchor text barely does, the mechanism is a vote. A citation is a quotation. The model reaches into a crowd of millions, pulls a handful of sources aside, asks them what the market says, and repeats what they said.
Nobody should care how many citations you have. Care about this: who got asked, and what did they say about you.
Four things the transcript shows that no link report will
Who speaks for your category
whether or not they mention you
The sources cited across every assistant. That is the list of outlets writing your market's account of itself, and most of it is press you have never pitched.
What they said
read the sentence, not the domain
The same source can call you the category leader or the cautionary tale. A citation count treats those identically. This is where the real damage hides: not missing, but present and described as the expensive enterprise option when you sell to startups.
Who you appear beside
the machine's competitive set
Co-citation is the comparison set the model builds for you. It is usually not the one in your deck. It is the one your buyers are handed.
Which questions have no good source
the whole game
The model answers them anyway — from something dated, thin or hostile, because nothing better exists.
When the model answers a real buying question from a weak source, the market just handed you a brief. This question. This claim. This evidence. Because the machine is currently reaching for something worse. You are not writing more content. You are replacing one bad source on one question.
And replacing it usually is not an article. It is fixing the comparison page the model keeps citing badly. Publishing the proof a claim needs so a skeptical source has something to cite. Correcting the third-party listing carrying a price you retired two years ago. Briefing the publisher who turns out to speak for your category. The diagnosis names the fix, which is something a keyword report has never once done for you.
One caution: being cited is not being used.
- An answer can list your page at the bottom while every claim in it came from somewhere else. Attached, not absorbed. Check whether the claims track the source before you celebrate.
05
What gets cited tells you what to make next.
The useful question is not who linked to you. It is what kind of page the machine reaches for when it answers your market's questions.
Sort the citations for one question cluster by the kind of page they are and a pattern falls out immediately. Comparison questions pull comparison pages. “How do I” questions pull documentation and specs. “Best tool for” questions pull roundups and listicles. And some questions pull a four-year-old forum thread, because nothing better exists — which is the most actionable finding available to you.
That pattern is not the same across sectors, and it is not intuitive. Regulated categories get cited from institutional and standards pages. Developer tools get cited from documentation and changelogs. Consumer categories get cited from reviews and video transcripts. You do not have to guess which one you are in — the answers name their sources, and the sources have kinds.
The format it keeps reaching for
make that, not a blog post
If eight of ten answers to your comparison questions cite comparison tables, the brief writes itself. Publishing another thought-leadership post against that pattern is choosing to lose.
Where it settles for something weak
the open goal
A cited forum thread or a stale roundup means the machine could not find an authoritative page. That is a question with no owner, and owning it is cheap.
What it can actually extract
specifics, structured
Pages that get cited tend to state things plainly — numbers, prices, limits, dates, comparisons in a table. Prose that buries the fact in a paragraph gets read past.
How fresh it prefers
half under thirteen weeks
Half of cited content is under thirteen weeks old (published by Profound). Freshness is not a ranking trick here; it is part of what makes a source look like the current answer.
The machine is telling you what to publish. Most teams read it as a scoreboard.
It is also telling you what your buyers are being told, which is the other half of the value. The sources that shape your category's answers are a reading list: the outlets, the analysts, the forums and the docs that the machine treats as the current account of your market. Read them and you know the frame your buyer arrives with, before they ever reach your site.
One caution about cadence. This set moves: at the domain level, 40 to 60% of the cited sources turn over month to month — 59.3% for Google AI Overviews, 54.1% for ChatGPT, 53.4% for Copilot, 40.5% for Perplexity, measured across roughly 80,000 prompts per platform a month apart (published by Profound). So the pattern you extract is a standing read, not a finding you file. And a page whose citations are fading is a refresh brief with a deadline on it, weeks before the traffic moves.
One question, thirteen weeks apart. Two sources survive, three fall away, three arrive — which is why the pattern is worth re-reading rather than filing.
06
“41% of what?” is the question that kills most visibility scores.
Somebody shows you a slide. AI visibility: 41%. Ask what it is 41% of.
If the answer is not a specific list of questions asked a specific number of times, you are looking at decoration. Three things make a percentage real, and none of them require a statistics degree.
How many answers is it counting?
Ten questions and a hundred questions produce equally confident charts and wildly different truth. At ten answers, reality could sit thirty points either side of your number. At four hundred, five.
The band the truth could sit in, by how many answers you collected. Repeated runs of similar questions lean on each other, so real precision is worse than the count suggests.
Which questions is it counting?
Fifty prompts run three times a day gives you a tighter estimate of those fifty questions. It tells you nothing about the fifty-first. Most teams buy repetition when what they need is coverage, because repetition is the thing dashboards make easy.
Would it say the same thing tomorrow?
Run the identical set twice before you trust any movement in it. The gap between two identical runs is your noise floor. Anything smaller than that is not news, and reporting it as news is how you lose credibility with your own executives.
Ten well-chosen questions make a great canary for the ten situations that pay your bills. They make a terrible estimate of a market. The fix is design, not volume.
Your CEO still wants one number. Give them one that survives contact.
- “We appear in 41% of 320 answers across four assistants for our top twenty buying questions, give or take five points, measured weekly since March.” That fits on a slide, it holds when somebody pushes on it, and next quarter it compares to itself. “Our AI visibility is 41%” does none of that. Same figure. One of them is defensible.
What we cannot do yet, since we are the ones demanding this standard. We cannot tell you how often real buyers ask a question — that needs a licensed panel of real prompts and we do not have one. Our runs go through provider APIs, which are close to but not identical to the consumer app. We cover four assistants. We read server logs from exactly one host, Vercel, which is not yours unless you happen to run there.
07
Attribution is the hard part. Almost nothing arrives labeled.
People do come. About two and a half percent of them tell you so.
In Profound's panel study pairing AI conversations with browsing behavior, people who saw an AI mention of a brand visited that brand's site at 1.5 to 2.5 times the baseline rate over the next seven days. Only about 2.5% of those visits carried a trackable AI referral parameter. Treat that as one vendor's proprietary estimate rather than a reproduced finding, and note what it counts: visits, not revenue.
Read that number as your evidence and you will conclude this channel does nothing, while being wrong by a factor of forty. The clean attribution model you want does not exist here. Nobody has one. Anyone selling you one is selling you a number.
If a referral parameter is your only evidence, you will conclude AI does nothing.
So stop trying to prove it with one number and read four at once. Each one is circumstantial. Together they are a case, and a case is what you actually have.
Two and a half percent of it is labeled. The four rows are how you read the rest.
Arrivals from the pages that cite you
referrer, not tag
The assistant will not label the visit. The cited page often does. When your traffic from a source climbs in the same month that source starts turning up in answers, that is not proof. It is corroboration, and you can have it tomorrow.
Arrivals onto the pages that get cited
landing, not source
You already know which of your URLs the machine quotes. Watch those specifically. A lift concentrated on exactly the cited pages is a different finding from the same lift spread evenly across the site.
How often the crawlers come back
cadence, and its direction
Crawlers refetch what they intend to keep using. Rising refetch frequency on a page is the machine telling you it now depends on that page. Falling frequency is a warning weeks before traffic shows you anything.
Agent or human on the other end
who is doing the fetching
An assistant sending an agent to read your page and a person opening it are both worth having, and they mean different things. Separate them in the logs or you average two stories into one flat line.
Instrumenting this takes both ends. A browser tag catches the visits that do carry a referral and what people do after they land. Server-log ingestion catches the crawlers and agents, which never run JavaScript and never appear in analytics at all. Miss the second one and you are not measuring a channel, you are measuring the part of it that happens to run your tracking script.
08
Start with your brand, then widen the questions.
Do not start with a hundred questions. Start with the ones where you already know the right answer.
Week one: ask about you, by name. What is [brand]? Is [brand] any good? Who is [brand] for? What does [brand] cost? Is [brand] legitimate? Ten to fifteen questions is plenty.
You know these answers. So every error is unambiguous: a wrong founding fact, a product you killed, a price from two years ago, a competitor's feature credited to you. Fix those first. They are the cheapest wins available and they compound, because the sources feeding those errors feed everything else.
You can run week one by hand this afternoon. Four assistants, fifteen questions, paste every answer and its links into a sheet. Sixty answers — enough to find broken facts, nowhere near enough to estimate a market.
| Stage | Question shape | Scale to start | What it reveals |
|---|---|---|---|
| Brand direct | “What is [brand]” · “Is [brand] any good” · “What does [brand] cost” | 10–15 questions | Whether they know you and whether the facts are right. First, because you can grade every answer. |
| Category and solution | “Best [category] for [segment]” · “Top tools for [job]” | 20–40 questions | Whether you show up when nobody names you. The biggest pool of real buyer questions, and where most brands find out they are invisible. |
| Comparative | “[You] vs [competitor]” · “Alternatives to [competitor]” | 2–3 per rival | How you stand against named rivals, and who gets recommended instead. Third, because it needs the competitive set stage two reveals — not the one in your deck. |
| Pricing and commercial | “How much does [category] cost” · “Is [brand] worth it” | 10–15 questions | The money narrative, wrong more often than any other. Buyers hit it late and it converts. |
| Objection and risk | “Problems with [brand]” · “Is [brand] secure” | 10–15 questions | What surfaces when a buyer turns skeptical. Hardest to move, most damaging to ignore, usually sourced from the oldest and angriest pages on the internet. |
Then add context, not more questions. Ask the same questions as a fifty-person company and as an enterprise. In every market you sell to. In the language your buyers actually use — not in English about a foreign market. Where the answers split by context is usually more useful than the headline number.
No research team? Cut in this order.
- Drop repetition before coverage.
- Drop personas before stages.
- Drop stages five and four before two and three.
- Never drop keeping the answers. One person can run stages one through three across four assistants with a weekly rerun. That is a real program.
What to produce. One page. The headline number with the questions and the count behind it. The three questions that moved most since the last run. Sources newly cited and newly gone. The one change you are making before the next run. Every line traceable to an answer you kept.
What to freeze. Your brand and category questions become a standing panel so the trend stays comparable. Rotate everything else. A panel you keep changing cannot show you change.
Design the study. Keep the evidence. Run it again.
AI Visibility runs the method in this report: you choose the questions, the personas and the assistants, every answer and citation is kept so you can reread them, and every number arrives with the count behind it.
