Quick Answer: What makes AI choose one business over another is not one judgement but two gates. A retrieval layer builds a candidate set, then the model picks from inside it. Most small businesses lose at retrieval, before any judgement occurs. No outsider can inspect the exact selection function these systems run.
"The question arrives in some version of why does it pick them and not me, and the truthful first move is to admit that nobody outside those companies can see the selection function. What we can see is inputs and outputs, measured at volume. The finding that repeats when I audit a business is that the owner assumes they lost a judgement call. Usually they were never in the running at all: the retrieval step never handed the model their name. Those are two different problems and they need two different repairs, and telling them apart is most of the work."
Matt Griffin, Founder, Formative DigitalThis page tries to be precise about a subject that invites confident nonsense. Everything asserted below is either measured in a published audit, stated by a platform in its own documentation, or explicitly flagged as inference. Where the honest answer is that nobody knows, that is what you will read.
One limitation belongs at the top rather than buried in the footnotes. Four of the ten sources cited here come from the same research group, and direct replication of that specific commercial-recommendation work is still thin. The mechanisms underneath it are on firmer ground: the popularity gradient and the position effects inside a retrieved context both appear in peer-reviewed work from unrelated labs, cited below, and the retrieval pipeline itself is documented by the platforms that run it. What remains unreplicated is the brand-level numbers. Read those as strong evidence about the audited conditions, not as physical constants.
What is actually known here, and what is guesswork
Very little about AI business selection is known with certainty, and a surprising amount is measurable anyway. The distinction matters because the two get blurred constantly in marketing copy. Nobody outside these companies has read the ranking code. Plenty of people have run tens of thousands of prompts through the production systems and counted what came out, which supports firm statements about aggregate behaviour and no statement at all about any individual answer.
Here is the honest accounting, sorted by how much weight each claim can bear.
| Claim | Evidence status | How much weight it holds |
|---|---|---|
| Retrieval happens before generation | Documented by Google | Firm. Google describes using retrieval-augmented generation to surface content from its Search index before composing an answer. |
| Smaller and regional businesses are frequently never retrieved | Measured externally | Firm for the audited conditions. Not proof about your specific business or category. |
| Being retrieved is not enough to be recommended | Measured externally | Firm. Retrieval and recommendation rates diverge sharply in the audit data. |
| Agreement across independent sources helps | Partly inferred | Reasonable inference. Supported by evidence that some models draw heavily on training priors, which reflect how consistently a brand appears across the corpus. |
| Schema markup causes AI to prefer you | Contradicted as stated | Weak. Google states structured data is not required for generative AI search and no special markup exists for it. |
| There is a stable AI ranking position to climb | Contradicted | Weak. Rewording, asker identity, and reruns all shift the output set measurably. |
| The exact selection function | Unknowable from outside | None. Closed weights, private retrieval indices, no published per-query logs. |
Anyone selling you certainty about row seven is selling you something. Rows one through three are where the useful work lives.
Selection resolves in two stages, not one
An AI recommendation is produced by at least two separable steps, and a business can fail either one for completely unrelated reasons. Step one is retrieval: a search layer runs behind the conversation and returns a set of documents. Step two is generation: the model reads those documents, weighs them against what it already holds from training, and writes an answer that names a few businesses. A company absent from step one cannot be chosen in step two no matter how good it is.
This is not a metaphor invented for a blog post. Google's own documentation on AI features describes the pipeline directly, using retrieval-augmented generation to pull relevant pages from its index and then reviewing that material to compose a response. Our explainer on where AI engines actually read from takes the retrieval side apart source by source.
The practical consequence is that most advice about AI visibility is answering the wrong question. Guidance about being persuasive, differentiated, or well reviewed addresses step two. If your problem is step one, none of it applies yet.
Gate one: does anything retrieve you at all?
For small and regional businesses, retrieval is the binding constraint, and the measured failure rate is severe. The 37,000-run audit published by Jack, Lehman, Maloney and Xu in May 2026 tested 215 commercially framed prompts across 19 sectors against a 533-brand reference catalogue, sorted into five prominence tiers. Long-tail specialists surfaced on 8% of relevant queries, and 48% of them never appeared in any of the roughly 37,000 runs. Regional players surfaced on 3% of queries, and 52% never appeared at all.
Read those two sentences again if you run a local business, because they describe you. Roughly half the regional brands in a curated catalogue of real companies were, for practical purposes, invisible to production AI assistants across an enormous sample.
That prominence gradient is older than the assistant era and does not rest on this research group at all. A 2023 ACL paper probing what language models actually retain found that they struggle with less popular factual knowledge, that scaling models up mainly improves recall of popular knowledge and fails to appreciably improve the tail, and that retrieval augmentation helps most precisely where the model's own memory is weakest. Set against the audit numbers, that is the same shape measured from a different direction. The smaller you are, the more completely your visibility rests on retrieval working, and the less you can count on the model simply knowing you.
| Prominence tier | Surfaced on a given query | Recommended once surfaced | Where it breaks |
|---|---|---|---|
| L1 category leaders | 77% | 25% to 41% | Retrieved, then passed over |
| L2 established challengers | Not separately reported | 37% to 52% | Providers disagree with each other |
| L3 mid-market | 23% | Not separately reported | Partial failure at every stage |
| L4 long-tail specialists | 8% | Not separately reported | Never retrieved: 48% never surfaced |
| L5 regional players | 3% | 12.8% | Never retrieved: 52% never surfaced |
Two cells say "not separately reported" because the published summary does not break those figures out. Filling them with plausible numbers would have made a tidier table and a dishonest one.
There is a further detail in that audit which deserves more attention than it has received. The researchers compared what a model's native web-search tool retrieved against what a dedicated neural retriever retrieved, and agreement between the two fell steadily as brands got smaller: roughly 0.83 at the category-leader tier, down to 0.50 at the regional tier. Between 56% and 60% of the smallest brands that surfaced at all were found by only the external retriever. The inference I draw from that, and it is an inference rather than a finding, is that built-in web search is systematically weaker at locating small businesses, so part of your visibility depends on retrieval plumbing you cannot see, cannot address, and did not choose. Our teardown of which sources Perplexity actually pulls for local queries shows how differently one engine's substrate behaves on exactly this kind of query.
Gate two: retrieved, and passed over anyway
Getting retrieved buys you a seat in the candidate set and nothing more. In the same audit, category leaders were retrieved on 77% of relevant queries but converted that into an actual recommendation only 25% to 41% of the time. Established challengers converted better, at 37% to 52%, despite lower visibility overall. Being the most findable company in a category is therefore not the same as being the most recommended one, and the gap between those two numbers is the entire second discipline.
What separates a retrieved-and-named business from a retrieved-and-skipped one is harder to establish from outside, and this is where I would encourage the most skepticism toward anyone who sounds certain. The strongest available evidence on the generation side is still the Princeton and Cornell study that coined the term generative engine optimization. Aggarwal and colleagues tested nine content modifications and reported visibility gains of up to 40% from adding statistics, quotations, and cited sources, while classic keyword repetition produced little to nothing. Their result is about content characteristics in generated answers rather than about business recommendation specifically, so treat it as adjacent evidence that transfers plausibly, not as a measurement of the thing this page is about. We work through what does and does not transfer in our analysis of how LLMs behave when choosing citations.
One qualification belongs here, because it cuts against the tidy version of gate two. Not all of the distance between being retrieved and being named is about the quality of your material. The peer-reviewed long-context literature reports a positional effect: models answer measurably better when the passage that matters sits near the start or the tail of the input context than when equivalent material sits somewhere in between. Some share of retrieved-and-skipped is therefore a function of where your document landed in a context window, which is not a thing you can write your way out of. How large that share is in production commercial answers I do not know, and nobody outside those companies does either.
If you want the short version of the second gate: once you are in the room, the machine is looking for material that can be quoted with a straight face. Vague positioning survives a human skim and dies in an extractive summary. Just do not read that as the whole explanation.
How much of the answer comes from retrieval, and how much from memory?
A meaningful share of AI recommendations are not grounded in anything the system just retrieved, which changes what "source authority" means in practice. The persona-conditioning audit published in May 2026 measured this directly and found that between 43% and 52% of one model family's brand recommendations had no supporting evidence in the retrieval layer, against 8% to 29% for the comparison configurations. Those recommendations came from the model's training priors: what it absorbed about the commercial world during pre-training.
Why that number reframes source agreement
If part of the answer comes from training rather than retrieval, then how broadly and how consistently your business was written about across the open web, years before the question was asked, is doing some of the work. That is the defensible version of the industry's claim that AI rewards consensus. The claim is usually asserted with no mechanism attached. The mechanism is parametric memory, and the size of its contribution varies by provider.
This also explains a pattern owners find maddening: fixing your website changes nothing for a while, and then changes something suddenly. Retrieval can pick up a new page within days. Training priors do not update until the next model does. You are writing to two clocks running at different speeds, and only one of them is visible to you.
OpenAI's crawler documentation makes those two clocks unusually literal. It runs separate crawlers for search and for training, with OAI-SearchBot used to surface websites in ChatGPT's search features and GPTBot used for foundation model training, and the two are controlled independently. A site can be legible to one and absent from the other. That makes one dull check worth running before any content work: confirm your robots.txt is not quietly declining the crawler that feeds the search side.
The honest qualifier: the 43% to 52% figure comes from one audit of specific model configurations on commercial prompts. It should not be generalized into a claim that half of all AI answers are ungrounded.
Entity clarity, and the part of it the industry oversells
Entity clarity matters for a narrow, mechanical reason, and the wider claims made about it are not supported. The narrow reason is disambiguation: these systems have to decide whether the business named in one document is the same business named in another. When your name, address, and phone number are recorded inconsistently across your site, your Google Business Profile, and the directories that syndicate from both, you are asking a matching process to guess, and a guess that resolves the wrong way splits your evidence across two half-built records.
What is not supported is the schema pitch. Google's documentation states outright that structured data is not required for generative AI search and that no special schema.org markup exists for it. The same documentation says Google Search ignores llms.txt files, and that there is no need to break content into small pieces for AI. Vendors still bill for all three. In the 12 Vectors framework this work sits under Vector 2, Anchor, and we scope it as insurance against misidentification rather than as a lever on preference, because that is what the documentation supports.
The version of this that does hold up
Consistency is cheap and its failure mode is expensive. A business whose details disagree across sources is not penalized by a rule somewhere; it simply presents a weaker, more contradictory record for any process trying to assemble a picture of it. That is a mundane claim, it costs nothing to act on, and it does not require believing anything unverifiable about how the models rank.
If you would rather see your own record before deciding whether any of this applies, we will run the checks and send you what we find, at no charge.
The query being answered is not the query that was typed
By the time a model composes an answer, the original question has been expanded, rewritten, and shaped by context the user never sees. Google documents the first part of this as query fan-out, describing a set of concurrent related queries the model generates to fetch additional results. Anthropic documents the same shape from the other side of the industry: its web search tool can run searches repeatedly within a single request, with simple factual questions typically using one to three searches and comparative research using ten or more. A buyer asking which supplier is best in their city is issuing comparative research. You are not competing for one question. You are competing across a spray of machine-written sub-questions you will never read.
Two 2026 audits measured how unstable that makes the output. The paraphrase-brittleness study found that rewording a query while holding the intent constant changed the recommendation set more than simply rerunning the identical query did, which means phrasing sensitivity exceeds the system's own baseline noise. The persona-conditioning study found that changing who appears to be asking reduced recommendation-set overlap by 0.12 to 0.20 on a Jaccard measure, with category leaders holding roughly 80% consistency across personas while mid-market brands swapped up to 75% of their recommendations.
Sit with that last figure. A mid-market business can be recommended to one buyer and omitted for the next, from the same model, on the same day, for the same underlying need, because the asker described themselves differently. We track this instability continuously in our AI answer variance work, and the practical lesson is measurement discipline: one spot-check of one phrasing tells you almost nothing. Sample the question several ways, several times, before concluding anything.
What cannot be known from outside these systems
Some questions about AI business selection are permanently closed to outside investigation, and it is worth naming them precisely. The model weights are private. The retrieval indices are private. No provider publishes per-query logs showing which documents were fetched and which were used. There is no ranking report, no equivalent of Search Console for a generated answer, and no appeal. Anyone claiming to know the weighting a system applies to reviews versus mentions versus schema is reporting a hunch in the grammar of a fact.
The systems are also non-stationary. Model versions change, indices refresh, and behaviour drifts, so a measurement taken in March describes March. Treat any single reading as one sample from a moving distribution.
There is, however, a genuinely useful finding sitting inside all this uncertainty, and it is the reason this page is not simply a list of things we cannot know. A June 2026 study by the same research group compared providers directly and reported that while the recommendations themselves diverge substantially, the diagnoses converge: when a brand is missing, the different systems tend to be missing it for the same underlying reasons. Recommendations differ. Failure modes rhyme.
That is the load-bearing conclusion here. Chasing a position in any one engine is chasing a number that moves when someone rephrases a sentence. Repairing the failure mode is durable, because the failure mode is shared across engines that otherwise agree on very little. The rest of our primary-source teardowns live in the research hub if you want the underlying papers rather than our reading of them.
Which gate are you failing? Run this check
You can identify your own failure gate in about twenty minutes with no tools and no budget. The point is not to produce a score. It is to tell a retrieval problem apart from a selection problem, because the repairs share nothing.
The five prompts
- 1. The category question. Ask an assistant for the best provider of what you sell, in your city, without naming your business. Record every company it names.
- 2. The direct question. Ask what it knows about your business by name. Record whether it knows you exist and whether the details are correct.
- 3. The grounding check. Run the category question again with web browsing explicitly turned on, then again with it off if the interface allows. Compare the two lists.
- 4. The rephrase. Ask the category question three more ways, changing the wording but not the intent. Note whether the named set holds steady.
- 5. The persona swap. Ask it once as a homeowner, once as a commercial buyer, or whichever two customer types you actually serve.
Now read the pattern rather than the individual answers.
Reading your results
- Absent from prompt 1, accurate on prompt 2: a retrieval problem. The system holds a record of you but nothing surfaces it for category queries. Work on being findable and quotable on the questions buyers actually ask, not on persuasion.
- Absent from prompt 1, wrong or blank on prompt 2: an entity problem, the deeper of the two. There is no clean record to retrieve. Start with consistency across every listing you control.
- Present with browsing on, absent with it off: you exist in retrieval and not in the model's priors. Expected for most small businesses, and a reason to keep earning independent mentions.
- Present in some phrasings, absent in others: you are marginal in the candidate set. You are being reached and then dropped, which puts you at gate two.
- Present consistently across all five: your problem is differentiation, not visibility. Recall the audit finding that leaders convert retrieval to recommendation only a quarter to two-fifths of the time.
Have us run the five prompts for you
We will run the check across ChatGPT, Perplexity, Gemini, and Google AI Overviews, sampling each question several ways rather than once. You get the raw prompts, the answers they returned, and a plain statement of which gate is closed. Usually inside a business day, and it costs nothing.
What the evidence actually supports doing
The defensible programme is short, and it follows the gates rather than the tactics. If retrieval is your constraint, the work is publishing pages that answer the specific questions buyers ask, in language a stranger could quote without editing, and earning mentions on sources that sit outside your own domain, since those are what an external retriever can find. If selection is your constraint, the work is evidence: dated figures, named people, real specifics that survive being compressed into three sentences by a machine. If entity clarity is your constraint, the work is janitorial and should be finished before anything else starts.
What none of it supports is a promise. This maps to Vector 1, Diagnose, and diagnosis is the honest deliverable: it tells you which gate is closed, not what revenue follows from opening it. Outcomes move with vertical, market density, and how long a domain has existed, and the compounding is slow at first. In one anonymized engagement with an Ontario shipping container dealer, the pattern over a 16-month window was single-digit daily clicks through most of the first year, then seventy-click days with several thousand daily impressions by early summer 2026 (Google Search Console, 16-month window, shared with permission). The flat stretch was not failure; it was what accumulation looks like before it is visible. If you want the mechanics behind that kind of build, our services page sets out how we structure the work.
Frequently Asked Questions
Can anyone actually see how AI chooses which business to recommend?
Not from outside. The model weights, the retrieval indices, and the ranking code are private to OpenAI, Google, Anthropic, and Perplexity, and none of them publish per-query logs. What outsiders can do is run the same prompts thousands of times and measure which brands appear, which is the method behind the 2026 audits cited on this page. That produces reliable statements about behaviour in aggregate and no reliable statement about why a specific answer named a specific company on a specific day.
Why does ChatGPT recommend my competitor but not me?
Usually because your competitor was in the candidate set and you were not, rather than because the model weighed you both and preferred them. In the 37,000-run audit by Jack and colleagues, regional players surfaced on only 3% of relevant queries and 52% of them never surfaced at all. Before assuming you lost a comparison, confirm you were in the comparison. Asking the assistant directly about your business by name tells you whether it holds a record of you at all.
Does schema markup make AI pick my business?
No, and Google says so plainly: structured data is not required for generative AI search, and there is no special schema.org markup you need to add. Schema is worth implementing for a narrower reason. It states your name, location, and phone number in a form that does not require interpretation, which reduces the chance a system merges you with a similarly named business. Treat it as disambiguation, not persuasion.
Why do I appear in one AI assistant but not another?
Because each assistant sits on a different retrieval substrate, and those substrates disagree most sharply about smaller businesses. The 2026 audit measured agreement between a native web-search tool and a neural retriever falling from 0.83 at the category-leader tier to 0.50 at the regional tier, with most low-prominence brands that surfaced at all being found by only one of the two. Cross-engine inconsistency is the expected result of that plumbing, not evidence that one engine dislikes you.
How long does it take to become visible to AI assistants?
Plan in quarters rather than weeks, and expect the curve to be flat before it is steep. Assistants recrawl on their own schedule and rebuild their picture of a business gradually, so changes you make today are read at some unknown later date. Timelines shift with vertical, market density, and how long the domain has existed. Anyone quoting a date for a first ChatGPT citation is quoting a number the mechanics do not support.
Sources
- Jack, W., Lehman, N., Maloney, K., & Xu, S. (2026, May 21). Prominence-Stratified Failure Modes in Retrieval-Augmented Commercial Recommendation: A 37,000-Run Audit. arXiv preprint arXiv:2605.27439. Link
- Jack, W., Lehman, N., Maloney, K., & Xu, S. (2026, May 28). Persona Conditioning of Brand Recommendations in Retrieval-Augmented Commercial Chat: A Prominence-Stratified Cross-Provider Audit. arXiv preprint arXiv:2605.30207. Link
- Jack, W., Lehman, N., Maloney, K., & Xu, S. (2026, May 28). Paraphrase Brittleness in Production Retrieval-Augmented Commercial Recommendation: Reproducibility Below the Rerun-Stability Baseline. arXiv preprint arXiv:2605.27440. Link
- Jack, W., Lehman, N., Maloney, K., & Xu, S. (2026, June 26). Divergent Recommendations, Convergent Diagnoses: Cross-Provider Failure-Mode Convergence in AI Commercial Recommendation. arXiv preprint arXiv:2606.26116. Link
- Google. AI features and your website: an AI optimization guide. Google Search Central documentation. Link
- Anthropic. Web search tool. Claude Platform documentation. Link
- OpenAI. Crawlers and user agents: OAI-SearchBot, GPTBot, and ChatGPT-User. OpenAI developer documentation. Link
- Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., & Deshpande, A. (2023). GEO: Generative Engine Optimization. arXiv preprint arXiv:2311.09735, KDD '24. Link
- Mallen, A., Asai, A., Zhong, V., Das, R., Khashabi, D., & Hajishirzi, H. (2023). When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Link
- Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2024). Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics, volume 12. Link
Find out which gate is closed
Formative Digital, Brantford, Ontario
Guessing at this is expensive, because the repair for a retrieval problem and the repair for a selection problem have almost nothing in common. We will tell you which one you have, and show you the prompts and answers the conclusion rests on.
Written by Matt Griffin, founder of Formative Digital, Brantford, Ontario. Published 2026-07-20.