Obenan BriefingFlagship Briefing
Making the list is not the same as being chosen
Being named by AI, named first and described correctly are separate problems. A seven-day study of two Bali submarkets helps measure each. No ranking recipe.
The one line
Separate problems, separate fixes, and one answer cannot tell you which one you have. A reply is a single draw from one system, one question, one place, one language and one moment.
The evidence
85.6%at least
of those venues were never recommended by any of the four systems.
SourceFull arXiv HTML paperFull arXiv paper, results. Checked 30 August 2026 and again on 2 September 2026.
- Published
- September 5, 2026
- Format
- Flagship Briefing
- Sources
- 2 primary, public
- Evidence basis
- Preregistered observational preprint, two Bali submarkets
In two Bali submarkets, most venues were never named by any of the four AI systems studied, and some answers still pointed to businesses that had closed. Those are separate problems, and neither is fixed by a ranking trick.
01The scene
A missing answer leaves no trace
Hypothetical sceneAn imagined situation used to explain the mechanism. No real business, customer or answer is described.
Picture a visitor two streets away, deciding where to eat. Instead of scrolling a map, they ask an AI assistant and read a short reply naming four places. If your venue is not one of them, nothing happens that you can see: no lost click, no abandoned booking, no complaint. The scene is hypothetical, but the shift behind it can be measured. An answer is a short list with room for a handful of names.
02The evidence
Three failures hide behind one short answer
If your venue is absent from an answer, or present but described wrongly, at least three different things could have gone wrong:
- 01
Entry
Was the venue named at all?
- 02
Rank
If it was named, where did it sit among the other names?
- 03
Freshness
Were the facts attached to it still true?
Separate problems, separate fixes, and one answer cannot tell you which one you have. A reply is a single draw from one system, one question, one place, one language and one moment.
03The evidence
Most venues were never named at all
A public preprint, a paper released before peer review, built a complete list of 4,776 cafes, restaurants and bars in greater Canggu and Ubud, two submarkets in Bali, then asked four AI systems where to eat and drink there. Within that market, and measured that way, at least 85.6% of those 4,776 venues were never recommended by any of the four systems. Among venues carrying at least 50 ratings, the share never recommended was 72.6%.
Both figures describe this registry, in these two submarkets, in English, through the search-grounded interfaces of these four systems, meaning developer access that answers with live web search rather than the consumer apps. They are not a benchmark for another city, language, category or app.
Most venues were never named
- cafes, restaurants and bars in the registry for greater Canggu and Ubud.
- 4,776
- of those venues were never recommended by any of the four systems.
- 85.6%at least
- never recommended, among venues carrying at least 50 ratings.
- 72.6%
This registry, these two submarkets, English queries, four systems reached through search-grounded interfaces, one seven-day window. Not a benchmark for another city, language, category or app.
SourceFull arXiv HTML paperFull arXiv paper, results. Checked 30 August 2026 and again on 2 September 2026.
04The evidence
Getting named and getting to the top are not one skill
In this study, what went with entering an answer was documentation: having its own website, review volume, a listed price, mentions elsewhere on the web. Star rating showed no association at that first step. Among venues that had already entered an answer, rating was associated with appearing in first position, and review volume still mattered there too.
Read narrowly: entry and rank behaved like two different outcomes and need measuring as two different things. None of these associations shows that changing a field causes a recommendation.
Entering an answer and reaching first place moved with different things
Odds ratios from the fitted model. Above 1 means higher odds inside that model. Association, not cause. Among venues that had already entered an answer, rating was associated with appearing in first position, and review volume still mattered there too.
Entering an answer at all
- Own website1.92
- Review volume1.64
- Listed price1.54
- Third-party web mentions1.44
- Star ratingread as no association at this step0.89
First position, among venues already named
- Star rating1.17
1.0 means no association
SourceFull arXiv HTML paperFull arXiv paper, entry and rank models. The author reads rating at 0.89 as no association for entry. Checked 2 September 2026.
05The evidence
Stale facts showed up more often than invented ones
Across 12,439 valid venue mentions, the systems recommended permanently closed venues 93 times, involving 14 venues confirmed as closed. Separately, one venue name was still judged likely invented after checking, appearing in 10 mentions, or 0.08% of those valid mentions.
The author's reading is that stale facts, not invented ones, were the practical failure in this collection. Read carefully, it is a comparison of two counts inside one seven-day window, not a ranking of harm. Both counts are small, and neither one was followed any further. This study observed what the systems said, not what anyone did next, so it cannot show what either kind of error costs a business.
The implication is modest. Keeping hours, address, prices, availability and closure status correct is worth doing on its own terms. This study does not show that correcting them causes a venue to enter an answer, rise to first place, or bring anyone through the door.
Stale facts against invented ones
valid venue mentions across 2,208 runs in the confirmatory wave
12,439
93
mentions of permanently closed venues14 venues confirmed as closed
10
mentions of one venue name judged likely invented0.08% of valid mentions
A comparison of two counts inside one seven-day window, not a ranking of harm. Neither count was followed any further, so the study cannot show what either kind of error costs a business.
SourceFull arXiv HTML paperFull arXiv paper, validation and error analysis. Checked 2 September 2026.
06The boundary
What this evidence does not settle
The study is observational. It cannot establish that a website, a review count, a price field, a rating or a factual correction causes entry or rank. It covers two Bali submarkets rather than Bali or any other market, English queries only, and four configured systems reached through search-grounded interfaces rather than the consumer apps people open, across one seven-day collection period. A recommendation is also not a visit, a booking, an order or a paying customer. Nothing here measures whether anyone walked in.
07What to do
What an operator can usefully do next
- 01Measure entry, rank and freshness separately. One combined visibility score hides which of them moved.
- 02Record the instrument: system, interface, configuration, question, language, place, date, the answer and the order of names.
- 03Repeat before concluding. Use more than one run, wording and system, and call small samples snapshots.
- 04Keep public facts current because customers use them, not because this evidence makes them a ranking lever.
- 05Prove customer outcomes somewhere else. Appearing in an answer is not evidence of a sale.
08In more detail
The evidence behind this, in more detail
The paper is a public preprint and has not been peer reviewed. Its confirmatory wave, the main preregistered collection, covered 2,208 runs over seven days and produced 12,439 valid venue mentions, using 96 English queries, each phrased as if from a particular type of visitor, put to ChatGPT, Claude, Gemini and Perplexity through search-grounded APIs rather than their consumer applications. The design, hypotheses and analysis plan were preregistered publicly before the confirmatory data were collected, which makes it easier to separate planned tests from explanations chosen after the results were seen.
The entry associations are reported as odds ratios from a fitted model: review volume 1.64, a website of its own 1.92, a listed price 1.54, third-party web mentions 1.44, and rating 0.89, which the author reads as no association at that step. Within answers, rating was associated with first position at 1.17. An odds ratio above 1 means higher odds inside that model. It is not a percentage-point gain, and it is not evidence of cause.
Repeatability was measured too, and the numbers are easy to misread. Asking the same system the same question again returned a substantially different set of venues: mean Jaccard similarity between repeated runs ran from 0.22 for Claude to 0.45 for Gemini. Jaccard similarity is the share of venues two lists have in common out of every venue either list names, so 1 means the two lists are identical and 0 means they share nothing. A mean of 0.30 therefore says the typical pair of runs overlapped by roughly a third of the names involved. It does not say the same answer came back 30% of the time. Measured the same way, overlap among the top 20 venues across different systems ran from 0.33 to 0.54. A retest two weeks later produced turnover comparable to reruns inside the original period, which the author attributes to sampling variation rather than change over time.
The instrument had limits of its own. Repetition counts differed by system, with Perplexity run 10 times per query, ChatGPT and Gemini 5 times, and Claude 3 times. Claude's route was capped at two searches per run. After fixes, 91.5% of runs passed the strict validity check, and matching the names in answers to real venues was roughly 98% to 99% accurate rather than perfect.
Norly funded and conducted the study, and Norly sells review-management and AI-visibility tools to hospitality businesses. The author discloses that competing interest and describes the preregistration and validation steps intended to insulate the work. Those steps do not make the commercial conflict disappear. Obenan works in the same category, so this evidence, including our reading of it, deserves the same care.
First run
Second run
- Named in both runs
- Named in one run only
2 / 6= 0.33
- Same system, same question, asked againmean Jaccard similarity, from Claude to Gemini
- 0.22 to 0.45
- Top 20 venues across different systemsmeasured the same way
- 0.33 to 0.54
SourceFull arXiv HTML paperFull arXiv paper, repeatability results. Checked 2 September 2026.
09Sources
Sources
Primary, public sources only. Publication dates and the dates we checked them are kept separate so the chronology stays inspectable.
- 01arXiv abstract and submission record, submitted 7 August 2026; checked 30 August 2026.
- 02Full arXiv HTML paper, methods, preregistration, results, validation, limitations, funding, and competing-interest disclosure; checked 30 August 2026 and again on 2 September 2026.