Source-linked AI summary
Does Rank Still Matter? Position Bias When AI Agents Shop on Our Behalf
Davood Wadi, Yu Ma
TL;DR
The paper asks whether search rankings retain their value when AI agents can read entire results pages despite scarce human attention making ordered placement influential. Using hotel-listing sessions across four major LLMs and human field data, it finds that rank still predicts inspection, but much more weakly than for humans and without a monotonic pattern.
Problem
The paper examines whether rankings remain valuable when AI agents can ingest entire results pages, unlike humans whose attention is scarce and ordered.
Method
The study observes AI agents inspecting and booking hotel listings through tool calls across 2,000 sessions involving four major LLMs from two providers, alongside a human field benchmark.
Results
5.83 inspections per session versus 1.12 for humans, while every AI agent booked; rank predicted inspection for every tested LLM, four to ten times more weakly than for humans.
Takeaways & Limitations
Rank remains relevant for AI-agent inspection, but its effect is weaker and non-monotonic rather than declining steadily across positions.
Takeaways & Limitations
The study fixed the query and destination in its hotel-booking setting, constraining the scope of its evidence.
Abstract
from arXiv · showhide
Search rankings are valuable because human attention is scarce and sequential. Higher-placed alternatives are easier to find, so they are examined and bought more often. Consumers are now delegating search to AI agents that can ingest an entire results page at once. Randomizing the order of one hundred hotel listings across 5,000 AI agent sessions, we compare four large language models against human field data. AI agents search more deeply than humans and never decline to buy. Position still predicts which listings are inspected, but weakly and non-monotonically: the middle of a results page has the lowest probability of inspection, not the bottom. Position reaches the choice stage for some models and not others, a heterogeneity that tracks neither provider nor capability. All models nonetheless converge on the same undominated listing. For agentic search, the attributes displayed on a results page matter more than placement within it.
1. Introduction
AI agents can read entire ranked results pages without human scrolling costs, challenging whether rank retains its influence. The study finds that rank still predicts inspection weakly, but does not generally determine booking because agents inspect more broadly and converge on the same strong option.
- Motivation: AI agents receive entire results pages at once and pay no marginal reading cost for lower-ranked listings.This removes the sequential attention constraint that traditionally makes top placement valuable for human search.
- Study design: 2,000 sessions across four major LLMs from two providers compare agent hotel search with a human field benchmark.Follow-up experiments vary reasoning effort and prompt wording.
- Findings: 1.63 to 5.83 inspections per session for AI agents versus 1.12 for humans, while every AI agent booked a hotel.Human consumers declined to book about one-third of the time; the booking pattern remained after removing recommendation instructions.
- Findings: Rank predicted inspection for every tested LLM, four to ten times more weakly than for humans, with a non-monotonic pattern.Except for Gemini 3.1 Pro, inspection declined from the top to around ranks 68 to 74 before rising again, making the bottom better than the middle.
- Findings: Chosen hotels averaged ranks 44.0 to 49.7 versus the position-neutral value of 50.5, while one undominated listing captured 78.2% of bookings.Thus, inspection effects did not systematically carry through to booking choice.
- Implications: Reasoning effort, rather than LLM identity, governed the remaining exposure to position, and greater effort reduced the position effect to insignificance.The authors argue that rank can affect inspection without affecting final choice for AI agents.
- Implications: Managers should adapt SEO strategies by emphasizing organic signals such as review scores, which remain primary heuristics and trust indicators when rankings shift.Top placement does not systematically determine choice, and the bottom is no longer necessarily the worst position.
2. Background
The background frames rank as a filter on which alternatives enter consideration, while AI agents may search differently because they computationally process full contexts. It motivates testing search depth, presentation-order effects, and whether inspection effects survive booking.
- Human search: Consumers typically use fast, frugal filters to assemble a small consideration set before choosing within it.Because the first stage governs what is evaluated, presentation order can affect demand and choice probabilities.
- Human search: Rank determines which alternatives enter the consideration set rather than what they are worth after entry.This distinction explains why position can matter strongly when consideration sets contain only one alternative.
- AI delegation: AI agents retrieve, compare, and complete purchase tasks end to end, potentially transforming the relationship between presentation order and choice.The paper contrasts pull-based results pages with conversational assistants.
- AI search: LLMs extract and synthesize attributes through explicit multi-step reasoning over text in context, unlike human reliance on visual heuristics and intuitive judgment.Their long-context retrieval is often more reliable at the beginning and end than in the middle.
- Research questions: The study asks whether AI search differs in depth from human search, whether presentation order shapes inspection, and whether any position effect survives booking.Reasoning effort is introduced as a possible additional influence on inspection and choice.
3. Experiment
The experiment places autonomous LLM agents in a hotel-ranking environment and evaluates their behavior against an established human benchmark. It reproduces the benchmark’s search context to study agents as surrogate consumers.
- Experiment: The study observes how an autonomous LLM agent navigates a search-ranking environment while acting as a surrogate consumer.Agent behavior is evaluated against a human benchmark.
- Experiment: The experiment replicates the context of Ursu’s randomized human field experiment to ensure comparability.The mapping from the consumer search environment to AI agents is described in Web Appendix A.
3.1. Experiment design
The design randomizes hotel presentation order and gives agents staged access to listing and detailed hotel information. Prompts and trip parameters align the agent task with the human benchmark.
- Randomization: Hotel information was partitioned into layers to simulate sequential search, and the 100-hotel order was fully randomized for each session.Randomized presentation isolates position from systematic placement of higher-quality options.
- Information layers: The initial listing page displayed hotel names, review scores, nightly and total prices, and aggregate review scores.Agents used an inspect tool to obtain additional hotel details.
- Information layers: Inspection returned star ratings, specific amenities, cleanliness and service sub-ratings, room availability, and detailed room prices.These fields provided the deeper information available after an agent selected a hotel to inspect.
- Task: Agents received a hotel-booking-assistant persona and were instructed to evaluate options and book using trip parameters.The itinerary matched the modal human benchmark: two adults, no children, one room, and two nights including Saturday.
3.2. Sample and procedure
The study evaluates proprietary LLMs from Google and Anthropic in repeated hotel-booking sessions, using default model settings and recorded tool-call sequences.
- Four proprietary LLMs from Google and Anthropic participated in the hotel-search experiment.The sample included Gemini 3.1 Pro, Gemini 3.7 Flash, Gemini 3.1 Flash Lite, and Claude Sonnet 5.
- The experiment used default reasoning effort and sampling parameters.The study notes that most modern LLM APIs reject manipulated non-default sampling parameters.
- 500 independent choice sessions were completed for each model.The repeated-sampling design followed behavioral evaluation paradigms for assessing LLM decision reliability.
- Each session presented a randomized list of 100 hotels for sequential inspection.The LLM could inspect listings as many times as it deemed necessary.
- Agents could book a hotel with the submit_choice tool or terminate without booking through the outside option.The chronological order of all tool calls was recorded throughout each experiment.
3.3. Measures
The analysis measures whether hotels are inspected and chosen as functions of randomized position, while controlling for observable listing attributes and clustering errors by session.
- Inspection records whether an agent invoked the inspect tool, serving as an analogue of a human click.Inspected equals 1 when the tool was invoked and 0 otherwise.
- Choice records whether a hotel was submitted as the final booking decision.Chosen equals 1 for the submitted hotel and 0 otherwise.
- Reduced-form model standard errors were clustered at the session level because hotels within a session were not independent.
- Position was indexed from 1 to 100 according to each hotel’s randomized placement on the listing page.Position served as the principal independent variable.
- The models control for listing-page attributes directly observable to the agent, including price and review score.Price captures the nightly hotel price, while review score captures the aggregate guest rating shown on the page.
- Additional controls capture chain affiliation and visible promotions as binary listing indicators.These variables approximate the closest direct overlaps with covariates from the original human study.
3.4. Results
AI agents inspect more listings than human consumers, yet listing position still predicts inspection weakly and in a non-monotone pattern. Position effects on choice vary across models, while choices concentrate on the same undominated hotel.
- Search behavior: Human consumers book 66% of sessions and take the outside option 34% of the time, whereas all AI agents book in 100% of sessions.The recommendation prompt may force conversion and rule out the outside option, a possibility tested through prompt variation.
- Position effect on inspection: Randomized rankings still predict inspection: AI position coefficients range from −0.0002 to −0.0005, compared with −0.0019 for humans.Moving a listing down ten ranks costs humans about 1.9 percentage points of inspection probability, versus 0.2–0.5 points for AI agents.
- Position effect on choice: Position effects on final choice vary by LLM, but every model concentrates 89.6%–100% of choices in the same five hotels and shares one modal choice.The modal hotel, citizenM New York Times Square, is undominated on price and review score, with a 4.7 review score at a $220 nightly price; pooled choice share is 78.2%.
4. Conclusion
In a hotel-search environment, AI agents searched more deeply than humans, and ranking affected inspection in a weak, non-monotonic way. Ranking affected final booking for only some models, while all four converged on the same undominated listing.
- AI agents searched more deeply than human consumers under default settings and never declined to book.
- For all LLMs except Pro, inspection declined from the top to a minimum near ranks 68 to 74 before rising again.
- Rank mattered at booking for Sonnet and Flash Lite, but not Flash or Pro, with heterogeneity unrelated to provider, capability tier, or search depth.
- All four LLMs converged on the same undominated listing, which captured 78.2% of all bookings.
- For humans, position effects reflect scrolling and sequential attention exhaustion; for AI agents, the U-shaped pattern resembles lost-in-the-middle retrieval.
- SEO strategy is conditional on the particular LLM used as a platform decision-maker, while review scores remain a primary heuristic and trust indicator even when rankings shift.
- The study fixed the query and destination in a hotel-booking task, so tasks with different attribute complexity could produce different results.
- The lost-in-the-middle interpretation remains provisional because identifying the mechanism would require attention-level access unavailable in the study.
5. Data availability statement
The authors provide data and code to reproduce the analysis.
- The data and code to reproduce the analysis are available at an anonymous location.