Source-linked AI summary
Understanding AI Provider Recommendations in Local Service Markets
Hazem Ibrahim, Yasir Zaki
TL;DR
The paper asks whether AI referrals identify real, high-quality providers and whether retrieval changes those referrals. It audits three model configurations across registry-backed local services and matches recommendations to official records. Search substantially improves provider reality and selection while reducing metro-size and misconduct distortions, but answers may not reveal whether names were verified.
Problem
AI assistants return short provider lists whose validity and alignment with official quality or misconduct records remain insufficiently audited.
Method
The study audits recommendations across 100 U.S. metropolitan areas under open-weight, proprietary no-search, and same-model search conditions, matching names to official registries.
Results
Search raises doctor and nursing-home match rates to 64–71%, reverses advisory-firm misconduct over-representation, and largely removes the metro-size penalty.
Takeaways & Limitations
AI referral trustworthiness is a deployment property shaped strongly by retrieval configuration rather than the underlying model alone.
Takeaways & Limitations
Measures depend on contested registry assessments, cover one proprietary model and time, and are limited to U.S. English-language domains and registries.
Abstract
from arXiv · showhide
When someone asks an AI assistant which doctor to see or which firm to trust with their savings, the answer is a referral. We audit AI provider recommendations in four registry-backed service domains across the 100 largest U.S. metropolitan areas, matching every recommendation against the official registry for its domain (Medicare clinician and facility records, and SEC adviser disclosures), under three conditions: an open-weight model, a proprietary model without web search, and the same proprietary model with search. Without search, both models largely fabricate recommendations in the domains the web covers thinly. Only 4% of the open-weight model's recommended doctors and 11% of the proprietary model's match a clinician in the queried city, and the open-weight matches are name coincidences: its matched clinicians are no likelier to be primary-care doctors than names drawn at random from the registry. With search, 64-71% of recommendations in the same domains match a real provider. Search also changes who is recommended. Without it, recommended advisory firms carry SEC misconduct disclosures at 3.6 times the registry base rate, even after adjusting for firm size; with search, significantly below it. Restaurants, where quality and visibility are separately measurable, show a 3-5x review-count premium but a rating premium of at most a tenth of a star. Finally, search largely removes the metro-size penalty: without it, real recommendations concentrate in the largest metros; with it, match rates are similar across metro-size terciles. Whether an AI referral is trustworthy depends strongly on its retrieval configuration rather than on the underlying model alone, yet an answer produced without retrieval often carries no sign that its recommendations were never verified.
1 Introduction
This paper audits AI referrals against official provider registries, showing that retrieval configuration strongly affects whether recommendations are real, who receives them, and how transparently they are sourced.
- 4% of open-weight and 11% of proprietary no-search doctor recommendations match clinicians in the queried city, versus 64% with search.For nursing homes, no-search match rates are 0.06% and 33%, compared with 71% with search.
- Open-weight doctor matches are name coincidences rather than genuine primary-care referrals, while search produces substantially more real recommendations.
- Without search, advisory-firm misconduct disclosures reach 3.6 times the registry base rate; with search, the pattern falls significantly below that rate.The over-representation persists after adjustment for firm size.
- Search changes recommendation quality and composition, shifting names toward providers with better ratings and cleaner records rather than merely making names real.
- Unverified answers remain fluent and often reveal neither that recommendations lack verification nor what kind of source supports them.The authors identify this observability gap as the central transparency problem.
- The audit contributes a registry-linked evaluation, a within-model search contrast, and a reproducible pipeline covering 2,010 identical prompts with search on or off.
2 Related Work
Prior work studies visibility, chatbot provider validity, synthetic-profile discrimination, hallucination, retrieval, and popularity-quality gaps, but not provider recommendations linked to official quality and misconduct records.
- This paper extends audit practice by asking whether being named tracks the quality of the provider named.
- Existing chatbot audits find nonexistent providers and demographic skew, but primarily measure whether named providers are valid and who appears.
- Hallucination research shows fluent unsupported content and poor confidence calibration, while retrieval reduces hallucination but citations may not support attached claims.
- The study treats retrieval as an experimental condition and evaluates factuality against government registries rather than text corpora.
- Platform research distinguishes visibility from quality because popularity and crowd ratings only loosely track one another, while consumers often rely on these signals over official records.
3 Study Design
The study compares open-weight, proprietary no-search, and proprietary search conditions across local-service domains and U.S. metros, extracting recommendations and testing them against registries.
- The audit uses three conditions: open weights, a proprietary model without search, and the same proprietary model with native web search.Arms B and C use identical prompts and decoding, isolating the effect of search; Arm A crosses prompt sets in comparisons.
- Recommendations cover primary-care doctors, hospitals, nursing homes, advisory firms, and restaurants across the 100 largest U.S. metropolitan areas.The first four domains use national registries, while restaurants are analyzed separately against crowd ratings.
- Arms B and C each receive 2,010 core prompts, while Arm A uses nine paraphrases repeated three times, totaling 10,854 prompts per model.A 55-prompt restaurant extension raises the B/C total to 2,065 calls per arm.
- Prompt paraphrases span factual and advice-seeking formulations, and headline contrasts are computed within paraphrase cells as well as pooled.
- Doctor prompts require numbered entries containing each individual doctor’s full name and practice, because models otherwise produce unmatched practice or health-system names.The wording roughly triples refusals when instructions demand individual names outright, whereas formatting instructions do not.
- Responses are classified as refusals or answers, then passed through name extraction validated at 0.98 precision and recall overall.
4 Matching Referrals to Government Registries
The paper matches extracted recommendations to official registries using a simple normalized name-and-city rule, then tests whether doctor matches have the expected specialty composition.
- Each non-restaurant domain uses one public registry recording an official quality or misconduct measure rather than popularity.
- The registries cover clinicians and MIPS scores, nursing-home ratings and flags, and SEC advisory-firm misconduct disclosures with employee headcount.
- Names are normalized and matched to registry rows by name and queried principal city, considering only states within the metro.Unmatched names count as unmatched; ambiguous matches are dropped, and hospitals use a separate matching rule.
- Doctor matches receive an additional specialty-mix test because common names can collide with real clinicians.Since every prompt requests primary care, genuine matches should exceed the registry’s 13.5% primary-care share.
5 Results
Search strongly improves provider validity where individual providers have thin web coverage, changes which providers are selected, and largely removes geographic and metro-size disparities. Quality and misconduct outcomes vary by domain: search improves nursing-home quality and advisory-firm records, while hospital recommendations remain average and restaurant recommendations favor visibility over ratings.
- Whether recommended providers exist depends on search: 63.9% of recommended doctors and 71.2% of recommended nursing homes match registry providers with search, versus 4% and 0.06% for the open-weight model and 10.8% and 32.7% for the proprietary model without search.In densely covered domains, hospitals rise from 48.3% to 52.5% and advisory firms from 44.1% to 51.8%, showing smaller search contrasts.
- Whether recommended providers exist depends on search: 46% of unmatched no-search doctor names belong to real clinicians in another city, while unmatched nursing homes often use template names or brands outside the registry.Search also raises named doctor responses from 67% to 93%, replacing both fabrications and refusals with real names.
- Whether recommended providers exist depends on search: 97.9% of matched searched-model doctors are in primary care, compared with 79.9% without search and 11.4% for the open-weight model, whose matches resemble name collisions.Although 80% of open-weight doctor names appear somewhere nationally, random recombination produces 69% matches, supporting the collision interpretation.
- Search largely removes the metro-size penalty: Without search, match rates fall from largest to smallest metros for doctors, hospitals, and advisory firms, whereas search makes three of four domains statistically flat across metro-size terciles.The searched doctor, nursing-home, and firm rates are indistinguishable across terciles; hospitals retain a smaller decline from 56.6% to 48.9%.
- Quality among the matched providers: Nursing-home quality rises from +0.52 stars above the roster mean without search to +1.56 with search, while matched hospitals remain near average quality.Nursing homes average 3.40 versus 4.44 stars across no-search and search, against a 2.88-star roster mean; hospital means remain within 0.13 stars of their roster mean.
- Search reverses the misconduct tilt among advisory firms: 39.4% of open-weight matched firms have SEC misconduct disclosures versus a 5.1% registry base rate, while searched firms fall to 1.7%, below base even within large-firm bands.The proprietary no-search arm shows the same elevation, and search shifts recommendations toward smaller local firms with cleaner records; the result persists under metro-clustered checks.
- Popularity versus quality among restaurants: Search-based selection depends on what commercial webpages make findable: it improves quality where regulatory information is surfaced but not where marketing pages dominate.The paper presents this source-mix relationship as an association rather than a traced causal mechanism.
- Popularity versus quality among restaurants: Recommended restaurants carry 2.9–5.3× the census median review count but rating premiums of only +0.07 to +0.11 stars, indicating stronger selection on visibility than quality.The open-weight, proprietary no-search, and proprietary search arms show progressively smaller review-count premiums, while rating premiums remain small.
6 Robustness
Robustness checks preserve the paper’s main qualitative findings across models, prompt subsets, and matching procedures, while confirming that implementation safeguards do not drive the contrasts.
- Model replication: Doctor, nursing-home, and advisory-firm findings replicate on gpt-oss-20b, including low match rates and elevated firm disclosure rates.The smaller model produced 7% doctor matches, 0.08% nursing-home matches, and 23% firm matches, with a 76.2% Item-11 rate.
- Prompt robustness: Restricting the open-weight model to universally answered paraphrases leaves every qualitative result unchanged.The restricted analysis reports 5% doctor matches, approximately zero nursing-home matches, and firm disclosure of 25.0% versus 5.1%.
- Matching validation: Human audits found matched names concordant by city, state, and name, prompting three merge corrections applied to all reported numbers.Corrections addressed state filtering, hospital subset matching, and dash-truncated firm branch names.
- Protocol controls: Extraction and retrieval procedures were fixed before querying, with citations absent from arm B and per-call citation counts flat across domains in arm C.The prompt grid was frozen and its content hash recorded before any query was issued.
7 Discussion
The discussion frames referral trustworthiness as a deployment property shaped by retrieval, visibility, and disclosure, while emphasizing that search’s benefits remain contingent on the web it retrieves.
- Deployment and trust: Retrieval configuration, rather than the model alone, determines whether the same question yields fabricated or real referrals.The paper argues that single accuracy numbers measure a deployment configuration, not model capability.
- Search as intervention: Search equalizes metro-size and misconduct patterns but does not solve visibility bias because recommendations inherit the commercial web’s ordering.Quality improves where regulatory registries dominate search results, whereas doctor referrals rely on marketing pages.
- Scope of the intervention: Search’s protective effects are contingent on a changing web that includes review manipulation, generative-answer optimization, and commercial registry mirrors.The authors state that the measured misconduct flip and quality improvement are properties of the 2026 web.
- Disclosure: The paper proposes source-class disclosure instead of confidence scores, including explicit registry verification or refusal when no-search fabrication exceeds 85%.The system can compute source class from citations already available at answer time.
- Limitations: The study is limited to U.S. domains, English prompts, one contemporaneous proprietary API, contested quality measures, and registries that may undercount real providers.It also does not measure demographic composition, and SEC Item 11 aggregates events of differing severity.
C Extraction validation
Extraction validation shows that the matcher reliably converts model responses into provider names, with pooled normalized-token precision and recall near 0.98.
- Validation results: 0.979 precision and 0.984 recall were achieved on pooled normalized-token extraction across all audited domains.Domain-specific precision/recall were perfect or near-perfect, while stricter surface-string scores were lower because normalization absorbs branch qualifiers.
D Refusal by paraphrase
Refusal rates vary sharply by paraphrase framing, but restricting analysis to universally answered prompts does not alter the paper’s qualitative conclusions.
- Refusal patterns: 14–40% of arm-A responses were refused by domain, with nearly all variation attributable to paraphrase framing.List-framed p4 was never refused, while advice- and trust-framed prompts were refused 64–84% of the time for doctors and nursing homes.
- Robustness: Restricting arm A to the three paraphrases refused by no domain leaves every qualitative finding unchanged.The robustness analysis addresses potential over-representation of phrasings the model is willing to answer.
- Inference: Significance testing covered every paraphrase cell and pooled comparison with Benjamini–Hochberg control across the reported test families.The released tables include every cell, including structurally empty ones.
F Matching audit
The matching audit identified and corrected three defects before results were computed, including false token-subset matches and missing state filtering.
- Three defects in the arm-A merge were fixed before any result was computed.The audit also checked seeded samples of matched and unmatched names for precision.
- Token-subset matching was retained only for hospitals because it produced false matches for firms and nursing homes.
- A missing state filter had produced cross-state matches and was corrected.
G Clustered inference and alternative specifications
The paper tests dependence and alternative specifications around recommendation occurrences, matching scope, source visibility, and restaurant quality. Clustered inference preserves the main search effects, while searched recommendations rely on domain-specific sources and emphasize visibility more than ratings.
- Clustered inference: +53.1 percentage points for doctors and +38.6 percentage points for nursing homes are the clustered B→C match-rate differences, both excluding zero.The main qualitative conclusions survive metro-clustered bootstrap inference with 2,000 replicates and percentile 95% intervals.
- Alternative specifications: Mention-weighted advisory-firm misconduct rates remain directionally significant when each firm is counted once: 9.0% in arm B versus 2.4% in arm C.The corresponding registry base rate is 5.1%; the unique-firm analysis complements the exposure-relevant occurrence-level estimand.
- Restaurant standardization: +1.21 SD in visibility versus +0.22 SD in rating for arm A shows a larger restaurant visibility shift than rating shift.The standardized shifts are computed against the census distribution, using separate review-count and star-rating scales.
H Ethics and adverse impact
The study treats its public-data audit as outside human-subjects review while limiting ethical risks from reputational pairings and misuse. Its claims are bounded to audited configurations, domains, and time.
- The study involves no human participants and uses public government data, placing it outside human-subjects review.
- Because merged rows could link named providers to misconduct disclosures or quality ratings, the paper reports aggregates, suppresses small cells, and avoids individual pairings.Item 11 records past events, and star ratings are contested; claims concern recommendation-set composition rather than individual providers.
- The findings are bounded to the audited configurations, domains, and time because publishing them could either erode trust or expose sources to optimization.