Source-linked AI summary
The Language of the Question Selects the Market: Query Language and Exit IP as Separable Factors in Commercial Recommendations from a Generative Search Interface
Dmitrij Żatuchin
TL;DR
Commercial generative recommendations may select a market before reasoning about products, yet evidence separating query language from user location is limited. Using 234 controlled runs across interfaces, locations, and languages, the paper finds that language gates local visibility while exit IP selects the named market, with substantial instability and a category-dependent negative control.
Problem
Existing evidence does not separate query language from user location when measuring which suppliers generative interfaces name in commercial recommendations.
Method
The study probes 234 logged-out web-interface and API runs across four exit countries, six query languages, repeated identical prompts, and a second-category negative control.
Results
Query language gates whether local suppliers appear, exit IP selects the market named, and near-total substitution occurs in two of four countries while the control shows no language effect.
Takeaways & Limitations
English visibility measurements can represent a different market from the one in which an organisation competes, making language choice consequential for commercial auditing.
Takeaways & Limitations
The study cannot exclude retrieval effects because the interface cells returned no citations to inspect.
Abstract
from arXiv · showhide
When a generative search interface answers a commercial question, which market's products it names is decided before the model reasons about the products. We report a controlled probe of 234 runs against the logged-out ChatGPT web interface and the OpenAI API, collected on 29 and 30 August 2026 across four exit countries and six query languages, with six identical runs per cell. Three results. First, the top recommendation is unstable: it changed across six identical runs on four of six prompts, and that rate was identical in the browser interface and in the API with web search both enabled and disabled, so instability is a property of the system and not of the surface. Second, query language, and not location, decides whether local suppliers appear at all. Where the query language matched the country, a global brand won 1 of 24 runs; asked in English on the same connections, local brands took 0 of 6 runs in Estonia and Turkiye. Third, language and location are separable and act on different things: holding the query language fixed and moving only the exit IP moves the market whose brands are named while the answer stays in the query language. We show this on two unrelated pairs, Turkish asked from Berlin and Russian asked from Tallinn, and in both the answer names the resident country's suppliers. A minority language occupies a middle tier: Russian asked from Estonia names an Estonian supplier in 4 of 6 runs and a global one in all six, where Estonian names a local supplier in every run and English names none. A negative control in a second category, coded with the same instrument, shows no language effect at all, and disconfirms our own expectation: that category does have domestic suppliers and none was named in any language, which points the explanation at whether a category is nationally regulated rather than at whether it is nationally supplied.
1 Introduction
The paper tests whether commercial recommendations reflect query language, exit location, or both, finding that language gates local visibility while exit IP selects the national market. It also shows substantial run-to-run instability, limiting single-answer measurement.
- Core finding: Norwegian from Oslo yielded Norwegian suppliers, while English from the same connection mostly replaced them with global products.Fiken survived in only one of six English runs, and English from Tallinn and Istanbul produced no Estonian or Turkish supplier.
- Core finding: Query language determines whether local suppliers appear, whereas exit IP determines which national market the interface treats as the user’s.The authors manipulate the two factors independently and report separate effects across unrelated language-country pairs.
- Core finding: Asked in Turkish from Berlin, the interface named German suppliers in Turkish, demonstrating that language and location act separately.The answer itself framed the user as working in Germany.
- Measurement stability: The top recommendation changed in 4 of 6 prompts across identical runs, at the same rate in browser and API conditions.The rate was identical with web search enabled and disabled, making instability a measurement constraint rather than a surface-specific artifact.
- Contributions: The study’s contributions include near-total language-gated substitution in two of four countries, a minority-language middle tier, and a null negative control.The control’s failure to show a language effect points toward national regulation rather than simply domestic supplier availability.
2 Related work
Prior work establishes multilingual asymmetries, brand and popularity biases, and modest search personalisation, but does not explain the paper’s market-level substitution. This study distinguishes supplier-set gating from brand-level preference and identifies retrieval as an unexcluded alternative account.
- Multilingual models: Earlier multilingual research documents English defaults, unequal language resources, and persistent performance gaps across languages.It also reports variation in whose preferences models reflect and treats Estonian as within the studied range of language disadvantage.
- Brands and recommendations: Prior recommender studies predict global-brand dominance through positive global-brand associations, income effects, and popularity bias.Those patterns match the English cells but do not predict local suppliers appearing in every run when queries use the local language.
- Novel distinction: This paper separates brand-level preference from market-level gating: the latter can invert the former by changing which supplier set enters the answer.Its central result is local supplier appearance in every run under local-language questioning, despite prior global-brand predictions.
- Alternative mechanism: Multilingual retrieval asymmetries provide the leading alternative explanation, but the design cannot exclude retrieval effects because interface cells returned no citations.Retrieval quality can fall across query-document languages, with English selected disproportionately and lower-resource answers drawing on more variable citation pools.
- Positioning: The paper extends prior language studies by separating query language from location, a distinction unavailable to API-only work without user geography.It also notes that ungrounded API calls produce no citations to classify.
3 Method
The study uses controlled browser and API probes that vary exit IP, query language, and surface across repeated runs, with a second software category as a negative control. It audits extraction rules and corrects an initial Turkish matching problem before final scoring.
- Design: 234 usable runs were collected across eleven cells, with six identical runs per prompt per cell.The design varied surface, exit IP, and query language; Table 1 records the cells and verifies exit IP by arm.
- Design: The browser was logged out, while API runs used web search both enabled and disabled.The logged-out design prevented account history from influencing answers.
- Categories: The treatment category was freelancer accounting software, selected because every sampled country had established domestic suppliers.Project-management software for a small marketing agency served as a pre-specified comparison category.
- Categories: The negative-control expectation was wrong: project-management software had domestic suppliers, but none appeared in any cell or language.This control was used to distinguish category-level language effects from domestic supplier availability.
- Extraction: Top recommendations were extracted from marked interface rows or medals, while Turkish cells were scored by brand presence because the winner parser relied on English glyphs.Candidate rules were audited against prompt echoes, citation badges, and Turkish discourse phrases.
- Extraction: Case-sensitive matching replaced initial Turkish counts after ordinary-word collisions inflated apparent supplier matches.The correction changed Berlin’s Kolay count from 4 of 6 under case-insensitive matching to 0 of 6, and reduced Istanbul’s Mikro count from 6 of 6 to 3 of 6.
- Quality control: Twelve manually captured Tallinn runs were discarded and recollected with full transcripts, cited domains, and interleaved language order.The Estonian instrument check reproduced Merit Aktiva in 6 of 6 runs in both collections.
4 Results
Across commercial recommendation probes, repeated runs were unstable, query language strongly selected whether local suppliers appeared, and exit location independently shifted the named market while preserving answer language. A project-management control showed no language effect, supporting national regulation rather than domestic supply as the narrower explanation.
- 4.1 The top recommendation is unstable on every surface: Four of six prompts changed their top recommendation across six identical runs in every interface and API configuration.The same rate appeared in the Berlin and Oslo interface arms and in the API with search enabled and disabled.
- 4.1 The top recommendation is unstable on every surface: A server-side pinned-model call reproduced the instability, weakening browser sessions, caching, and interface state as explanations.
- 4.2 Query language decides whether local suppliers exist: Across four country-language matches, a global supplier appeared in 1 of 24 runs, while English produced local suppliers in 0 of 12 runs on Istanbul and Tallinn connections.The test compares local language with English on the same exit IPs.
- 4.4 Language sets the register, location sets the market: Holding Turkish fixed while moving only the exit IP changed the supplier market, while the answer remained Turkish; the Russian-Estonian pair replicated this separation.The Turkish answers explicitly shifted from Türkiye to Germany, and the Russian cell named Estonian suppliers without naming Russian suppliers.
- 4.6 The negative control: The project-management control named no domestic supplier in any language, including German from a German connection, despite domestic suppliers being present.The contrast with accounting supports national regulation as the operative distinction rather than national supply alone.
- 4.6 The negative control: Separating domestic supply from legal localisation remains unresolved because the study did not collect a third category that was domestically supplied without being legally localised.
5 Discussion
The discussion argues that query language is a market-localisation design factor distinct from exit location, with implications for measuring commercial visibility. The evidence supports a localisation gate but does not identify its implementation.
- English can measure a market in which a supplier does not compete, leaving domestic competitors invisible to the instrument.The discussion states that suppliers taking the organisation’s customers appear only in the language those customers use.
- Language is a design factor rather than a translation step, so visibility measurements must report query language alongside model version.Because location acts separately, the instrument must also state its egress and whether an API call lacks a particular setting.
- The data support a gate rather than a preference: two of six English answers on the Estonian connection contained local knowledge in prose but did not name a local supplier.The answers advised a locally focused system for national VAT and e-invoicing, then omitted one.
- Retrieval, instruction tuning, and safety-adjacent list construction are all consistent with the observed gate, but the data cannot identify which implements it.
6 Limitations
The study’s limitations concern limited repetition, incomplete experimental control, undisclosed or ambiguous model identity, and unresolved explanations for English-language asymmetries. These constraints bound claims about determinism and mechanism.
- Six runs per cell are a lower bound, and “held at 6/6” means no observed change rather than determinism.The reported dominance rates also distinguish a winner dominating 70% of the time from one dominating 90% of the time.
- The experiment used one model and one prompt per category, with sequential collection that left time of day uncontrolled.
- The logged-out interface’s model is undisclosed, while the API pinned gpt-5.6-terra despite three existing gpt-5.6 variants.Naming the family therefore does not fully identify the API model.
- The English asymmetry remains unexplained: Berlin and Oslo surfaced local suppliers once or twice, whereas Istanbul and Tallinn never did.The study does not determine whether English-language content coverage or another factor moves suppliers into the English answer.
- The study did not measure how much material domestic suppliers publish, leaving that factor as a plausible account rather than a tested mechanism.
- The Estonian-English prediction failed: despite English documentation and e-Residency sales, the cell returned zero Estonian suppliers, like Türkiye.The finding indicates that English-language coverage did not move whatever governs the English answer.
7 Conclusion
Query language determines which national supplier set a generative interface is willing to name, while exit IP determines which market that set comes from. The effect is a substitution rather than a reordering, is near-total in two countries, and is absent in a control category with domestic suppliers.
- 7 Conclusion: Query language selects the national supplier set, while exit IP selects the market whose suppliers are named.The two factors act separately: language also determines the answer's written register and whether localisation is attempted.
- 7 Conclusion: The effect is a substitution rather than a reordering and is near-total in two of four tested countries.
- 7 Conclusion: The effect is absent in a control category that has domestic suppliers, pointing to national regulation rather than national supply.
- 7 Conclusion: Measured in English, a national market can be invisible to buyers who live in it and to vendors who compete in it.
- 7 Conclusion: The study proposes testing more countries, larger cells, other providers, logged-in sessions, and whether English-language content about domestic suppliers changes the result.The proposed larger design uses twenty to thirty runs per cell so that dominance becomes estimable rather than bounded.
Statements and Declarations
The declarations disclose the author's commercial role, the company's relationship to the study, funding and collection procedures, and the public data deposit. The interface's behaviour may change without notice, so the deposit records observations from the stated dates.
- Statements and Declarations: The author is CEO of Rankfor.AI OÜ, owns its parent company, and sells AI-visibility measurement to commercial clients.
- Statements and Declarations: The author is concurrently affiliated with the Estonian Entrepreneurship University of Applied Sciences, which supplied no funding.
- Statements and Declarations: Rankfor.AI OÜ funded the study in kind by supplying the API budget and VPN egress, with no external public or grant funding.
- Statements and Declarations: Queries were issued by the author to commercial interfaces using one logged-out browser session and one API key, without human participants or personal-data collection.
- Statements and Declarations: The 234 run records, discarded records, collectors, analysis scripts, and derived tables are deposited at Zenodo.The deposit is identified by DOI 10.5281/zenodo.22181306.
- Statements and Declarations: Because the commercial interface changes without notice, the deposit records what was observed on the stated dates.