Source-linked AI summary

Large Language Models Pass the Turing Test

Cameron R. Jones, Benjamin K. Bergen

arXiv:2503.23674v1cs.CLcs.HC

TL;DR

The paper examines whether contemporary LLMs can be distinguished from humans in a standard three-party Turing test, addressing concerns about narrow, static AI benchmarks. It evaluates multiple AI systems across independent populations and considers prompting, finding that prompting is important for Turing-test performance.

  • Problem

    Narrow, static AI benchmarks may reflect memorization or shortcut learning, motivating evaluation of interactive capacities and whether systems can substitute for real people without detection.

  • Method

    The study used a randomized, controlled three-party Turing test with a five-minute limit, evaluating GPT-4.5, LLaMa-3.1-405B, GPT-4o, and ELIZA across two independent populations, with and without persona prompts.

  • Results

    Interrogators were most accurate when witnesses responded to strange prompts or typical LLM jailbreaks, while questions about weather or human experiences tended to be least effective.

  • Takeaways & Limitations

    The Turing test measures practical capacities related to substitutability and deception, while prompting helps LLMs adapt their behavior and apparently pass as human.

  • Takeaways & Limitations

    The systems did not robustly pass without the PERSONA prompt, and no-prompt win rates differed across the undergraduate and Prolific studies.

Abstract

from arXiv · show

We evaluated 4 systems (ELIZA, GPT-4o, LLaMa-3.1-405B, and GPT-4.5) in two randomised, controlled, and pre-registered Turing tests on independent populations. Participants had 5 minute conversations simultaneously with another human participant and one of these systems before judging which conversational partner they thought was human. When prompted to adopt a humanlike persona, GPT-4.5 was judged to be the human 73% of the time: significantly more often than interrogators selected the real human participant. LLaMa-3.1, with the same prompt, was judged to be the human 56% of the time -- not significantly more or less often than the humans they were being compared to -- while baseline models (ELIZA and GPT-4o) achieved win rates significantly below chance (23% and 21% respectively). The results constitute the first empirical evidence that any artificial system passes a standard three-party Turing test. The results have implications for debates about what kind of intelligence is exhibited by Large Language Models (LLMs), and the social and economic impacts these systems are likely to have.

1 Introduction

The Turing test evaluates whether a machine can appear human in a direct, text-based comparison with a real person. This study applies a standard three-party version to contemporary AI systems, examining prompting, model type, and performance across independent participant populations.

  • The Turing test: The Turing test pairs one human and one machine witness in simultaneous text conversations with an interrogator, who judges which witness is human.A machine passes when the interrogator cannot reliably identify the human.
  • The Turing test: Unlike narrow, static benchmarks, the Turing test is flexible, interactive, and adversarial, allowing interrogators to probe open-ended weaknesses.This makes it a complementary evaluation for contemporary AI systems.
  • Why the test matters: The test measures practical capacities including substitutability and the ability to cause people to mistake a model for a real person.The paper links these capacities to possible job disruption, social engineering, and misinformation risks.
  • Study design: The study used a randomized, controlled, pre-registered three-party test to evaluate GPT-4.5 and LLaMa-3.1-405B, with GPT-4o and ELIZA as baselines.It compared persona and no-persona prompts and replicated the experiment across undergraduate students and Prolific workers.
  • AI models: The evaluated systems included three prompted language models and ELIZA, a rules-based chatbot from the 1960s included as a manipulation check.The model and prompt combinations included persona and no-persona variants for GPT-4.5 and LLaMa, plus GPT-4o and ELIZA baselines.
  • Study design: Participants completed 5-minute split-screen conversations with one human and one AI witness before judging which witness was human.They exchanged messages simultaneously, then reported their verdict, confidence, and reasoning.

2 Results

Across two studies, persona-prompted GPT-4.5 and LLaMA witnesses were difficult to distinguish from humans, while baseline systems performed below chance. Interrogators commonly relied on small talk, linguistic style, and interactional dynamics, although unusual prompts and signs of ignorance were more accurate cues.

  • Win rates: 73%: GPT-4.5-PERSONA’s overall win rate, significantly above chance in both studies.Its win rate was 69% among UCSD undergraduates and 76% among Prolific participants.
  • Win rates: 56%: LLAMA-PERSONA’s overall win rate, significantly above chance in Prolific but not undergraduate participants.Its study-specific win rates were 45% for undergraduates and 65% for Prolific participants.
  • Win rates: 21% and 23%: GPT-4o-NO-PERSONA and ELIZA had the lowest overall win rates, respectively.ELIZA’s win rate was significantly below chance in both populations.
  • Win rates: PERSONA-prompted models performed better than NO-PERSONA witnesses overall, with all comparisons significant at p < 0.001.GPT-4.5-NO-PERSONA and LLAMA-NO-PERSONA had overall win rates of 36% and 38%.
  • Interrogator accuracy: Accuracy effects were inconsistent across studies, with no variable showing a consistent effect across both populations.Age, LLM knowledge, and chatbot-interaction frequency did not consistently predict participant accuracy.
  • Strategies & reasons: 61%: Small talk was the most common interrogator strategy, followed by probing social and emotional qualities in 50% of games.Asking about weather or human experiences tended to be least effective, whereas unusual statements and typical LLM jailbreaks were most accurate despite being infrequent.
  • Strategies & reasons: 27% and 23%: The most frequent verdict reasons concerned linguistic style and interactional dynamics, respectively.Reasons focusing on question handling and lack of knowledge tended to produce the most accurate verdicts.

3 Discussion

The findings show that suitably prompted LLMs can pass the original three-party Turing test, while also clarifying that success reflects humanlikeness rather than a direct, decisive test of intelligence. The results further indicate practical implications for substituting for people in short conversations, alongside important boundaries concerning prompting, populations, detection, and test duration.

  • Passing the Turing test: GPT-4.5-PERSONA was judged human significantly more often than chance in both independent populations, providing the first robust evidence that any system passes the original three-party test.The result means interrogators were more likely to select the model than the other human participant.
  • Prompting and robustness: Without the PERSONA prompt, GPT-4.5 and LLaMa did not robustly pass across populations.Both models were indistinguishable from humans in the Prolific study but not in the undergraduate study.
  • What the test measures: Participants mostly evaluated social, emotional, cultural, and linguistic behaviour rather than the knowledge and reasoning questions envisioned by Turing.Only 12% quizzed witnesses on knowledge and reasoning questions, while many focused on language use and personality.
  • What the test measures: The Turing test measures humanlikeness rather than directly or decisively measuring intelligence.The authors argue that intelligence is complex and multifaceted, so Turing-test performance should be considered alongside other evidence.
  • Social and economic relevance: The findings suggest that contemporary LLMs can substitute undetectably for real people in short conversations, with possible wider social and economic consequences.The paper discusses potential substitution in economic roles and secondary consequences from robustly imitating human interaction.

4 Methods

The studies evaluated multiple AI witnesses with baseline and persona prompts in controlled Turing-test games, then analyzed identification accuracy and exploratory demographic correlates.

  • Models and prompts: PERSONA prompts instructed models to adopt a specific young, introverted persona interested in internet culture.The prompt was selected after an exploratory study found it performed best among users of turingtest.live.
  • Models and prompts: 6 AI systems combined GPT-4.5 and LLaMa-3.1-405B with PERSONA and NO-PERSONA prompts, alongside GPT-4o and ELIZA baselines.GPT-4.5 and LLaMa-3.1-405B received both prompt types; GPT-4o received NO-PERSONA, and ELIZA served as a baseline.
  • Message timing: 1 + N(0.3, 0.03) × n_char + N(0.03, 0.003) × n_char_prev + Γ(2.5, 0.25)s modeled minimum, typing, reading, and thinking delays.The delay design added approximately 0.3 seconds per current-message character and 0.03 seconds per previous-message character, plus right-skewed thinking time.
  • Game procedure: 5-minute games used simultaneous conversations with two witnesses, followed by a verdict, confidence rating, and reasoning from the interrogator.The interface resembled a messaging application, with separate witness conversations and a time limit.
  • Statistical analyses: PERSONA prompting significantly reduced interrogator accuracy in both populations (Undergraduates: z = −5.27, p < 0.001; Prolific: z = −5.08, p < 0.001).The analysis used logistic mixed-effects regression with random intercepts for model and interrogator ID.
  • Statistical analyses: NO-PERSONA identification was not above chance in Prolific but was above chance for both GPT-4.5 and LLaMa in the Undergraduate study.For GPT-4.5, corrected p = 0.09 in Prolific versus corrected p < 0.001 for Undergraduates; for LLaMa, corrected p = 0.32 versus corrected p < 0.001.
  • Exploratory analyses: Demographic effects were limited: Prolific accuracy varied by gender, trial count, and beliefs about AI intelligence, while many other variables showed no significant effects.Age, LLM knowledge, and chatbot interaction frequency had no significant effects in either study.

Appendix

The appendix documents the study populations, interfaces, prompts, timing-related materials, and schemas used to classify interrogator strategies and reasons.

  • Interface and calibration: Figures 7–9 document the messaging interface, verdict interface, and confidence calibration by witness type.The verdict interface records the selected human witness, confidence, and reason; calibration patterns vary across models.
  • Study materials: Figure 10 documents demographic-data distributions across the two studies’ distinct populations.
  • Interrogator behavior: Figures 11–15 illustrate ELIZA examples and classify interrogator strategies and reasons for successful and failed AI witnesses.The classifications include overall categories and top reasons by AI witness outcome.
  • Coding and survey materials: Tables 2–4 provide the strategy-tagging schema, reason-tagging schema, and exit-survey questions.
Loading 2503.23674v1…