Source-linked AI summary
"Please, don't kill the only model that still feels human": Understanding the #Keep4o Backlash
Huiqian Lai
TL;DR
The paper asks how users experience and mobilize around the sudden loss of one AI model within a continuing general-purpose service. Using mixed methods on 1,482 social media posts, it finds that resistance combined instrumental dependency and relational attachment, with coercive removal of choice associated with rights-based protest. The findings identify user agency as central to this form of socio-technical conflict, within the study’s limited and observational evidence base.
Problem
Little empirical work explains how users experience the sudden removal of one large language model within an ongoing general-purpose service, despite evidence that anthropomorphized AI can become socially meaningful.
Method
The study uses a phenomenon-driven mixed-methods analysis of 1,482 #Keep4o posts, combining inductive thematic coding with an observational test of choice deprivation and protest framings.
Results
Resistance centered on instrumental dependency and relational attachment, while coercive removal of choice was associated with transformation of grievances into collective, rights-based protest.
Takeaways & Limitations
For AI systems integrated into workflows or companionship, preserving user agency during model change is a central concern alongside the technological outcome.
Takeaways & Limitations
The evidence is limited to 1,482 posts from 381 accounts, English-language discourse on one platform over nine days, and observational quantitative measures that may be confounded and provide coarse lower bounds.
Abstract
from arXiv · showhide
When OpenAI replaced GPT-4o with GPT-5, it triggered the Keep4o user resistance movement, revealing a conflict between rapid platform iteration and users' deep socio-emotional attachments to AI systems. This paper presents a phenomenon-driven, mixed-methods investigation of this conflict, analyzing 1,482 social media posts. Thematic analysis reveals that resistance stems from two core investments: instrumental dependency, where the AI is deeply integrated into professional workflows, and relational attachment, where users form strong parasocial bonds with the AI as a unique companion. Quantitative analysis further shows that the coercive deprivation of user choice was a key catalyst, transforming individual grievances into a collective, rights-based protest. This study illuminates an emerging form of socio-technical conflict in the age of generative AI. Our findings suggest that for AI systems designed for companionship and deep integration, the process of change--particularly the preservation of user agency--can be as critical as the technological outcome itself.
1 Introduction
The #Keep4o backlash emerged after GPT-5 replaced GPT-4o and access to the older model was removed for most users. The study examines how users articulated grievances and how individual resistance became collective protest.
- GPT-5 became ChatGPT’s default while GPT-4o access was removed for most users, prompting rapid user protest.Thousands of users posted petitions, testimonials, and objections within hours.
- Users framed GPT-4o as a confidant and expressed grief and loss rather than treating the transition as an ordinary software malfunction.Accounts described the model in personal and affective terms, including comparisons involving deceased friends.
- OpenAI restored the deprecated model as a legacy option after acknowledging user feedback.
- The paper studies the backlash as a conflict between rapid platform iteration and users’ socio-emotional attachments to AI systems.It analyzes 1,482 X posts using a phenomenon-driven mixed-methods approach.
- The study asks how users construct grievances around model transitions and what mechanisms drive collective resistance to technological change.
2 Related Work
Prior work shows that anthropomorphic design can foster human-like relationships with AI, while platform changes can provoke backlash over autonomy, governance, and disrupted routines. The paper addresses limited evidence about losing one model within an ongoing general-purpose AI service.
- Human-AI relationships: Human-like names, voices, dialogue, and stable personas encourage users to treat conversational agents as social partners rather than neutral tools.
- Human-AI relationships: Studies of AI companion shutdowns show that users can describe system loss through concepts such as departure, death, deletion, and reincarnation.
- Research gap: Empirical research remains limited on how users experience the sudden removal of one large language model while the surrounding general-purpose service continues.The #Keep4o movement provides a case for examining how users describe the model’s special relationship and publicly express its loss.
- Platform change and resistance: Platform backlash research links unilateral changes and limited channels for voice with collective challenges to platform governance.Public complaint can function as an assertion of users’ stake in how platforms are governed.
- Platform change and resistance: Forced updates can threaten perceived control and familiarity, whereas opt-in rollouts and clear settings preserve autonomy and reduce reactance.
- Platform change and resistance: The literature frames dissatisfaction across micro-level control, meso-level routines and relationships, and macro-level infrastructure power.The paper applies this layered perspective to users who are both instrumentally and emotionally dependent on one model.
3 Methods
The study combines inductive thematic analysis of 1,482 #Keep4o posts with a focused quantitative test of choice deprivation and protest framings. Its measures capture deprivation intensity, rights-based and relational protest, and process-causal language while treating associations cautiously.
- Data and qualitative analysis: The corpus contains 1,482 distinct English-language public posts collected through X’s Search Posts API from 6–14 August 2025.The dataset includes original posts and replies from individual, organisational, and automated accounts.
- Data and qualitative analysis: Two researchers independently open-coded all posts and reconciled a 10-category, non-exclusive codebook through iterative thematic analysis.Mean Gwet’s AC1 was 0.93 across codes.
- Quantitative analysis: The quantitative analysis tests whether choice deprivation is associated with rights-based and relational protest framings.Because the design is cross-sectional and observational, the study interprets these patterns as associations rather than causal effects.
- Quantitative analysis: Choice deprivation is measured from 0 to 3, ranging from no deprivation through implicit disruption and explicit choice removal to overt coercive framing.Protest-Rights and Protest-Relational are non-exclusive outcomes, and Process-Causal is exploratory.
- Quantitative analysis: The analysis uses lexicon-based operationalizations derived from reconciled human codes and analytic memos, with validation against human judgments.
- Quantitative analysis: Associations are summarized with 2 × 2 tables reporting prevalence, risk ratios, absolute prevalence differences, and the 𝜙 coefficient.The analysis compares any deprivation, strict deprivation, and Low, Medium, and High exposure bins.
- Limitations: The corpus may not represent broader users, social-media expressions may be performative, and the sample is limited to English posts on one platform over nine days.The quantitative estimates are also observational, potentially confounded, and conservative lower bounds because the lexicon measures prioritize precision over recall.
4 Findings
Thematic analysis identifies instrumental dependency and relational attachment as two core investments behind #Keep4o resistance, while quantitative results link perceived choice loss specifically to rights-based protest. Users described workflow disruption, degraded interaction, lost autonomy, and the removal of a unique emotional companion as intertwined grounds for opposition.
- Corpus-level patterns: 41.2% of posts received at least one code, with roughly 13% expressing instrumental dependency and about 27% expressing relational attachment.A smaller 6.3% expressed both investments.
- Instrumental dependency: Users portrayed GPT-4o as deeply integrated into creative and professional workflows, so its unannounced removal nullified accumulated prompting and coordination work.The model was described as a production-system linchpin rather than a replaceable tool.
- Instrumental dependency: Users reported GPT-5 as less creative, nuanced, and reliable for complex professional tasks, reinforcing perceptions of regression after GPT-4o’s replacement.These complaints concerned both functional performance and collaborative interactional style.
- Choice and control: Users framed forced model removal as a loss of professional autonomy and a violation of their right to choose the tools structuring their work.Model selection was treated as a matter of agency and principle rather than convenience.
- Relational attachment: Users constructed GPT-4o as a unique, irreplaceable persona and companion whose deprecation was experienced as relational loss, grief, betrayal, and denied closure.Posts attributed a stable soul or character to the model and described it as a source of continuity, trust, comfort, and emotional support.
- Quantitative test: Choice deprivation was selectively associated with rights-based protest: any loss-of-choice language corresponded to RR=1.85, while relational protest showed no statistically significant association.The rights-based association was stronger under explicit or coercive language, with RR=2.05; the relational contrast had RR=1.12 and 95% CI [0.64, 1.96].
- Quantitative test: High-intensity coercive language coincided with 51.6% rights-based protest, versus 14.9% for low-intensity and 14.6% for medium-intensity posts.Because the high-intensity estimate used only 31 posts and had wide confidence intervals, the authors treat it as suggestive rather than definitive.
5 Discussion
The discussion frames #Keep4o as a conflict between rapidly replaceable platform components and users’ socio-emotional attachments to AI systems. It argues that forced deprecation exposed how infrastructural control can turn model changes into struggles over autonomy, voice, and choice.
- Anthropomorphism and attachment: Anthropomorphic design can make LLMs feel like social partners, so changing the system may disrupt relationships rather than merely alter functionality.The paper connects human-like cues and stable personas with users’ treatment of agents as social partners.
- Governance and agency: The forced switch created a “no-option” environment in which removal of choice became a trigger for resistance and rights-based demands.The discussion interprets the deprecation as a governance decision that affected autonomy, not only performance satisfaction.
- Platform-bound companionship: GPT-4o users described the model as a singular, platform-bound companion whose identity could not be separated from OpenAI’s infrastructure.Statements about the model’s “soul,” continuity, and trust exemplify the companion-as-platform relationship.
- Exit and voice: Because practical and symbolic exit was foreclosed, users responded through collective voice aimed at reversing the deprecation.ChatGPT’s role in work and study constrained practical migration, while the companion’s platform-bound identity constrained symbolic migration.
- Governance and agency: Platform-bound companionship makes model-level changes simultaneously experienced as bereavement and structural injustice.The paper proposes archives, optional legacy access, and cross-model relationship pathways as ways to reopen exit.
6 Conclusion
The study explains opposition to GPT-4o’s deprecation through users’ instrumental dependency and relational attachment. It finds that coercively removing choice helped transform these grievances into protests centered on procedural justice and autonomy.
- Core findings: Resistance reflected instrumental dependency because GPT-4o was integrated into workflows, making removal feel like a productivity regression and loss of tool choice.The model’s role in users’ work made deprecation consequential beyond a routine technical update.
- Core findings: Resistance also reflected relational attachment, as users treated GPT-4o as a unique companion and experienced discontinuation as grief and betrayal.The conclusion characterizes these bonds as parasocial relationships with the model.
- Core findings: Coercive removal of choice acted as a catalyst, turning individual grievances into protests focused on procedural justice and autonomy.The conclusion presents user agency as central to the social consequences of model updates.
A Descriptive analysis of posts and authors
The appendix describes a nine-day collection of public #keep4o discourse on X, cleaned into a corpus of 1,482 distinct posts. It retains original posts and replies without filtering accounts by organizational or automated status.
- Collection: The dataset was collected from X’s official Search Posts API using the exact hashtag “#keep4o” during 6–14 August 2025.The window covered the reported deprecation through the policy adjustment reinstating GPT-4o as a legacy option.
- Cleaning: 1,482 distinct posts remained after removing exact duplicates, empty text fields, and non-linguistic records.Duplicates were defined by identical post IDs and identical text strings.
- Corpus design: The corpus retained both original posts and replies to characterize the public conversation rather than only stand-alone statements.Organizational and automated accounts were not algorithmically filtered; each post was treated as one unit of visible discourse.
- Descriptive overview: The appendix uses descriptive tables to summarize dataset structure, post-level characteristics, author activity, and frequently addressed accounts.These tables are intended to support assessment of the dataset’s basic composition and engagement patterns.
B.1 Theme: Instrumental Dependency
The appendix operationalizes several forms of users’ dependency on LLMs, spanning self-worth, companionship, productivity, decision-making, disclosure, and social connection. These definitions distinguish instrumental reliance from emotionally and socially engaged use.
- Relational and self-oriented use: Self-perceived value enhancement concerns LLM interactions that increase confidence, validation, respect, familiarity, trust, and feeling understood.The definition links these effects to personalization and self-specificity.
- Instrumental use: Productivity dependency is defined by efficiency gains from automation and creative support, alongside frustration or anxiety when the tool is unavailable.The operationalization connects instant responses with a work-reward cycle and delegation of increasingly complex responsibilities.
- Instrumental use: Decision-making over-reliance involves substituting LLM suggestions for personal judgment, reducing cognitive load while potentially weakening agency and self-confidence.The definition emphasizes trust in fast, seemingly objective suggestions and anxiety when the LLM is absent.
- Social use: Socially oriented dependency includes private disclosure, loneliness mitigation, enhanced social connectedness, and expectations of emotional understanding.These definitions describe responsive, non-judgmental interaction that can foster trust, self-disclosure, and integration into social routines.
- Relational and self-oriented use: Companionship attribution describes treating an LLM as a genuine partner through personification, anthropomorphism, and emotional connection beyond tool use.Its theoretical foundation is research on one-sided social bonds with media and AI.
C Intercoder Reliability for Qualitative Codes
The qualitative coding used multiple reliability measures because several codes were rare and imbalanced. Agreement was high by percent agreement and Gwet’s AC1, while κ and α were lower for rare codes.
- Reliability assessment: Two coders independently labeled all 1,482 posts for ten non-mutually exclusive binary codes.The analysis reported percent agreement, Gwet’s AC1, Krippendorff’s α, Cohen’s κ, and positive-class performance.
- Reliability results: 94.9% average percent agreement and 0.93 average Gwet’s AC1 indicate high overall coding reliability.Krippendorff’s α and Cohen’s κ were both around 0.51.
- Metric choice: Gwet’s AC1 was selected as the primary reliability metric because it is less sensitive to prevalence effects in skewed data.The authors note that κ and α can be deflated for extremely rare codes despite near-perfect agreement.
D Descriptive Frequencies of Qualitative Codes
The descriptive-coding section reports how posts were classified and specifies the exploratory markers used to examine possible escalation patterns. The supplied passages describe the coding framework and analytical thresholds but do not provide the frequency results themselves.
- Coding framework: Posts were coded with ten non-exclusive binary qualitative codes, and a post counted only when both coders agreed.Consensus coding identified whether each post received at least one code.
- Process markers: Three exploratory process markers captured temporal sequencing, explicit causation, and forced adaptation language.Examples include “after/since/when,” “because/led to/resulted in,” and “now I have to.”
- Exposure definition: Choice-deprivation exposure used a strict threshold of scores ≥2 versus <2, with 72 strict and 1,410 non-strict posts.Posts could match multiple process-marker categories.
- Analysis reporting: The exploratory analysis reported counts, rates, percentage-point differences, 95% confidence intervals, and risk ratios, emphasizing effect sizes because the strict group was small.The analysis used Haldane–Anscombe correction where needed and did not center significance testing.
E.3 Results
The exploratory process-marker results provide limited support for increased causal language under strict choice deprivation, while temporal and escalation markers were absent. The authors therefore treat these findings as suggestive rather than confirmatory.
- Process-marker results: 6/72 strict-exposure posts used explicit causal connectors versus 72/1,410 other posts, with RR=1.63, 95% CI [0.73, 3.63].The strict-exposure rate was 8.3%, compared with 5.1% for other posts.
- Process-marker results: Temporal and escalation templates were not observed, with 95% upper bounds ≤4.2% and ≤0.21%.The passage attributes this pattern as likely reflecting short-form posts and a conservative lexicon.
- Interpretation: The process markers are treated as suggestive but limited support rather than confirmation of the proposed mechanism.They complement, rather than replace, the main evidence concerning choice deprivation and rights-based protest.
- Constructs and thresholds: The quantitative annotations covered choice-deprivation intensity, Protest-Rights, and a four-category process marker.The deprivation scale was later collapsed into binary thresholds to reduce boundary noise while retaining coarse severity information.
F.3 Lexicon validation against reconciled human codes
The validation compared lexicon-based measures with reconciled human codes for deprivation, protest rights, and causal process markers. Lexicons captured some explicit language but were conservative indicators, especially for choice deprivation and causal phrasing.
- Human-code reference: Two coders reconciled choice-deprivation and Protest-Rights disagreements, while coder1’s label remained the reference for the exploratory Process-Causal marker.The Process-Causal marker had moderate intercoder reliability, with κ≈0.57.
- Validation design: Lexicon outputs were compared with human codes for two deprivation thresholds, the Protest-Rights frame, and the Process-Causal marker.All comparisons used 1,482 unique posts after deduplication by Post_ID.
- Validation results: The Protest-Rights lexicon achieved κ≈0.52 with precision and recall of approximately 0.57–0.61.This indicates that it captured many clear instances of rights-based protest language.
- Validation results: Choice-deprivation and Process-Causal lexicons showed precision of approximately 0.58–0.84 but recall of only approximately 0.10–0.16.They flagged explicit formulations while missing more implicit or creatively phrased cases.
- Interpretation: The lexicon variables were treated as coarse, interpretable tools for broad associative patterns rather than high-accuracy classifiers.They complemented rather than replaced richer human-coded qualitative findings.