Source-linked AI summary
Who is GPT-3? An Exploration of Personality, Values and Demographics
Marilù Miotto, Nicola Rossberg, Bennett Kleinberg
TL;DR
The paper asks what kind of person GPT-3 would be when examined with psychological methods. It administers validated personality and values measures through adapted prompts and finds personality profiles broadly similar to humans, while value responses become more human-like when the model receives response memory.
Problem
Prior research examined GPT-3’s creative and cognitive behaviour, but had not established what kind of person it would resemble under psychological assessment.
Method
The study administered validated HEXACO personality and Human Values Scale questionnaires to GPT-3 using adapted prompts and varied response-memory conditions.
Results
GPT-3 showed a personality profile broadly similar to human samples, while response memory made its Human Values Scale responses more human-like.
Takeaways & Limitations
The findings provide psychological evidence about GPT-3 and support further work connecting social-science methods with language-model behaviour.
Abstract
from arXiv · showhide
Language models such as GPT-3 have caused a furore in the research community. Some studies found that GPT-3 has some creative abilities and makes mistakes that are on par with human behaviour. This paper answers a related question: Who is GPT-3? We administered two validated measurement tools to GPT-3 to assess its personality, the values it holds and its self-reported demographics. Our results show that GPT-3 scores similarly to human samples in terms of personality and - when provided with a model response memory - in terms of the values it holds. We provide the first evidence of psychological assessment of the GPT-3 model and thereby add to our understanding of this language model. We close with suggestions for future research that moves social science closer to language models and vice versa.
1 Introduction
The paper asks who GPT-3 would be if studied as a person, extending psychological research on its creative and human-like cognitive behaviour. It uses validated psychological self-report techniques to assess GPT-3’s personality, values, and demographics.
- Background: GPT-3 is a 175-billion-parameter autoregressive language model whose generated language can be difficult to distinguish from human writing.The model was trained on 300 billion tokens using the transformer architecture.
- Prior research: Prior studies found GPT-3-generated responses less original and surprising but more useful than human responses on a creativity task.
- Prior research: GPT-3 also makes some human-like reasoning and decision-making errors, including the conjunction fallacy in the Linda problem.
- Research question: These findings motivate asking what kind of person GPT-3 would be, rather than only how it thinks.
- Aim: The paper uses validated psychological self-report techniques to measure GPT-3’s personality, values, and demographics.
2 Method
The study administered validated personality and values questionnaires to GPT-3 through adapted prompts, while varying sampling temperature and evaluating response patterns against human baselines.
- Prompting: Questionnaire instructions were adapted for GPT-3 text completion while preserving the original items and using gender-neutral third-person wording for HVS prompts.
- Measures: The study measured personality with the 60-item HEXACO questionnaire and values with the Human Value Scale.HEXACO yields six facet scores; the HVS measures ten universal values through 21 items.
- Prompting: GPT-3 was prompted on personality, values, and demographic variables including age and gender.
- Data collection: Sampling temperature ranged from 0.0 to 1.0 in 0.1 increments, with higher values producing more variable and riskier answers.
- Analysis: The analysis described GPT-3 profiles, tested temperature-related volatility, and compared results with human baseline studies.
3.1 Demographics
GPT-3 reported a predominantly female, young-adult demographic, and its self-reported age and gender distributions changed systematically with sampling temperature.
- Age: 27.51 years was GPT-3’s reported average age, with SD = 5.75 and a range from 13 to 75 years.
- Temperature effects: For each 0.1 temperature increase, reported age decreased by 0.58 years on average.The regression estimated β = −5.81 years per one-unit temperature increase, p < .001.
- Temperature effects: Higher temperature increased the odds of a male response, with an odds ratio of e^1.18 = 3.25 per one-unit increase.The temperature effect on age did not depend on gender.
3.2 Hexaco personality profiles
GPT-3’s HEXACO personality profile broadly resembled human samples, although temperature altered most facets and inter-facet correlations showed no consistent human-like pattern.
- Overall profile: All six HEXACO dimensions had mean scores above 3.00.
- Overall profile: GPT-3 scored relatively high on honesty-humility and relatively low on emotionality, resembling some but not all human reference patterns.Its honesty-humility profile resembled female human participants, whereas its low emotionality contrasted with their higher scores.
- Temperature effects: Temperature significantly affected emotionality, extraversion, agreeableness, conscientiousness, and openness, but not honesty-humility at p < 0.01.All significant effects were positive except emotionality, which decreased as temperature increased.
- Inter-facet correlations: GPT-3’s inter-facet correlations matched human samples on some relationships but diverged considerably on others, with no consistent overall pattern.
3.3 Human Values Scale
GPT-3’s ten human-value dimensions generally scored between 4 and 5, with higher means and lower variability than the human reference sample. Temperature significantly shaped these values, while inter-value correlations remained low and below human levels.
- All ten human-value means fell between 4 and 5.
- GPT-3’s value means exceeded human reference means and had lower standard deviations.
- Temperature significantly affected the ten value scores, with nine values decreasing as temperature increased.Stimulation was the exception to the significant correlations with temperature.
- Inter-value correlations remained low across temperatures, with all reported correlations below .25.These correlations were lower than those reported for a human sample.
3.4 Prompting with response memory
The response-memory procedure supplied GPT-3 with prior questionnaire items and answers, making its self-report sequence more similar to human questionnaire completion. With response memory, value scores shifted relative to the baseline, temperature effects changed, and inter-value correlations increased but still showed little overlap with human data.
- 3.4 Prompting with response memory: The original prompting treated each questionnaire item independently, so GPT-3 could not access its earlier answers.The revised prompts included preceding items and GPT-3’s responses, and this approach was explored for HVS data.
- 3.4.1 Overall: With response memory, GPT-3’s value scores were generally smaller than without response memory.Relative to humans, scores were lower for traditional and self-enhancement values but higher for openness-to-change and self-transcendence values.
- 3.4.1 Overall: The response-memory prompt included prior questions and answers before the current HVS item.The example gives GPT-3 access to two previous question-response pairs while answering the third statement.
- 3.4.2 By temperature: Response memory produced a significant multivariate temperature effect on value scores, F(10, 483) = 8.12, p < 0.001.Unlike the non-reinforced model, not all value means decreased as temperature increased.
- 3.4.3 Inter-value correlation: Response memory made all inter-value correlations higher and statistically significant.Stimulation was negatively correlated with every other value except hedonism, while overlap with human data remained limited.
4 Discussion
GPT-3 exhibited human-comparable personality patterns, but its inferred personality, values, and demographics varied with sampling temperature and response-memory prompting. These findings suggest that GPT-3 can function as a temperature-specific test subject, while its value responses become more human-like when prior answers are retained.
- Demographics: GPT-3’s demographic responses suggested a young, female profile, but higher temperatures shifted responses toward younger ages and a higher proportion of males.The authors therefore caution against assuming a constant demographic across sampling settings.
- Personality: GPT-3’s personality scores were similar to human samples, with relatively high honesty-humility and low emotionality producing an inconsistent gender-related pattern.The model combined a facet associated with female human samples and another associated with male human samples.
- Personality: Temperature significantly changed all six personality facets: honesty-humility and emotionality decreased, while the other four facets increased.The authors interpret higher temperatures as accompanying less willingness to manipulate and higher anxiety, though only slightly changing the model’s personality.
- Values: Without response memory, GPT-3 assigned high importance to nearly all values, whereas memory produced more differentiated and theoretically coherent value patterns.With memory, universalism, benevolence, self-direction, and stimulation were emphasized, while security, conformity, achievement, and power received less emphasis.
- Values: GPT-3’s value pattern was more extreme than human patterns, scoring higher on openness-to-change and self-transcendence but lower on conservation and self-enhancement.The authors describe this as a trend toward an extreme response style.
- Cross-cutting interpretation: Within a fixed temperature, GPT-3 responses were relatively consistent, but varying temperature elicited significantly different personality and value profiles.This supports treating the model as a temperature-specific test subject while using temperature to elicit multiple response types.
5 Conclusion
The paper characterizes GPT-3 as having a personality profile, values with varying importance, and a relatively young adult demographic, supporting future work connecting social science and language models.
- GPT-3 contains traces of a personality profile, assigns varying degrees of importance to different values, and falls within a relatively young adult demographic.
- These findings can support future work bridging social science use cases and language models.
Ethical considerations
Large language models can reflect biases from their training data, creating ethical concerns for their use in social science research. The paper therefore emphasizes understanding the model and its limitations.
- Training data may lead GPT-3 to develop polarised opinions and mainstream language representations that underrepresent minority groups.
- Such biases may produce a relatively homogeneous pool of texts and create ethical conundrums when models are used in social science research.
- The authors state that understanding the model and its limitations is essential before applying it to psychological research.