Source-linked AI summary
The Role of User Profile for Fake News Detection
Kai Shu, Xinyi Zhou, Suhang Wang, Reza Zafarani, Huan Liu
TL;DR
Fake-news detection is difficult because fake news mimics true news and content-only methods are limited. The paper identifies representative fake- and real-news sharers, compares their explicit and implicit profiles, and tests those features for classification. User profile features consistently outperform state-of-the-art news-content features, with average F1 above 0.90 and implicit features performing better than explicit ones.
Problem
The paper addresses the limited effectiveness of content-only fake-news detection and the lack of principled understanding of how user profiles relate to fake and real news sharing.
Method
The authors measure sharing behavior, form representative user groups, compare explicit and implicit profile features, and evaluate those features with classification and importance analyses.
Results
User profile features consistently outperform state-of-the-art news-content features, achieve average F1 above 0.90, and show stronger performance for implicit than explicit features.
Takeaways & Limitations
The findings lay a foundation for deeper exploration of social-media user profiles in fake-news detection.
Abstract
from arXiv · showhide
Consuming news from social media is becoming increasingly popular. Social media appeals to users due to its fast dissemination of information, low cost, and easy access. However, social media also enables the widespread of fake news. Because of the detrimental societal effects of fake news, detecting fake news has attracted increasing attention. However, the detection performance only using news contents is generally not satisfactory as fake news is written to mimic true news. Thus, there is a need for an in-depth understanding on the relationship between user profiles on social media and fake news. In this paper, we study the challenging problem of understanding and exploiting user profiles on social media for fake news detection. In an attempt to understand connections between user profiles and fake news, first, we measure users' sharing behaviors on social media and group representative users who are more likely to share fake and real news; then, we perform a comparative analysis of explicit and implicit profile features between these user groups, which reveals their potential to help differentiate fake news from real news. To exploit user profile features, we demonstrate the usefulness of these user profile features in a fake news classification task. We further validate the effectiveness of these features through feature importance analysis. The findings of this work lay the foundation for deeper exploration of user profile features of social media and enhance the capabilities for fake news detection.
I. INTRODUCTION
Social media’s growing role in news consumption has increased exposure to fake news, while content-only detection remains difficult because fake news mimics true news. This paper investigates user profiles as complementary signals for understanding and detecting fake news.
- 62% of US adults got news from social media in 2016, compared with 49% in 2012.
- Fake news can cause people to accept deliberate lies, alter responses to legitimate news, and lose trust in the broader news ecosystem.
- Content-only detection is challenging because fake news is intentionally written to mislead readers and social-media data is large-scale, multimodal, noisy, and sometimes anonymous.
- The study asks which users share fake or real news, how their profiles differ, and whether profile features can detect fake news.
- The authors compare explicit and implicit profile features between representative user groups and evaluate their usefulness for fake-news classification.
- Average F1 exceeded 0.90, and implicit features such as political bias performed better than explicit features.
II. ASSESSING USERS’ SHARING BEHAVIORS
The study assesses sharing behavior to identify representative users and uses FakeNewsNet social-context data after filtering likely bot accounts. The resulting user groups support analysis of profile differences linked to fake and real news sharing.
- User sharing behavior is measured to identify users more likely to share fake or real news for subsequent profile characterization.
- FakeNewsNet contains Politifact and Gossipcop datasets with fact-checker labels, news content, and Twitter social engagements.
- Botometer filters users scoring above 0.5, retaining the remaining accounts as authentic human users.
- 14.2% and 13.7% of Politifact users, and 21% and 18.9% of Gossipcop users, were filtered for fake and real news respectively.
- A bigger ratio of bot users was found among users spreading fake news than among users spreading real news.
1) Absolute Measure:
The study uses absolute sharing counts and the Fake news Ratio to characterize users according to both volume and the relative share of fake news in their histories.
- Absolute sharing counts identify users who share the most fake or real news items relative to other users.
- The Fake news Ratio measures a user’s total fake-news shares divided by total shares of all news items.
- The relative measure captures users who share more fake news in their own histories even when their absolute fake-news count is not high overall.
- A larger Fake news Ratio indicates that a higher percentage of the user’s shared news is fake.
3) User Groups:
Users are grouped by sharing patterns and selected into representative fake- and real-news sharer sets. Their explicit and implicit profiles are then compared to identify differences relevant to fake-news detection.
- 3) User Groups:: Users are divided into “Only Fake,” “Only Real,” and “Fake and Real” groups according to whether they share fake news, real news, or both.
- 3) User Groups:: Top users from the “Only Fake” and “Only Real” groups are selected by the numbers of fake or real news items they share.
- 3) User Groups:: Users with lower Fake news Ratio scores are added to the real-news set using threshold t, which is set to 0.2 to reduce noise.
- 3) User Groups:: The selected users are equally sampled into U(f) and U(r), representing users more likely to share fake and real news respectively.
- 3) User Groups:: The analysis uses implicit features inferred from behavior or metadata and explicit features obtained directly from social-media APIs.
A. Implicit Profile Features
The paper compares implicit user-profile features between users more likely to share fake versus real news. Predicted age, personality, location, profile images, and political bias reveal differences with potential to differentiate the two groups.
- Feature Construction: Implicit features are predicted with widely used unsupervised tools for comparative analysis, but their prediction accuracy is not guaranteed and is not the paper’s focus.The stated goal is fair comparison of predicted features rather than validating the prediction tools.
- Age: Predicted ages differ significantly between fake- and real-news sharers, with younger users more represented among fake-news sharers and older users among real-news sharers.A t-test reports p-value<0.05 in both datasets.
- Personality: Personality profiles differ across the groups: fake-news sharers tend to have relatively low Neuroticism, suggesting less anxiety in their predicted online personalities.Personality is measured using the Five Factor Model and predicted from historical tweets with Pear.
- Profile Image: Profile-image class distributions differ consistently, with “wig” and “mask” prominent among fake-news sharers and “website” and “envelope” among real-news sharers.Object types are classified using the pre-trained VGG16 model.
- Political Bias: Political-bias distributions differ across datasets: fake-news sharers tend to be more ideologically biased, whereas real-news sharers tend to be neutral-biased.Political-bias scores are inferred from user interests and range from −1 to 1.
B. Explicit Profile Features
The study compares explicit profile attributes between users more likely to share fake versus real news. These groups differ significantly across profile, activity, and network-related features.
- B. Explicit Profile Features: The comparison covers verification status, registration time, posting and favoriting activity, followers, and following counts.Categorical features are compared by category ratios, while numerical features use box-and-whisker diagrams and statistical tests.
- B. Explicit Profile Features: 938 and 188 more verified users appear in U(r) than U(f) on PolitiFact and GossipCop, respectively.The results indicate that verified users are more likely to share real news.
- B. Explicit Profile Features: Registration-time distributions differ significantly between U(f) and U(r), according to the box plot and two-tailed t-test.The reported t-test p-value is less than 0.05.
- B. Explicit Profile Features: Users in U(f) generally publish fewer posts but perform more favor actions than users in U(r) on both datasets.The authors associate higher posting activity with users sharing more real news and greater favoring with willingness to reach out to others.
- B. Explicit Profile Features: Users in U(f) have significantly fewer followers and more following counts than users in U(r) on both datasets.For example, U(f) has 46 and 579 fewer followers on PolitiFact and GossipCop, respectively.
- B. Explicit Profile Features: Users in U(f) and U(r) show different distributions across most explicit and implicit feature fields.The authors identify these differences as potentially useful for guiding fake news detection.
IV. EXPLOITING USER PROFILES
This section investigates whether user profile features can improve fake news detection and how to construct effective models from them. It also examines feature importance and robustness across learning algorithms.
- IV. EXPLOITING USER PROFILES: The study explores user-profile-based fake news detection through feature importance and model robustness analyses.The section specifically addresses RQ3 concerning whether and how user profile features can detect fake news.
A. Experimental Settings
The experiments represent each news item by aggregated user profile features and compare them with content-based and combined representations. Evaluation uses repeated train-test splits and several classifier metrics.
- A. Experimental Settings: For each news item, the method averages concatenated profile-feature vectors from all users who shared it.Profile image features are reduced from 1000 dimensions to 10 using Principal Component Analysis.
- A. Experimental Settings: The proposed User Profile Feature vector is denoted UPF.UPF represents the aggregated user-profile representation for a news item.
- A. Experimental Settings: Performance is evaluated with Accuracy, Precision, Recall, and F1 using 80% training data and 20% testing data across five repetitions.The reported results are averages over the five runs.
- A. Experimental Settings: The comparison includes UPF, RST, LIWC, and concatenated RST-UPF and LIWC-UPF representations.RST and LIWC provide content-based features, while the combined representations test complementary information from content and profiles.
B. Fake News Detection Performance Comparison
User profile features provide effective fake news detection signals across datasets and classifiers. Combining profile and content features improves performance, while feature groups differ in effectiveness and construction cost.
- B. Fake News Detection Performance Comparison: LIWC outperforms RST among news-content-based methods, suggesting word choice captures differences between fake and real news.The authors interpret this difference from a psychometric perspective.
- B. Fake News Detection Performance Comparison: UPF achieves good performance on both datasets across all reported metrics.The result supports differences in the demographics and characteristics of users sharing fake versus real news.
- B. Fake News Detection Performance Comparison: RST-UPF performs better than either RST or UPF, indicating complementary information from news content and user profiles.The paper describes these representations as coming from orthogonal information spaces.
- B. Fake News Detection Performance Comparison: Random Forest achieves the best overall UPF performance across both datasets among the evaluated learning algorithms.The comparison includes Random Forest, Support Vector Machine, Decision Trees, and Logistic Regression.
- D. Feature Importance Analysis: RegisterTime has the highest feature-importance score at 0.937, followed by Verified at 0.099 and Political Bias at 0.063.Personality and StatusCount follow with scores of 0.036 and 0.035.
- D. Feature Importance Analysis: Feature importance analysis links RegisterTime, verification, political bias, personality, and StatusCount to distinctions between fake- and real-news sharers.The paper relates these features to account age, verification, ideological bias, personality, and user activeness.
- D. Feature Importance Analysis: Using all profile features raises PolitiFact F1 by 4.51% versus explicit features and 9.84% versus implicit features.The authors report that explicit and implicit features contain complementary information; implicit features are more effective than explicit features on GossipCop for Accuracy and F1.
V. RELATED WORK
Fake news detection methods use either news content or social context, while user-profile approaches have lacked systematic understanding and interpretability.
- Detection Approaches: Fake news detection approaches generally use news content or social context as their primary information source.Content-based methods include linguistic, visual, knowledge-based, and style-based features; social-context methods incorporate user profiles and related signals.
- Content-Based Approaches: Content-based methods extract linguistic features such as writing styles and sensational headlines that commonly occur in fake news.They also use visual features for identifying intentionally created or characteristic fake-news images.
- Content-Based Approaches: Knowledge-based models fact-check claims against external sources, whereas style-based models capture deceptive or non-objective writing patterns.
- User-Profile Approaches: Existing user-profile approaches train classifiers from extracted features without systematic understanding, making them difficult to interpret.The paper addresses this gap through an in-depth investigation of user profiles for fake-news detection.
B. Measuring User Profiles on Social Media
The paper examines explicit and implicit user profiles by identifying representative sharers, comparing their characteristics, and evaluating profile features for fake-news detection.
- Profile Features: The study extracts both provided explicit and inferred implicit user-profile features to capture different user demographics for fake-news detection.
- User Groups: Two real-world datasets with news content, social context, and reliable ground truth support analysis of users’ fake- and real-news sharing behaviors.
- User Groups: Absolute and relative sharing measures identify representative user sets more likely to share fake or real news.
- Profile Comparisons: Most explicit and implicit profile features show distinct values or distributions between users more likely to share fake and real news.The authors use detailed statistical comparisons to characterize these differences.
- Detection Evaluation: User-profile features significantly contribute to fake-news detection and remain broadly robust across different learning algorithms.The paper reports consistently good results compared with several state-of-the-art baselines.
- Detection Evaluation: A limited feature set can provide reasonably good performance when time or computational resources are constrained.