Source-linked AI summary
You Tweet What You Eat: Studying Food Consumption Through Twitter
Sofiane Abbar, Yelena Mejova, Ingmar Weber
TL;DR
Traditional dietary studies rely on intrusive, expensive diaries and questionnaires, motivating the search for broader evidence. The paper analyzes Twitter food mentions alongside nutrition, demographics, interests, and social networks, finding associations with state health statistics and user characteristics. It concludes that Twitter has potential for US-wide nutritional monitoring, while warning that food mentions do not necessarily indicate consumption.
Problem
Questionnaires and food diaries for large-scale dietary studies can be intrusive and expensive, while existing recipe studies lack demographic information and use coarse geographic analysis.
Method
The study analyzes 210K US Twitter users by linking food mentions and estimated calories with demographic, interest, and social-network information.
Results
The study reports predictive relationships between tweeted foods and national obesity and diabetes statistics, with correlations of r = .77 and r = .66, and finds demographic and social dietary pattern differences.
Takeaways & Limitations
The findings support Twitter as a potential source for monitoring US-wide nutritional behavior and examining demographic and social patterns in food interests.
Takeaways & Limitations
Food mentions do not necessarily represent consumption, and dining-experience detection remains difficult because the classifier produces noisy output.
Abstract
from arXiv · showhide
Food is an integral part of our lives, cultures, and well-being, and is of major interest to public health. The collection of daily nutritional data involves keeping detailed diaries or periodic surveys and is limited in scope and reach. Alternatively, social media is infamous for allowing its users to update the world on the minutiae of their daily lives, including their eating habits. In this work we examine the potential of Twitter to provide insight into US-wide dietary choices by linking the tweeted dining experiences of 210K users to their interests, demographics, and social networks. We validate our approach by relating the caloric values of the foods mentioned in the tweets to the state-wide obesity rates, achieving a Pearson correlation of 0.77 across the 50 US states and the District of Columbia. We then build a model to predict county-wide obesity and diabetes statistics based on a combination of demographic variables and food names mentioned on Twitter. Our results show significant improvement over previous CHI research (Culotta'14). We further link this data to societal and economic factors, such as education and income, illustrating that, for example, areas with higher education levels tweet about food that is significantly less caloric. Finally, we address the somewhat controversial issue of the social nature of obesity (first raised by Christakis & Fowler in 2007) by inducing two social networks using mentions and reciprocal following relationships.
INTRODUCTION
The paper asks whether Twitter can provide large-scale dietary insight beyond intrusive, expensive surveys, linking food mentions to health, demographic, interest, and social-network patterns. It reports state-level health associations and personal dietary differences, while positioning social media as a possible public-health monitoring resource.
- Large-scale dietary studies traditionally use questionnaires and food diaries, which can be intrusive and expensive.
- The study analyzes 210K US Twitter users and 502M tweets, combining food mentions with nutritional, demographic, interest, and interaction data.
- 0.77 and 0.66 are the reported correlations between food mentions and state-wide obesity and diabetes rates, respectively.
- Women generally tweet about less caloric foods than men, while areas with higher education levels mention fewer calories.
- Users mentioning cooking had a 1.5% lower chance of being obese, and food choices differed between urban and rural settings.
- Networks based on reciprocal communication and following relationships showed dietary assortativity beyond chance.
- The authors frame social media as potentially useful for monitoring dietary public-health trends and informing targeted awareness campaigns.
RELATED WORK
Prior work used social media and recipe-related data for health and food-culture analysis, but often lacked personal information or fine-grained dietary context. This paper builds a Twitter food dataset, enriches it with nutritional information, and validates food-mention detection.
- Earlier Twitter research monitored flu-like symptoms, drug side effects, tobacco use, and county-level health statistics.
- Recipe and food-search studies identified ingredient, temporal, geographic, and climatic patterns in food culture.
- Christakis and Fowler reported a 57% increase in obesity risk when a friend became obese, whereas later work contested the persistence of social-network effects.
- The dataset began with 50M keyword-filtered tweets collected through the Twitter Streaming API during one month in 2013.
- Crowdsourced labeling used 2,157 tweet examples and achieved 95.9% agreement for training food-related tweet classification.
- The authors assigned caloric values by averaging per-serving values from the top 25 nutritional search results for each food keyword.
- Food mentions exhibited weekly periodicity, peaking on Saturdays and reaching their lowest level on Mondays.
- Manual examination found 70% of sampled food-containing tweets mentioned food, including 63% about consumption and 12.5% expressing a wish.
STATE-LEVEL CORRELATIONS
The paper compares tweet caloric density with CDC obesity and diabetes statistics across the 50 states and Washington, DC. The strongest associations use all foods, while geographic clustering suggests regional food-culture patterns.
- The analysis correlates state-level tweet caloric values with CDC obesity rates from 2012 and diabetes incidence from 2005–2007.
- 0.772 is the Pearson correlation between all-food tweet caloric density and obesity, while 0.658 is the corresponding correlation with diabetes.
- Beverage caloric value correlates more strongly with both ailments than solid-food caloric value alone.
- Alcoholic-beverage caloric value correlates with obesity at 0.445 but has no statistically significant relationship with diabetes.
- Southern states cluster in the upper-right of the caloric-value versus obesity plot, with Louisiana and Arkansas at the extreme right.
- Regional clustering is interpreted as suggesting shared food culture among geographically proximate populations, while Washington, DC appears somewhat separated from nearby states.
COUNTY-WIDE MODEL FITTING
The study aggregates Twitter-derived food, language, caloric, and demographic features to predict county-level obesity and diabetes. Food-Demog outperforms the prior LIWC-demographic approach, while simpler Calories modeling remains interpretable.
- County-level correlations: 0.501 obesity and 0.447 diabetes Pearson correlations were found for counties with at least 100 users.Among counties with at least 200 users, correlations increased to 0.605 for obesity and 0.498 for diabetes.
- Feature construction: Each user is represented by LIWC categories, food names, average caloric value, and five census-derived demographic variables.The demographic variables are age-group proportions, female proportion, Afro-Hispanic proportion, and log median household income.
- County aggregation: County vectors contain 64 LIWC categories, 461 food names, and five demographic variables, weighted by the proportion of users expressing each feature.Only counties with at least 100 users are retained, unlike the prior work’s focus on the 100 most populous counties.
- Model evaluation: Six regression models predict obesity and diabetes across 346 counties using five-fold cross-validation.The models are Demog, Liwc, Calories, Food, Liwc-Demog, and Food-Demog.
- Model comparison: Food-Demog outperforms Culotta’s Liwc-Demog model, while Food significantly outperforms the aggregated Calories model for both outcomes.Calories uses only avgCal, requires no model fitting, and is described as simpler and more interpretable.
Income, Education and Gender
The analysis connects Twitter food descriptions with education, income, gender, and rural-urban context. Higher-level comparisons show meaningful differences in tweeted caloric values and food vocabularies across populations.
- Income and education: Education and income are mapped from the 2010 US Census to the zip codes associated with Twitter users.The analysis is motivated by documented variation in obesity by income, education, and gender.
- Gender: 37.2% of users were classified as female, 32.1% as male, and 30.7% as none by the Genderize API.The study supplemented name-based classification with a crowdsourced experiment for users whose gender was initially undetected.
- Robustness: 0.772 decreased to 0.741 for the caloric-value–obesity Pearson correlation after removing users without detected gender.The passage characterizes the excluded users as having limited impact on the results.
- Gender, education, and income: Gender differences appeared in the heaviness of tweeted foods but not in estimated obesity rates.Figure 3 compares predicted obesity and tweet caloric values across education and income quartiles with 95% confidence intervals.
- Rural and urban users: At p < 0.001, rural users averaged 164.8 calories in their tweets, significantly differing from urban users.Figure 4 further distinguishes rural and urban users by the foods they mention.
Interests
The study detects declared food-related interests, self-reported overweight identities, and following-based interests to examine associations with estimated obesity. It also compares these observations with prior Facebook findings and identifies a sports-related interpretive boundary.
- Declared interests: Keyword lists detect users mentioning cooking, dieting, organic food, health, or family-related identities.These lists target interests users feel comfortable declaring in their profiles or text.
- Overweight self-reference: 10,797 users were detected using at least one overweight-related hashtag such as #fatgirlproblems or #fatguyproblems.The hashtags are typically used in self-reference.
- Profile-factor associations: Table 2 reports obesity-rate differences between users with and without detected profile factors, using 128,487 users with identified gender.Positive differences indicate higher estimated obesity for the factor-present group.
- Following-based interests: Interest scores are computed from the prominence of WeFollow users followed across 61 areas, then entered into a linear regression model.A user counts as interested when the aggregate prominence score reaches the mean among potentially interested users in that area.
- Comparison and limitation: The findings partly confirm television-related obesity associations but diverge from prior Facebook results for activity-related sports interests.The paper notes that separating watching sports from participating in sports could clarify the discrepancy.
SOCIAL NATURE OF FOOD
The paper examines whether social interactions relate to eating behavior by using tweet text and follower-network relationships. It operationalizes eating behavior through predicted obesity and diabetes, food-mention frequency, and tweet caloric value.
- Social relationships: Social circumstances are examined as potential influences on food consumption and food choice.The analysis considers relationships expressed both in tweet text and in users’ follower networks.
- Network construction: Two relationship networks are used to study associations with users’ eating behavior.The passage introduces networks based on social interactions and follower-network structure.
- Behavioral measures: Eating behavior is operationalized through predicted obesity and diabetes probabilities, food-mention frequency, and tweet caloric value.These measures connect social-network relationships to multiple dietary and health-related outcomes.
User-level obesity, diabetes and food frequency
The analysis examines whether social relationships correspond to users’ estimated obesity, diabetes, and food-related tweeting, using Friendship and Mention networks. Activation probabilities rise with more high-scoring friends, and removing replies and retweets leaves the activation curve unchanged.
- User-level obesity, diabetes and food frequency: Users’ obesity and diabetes scores are estimated from food names, while county rates are assigned for model training and prediction.A Ridge regression learns individual-level scores from food-related tweeting; demographic variables are deliberately excluded from the social analysis.
- User-level obesity, diabetes and food frequency: Friendship Network links reciprocal followers, whereas Mention Network links users when one has mentioned the other.The two networks represent structural and behavioral aspects of Twitter, respectively.
- Obesity and Diabetes Activation (Spread): Users above the 90th percentile of obesity or diabetes scores are labeled active, and each user’s number of active friends is counted.This operationalizes a threshold-style social activation analysis.
- Obesity and Diabetes Activation (Spread): Activation probability increases with the number of active friends, particularly through four active friends, although uncertainty grows beyond four.The outcome is based on strongly increased obesity probability estimated from tweeted food names.
- Obesity and Diabetes Activation (Spread): Removing replies and retweets produces the same activation curve, although offline motivational effects cannot be excluded.This test addresses the possibility that users mention foods because their friends posted them.
Clique-ness Analysis
The clique-ness analysis measures social connection strength through shared-friend overlap and tests whether stronger ties correspond to more similar food-related tweeting. Highly related users are often located in shared social settings, such as schools or universities.
- Clique-ness Analysis: Link strength is measured with Jaccard similarity of users’ friend sets in both Friendship and Mention networks.The analysis uses a user-friend graph containing approximately 180M links from up to 5,000 collected friends per user.
- Clique-ness Analysis: More than 71.6% of Mention Network links have strength scores below 0.1, showing a heavily skewed similarity distribution.Users and links are assigned to Jaccard-score bins before comparing food-related tweeting.
- Clique-ness Analysis: Figure 6 bins user pairs from least to most similar by ego-network overlap and correlates each user’s food-tweet fraction with that of friends.The final [0.2, 1.0] bin contains only 5,428 Mention users and 5,076 Friendship users, so intervals are not normalized.
- Clique-ness Analysis: Manual examination finds highly related users in shared social locales, including students at the same university or school.The authors also relate the pattern to an exposure curve in which interest rises with exposures and later decreases after overexposure.
DISCUSSION & FUTURE WORK
The discussion presents social media as a low-cost, timely source for dietary and public-health monitoring, while emphasizing sampling, measurement, representativeness, and causal limitations. Future work includes more accurate user characterization and targeted interventions.
- DISCUSSION & FUTURE WORK: Social-media data can be collected and analyzed in days or hours, and can provide rich information about users’ hobbies and interests.These advantages are contrasted with the higher cost and slower timelines of traditional surveys.
- DISCUSSION & FUTURE WORK: Linking inferred obesity likelihood with interests could help target population segments and inform context-aware public-health messages.The paper presents automated dietary assessment as a prerequisite and describes this work as an initial step.
- DISCUSSION & FUTURE WORK: The sample overrepresents affluent and tech-savvy neighborhoods, with average household income of 85,117 versus the US average of 51,017.Bachelor-degree attainment is also slightly higher in the sample, 23.71% versus 22.23% nationwide, although aggregation may affect the comparison.
- DISCUSSION & FUTURE WORK: Hand-crafted keyword filters offer high precision but may be brittle under vocabulary changes and have low recall.More accurate identification of overweight users would enable finer-grained validation beyond the current state-level validation.
- DISCUSSION & FUTURE WORK: Food mentions do not establish consumption, and identifying actual dining experiences is difficult: annotator agreement was 78% and the classifier was noisy.Determining the exact nature of food mentions is left for future work.
- DISCUSSION & FUTURE WORK: Users tweet about food frequently enough to suggest they are not focused only on special occasions, with a 1.2-tweet weekly overall median.The top 1% have a median of 18.8 food tweets per week.
- DISCUSSION & FUTURE WORK: The analysis reports correlations rather than causation, so controlled experiments are needed before stronger causal conclusions.The authors identify targeted Twitter interventions as a possible direction for public-health awareness campaigns.
- DISCUSSION & FUTURE WORK: Direct social influence through verbal and non-verbal interactions remains beyond the study’s scope and is identified as future work.Related work examined dissemination of pro- and anti-anorexia content across emerging social networks.
CONCLUSIONS
The paper uses Twitter to monitor US-wide nutritional behavior and relates tweeted foods to health, demographic, interest, and social-network measures. Food mentions predict state obesity and diabetes statistics, while users with more shared friends show more similar food interests.
- CONCLUSIONS: Foods mentioned in daily tweets predict national obesity and diabetes statistics, with r = .77 and r = .66 across 50 states and Washington DC.The paper presents these associations as evidence of Twitter’s potential for public-health research.
- CONCLUSIONS: Tweeted calories are linked to user interests and demographic indicators, and users sharing more friends are more likely to show similar food interests.The conclusion calls for more sensitive and accurate tools using textual and social-network information.