Source-linked AI summary

Disentangling Threads: Exploring the Potential of LLM-Supported Discussion Forum Analysis for Community Insight

Tony W. Li, Zhiqing Wang, Thanh-Nha Tran, Yu-Chun Grace Yen, Steven P. Dow

arXiv:2608.20591v1cs.HC

TL;DR

Online forums offer diverse, interactive community perspectives, but their tangled structure and freeform comments complicate analysis, while LLMs can miss nuance and misalign with researchers’ goals. The authors manually analyzed a forum, developed the Topic-Discourse-Reasoning Framework, built a grounded probe, and interviewed 21 researchers. Participants saw opportunities for exploratory and targeted research but wanted source-grounded, flexible analysis and follow-up support amid concerns about credibility, semantic accuracy, anonymity, and omitted detail.

  • Problem

    Forum discussions can reveal collective viewpoints, but their tangled structure and context-dependent comments make analysis difficult, while LLMs may miss semantic nuance or produce biased outputs.

  • Method

    The authors manually analyzed a forum, synthesized the Topic-Discourse-Reasoning Framework, explored LLM extraction, built a design probe, and interviewed 21 researchers.

  • Results

    Participants saw forums as useful for exploratory and targeted research, desired grounded and flexible analysis, and raised concerns about comment quality, credibility, semantic accuracy, and anonymity.

  • Takeaways & Limitations

    Community-sensemaking tools should connect flexible analytical goals to raw user data, support follow-up research, and balance commenter context with anonymous expression.

  • Takeaways & Limitations

    Findings are limited by a single subreddit thread, primarily local participants, a static human-labeled prototype, and no systematic evaluation of modern LLM workflows or consistency.

Abstract

from arXiv · show

Online discussion forums enable people from diverse backgrounds to share ideas, feedback, and perspectives. These organic discussions can help researchers understand communities' collective viewpoints, but insights are often difficult to uncover given their freeform reply structure. Large language models (LLMs) support qualitative text analysis but can misalign with researchers' analytical intent and miss key insights. To inform design considerations for forum sensemaking tools, we manually analyzed a forum discussion, synthesized an exploratory analysis framework from relevant literature, built a design probe, and interviewed 21 researchers to uncover perceived opportunities and barriers with LLM representations of collective discussions. We provide recommendations for community sensemaking tools to support flexible analytical goals grounded in raw user data and enable follow-up research processes, while balancing anonymous free expression with the desire for contextual information on commenters.

1 INTRODUCTION

Online forums offer researchers access to diverse, interactive community perspectives, but their tangled structure and freeform comments make analysis difficult. This study examines opportunities and barriers for LLM-assisted forum analysis.

  • Motivation: Forums complement interviews and surveys by exposing diverse opinions and interactive community sensemaking.They provide an accessible snapshot of evolving topics, viewpoints, and interaction dynamics.
  • Motivation: LLMs may make text analysis more efficient, but they can miss semantic nuance and produce biased outputs.These limitations are especially relevant when comments are context-dependent and shaped by community norms.
  • Motivation: Researchers must navigate related ideas scattered across sub-threads, off-topic drift, and sprawling reply hierarchies.Freeform comments may also be irrelevant or meaningful only alongside other comments.
  • Research questions: The study asks what opportunities and barriers researchers perceive for LLM-assisted discussion-forum analysis.The research questions focus on LLM support for understanding communities through forums.
  • Approach: The authors manually analyzed a forum thread, synthesized a Topic-Discourse-Reasoning Framework, assessed LLM extraction feasibility, and interviewed 21 researchers.The framework draws on prior work concerning topics, discourse roles, and reasoning types.
  • Findings: Participants valued grounded data, flexible follow-up research, and both exploratory and targeted uses, while raising concerns about comment quality, credibility, and LLM accuracy.They also identified a tension between contextual commenter information and preserving anonymous expression.

2 RELATED WORK

Prior work establishes forums as rich but labor-intensive sources of collective insight, while existing tools and LLMs provide incomplete support for nuanced, interactive analysis.

  • Forum analysis: Forums can complement interviews and surveys by offering low-cost access to diverse, unfiltered perspectives and interactive problem-solving.However, large discussions distribute information across threads and duplicate ideas.
  • Forum analysis: Commercial social-listening tools emphasize aggregate trends such as sentiment and keyword prevalence rather than collective intelligence research.They do not directly support understanding perspective ranges or designing responses to emergent community needs.
  • LLM-assisted analysis: LLMs are increasingly used for thematic analysis, codebook creation, and collaborative qualitative coding.They appear especially suited to deductive coding with predetermined codebooks.
  • LLM-assisted analysis: LLMs often struggle with semantic context and expressive language, can produce biased outputs, and have prompted calls for human interpretative judgment.These concerns make their use with freeform, subjective, and tangled forum data uncertain.
  • Research gap: The paper therefore investigates researchers’ analytical needs and perceived limitations of LLM technologies for community forum analysis.Its stated goal is to support, rather than replace, human researchers in disentangling complex discussions.

3 DEVELOPING OUR GROUNDED PROBE

The authors developed a grounded probe by manually analyzing a real forum, synthesizing the Topic-Discourse-Reasoning Framework, testing LLM extraction, and building a multi-view prototype.

  • Probe development: The probe combines manual forum analysis, literature synthesis, an LLM feasibility exploration, and a prototype interface.It was designed to ground researchers’ reflections in a concrete analytical task.
  • Analytical framework: The Topic-Discourse-Reasoning Framework organizes comments by topics, discourse roles, and reasoning types.Its discourse roles include Problem Statement, Design Proposal, and Agreement/Disagreement; reasoning types include Logos, Ethos, and Pathos.
  • Data: The study analyzed a public Reddit thread about improving the US K-12 education system, containing roughly 320 comments across several years.The data collection followed Reddit’s developer and data policies and received exempt institutional review.
  • Data preparation: Two authors inductively coded Topics and deductively labeled Discourse Role and Reasoning Type, resolving disagreements through iterative discussion.The manually labeled data populated the probe while participants imagined it had been LLM-analyzed.
  • Prototype design: The prototype presented Topic Selection, Topic Overview, Subtopic View, and Comment View for complementary levels of analysis.The interface included aggregate distributions, reply-tree structure, hierarchical replies, TDR tags, and an optional AI-summary toggle.
  • Technical feasibility: Bottom-up LLM topic generation covered 89% of manual topics, while TDR classification exceeded BERT, RoBERTa, and logistic-regression baselines on macro-F1 across all dimensions.Classification still left room for improvement relative to manual labels.

4 METHOD

The study used a one-hour semi-structured online interview and a grounded prototype to examine researchers’ interactions with forum-analysis features, followed by inductive thematic analysis.

  • Participants: The interviews involved 21 researchers with qualitative data-analysis experience, recruited through institutional channels, a subject pool, and personal networks.Most participants had also posted in online discussion threads.
  • Probe materials: The probe represented forum comments with Topic, Discourse Role, and Reasoning Type labels.The sample interface included topic and commenter summaries alongside comment-level information.
  • Probe materials: The sample discussion included policy-oriented topics such as curriculum reform, parental involvement, school integration, scheduling, and assessment metrics.The displayed topics varied in comment and commenter counts.
  • Procedure: The one-hour protocol comprised a questionnaire and warm-up, a Reddit baseline exploration, structured probe tasks, and a post-study interview.Participants thought aloud while exploring the thread and assessed visualizations, subtopics, summaries, and information needs.
  • Analysis: Two authors developed an inductive codebook by independently coding interviews, comparing interpretations, and continuing until code saturation after four iterations.They then organized roughly 350 codes into 12 groups and reached consensus on final themes through reflexive discussion.

5 RESULTS

Researchers saw forums and LLM-supported analysis as useful for exploratory and targeted community insight, but wanted flexible, source-grounded control. They also identified tensions around commenter context, anonymity, and preserving human nuance in AI representations.

  • Research opportunities: Participants valued forums for exploratory research, targeted questions, diverse perspectives, and follow-up analysis.They described using forum insights to guide hypothesis testing, revisit prior analyses, prioritize next steps, and conduct further research.
  • Analytical representations: The Topic-Discourse-Reasoning Framework helped researchers browse topics, identify discourse roles, and interpret comments through reasoning types.Participants said combining framework dimensions supported sequential analysis and discovery of useful data slices.
  • Follow-up research: Participants saw opportunities for iterative interaction loops in which researchers ask commenters targeted follow-up questions and questionnaires.They also identified the probe as a way to determine where to pursue analysis outside the forum, including related research literature.
  • Trust and control: Participants wanted analytical control and source-grounded verification because forum data is noisy and LLM outputs may miss semantic nuance.They favored representative examples, raw comments, and within-comment highlighting to validate tags and support analysis without relying entirely on summaries.
  • Context and anonymity: Researchers wanted demographic context about commenters to assess credibility and representativeness, while recognizing that identity markers could undermine anonymity and authentic expression.The desired compromise reflects tension between contextualizing contributions and preserving forums’ low barrier to entry and freedom of expression.

6 DISCUSSION AND FUTURE WORK

The study identifies opportunities for LLM-assisted forum analysis while emphasizing grounded verification, flexible research workflows, and ethical balancing of anonymity with contextual information.

  • AI for Human Understanding, Beyond Summarization: Participants valued TDR Framework breakdowns but distrusted AI summaries without easy verification against human voice and context.Tagging supported faster validation against raw comments, whereas summaries often prompted participants to read full comments.
  • AI for Human Understanding, Beyond Summarization: Future tools should keep researchers connected to raw data through relevant passages, representative quotes, and transparent, evaluable LLM criteria.LLM functionality should focus and validate researcher analysis rather than replace it.
  • Tensions Between Open Communication and Analytical Utility: Participants saw forums as useful for diverse collective discourse but wanted contextual information about commenters’ expertise, representativeness, and good-faith intent.Anonymity makes it difficult to verify commenters’ community standing and motivations.
  • Tensions Between Open Communication and Analytical Utility: Future platforms should contextualize expertise and intent without exposing identities, potentially by surfacing contribution patterns or verified professions.The authors note that communities differ in their values and require situated ethical approaches to analytical transparency.
  • Follow-up Research: Tools should support follow-up research through conversational loops, researcher comments, and triangulation with complementary sources such as interviews.Participants valued both exploratory breadth and targeted depth in forum analysis.
  • Ethical Considerations: Forum analysis must account for observer effects, resource costs, confidentiality, informed consent, and potential social consequences in public sharing spaces.These concerns become especially relevant as AI research tools become more popular.

7 LIMITATIONS

The study’s findings are limited by its narrow empirical setting, framework validation, and prototype-based treatment of LLM analysis.

  • Empirical Scope: The study analyzed one subreddit thread with 21 primarily local participants, limiting the generalizability of its findings.Broader samples and more varied discussions are needed to assess whether the findings generalize.
  • Framework Validation: The Topic-Discourse-Reasoning Framework was useful in this study but requires broader-sample validation as an analytical paradigm.The study did not establish the framework’s validity beyond the investigated setting.
  • LLM Evaluation: The prototype used human-labeled data and elicited imagined LLM use, rather than evaluating a specific model or systematically comparing modern LLM workflows.The authors demonstrated feasibility with GPT-5-mini but did not assess model consistency across workflows.

8 CONCLUSION

The paper combines forum analysis, an exploratory Topic-Discourse-Reasoning Framework, LLM feasibility work, and interviews to study LLM-assisted community sensemaking.

  • Conclusion: The study analyzed a real-world forum thread, synthesized an exploratory framework, tested preliminary LLM extraction feasibility, and interviewed 21 researchers.The interviews elicited reflections on LLM-assisted forum analysis.
  • Conclusion: The Topic-Discourse-Reasoning Framework organizes forum analysis beyond simple topic extraction by incorporating discourse and reasoning dimensions.The framework was developed from related work and used to structure the design probe.
  • Conclusion: The findings identify tensions between anonymous participation and demographic understanding, LLM semantic limitations and trust, and flexible analysis needs grounded in sources.These empirical results inform design implications for understanding communities’ collective intelligence.

A PROBE STUDY INTERVIEW GUIDE

The probe study used semi-structured online interviews to examine researchers’ forum use, qualitative-analysis experience, and desired information from a policymaking scenario.

  • Interview Format: Interviewers conducted semi-structured one-hour online interviews and asked follow-up questions beyond the interview outline.The guide covered forum use, posting behavior, qualitative data, and analysis tools.
  • Forum Experience: Participants were asked which online discussion forums they had used, why they used them, and what kinds of threads they posted.The guide also asked how many comments those threads received when applicable.
  • Sensemaking Task: A policymaker scenario asked participants what information they would seek from target users’ subreddit comments and how they would obtain it.Participants were given a thread URL and asked to explain their approach.
  • Think-Aloud Procedure: Participants were encouraged to inspect the data briefly and think aloud about what they did and observed.This elicited their immediate sensemaking strategies while working with the forum material.
  • Reflection: Interviewers asked whether participants found the desired information, what challenges arose, and what else informed their potential decision.These follow-ups probed both successful discovery and additional observations.

A.3 Probe Tasks (30 min)

The probe presented topic-based and feature-based overviews alongside filtered raw comments, then asked participants to assess their usefulness, organization, and implications for policy decisions.

  • Topic organization: The probe asked participants whether topics were meaningfully organized and how they would change or select topics across granularity views.This focused evaluation on the organization and navigability of topic information.
  • Topic and feature overviews: Participants viewed high-level topic overviews broken down by discourse role or reasoning type for each topic.These overviews were paired with feature-specific topic views.
  • Raw comment exploration: The comment view exposed individual comments as raw text and supported filtering by topic, discourse role, and reasoning type.This provided a low-level complement to the overviews.
  • Participant evaluation: Participants were asked how useful the presented information might be for policy design decisions and what they hoped to uncover.The probe also elicited desired additional functionality and decision-relevant information.

B TOPIC CODEBOOK

The topic codebook organized the analyzed subreddit discussion into categories with defined inclusion criteria. Two researchers independently developed and refined the codebook through iterative discussion.

  • Codebook development: Two researchers independently developed and refined the codebook through iterative discussion.The refinement process is described in the paper’s codebook-development subsection.
  • Topic categories: The manual codebook defined topic categories and their inclusion criteria for a subreddit thread about education policy.Table 2 presents these categories and criteria.
  • Analytical structure: The codebook served as a manual categorization structure for analyzing the education-policy subreddit discussion.Its categories were accompanied by criteria specifying what each category included.
Loading 2608.20591v1…