Source-linked AI summary
It Takes Three to Converse: Empirical Observations on How the Developer, the Convener and the Participant Shaped 119 Polis Conversations
Lodewijk Gelauff
TL;DR
The paper examines how platform configurations and convener choices shape participant behavior and Polis outputs. Using exported data from Polis conversations, it finds distinct participation patterns and reports that exported cluster assignments cannot be exactly reproduced.
Problem
Polis platforms offer configurations, including seeding and moderation policies, that may shape statement growth and participants’ experience.
Method
The paper analyzes anonymous Polis datasets downloaded as CSV files from publicly available minted reports.
Results
Participants completed more of smaller, exhaustible statement sets, while reported cluster assignments were affected by platform design choices and could not be exactly reproduced.
Takeaways & Limitations
Conveners should consciously document statement-admission policies and conduct their own clustering analysis using complete data after the process ends.
Takeaways & Limitations
Without controlled trials, the observations cannot establish whether statement availability or dosing causally changes engagement, and participant behavior may differ across cultures.
Abstract
from arXiv · showhide
Polis is a popular democratic innovation tool that allows asynchronous citizen engagement through atomic statements: short statements that together describe a complex question, inviting the citizen to vote Agree or Disagree on each. This paper uses 119 conversations with 100 or more participants and an extensive data export, drawn from a wider set of 271 collected processes. The paper asks what determines the output of such a process. Three parties shape the result. The developer of the platform has made important design choices that restrict the outcome: the number of groups the platform is able to report (restricted to 2--5) and which statements are prioritized. The convener defines the assignment: the initial statements that set the tone, the policy that accepts or rejects new statements and who can be invited. Finally, the participant works within these boundaries. With access to less than half of the generated statements, they end up responding to more statements when their conversation seems to have an achievable number of statements to complete, than when they are presented with more statements. Due to choices such as warm path clustering, the exported resulting clustering cannot be reproduced based on the voting data. Conveners may want to re-analyse their own conversations once the process is closed, to consider the data in its entirety, and make their own analysis priorities explicit.
1 Introduction
Polis enables asynchronous citizen engagement through atomic statements, while platform design, convener choices, and participant behavior jointly shape the process and its results. This paper analyzes 119 conversations with complete vote records to examine those influences and argue for post-process reanalysis.
- Three parties shape the process: Platform configurations let conveners seed statements, admit participant submissions, and choose moderation policies that affect statement growth and participant experience.Strict or permissive moderation may determine how rapidly available statements grow.
- Dataset and contribution: 119 conversations with complete vote records provide scale and documented implementation for analyzing Polis dynamics.The collection comes from 271 usable processes identified within a wider, non-exhaustive search.
- Research focus: The study examines how software design choices, convener decisions, and participant behavior affect Polis outputs.Its contribution includes examining the effects of design choices such as an exhaustible stack of statements.
- Interpreting outputs: The platform’s exported clusters represent the analysis engine’s state at export, not necessarily the only or best representation of the full conversation.The paper recommends that conveners clean the complete data and conduct their own analysis after closure.
2 Background
Polis is an open-source asynchronous consultation system in which participants vote on atomic statements and may submit new ones, producing clustered opinion landscapes. The paper situates Polis among related platforms and prior work while focusing on what consultation datasets contain and reveal about the process.
- Applications: Polis has been used in citizens’ assemblies, party platform-building, and government consultation, including deployments such as vTaiwan, Klimarat, and Aufstehen.The cited examples span governmental, movement, and broader public-consultation contexts.
- How Polis works: Conveners configure Polis conversations by selecting seeds and meta-statements, setting participation and moderation rules, prioritizing statements, and controlling dissemination.These settings determine important aspects of the participant experience and available data.
- How Polis works: Polis presents atomic statements sequentially for Agree, Disagree, or Pass responses, while participants may submit alternative statements.The platform then clusters participants based on their responses and presents the resulting opinion landscape.
- Platform family: Polis+ platforms use the same software or forks, while Polis++ platforms reimplement the method in other code.The paper distinguishes these families from interfaces that alter what participants see without changing computation.
- Previous work: Prior work describes Polis’s scaling challenges, iterative clustering mechanics, and related uses, but comparisons of what consultation datasets contain are rare.This paper therefore asks what can be learned from available datasets about the process itself.
3 Data
The paper assembles and analyses 119 Polis conversations selected from a broader collection, using exported vote, statement, participant, and clustering data. The corpus is heterogeneous, long-tailed, and analysed within explicitly defined conversation windows and bases.
- Corpus construction: The collection began with 329 raw export directories, while the wider identified corpus contains 271 usable conversations.The datasets were located through iterative online searches and archived to establish provenance for the downloads.
- Corpus construction: 119 conversations with at least 100 voters remained after excluding datasets with integrity, event-window, or participation problems.The final analysis cohort contains 6,883,979 votes; analyses use 6,540,474 votes within established voting windows.
- Analytical bases: The paper distinguishes four non-interchangeable analytical bases, including the census, collected corpus, and 119-conversation analysis cohort.The census includes leads for which no data was obtained, so it is not a superset of the corpus.
- Voting patterns: 81.7% of 21,432 statements receiving at least ten non-author votes drew more agreement than disagreement.The agreement share remained between 80.2% at one vote and 83.6% at fifty votes.
4 The developer’s choices
Polis’s developer choices shape which statements participants see and how their opinions are grouped. Routing priorities favor novelty and group-separating statements, while warm-path clustering and constrained K selection make exported groupings dependent on engine history.
- 4.1 Routing and exposure: The routing formula gives fresh statements a strong novelty advantage that decays exponentially as votes accumulate.The novelty term starts at 9 for an unvoted statement and approaches 1; the priority is then squared, with a smaller net advantage after accounting for statement importance.
- 4.1 Routing and exposure: Passes by the first few participants can sharply reduce a statement’s priority for a long time, limiting its subsequent exposure.The paper identifies this as a routing consequence of the priority formula.
- 4.1 Routing and exposure: At 10 votes, an all-agree statement’s priority outweighs an all-disagree statement’s by 121x; at 30 votes, the ratio is 961x.The priority formula is asymmetric in agrees and disagrees when extremity is held constant.
- 4.2 Clustering: The clustering engine uses a warm path, carrying forward projection axes, k-means centroids, and the group-count state from one tick to the next.The reported number of groups therefore depends on the warm path, math engine, and permitted K range.
- 4.2 Clustering: K is restricted to min(5, 2 + floor(count/12)) at both group and subgroup levels.The engine’s candidate range begins at two, so its selection cannot return fewer than two groups.
- 4.2 Clustering: The engine reproduced its own random-start result byte-identically in 97.8% of six-draw cases across 180 conversations, but this does not establish reproducibility from exported voting data.This separates invariance to the engine’s starting seed from recoverability of the reported clustering using the export.
5 The convener’s choices
Conveners shape Polis conversations through seeding, moderation, and the handling of participant statements. Their choices alter statement availability and composition, while the observed effects sometimes remain difficult to attribute uniquely to conveners or participants.
- 5.2 Seeding: Conveners seeded a median of 23 statements, representing 17.6% of the eventual statement set; every analysed conversation except one had at least one seed.The interquartile range was 16–32 seeds and 4–32% of the eventual set, with maxima of 142 seeds and 77% seeded.
- 5.2 Seeding: Opening seeds received majority agreement in 70.1% of cases, while 81.6% of statements added during the conversation reached majority agreement.At the vote level, disagreements accounted for 29.4% of opening-seed votes versus 17.0% for participant statements.
- 5.2 Seeding: The higher agreement rate for conveners’ statements cannot immediately be attributed to either conveners’ framing or participant pressure.The paper identifies the 2025–2026 uniform-prioritization software error as a possible natural experiment for examining asymmetric routing effects.
- 5.3 Statement moderation: Strict moderation was recorded for 59 of 119 conversations, or 51% of the 115 conversations with a known policy.Strict moderation acts as a whitelist, whereas permissive moderation acts as a blacklist; participants receive no feedback about moderation outcomes.
- 5.3 Statement moderation: Exhaustible conversations had 85 statements versus 208 open-ended ones, yet first-session participants cast 33 versus 27 votes in the same median 226 seconds.Participants facing fewer statements therefore voted more, and completed a much larger share of the statements available to them.
- 5.4 Mid-conversation seeding: In 61% of conversations, conveners added a median of 9 statements after the conversation began.The mean was 34 because one conversation added 1,247 statements; these mid-conversation additions received 75.6% agreement on average.
- 5.4 Mid-conversation seeding: Among matched mid-conversation seeds, 98.4% resembled exactly one earlier participant statement, but the export cannot establish that any pair was a replacement.The pattern appeared in 26 of 70 conversations that seeded mid-run, and the evidence supports a corpus-level direction rather than individual replacement cases.
6 The participant’s choices
Participants’ engagement is shaped by authentication limits, statement availability, and conversation design. Most participate briefly and once, while a minority return and spend longer on statements opposing the majority.
- Definitions: Participants are defined by identification cookies, and sessions by vote series separated by at least one hour.A vote is an opinion on one statement; a statement submission is a participant’s submitted comment or statement.
- Voting behavior: 0.190 elasticity means 10% more statements correspond to about 2% more votes per participant.The reported interval is [0.122, 0.264].
- Participation patterns: The vast majority of participants contribute only one voting session, averaging about three minutes.The paper notes that poor authentication may cause some returning participants to be recorded as new ones.
- Voting behavior: 10.1% longer response times occur for statements where participants disagree with the majority.The effect holds in 153 of 155 individually measured conversations, with similar statement controversy and length.
- Returning participants: 8.5% of participants return for a second session on average, with a median gap of 15 hours.Sessions are separated by at least one hour; the 1–3 hour returns may represent extended breaks rather than distinct sessions.
- Returning participants: Returns cluster at one-, two-, three-, and four-day gaps at 1.46, 1.60, 1.58, and 1.66 times the surrounding rate.The paper attributes these peaks sufficiently to participants returning at similar times of day, because voting activity is concentrated by hour.
7 Interpreting Polis Cluster Outputs
Polis’s exported clustering is path-dependent and constrained by the platform’s candidate group range, so the reported clusters are not necessarily reproducible or uniquely representative of final opinions.
- Limitations: The exported clustering cannot be reproduced from the export alone because only the final state, not intermediate states or computation times, is preserved.The clustering process begins each tick from the previous state, and the author was unable to recover the reported clusters meaningfully.
- Interpreting outputs: The reported clustering is where the computational path arrived, rather than the fit selected by the final votes alone.This makes the export a representation of the engine’s trajectory, not necessarily the best representation of the final opinion set.
- Platform constraints: 2–5 candidate groups bound the engine’s reported clustering, and the floor of two is the only binding constraint in the analysed conversations.The upper bound was five in every conversation analysed.
- Independent analysis: Independent reanalysis must account for differences between clustering data and exported data, including zeroed statements and reversed vote encoding.The clustering process excludes or zeroes some meta-statements and moderated statements, while database and export encodings differ.
- Independent analysis: Only participants with at least 7 votes or complete statement responses are included in the clustering results, except in very small conversations.This inclusion rule further distinguishes the reported clustering population from the full participant record.
8 Discussion and limitations
The discussion translates findings into recommendations for conveners and platform designers while stressing that the observational dataset cannot establish causality. It emphasizes documenting admission policies, recording platform behavior, and re-analyzing complete conversations with explicit clustering choices.
- Recommendations to practitioners: Admission policies should be chosen consciously, preferably in advance, and documented because fixed and automatically accepted statement sets are associated with different participant behavior.The paper states that causality remains unconfirmed, though the pattern appears plausible.
- Recommendations to practitioners: Conveners should consider additional clustering analyses after collecting the complete set of statements and votes, making their analytical priorities explicit.The exported clustering reflects design choices, so post-process analysis can support more deliberate interpretation.
- Recommendations to platform designers: Recording what participants were shown and the platform’s assignment probabilities supports outcome reproducibility, fairness checks, and evaluation of incremental improvements.The paper also recommends independently recording settings during the process to reduce uncertainty about last-minute changes.
- Recommendations to platform designers: Tracking statement lineage may improve clustering mechanisms and provide conveners with useful structured information.Initial simulations suggest that clustering results can change substantially when lineage is considered.
- Limitations: The dataset cannot support strong causal claims because it was not collected as a randomized controlled trial and includes conversations involving repeated or connected conveners.The spectrum of observed possibilities is therefore more informative than exact percentages.
- Limitations: Participant behavior also requires controlled research, and responses to the same interface may differ across cultures.The paper specifically leaves unresolved whether participants return when more statements are available or whether statement dosing improves engagement.
9 Conclusion
The conclusion shows that Polis outcomes reflect platform and convening choices as well as participant voting. Participants responded more completely in exhaustible conversations, while exported cluster assignments cannot be exactly reproduced and exports often lack timing information.
- Conclusion: Platform design choices shaped the 119 conversations’ clusters as much as participant behavior, and exact reported cluster assignments cannot be reproduced.Warm path clustering and hard limits on the number of clusters affect both cluster counts and assignments.
- Conclusion: 93% versus 15%: median participants completed this share of their available statements in exhaustible versus open-ended conversations.Median votes were 33 versus 27, with no greater time expenditure or effort reported; causal interpretation remains unresolved.
- Conclusion: 23 statements: the median convener seeded conversations with this many statements, and 70.7% of opening seeds were agreed with by a majority of voters on average.This describes the typical starting configuration and its reception by participants.
- Conclusion: Exports usually record what happened but not when, limiting conclusions about moderation, recomputation, and the statements shown to each participant.Timing-rich exports would turn many unresolved questions into measurable ones.
Declarations
The declarations describe the author’s use of contemporary LLM tools during preparation and state that the dataset will be released at final publication.
- Declarations: Contemporary Claude Code and supporting LLM tools assisted with scripts, tables, figures, and draft prose, while the author selected methods and verified the analysis pipeline.The stated workflow presents the analysis as human-directed and reproducible.
- Declarations: The dataset is planned for release at final publication.The declaration does not provide a publication date or access location.