Source-linked AI summary
Rules Before Oracles: Auditable, User-Configurable Argument Selection for Deliberative Polling
Muntaser Syed, Markus Zanker, Marius Silaghi
TL;DR
Large deliberative-poll corpora require a mechanism to choose each voter’s argument slate, but opaque rankers make that exposure difficult to recompute or contest. This paper proposes a published, voter-configurable rule over public argument links, evaluates it across multiple slate measures and attacks, and finds bounded selection gains with stronger benefits on ordering, endorsement mass, and realistic authoring.
Problem
The paper asks whether argument exposure can use a published rule over publicly recomputable evidence, because opaque learned rankers prevent voters from recomputing or contesting the slate shaping their vote.
Method
The paper formalizes bipolar justification sets and three slate measures, defines seven admissibility criteria, and evaluates a one-hop reversed endorsement-flow rule in seed-controlled agentic simulations and coordinated attacks.
Results
Across roughly 17,000 seeded runs, the rule falls 0.035 short of a label-reading ceiling, matches random selection on non-degenerate set coverage, and leads on order-sensitive exposure and endorsement mass.
Takeaways & Limitations
The results make policy configuration a voter-facing coverage-versus-mass choice while locating robustness in a published author-count normalization rather than detection.
Takeaways & Limitations
The electorate is simulated, so the reported quantities describe modelled agents rather than human populations; the degenerate-authoring result also depends on the simulator’s adoption rates.
Abstract
from arXiv · showhide
In a deliberative poll, once submissions outnumber what anyone will read, some mechanism chooses which arguments each voter sees, acquiring much of the decision; practice delegates it to opaque learned rankers, so a voter cannot recompute or contest the exposure that shaped their vote. We ask whether it can be a published rule over publicly recomputable evidence with parameters held by the voter, treating legibility as an admissibility condition on usable mechanisms, not an objective traded against accuracy. We formalise a poll over bipolar justification sets, judging a slate by reason coverage, the order it arrives in, and captured endorsement mass; we give seven checkable criteria for a civic recommender and a rule meeting them: a one-hop reversed endorsement flow parameterised by a relation-weight function. An agentic simulator records every slate at every vote, over about 17,000 seed-paired runs. Served slates fall 0.035 short of a label-reading ceiling upper-bounding every selection procedure, opaque ones included: any unconstrained ranker's advantage is bounded and small. On coverage alone, with non-degenerate authoring, the rule is indistinguishable from a random slate, a null due to an order-blind, charity-blind instrument; on the other two it leads at every prefix by a margin widening with adversarial pressure and dominates on mass by a factor of 3.3. Once a realistic fraction of submissions carries no reasons, the coverage margin returns and grows. Label-homogeneous flooding collapses completeness from 0.81 to 0.34 under a flat weight policy, only to 0.44 under author-count normalisation, making the weight function a security control worth 10% of completeness. The choice between ranking arms is a position on a coverage-versus-mass frontier, not a fact, the kind of choice only a legible rule can hand to the person it affects. It maps onto an open-source peer-to-peer platform.
1 Introduction
The paper frames argument selection as a democratic-procedure problem: civic slates should be produced by published, recomputable rules over public evidence, not opaque ranking. It formalizes three ways to judge slates and previews results showing that coverage alone can miss benefits visible in ordering, endorsement mass, and realistic authoring.
- The exposure problem: Every electorate needs a selection step because voters cannot read all submitted reasons, making exposure part of the decision process.The paper distinguishes the public, checkable tally from the usually opaque slate of arguments that shaped votes.
- The position taken: The proposed standard treats legibility as an admissibility condition: rules should be published, recomputable from retrievable evidence, and configurable by each voter.The paper explicitly does not claim that opaque rankers perform badly; it argues that opaque ranking is unsuitable for a reconstructible public record.
- The formal approach: The framework evaluates slates by coverage of live reasons, prefix-sensitive ordering, and captured endorsement mass rather than by one objective alone.It models an alternative-based poll over bipolar justification sets and uses a greedy cover only as an evaluation ceiling.
- The exposure problem: Three harms motivate the problem: temporal inequity, silent narrowing, and unfalsifiable influence from opaque learned selectors.At the reference configuration, within-run completeness varies about 0.21, five times the variation between runs.
- Previewed findings: On non-degenerate authoring, coverage does not differ significantly from random selection, while ordering and other measures show strong benefits over controls.The paper attributes the coverage null to the order-blind, charity-blind instrument and reports benefits under realistic degenerate authoring.
- Implementation direction: The design maps onto DDP2P, an open-source peer-to-peer platform whose independent replicas can turn several procedural promises into structural properties.The paper presents this architecture as a way to reduce dependence on an operator-controlled server.
2 Background and Related Work
The paper connects deliberative polling, argumentation, social choice, recommendation, civic platforms, and peer-to-peer systems. Its distinctive focus is upstream of aggregation: determining what voters see before they form positions, while using agentic simulation with an explicit behavioral caveat.
- Research junction: The work joins five literatures: deliberative polling, bipolar argumentation, computational social choice, recommendation manipulation, and peer-to-peer systems.Large language models used as simulated populations provide the evaluation register rather than a separate substantive literature claim.
- Deliberative polling: Deliberative polling demonstrates opinion change after balanced exposure and discussion, but facilitated cohorts of a few hundred do not obviously scale nationally.Digital deployments remove the facilitation bottleneck while creating the exposure problem for unread corpora.
- Bipolar argumentation: Bipolar argumentation extends attack graphs with support relations, providing the formal language for the paper’s typed justification links.The related literature makes alternative readings of support explicit through labelled bipolar frameworks.
- Computational social choice: Unlike downstream aggregation work, this paper studies what a voter is shown before forming a position; existing aggregation rules can sit downstream of its slate mechanism.This establishes complementarity with liquid-democracy and representative-committee approaches.
- Argument mining: Argument-mining methods would produce the paper’s reason labels and relations in deployment, but the serving rule assumes no better semantic interpretation than those stored labels.The paper keeps text interpretation off the serving path.
- Civic platforms: Existing civic platforms surface representative or common-ground statements, yet the paper’s objection is categorical: unreconstructible mediation is unsuitable for binding processes regardless of measured quality.This contrasts operator-trusted clustering and interactive value inference with the paper’s admissibility requirement.
3 The Object: An Alternative-Based Poll and Three Ways to Judge a Slate
The paper models a binary poll whose justifications are connected by typed rebuttal and reinforcement links, then evaluates each voter’s slate by coverage, ordering, and captured endorsement mass. It also frames selection as a reproducible, configurable civic procedure and proposes an endorsement-based rule whose limits include NP-hard coverage optimization and cold-start invisibility.
- 3.1 Polls, sides, and justifications: The poll contains finite supporting and opposing justification sets, typed rebuttal and reinforcement relations, civic weights, and link weights.Civic weight is defined as public endorsement count, while relation weights let participants configure the relative importance of rebuttals and reinforcements.
- 3.2 Reason vocabularies and the live denominator: Reason labelling assigns each justification a non-empty subset of a finite vocabulary, with labels stored at authoring time for evaluation.The denominator is restricted to reasons that existed at the relevant time, separating recommender scoring from temporal inequity.
- 3.3 Three instruments: coverage, order, and endorsement mass: Coverage is order-blind and charity-blind, so it cannot distinguish rankings that return the same items and may resemble random selection when the live reason universe is only slightly larger than the budget.When every submission is a well-formed argument, items have similar coverage value and the measure offers little discrimination.
- 3.3 Three instruments: coverage, order, and endorsement mass: The endorsement-mass instrument measures the share of maximum achievable endorsement weight captured by a slate, and the paper treats the instruments as a configurable frontier.The reported comparison gives the random baseline a 3.3-fold disadvantage on this axis.
- Coverage optimization and admissibility: Coverage optimization is NP-complete, while greedy cover achieves at least (1−e^-1) ≈0.632 of optimum but violates semantic abstinence and evidence locality.The greedy procedure reads labels and inspects every candidate, so it serves as an evaluation oracle rather than an admissible mechanism.
- Reach and rule evaluation: The endorsement rule can be evaluated from local endorsement information, but brand-new items with no endorsements or endorsed incoming links remain invisible until authoring-side exploration surfaces them.This partial-replica property lets participants recompute scores themselves while defining a concrete cold-start boundary.
- A worked example: On the worked example, the rule captures endorsement mass 77/77 = 1.00, compared with 49/77 = 0.64 and 15/77 = 0.19 for the alternative slate.The example also shows that link terms may leave adjacent endorsement rankings unchanged on a uniformly competent corpus.
4 The Charter: Criteria Before Objectives
The section treats legibility as an admissibility condition for civic selection mechanisms, requiring published, recomputable, contestable procedures before ranking quality is compared. It specifies seven operational criteria and presents a simple endorsement-flow rule with voter-held parameters.
- The charter: The charter defines seven admissibility criteria: determinism, evidence locality, author blindness, semantic abstinence, reproducibility, contestability, and configurability.Each criterion has an implementation test and an observable failure condition.
- Why opacity fails: Recomputation, attribution, explanation, contest, and historical reconstruction are concrete operations enabled by published weights over public evidence.Under the rule, disputes can terminate in a tally error, signed-link error, or openly contestable policy choice.
- Why opacity fails: Opaque learned rankers fail four criteria by construction, rather than because of implementation quality.The section argues that better ranking cannot compensate for missing recomputation, attribution, explanation, contest, or reconstruction.
- The rule: The proposed rule scores items from endorsement counts and one-hop reversed links, crediting endorsements attached to items they point toward.Rebuttals to heavily endorsed objections and reinforcements of heavily endorsed claims are thereby promoted.
- The rule: Changing the relation-weight function alters completeness by 0.08–0.12 under coordinated attack, while changing α and β is indistinguishable from noise without attack.The relation-weight policy is therefore the rule’s security-sensitive design choice.
- Configurability: Voter-held policy vectors let participants choose coverage-versus-endorsement-mass positions while preserving a shared corpus, link set, labels, and rule.Persisting each policy vector with each slate makes different configurations reproducible and contestable.
5 Abas: An Agentic Instrument for Auditable Deliberation
Abas is an agent-based simulator and persistent audit record for studying deliberative exposure under a real poll protocol. It reconstructs every served slate and its score decomposition, while explicitly limiting claims to mechanism behavior under the stated agent model.
- Instrument: Abas instantiates constituent agents, runs the poll protocol to completion, and persists enough state to reconstruct every served slate.The simulator records the electorate, submissions, links, ballots, parameters, seed, policy vector, and served order.
- Agent model: Agents hold opinions in [−1, 1], generate position texts, and use TF–IDF cosine similarity to judge which served material speaks for them.Their ballot direction is stochastic, so agents near the centre remain uncertain.
- Agent model: The simulator measures exposure rather than persuasion because agents do not update opinions in response to served material.Results are claims about mechanism behavior under a stated constituent model, not evidence about human constituent behavior.
- Protocol: Adoption is restricted to served items, coupling recommendations to endorsement tallies and making the system a feedback loop.New corpus labels are drawn independently of the slate, allowing ranking arms to see bit-identical corpus growth in the ablations.
- Comparators: The experiments compare ranking arms that differ only in the ranking step, including the published rule, endorsement-only, enhancement-only, attack-only, and random selection.Random selection uses an independent generator so stances, actions, and link choices remain seed-paired across arms.
- Audit record: Persisted slate order and score decompositions make reproducibility, contestability, ordering, and endorsement-mass measurements directly testable.The browser exposes served slates, component scores, contributing links and authors, ballots, and counterfactual policy outputs.
6 Experimental Setup
The experimental setup fixes a reference poll, varies structural and adversarial factors in seed-paired simulations, and defines the attack capabilities and statistical conventions in advance. It also makes explicit that authentication is assumed and that multiple-testing correction is incomplete.
- Reference configuration: The reference configuration uses N=1000, K=20 per side, πown=0.10, πreuse=0.50, πabst=0.40, πlnk=0.60, and author-normalised ωr.The resulting regime is non-degenerate: vocabulary saturates, the corpus is roughly 50 items per side against 30 labels, and the ballot split is near even.
- Experimental factors: The sweeps vary authoring rate, slate size, link rate, and electorate size one at a time around the reference point.The ranking ablation compares five arms across K∈{5,10,20,30} using 4000 seed-paired runs.
- Adversaries: Attack experiments assume valid authentication, while coalitions control what members write, where they link, and when they act.Coalition sizes range from 0% to 25% of N, with members spread evenly through processing order.
- Adversaries: The attack families include hub-riding, label flooding, heterogeneous controls, and co-signed flooding.Co-signed flooding varies group size C∈{1,2,5,10,25,50} under both relation-weight policies.
- Adversaries: The relation store permits repeated assertions of the same link by distinct authors, so co-signing remains an available attack strategy.The setup does not defend against this strategy by excluding repeated assertions from representation.
- Statistical conventions: Comparisons use seed-paired tests for shared seeds and Welch’s test otherwise, with intervals reported as across-seed standard deviations.The manuscript reports raw p-values for many contrasts and treats p≈10^-2 cautiously because no manuscript-wide multiplicity correction was applied.
- Statistical conventions: At πown=0.02 and πown=0.30, vocabulary saturation is anchored near 800 and 150 constituents, respectively, in runs of N=1000.The figure’s intermediate curve shape is illustrative; raising authoring rate reduces voting against an unfinished reason space.
7 Results I: An Attack-Free Electorate
Under non-degenerate authoring, completeness is driven mainly by corpus growth and voter timing, while label-free selection loses only a small amount to a label-reading ceiling.
- Reference performance: 0.806 ± 0.038 mean combined completeness on BRA and 0.804 ± 0.032 on UBI show that reference runs cover about four-fifths of existing reasons.These are averages over the reasons available when constituents voted.
- Reference performance: 0.213 within-run completeness spread is more than five times the between-run variation, making arrival time the dominant source of exposure differences.Early voters encounter a reason vocabulary that is still assembling.
- Levers affecting coverage: 0.44 at K=5 to 0.83 at K=30 is the completeness range for slate size, while slots beyond the twentieth add only two to three points per block.Coverage is bounded by the authored corpus, and later gains arise from lower-ranked, reason-novel items.
- Levers affecting coverage: 0.70 to 0.88 is the completeness increase when π_own rises from 0.02 to 0.30, as higher authoring rates saturate the vocabulary sooner.Electorate size shows the same direction, rising from 0.69 at N=200 to 0.90 at N=5000.
- Levers affecting coverage: 0.004 is the completeness movement across a fivefold change in link rate on both propositions under competent authoring.Links matter for discrimination once the corpus contains material worth distinguishing, but not for completeness in this regime.
- Selection ceiling: 0.031 ± 0.014 separates the endorsement rule from the label-reading ceiling on end-of-round completeness, with the gap positive in all ten seeds.Against the live vocabulary, the ceiling reaches 1.000 and served slates reach 0.968 ± 0.015; the remaining 0.163 shortfall is temporal rather than selectable.
8 Results II: What the Coverage Measure Could Not See
The coverage measure initially finds no advantage over random selection because it ignores order and every authored item is useful, but other instruments reveal meaningful ordering and mass gains. These gains depend on corpus heterogeneity and inherit the simulator’s assumption that degenerate items receive fewer endorsements.
- 8.1 The null result, for non-degenerate authoring: Under non-degenerate authoring, ranked arms differ by no more than 0.004, and the full rule often produces the same slate as pure endorsement count.Only one of twenty-four ranked-arm contrasts survives Bonferroni correction; at K=5, enhance-only leads pure endorsement count by 0.0039 (p=0.001), while 58 of 100 BRA seeds are bit-identical at K=5.
- 8.1 The null result, for non-degenerate authoring: The null occurs because completeness ignores slate order, while link bonuses are too small to change slate membership at the K-boundary.The largest instrumented bonus was 0.96 versus a raw-endorsement boundary gap of 3, so links reorder served items without changing their union.
- 8.1 The null result, for non-degenerate authoring: Under attack, the rule trails random coverage: −0.0385 at a tenth-electorate coalition and −0.0706 at a quarter-electorate coalition.The passage attributes this to ranked arms reading coalition-inflated endorsement tallies, allowing flooded clones into the slate.
- 8.2 Order: what a ranking rule is actually for: At every prefix before the full slate, the endorsement rule leads uniform selection, with larger gains under stronger coalitions.At a quarter-electorate coalition, gains include p1 +0.0335, p5 +0.1387, AUC +0.0796, RDC +0.1788, and e90 −5.69 positions.
- 8.2 Order: what a ranking rule is actually for: At a quarter-electorate coalition, ranked reading reaches nine-tenths coverage after 8.9 items versus 14.6 for uniform selection, despite equal final coverage.The coalition’s material fills the corpus but does not occupy early positions; finite-attention readers therefore experience different slates.
- 8.3 Degenerate authoring: giving the rule something to discriminate against: With degenerate submissions, the coverage advantage returns and grows: the rule leads random by 0.0059 at G=0.1 and 0.0471 at G=0.7.At G=0.7, the rule serves 66.6% degenerate material versus 70.5% for a uniform draw, while its ordering metrics also favor the rule.
- 8.4 Why the link terms square the discrimination: Link terms contribute +0.0617 at G=0.7, representing 93% of the full margin over uniform selection, because they amplify endorsement-based discrimination.Genuine items lead reason-free items by factors of 7.7 on endorsements and 7.2 on link credit; the factors multiply through target endorsements and link patterns.
- 8.4 Why the link terms square the discrimination: The degenerate-authoring result depends on the simulator’s adoption model, where degenerate items receive 1.43 endorsements versus 11.04 for other items.The paper states that this is a property of modeled agents rather than the rule, and does not estimate how human electorates discriminate.
9 Results III: A Coalition in the Electorate
Coalition behavior depends sharply on what it authors and how the rule weights authorship. Ordinary hub-riding leaves completeness unchanged, whereas label-identical flooding substantially reduces it, with author normalisation retaining more coverage than flat weights.
- Hub-riding: Ordinary hub-riding leaves completeness unchanged across coalition sizes and propositions.With 200 seeds per cell, changes from the zero-attacker baseline were statistically indistinguishable from zero, and the two weight policies differed by less than 0.002.
- Hub-riding: The one-hop rule gives hub-riding a bounded visibility bonus without compounding effects.Because the rule computes one hop and no fixed point, coordinated endorsement of a hub produces no farmable eigenvector-style escalation.
- Label flooding: Under label-identical flooding, author-normalised completeness falls from 0.816 to 0.377 on BRA and from 0.815 to 0.372 on UBI.The same coalition drives completeness to 0.228 and 0.219 under flat weights; all non-zero cells and policy comparisons are significant at p< .001.
- Label flooding: The defence margin from author normalisation is 0.117 at 5% coalition, widens to 0.181 at 15%, and is 0.151 at 25%.This supports treating the relation-weight function as a security control rather than merely a tuning parameter.
- Measurement and refresh: Flooding creates side imbalance: the within-voter gap rises from 0.085 without a coalition to 0.15–0.16 at the largest coalitions.The attack lowers coverage unevenly, so mean completeness can conceal damage concentrated on one side.
- Measurement and refresh: With a coalition, refreshing every 200 votes instead of every vote raises the ranked arm’s completeness by 0.013 while leaving uniform random selection unchanged.Without a coalition, the rule’s full-slate contrast remains null at both refresh intervals; under attack, lag changes the result in the unhelpful direction.
10 Peer-to-Peer Realisation on DDP2P
The paper maps its auditable slate rule onto DDP2P by combining signed, self-describing records with local computation. Peer evaluation remains verifiable but can operate on incomplete evidence, whose effects are bounded and visible.
- Platform mapping: DDP2P supplies an open-source decentralised platform whose self-contained, globally identified records fit the civic recommender charter.Its item types cover peers, constituents, motions, justifications, signatures, votes, witnessing statements, news, and translations.
- Platform mapping: The model maps closely onto DDP2P item types, but decentralised label assignment is needed to avoid concentrating semantic power.Table 16 identifies what already exists and what a deployment would have to add; centrally assigned labels would undermine the charter’s protection against semantic control.
- Identity: A witnessed decentralised census addresses identity eligibility, but the platform supplies verifiable evidence rather than a single eligibility verdict.Different observers may apply different eligibility criteria to the same signed census data.
- Partial replicas: On a partial replica, the same rule computes the full-pool rule restricted to held items, so missing items can only lower achievable completeness.Missing votes and links also lower computed item scores, making peer scores lower bounds on true scores.
- Partial replicas: Local evaluation prioritises votes and links because partial non-negative evidence converges faster than the justification corpus.The interface can expose incompleteness while synchronisation proceeds.
- Local computation: Algorithm 4 yields a slate and a checkable statement of the evidence supporting it.The local cycle merges signed items, discards unverifiable records, and reports held votes, links, justifications, and the synchronisation horizon.
11 Discussion
The discussion argues that legibility imposes little measurable quality cost, while policy choices and threat assumptions matter more than the ranking mechanism itself. It treats robustness and coverage-versus-mass trade-offs as participant-visible design choices.
- Discussion: The charter’s measurable price is 0.031 ± 0.014 against a ceiling that upper-bounds every procedure on the same pool.Four fifths of observed incompleteness came from vocabulary that did not yet exist, which no selector could recover.
- Discussion: An unconstrained learned ranker has at most three points of completeness advantage, versus 0.10 from raising authoring and 0.14 from doubling slate size.The comparison bounds the quality cost of legibility relative to deployment-controlled levers.
- Discussion: Across 4000 seed-paired runs, the rule was indistinguishable from random selection on coverage because the instrument is order-blind.The discussion presents this as a measurement consequence predicted by the definition, not as evidence that ranking never matters.
- Discussion: Author normalisation supplies a publishable, verifiable defence worth 0.151–0.181 completeness against a coalition holding a fifth of the electorate.The defence uses distinct-author counts rather than classifiers, trust scores, or material removal.
- Discussion: At G=0.5, link summands buy 2.8 points of coverage for 2.4 points of endorsement mass, defining a participant-selectable frontier.Choosing a point on this frontier is a substantive trade-off between breadth of represented reasons and fidelity to endorsement uptake.
- Discussion: Deployment priorities are to improve authoring and publish the weight policy as a security control.The discussion identifies authoring speed as the dominant lever and warns that a flat default leaves the main defence unarmed.
12 Limitations
The conclusions are bounded by a simulated, one-round electorate, untested identity resistance and partial-replica behavior, and an unexamined fairness question. Several statistical contrasts are also explicitly treated as suggestive rather than decisive.
- Scope and model: The electorate is simulated, so reported quantities describe modelled agents rather than human voters.The simulator does not model opinion updating, boredom, persuasion, or strategy beyond specified coalitions.
- Scope and model: Degenerate-authoring results depend on cosine-similarity adoption, where degenerate items receive 1.45 endorsements versus 10.94 for ordinary items.The rule inherits and amplifies this modelled adoption pattern; its behavior under different human endorsement patterns is not established.
- Scope and model: The study evaluates one voting round, leaving the rule’s ordering advantage under opinion revision untested.Multiple rounds could change adoption dynamics, endorsement distributions, and link structure.
- Statistical scope: Twenty-four ordering contrasts are uncorrected for multiplicity, and two link-term effects near p≈10^-2 are only suggestive.The conclusions rely instead on effects reported at p<10^-10, which the paper says would survive reasonable correction.
- Deployment boundaries: Sybil resistance is assumed rather than demonstrated; if the census is broken, the Section 9 results fail because identities can mint endorsements.The intended witnessed-census remedy was not evaluated, making this the largest deployment gap.
- Deployment boundaries: Partial replicas are analysed but not measured, so the prediction that ordering matters more under incomplete replication remains untested.The paper provides a monotone degradation proposition and an evidence-reporting algorithm, not a divergent-replica experiment.
- Deployment boundaries: Fairness across minority positions remains unexamined because endorsement mass may disadvantage positions represented by fewer items.The weighted variant anticipates the issue syntactically but does not answer it.
13 Work, Meaning, and the E-Citizen
The paper frames civic procedure as a way to make collective reasoning possible, while arguing that artificial intelligence could release citizen time without taking over citizens’ informational choices. It therefore places language models at authoring time, not at the unaccountable serving-time selection step.
- Civic procedure converts collective volume into a decision by constraining the form of assembly.
- Historical self-government was constrained by the citizen-hours required for assembly attendance, jury service, and office rotation.The paper argues that Athens addressed this constraint through slavery, leaving later franchise expansions without that solution.
- Artificial intelligence could release time for self-government by discharging productive labour, if human time is not valued only for producing goods and services.
- Delegating argument assembly and exposure decisions to the same technology would leave citizens with leisure but the informational position of spectators.
- Language models may help citizens express their meanings during authoring, but must not decide unaccountably which meanings voters encounter.The paper treats this as the boundary between assisting a citizen and replacing one.
14 Conclusions
The paper presents a published, parameterized selection rule for deliberative polls and evaluates it across coverage, order, and endorsement mass. Its results show small overall cost, strong performance on order and mass, conditional coverage gains, and important assumptions about modeled discrimination and census integrity.
- The proposed poll uses bipolar justification sets, three slate instruments, and a one-hop reversed endorsement flow with voter-held parameters.
- Across roughly 17,000 seeded runs, semantic abstinence costs 0.031 ± 0.014 against a ceiling bounding every mechanism.Four fifths of the residual shortfall is temporal rather than algorithmic.
- On non-degenerate authoring, set coverage cannot distinguish the rule from a random draw, while order-sensitive performance leads at every prefix and widens under attack.The order margin reaches −8.3 positions of 𝑒90 at a quarter-electorate coalition.
- On endorsement mass, the random baseline loses by a factor of 3.3, while realistic unjustified submissions restore and grow the rule’s coverage margin.At 𝐺 = 0.7, link summands supply 93% of the coverage margin.
- The conclusions depend on modeled human discrimination against unjustified submissions and on census integrity for adversarial results.
- The paper’s strongest claim is procedural: voters should be able to answer why they saw an argument with a recomputable table of numbers.The constrained system costs about three points of coverage and is easier to understand when wrong.
A note on scope and provenance
The manuscript is presented as an independently written extended treatment rather than a camera-ready conference version, with its definitions, analyses, and visual materials composed for this document.
- The manuscript reproduces no text, figure, table, algorithm, or example from a conference paper.
- Its notation, three-instrument framing, seven criteria and tests, worked example, and diagrams are original to this document.
Tools, data and reproducibility
The simulations are fully seeded and recorded, and the authors make the run databases and analysis scripts available for reproducing the reported tables and figures.
- Every simulation run records its seed, full parameter set, served slates in order, and policy vector in force.
- The run databases and analysis scripts producing every table and figure are available at the project’s GitHub repository.
- Language-model assistance was limited to drafting and editing prose; the authors conducted the technical work, experimental design, analysis, and conclusions.
A The Comparison This Manuscript Declines to Run
The manuscript declines to compare its rule with a learned ranker because such a comparison cannot test the paper’s admissibility claim about legibility. Instead, it reports an upper-bound comparison and identifies a human-participant experiment as the evidence that could change its position.
- A learned-ranker comparison cannot adjudicate the manuscript’s normative claim that civic mechanisms must be recomputable, attributable, contestable, and reconstructable.Performance differences would not resolve whether opacity makes a mechanism unsuitable for a binding civic process.
- 0.035 ± 0.013 is the gap between served slates and a label-reading greedy-cover ceiling that upper-bounds every selection procedure, including learned rankers.The ceiling saturates the live vocabulary in every seed, bounding the advantage available to an unconstrained mechanism.
- The manuscript treats procedural legibility like other procedural properties whose value is not reducible to measurement quality.Its examples include the secret ballot, double-entry bookkeeping, and rules of order.
- The comparison would be uninformative because a learned baseline’s behaviour depends on non-inspectable training data, objectives, and tuning choices.Both a result favouring the rule and one favouring the ranker would remain unconvincing under the manuscript’s framing.
- The proposed informative experiment would compare matched populations receiving identical items but differing in whether slate selection is recomputable and decomposable.It would measure contestation, contest termination, and reported trust, but requires human participants and is outside this simulation study’s scope.