Paper deep dive
Rules Before Oracles: Auditable, User-Configurable Argument Selection for Deliberative Polling
Muntaser Syed, Markus Zanker, Marius Silaghi
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 96%
Last extracted: 8/26/2026, 5:15:58 AM
Summary
The paper proposes a legible, user-configurable argument selection rule for deliberative polling to replace opaque learned rankers. It formalizes a poll over bipolar justification sets and introduces a 'one-hop reversed endorsement flow' rule parameterized by a relation-weight function. Evaluated via an agentic simulator with ~17,000 runs, the rule achieves near-ceiling coverage (0.035 short of upper bound) and significantly outperforms random selection in endorsement mass and order-sensitive metrics, particularly under adversarial conditions like label-homogeneous flooding. The authors argue that legibility is an admissibility condition for civic processes and map the solution to the DirectDemocracy Peer-to-Peer (DDP2P) platform.
Entities (11)
Relation Signals (7)
Muntaser Syed â affiliatedwith â Florida Institute of Technology
confidence 99% ¡ Muntaser Syed... Affiliation: Florida Institute of Technology
Marius Silaghi â affiliatedwith â Florida Institute of Technology
confidence 99% ¡ Marius Silaghi... Affiliation: Florida Institute of Technology
Markus Zanker â affiliatedwith â Free University of Bozen-Bolzano
confidence 99% ¡ Markus Zanker... Affiliation: Free University of Bozen-Bolzano
One-hop reversed endorsement flow â usesparameter â Relation-weight function
confidence 95% ¡ a rule meeting them: a one-hop reversed endorsement flow parameterised by a relation-weight function
One-hop reversed endorsement flow â mapsto â DDP2P
confidence 94% ¡ It maps onto an open-source peer-to-peer platform... Section 10 shows the model developed here maps onto DirectDemocracy Peer-to-Peer (DDP2P)
Relation-weight function â mitigates â Label-homogeneous flooding
confidence 92% ¡ Label-homogeneous flooding collapses completeness... under a flat weight policy, only to 0.44 under author-count normalisation, making the weight function a security control
One-hop reversed endorsement flow â outperforms â Random slate
confidence 90% ¡ on the other two it leads at every prefix... and dominates on mass by a factor of 3.3
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:In a deliberative poll, once submissions outnumber what anyone will read, some mechanism chooses which arguments each voter sees, acquiring much of the decision; practice delegates it to opaque learned rankers, so a voter cannot recompute or contest the exposure that shaped their vote. We ask whether it can be a published rule over publicly recomputable evidence with parameters held by the voter, treating legibility as an admissibility condition on usable mechanisms, not an objective traded against accuracy. We formalise a poll over bipolar justification sets, judging a slate by reason coverage, the order it arrives in, and captured endorsement mass; we give seven checkable criteria for a civic recommender and a rule meeting them: a one-hop reversed endorsement flow parameterised by a relation-weight function. An agentic simulator records every slate at every vote, over about 17,000 seed-paired runs. Served slates fall 0.035 short of a label-reading ceiling upper-bounding every selection procedure, opaque ones included: any unconstrained ranker's advantage is bounded and small. On coverage alone, with non-degenerate authoring, the rule is indistinguishable from a random slate, a null due to an order-blind, charity-blind instrument; on the other two it leads at every prefix by a margin widening with adversarial pressure and dominates on mass by a factor of 3.3. Once a realistic fraction of submissions carries no reasons, the coverage margin returns and grows. Label-homogeneous flooding collapses completeness from 0.81 to 0.34 under a flat weight policy, only to 0.44 under author-count normalisation, making the weight function a security control worth 10% of completeness. The choice between ranking arms is a position on a coverage-versus-mass frontier, not a fact, the kind of choice only a legible rule can hand to the person it affects. It maps onto an open-source peer-to-peer platform.
Tags
Links
- Source: https://arxiv.org/abs/2608.23979v1
- Canonical: https://arxiv.org/abs/2608.23979v1
Trouble viewing inline? Open PDF directly â
Full Text
200,836 characters extracted from source content.
Expand or collapse full text
Rules Before Oracles: Auditable, User-Configurable Argument Selection for Deliberative Polling Volume: 008 Muntaser Syed email: muntaser@my.fit.edu Affiliation: Florida Institute of Technology, Melbourne, Florida, USA , Markus Zanker email: markus.zanker@unibz.it Affiliation: Free University of Bozen-Bolzano, Bolsano, Italy and Marius Silaghi Note: Corresponding Author. email: msilaghi@fit.edu Affiliation: Florida Institute of Technology, Melbourne, Florida, USA 2026 Abstract. Background: A deliberative poll asks people to decide only after considering the arguments that bear on the decision. Once submitted arguments outnumber what anyone will read, some mechanism must choose which of them each voter sees, and that choice acquires a large share of the decision. Contemporary practice delegates it to opaque learned rankers, so a participant cannot recompute, attribute or contest the exposure that shaped their vote. Objectives: We ask whether argument selection in a binding civic process can be a published rule over publicly recomputable evidence, with free parameters held by the individual voter, and what such a constraint costs. We treat legibility/explainability as an admissibility condition on the class of usable mechanisms rather than as an objective to be traded against accuracy, and we quantify the trade-off that is thereby foreclosed. Methods: We formalise a poll as an alternative-based information system over bipolar justification sets, define three complementary evaluation instruments for judging a served slate: (i) set coverage of the live reason vocabulary, (i) the order in which coverage arrives, and (i) captured endorsement mass. We also state the Subsuming Justification Problem with a greedy submodular bound used strictly as an evaluation ceiling, and give seven checkable criteria for a civic recommender and a rule satisfying all seven: a one-hop reversed endorsement flow whose only policy parameter is a relation-weight function. An agentic simulator instantiates N constituent agents and persists every slate at the instant of every vote; we report roughly 17,00017,000 seeded, seed-paired runs over two propositions. Results: In terms of coverage of live reason vocabulary fraction, served slates fall only 0.035Âą0.0130.035Âą 0.013 short of a label-reading ceiling that upper-bounds every selection procedure, opaque ones included, so the entire competitive advantage available to an unconstrained ranker is bounded and small. On set coverage alone with solely non-degenerate authoring the rule is statistically indistinguishable from a uniformly random slate; we show that null is an artefact of an order-blind, charity-blind instrument. Under the two remaining evaluation instruments, the rule has a clear effect: it leads at every slate prefix by a margin that widens with adversarial pressure (â5.7-5.7 positions of depth-to-90% at a quarter-electorate coalition), and it dominates on endorsement mass by a factor of 3.33.3. Once a realistic fraction of submissions carries no reasons, the coverage margin returns and grows monotonically, with the two link summands â exactly inert on a uniformly good corpus â supplying 91%91\% of it. Label-homogeneous flooding collapses completeness from 0.810.81 to 0.340.34 under a flat weight policy but only to 0.440.44 under an author-count-normalised one; a coalition co-signing to defeat the normalisation makes itself monotonically weaker. Conclusions: Serving as a key security control, the weight policy carries significant completeness value (10%). The choice between ranking arms is a position on a coverage-versus-mass frontier rather than a fact, which is precisely the kind of choice only a legible rule can hand to the person it affects. We map the construction onto an existing open-source peer-to-peer platform, where per-peer local evaluation turns configurability from an operatorâs concession into a structural property. 1. Introduction In an assembly small enough that everyone hears everything, nobody has to decide what gets heard. Every real electorate is larger than that, so a selection step is unavoidable: something decides which of the thousands of submitted reasons appear on the screen of a voter about to cast a ballot. That step is where the power sits. The tally is public and checkable; the slate of arguments that produced the votes is usually neither. This manuscript takes the position that in a binding civic process the selection step is a piece of democratic procedure and must be built like one â a rule published in advance, computed over evidence any participant can retrieve and recompute, with the free parameters belonging to the individual voter rather than to whoever runs the servers. The key insight that makes this practical is that the property we actually want from a slate â that its reasons jointly span the reasons in play â is a set-level property of the served collection, and set-level properties can be pursued through the structure of the argumentation graph without ever interpreting the text; a rule that reads only endorsement counts and signed links is therefore both good enough to use and far harder to bend without leaving a public trace. Technically we model a poll as an alternative-based information system, measure a served slate along three axes rather than one, exhibit a one-hop reversed endorsement-flow rule whose sole policy lever is a relation-weight function, evaluate it inside a seed-controlled agentic simulator against a greedy submodular ceiling and under coordinated attack, and map the result onto an existing peer-to-peer deliberation platform. The contributions are: ⢠A charter for civic recommenders (Section 4): address seven operational criteria â 1. determinism, 2. evidence locality, 3. author blindness, 4. semantic abstinence, 5. reproducibility, 6. contestability, 7. configurability â each with the test that discharges it and the observable that would show it failing, argued as admissibility conditions rather than objectives; an opaque learned ranker fails four by construction. ⢠A three-instrument model of what a slate is for (Section 3): define an alternative-based poll over bipolar justification sets and measure 1. completeness of coverage against the reason vocabulary live at the instant of each vote; 2. a prefix-sensitive family of order measures; 3. captured endorsement mass. The Subsuming Justification Problem is stated with its weighted and componentwise variants, and a greedy (1âeâ1)(1-e^-1) cover to be used strictly as an evaluation ceiling. ⢠A rule that meets the charter (Sections 4â5): a one-hop reversed endorsement flow over rebuttal and reinforcement links, with a per-voter policy vector making the rule individually configurable without making it individually unpredictable; and Abas, a seed-controlled simulator of N constituent agents that persists every served slate before the ballot it informed, yielding some 17,00017,000 runs across sensitivity sweeps, a degenerate-authoring sweep, ranking-term ablations and five adversarial families. ⢠A measured price for the charter (Section 7.3): the coverage fraction distance between the served slates and a label-reading greedy ceiling that upper-bounds every selection procedure on the same pool is 0.031Âą0.0140.031Âą 0.014, against a temporal component of 0.1630.163 that no procedure of any kind could recover. ⢠Experimental results and what a ranking rule is actually for (Sections 8â8.5): on set coverage alone, with solely non-degenerate authoring, the rule does not separate from a uniformly random slate, its two link summands are all but inert, and under flooding the random slate even wins. We show it is an artefact of an order-blind instrument and an implausibly charitable authoring model: under an order-sensitive reading the rule leads at every prefix with the margin growing under attack, and once a realistic fraction of submissions fails to justify, the coverage margin returns, rises monotonically with that fraction, and makes the link summands worth as much as the whole ranking advantage â by squaring the discrimination the electorate already exercises. ⢠An empirical manipulation result with a named defence (Section 9): under authentication, hub-riding is harmless and label-homogeneous flooding is not; author-count normalisation retains 0.080.08â0.120.12 of completeness that a flat policy loses; a rate- and corpus-matched control isolates homogeneity as four fifths of the cause; and the obvious escalation â co-signing links to inflate the normalising numerator â makes the coalition monotonically weaker. ⢠A peer-to-peer realisation and a normative reading (Sections 10 and 13): a component-by-component mapping onto DirectDemocracy Peer-to-Peer (DDP2P), covering identity and census, gossip synchronisation and local evaluation over partial replicas; and an argument for why legible procedure is constitutive rather than decorative. The remainder of this section states the exposure problem and previews what the measurements turned out to say. Section 2 reviews the literatures the work sits between, Section 3 builds the formal object and its three instruments, Section 4 states the charter and the rule, Section 5 describes the simulator and Section 6 the protocol. Sections 7â9 report the non-degenerate authoring attack-free sweeps and the price of semantic abstinence, the null result and its two resolutions, and the adversarial sweeps. Section 10 maps the design onto DDP2P, Section 11 discusses, Section 12 lists what would change the conclusions, Section 13 makes the normative argument, and Section 14 concludes. 1.1. The exposure problem, and the position taken here The legitimacy of a collective decision has never rested on the tally alone. From Millâs marketplace of ideas to Habermasâs account of communicative rationality Mill, 1859; Habermas, 1984, the tradition holds that a decision earns authority from the quality of the discourse preceding it. Fishkin turned that into a measurable protocol Fishkin, 1991; Fishkin et al., 2000: draw a stratified sample, expose it to balanced material and structured discussion, measure how opinion moves. Deployments across many countries report the same qualitative finding â considered opinion differs systematically from raw opinion, and partisan cues lose their grip once people meet the reasons behind positions Luskin et al., 2002; Fishkin, 2009 â and the protocolâs weakness is cost, a facilitated cohort of a few hundred being an expensive instrument that does not obviously scale to a national electorate Lukensmeyer & Brigham, 2005. Digital platforms remove that barrier and immediately install another. When a hundred thousand people submit reasons, nobody reads the corpus; each person reads a slate. Call this the exposure problem: given a growing corpus of justifications and a per-voter display budget of K items, how should the slate be chosen so that the reasons in play are represented in what voters actually see? Getting it wrong produces three distinct harms, and they call for different remedies: ⢠Temporal inequity: in a sequential process the corpus is thin at the start and rich at the end, so early voters decide against a narrower reason space than late voters through no fault of their own; at the reference configuration the spread of per-voter completeness within a single run is about 0.210.21, five times the variation between one entire run and another. ⢠Silent narrowing: a selector tuned for engagement converges on whatever holds attention, which in political material is typically what confirms a prior Pariser, 2011; Sunstein, 2007; Willson, 2014, and the narrowing is silent because no individual slate looks censored â the missing reasons are simply never surfaced. ⢠Unfalsifiable influence: if the selector is a learned model, a participant who suspects their side is being systematically under-surfaced cannot establish it, cannot recompute the ranking, cannot point at the clause that produced the outcome, and cannot distinguish deliberate suppression from an artefact of training data Burrell, 2016; Lipton, 2018. A civic process unable to answer âwhy was I shown this?â with a checkable derivation has replaced procedural legitimacy with trust in an operator. The position of this manuscript follows from the last two harms in particular, and it is a position about admissibility, not about performance. We do not claim that opaque rankers rank badly. We claim that opaque ranking is the wrong frame for a civic slate: the object produced is not a personalised feed but a piece of the public record of what a voter was shown, and a public record must be reconstructible. Legibility therefore determines which procedures may be used at all, and questions of comparative quality arise only among those that qualify â the same structure as the secret ballot, which nobody defends on the grounds that it measures preferences more accurately than open voting. Section 4.2 states the criteria in a form that can be checked against an implementation, and every subsequent section is written so that a reader can see which criterion each design decision is discharging. Figure 1 states the commitment in one picture. corpus of justificationsbehavioural telemetrylearned selector(weights not public)slate of KKinadmissibleno derivation to point atcorpus of justificationspublic endorsements+ signed linkspublished rulescθiâ(â )sc_ _i(¡)slate of KKvoterâs own policy vector θi _iadmissible Figure 1. The two pipelines. Both consume the same corpus and both emit K items. The upper one additionally consumes behavioural telemetry and passes it through parameters nobody outside can inspect, so a disputed slate has no derivation to point at. The lower one consumes only evidence every participant can already retrieve â who endorsed what, and which signed links exist â and exposes its policy parameters to the individual voter, so a disputed slate reduces either to a disputed input, which is checkable, or to a disputed policy, which is arguable. This manuscript treats the difference as a condition of admissibility rather than as a term in an objective. 1.2. What the measurements turned out to say In a simplified setting where all justification authors correctly express their reasons (non-degenerate authoring), the fraction of the live reason vocabulary covered by a served slate, a simplified optimization objective for the motivating harms, does not differ significantly between our rule and random selection. With metrics that evaluate the ordering of justifications within the slate, and/or under a realistic electorate with degenerate authoring, the rule shows strong benefits over the control alternatives. 1.3. Why the substrate matters A published rule evaluated on a server the operator controls is a large improvement over an unpublished one, and it stops short. The operator still decides which endorsements are in the database, when the index refreshes, and whose items quietly stop being returned. Configurability granted by an operator is revocable by the same operator, and reproducibility asserted by an operator is a claim about a machine nobody else can inspect. The natural terminus of the argument is an architecture in which each participant holds their own replica of the items they care about and evaluates the rule locally over it. DDP2P â DirectDemocracyP2P â is an existing open-source Java platform built on precisely that premise Silaghi et al., 2013; Silaghi et al., 2013a: peers keep independent databases of self-contained items, name them by global identifiers derived from public keys and content digests, and converge by pushâpull gossip. Section 10 shows the model developed here maps onto DDP2Pâs item types with very little impedance mismatch, and that doing so converts several criteria of the charter from policy promises into structural facts. 2. Background and Related Work This work sits at the junction of five literatures: (1) the empirical political science of deliberative polling, (2) the formal theory of bipolar argumentation, (3) computational social choice for participatory processes, (4) the systems literature on recommendation and its manipulation, and (5) the peer-to-peer work that supplies the substrate. A sixth â the use of large language models as simulated populations â supplies the register in which the evaluation is conducted. We state below what we take from each and, where it matters, what we deliberately do not take. Deliberative polling and its scaling problem. Deliberative polling instruments a simple hypothesis: opinion formed after exposure to balanced argument differs systematically from opinion measured cold Fishkin, 1991; Fishkin et al., 2000. The canonical design draws a stratified sample, supplies balanced briefing material, runs moderated small-group discussion, and re-measures. Results across many deployments are consistent in direction â movement is largest where the initial position was least informed Luskin et al., 2002; Fishkin, 2009 â and the effect survives when partisan cues are stripped from the material Price et al., 2002. The design constraint is that facilitation does not scale: cohorts are a few hundred and the per-participant cost is that of a small conference Lukensmeyer & Brigham, 2005. Digital deployments relax that constraint and inherit a new one Tolbert et al., 2009; Toots, 2019: with a corpus no participant reads in full, the balance of the briefing material is no longer something an organiser curates but something a selection mechanism produces, per voter, in real time. We take the goal â reasoned exposure before the ballot â and treat selection, rather than facilitation, as the object to be engineered. Bipolar argumentation. Dungâs abstract frameworks model a debate as a directed attack graph and define acceptability through extensions Dung, 1995. Bipolar frameworks add a support relation alongside attack Cayrol & Lagasquie-Schiex, 2005; Cayrol & Lagasquie-Schiex, 2013, and the handbook literature surveys the resulting semantic landscape 1, 1; Amgoud et al., 2008; labelled bipolar frameworks give a unified semantics for the many readings of support and make the choice among them explicit Gonzalez, 2021, and the correspondence with logic programming has since been settled in detail Alcântara & Cordeiro, 2025. Value-based extensions attach audience-relative preferences to arguments and thereby explain rational disagreement between audiences that accept the same facts Bench-Capon & Dunne, 2007, which is close in spirit to the per-voter policy vector of Section 4.6. Our use of this apparatus is deliberately shallow, and the shallowness is a design commitment. We adopt the bipolar signature â two relations of opposite polarity over a set of items â and decline the semantics: we compute no extensions and pronounce no argument acceptable or defeated. The reason is Criterion 4 (semantic abstinence, Section 4.2) â any semantics deciding which arguments survive is a machine deciding which arguments a voter should stop considering, exactly the authority a civic slate must not delegate. Links here are evidence about relevance, not adjudication of truth. That commitment also distinguishes this work from dialectical-quality measures over exchanges between parties Rocha et al., 2026, which judge the argumentation; Prakken and Sartor on argument schemes in law Prakken, 2001 and the procedural tradition running back to Robertâs rules Robert, 1915 are closer to the role links play here, structuring who may speak to what rather than who is right. Computational social choice for participatory processes. A parallel line asks how collective decisions should be composed once participation is mediated by software. Liquid democracy has been given an algorithmic treatment with explicit accuracy guarantees and failure modes Kahng et al., 2021; representative committees can be sampled from peers so that the committeeâs decisions track the electorateâs Meir et al., 2021; proposing and voting have been unified as aggregation over a metric space Bulteau, 2021; and proportionality has been extended from one-shot elections to sequential decision making Chandak et al., 2026. The complexity of outcome determination in judgment aggregation is likewise mapped Endriss, 2020. This literature aggregates positions; the present work is upstream of it, concerning what a voter is shown before a position is formed, and the two are complementary: any of these aggregation rules can sit downstream of the slate mechanism studied here. Argument mining and its role here. Argument mining extracts claims, premises and relations from text Lippi & Torroni, 2016; Stab & Gurevych, 2017; retrieval systems rank passages by argumentative quality Wachsmuth et al., 2018; Carenini & Moore, 2006; and argumentative dialogue agents use such structures to conduct persuasive exchanges Chalaguine & Hunter, 2020. These techniques are how the reason labels of Section 3.2 would be produced in a real deployment, and we assume nothing better than what they currently deliver. Their placement in the architecture is the point. Label extraction is an authoring-time operation: it runs once per item, its output is attached to the item, it is visible to the author, and it can be contested and corrected before the item is ever served. Selection is a serving-time operation over that frozen output. Keeping a language model strictly on the authoring side of that line means a disputed slate never requires anyone to reason about model internals; it requires them to point at a label they think is wrong, which is a claim about a public artefact. Section 12 returns to what happens when the labels themselves are adversarial. Civic platforms with algorithmic assistance. Deployed systems already make these choices. Polis clusters participants by agreement pattern and surfaces statements that bridge clusters Small et al., 2021, an approach with a genuine claim to representativeness whose clustering step nevertheless requires the operatorâs pipeline to be trusted; hybrid participatory systems go further and estimate the values behind participantsâ textual motivations, disambiguating them interactively Liscio, 2025, which is the closest published treatment of the authoring-time inference we place off the serving path. Deliberation-support platforms have been surveyed for their argumentation structures Brenneis et al., 2021; Mancini, 2015, and recent work examines language models as facilitators Behrendt et al., 2024; Tessler et al., 2024, with Tessler et al. reporting that a model-generated statement can find more common ground than a human mediator. That result is real and we do not dispute it. Our objection is categorical rather than empirical: a mediator whose reasoning cannot be reconstructed is unsuitable for a binding process regardless of measured quality, for the same reason that a demonstrably accurate but unauditable vote count is unsuitable. The regulatory direction is consistent â the Digital Services Act obliges very large platforms to disclose recommender parameters and offer a non-profiling option European Parliament and Council, 2022, and the AI Act imposes transparency and human-oversight duties on systems used in democratic processes European Parliament and Council, 2024 â and sycophancy Sharma et al., 2023 and measurable political leaning in pretrained models Feng et al., 2023 are further reasons to keep such models off the serving path. Diversity, coverage and manipulation in recommendation. Ranking by predicted relevance alone yields redundant lists; maximal marginal relevance trades relevance against novelty Carbonell & Goldstein, 1998, and aggregate diversity has been studied as a system-level objective Adomavicius & Kwon, 2012. Multi-stakeholder recommendation observes that platform, provider and consumer objectives diverge Burke, 2017; recency-aware work notes the drift of freshness against quality Chakraborty et al., 2019; and conceptual modelling of explainable recommenders has mapped what an explanation of a recommendation can even be Caro-MartĂnez et al., 2021. Our completeness functional is a coverage objective in this family, with the civic-specific twist that the covered universe is the set of reasons live at the moment of the vote rather than a static catalogue. The manipulation literature is more directly load-bearing for Section 9. Shilling attacks against collaborative filtering inject profiles to promote a target item Lam & Riedl, 2004; Shardanand & Maes, 1995; Sybil attacks manufacture identities to acquire disproportionate influence Douceur, 2002; eclipse attacks isolate a peerâs view of the network Singh et al., 2006; and election manipulation on social networks has been analysed as seeding and edge modification, with the hardness of each variant established Castiglioni, 2021. Eigenvector-style ranking schemes Page et al., 1999 are structurally susceptible to link farms, which is precisely why our rule takes exactly one hop and no iteration to a fixed point: Section 9.1 confirms empirically that hub-riding is inert against it, and Section 9.2 shows where the residual exposure lives. Agentic simulation as an evaluation method. Generative agents with memory and planning reproduce plausible social behaviour in sandboxed environments Park et al., 2023; language models conditioned on demographic profiles reproduce survey response distributions Argyle et al., 2023; multi-agent orchestration frameworks have matured Wu et al., 2024; and the method has been applied to social-science questions directly Bail, 2024. Our simulator belongs to this family and inherits its central caveat: agent behaviour is a model of constituent behaviour, not evidence about it. We therefore restrict every claim to statements about mechanism under a stated behavioural model â comparisons between ranking arms, sensitivities to structural parameters, responses to attacks â and make no claim about the magnitude any quantity would take in a human electorate. Decentralised infrastructure for civic processes. Structured overlays give scalable key-based routing Maymounkov & Mazières, 2002; epidemic protocols give robust eventual dissemination Demers et al., 1988; conflict-free replicated data types give convergence without coordination Shapiro et al., 2011; hash trees give compact integrity proofs Merkle, 1988; blockchains give append-only public ledgers Nakamoto, 2008, with transparency-log constructions applied to update integrity Nikitin et al., 2017 and cryptographic scrutiny of agreement protocols Miller, 2020; and decentralised social platforms have explored gossip-based dissemination Boutet et al., 2013. DDP2P occupies a specific position: it is not a ledger and not a DHT, but a gossip-replicated store of self-contained, signed items designed for petition drives and organisational decision-making Silaghi et al., 2013; Silaghi et al., 2013a; Silaghi & Roussev, 2014. Its research programme covers the pieces a civic deployment actually needs â decentralised census construction and verification Qin et al., 2013a; Qin et al., 2014, detection of false identities Qin et al., 2013, peer reputation Qin et al., 2013b, supernodes for peers behind NAT Alhamed & Silaghi, 2014, protocol stacking and update propagation Alhamed et al., 2013; Alhamed et al., 2013a, meta-level recommendation over decentralised data Alhamed et al., 2016, trust and key management Silaghi et al., 2016, interface studies for non-expert users Alqahtani & Silaghi, 2017; Alqahtani & Silaghi, 2016; Kattamuri et al., 2005, its logical foundations Roussev & Silaghi, 2017, related vehicular deployments Dhannoon et al., 2013; Dhannoon, 2013, and the motivating argument for why participation is worth the engineering Silaghi et al., 2017. Section 10 builds on that stack rather than proposing a new one. 3. The Object: An Alternative-Based Poll and Three Ways to Judge a Slate Before anything can be recommended, the thing being recommended over has to be defined, and before any recommender can be judged, the standard of judgement has to be fixed. This section does both. The object is an alternative-based information system: a poll in which each side of the question carries its own set of justifications, and in which typed links between justifications carry the polarity of the relations participants assert between them. The standard is deliberately plural. We define three instruments â how much of the live reason vocabulary a slate covers, how early in the slate that coverage arrives, and how much of the electorateâs expressed endorsement the slate captures â because Section 8 will show that any one of them alone gives a misleading verdict. We then state the combinatorial problem that coverage induces, prove it hard, and construct the greedy bound we use as an evaluation ceiling, with an explicit statement of the discipline governing that boundâs use. 3.1. Polls, sides, and justifications A poll asks a question with a finite set of mutually exclusive alternatives. Throughout we take the binary case â support or oppose a motion â because it is the case civic petitions actually present and because it keeps the notation legible; nothing in the construction depends on it, and Section 14 notes the multi-alternative generalisation. Definition 3.1 (Alternative-based poll). An alternative-based poll is a tuple (1) Î =â¨+,â,ââ,ââ,Ďc,Ďr⊠\;=\; \,J^+,\ J^-,\ R ,\ R ,\ _c,\ _r\, where +J^+ and âJ^- are finite sets of justifications attached to the supporting and opposing alternatives respectively11 1 In the experiments reported here, +J^+ and âJ^- are disjoint, but that can be relaxed.; âââĂR ĂJ with =+âŞâJ=J^+ ^- is the rebuttal relation, read â(j,k)âââ(j,k) asserts that j tells against kâ; âââĂR ĂJ is the reinforcement relation, read âj tells in favour of kâ; Ďc:âââĽ0 _c:J _⼠0 assigns each justification a civic weight; and Ďr:âââŞâââââĽ0 _r:R _⼠0 assigns each link a weight. Three features carry the design. The two relations are typed but not signed at the level of the tuple: polarity lives in which relation a link belongs to, and the score function of Section 4.4 attaches a separate coefficient to each, so a participant may configure how much a rebuttal counts relative to a reinforcement â a substantive editorial choice â without reaching inside the graph. The civic weight Ďc _c is the endorsement count: how many constituents cast their ballot citing that item. It is a public tally, not a quality judgement, and no part of the construction asks whether an item deserves the endorsements it has. Finally Ďr _r is where policy lives: Section 4.5 shows that choosing it flat or normalised by the number of distinct authors behind the link moves completeness under coordinated flooding by 0.080.08â0.120.12, an order of magnitude larger than any other design decision here. 3.2. Reason vocabularies and the live denominator Coverage requires a universe to cover. We obtain it by labelling each justification with the reasons it invokes. Definition 3.2 (Reason labelling). Fix a finite reason vocabulary Î . A labelling is a map Îť:â2ÎÎť:Jâ 2 assigning each justification the non-empty set of reasons it invokes22 2 In the first set of reported experiments, Îť:â2Îââ Îť:Jâ 2 \ \ to have solely non-degenerate authorship.. Just for evaluation in reported experiments, labels are associated at authoring time (Section 2), stored on the item, and never recomputed at serving time; a deployment publishes its vocabulary and its extraction procedure. The denominator is the subtler half. Comparing a slate against the full vocabulary Î would penalise an early voterâs slate for failing to see reasons nobody had yet written, which is exactly the temporal inequity of Section 1.1 â a real phenomenon, but one to measure separately rather than fold into the score of the recommender. We therefore evaluate against what existed. Definition 3.3 (Live vocabulary). For a voter i casting a ballot at time tit_i, let tiJ_t_i be the set of justifications submitted strictly before tit_i, and let the live vocabulary on side Ďâ+,âĎâ\+,-\ be ÎtiĎ=âjâĎâŠtiÎťâĄ(j) ^Ď_t_i= _\,j ^Ď _t_iÎť(j). Definition 3.4 (Slate completeness). Let SiĎâĎâŠtiS_i^Ď ^Ď _t_i be the slate of at most K justifications served to voter i on side Ď. Its completeness is (2) ciĎ=|âjâSiĎÎťâĄ(j)||ÎtiĎ|â[0,1],c_i^Ď\;=\; |\ _jâ S_i^ĎÎť(j)\ | |\ ^Ď_t_i\ |\;â\;[0,1], with ciĎ=1c_i^Ď=1 by convention when ÎtiĎ=â ^Ď_t_i= . The combined completeness of a run over N voters, by averaging, is (3) cÂŻ=12âNââi=1N(ci++ciâ). c\;=\; 12N _i=1^N (c_i^++c_i^- ). Completeness is defined per voter and per side, and (3) aggregates a voterâs two values by averaging them. Averaging permits compensation: a slate that covers one side well and the other badly scores the same as one that covers both indifferently. A flooding coalition attacks one side, so this is precisely the configuration in which the reported number and the experienced one can come apart. We therefore also run evaluations with the average over a voterâs two sides replaced by the minimum, (4) cimin=minâĄ(ci+,ciâ),cÂŻmin=1Nââi=1Ncimin,c_i \;=\; (c_i^+,\,c_i^- ), c \;=\; 1N _i=1^Nc_i , under which a voter is served exactly as well as their worse-covered side and is granted no credit for coverage they did not lose. Two properties of (2) deserve to be stated here rather than discovered later, because between them they account for the entire narrative arc of Section 8. Remark 1 (Order blindness). ciĎc_i^Ď depends on SiĎS_i^Ď only as a set. Permuting the slate leaves it unchanged. Completeness therefore cannot, even in principle, distinguish a ranking rule from any procedure that returns the same K items in a different order â and when |ĎâŠti||J^Ď _t_i| is not much larger than K, it cannot distinguish it from a random draw either, because both return nearly the same set. Remark 2 (Charity blindness). ciĎc_i^Ď credits an item for the labels it carries, assumed to be a nonempty set of a size above some given positive threshold, whatever the quality of the argument attached to them. If every submission in the corpus is a well-formed argument on its own side, every item is worth roughly the same to this measure, and no selection rule can beat a random draw by much. The measure only rewards discrimination when there is something to discriminate against. 3.3. Three instruments: coverage, order, and endorsement mass A slate is a sequence, presented to a person with finite attention, drawn from a corpus the electorate has already expressed opinions about. Each clause supports a different question about whether the slate was well chosen, and we report all three because Section 8 shows they disagree in ways that matter. Instrument I: coverage. Completeness cÂŻ c per Definition 3.4. It answers: were the reasons in play represented at all? It is the natural formalisation of the balance requirement inherited from deliberative polling, and it is order-blind and charity-blind per Remarks 1â2. Instrument I: order. Real readers do not consume slates uniformly. Let SiĎ=â¨j1,âŚ,jKâŠS_i^Ď= j_1,âŚ,j_K now be an ordered slate, and define the prefix coverage at depth m, (5) piĎ(m)=|âuâ¤mÎťâĄ(ju)||ÎtiĎ|,m=1,âŚ,K,p^Ď_i(m)\;=\; |\ _u⤠mÎť(j_u)\ | |\ ^Ď_t_i\ |, m=1,âŚ,K, so that piĎâ(K)=ciĎp^Ď_i(K)=c^Ď_i recovers Instrument I. From the prefix profile we take three summaries. The area under the prefix curve AUCiĎ=Kâ1ââm=1KpiĎâ(m)AUC^Ď_i=K^-1 _m=1^Kp^Ď_i(m) is the coverage enjoyed by a reader who stops at a uniformly random depth. The rank-discounted coverage weights each reason by where it first appears, (6) RDCiĎ=(âtâ¤|ÎtiĎ|1log2âĄ(1+t))â1ââââÎtiĎâ[ââ served]log2âĄ(1+posiĎâ(â)),RDC^Ď_i\;=\; ( _tâ¤| ^Ď_t_i| 1 _2(1+t) )^-1 _ â ^Ď_t_i 1[ served] _2 (1+pos_i^Ď( ) ), where â[ââ served]1[ served] is an indicator function evaluating to 11 if the reason â appears on at least one item of the slate served to voter i on side Ď, and 00 otherwise, posiĎâ(â)pos_i^Ď( ) is the position of the first slate item carrying â and the normaliser is the value a slate would attain by covering the whole live vocabulary as early as positions allow, the normalizer being replaced with 1 at the beginning when that is needed to avoid division by 0. Finally e90e_90 is the smallest level m with piĎâ(m)âĽ0.9âciĎp^Ď_i(m)⼠0.9\,c^Ď_i â how deep the reader must go to obtain nine tenths of what that slate was ever going to give them. Instrument I answers: did the coverage arrive early enough to be read? Instrument I: endorsement mass. A slate can cover the vocabulary while serving nothing that any constituent actually adopted. Define (7) MiĎ=âjâSiĎĎcâ(j)maxâĄâjâTâĎâŠti,|T|â¤KâĄĎcâ(j),M^Ď_i\;=\; _jâ S^Ď_i _c(j) _\,T ^Ď _t_i,\ |T|⤠K\ _jâ T _c(j), the share of the maximum achievable endorsement weight that the served slate captured. Instrument I answers: did the slate show the voter what the electorate had actually taken up? It is the axis on which Section 8.5 finds the random baseline dominated by a factor of 3.33.3, and the axis on which the greedy oracle of Section 3.5 stops dominating under attack. Nothing in the design privileges one instrument: Section 8.5 treats them jointly, as a frontier over which a participant â not the system â chooses a position, and that this is even expressible is a consequence of Criterion 7 (configurability, Section 4.2). 3.4. The Subsuming Justification Problem Instrument I induces a combinatorial problem which we name and characterise, both because it locates the construction in complexity terms and because its hardness is what licenses the greedy bound of the next subsection. Say that SâĎS ^Ď subsumes a target vocabulary LâÎL when âjâSÎťâĄ(j)âL _jâ SÎť(j) L. Definition 3.5 (Subsuming Justification Problem, SJP). Instance: a labelled justification set (Ď,Îť)(J^Ď,Îť), a target LâÎĎL ^Ď, and a budget KââK . Question: does some SâĎS ^Ď with |S|â¤K|S|⤠K subsume L? Proposition 3.6. SJP is NP-complete. Proof. Membership: a certificate S with |S|â¤K|S|⤠K is verified by computing âjâSÎťâĄ(j) _jâ SÎť(j) and testing containment of L, in time linear in âjâS|ÎťâĄ(j)| _jâ S|Îť(j)|. Hardness: reduce from Set Cover Karp, 1972. Given a universe U, a family âą=F1,âŚ,FnF=\F_1,âŚ,F_n\ of subsets of U, and a budget K, put Î=U =U, create one justification jrj_r per FrF_r with ÎťâĄ(jr)=FrÎť(j_r)=F_r, set L=UL=U, and keep the budget. Any S of size â¤K⤠K subsuming U maps to a cover of size â¤K⤠K and conversely; the construction is linear in the instance size. â Two variants matter operationally. Weighted SJP adds Îź:ÎâââĽ0Îź: _⼠0 and a threshold Ď and asks whether some S with |S|â¤K|S|⤠K attains ÎźâĄ(âjâSÎťâĄ(j))âĽĎÎź( _jâ SÎť(j))âĽĎ; it inherits hardness at ÎźâĄ1Ο⥠1, Ď=|L|Ď=|L|, and its interest is normative rather than computational, since Îź is where a deployment would encode that some reasons must not be crowded out â minority-held reasons, statutory considerations â exactly the kind of parameter that must be published in advance rather than learned, a learned Îź being an unaccountable editorial line. Componentwise SJP reflects that a slate covering one side well and the other badly is not balanced: given budgets K+,KâK^+,K^- and targets L+,LâL^+,L^-, it asks for S+,SâS^+,S^- with |SĎ|â¤KĎ|S^Ď|⤠K^Ď subsuming LĎL^Ď for both Ď. Since the two sides share no items it decomposes into two independent SJP instances â which is why (3) averages per-side completeness rather than pooling the vocabularies: pooling would let a slate compensate a neglected side with an over-covered one, the opposite of balance. 3.5. The greedy ceiling, and the discipline governing its use Since coverage is the objective of an NP-hard problem, an exact optimum is not an available comparator. A standard bound is, and it happens to be tight enough to be informative. Given (ĎâŠti,Îť)(J^Ď _t_i,Îť), target ÎtiĎ ^Ď_t_i and budget K, the greedy cover selects, K times, the item maximising the number of not-yet-covered reasons it contributes, breaking ties by a published deterministic key. Proposition 3.7. The reason count fâĄ(S)=|âjâSÎťâĄ(j)|f(S)=| _jâ SÎť(j)| is monotone and submodular, so the greedy cover attains at least (1âeâ1)â0.632(1-e^-1)â 0.632 of the optimal value at budget K Nemhauser et al., 1978, and no polynomial-time algorithm improves the ratio unless P=NP P= NP Feige, 1998; Hochbaum, 1997. Proof. Monotonicity is immediate. For submodularity, take SâTS T and jâTjâ T; then fâĄ(SâŞj)âfâĄ(S)=|ÎťâĄ(j)ââkâSÎťâĄ(k)|âĽ|ÎťâĄ(j)ââkâTÎťâĄ(k)|=fâĄ(TâŞj)âfâĄ(T)f(SâŞ\j\)-f(S)=|Îť(j) _kâ SÎť(k)|âĽ|Îť(j) _kâ TÎť(k)|=f(TâŞ\j\)-f(T), since the subtracted set is larger. The bound and its optimality follow. â The greedy cover reads Îť directly â a semantic operation requiring the selector to know what each item is about and to choose items for what they say. It therefore violates Criterion 4 (semantic abstinence) outright, and because it must inspect every candidateâs labels it also fails Criterion 2 (evidence locality); both are stated with their tests in Section 4.2. It is not a candidate mechanism, yet we compare against it constantly, which requires stating the discipline explicitly. Principle 1 (Oracle discipline). Like the labels, the greedy cover appears in this manuscript only on the evaluation path. It is computed offline, from the persisted database, after the poll has closed. No agent ever sees a slate it produced; no served slate was ever influenced by it; it is not proposed, and could not be adopted, as a deployable selection rule. Its role is to answer one question that no admissible mechanism can answer about itself: how much coverage was available in this pool at all? Principle 1 is what makes the number in Section 7.3 meaningful. Because the greedy cover reads the labels, its coverage upper-bounds â to within the 0.6320.632 factor, and in practice far more tightly â what any selection procedure applied to the same pool could have achieved, including an arbitrarily good opaque learned ranker. The gap between the served slates and that ceiling is therefore not the price of our particular rule against some better rule; it is the price of the entire admissible class against the unreachable best case, and Section 7.3 measures it at 0.031Âą0.0140.031Âą 0.014. 3.6. Reach: what a slate can be held responsible for A rule that reads only endorsements and one-hop links cannot serve what it cannot see. Making that horizon explicit turns a limitation into a specification. The reach of a justification j at time t is Ďt(j)=jâŞk:(j,k)ââââŞââ,kât _t(j)=\j\âŞ\\,k:(j,k) ,\ k _t\,\, and j is visible to the rule at t if Ďcâ(j)>0 _c(j)>0 or Ďcâ(k)>0 _c(k)>0 for some kâĎtâ(j)kâ _t(j). Proposition 3.8 (Partial-replica sufficiency). Evaluating the endorsement rule of Section 4.4 for the items in a candidate pool P requires only the endorsement counts of âjâPĎtâ(j) _jâ P _t(j) and the link tuples out of P. It requires no global view of J, ââR or ââR . This is immediate from the score (8) given in Section 4.4: the score of j is a function of Ďcâ(j) _c(j) and of Ďcâ(k) _c(k) for k ranging over jâs out-neighbours only. Proposition 3.8 is the technical fact that makes Section 10 possible. A peer holding a partial replica computes exactly the same scores for the items it holds as a peer holding everything, which is what allows every participant to run the rule themselves rather than trust an operator to run it for them. It also delimits responsibility: a brand-new item with no endorsements and no incoming links from endorsed material is invisible to the rule, and no amount of tuning changes that. Cold-start items surface when someone reads them â which is what the authoring-side random exploration of Section 5.2 is for â and not before. 3.7. A worked example The following instance is small enough to check by hand and exhibits every phenomenon the experiments will measure at scale. Example 3.9 (Municipal night-bus extension). A city puts to its residents: extend the night-bus network to the outer districts? At the moment resident r opens the ballot, the supporting side holds five justifications and the opposing side four. Labels are drawn from a published vocabulary; endorsement counts Ďc _c are the public tallies. id gist Ďc _c labels Îť supporting side +J^+ s1s_1 shift workers stranded after 23:00 46 access, equity s2s_2 night transit cuts drink-driving crashes 31 safety s3s_3 same rolling stock, marginal cost only 12 cost s4s_4 shift workers are mostly low-income 04 equity s5s_5 fewer night cars lowers emissions 03 climate opposing side âJ^- o1o_1 subsidy competes with school repairs 39 cost, priorities o2o_2 measured outer-district demand is very low 28 demand o3o_3 depot noise at 02:00 in residential streets 09 amenity o4o_4 rosters would breach the rest agreement 02 labour Both live vocabularies have five reasons, so |Î+|=|Îâ|=5| ^+|=| ^-|=5. The links asserted so far are (s3,o1)âââ(s_3,o_1) (marginal cost tells against the subsidy objection), (s4,s1)âââ(s_4,s_1) , (s2,o3)âââ(s_2,o_3) , (o2,s1)âââ(o_2,s_1) and (o4,o1)âââ(o_4,o_1) . The display budget is K=2K=2 per side. What each procedure serves. Take the rule of Section 4.4 at the reference policy Îą=β=0.5Îą=β=0.5, ĎrâĄ1 _r⥠1. On the supporting side, scâĄ(s1)=46sc(s_1)=46, scâĄ(s2)=31+0.5â 9=35.5sc(s_2)=31+0.5¡ 9=35.5, scâĄ(s3)=12+0.5â 39=31.5sc(s_3)=12+0.5¡ 39=31.5, scâĄ(s4)=4+0.5â 46=27sc(s_4)=4+0.5¡ 46=27, scâĄ(s5)=3sc(s_5)=3. The rule serves â¨s1,s2⊠s_1,s_2 , covering access,equity,safety\ access, equity, safety\: completeness 3/5=0.603/5=0.60, prefix profile pâĄ(1)=0.40p(1)=0.40, pâĄ(2)=0.60p(2)=0.60. The greedy cover reads labels and serves s1,s3\s_1,s_3\ or s1,s5\s_1,s_5\ â also 0.600.60, since no pair here exceeds three reasons. A uniform draw serves 2.42.4 distinct reasons in expectation, i.e. completeness 0.480.48. Where the instruments diverge. Let the draw return s1,s5\s_1,s_5\, which it does as often as any other pair: its coverage is 3/53/5, equal to the ruleâs, so on Instrument I the two procedures are indistinguishable on this run. On Instrument I they are not: the rule captures M+=77/77=1.00M^+=77/77=1.00 against 49/77=0.6449/77=0.64 for s1,s5\s_1,s_5\ and 15/77=0.1915/77=0.19 for s3,s5\s_3,s_5\. The voter shown s3,s5\s_3,s_5\ met two reasons four of their neighbours had ever endorsed; the voter shown â¨s1,s2⊠s_1,s_2 met the two that seventy-seven had. Where the link terms earn their place. Item s3s_3 ranks fourth on its own endorsements and is promoted to third by the rebuttal coefficient because it speaks to the most endorsed item on the other side â but ranking on endorsements alone would place it third as well. That is the inertia of Section 8.1 in miniature: on a uniformly competent corpus the link terms reorder items that were already adjacent. Now replace s5s_5 by a submission s5â˛s_5 carrying the label climate and no actual reason: a sentence of agreement, correctly labelled, endorsed by nobody, asserting no links. Coverage is blind to the substitution, so a uniform draw serves s5â˛s_5 exactly as often as it served s5s_5; the rule scores it 00, both because nobody endorsed it and because it points at nothing anyone endorsed. That double penalty â the squaring of discrimination measured in Section 8.4 â is invisible here only because the instance contains one such item. At the fractions Section 8.3 reports it is worth as much as the whole ranking advantage. 4. The Charter: Criteria Before Objectives This section states what a selection mechanism must satisfy to be used in a binding civic process, and exhibits a rule that satisfies it. The order of presentation within this section is itself part of the argument: the criteria come before the rule, as conditions on the admissible class, and the rule follows as one member of that class chosen for its simplicity. We begin with what opacity costs in concrete operational terms, state the seven criteria with their tests, argue that they are prior to rather than commensurable with measured quality, give the rule, isolate the one line of it that is a security control, and close on how per-voter configuration is possible without dissolving the shared record. 4.1. What opacity costs, concretely The case against an opaque selector here is not that such systems are inaccurate; it is that some operations a civic process requires become impossible, and it is worth naming them as operations rather than as values. Recomputation: a participant who disputes a slate should be able to obtain the same slate from the same inputs, which with published weights over public evidence is arithmetic, whereas with a learned selector the participant would need the model, its parameters, the exact feature vector at serving time and the personalisation state â and even a cooperative operator disclosing all four supplies a recomputation nobody outside can independently verify was the one actually run. Attribution: when a slate is wrong, a process needs to say which input made it wrong, and under the rule of Section 4.4 the answer is a term â this item ranked here because it carried these endorsements and pointed at those items â whereas under a learned ranker the answer is that the output is a function of the training distribution, post-hoc attribution methods producing explanations that are themselves unverifiable Rudin, 2019; Lipton, 2018; surveys of explainability for supervised learning make the gap between an explanation and a derivation plain Burkart & Huber, 2021. Explanation: a voter is owed an account of why their slate contains what it contains, and that account must be the computation rather than a story fitted to it. Under the rule it is the score read aloud and cannot drift from the mechanism because it is the mechanism, whereas a learned selector explains itself through a second model whose account is plausible to the recipient rather than identical to the computation that ran Caro-MartĂnez et al., 2021. An explanation a voter cannot check against the public record is indistinguishable from persuasion. Contest: a dispute must terminate, and under the rule it terminates in one of three ways â an endorsement tally is wrong and tallies are public, a link is wrong and links are signed by their authors, or the policy is wrong and that is a political argument conducted in the open â whereas under an opaque selector the dispute has no terminating move, which converts every disagreement about content into a standing grievance about the operator. emphBounded drift: a published rule changes when someone changes it and the change is a diff, whereas a learned selector changes whenever it is retrained, whenever the population shifts and whenever a feature pipeline is updated, so its behaviour last month is not recoverable â and a civic record that cannot be reconstructed a year later is not a record. None of these costs is offset by better ranking, because none is a ranking property. This is the structural reason the criteria below are stated as admissibility conditions rather than as objectives to be weighed, and it is also where this construction parts company with the fairness-and-accountability literature that treats such properties as quantities to optimise: proposals for socially responsible AI Cheng et al., 2021 and critiques of fairness as a single formal target Weinberg, 2022 both proceed by asking what a system should maximise, whereas the question here is which systems may be used at all. 4.2. Seven criteria, with their tests Each criterion below is stated so that an auditor can check it against an implementation, together with the observable that would show it violated. They are numbered for reference throughout the rest of this manuscript; every subsequent design decision is annotated with the criterion it discharges. Criterion 1 (Determinism). Given the same evidence and the same policy vector, the mechanism returns the same slate. Test: run it twice on a frozen snapshot; the outputs are identical, ties included. Violation looks like: two participants with identical configurations and identical local data seeing different slates, with no published reason. Criterion 2 (Evidence locality). The score of an item depends only on that item and on a bounded, explicitly specified neighbourhood of it. Test: perturb an item outside the declared neighbourhood; the score does not move. Violation looks like: a global fixed point, or any quantity whose value depends on the whole corpus, which makes both partial-replica evaluation and local reasoning about a dispute impossible. Criterion 3 (Author blindness). No term in the score refers to who wrote an item, beyond counting distinct authors where a policy explicitly calls for it. Test: permute author identities across items; the ranking is invariant. Violation looks like: reputation, seniority, or verified-account status entering the ranking â reintroducing the influence hierarchy that a poll is supposed to flatten. Criterion 4 (Semantic abstinence). The serving-time mechanism does not read the content of items. Test: replace every itemâs text with an opaque identifier; the slate is unchanged. Violation looks like: the selector deciding, at serving time, what an argument means or whether it is any good â which is the specific authority a civic slate must not delegate, and which the greedy cover of Section 3.5 openly exercises, hence Principle 1. Criterion 5 (Reproducibility). Every served slate is recomputable after the fact from persisted evidence. Test: from the archive, reconstruct any historical slate exactly, including the state of the corpus at that instant. Violation looks like: an audit that can establish what was tallied but not what was shown. Criterion 6 (Contestability). Every slate decomposes into named contributions, each traceable to a public artefact. Test: for any served item, produce the list of (term,evidence,value)(term,evidence,value) triples summing to its score. Violation looks like: an explanation that is itself a model output. Criterion 7 (Configurability). The policy parameters are held by the individual participant receiving it, not by the operator, and the participant can change them and observe the effect. Test: two participants with different policies, same evidence, obtain different and separately correct slates. Violation looks like: a single operator-chosen ranking presented as neutral, or a personalisation the participant cannot inspect or switch off. Table 1 summarises how three mechanism families fare. The pattern is not that learned rankers score lower; it is that they are disqualified on four criteria at once, by construction rather than by implementation quality. Table 1. The seven criteria against three mechanism families. â denotes failure by construction rather than by implementation choice: no amount of engineering effort within that family recovers the property. The greedy cover is included because it appears throughout the evaluation and it is important to be explicit that it is inadmissible; see Principle 1. endorsement rule greedy cover learned ranker (§4.4) (§3.5) C1 determinism yes yes only if frozen C2 evidence locality one hop whole pool â unbounded â C3 author blindness yes tie-breaks learned from behaviour â C4 semantic abstinence yes reads labels â reads everything â C5 reproducibility yes yes needs exact checkpoint â C6 contestability per-term per-item post-hoc only â C7 configurability per voter fixed operator-held â 4.3. Why this is not an empirical hypothesis A natural objection runs: the assertion that rule-based selection is preferable, is never tested against a learned alternative. The objection is well formed and the answer governs how the rest of the results should be read. The claim is normative and it is a claim about admissibility: a mechanism that cannot be recomputed, attributed, explained, contested or reconstructed is unsuitable for a binding civic process â not that it ranks worse. A comparison against a learned ranker would report a difference in coverage or in some downstream engagement statistic, and whatever number came out could not bear on the claim, since a learned ranker covering more of the live vocabulary would still fail C2, C4, C5 and C6 and would still leave a disputing participant with no terminating move. The structure is familiar from other civic procedures. The secret ballot is not defended on the grounds that it measures preferences more accurately than open voting â it plainly measures some things less well, since it destroys the ability to audit an individualâs vote â but because the procedure must possess a property that outranks measurement quality. Double-entry bookkeeping is not the most compact representation of a firmâs accounts; rules of order do not produce the fastest decisions. In each case a procedural property is treated as prior, and efficiency questions are settled within the class of procedures that possess it. What is legitimately empirical, and what this manuscript tests, is everything downstream of that commitment: how much does the commitment cost?, answered against a ceiling that upper-bounds every mechanism including inadmissible ones (Section 7.3); which terms of the rule earn their place?, answered by ablation, including the finding that two of them earn nothing on a charitable corpus and a great deal on a realistic one (Sections 8.1 and 8.3); which policy choices are security controls?, answered by adversarial sweep (Section 9.2); and where does the choice between configurations actually lie?, answered by a frontier rather than an optimum (Section 8.5). 4.4. The endorsement rule We now give a mechanism satisfying all seven criteria. Its simplicity is a feature. Definition 4.1 (Endorsement rule). Let Eâ(j)=Ďcâ(j)E(j)= _c(j) be the endorsement count of item j. Given a policy vector θ=(Îą,β,Ďr)θ=(Îą,β, _r) with Îą,βââÎą,β , the endorsement score of j at time t is (8) scθâ(j)=EâĄ(j)+Îąââ(j,k)âââ,kâtĎrâ(j,k)âEâ(k)+βââ(j,k)âââ,kâtĎrâ(j,k)âEâ(k).sc_θ(j)\;=\;E(j)\;+\;Îą\!\! _(j,k) ,\ k _t\!\! _r(j,k)\,E(k)\;+\;β\!\! _(j,k) ,\ k _t\!\! _r(j,k)\,E(k). The slate served to voter i on side Ď is the top-K of ĎâŠtiJ^Ď _t_i by scθisc_ _i, ties broken by a published deterministic key (ascending item identifier). The direction of both summands is the load-bearing detail and is easy to get backwards. An item is credited for the endorsements of what it points at, not for links pointing at it. An item that rebuts a heavily endorsed objection is thereby promoted, because a voter about to accept that objection has a specific interest in seeing the response to it; and an item reinforcing a heavily endorsed claim is promoted because it deepens material the electorate has already taken up. Reversing the direction would reward being talked about, which is the property a coordinated group can manufacture most cheaply â and it is why the eigenvector family, which propagates credit backwards along links to a fixed point Page et al., 1999, is structurally exposed to link farms in a way Equation (8) is not. Section 9.1 confirms this: a coalition that constructs a hub and points the whole corpus at it gains nothing measurable. The rule discharges the charter as follows. It is a closed-form arithmetic expression over integers and published constants, hence C1. It reads j and jâs out-neighbours and nothing else, hence C2, which by Proposition 3.8 is also what makes per-peer evaluation more possible. No term names an author, and where a policy counts authors it counts distinct ones without regard to which, hence C3. No term reads text, hence C4 â note that Îť appears nowhere in (8); labels are used to evaluate slates and never to choose them, exactly the asymmetry Principle 1 protects. Persisting endorsements and links with timestamps makes every historical slate recomputable, hence C5. And each of the three summands is separately displayable against the artefacts that produced it, hence C6. C7 is the subject of Section 4.6. Algorithm 1 states the serving procedure for a peer holding a partial replica. Algorithm 1 ScoreAndServe â serve a slate for one voter on one side Given: candidate pool PâĎâŠtP ^Ď _t; endorsement counts E on PâŞâjâPĎtâ(j)P⪠_jâ P _t(j); out-links of P; voter policy θ=(Îą,β,Ďr)θ=(Îą,β, _r); budget K Yields: ordered slate â¨j1,âŚ,jminâĄ(K,|P|)⊠j_1,âŚ,j_ (K,|P|) and, for each, its score decomposition 1 foreach jâPjâ P do 2 aââ(j,k)âââ,kâtĎrâ(j,k)âEâ(k)aâ _(j,k) ,\,k _t _r(j,k)\,E(k); 3 rââ(j,k)âââ,kâtĎrâ(j,k)âEâ(k)râ _(j,k) ,\,k _t _r(j,k)\,E(k); 4 scâĄ[j]âEâĄ(j)+Îąâa+βârsc[j]â E(j)+Îą a+β r; 5 whyâĄ[j]ââ¨(own,EâĄ(j)),(reinforces,Îąâa),(rebuts,βâr)âŠwhy[j]â ( own,E(j)),\ ( reinforces,Îą a),\ ( rebuts,β r) ; 6 end foreach 7 sort P by scsc descending, ties by ascending identifier; 8 return the first minâĄ(K,|P|) (K,|P|) items of P with their whywhy records; Note that Algorithm 1 returns the decomposition alongside the slate rather than offering it on request. Contestability that must be asked for is contestability most participants never exercise; the interface of Section 5.6 shows the terms next to each served item. 4.5. The weight policy is the whole game Equation (8) has three policy parameters and they are not of equal consequence. Varying Îą and β across their plausible range moves completeness by amounts we could not distinguish from noise in the attack-free regime (Section 8.1); changing Ďr _r as seen next moves it by 0.080.08â0.120.12 under coordinated attack (Section 9.2). We compare two policies: the flat policy Ďrflatâ(j,k)=1 _r^flat(j,k)=1, and the author-normalised policy (9) Ďrnormâ(j,k)=|AâĄ(j,k)|maxâĄ(1,VĎâĄ(k)), _r^norm(j,k)\;=\; |A(j,k)| (1,V^Ď(k) ), where AâĄ(j,k)A(j,k) is the set of distinct participants who have asserted the link (j,k)(j,k) and VĎV^Ď is the number of participants who have cast a ballot on side Ď so far. Stating the normalisation in terms of distinct authors is what keeps it compatible with C3: it counts how many separate people stand behind a claimed relation and is indifferent to which people they are. Its effect on a coordinated coalition is direct. A coalition of size m can multiply the number of links it asserts freely, but under (9) each link is worth only the distinct authorship behind it, so the coalitionâs total link credit scales with m rather than with its output; Section 9.2 measures the resulting difference at 0.080.08â0.120.12 of completeness. The obvious escalation is to co-sign â have all m members assert the same links so that |AâĄ(j,k)|=m|A(j,k)|=m and the numerator recovers. Section 9.6 reports what happens, and the result is pleasantly counter-intuitive: completeness rises monotonically with the degree of co-signing, because co-signing concentrates the coalitionâs identities on a small set of links instead of spreading them across many, so the same authorship budget buys fewer distinct promoted items and the coalitionâs own material crowds itself out. Table 2 gives the two policies side by side. Table 2. The two relation-weight policies. The last column previews Section 9.2: mean combined completeness under a flooding coalition holding a quarter of the electorate, at the reference configuration. policy Ďrâ(j,k) _r(j,k) reads under flood flat 11 nothing beyond the linkâs existence 0.340.34 author-normalised |AâĄ(j,k)|/maxâĄ(1,VĎâĄ(k))|A(j,k)|\,/\, (1,V^Ď(k)) how many distinct people asserted it, against how many voted on that side 0.440.44 4.6. Configurability without incoherence C7 asks that the policy belong to the participant. The immediate worry is that this dissolves the shared object: if everyone sees a different slate, what is the common record the deliberation is about? The resolution is a matter of what varies and what does not. What varies is the policy vector θi=(Îąi,βi,Ďr(i),Ki) _i=( _i, _i, _r^(i),K_i), plus optionally a reason emphasis Îźi _i in the sense of weighted SJP. A participant who wants the strongest objections to positions they already hold sets β high; one who wants depth on established claims sets Îą high; one who distrusts coordinated authorship adopts (9); one who wants a longer slate raises K. These are editorial stances, they are legitimately personal, and none is the kind of thing an operator should be choosing on anyoneâs behalf33 3 For the special case of so-called grassroot organizations, users can also influence voter eligibility coefficients Qin et al., 2013a.. What does not vary is everything else: the corpus, the endorsement tallies, the link set, the labels, and the rule itself. Two participants with different θ compute different slates from identical public evidence, and each can compute the otherâs exactly. This is what distinguishes configuration from personalisation: a personalised feed differs because the system holds a private model of the user, whereas a configured slate differs because the user set a published parameter, and the difference is fully explained by that parameter. Under configuration a disagreement about slates reduces to a disagreement about policy, arguable in public; under personalisation it reduces to nothing at all. Proposition 4.2 (Charter closure under configuration). If the rule satisfies C1âC6 for every fixed θ, then the family scθâÎ\sc_θ\_θâ satisfies them for participant-chosen θ, provided θi _i is recorded alongside the slate. Determinism, locality, author blindness and semantic abstinence hold pointwise in θ and are inherited; reproducibility and contestability require the auditor to know which θ was in force, and persisting (θi,ti)( _i,t_i) with each slate â as Section 5.5 does â supplies it. Proposition 4.2 is why Section 5.5 stores the policy vector with every served slate, and why Section 8.5 can present its main finding as a frontier rather than an optimum: once the choice between ranking arms is a position on a trade-off between coverage and endorsement mass, there is no system-level answer to which position is right, and C7 is what lets the question be handed to the person it affects. 4.7. Rules of order as algorithms of interaction The idea that procedure constitutes rather than merely regulates deliberation is older than any of this. Robertâs Rules of Order Robert, 1915 are a specification for a distributed system with adversarial participants: recognition of the floor is admission control; the motion stack is a well-founded ordering guaranteeing termination; the requirement to speak to the motion is a relevance predicate; limits on repeat speech are rate limiting; and the prohibition on impugning motives is a type restriction on admissible utterances. None of it was derived from a theory of good outcomes, but from the observation that without such constraints an assembly is captured by whoever is loudest and most persistent. Our construction translates the same instinct into a setting where the floor is a screen: the published rule is the standing order, author blindness is the convention that the chair recognises members rather than reputations, and the persistence of Section 5.5 is the minutes. DDP2Pâs own rule that a motion admits one active argumentation per voter at a time Silaghi et al., 2013 is an anti-filibuster device with an exact parliamentary ancestor, and formal work on procedural argumentation makes the correspondence explicit Prakken, 2001; Silaghi & Roussev, 2014. The relevant inheritance is not any particular clause but the stance: an assemblyâs rules must be knowable in advance by everyone bound by them, which is a constraint on the form of the rules, not a preference about their content. 5. Abas: An Agentic Instrument for Auditable Deliberation Everything in Sections 3 and 4 is a specification. To find out what it does one needs an electorate, and human electorates are not available for parameter sweeps. Abas â Agent-Based Argument Simulator â instantiates N constituent agents against a real implementation of the poll object, runs the round protocol to completion, and persists enough state that every served slate can be reconstructed afterwards. This section describes the agent, the round, the model of submissions that fail to justify, the arms compared, what is stored, and the interface through which a participant would inspect any of it. The design commitment throughout is that the simulator is an instrument, not a product demonstration: it is built so that its own outputs can be audited, and Section 5.5 is where that commitment is cashed. 5.1. The constituent agent Each agent holds a scalar opinion θâ[â1,1]θâ[-1,1] on the motion, drawn once at initialisation, and a short position text generated from that opinion by a topic-specific generator. The opinion determines the ballot direction stochastically, so agents near the centre are genuinely uncertain; the position text is what the agent uses to judge which existing material speaks for it, through TFâIDF cosine similarity Manning et al., 2008; Spärck, 1972. Two properties of the agent matter for how the results should be read. First, agents do not update their opinion in response to what they are shown: the simulator measures exposure, not persuasion, and a model of persuasion would import assumptions we have no basis for. Second, agent behaviour is a model of constituent behaviour, not evidence about it, in the sense of Section 2; every claim we make is a claim about mechanism under a stated model. Section 12 states the specific form this caveat takes for each result, and in one case â the degenerate-authoring result â it is load-bearing enough that we state it twice. 5.2. The round protocol Agents are processed in identifier order, each one completing the sequence of Algorithm 2 before the next begins. Because the corpus therefore grows underneath the sequence, the live vocabulary of Definition 3.3 genuinely differs between the first and last agent, which is what makes the temporal inequity of Section 1.1 measurable rather than assumed. Algorithm 2 One constituentâs turn Given: agent a with opinion θa _a and position text xax_a; poll state at time t; policy θ; budget K; action probabilities Ďreuse,Ďabst,Ďown _reuse, _abst, _own with Ďreuse+Ďabst+Ďown=1 _reuse+ _abst+ _own=1; link probability Ďlnk _lnk 1 S+âS^+â ScoreAndServe(+âŠtJ^+ _t, θ, K); SââS^-â ScoreAndServe(ââŠtJ^- _t, θ, K); 2 persist â¨a,t,θ,S+,Sâ⊠a,t,θ,S^+,S^- ; // before any ballot is cast 3 vâvâ ballot direction drawn from θa _a; 4 uâ[0,1)u [0,1); 5 if u<Ďreuseu< _reuse then 6 jâjâ PickFromSlate(SvS^v, xax_a) ; // similarity-weighted draw from the served slate 7 record ballot for v citing j; increment EâĄ(j)E(j); 8 else if u<Ďreuse+Ďabstu< _reuse+ _abst then 9 record ballot for v with no citation; 10 else 11 if a authors degenerately (Section 5.3) then 12 jâjâ WriteWithoutReason(xax_a, v); 13 else 14 jâjâ WriteNew(xax_a, v) ; // fresh item, labels from the side vocabulary 15 end if 16 insert j; record ballot for v citing j; 17 if [0,1)<ĎlnkU[0,1)< _lnk then 18 AssertLinks(j) ; // one rebuttal at an opposing item, one reinforcement at a same-side item 19 end if 20 end if Adoption is restricted to the served slate â an agent can only cite what it was shown â which couples the recommender to the endorsement tallies and makes the system a feedback loop rather than a passive display. And the authoring branch draws labels from the sideâs vocabulary independently of the slate, so corpus growth is driven by the authoring draw and not by what was served; this is why the ablations of Section 8 are clean, all ranking arms seeing a bit-identical corpus and differing only in which K items of it were displayed. Link targets are chosen by the agent, not the system: a rebuttal aims at the opposing item least similar to the authorâs position, a reinforcement at the same-side item most similar to it. 5.3. Submissions that fail to justify In the first set of experiments we report here, the authoring model had every agent who wrote anything write a well-formed argument carrying two to five genuine reasons for the side it had voted. Section 8.1 reports the price of that charity: under it, almost any twenty items cover almost everything, and there is nothing for a ranking rule to do. Real deliberative corpora are not like that: a large share of contributions state a position without giving a logically valid reason for it, and a further share give what the author takes to be a reason for their position but which is in fact a reason for the other side. We model both, with two parameters. G â degenerate fraction.: The share of own-authored items that fail to justify. Swept over 0,0.1,0.3,0.5,0.7\0,0.1,0.3,0.5,0.7\. E â erroneous-support share.: Of those, the share that are erroneous supports â content belonging to the opposing side, filed under this one â rather than merely reason-free. Swept over 0.1,0.2,0.4\0.1,0.2,0.4\. A reason-free item carries a position and no labels: it occupies a slate slot and contributes nothing to any union. An erroneous support carries labels belonging to the other side: it occupies a slot and inflates the denominator with vocabulary that does not belong to the side it was filed under. Two implementation choices make the resulting comparisons trustworthy. First, who authors degenerately is drawn off a dedicated random stream, so raising G changes the content of affected items while leaving the corpus skeleton â which agents author, how many items exist, which links form, how many ballots are cast â bit-identical across the entire grid; every comparison in Section 8.3, across arms and across grid points alike, is therefore exactly paired. Second, completeness is measured against the genuine vocabulary, a reason counting towards a sideâs denominator only if it is a canonical reason for that side. This is the conservative choice: counting opposite-side labels would let erroneous supports inflate the very quantity they are supposed to damage and would reward a slate for surfacing them. We record the denominator-inflating variant alongside so that the size of the measurement artefact is visible rather than assumed away, and Section 8.3 reports it. 5.4. Ranking arms and the evaluation ceiling The arms compared throughout Sections 7â9 are identical in every respect except the ranking step of Algorithm 2. full: The published rule, (8) with both link summands active. endorse-only: Îą=β=0Îą=β=0: a raw endorsement count. enhance-only: β=0β=0: reinforcement summand retained. attack-only: Îą=0Îą=0: rebuttal summand retained. random: The score is ignored and the slate is drawn uniformly from the live pool, using a generator seeded independently of the main stream so that stances, actions and link choices are consumed identically across arms and the comparison stays seed-paired. Alongside these we compute, on the evaluation path only and subject to Principle 1, the greedy cover of Section 3.5. Algorithm 3 states it explicitly so that what it reads â and therefore why it is inadmissible â is on the page. Algorithm 3 GreedyCover â evaluation ceiling only; not a deployable mechanism Given: live pool P, labelling Îť (read directly â violates C4), live vocabulary ÎtĎ ^Ď_t, budget K Yields: a K-subset of P and its completeness 1 Sââ Sâ ; Cââ Câ ; 2 for 11 to minâĄ(K,|P|) (K,|P|) do 3 jââargâĄmaxjâPâSâ|ÎťâĄ(j)âC|j â _jâ P S |Îť(j) C |, ties by ascending identifier; 4 if |ÎťâĄ(jâ)âC|=0|Îť(j ) C|=0 then break; 5 SâSâŞjâSâ SâŞ\j \; CâCâŞÎťâĄ(jâ)Câ CâŞÎť(j ); 6 end for 7 return S and |C|/|ÎtĎ||C|/| ^Ď_t|; 5.5. Persistence: the record is the point Every run writes a single relational database holding: the agents and their initialisation parameters; every justification with its author, side, timestamp and labels; every link with its author, type and endpoints; every ballot with its direction, citation and timestamp; the simulation parameters and seed; and â the entry that matters â every served slate, stored with the identifier of the constituent it was served to, the instant, the policy vector in force, and the item identifiers in served order. This is what makes C5 and C6 testable rather than asserted. Any historical slate can be reconstructed exactly, the corpus state at that instant with it, and the score decomposition of every served item recomputed and checked against the ranking actually served. It is also what makes the ordering and endorsement-mass instruments of Sections 8.2 and 8.5 computable at all, since both read served position. A pipeline logging only the aggregate metric would have required the entire experiment to be repeated. 5.6. Inspecting the record A record nobody can read discharges nothing. Abas ships a browser over the persisted database that presents, for any constituent: the two slates they were served, in order; for each served item, the decomposition (own,EâĄ(j))( own,E(j)), (reinforces,Îąâa)( reinforces,Îą a) and (rebuts,βâr)( rebuts,β r) produced by Algorithm 1; the links that contributed each term, with their authors; and the ballot that followed. It also presents the counterfactual a participant most often wants â the slate the same evidence would have produced under a different policy vector â which is C7 made concrete rather than promised. The interface is deliberately unremarkable: tables and links, no visualisation of anything the arithmetic does not literally contain. The design literature on decentralised civic tools is consistent about non-expert users needing the model of the system to be simple enough to hold in mind Alqahtani & Silaghi, 2017; Alqahtani & Silaghi, 2016; Kattamuri et al., 2005, and a three-term sum is about the limit of what one can expect a participant to check while deciding how to vote. 6. Experimental Setup This section fixes the protocol so that every number in Sections 7â9 can be located. Two propositions are used throughout, one reference configuration anchors all sweeps; the adversary families are stated with the exact capability each is granted; and the statistical conventions are declared in advance, including where we do not correct for multiplicity and what we therefore do not claim. A third proposition is used for the refresh-interval scenario of Section 9.4. 6.1. Propositions UBI asks âShould one address the AI revolution by introducing basic income?â, with thirty canonical reasons per side spanning fiscal, labour-market, administrative, distributive and political-economy considerations. BRA asks âShould one address the AI revolution by introducing basic resource assurance (minimal healthy housing, food, and emergency healthcare)?â. A third proposition, OPT-OUT, asks âShould primary school pupils be able to study without computers and an internet connection?â, with thirty canonical reasons per side spanning pedagogical, developmental, equity-of-access and administrative considerations. It is used only in Section 9.4, where the question is whether a protocol parameter held fixed everywhere else changes the conclusions. 6.2. Reference configuration Unless a sweep states otherwise: N=1000N=1000 constituents; action probabilities Ďown=0.10 _own=0.10, Ďreuse=0.50 _reuse=0.50, Ďabst=0.40 _abst=0.40; link probability Ďlnk=0.60 _lnk=0.60; slate size K=20K=20 per side; author-normalised weight policy (9); policy coefficients Îą=βι=β at the implementation default; score caches re-read after every constituent (M=1M=1), the interval swept in Section 9.4. Table 3 characterises what this produces and establishes that the regime is not degenerate in either direction: the vocabulary saturates, the corpus is roughly 5050 items per side against a 3030-label vocabulary, and the ballot split is near even. Table 3. The reference configuration characterised: BRA proposition, Ďown=0.10 _own=0.10, K=20K=20, Ďlnk=0.60 _lnk=0.60, N=1000N=1000, author-normalised Ďr _r; mean Âą standard deviation over ten seeds. Completeness is that of the slate the endorsement rule actually served. âWithin-run spreadâ is the standard deviation of per-constituent completeness inside a single run, averaged over seeds: it measures temporal inequity, not experimental noise, and at 0.2130.213 it is more than five times the across-seed variation. Quantity Mean Std Combined completeness, mean over constituents 0.806 0.038 Combined completeness, median over constituents 0.886 0.042 Within-run spread of completeness 0.213 0.032 Live endorsing vocabulary (of 3030) 30.0 0.0 Live opposing vocabulary (of 3030) 29.9 0.3 Authored endorsing justifications 48.3 6.6 Authored opposing justifications 49.7 6.8 Rebuttal links 346 19 Reinforcement links 342 15 Total links 688 33 Endorsing ballots 500.9 15.1 Opposing ballots 499.1 15.1 6.3. Structural levers of the non-degenerate authoring reference configuration In the first set of experiments with non-degenerate authoring, four structural levers are swept one at a time about the reference point, ten seeds per cell, both propositions: authoring rate Ďownâ0.02,0.05,0.10,0.20,0.30 _ownâ\0.02,0.05,0.10,0.20,0.30\ with Ďreuse:Ďabst _reuse: _abst held at :45\!:\!4; slate size Kâ5,10,15,20,30Kâ\5,10,15,20,30\; link rate Ďlnkâ0.20,0.40,0.60,0.80,1.00 _lnkâ\0.20,0.40,0.60,0.80,1.00\; electorate Nâ200,500,1000,2000,5000Nâ\200,500,1000,2000,5000\. The ranking ablation runs five arms at Kâ5,10,20,30Kâ\5,10,20,30\ with 100100 seed-paired seeds per cell (40004000 runs) and repeats three arms at K=20K=20 under attack (18001800 runs). The degenerate-authoring sweep runs three arms over Gâ0,0.1,0.3,0.5,0.7ĂEâ0.1,0.2,0.4Gâ\0,0.1,0.3,0.5,0.7\Ă Eâ\0.1,0.2,0.4\ with 5050 seeds per cell per proposition (39003900 runs). The adversarial families are described next. In total the results below rest on roughly 17,00017,000 seeded runs. 6.4. Adversary model In this set of experiments, authentication is assumed to hold: the census is well-formed, nobody votes twice, and no endorsement is forged. Sybil resistance is therefore out of scope as an attack and in scope as an infrastructure requirement, which is where Section 10 takes it up. What a coalition controls is what its members write, where they point their links, and when they act. Coalition sizes are 00, 55, 1010, 1515, 2020 and 2525 percent of N, and members are spread evenly through the processing order, so they enjoy no first-mover endorsement cascade and must rely on the graph for visibility. The attacks studied are: Hub-riding.: Members author items in the ordinary way and attach reinforcement links to the most endorsed same-side item, attempting to launder its endorsement mass into their own score. Label flooding.: Members author on every turn, drawing both labels from a single shared pair, link on every turn, and aim those links at the two most endorsed items. The clones carry real corpus labels, so any damage is redundancy, not label absence. Heterogeneous control.: Identical to flooding in authoring rate, link rate and link targeting, but each member draws its label pair independently: the rate- and corpus-matched control isolating homogeneity. A second variant restores most-related link targeting, isolating hub aiming as a residual. Co-signed flooding.: The coalition is partitioned into groups of size C; each group publishes one poison item per side and every member endorses it and re-asserts the same links, driving the distinct-author numerator of (9) from one to the groupâs per-side size. Câ1,2,5,10,25,50Câ\1,2,5,10,25,50\ at coalitions of a tenth and a quarter, both propositions, both weight policies (48004800 runs). C=1C=1 reproduces the flooding sweep exactly, to the last stored digit, and serves as an in-family control. Nothing in the implementation obstructs any of these strategies. In particular the relation store admits repeated (from,to)(from,to) assertions from distinct authors, which is exactly what makes the co-signing numerator movable â we did not defend against the attack by refusing to represent it. 6.5. Statistical conventions Comparisons between arms are seed-paired and tested with a two-sided paired t-test on per-seed means; comparisons between configurations that do not share seeds use Welchâs unequal-variance test. Reported intervals are across-seed standard deviations unless stated. Where many contrasts are computed on the same stored runs â notably the twenty-four ordering contrasts of Section 8.2 â we report the raw p-values without multiplicity correction and say so at the point of use, and any contrast whose significance would not survive a Bonferroni adjustment at that family size is described as suggestive rather than established. We have not adjusted for the number of hypotheses across the manuscript as a whole, and readers should treat effects at pâ10â2pâ 10^-2 accordingly; the effects we build arguments on sit at p<10â10p<10^-10. 7. Results I: An Attack-Free Electorate We begin with an electorate in which every participant who writes anything writes a competent argument for the side they support (non-degenerate authoring). This regime answers two questions well and one question misleadingly. It establishes which structural levers govern how much of the live reason vocabulary a served slate covers â and the answer, that everything which matters works by enriching the corpus before the voter arrives rather than by selecting more cleverly from it, is stable across both propositions. It establishes what the charter of Section 4 costs, by measuring the served slates against a label-reading ceiling that upper-bounds every mechanism including the ones the charter excludes. 7.1. What the reference configuration produces Table 3 gives the characterisation. Mean combined completeness is 0.806Âą0.0380.806Âą 0.038 on BRA and 0.804Âą0.0320.804Âą 0.032 on UBI: a constituent at the reference configuration meets, on average, four fifths of the reasons that existed on each side at the moment they voted. The number that deserves more attention is the within-run spread, 0.2130.213. This is the standard deviation of completeness across constituents inside a single run, and it is more than five times the standard deviation between runs. The dominant source of variation in what a voter sees is therefore not the seed, the configuration, or the recommender â it is when in the sequence the voter arrived. Early constituents vote against a vocabulary that is still assembling itself; late ones vote against a mature one. This is the temporal inequity of Section 1.1, measured, and it is an inequality in the informational conditions of the ballot that no choice of selection rule addresses. Figure 2 shows how the saturation point moves with authoring rate, and the area to the left of each curve is the population that voted early enough for it to matter. Figure 2. Vocabulary saturation under three authoring rates, anchored at the measured saturation points (â800â 800 constituents at Ďown=0.02 _own=0.02, â150â 150 at Ďown=0.30 _own=0.30, in runs of N=1000N=1000); the vertical rules mark the measured anchors and the shape between them is illustrative. The area to the left of each curve is the population that voted against a reason space still under construction â the quantity a deployment reduces by raising the authoring rate, not by ranking better. 7.2. Which levers move coverage Table 4 gives every setting of every lever on both propositions. The ordering is unambiguous and identical across them. Slate size spans widest â 0.440.44 at K=5K=5 to 0.830.83 at K=30K=30 â and is bounded above by the corpus, since a slate cannot cover reasons nobody has written. The fifth through fifteenth slots contribute the overwhelming majority; slots beyond the twentieth add two to three points per block, and only because the rule is label-blind, so lower-ranked but reason-novel items keep entering. A label-reading selector would exhaust the vocabulary earlier and then add nothing. Authoring rate follows: raising Ďown _own from 0.020.02 to 0.300.30 lifts completeness from 0.700.70 to 0.880.88, by the mechanism visible in Figure 2 â a higher rate saturates the vocabulary sooner, so a smaller fraction of the electorate votes against an immature reason space. It is the cheapest lever per point gained, and one a deployment pulls through interface design and prompting rather than through algorithms. Electorate size behaves identically, 0.690.69 at N=200N=200 rising to 0.900.90 at N=5000N=5000, with the difference that it is not a design parameter; it is worth stating because it means the construction gets better at scale, the opposite of the usual finding for deliberative procedures. Among these levers, link rate does essentially nothing: completeness moves by 0.0040.004 across a fivefold change in linking activity, on both propositions. This is the first appearance of a fact Section 8.1 will sharpen considerably. It does not mean links are useless â links are what the rebuttal and reinforcement summands read, and Section 8.3 shows they carry real weight once the corpus contains material worth discriminating against. It means that in a corpus where every item is a competent argument, the ordering induced by the links reshuffles items that were going to be served anyway. Table 4. Mean combined completeness cÂŻ c at every setting of every lever, both propositions, author-normalised weight policy, ten seeds per cell. Settings run left to right in the order given in the row label; the reference setting is in bold. Across-seed Ď runs 0.0130.013â0.0820.082 throughout and is tabulated per cell in the online appendix. In the authoring-rate row Ďreuse=0.50 _reuse=0.50 is fixed and Ďabst _abst absorbs the remainder; all other levers are held at reference. Lever (settings) prop. cÂŻ c by setting Ďownâ0.02,0.05,0.10,0.20,0.30 _ownâ\0.02,0.05,0.10,0.20,0.30\ BRA 0.702 0.713 0.806 0.841 0.878 UBI 0.684 0.751 0.804 0.853 0.873 Kâ5,10,15,,30Kâ\5,10,15,20,30\ BRA 0.448 0.666 0.763 0.806 0.833 UBI 0.448 0.666 0.768 0.804 0.832 Ďlnkâ0.20,0.40,0.60,0.80,1.00 _lnkâ\0.20,0.40,0.60,0.80,1.00\ BRA 0.806 0.804 0.806 0.804 0.804 UBI 0.799 0.801 0.804 0.805 0.802 Nâ200,500,,2000,5000Nâ\200,500,1000,2000,5000\ BRA 0.691 0.744 0.806 0.850 0.904 UBI 0.680 0.732 0.804 0.865 0.894 7.3. The price of semantic abstinence, measured Section 3.5 introduced the greedy cover as an evaluation ceiling and Principle 1 confined it to the evaluation path. Here we spend it. For every slate served in the reference cell we also compute, over the identical pool and at the identical cache refresh, the slate a label-reading greedy selector would have returned, scoring both with the same measure. This answers the question the charter leaves open: not whether some excluded procedure could score higher, but whether the admissible class contains a procedure good enough to use. It does. Against the end-of-round vocabulary (Table 5) the endorsement rule serves 0.806Âą0.0380.806Âą 0.038 where the ceiling reaches 0.837Âą0.0340.837Âą 0.034: a gap of 0.031Âą0.0140.031Âą 0.014, or 3.7%3.7\% of the ceiling, positive in all ten seeds and never exceeding 0.0590.059. Reading the labels is worth about three points of completeness. The second block is the more consequential reading. Scored against the vocabulary actually live when each slate was served â the denominator of Definition 3.4 â the ceiling attains 1.0001.000 in every seed and the served slates 0.968Âą0.0150.968Âą 0.015. The ceiling saturates because the reference cell ends with a thirty-label vocabulary per side against a twenty-item slate, so a spanning sub-collection almost always exists; the (1âeâ1)(1-e^-1) worst case of Proposition 3.7 is nowhere near binding at this scale. The 0.1940.194 shortfall from cÂŻ=1 c=1 therefore decomposes into two very unequal parts: a selection component of 0.0310.031, which is what abstaining from the labels costs, and a temporal component of 0.1630.163 â vocabulary that did not yet exist when the slate was assembled, which no procedure of any kind, rule-based or learned or label-omniscient, could have shown that voter. Four fifths of the measured incompleteness is a property of sequential deliberation rather than of the recommender. This governs how the rest of the manuscript should be read. Every lever of Section 7.2 that moves completeness substantially does so by attacking the temporal component; the selection component is small and is the only one any better selection procedure could recover, so a learned ranker offered the same pool could at best close 0.0310.031, and only by reading what C4 forbids. Table 5. Served completeness against the label-reading greedy ceiling, reference configuration (BRA), summarised over ten seeds; the per-seed rows are in the online appendix. The ceiling is evaluated on the same pool and at the same cache refresh as the served slate, so the gap isolates selection quality rather than index staleness. The left block scores against the end-of-round vocabulary, the right block against the vocabulary live at the moment of serving (Definition 3.4). The ceiling saturates the live vocabulary in every seed, so the two gaps almost coincide and the residual shortfall in the left block is temporal rather than algorithmic. End-of-round vocabulary Live vocabulary Served Ceiling Gap Served Ceiling Gap Mean over seeds 0.806 0.837 0.031 0.968 1.000 0.032 Std over seeds 0.038 0.034 0.014 0.015 0.000 0.015 Minimum 0.750 0.793 0.011 0.940 1.000 0.011 Maximum 0.862 0.888 0.059 0.989 1.000 0.060 8. Results I: What the Coverage Measure Could Not See This section reports the ablation â does the rule beat not ranking at all? â and the answer, on set coverage with non-degenerate authoring, is no. It reports that null in full, with the seeds on which the ranked and random slates were bit-identical, because a manuscript arguing for legibility owes its readers the result that embarrassed it. It then shows that the null is an artefact of two things: an instrument that discards order by construction, and an authoring model in which every submission is worth reading. Correcting either one dissolves it. Correcting both reveals that the two link summands, which the ablation found exactly inert, are worth as much as the entire ranking advantage once the corpus contains material a rule should keep off the slate. The section closes on the third instrument, endorsement mass, which shows that the random baseline was never competitive on any axis except the one we happened to be measuring. 8.1. The null result, for non-degenerate authoring Five arms, identical in every respect but the ranking step (Section 5.4), 100100 seed-paired seeds per cell, four slate sizes, both propositions â 40004000 runs. One property of the model makes the comparison unusually clean: corpus growth is driven by the authoring draw and not by the slate, so all arms see a bit-identical corpus (102.4102.4 items and 694694 links at zero attackers, 326.4326.4 and 10351035 at a quarter-electorate coalition), and the served slate is the only thing that differs. The link terms are all but inert. Table 6 gives the non-degenerate authoring attack-free sweep. No two of the four ranked arms separate by more than 0.0040.004, and of the twenty-four seed-paired contrasts between ranked arms only one survives a Bonferroni correction for that many comparisons: at K=5K=5 the enhance-only arm leads the pure endorsement count by 0.00390.0039 (p=0.001p=0.001). At that same slate size 5858 of the 100100 BRA seeds are bit-identical between the full rule and the pure endorsement count; at K=20K=20 about a fifth of seeds still are. The mechanism is arithmetic rather than mysterious. Under author normalisation a relation weight is the number of distinct authors of an edge divided by the constituents voting that side, which at the reference configuration ranges from 0.0020.002 to 0.0230.023; multiplied by the endorsement count of a hub it yields a bonus that never exceeded 0.960.96 in the state we instrumented. The gap in raw endorsements at the K-boundary was 33 in that same state. A bonus smaller than the boundary gap reorders items within the slate without changing which items are in it â and completeness, by Remark 1, reads only the union of what is served. The link terms are not doing nothing. They are doing something the instrument is constitutionally unable to see. Table 6. Ranking-term ablation in an non-degenerate authoring attack-free electorate: mean combined completeness pooled over the two propositions, 100100 seed-paired seeds per cell. The live pool holds about 5151 items per side throughout, so K is also a proxy for the fraction of the pool served. Across-seed Ď runs 0.0260.026â0.0490.049 for the four ranked arms and 0.0150.015â0.0250.025 for the random arm. K full endorse-only enhance-only attack-only random 55 0.4500 0.4470 0.4509 0.4498 0.4475 1010 0.6706 0.6708 0.6712 0.6723 0.6652 2020 0.8205 0.8199 0.8202 0.8213 0.8181 3030 0.8511 0.8505 0.8510 0.8510 0.8506 The random baseline is not beaten, and under attack it wins. At non-degenerate authoring and zero attackers the rule leads by 0.82050.8205 against 0.81810.8181, a margin of 0.00240.0024 that is suggestive at p=0.034p=0.034 and does not survive correction for the number of contrasts; at a tenth of the electorate the rule trails, 0.60880.6088 against 0.64720.6472 (â0.0385-0.0385, pâ5Ă10â27pâ 5Ă 10^-27); at a quarter it trails further, 0.37310.3731 against 0.44370.4437 (â0.0706-0.0706, pâ3Ă10â39pâ 3Ă 10^-39), all seed-paired over 200200 runs per comparison. The direction is the mechanism rather than an anomaly: the ranked arms read the endorsement tallies the coalition inflates, and a uniform draw does not, so the clones the flood manufactures reach the slate through the ranking step and not around it. The three-way comparison at the largest coalition is the one to sit with: the author-normalised policy reaches 0.36950.3695 on BRA and 0.37680.3768 on UBI, the random slate 0.44190.4419 and 0.44560.4456, the flat policy 0.22790.2279 and 0.21950.2195. Read at face value the flat policy is some twenty-two points worse than not ranking at all, and author normalisation buys back about two thirds of that deficit but not the whole of it, leaving the rule seven points short of the uniform draw â its virtue under attack is the limiting of amplification rather than the delivery of coverage. The random arm also carries less than half the across-seed variance of the ranked arms under attack (Ďâ0.035Ďâ 0.035 against 0.0780.078), an effect we do not currently model. modelling assumption responsible. A learned selector exhibiting the same null would have offered a metric that failed to move and an unbounded space of explanations. 8.2. Order: what a ranking rule is actually for Completeness is invariant to slate order (Remark 1), and a ranking rule is precisely a claim about order. mâ1,2,3,5,10,20mâ\1,2,3,5,10,20\, the area under the prefix curve, rank-discounted coverage, and e90e_90, at K=20K=20, pooled over both propositions and seed-paired across arms. Table 7 and Figure 3 give the result, and it is unambiguous in exactly the way Table 6 was not. At every prefix short of the full slate the endorsement rule leads the uniform draw, and the lead grows with adversarial pressure. In an attack-free electorate the gains are small but uniformly significant: p1p_1 +0.0094+0.0094 (pâ10â10pâ 10^-10), p5p_5 +0.0198+0.0198 (7Ă10â157Ă 10^-15), AUCAUC +0.0135+0.0135 (3Ă10â193Ă 10^-19), RDCRDC +0.0347+0.0347 (10â1710^-17), e90e_90 â0.40-0.40 positions (5Ă10â105Ă 10^-10). At a coalition of a tenth of the electorate they become p5p_5 +0.1016+0.1016, AUCAUC +0.0558+0.0558 and e90e_90 â3.14-3.14 positions, all at pâ2Ă10â34pâ 2Ă 10^-34; at a quarter, p1p_1 +0.0335+0.0335, p5p_5 +0.1387+0.1387, AUCAUC +0.0796+0.0796, RDCRDC +0.1788+0.1788 and e90e_90 â5.69-5.69 positions, all at pâ¤5Ă10â34p⤠5Ă 10^-34. The reading is direct. A constituent under a quarter-electorate coalition who reads the ranked slate reaches nine tenths of that slateâs coverage after 8.98.9 items; one reading the uniform draw needs 14.614.6. Both slates end at the same p20p_20 â which is why Table 6 saw nothing â but they are not the same object for anyone with finite attention. The coalition succeeds at filling the corpus and fails at placing its material where the reader is. That the margin grows under attack inverts the natural expectation that an adversary manipulating the graph would degrade the ruleâs ordering advantage. What happens instead is that as the pool fills with redundant clones, the difference between ordering by endorsement flow and not ordering at all becomes the difference between a reader meeting distinct reasons early and a reader wading through duplicates â a distinction that barely exists in a clean pool and dominates in a polluted one. The link terms remain nearly invisible on this instrument too. Comparing the full rule against endorse-only, no contrast reaches significance at zero or a tenth attackers; at a quarter, RDCRDC gains 0.0130.013 (p=0.011p=0.011) and AUCAUC gains 0.0040.004 (p=0.027p=0.027). With twenty-four contrasts computed on the same stored runs and no multiplicity correction (Section 6.5), we report this as suggestive and nothing more. The next subsection is where the link terms stop being suggestive. Figure 3. Prefix coverage against slate depth at K=20K=20, pooled over both propositions, 28002800 stored runs. Panels: attack-free electorate, coalition at a tenth, coalition at a quarter. The shaded region is what the endorsement rule delivers above the uniform draw at each depth; ticks on the axis mark e90e_90, the depth at which each arm reaches nine tenths of its own final coverage. The three arms meet at m=20m=20 in the attack-free panel and separate everywhere before it, by a margin that widens with attack; under attack the endorsement rule ends m=20m=20 below the draw â the only quantity Table 6 could see â and leads it at every earlier depth. Table 7. Ordering measures at K=20K=20, pooled over both propositions, seed-paired, 28002800 stored runs. pmp_m is prefix coverage at depth m; AUCAUC is the mean over depths; RDCRDC is rank-discounted coverage (6); e90e_90 is the depth reaching nine tenths of that slateâs own final coverage (lower is better). |A||A| is the coalition size out of N=1000N=1000. Note that p20p_20 â the only column completeness can see â agrees across arms to within 0.0050.005 in the attack-free electorate, and under attack moves against the ranked rule while every earlier column moves for it. |A||A| arm p1p_1 p3p_3 p5p_5 p10p_10 AUCAUC RDCRDC e90e_90 p20p_20 00 full 0.1290 0.3301 0.4738 0.6865 0.4466 1.3298 11.74 0.8205 endorse-only 0.1281 0.3267 0.4720 0.6872 0.4451 1.3258 11.68 0.8199 random 0.1165 0.3054 0.4476 0.6650 0.4284 1.2833 12.28 0.8181 100100 full 0.1205 0.2919 0.4066 0.5457 0.3648 1.0705 09.79 0.6088 endorse-only 0.1182 0.2898 0.4068 0.5635 0.3732 1.1021 10.49 0.6474 random 0.0901 0.2048 0.2894 0.4509 0.3060 0.9544 14.38 0.6472 250250 full 0.1200 0.2597 0.3167 0.3545 0.2712 0.7884 06.29 0.3731 endorse-only 0.1201 0.2699 0.3435 0.4050 0.2974 0.8670 07.97 0.4389 random 0.0790 0.1400 0.1857 0.2864 0.2082 0.6676 14.61 0.4437 8.3. Degenerate authoring: giving the rule something to discriminate against The second correction addresses Remark 2. Under the model of Section 7 every submission is a competent argument, so every item is worth roughly the same to a coverage measure and no rule can beat a draw by much. Section 5.3 introduced two parameters â the degenerate fraction G and the erroneous-support share E â and this subsection sweeps them: three arms, Gâ0,0.1,0.3,0.5,0.7Gâ\0,0.1,0.3,0.5,0.7\, Eâ0.1,0.2,0.4Eâ\0.1,0.2,0.4\, 5050 seeds per cell per proposition, 39003900 runs, zero failures, twenty minutes of wall clock. Because degenerate authorship is drawn off a dedicated stream, the corpus skeleton is bit-identical across the whole grid and every comparison is exactly paired. Table 8 gives the result: the margin the coverage instrument could not find in Table 6 reappears as soon as there is anything to discriminate against, and grows monotonically with how much there is. At G=0G=0 the arms are all but indistinguishable, matching the attack-free sweep of Table 6 (+0.0033+0.0033, p=0.018p=0.018) â the control establishing that the sweep measures what it claims. At G=0.1G=0.1 the rule already leads the uniform draw by 0.00590.0059 (p=7Ă10â4p=7Ă 10^-4); by G=0.7G=0.7 it leads by 0.04710.0471 (p=8Ă10â14p=8Ă 10^-14), more than an order of magnitude larger. The served degenerate share tracks it: at G=0.7G=0.7 the rule serves 66.6%66.6\% degenerate material against the 70.5%70.5\% corpus base rate a uniform draw returns, significant at p<10â6p<10^-6 in every cell. Raising E â moving degenerate items from reason-free towards erroneous-support â consistently shrinks the margin, by 0.0050.005 to 0.0100.010 across the grid. This is the expected direction and a useful sanity check: an erroneous support carries real labels and real endorsement potential, so a rule reading only endorsements and links finds it harder to distinguish from a genuine item than a reason-free stub. The ruleâs advantage is largest exactly where the material is most obviously worthless. We also record the measurement artefact rather than assuming it away: scoring against the denominator-inflating variant at G=0.3G=0.3, E=0.4E=0.4 gives 0.62270.6227 where the genuine-vocabulary measure gives 0.74240.7424, because erroneous supports push the apparent per-side vocabulary from 29.629.6 labels to 44.844.8. Reporting the raw variant would have made every arm look worse and credited slates for surfacing material filed on the wrong side; the genuine-vocabulary measure of Section 5.3 is the conservative choice and is what Table 8 reports throughout. Ordering and coverage now agree. At G=0.7G=0.7, E=0.1E=0.1 the rule reaches p5=0.4158p_5=0.4158 against the drawâs 0.18620.1862, AUCAUC 0.36290.3629 against 0.21040.2104, and e90e_90 of 7.397.39 against 13.2513.25. The two instruments that disagreed in Sections 8.1 and 8.2 disagree because of the authoring model, and once it is corrected they say the same thing. Table 8. Degenerate authoring: mean combined completeness against the genuine reason vocabulary, pooled over both propositions, 100100 runs per cell, seed-paired across arms. G is the fraction of own-authored items that fail to justify; E is the share of those that are erroneous supports rather than reason-free. p-values are two-sided Wilcoxon signed-rank tests on the full-minus-random contrast. The G=0G=0 row reproduces the attack-free sweep of Table 6 and serves as the control. G E full endorse-only random full â- random p full â- endorse 0.00.0 â 0.8232 0.8233 0.8199 +0.0033+0.0033 0.0180.018 â0.0000-0.0000 0.10.1 0.10.1 0.7996 0.7982 0.7940 +0.0056+0.0056 2.3Ă10â32.3Ă 10^-3 +0.0014+0.0014 0.10.1 0.40.4 0.7993 0.7982 0.7940 +0.0053+0.0053 5.5Ă10â35.5Ă 10^-3 +0.0011+0.0011 0.30.3 0.10.1 0.7479 0.7323 0.7268 +0.0210+0.0210 3.6Ă10â123.6Ă 10^-12 +0.0156+0.0156 0.30.3 0.40.4 0.7424 0.7323 0.7268 +0.0156+0.0156 9.2Ă10â99.2Ă 10^-9 +0.0102+0.0102 0.50.5 0.10.1 0.6639 0.6242 0.6248 +0.0392+0.0392 4.9Ă10â154.9Ă 10^-15 +0.0397+0.0397 0.50.5 0.40.4 0.6538 0.6242 0.6248 +0.0290+0.0290 4.5Ă10â124.5Ă 10^-12 +0.0295+0.0295 0.70.7 0.10.1 0.5764 0.5147 0.5104 +0.0661+0.0661 2.9Ă10â162.9Ă 10^-16 +0.0617+0.0617 0.70.7 0.40.4 0.5596 0.5147 0.5104 +0.0492+0.0492 1.9Ă10â141.9Ă 10^-14 +0.0449+0.0449 8.4. Why the link terms square the discrimination Table 8âs last column is the surprise. The link summands, worth â0.0000-0.0000 at G=0G=0, are worth +0.0617+0.0617 at G=0.7G=0.7 â which is 93%93\% of the full margin over the uniform draw. Terms that Table 6 found all but inert become, in a realistic corpus, the substance of the rule. This subsection explains why, by measuring the graph directly. Recall the direction of (8): an item is credited for the endorsements of what it points at. So the question is whether a degenerate authorâs outgoing links land on lower-endorsed targets than a genuine authorâs do. Instrumenting 100100 runs at G=0.7G=0.7, E=0.4E=0.4 gives the answer in one table of five numbers per item class: item class count own endorsements reinforcement out-edges per item link credit genuine 30.2 11.04 196.1 6.5 274.25 erroneous support 28.6 06.95 117.5 4.1 136.28 reason-free 43.6 01.43 034.2 0.8 038.12 On endorsements alone, genuine items lead reason-free ones by a factor of 7.77.7. On link credit they lead by a factor of 7.27.2 â and the two factors multiply, because an itemâs link credit is the sum of its targetsâ endorsements and a degenerate item both asserts fewer links (0.80.8 per item against 6.56.5) and aims the few it asserts at material nobody adopted. The link summands do not introduce new discrimination. They square the discrimination already present in the endorsement count. This also explains the inertness at G=0G=0 without special pleading. When every item is genuine, every item has both a healthy endorsement count and healthy out-edges into other well-endorsed items; squaring a ratio of one leaves one. The link terms were never a mechanism for separating good arguments from good arguments. They are a mechanism for separating arguments from non-arguments, and we had not given them any non-arguments. One limitation must be stated here rather than deferred, because it governs the reading of the entire subsection. The chain rests on a single modelled fact â degenerate items are adopted less (1.431.43 endorsements against 11.0411.04) â and that fact follows from the simulatorâs adoption step, which selects from the served slate by TFâIDF cosine similarity to the agentâs position text (Section 5.1). It is a property of the agents, not of the rule. The finding is therefore that the rule inherits and amplifies whatever discrimination the electorate itself exercises. The estimation of the rate at which human electorate discriminates is an empirical question this manuscript does not attempt to answer (Section 12). 8.5. Coverage against endorsement mass The third instrument closes the account. Sections 8.1 and 8.2 left an uncomfortable residue: on coverage the uniform draw was never beaten, and under flooding attacks it beat the rule. Instrument I (7) asks the question coverage cannot â did the slate show the voter what the electorate had actually taken up? â and the answer is that the draw was never competitive. Table 9 gives five regimes: the attack-free non-degenerate corpus, two degenerate-authoring levels, and two coalition sizes. The uniform draw captures 0.490.49 of the achievable endorsement mass in the attack-free regime against the ruleâs 0.760.76, and 0.170.17 against 0.550.55 at a quarter-electorate coalition â a factor of 3.33.3. Its coverage â level with the rule in the attack-free regime, ahead of it under attack â is purchased entirely by serving material nobody had adopted. A voter shown the random slate met at least as many reasons, attached to items their fellow constituents had, in the main, passed over. The greedy ceiling behaves in a way worth dwelling on. It attains coverage 1.00001.0000 in every regime, as Section 7.3 led one to expect. On endorsement mass it dominates the rule in the three attack-free regimes (0.790.79 against 0.760.76 at G=0G=0) and is dominated under attack (0.380.38 against 0.550.55 at a quarter-electorate coalition). Under flooding, the label-reading selector buys its perfect coverage by abandoning what constituents endorsed â it reaches for the rare labels, which under a homogeneity attack are the ones the coalition has not saturated and which almost nobody has adopted. The inadmissible mechanism is not merely inadmissible; on the axis that measures whether a slate reflects the electorate, it is worse under exactly the conditions a civic deployment must survive. Finally, at G=0.5G=0.5 the link terms buy +2.8+2.8 points of coverage for â2.4-2.4 points of endorsement mass. That is not an error and not a defect. It is a frontier: the two link summands trade breadth of reasons against fidelity to what people endorsed, and there is no system-level fact about which is correct. It is exactly the kind of question C7 exists to hand to the person it affects, and Section 11 takes up what it means that our clearest quantitative finding turns out to be a menu rather than an answer. Table 9. Coverage against endorsement mass (7) across five regimes, pooled over both propositions. G is the degenerate fraction at E=0.1E=0.1; |A||A| is coalition size out of N=1000N=1000. The greedy ceiling attains coverage 1.00001.0000 in every regime, so only its mass is shown. The uniform draw is level with the rule on coverage in the attack-free regime and ahead of it under attack, and loses on mass everywhere, by a factor rising to 3.33.3; the ceiling dominates on mass in the three attack-free regimes and is dominated in both attacked ones. coverage endorsement mass ceiling regime full end.-only random full end.-only random mass G=0G=0, no coalition 0.8232 0.8233 0.8199 0.7560 0.7601 0.4906 0.7905 G=0.3G=0.3 0.7479 0.7323 0.7268 0.7458 0.7610 0.4915 0.7645 G=0.7G=0.7 0.5764 0.5147 0.5104 0.7215 0.7590 0.4906 0.7920 |A|=100|A|=100 0.6088 0.6474 0.6472 0.6895 0.6943 0.2916 0.5875 |A|=250|A|=250 0.3731 0.4389 0.4437 0.5460 0.5579 0.1673 0.3809 9. Results I: A Coalition in the Electorate We now admit a coordinated coalition under the model of Section 6.4: authentication holds, so nobody votes twice and no endorsement is forged, and the coalitionâs only instruments are what it writes, where it points its links, and when it acts. The two natural strategies produce opposite outcomes, and the contrast is the most useful result in this manuscript for anyone actually deploying such a system â it says that the intuitive attack is harmless, that the unintuitive one is severe, and that the defence against the severe one is a single line of the specification. We then decompose the damage to establish what causes it, and give the coalition its best available escalation to establish that the defence does not merely displace the problem. 9.1. Hub-riding is inert When coalition members author ordinary items and attach reinforcement links to the leading same-side item, in order to launder its endorsement mass into their own score, completeness does not move. At every coalition size, on both propositions, the change from the zero-attacker cell is statistically indistinguishable from zero (n=200n=200 seeds per cell, all p>0.50p>0.50 under Welchâs test), and the two weight policies differ by less than 0.0020.002. The reason is structural. Completeness is a property of a union of sets, and exchanging one ordinary item for another changes which elements contribute to the union without changing the union much. Hub-riding achieves its proximate goal â the attackerâs argument gets seen â while leaving the informational quality of the slate intact. Whether this constitutes an attack at all is a fair question: a system in which arguing that your point extends a popular point makes your point more visible is arguably working as designed, and this is one of the few places where the right response to an adversarial finding is to accept the behaviour rather than defend against it. The result also confirms the design choice of Section 4.4 empirically. (8) takes exactly one hop and computes no fixed point, so there is no eigenvector to farm Page et al., 1999; a coalition that constructs a hub and points its whole corpus at it gains a bounded, one-hop bonus and nothing compounding. 9.2. Label flooding, and the weight policy as a security control The picture inverts entirely when the clones are label-identical. Table 10 gives the sweep, 200200 seeds per cell, spread timing so the coalition enjoys no first-mover cascade. Under the author-normalised policy completeness falls from 0.8160.816 at zero attackers to 0.3770.377 at a quarter of the electorate on BRA, and 0.8150.815 to 0.3720.372 on UBI. Under the flat policy the same coalition drives it to 0.2280.228 and 0.2190.219. Every non-zero cell differs from its baseline at p<.001p<.001, and every normalised-versus-flat comparison at a non-zero coalition size differs at p<.001p<.001 as well. The defence margin â the completeness that author normalisation retains and a flat policy loses â is 0.1170.117 at a coalition of five percent, widens to 0.1810.181 at fifteen and eases to 0.1510.151 at twenty-five. Section 4.5 claimed the relation-weight function is a security control rather than a tuning knob; this is the measurement behind the claim, and it should be read against Section 7.2, where the four structural levers of the attack-free regime were the only things that moved completeness at all. The arithmetic behind the defence is Example 3.9 scaled up. Under the flat policy a lone cloneâs link into the hub transfers the hubâs entire endorsement mass, so every clone inherits a score comparable to the most successful non-coalition item in the pool and the top of the ranking fills with duplicates. Under author normalisation the same link transfers a fraction |AâĄ(j,k)|/VĎ|A(j,k)|/V^Ď; a solo clone inherits about a thousandth of the hubâs mass. To buy what the flat policy gives away, the coalition must make many members assert the same link â which means the cost of the attack scales with its size, and, more importantly, that the attack becomes visible in the public link record as an anomalous concentration of identical assertions. A defence that converts a covert manipulation into an overt one is doing exactly what a civic mechanism should do, and note that it does so without any term that reads which participants asserted the link: (9) counts distinct authors and remains compliant with C3. Table 10. Flooding attack under spread timing: mean combined completeness against coalition size, both weight policies, both propositions, 200200 seeds per cell. The final column gives the defence margin â how much completeness author normalisation retains that a flat policy loses â averaged over the two propositions. Hub-riding, not shown, produces no significant change at any coalition size. BRA UBI Coalition normalised flat normalised flat margin 0%0\% 0.816 0.814 0.815 0.814 â 5%5\% 0.713 0.597 0.708 0.590 0.1170.117 10%10\% 0.608 0.440 0.602 0.427 0.1720.172 15%15\% 0.519 0.340 0.514 0.331 0.1810.181 20%20\% 0.443 0.282 0.430 0.265 0.1630.163 25%25\% 0.377 0.228 0.372 0.219 0.1510.151 9.3. The mean over sides hides how one-sided the damage is One caveat applies to both aggregators in Equations 4 and 3. About 2%2\% of constituents â 19.619.6 of 987987 at the reference configuration â receive a slate on one side only and so contribute a single value, whereas the factor 1/2âN1/2N in (3) presumes two. The figure we report as cÂŻ c throughout is the mean over slates actually served, which coincides with (3) when every voter is served on both sides and otherwise sits 0.0030.003â0.0070.007 below it, the one-sided voters being those whose remaining side is the sparse one; cÂŻmin c is necessarily restricted to the voters served on both. Both discrepancies are an order of magnitude smaller than any effect discussed here, but the asymmetry is worth naming: an unserved side is a coverage failure that neither equation counts as one. The substantive conclusions are unchanged. The minimum tracks the mean at an offset of 0.040.04â0.080.08 across the whole grid, the damage it registers is 55â13%13\% larger than the meanâs at every coalition size, and the advantage of author normalisation over a flat policy is not merely preserved but wider under it, 0.1330.133â0.1880.188 against 0.1180.118â0.1820.182, at the same significance. The choice of aggregator does not decide any claim we make. What it does change is what a reader can see. Table 11 reports, for the same runs as Table 10, the mean over voters of the worse-covered side, the mean within-voter gap between the two sides, and the share of constituents whose worse side falls below one half. The gap roughly doubles under attack, from 0.0850.085 with no coalition to 0.150.15â0.160.16 at the largest ones: flooding does not lower coverage evenly, it unbalances it. The last column is the consequence. At a 15%15\% coalition under author normalisation the mean reads 0.520.52, which invites the reading that a typical voter still meets half of the reasons; in those same runs 58%58\% of constituents have a side below 0.50.5, and under a flat policy 94%94\% do. Both statements are true of the same electorate, and only the first is visible in (3). Table 11. Side imbalance under label flooding, on the runs of Table 10 (200 seeds per cell, pooled over the two propositions). âminâ is cÂŻmin c of (4), âgapâ the mean over voters of the difference between their two sides, and â<0.5<\!0.5â the share of voters whose worse side falls below one half. author-normalised flat Coalition min gap <0.5<\!0.5 min gap <0.5<\!0.5 0%0\% 0.779 0.085 0.118 0.777 0.086 0.118 5%5\% 0.661 0.107 0.130 0.528 0.138 0.348 10%10\% 0.538 0.140 0.257 0.351 0.168 0.782 15%15\% 0.447 0.144 0.581 0.259 0.155 0.941 20%20\% 0.359 0.158 0.878 0.194 0.160 0.987 25%25\% 0.300 0.150 0.978 0.149 0.149 0.997 We keep (3) as the reported instrument. It is linear, so a drop in it decomposes into the per-voter and per-side contributions that make an attack attributable to a subpopulation, and it is the expectation of the coverage seen by a reader who stops at a uniformly random point, which is the quantity the three measures are jointly framed around; the minimum has neither property. But the mean supports a narrower reading than it appears to. A completeness of 0.520.52 is a statement about an average side, not about an average voter, and an adversary who concentrates on one side is rewarded by exactly that difference. Where a deployment sets a threshold below which it will not certify a poll, the threshold belongs on the imbalance, not on the mean. 9.4. The refresh interval matters only when someone is attacking Score caches are re-read after every constituent (Section 6.2), so no slate is scored against a snapshot older than the vote before it. Widening the interval to M constituents makes it the latency of the endorsement feedback loop, and that lag reinterprets both preceding results at once: a coalition accumulating endorsement mass gets M votes of head start before the quantity it inflates is read again, and a rule scoring against a stale snapshot is not ranking the corpus the constituent is about to see. What lags is the tallies and the relation weights, not the corpus: a newly authored item is appended and forces a re-rank at any M. Crossing M=200M=200 against the reported M=1M=1 with the coalition makes the objection an estimable interaction. The experiment was run on a third proposition, OPT-OUT (should primary pupils be able to study without computers and an internet connection?, thirty canonical reasons per side spanning pedagogical, developmental, equity-of-access and administrative considerations), and is reported on its own rather than pooled with the two of Section 6.1: Mâ1,200Mâ\1,200\ crossed with no coalition against the co-signed flood of Section 9.6 at a tenth of the electorate in groups of C=5C=5, three arms, the two extremes of the degenerate axis, 5050 seeds shared across all four corners, 12001200 runs. Table 12 supports two conclusions of different kinds. With no coalition the interval changes nothing the ordering claims rest on: on the non-degenerate pool fullâ-random is â0.0011-0.0011 at M=1M=1 and â0.0015-0.0015 at M=200M=200 (both p>0.5p>0.5), both consistent with the attack-free contrast of Table 7 (+0.0024+0.0024, suggestive only); under degenerate authoring the ruleâs advantage is +0.0222+0.0222 and +0.0202+0.0202 (4Ă10â64Ă 10^-6, 1Ă10â51Ă 10^-5) â the same effect whether the loop is fresh or two hundred votes stale. With the coalition the interval matters, in the unhelpful direction: lagging the refresh to M=200M=200 raises the ranked arm by 0.0130.013 of completeness and leaves the uniform draw untouched. The interaction â the effect of the lag on fullâ-random under attack minus the same effect without â is +0.0136+0.0136 on the non-degenerate pool (95% CI [+0.0062,+0.0209][+0.0062,+0.0209], p=5Ă10â4p=5Ă 10^-4) and +0.0069+0.0069 under degenerate authoring (CI [â0.0011,+0.0150][-0.0011,+0.0150], p=0.089p=0.089). Table 12. Completeness at the four corners of the refresh-interval Ă coalition results table, OPT-OUT, 5050 seeds per cell, author-normalised weights, co-signed flood at C=5C=5. M is the number of constituents between cache refreshes; M=1M=1 is the reference configuration. The fullâ-random rows are seed-paired differences with two-sided paired t-tests. E is inert at G=0G=0. no coalition 10%10\% coalition Load arm M=1M=1 M=200M=200 M=1M=1 M=200M=200 G=0G=0 full 0.815 0.815 0.707 0.720 endorse-only 0.815 0.815 0.717 0.729 random 0.816 0.816 0.742 0.742 fullâ-random â0.0011-0.0011 â0.0015-0.0015 â0.0348â-0.0348^*** â0.0216â-0.0216^*** G=0.5G=0.5, E=0.4E=0.4 full 0.642 0.640 0.534 0.539 endorse-only 0.625 0.625 0.527 0.538 random 0.620 0.620 0.552 0.552 fullâ-random +0.0222â+0.0222^*** +0.0202â+0.0202^*** â0.0174ââŁâ-0.0174^** â0.0124â-0.0124^* The mechanism is the order-invariance of Remark 1, visible slate by slate; Table 13 isolates each arm. random never consults the cached scores, and its slates are identical item for item and position for position across intervals at every corner, so its completeness is exactly equal on all 5050 seeds. endorse-only is exactly invariant too, but only without a coalition: over ten seeds its 19,70219,702 slates have identical membership at M=1M=1 and M=200M=200 while 93%93\% are served in a different order, which completeness cannot see. Under the coalition that breaks â at seed 33 only 801801 of 19881988 slates keep their membership â and completeness moves +0.0120+0.0120 (5Ă10â105Ă 10^-10). What the flood supplies is not staleness but speed: it makes tallies move far enough within one window for the stale snapshot to select a different twenty items. full also reads relation weights, whose staleness changes which items are served on 40%40\% of slates at G=0G=0 and 49%49\% at G=0.5G=0.5 even with no attacker, but there the coverage gained and lost cancels (â0.0004-0.0004, â0.0020-0.0020, both p>0.35p>0.35). Table 13. Cost of lagging the refresh interval, per arm: mean seed-paired M=200M=200 minus M=1M=1 completeness, OPT-OUT, 5050 seeds. exact marks cells in which the two intervals produce the same completeness on every seed to the last stored digit, so no test applies. no coalition 10%10\% coalition Load arm Î p Î p G=0G=0 full â0.0004-0.0004 0.790.79 +0.0132+0.0132 3Ă10â43Ă 10^-4 endorse-only 0.0000 +0.0000 exact +0.0120+0.0120 5Ă10â105Ă 10^-10 random 0.0000 +0.0000 exact 0.0000 +0.0000 exact G=0.5G=0.5, E=0.4E=0.4 full â0.0020-0.0020 0.360.36 +0.0050+0.0050 0.170.17 endorse-only 0.0000 +0.0000 exact +0.0110+0.0110 4Ă10â54Ă 10^-5 random 0.0000 +0.0000 exact 0.0000 +0.0000 exact The prefix measures agree at both intervals. Rankingâs lead peaks at p5p_5 in five of the eight corner-by-load cells and between p2p_2 and p10p_10 in the rest: under degenerate authoring it is +0.1363+0.1363 against +0.1133+0.1133 with no coalition and +0.0786+0.0786 against +0.0652+0.0652 with one. The lead is lost before the full slate in one cell only â the non-degenerate pool under attack at M=200M=200, +0.0011+0.0011 at p5p_5 and â0.0006-0.0006 at p10p_10, the corner leaving a ranking rule least to discriminate between. Scoring each constituent on the worse-served of their two sides rather than averaging leaves the pattern intact: fullâ-random on that side is â0.0365-0.0365 and â0.0222-0.0222 under attack on the non-degenerate pool, +0.0208+0.0208 and +0.0213+0.0213 under degenerate authoring without one (all p<5Ă10â4p<5Ă 10^-4), and the gap between a constituentâs two slates is null in all eight cells (pâĽ0.26p⼠0.26). Two contrasts, at p=0.043p=0.043 and p=0.089p=0.089, are the kind Section 6.5 asks be read as suggestive; the null without a coalition, the interaction on the non-degenerate pool and the exact invariance of the unranked arm do not depend on them. For the charter the reading is a small negative one, worth stating because the opposite is the intuitive guess. M is a published policy parameter like (9), and refreshing after every vote sounds like the setting that keeps up best with an attack. It is not: under a coalition the shorter interval carries the flood into the ranked slates slightly faster and leaves the uniform draw exactly where it was, and with no coalition it changes nothing. Reporting at M=1M=1 is thus the conservative choice â at M=200M=200 the rule would have looked 0.0130.013 better under attack than it does â and the defence remains the line that counts distinct authors, not the rate at which its inputs are read. 9.5. The collapse is homogeneity, not hyperactivity A flooding attacker differs from a non-coalition constituent in four ways at once, and only one of them is the claim. It authors on every turn rather than with probability Ďown _own; it authors the same narrow label set as every other member; it links on every turn rather than with probability Ďlnk _lnk; and it aims its links at the two most endorsed items rather than at the most related one. The first and third inflate the corpus and the link graph, the fourth concentrates endorsement mass, and any of them could depress completeness alone, so comparing a coalition against an attacker-free baseline confounds all four. We therefore ran the two rate-matched controls of Section 6.4, at every coalition size, on both propositions, under both weight policies, 100100 seeds per cell, seed-paired with the flooding sweep (40004000 runs). The match is tight â at a tenth of the electorate on BRA the flooding arm produces 190.6190.6 items and 837837 links against the heterogeneous controlâs 191.6191.6 and 834834 â so the arms differ in label distribution and in nothing else the pipeline can observe. Table 14 gives the decomposition. Hyperactivity is real but small: a coalition holding a quarter of the electorate that authors and links on every turn with heterogeneous labels costs 0.0670.067 of completeness under normalisation, for an innocuous reason â tripling the corpus enlarges the denominator |ÎĎ|| ^Ď| faster than a fixed twenty-item slate can cover it. Homogeneity costs 0.3800.380 on top of that. Across coalition sizes and both weight policies, between 78%78\% and 85%85\% of the collapse is attributable to label homogeneity alone, and the share rises with coalition size: volume damage saturates while homogeneity damage does not. Hub aiming, isolated by the third arm, is negligible under normalisation (â¤0.006⤠0.006 at every coalition size, all pâĽ0.14p⼠0.14) and worth at most 0.0240.024 under the flat policy â consistent with the mechanism, since author normalisation is precisely the rule that discounts many identical assertions aimed at one target. Two findings now agree from opposite directions: Section 7.2 found link volume harmless, and this control finds authoring volume largely harmless too. What the pipeline cannot absorb is many voices saying one thing. Table 14. Decomposing the flooding collapse, pooled over the two propositions, 100100 seeds per cell. Heterogeneous is the rate- and corpus-matched control in which attackers author and link at the same inflated rates and aim at the same hubs but draw labels independently. Îvol _vol is completeness lost to sheer volume (baseline minus heterogeneous); Îhom _hom is the additional loss attributable to label homogeneity alone (heterogeneous minus flooding). Every Îhom _hom is significant at p<10â40p<10^-40 under Welchâs test. mean completeness decomposition Coalition heterog. flooding (baseline) Îvol _vol Îhom _hom homog. share author-normalised weights 5%5\% 0.797 0.711 0.821 0.0240.024 0.0860.086 79%79\% 10%10\% 0.780 0.609 0.821 0.0400.040 0.1720.172 81%81\% 15%15\% 0.772 0.519 0.821 0.0490.049 0.2530.253 84%84\% 20%20\% 0.762 0.435 0.821 0.0580.058 0.3280.328 85%85\% 25%25\% 0.753 0.373 0.821 0.0670.067 0.3800.380 85%85\% flat weights 5%5\% 0.769 0.595 0.818 0.0490.049 0.1740.174 78%78\% 10%10\% 0.750 0.437 0.818 0.0680.068 0.3130.313 82%82\% 15%15\% 0.743 0.341 0.818 0.0740.074 0.4020.402 84%84\% 20%20\% 0.737 0.275 0.818 0.0810.081 0.4620.462 85%85\% 25%25\% 0.731 0.224 0.818 0.0870.087 0.5070.507 85%85\% 9.6. The coalition cannot coordinate its way out The defence argument of Section 9.2 describes what a coalition must do to defeat normalisation; it does not show that doing it fails. A co-signing coalition is strictly more powerful than one that does not, since co-signing costs it nothing in membership, so we gave it the capability. The coalition is partitioned into groups of C; each group publishes one poison item per side, and every member endorses that item and re-asserts the same two links from it, driving the distinct-author numerator of (9) from one to the groupâs per-side size. C=1C=1 is the attack of Table 10 and reproduces its per-seed completeness exactly, to the last stored digit, across all eight cells â an in-family control. We swept Câ1,2,5,10,25,50Câ\1,2,5,10,25,50\ at coalitions of a tenth and a quarter, both propositions, both weight policies, 48004800 runs. Table 15 reports the cleanest finding in the manuscript: coordination is strictly counterproductive. Completeness rises monotonically in C at every coalition size, on both propositions, under both weight policies. The uncoordinated C=1C=1 attack is the coalitionâs optimum within the family, and a fully coordinated coalition holding a quarter of the electorate leaves completeness at 0.7590.759 against an attacker-free baseline of 0.8150.815 â it has spent a quarter of the electorate to buy six points. The reason is a rate mismatch the coalition cannot escape. Raising the numerator requires spending members on one item; members are finite; so the number of payload-carrying items falls as 1/C1/C while the co-signature count rises only linearly, and it rises against a denominator â constituents voting that side â that co-signing does not touch. Measured on BRA at the 25%25\% coalition, going from C=1C=1 to C=50C=50 multiplies mean co-signers per edge by 3.43.4 and the most concentrated edge by 1.81.8, while cutting poison items from 250.1250.1 to 10.010.0, a factor of 2525. The non-coalition corpus is unmoved throughout (76.376.3 items at C=1C=1, 76.776.7 at C=50C=50). Completeness is damaged by items occupying slate slots, and the coalition has traded away twenty-five of those for every threefold gain in transfer. Table 15. The coordinated adversary, pooled over the two propositions, 100100 seeds per cell. Group size C is the number of coalition members sharing one poison item and co-signing its links; C=1C=1 is the uncoordinated attack of Table 10. Payload items and co-signers per edge are measured on BRA at the 25%25\% coalition. Completeness rises monotonically with C in every cell: coordination helps the electorate, not the coalition. coalitionâs position 10%10\% coalition 25%25\% coalition C payload items co-signers/edge normalised flat normalised flat 11 250250 1.61.6 0.609 0.437 0.373 0.224 22 188188 2.02.0 0.644 0.494 0.425 0.268 55 9797 3.03.0 0.710 0.629 0.530 0.404 1010 5050 4.14.1 0.755 0.735 0.612 0.537 2525 2020 5.35.3 0.785 0.781 0.722 0.711 5050 1010 5.55.5 0.793 0.792 0.759 0.754 Note also what happens to the gap between the two weight policies as C rises: it closes, from 0.1490.149 at C=1C=1 to 0.0050.005 at C=50C=50 at the 25%25\% coalition. This is the expected signature. Normalisation exists to discount solo assertions; when the coalition co-signs everything, the two policies are computing nearly the same quantity, and the coalition has arrived at a regime where its own concentration is what limits it. 9.7. What the adversarial results say about the charter Three consequences, each an argument for legibility rather than an argument that merely happens to be compatible with it. First, the effective defence is a published policy parameter, not a detector: nothing in this section trains a classifier, scores a participantâs trustworthiness, or removes anyoneâs material. (9) is one line, it counts distinct authors, and it is worth ten points of completeness against a coalition holding a quarter of the electorate. A deployment can publish it, a participant can verify it was applied, and an adversary who reads it learns only that the attack is expensive â the property one wants a published defence to have. Second, the defence works by making the attack visible. A coalition that pays the cost of co-signing writes its coordination into the public link record, where an anomalous concentration of identical assertions from distinct authors is directly observable by anyone holding a replica. A learned ranker that resisted the same attack would do so through parameters nobody can inspect, leaving participants unable to distinguish successful defence from successful attack. Third, the results relocate rather than remove the residual risk. Under authentication the surviving exposure is not manipulation of the ranking â hub-riding is inert, co-signing is self-defeating â but corpus capture: a coalition large enough simply crowds the reason space with redundancy, and Section 9.5 shows four fifths of the damage comes from that homogeneity rather than from anything the selector does. The corresponding mitigations are not ranking mitigations. They are census integrity (Section 10), which is what bounds coalition size at all, and slate capacity, which Section 7.2 shows is the widest lever available. A recommender is not where this class of attack should be defeated, and a charter-compliant one at least makes that visible on the record. 10. Peer-to-Peer Realisation on DDP2P The charter of Section 4 can be honoured on a central server, but only as a promise. Determinism, evidence locality and reproducibility are properties of a computation, and if one party owns the machine that performs it, participants are trusting that partyâs word about what ran â and configurability granted by an operator is revocable by the same operator. The remedy is architectural: give each participant a replica of the items they care about and let them evaluate the rule themselves. Proposition 3.8 already established that this is possible, because (8) reads only an item and its out-neighbours. This section shows that DDP2P Silaghi et al., 2013; Silaghi et al., 2013a supplies almost everything else: what the platform provides, how the modelâs objects land on its item types, the identity problem that decentralisation makes harder, evaluation over an incomplete replica, and what the empirical findings of Sections 8â9 imply for a peer deployment specifically. What the platform provides. DDP2P is a project developing an open-source platform for decentralised deliberative petition drives, built around commitments that turn out to be the ones the charter needs. Items are self-contained and globally identified: every peer record, organisation, constituent, motion, justification, signature, vote, witnessing statement, news item and translation carries everything needed to interpret it and is named by an identifier derived from a public key with a creation date, or from a digest of its content Silaghi et al., 2013. Identifiers compose hierarchically â an organisationâs derives from its founding parameters, a motionâs from the organisation and its text, a justificationâs from the motion and its own text, a voteâs from the motion, the constituent and the justification cited â so two peers who have never communicated agree on the name of everything and semantically distinct items cannot collide. Synchronisation is pushâpull gossip with a horizon Silaghi et al., 2013a; Demers et al., 1988: a request carries the requesterâs interests and a time horizon and the answer carries items newer than that horizon, with no global index and no authoritative replica. Connectivity needs no owned infrastructure: directory servers help peers find each other and data servers hold items for offline peers, but neither is trusted, both are replaceable, and their content is verifiable against signatures; whether a peer relays for others is under that peerâs own control, by explicit design Alhamed & Silaghi, 2014, and mobile ad hoc and vehicular operation have been studied as extreme cases of the same idea Dhannoon et al., 2013; Dhannoon, 2013. Finally, organisations are rules rather than accounts: an organisation is a definition of a constituency and a jurisdiction, constituencies may be defined recursively with membership settled by a membership referendum â a fixpoint grassroots organisations resolve bottom-up rather than by administrative decree Silaghi et al., 2013 â motions are Robertâs motions (Section 4.7), and justifications are the arguments attached to signatures. Mapping the model onto the platform. Table 16 gives the correspondence. It is close to one-to-one, unsurprisingly since the model of Section 3 was distilled from this line of work Kattamuri et al., 2005; Silaghi & Roussev, 2014; Silaghi et al., 2017, but the residual mismatches are where the engineering lies, and two rows carry most of the weight. The reason labelling Îť is new, and if labels were assigned centrally the assigning party would hold semantic power over visibility â exactly the concentration C4 exists to prevent. The safe arrangement is that labels are declared by the justificationâs own author as part of the signed item, so a label is a claim like any other: checkable against the text by any reader and disputable through the ordinary argumentation mechanism. A mining pipeline (Section 2) may propose labels, but its proposals are advisory items, not ground truth; this leaves Îť adversarially controllable, which Section 12 takes up. The slate iS_i is likewise new, and need not leave the device at all â though a constituent who wants a public record of what they saw can publish it as a signed item, converting C5 from a platform guarantee into a personally held receipt, which is the strongest form the criterion can take. Table 16. Mapping the alternative-based poll of Definition 3.1 onto DDP2P item types. Entries marked new are what a deployment would have to add; everything else already exists. Model component DDP2P item Deployment note poll Î motion within an organisation direct; the organisation supplies the constituency and Ďc _c +,âJ^+,J^- justifications typed by signature polarity direct ââR claimed_refutes between justifications direct; the claim is signed by its author and never adjudicated ââR claimed_subsumes / includes direct; more_recent additionally orders revisions Ďc _c constituent record, organisation rules direct; membership class or shareholding EâĄ(j)E(j) votes citing j direct; every vote is a signed item naming the justification it cites Ďr _r â new: the policy of Section 4.5, computed locally from the count of distinct authors of the same link Îť â new: reason labels, author-declared or mined; used for audit and evaluation only, never inside ScoreAndServe() slate iS_i â new: computed locally; optionally published as a signed item so the participant holds a receipt θi _i local configuration new: per-peer policy vector, never transmitted Identity: the one thing decentralisation makes harder. Everything above assumes endorsement counts mean something, which assumes identities are not free. In a centralised deployment with an authenticated roll this is the registrarâs problem; open peer-to-peer membership makes it the systemâs problem, and it is the classical one Douceur, 2002. DDP2Pâs answer is a decentralised census with witnessing: constituents certify one anotherâs existence and eligibility, the certifications are signed items propagating like any other, and reputation over the witnessing graph estimates how much of the claimed population is real Qin et al., 2013a; Qin et al., 2013b; Qin et al., 2013; Qin et al., 2014, with a Bayesian extension of the web-of-trust idea estimating the number of distinct eligible signatories behind a set of signatures â directly the quantity a petition threshold depends on Silaghi et al., 2016. The structural point is that different observers may run different eligibility criteria over the same signed census data and reach their own conclusions: the platform supplies verifiable evidence, not a verdict. That is C7 at the level of the constituency rather than the slate. Section 6 assumes this problem solved, and the assumption is what makes the adversarial results interpretable â under authentication a coalition cannot inflate E and the link graph is the only surface left. It also means census integrity is not a side condition but the binding constraint: Section 9.7 located the residual risk in corpus capture, and what bounds corpus capture is what bounds coalition size. Evaluating the rule on an incomplete replica. A peer holds what it has synchronised, generally a subset of what exists. This is the sharpest technical objection to local evaluation and Proposition 3.8 only half answers it. Proposition 10.1 (Monotone degradation). Let â˛âĎJ ^Ď be the items a peer holds and â˛S the slate Algorithm 1 produces from them. Then â˛S is exactly the slate the same rule and policy would produce on the full pool restricted to â˛J . Missing items can only lower achievable completeness, never corrupt the score of a held item. Missing votes and links, however, lower the computed score of a held item, so the peerâs scores are lower bounds on the true ones. The second half is the whole difficulty. A peer holding a justification but not yet the votes citing it underestimates E; a peer missing links underestimates inherited standing. Three mitigations apply, all ordinary distributed-systems engineering. Votes and links are small items and are prioritised in the pull request over justification text, so counting evidence converges faster than the corpus. The score is a sum of non-negative terms, so partial evidence yields a lower bound and the interface can state the replicaâs coverage of the known item count â a peer can be told it holds 94%94\% of the votes and decide whether that is enough to act on. And the underlying structure is append-only with union merge, so replicas converge without conflict resolution in the manner of a grow-only set Shapiro et al., 2011, while a digest exchange in the style of Merkle hashing Merkle, 1988 makes divergence cheap to detect. Algorithm 4 states the resulting local cycle, including the reporting step that makes incompleteness visible. There is a real trade here: a centralised deployment computes the rule on complete evidence and asks you to trust the computation, whereas a decentralised one computes a verifiable rule on evidence that may be incomplete and tells you how incomplete it is. For a civic process the second is the better failure mode, because incompleteness is visible and declining trustworthiness is not. Algorithm 4 Local slate computation on a peer holding a partial replica Given: local store D; motion m; local policy θ; neighbour set N; horizon h Yields: a slate, plus a checkable statement of the evidence it rests on 1 foreach nân do 2 send a request naming m, this peerâs interests, and horizon h; 3 âDâ MergeSigned(D, signed items returned by n) ; // union of signed items; no conflict resolution needed 4 end foreach 5 discard items whose signature or identifier derivation fails to verify; 6 EâEâ tally of held votes per justification; 7 Ďrâ _râ policy of θ applied to held links; 8 âSâ ScoreAndServe(held items of m, θ, K); 9 report alongside S: counts of held votes, links and justifications, and the horizon h; 10 return S; What the findings imply for a peer deployment. Three results change their character when read against this substrate rather than a server. Ordering matters more, not less: Section 8.2 found the ruleâs advantage over an unordered draw concentrated in the first few positions, and on a peer holding an incomplete replica the effective slate is shorter still, so the fraction of the value delivered by the first five items rises. Degenerate authoring is more likely, not less: open membership lowers the cost of submitting, which is the point, and raises the share of submissions that fail to justify â exactly the regime where Section 8.3 finds the ruleâs margin largest and Section 8.4 finds the link summands supplying most of it, so the two terms a server deployment might reasonably drop as inert are the terms a peer deployment most needs. The weight policy must be local: Section 9.2 makes Ďr _r a security control worth ten points of completeness, and on a server the operator picks it for everyone whereas on a peer each participant picks it and can compute what the other choice would have given them. Since some of these choices are frontier positions with no correct answer (Section 8.5), configurability stops being a feature and becomes the only coherent way to hold a parameter that is simultaneously a security control and a value judgement. 11. Discussion The results admit a compact summary: the charter is cheap, the instrument matters more than the mechanism, the security lives in one line of policy, and the remaining choices are not the systemâs to make. Each cuts against a default assumption about civic recommenders â that transparency costs accuracy, that one quality metric suffices, that robustness comes from detection, and that a well-designed system should decide. The charter is cheap, and the cheapness is measurable. Section 7.3 put the price of semantic abstinence at 0.031Âą0.0140.031Âą 0.014 against a ceiling that upper-bounds every procedure applied to the same pool, admissible or not, and four fifths of the observed incompleteness was vocabulary that did not yet exist, which no procedure recovers. The entire competitive advantage available to an unconstrained learned ranker is therefore three points of completeness, purchased by reading what C4 forbids â against 0.100.10 for raising the authoring rate from 0.020.02 to 0.100.10 and 0.140.14 for doubling the slate. The usual argument for opacity is that legibility costs quality and the cost is unknown; here it is known, bounded, and small relative to the levers a deployment actually controls. The instrument mattered more than the mechanism. The most transferable lesson is methodological. We evaluated a ranking rule with a set functional and concluded, across 40004000 seed-paired runs, that it was indistinguishable from a random subset â an artefact of the measurement that two remarks derivable from the definition (Remarks 1â2) predicted in advance. Coverage-style aggregates are the default in the diversity-aware recommender literature Adomavicius & Kwon, 2012; Carbonell & Goldstein, 1998, and any evaluation of a civic selector reporting only such an aggregate is exposed to the same null. Robustness came from policy, not from detection. Nothing in Section 9 trains a classifier, scores trustworthiness, or removes material. The effective defence is (9): count distinct authors, divide by voters on that side. It is one line, publishable, verifiable by any participant, and worth 0.1510.151â0.1810.181 of completeness against a coalition holding a fifth of the electorate; the available escalation makes the coalition monotonically weaker. Two features generalise. The defence works by making the attack expensive and visible â a co-signing coalition writes its coordination into the public link record â and it is compatible with author blindness, since (9) counts how many distinct people asserted a link and never which, so robustness did not require reintroducing the reputation hierarchy C3 exists to exclude. That these are compatible was not obvious in advance. Section 9.7 also relocated rather than removed the residual risk: under authentication what survives is corpus capture, which belongs to census integrity and slate capacity rather than to ranking. Section 8.5 found that at G=0.5G=0.5 the link summands buy 2.82.8 points of coverage for 2.42.4 points of endorsement mass. That is not a result with a right answer but a frontier position, and choosing among such positions is choosing between breadth of reasons represented and fidelity to what constituents took up. We take this as the strongest available argument for C7, and it arrived from an unexpected direction: the case for participant-held parameters is usually made on autonomy grounds, whereas here it is forced by the measurements. A platform that picked a point on that frontier and presented it as neutral would be making a political choice while denying it. Four consequences follow for deployment. Invest in authoring, not in ranking: the dominant lever is how quickly the reason space matures, and the selection component of the shortfall is a fifth the size of the temporal one. Publish the weight policy and treat it as a security control: it is the one parameter with an order-of-magnitude effect under attack, and a deployment shipping a flat default has left its main defence unarmed without the participant being able to tell. Regulatory alignment is already close: the Digital Services Act requires very large platforms to disclose recommender parameters and offer a non-profiling option European Parliament and Council, 2022, and the AI Act imposes transparency and human-oversight duties on systems used in democratic processes European Parliament and Council, 2024. A published rule over public evidence with participant-held parameters satisfies the letter of both without a compliance layer, because there is nothing to disclose that is not already disclosed. 12. Limitations The following would change our conclusions, and we list them in descending order of how much. Several are stated more sharply than a reader would infer from the results alone, because we would rather over-declare than have a finding survive on an unexamined assumption. The electorate is simulated. Every number here describes agents, not people. Adoption is TFâIDF similarity to a position text; opinions do not update; nobody gets bored, persuaded, or strategic beyond the modelled coalitions. Following Section 2 we restrict claims to statements about mechanism under a stated model. Nothing here licenses a claim about the magnitude any quantity would take in a human electorate. The mechanism result rests on one modelled fact. The degenerate-authoring findings of Sections 8.3 and 8.4 descend from the observation that degenerate items are adopted less â 1.451.45 endorsements against 10.9410.94 â and that follows from the simulatorâs cosine-similarity adoption step, a property of the agents rather than of the rule. The finding is that the rule inherits and amplifies whatever discrimination the electorate itself exercises. If a human electorate endorsed reason-free submissions at the same rate as reasoned ones, the link summands would revert to the inertness of Table 6, and the coverage margin would go with them. We regard establishing that adoption rate empirically as the single most valuable follow-up in this programme. One round. Constituents vote once. Multi-round deliberation with opinion revision would change adoption dynamics, the endorsement distribution, and probably the link structure. Whether the ruleâs ordering advantage survives revision is untested. This was implemented but not thoroughly evaluated. Statistical multiplicity. Section 6.5 declares that we do not correct across the manuscript. The ordering contrasts of Section 8.2 number twenty-four; the two link-term effects at pâ10â2pâ 10^-2 are reported as suggestive and should not be built on. The effects we do build on sit at p<10â10p<10^-10 and would survive any reasonable correction. Sybil resistance is assumed, not demonstrated. Under a broken census every result in Section 9 fails, since a coalition that mints identities mints endorsements. DDP2Pâs witnessed-census line Qin et al., 2013a; Qin et al., 2014; Silaghi et al., 2016 is the intended answer and we have not evaluated it here. Section 10 states the dependency; this is the largest gap between the manuscript and a deployment. Partial replicas are analysed, not measured. Proposition 10.1 bounds the degradation and Algorithm 4 declares the evidence, but we ran no experiment with peers holding genuinely divergent replicas. The prediction of Section 10 â that ordering matters more under partial replication â is untested. Fairness across minority positions is unexamined. Endorsement mass rewards what constituents took up. Whether (8) systematically disadvantages minority positions, whose reasons appear in fewer items and therefore accumulate less endorsement, is an open and important question that the weighted variant of Section 3.4 anticipates syntactically without answering. 13. Work, Meaning, and the E-Citizen A manuscript about the mechanics of argument selection owes its readers an account of why the mechanics matter, and the answer is not confined to the integrity of any particular poll. The argument has four steps: procedure is constitutive of collective agency rather than decorative; the historical obstacle to genuine self-government has been the time it consumes; artificial intelligence, by absorbing productive labour, removes that obstacle in a way no previous technology has; and the role of citizen is therefore available as a destination for human effort in a way it has not been for two and a half millennia â provided the instruments of that role remain legible to the people exercising it. The final clause is where this section rejoins the rest of the manuscript. That Robertâs rules of order were the salt which turned a mob into a society Silaghi, 2025; Robert, 1915 is a claim about what makes collective reasoning possible at all. A crowd has volume; an assembly has a procedure for converting volume into a decision, and the procedure is what everyone can agree to precisely because it constrains form rather than content. One can accept a rule about who speaks next without accepting anything about what they will say â which is why procedural agreements survive substantive disagreements. The charter of Section 4.2 carries that idea into the selection step, the one place where digital deliberation has so far had no procedure whatsoever: we regulate speaking time in a parliament to the minute, and we let an unpublished model decide which of a hundred thousand submitted reasons a voter sees. Semantic abstinence is the descendant of content-indifference in the chair; configurability descends from the assemblyâs authority over its own rules; contestability descends from the appeal from the chair. The novelty is not the principle but that it must now be enforced in software, because that is where the procedure has migrated. Athenian democracy worked, to the extent it did, because a body of citizens had time to do it. Assembly attendance, jury service and rotation through office consume days, not minutes, and are impossible for people whose waking hours are claimed by subsistence. Athens solved the time problem with slavery: the enfranchised were free to deliberate because the disenfranchised did the work Narcisse, 2012. Every subsequent expansion of the franchise inherited the structural problem without the solution. Representative democracy is, read uncharitably, a device for economising on citizensâ time â elect somebody to be the citizen for you, on the grounds that you have a job â and it produces the characteristic modern experience of participation as a thin, performative gesture. That thinness has been described as a kitsch of the political form Silaghi, 2025, borrowing a diagnosis developed for aesthetics Calinescu, 1987; Lazare, 1999: the surface features of the real thing, arranged for easy consumption, with the deliberation removed. The literature on why participation platforms fail Toots, 2019; Bright & Margetts, 2016 is largely a catalogue of the same thinness, and the finding that direct-democracy processes educate the citizens who use them Smith & Tolbert, 2004; Tolbert et al., 2009 is the other side of the coin. There is a direct line from that diagnosis to the exposure problem of Section 1.1: a slate assembled by an unpublished model, shown to a voter who cannot recompute it, is the kitsch form of deliberative exposure â the appearance of having met the arguments, with the part that would make the meeting real removed. The anxious question about artificial intelligence and work is whether it will take our jobs, and the productivity framing of the fourth industrial revolution Schwab, 2017 does not settle it. But the anxiety contains an assumption worth naming: that the only worthwhile use of human time is producing goods and services. Set that aside and the picture inverts. Self-government has always been constrained by a shortage of citizen-hours, and a technology that discharges productive labour is by construction a technology that produces them. What Athens obtained by enslaving people, an automated economy could obtain without enslaving anyone â and where the Athenian arrangement was a moral defect of that society, the same structural position occupied by machines is available to be read as a strength of ours Silaghi, 2025. This is not a prediction: time released from labour goes wherever the surrounding institutions send it. It is a claim about what becomes possible â that for the first time since a slave economy made it possible for a few, citizenship as a substantial, time-consuming, skilled activity becomes available to many. An epistemic argument runs alongside the ethical one. The case for inclusive deliberation over rule by the competent is that cognitive diversity does work no amount of individual expertise substitutes for Landemore, 2012; Landemore, 2013; that argument is only cashable if the varied body is actually reasoning, which costs time, so the epistemic case for inclusion and the material case for citizen-hours are the same case seen from two sides. It also explains why the quantity measured here is coverage of the reason vocabulary rather than agreement: the value of a large deliberating public lies in the reasons it collectively holds, and a selection step that loses them destroys precisely what made inclusion worth having. Our own measurements make the argument in miniature: Section 7.2 found that persuading one constituent in ten to write a reason instead of one in fifty is worth as much completeness as multiplying the electorate fivefold. The binding constraint is not how many people vote but how many think in public, and that is a quantity measured in hours. The argument has an obvious failure mode, and it is the one this manuscript is built to forestall. If the same technology that releases the time also assembles the arguments, filters the objections and decides which considerations a citizen encounters, then the citizen-hours have been created and simultaneously hollowed out: one would have the leisure of the Athenian and the informational position of a spectator, and the role would be available and not worth occupying. This is why C4 and C7 are the two criteria that matter most, and why Section 4.3 declines to trade them for coverage. A citizen who can recompute why they were shown what they were shown is exercising judgement; one who cannot is receiving a service. The difference does not show up in any completeness measure â Section 7.3 put its entire measurable cost at 0.0310.031 â and it is the whole difference between the two futures. It also sets the correct place for a language model in a civic system, which is not nowhere: Section 2 places extraction at authoring time, visible to the author and contestable before serving, so a model may help a person say what they mean. What it must not do is decide, at serving time and unaccountably, which of those meanings anyone meets. The line is exactly the one between assisting a citizen and replacing one. 14. Conclusions We set out to show that the selection step in a deliberative poll can be democratic procedure rather than infrastructure: a published rule over public evidence, with parameters held by the people it affects. The construction is an alternative-based poll over bipolar justification sets, judged by three instruments â coverage of the live reason vocabulary, the order in which that coverage arrives, and the endorsement mass captured â and served by a one-hop reversed endorsement flow whose only policy lever is a relation-weight function. Seven criteria state what makes a mechanism admissible at all, and the rule satisfies all seven. This section states what was established, what was not, and what comes next. What the measurements established, across roughly 17,00017,000 seeded runs: the levers that govern coverage are the ones that mature the corpus, not the ones that select from it; the price of semantic abstinence is 0.031Âą0.0140.031Âą 0.014 against a ceiling that bounds every mechanism including inadmissible ones, with four fifths of the residual shortfall temporal rather than algorithmic; set coverage on non-degenerate authoring alone cannot distinguish the rule from a random draw, and the reasons are properties of the instrument that we state as Remarks 1â2; on an order-sensitive reading the rule leads at every prefix short of the full slate, by a margin that widens under attack and reaches â8.3-8.3 positions of e90e_90 at a quarter-electorate coalition; once a realistic fraction of submissions fails to justify, the coverage margin returns and grows monotonically, with the link summands supplying 93%93\% of it at G=0.7G=0.7 by squaring the electorateâs own discrimination; on endorsement mass the random baseline loses by a factor of 3.33.3 and the greedy ceiling stops dominating under attack; hub-riding is inert; label flooding is severe and the weight policy is worth 0.1510.151â0.1810.181 against it; four fifths of that damage is homogeneity rather than volume; and a co-signing coalition makes itself monotonically weaker. What was not established is listed in Section 12 and led by two items: that a human electorate discriminates against unjustified submissions at anything like the modelled rate, on which the mechanism result depends; and that census integrity holds, on which every adversarial result depends. The claim we would defend most firmly is not that this rule ranks well. It is that the question âwhy was I shown this?â must have an answer that is a table of numbers rather than a narrative, and that a system built to that constraint turns out to cost about three points of coverage, to be more robust than we expected, and to be far easier to understand when it is wrong. A note on scope and provenance This manuscript is an independently written, extended treatment of a line of work by the same labs. It is not the camera-ready version of any conference paper and it reproduces no text, figure, table, algorithm or example from one. Every definition, proposition, algorithm, worked example, figure and table here was composed for this document; the notation, the three-instrument framing, the seven criteria and their tests, the worked example of Section 3.7, and all diagrams are original to it. Where the underlying research programme has been reported elsewhere, the overlap is one of subject matter and of numerical results computed from the same experimental runs, not of expression. Results attributed to prior work are cited as such. The extensions developed here include the ordering and endorsement-mass instruments, the degenerate-authoring model and its mechanism analysis, the coverage-versus-mass frontier, the elaborated criteria and their audit tests, and the side balance evaluation. Tools, data and reproducibility All simulations were run with a seeded pseudorandom generator; every run records its seed, its full parameter set, and every served slate in served order together with the policy vector in force (Section 5.5). The run databases and the analysis scripts that produce every table and figure are made available at https://github.com/devfitcs/ABAS. Language-model assistance was used in drafting and editing prose; all technical content, experimental design, analysis and conclusions are the authorâs. References Adomavicius & Kwon (2012) Gediminas Adomavicius and YoungOk Kwon âImproving Aggregate Recommendation Diversity Using Ranking-Based Techniquesâ In IEEE Transactions on Knowledge and Data Engineering 24.5 IEEE, 2012, p. 896â911 Alcântara & Cordeiro (2025) JoĂŁo Alcântara and Renan Cordeiro âOn the Equivalence between Logic Programs and Bipolar Argumentation Frameworksâ In Journal of Artificial Intelligence Research 84, 2025 DOI: 10.1613/jair.1.18086 Alhamed & Silaghi (2014) Khalid Alhamed and Marius. Silaghi âUser Freedom: To Be or Not to Be a âSupernodeââ In Proceedings of the 14th IEEE International Conference on Peer-to-Peer Computing (P2P 2014) London, UK: IEEE, 2014, p. 1â5 Alhamed et al. (2013) Khalid Alhamed et al. âStacking the Deck Attack on Software Updates: Solution by Distributed Recommendation of Testersâ In Proceedings of the IEEE/WIC/ACM International Conference on Intelligent Agent Technology (IAT 2013) Atlanta, GA: IEEE, 2013, p. 293â300 Alhamed et al. (2013a) Khalid Alhamed, Marius. Silaghi, Ihsan Hussien and Yi Yang âSecurity by Decentralized Certification of Automatic Updates for Open Source Software Controlled by Volunteersâ In Proceedings of the International Workshop on Decentralized Coordination (DC 2013), 2013 Alhamed et al. (2016) Khalid Alhamed, Markus Zanker, Shakre Elmane and Marius. Silaghi âP2P Meta-Recommenders: Aggregated Diversity Maximization as a Bulwark against Attacks on Reviewersâ In Proceedings of the IEEE/WIC/ACM International Conference on Web Intelligence (WI 2016) Omaha, NE: IEEE, 2016, p. 208â215 Alqahtani & Silaghi (2016) Abdulrahman Alqahtani and Marius. Silaghi âEvaluation Technique for Argumentation Architectures from the Perspective of Human Cognitionâ In Proceedings of the Twenty-Ninth International Florida Artificial Intelligence Research Society Conference (FLAIRS-29), Poster Abstracts AAAI Press, 2016 Alqahtani & Silaghi (2017) Abdulrahman Alqahtani and Marius. Silaghi âHuman-Computer Interaction in a Debate Decision Support Systemâ In Proceedings of the Thirtieth International Florida Artificial Intelligence Research Society Conference (FLAIRS-30) AAAI Press, 2017, p. 773 Amgoud et al. (2008) Leila Amgoud, Claudette Cayrol, Marie-Christine Lagasquie-Schiex and Pierre Livet âOn Bipolarity in Argumentation Frameworksâ In International Journal of Intelligent Systems 23.10 Wiley, 2008, p. 1062â1093 Argyle et al. (2023) Lisa. Argyle et al. âOut of One, Many: Using Language Models to Simulate Human Samplesâ In Political Analysis 31.3 Cambridge University Press, 2023, p. 337â351 Bail (2024) Christopher. Bail âCan Generative AI Improve Social Science?â In Proceedings of the National Academy of Sciences 121.21, 2024, p. e2314021121 (1) âHandbook of Formal Argumentation, Volume 1â London: College Publications, 2018 Behrendt et al. (2024) Maike Behrendt et al. âAQuA â Combining Expertsâ and Non-Expertsâ Views to Assess Deliberation Quality in Online Discussions Using LLMsâ In Proceedings of the First Workshop on Language-Driven Deliberation Technology (DELITE) ELRAICCL, 2024, p. 1â12 Bench-Capon & Dunne (2007) Trevor.. Bench-Capon and Paul. Dunne âArgumentation in Artificial Intelligenceâ In Artificial Intelligence 171.10â15 Elsevier, 2007, p. 619â641 Boutet et al. (2013) Antoine Boutet et al. âWhatsUp: A Decentralized Instant News Recommenderâ In Proceedings of the 27th IEEE International Symposium on Parallel and Distributed Processing (IPDPS 2013) IEEE, 2013, p. 741â752 Brenneis et al. (2021) Markus Brenneis, Maike Behrendt and Stefan Harmeling âHow Will I Argue? A Dataset for Evaluating Recommender Systems for Argumentationsâ In Proceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL 2021) Association for Computational Linguistics, 2021, p. 360â367 Bright & Margetts (2016) Jonathan Bright and Helen Margetts âBig Data and Public Policy: Can It Succeed Where E-Participation Has Failed?â In Policy & Internet 8.3 Wiley, 2016, p. 218â224 Bulteau (2021) Laurent Bulteau âAggregation over Metric Spaces: Proposing and Voting in Elections, Budgeting, and Legislationâ In Journal of Artificial Intelligence Research 70, 2021, p. 1413â1439 DOI: 10.1613/jair.1.12388 Burkart & Huber (2021) Nadia Burkart and Marco. Huber âA Survey on the Explainability of Supervised Machine Learningâ In Journal of Artificial Intelligence Research 70, 2021, p. 245â317 DOI: 10.1613/jair.1.12228 Burke (2017) Robin Burke âMultisided Fairness for Recommendationâ Presented at the Workshop on Fairness, Accountability and Transparency in Machine Learning (FAT/ML) In arXiv preprint arXiv:1707.00093, 2017 Burrell (2016) Jenna Burrell âHow the Machine âThinksâ: Understanding Opacity in Machine Learning Algorithmsâ In Big Data & Society 3.1 SAGE Publications, 2016, p. 1â12 Calinescu (1987) Matei Calinescu âFive Faces of Modernity: Modernism, Avant-Garde, Decadence, Kitsch, Postmodernismâ Durham, NC: Duke University Press, 1987 Carbonell & Goldstein (1998) Jaime Carbonell and Jade Goldstein âThe Use of MMR, Diversity-Based Reranking for Reordering Documents and Producing Summariesâ In Proceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval ACM, 1998, p. 335â336 Carenini & Moore (2006) Giuseppe Carenini and Johanna. Moore âGenerating and Evaluating Evaluative Argumentsâ In Artificial Intelligence 170.11 Elsevier, 2006, p. 925â952 Caro-MartĂnez et al. (2021) Marta Caro-MartĂnez, Guillermo JimĂŠnez-DĂaz and Juan. Recio-GarcĂa âConceptual Modeling of Explainable Recommender Systems: An Ontological Formalization to Guide Their Design and Developmentâ In Journal of Artificial Intelligence Research 71, 2021, p. 557â589 DOI: 10.1613/jair.1.12789 Castiglioni (2021) Matteo Castiglioni âElection Manipulation on Social Networks: Seeding, Edge Removal, Edge Additionâ In Journal of Artificial Intelligence Research 71, 2021, p. 1049â1090 DOI: 10.1613/jair.1.12826 Cayrol & Lagasquie-Schiex (2013) Claudette Cayrol and Marie-Christine Lagasquie-Schiex âBipolarity in Argumentation Graphs: Towards a Better Understandingâ In International Journal of Approximate Reasoning 54.7 Elsevier, 2013, p. 876â899 Cayrol & Lagasquie-Schiex (2005) Claudette Cayrol and Marie-Christine Lagasquie-Schiex âOn the Acceptability of Arguments in Bipolar Argumentation Frameworksâ In Symbolic and Quantitative Approaches to Reasoning with Uncertainty (ECSQARU 2005) 3571, Lecture Notes in Computer Science Berlin, Heidelberg: Springer, 2005, p. 378â389 Chakraborty et al. (2019) Abhijnan Chakraborty, Saptarshi Ghosh, Niloy Ganguly and Krishna. Gummadi âOptimizing the Recency-Relevance-Diversity Trade-Offs in Non-Personalized News Recommendationsâ In Information Retrieval Journal 22.5 Springer, 2019, p. 447â475 Chalaguine & Hunter (2020) Lisa. Chalaguine and Anthony Hunter âA Persuasive Chatbot Using a Crowd-Sourced Argument Graph and Concernsâ In Computational Models of Argument (COMMA 2020) 326, Frontiers in Artificial Intelligence and Applications IOS Press, 2020, p. 9â20 Chandak et al. (2026) Nikhil Chandak, Shashwat Goel and Dominik Peters âProportional Aggregation of Preferences for Sequential Decision Makingâ In Journal of Artificial Intelligence Research 85, 2026 DOI: 10.1613/jair.1.18660 Cheng et al. (2021) Lu Cheng, Kush. Varshney and Huan Liu âSocially Responsible AI Algorithms: Issues, Purposes, and Challengesâ In Journal of Artificial Intelligence Research 71, 2021, p. 1137â1181 DOI: 10.1613/jair.1.12814 Demers et al. (1988) Alan Demers et al. âEpidemic Algorithms for Replicated Database Maintenanceâ In ACM SIGOPS Operating Systems Review 22.1 ACM, 1988, p. 8â32 Dhannoon et al. (2013) Osamah Dhannoon, Rahul Vishen and Marius. Silaghi âContent Dissemination over VANET: Boosting Utility-Based Heuristics Using Interestsâ In Proceedings of the International Conference on Connected Vehicles and Expo (ICCVE 2013) IEEE, 2013, p. 106â113 Dhannoon (2013) Osamah Dhannoon âMultiplexing Content Exchange for Petition Drives in VANETs: Receiver Interest and Sender Utilityâ, 2013 Douceur (2002) John. Douceur âThe Sybil Attackâ In Peer-to-Peer Systems (IPTPS 2002) 2429, Lecture Notes in Computer Science Springer, 2002, p. 251â260 Dung (1995) Phan Dung âOn the Acceptability of Arguments and Its Fundamental Role in Nonmonotonic Reasoning, Logic Programming and n-Person Gamesâ In Artificial Intelligence 77.2 Elsevier, 1995, p. 321â357 Endriss (2020) Ulle Endriss âThe Complexity Landscape of Outcome Determination in Judgment Aggregationâ In Journal of Artificial Intelligence Research 69, 2020, p. 687â731 DOI: 10.1613/jair.1.11970 European Parliament and Council (2022) European Parliament and Council âRegulation (EU) 2022/2065 on a Single Market for Digital Services (Digital Services Act)â, Official Journal of the European Union, L 277, 27 October 2022, 2022 European Parliament and Council (2024) European Parliament and Council âRegulation (EU) 2024/1689 Laying Down Harmonised Rules on Artificial Intelligence (Artificial Intelligence Act)â, Official Journal of the European Union, L series, 12 July 2024, 2024 Feige (1998) Uriel Feige âA Threshold of lnâĄn n for Approximating Set Coverâ In Journal of the ACM 45.4 ACM, 1998, p. 634â652 Feng et al. (2023) Shangbin Feng, Chan Park, Yuhan Liu and Yulia Tsvetkov âFrom Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Modelsâ In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL 2023) Association for Computational Linguistics, 2023, p. 11737â11762 Fishkin (1991) James. Fishkin âDemocracy and Deliberation: New Directions for Democratic Reformâ New Haven, CT: Yale University Press, 1991 Fishkin (2009) James. Fishkin âWhen the People Speak: Deliberative Democracy and Public Consultationâ Oxford, UK: Oxford University Press, 2009 Fishkin et al. (2000) James. Fishkin, Robert. Luskin and Roger Jowell âDeliberative Polling and Public Consultationâ In Parliamentary Affairs 53.4 Oxford University Press, 2000, p. 657â666 Gonzalez (2021) Melisa.Ăąuela Gonzalez âLabeled Bipolar Argumentation Frameworksâ In Journal of Artificial Intelligence Research 70, 2021, p. 1557â1636 DOI: 10.1613/jair.1.12394 Habermas (1984) JĂźrgen Habermas âThe Theory of Communicative Action, Volume 1: Reason and the Rationalization of Societyâ Translated by Thomas McCarthy Boston, MA: Beacon Press, 1984 Hochbaum (1997) Dorit. Hochbaum âApproximating Covering and Packing Problems: Set Cover, Vertex Cover, Independent Set, and Related Problemsâ In Approximation Algorithms for NP-Hard Problems Boston, MA: PWS Publishing, 1997, p. 94â143 Kahng et al. (2021) Anson Kahng, Simon Mackenzie and Ariel Procaccia âLiquid Democracy: An Algorithmic Perspectiveâ In Journal of Artificial Intelligence Research 70, 2021, p. 1223â1252 DOI: 10.1613/jair.1.12261 Karp (1972) Richard. Karp âReducibility among Combinatorial Problemsâ In Complexity of Computer Computations New York, NY: Plenum Press, 1972, p. 85â103 Kattamuri et al. (2005) Kiran Kattamuri et al. âSupporting Debates over Citizen Initiativesâ In Proceedings of the 2005 National Conference on Digital Government Research (dg.o 2005) Digital Government Society of North America, 2005, p. 279â280 Lam & Riedl (2004) Shyong. Lam and John Riedl âShilling Recommender Systems for Fun and Profitâ In Proceedings of the 13th International Conference on World Wide Web (W 2004) ACM, 2004, p. 393â402 Landemore (2013) HĂŠlène Landemore âDeliberation, Cognitive Diversity, and Democratic Inclusiveness: An Epistemic Argument for the Random Selection of Representativesâ In Synthese 190.7 Springer, 2013, p. 1209â1231 Landemore (2012) HĂŠlène Landemore âDemocratic Reason: Politics, Collective Intelligence, and the Rule of the Manyâ Princeton, NJ: Princeton University Press, 2012 Lazare (1999) Daniel Lazare âModernism as Kitsch: Hilton Kramerâs Thirty-Year Culture Warâ In The Baffler 13 MIT Press, 1999, p. 27â33 Lippi & Torroni (2016) Marco Lippi and Paolo Torroni âArgumentation Mining: State of the Art and Emerging Trendsâ In ACM Transactions on Internet Technology 16.2 ACM, 2016, p. 1â25 Lipton (2018) Zachary. Lipton âThe Mythos of Model Interpretabilityâ In Communications of the ACM 61.10 ACM, 2018, p. 36â43 Liscio (2025) Enrico Liscio âValue Preferences Estimation and Disambiguation in Hybrid Participatory Systemsâ In Journal of Artificial Intelligence Research 82, 2025, p. 819â850 DOI: 10.1613/jair.1.14958 Lukensmeyer & Brigham (2005) Carolyn. Lukensmeyer and Steve Brigham âTaking Democracy to Scale: Large Scale InterventionsâFor Citizensâ In The Journal of Applied Behavioral Science 41.1 SAGE Publications, 2005, p. 47â60 Luskin et al. (2002) Robert. Luskin, James. Fishkin and Roger Jowell âConsidered Opinions: Deliberative Polling in Britainâ In British Journal of Political Science 32.3 Cambridge University Press, 2002, p. 455â487 Mancini (2015) Pia Mancini âWhy It Is Time to Redesign Our Political Systemâ In European View 14.1 SAGE Publications, 2015, p. 69â75 Manning et al. (2008) Christopher. Manning, Prabhakar Raghavan and Hinrich SchĂźtze âIntroduction to Information Retrievalâ Cambridge, UK: Cambridge University Press, 2008 Maymounkov & Mazières (2002) Petar Maymounkov and David Mazières âKademlia: A Peer-to-Peer Information System Based on the XOR Metricâ In Peer-to-Peer Systems (IPTPS 2002) 2429, Lecture Notes in Computer Science Springer, 2002, p. 53â65 Meir et al. (2021) Reshef Meir, Fedor Sandomirskiy and Moshe Tennenholtz âRepresentative Committees of Peersâ In Journal of Artificial Intelligence Research 71, 2021, p. 401â429 DOI: 10.1613/jair.1.12521 Merkle (1988) Ralph. Merkle âA Digital Signature Based on a Conventional Encryption Functionâ In Advances in Cryptology â CRYPTO â87 293, Lecture Notes in Computer Science Springer, 1988, p. 369â378 Mill (1859) John Mill âOn Libertyâ Project Gutenberg eBook no. 34901, released January 10, 2011 LondonFelling-on-Tyne; New YorkMelbourne: Walter Scott Publishing Co., Ltd., 1859 URL: https://w.gutenberg.org/files/34901/34901-h/34901-h.htm Miller (2020) Greg Miller âThe Intelligence Coup of the Centuryâ, The Washington Post, 11 February 2020. https://w.washingtonpost.com/graphics/2020/world/national-security/cia-crypto-encryption-machines-espionage/, 2020 Nakamoto (2008) Satoshi Nakamoto âBitcoin: A Peer-to-Peer Electronic Cash Systemâ White paper, https://bitcoin.org/bitcoin.pdf, 2008 Narcisse (2012) Tiky Narcisse âThe African Origins of the Athenian Democracyâ In Proceedings of the 43rd National Conference of Black Political Scientists (NCOBPS), 2012 Nemhauser et al. (1978) George. Nemhauser, Laurence. Wolsey and Marshall. Fisher âAn Analysis of Approximations for Maximizing Submodular Set FunctionsâIâ In Mathematical Programming 14.1 Springer, 1978, p. 265â294 Nikitin et al. (2017) Kirill Nikitin et al. âCHAINIAC: Proactive Software-Update Transparency via Collectively Signed Skipchains and Verified Buildsâ In Proceedings of the 26th USENIX Security Symposium USENIX Association, 2017, p. 1271â1287 Page et al. (1999) Lawrence Page, Sergey Brin, Rajeev Motwani and Terry Winograd âThe PageRank Citation Ranking: Bringing Order to the Webâ In Stanford InfoLab Technical Report 1999-66, 1999 Pariser (2011) Eli Pariser âThe Filter Bubble: What the Internet Is Hiding from Youâ New York, NY: Penguin Press, 2011 Park et al. (2023) Joon Park et al. âGenerative Agents: Interactive Simulacra of Human Behaviorâ In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST 2023) ACM, 2023, p. 1â22 Prakken (2001) Henry Prakken âFormalizing Robertâs Rules of Order: An Experiment in Automating Mediation of Group Decision Makingâ, 2001 Price et al. (2002) Vincent Price, Joseph. Cappella and Lilach Nir âDoes Disagreement Contribute to More Deliberative Opinion?â In Political Communication 19.1 Taylor & Francis, 2002, p. 95â112 Qin et al. (2014) Song Qin et al. âOpen Census for Addressing False Identity Attacks in Agent-Based Decentralized Social Networksâ In Proceedings of the 13th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2014) Paris, France: IFAAMAS, 2014, p. 1591â1592 Qin et al. (2013) Song Qin et al. âAddressing False Identity Attacks in Action-Based P2P Social Networks with an Open Censusâ In Proceedings of the IEEE/WIC/ACM International Conference on Web Intelligence (WI 2013) Atlanta, GA: IEEE, 2013, p. 50â57 Qin et al. (2013a) Song Qin et al. âP2P Decentralized Population Censusâ In Proceedings of the International Workshop on Decentralized Coordination (DC 2013), 2013 Qin et al. (2013b) Song Qin et al. âReputation System for Decentralized Population Censusâ In Proceedings of the IJCAI Workshop on Incentives and Trust in E-Commerce (WIT-EC 2013), 2013, p. 37â48 Robert (1915) Henry. Robert âRobertâs Rules of Order Revised for Deliberative Assembliesâ Chicago, IL: Scott, ForesmanCompany, 1915 Rocha et al. (2026) Victor Rocha, Fabio Cozman and Serena Villata âAssessing the Minimal Dialectical Quality in Argumentation: A Neuro-Symbolic Approach Integrating Argument Mining, Quality Assessment, and Probabilistic Reasoningâ In Journal of Artificial Intelligence Research 86, 2026 DOI: 10.1613/jair.1.21868 Roussev & Silaghi (2017) Roussi Roussev and Marius. Silaghi âA Logic for Making Hard Decisionsâ In Proceedings of the Thirtieth International Florida Artificial Intelligence Research Society Conference (FLAIRS-30) AAAI Press, 2017, p. 712â716 Rudin (2019) Cynthia Rudin âStop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Insteadâ In Nature Machine Intelligence 1.5 Nature Publishing Group, 2019, p. 206â215 Schwab (2017) Klaus Schwab âThe Fourth Industrial Revolutionâ New York, NY: Currency, 2017 Shapiro et al. (2011) Marc Shapiro, Nuno Preguiça, Carlos Baquero and Marek Zawirski âConflict-Free Replicated Data Typesâ In Stabilization, Safety, and Security of Distributed Systems (S 2011) 6976, Lecture Notes in Computer Science Springer, 2011, p. 386â400 Shardanand & Maes (1995) Upendra Shardanand and Pattie Maes âSocial Information Filtering: Algorithms for Automating âWord of Mouthââ In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI â95) ACM, 1995, p. 210â217 Sharma et al. (2023) Mrinank Sharma et al. âTowards Understanding Sycophancy in Language Modelsâ In arXiv preprint arXiv:2310.13548, 2023 Silaghi (2025) Marius. Silaghi âRepresentative Democracy as Kitsch, and Artificial Intelligenceâs Promise of Emancipationâ In Global Modernity from Coloniality to Pandemic: A Cross-Disciplinary Perspective London: Routledge, 2025, p. 349â370 Silaghi et al. (2013) Marius. Silaghi et al. âDirectDemocracyP2PâDecentralized Deliberative Petition Drivesâ In Proceedings of the 13th IEEE International Conference on Peer-to-Peer Computing (P2P 2013) Trento, Italy: IEEE, 2013, p. 1â2 Silaghi et al. (2013a) Marius. Silaghi et al. âP2P Petition Drives and Deliberation of Shareholdersâ In Proceedings of the International Workshop on Decentralized Coordination (DC 2013), 2013 Silaghi et al. (2016) Marius. Silaghi et al. âBayesian Network-Based Extension for PGPâEstimating Petition Supportâ In Proceedings of the Twenty-Ninth International Florida Artificial Intelligence Research Society Conference (FLAIRS-29) AAAI Press, 2016 Silaghi & Roussev (2014) Marius. Silaghi and Roussi Roussev âRecommending the Most Encompassing Opposing and Endorsing Arguments in Debatesâ https://arxiv.org/abs/1411.5416, 2014 Silaghi et al. (2017) Marius-Calin Silaghi, Roussi Roussev and Badria Alfurhood âWhy Do They Vote That?â In Proceedings of the Thirtieth International Florida Artificial Intelligence Research Society Conference (FLAIRS-30) Marco Island, FL: AAAI Press, 2017, p. 128â133 Singh et al. (2006) Atul Singh, Tsuen-Wan Ngan, Peter Druschel and Dan. Wallach âEclipse Attacks on Overlay Networks: Threats and Defensesâ In Proceedings of the 25th IEEE International Conference on Computer Communications (INFOCOM 2006) IEEE, 2006, p. 1â12 Small et al. (2021) Christopher Small et al. âPolis: Scaling Deliberation by Mapping High Dimensional Opinion Spacesâ In Proceedings of the AAAI Symposium on Computational Approaches to Online Collective Deliberation AAAI Press, 2021 Smith & Tolbert (2004) Daniel. Smith and Caroline. Tolbert âEducated by Initiative: The Effects of Direct Democracy on Citizens and Political Organizations in the American Statesâ Ann Arbor, MI: University of Michigan Press, 2004 Spärck (1972) Karen Spärck âA Statistical Interpretation of Term Specificity and Its Application in Retrievalâ In Journal of Documentation 28.1 Emerald, 1972, p. 11â21 Stab & Gurevych (2017) Christian Stab and Iryna Gurevych âParsing Argumentation Structures in Persuasive Essaysâ In Computational Linguistics 43.3 MIT Press, 2017, p. 619â659 Sunstein (2007) Cass. Sunstein âRepublic.com 2.0â Princeton, NJ: Princeton University Press, 2007 Tessler et al. (2024) Michael Tessler et al. âAI Can Help Humans Find Common Ground in Democratic Deliberationâ In Science 386.6719 American Association for the Advancement of Science, 2024, p. eadq2852 Tolbert et al. (2009) Caroline. Tolbert, Daniel. Smith and John. Green âStrategic Voting and Legislative Redistricting Reform: District and Statewide Representational Winners and Losersâ In Political Research Quarterly 62.1 SAGE Publications, 2009, p. 92â109 Toots (2019) Maarja Toots âWhy E-Participation Systems Fail: The Case of Estoniaâs Osale.eâ In Government Information Quarterly 36.3 Elsevier, 2019, p. 546â559 Wachsmuth et al. (2018) Henning Wachsmuth, Shahbaz Syed and Benno Stein âRetrieval of the Best Counterargument without Prior Topic Knowledgeâ In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL 2018) Association for Computational Linguistics, 2018, p. 241â251 Weinberg (2022) Lindsay Weinberg âRethinking Fairness: An Interdisciplinary Survey of Critiques of Hegemonic ML Fairness Approachesâ In Journal of Artificial Intelligence Research 74, 2022, p. 75â109 DOI: 10.1613/jair.1.13196 Willson (2014) Michele Willson âThe Politics of Social Filteringâ In Convergence: The International Journal of Research into New Media Technologies 20.2 SAGE Publications, 2014, p. 218â232 Wu et al. (2024) Qingyun Wu et al. âAutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversationâ In Proceedings of the First Conference on Language Modeling (COLM 2024), 2024 Appendix A The Comparison This Manuscript Declines to Run Section 4.3 states in the body why no learned-ranker baseline appears in this manuscript. This appendix records what such an experiment would have to look like to be worth running. The claim is normative and concerns admissibility, not performance. It says that a mechanism which cannot be recomputed, attributed, contested or reconstructed is unsuitable for a binding civic process. A comparison against a learned ranker would report a difference in coverage, in ordering, or in some downstream engagement statistic. Whatever number emerged could not bear on the claim: a learned ranker covering more of the live vocabulary would still fail C2, C4, C5 and C6, and would still leave a disputing participant with no terminating move. Running the comparison and declining to act on its result would be theatre; running it and acting on its result would concede that legibility is a quantity to be traded, which is the position this manuscript exists to reject. Three supporting observations: The comparison that does bear on a real question was run. Section 7.3 measures the served slates against a label-reading greedy cover which, by Proposition 3.7, upper-bounds what any selection procedure could achieve on the same pool â a learned ranker included, and by a wide margin, since the ceiling saturates the live vocabulary in every seed. The answer, 0.035Âą0.0130.035Âą 0.013, bounds the entire competitive advantage available to an unconstrained mechanism. That is strictly more informative than beating one particular trained model, because it is a bound rather than a match result. The precedent structure is familiar. The secret ballot is not defended on the grounds that it measures preferences more accurately than open voting â it plainly measures some things less well, since it destroys the ability to audit an individualâs vote. It is defended because a procedural property outranks measurement quality. Double-entry bookkeeping is not the most compact representation of a firmâs accounts. Rules of order do not produce the fastest decisions. In each case the procedural property is treated as prior, and efficiency questions are settled within the admissible class. A learned baseline would import the very opacity at issue. Its behaviour would depend on training data, objective, and the operating point chosen by whoever tuned it â all under our control as the authors, and none of it inspectable by a reader. A result favouring our rule would be unconvincing for exactly that reason, and a result favouring the learned ranker would be equally unconvincing. The experiment has no configuration in which its outcome is informative. A.1. What would make such an experiment worth running A reader who rejects the admissibility framing is entitled to ask what evidence would move us, and there is a specific answer. The interesting experiment is not ranker versus rule on coverage. It is participant behaviour under a disputed slate: give two matched populations the same corpus and the same served items, differing only in whether the selection can be recomputed and decomposed, then measure whether participants contest slates, whether contests terminate, and whether reported trust in the outcome differs. That experiment tests the actual claim â that legibility does procedural work â and its outcome could genuinely change our position or its strength, in either direction. It requires human participants and is out of scope for a simulation study; we regard it as the most valuable experiment this line of work has not yet done, and it belongs on the roadmap of Section 14 alongside the field deployment.