Paper deep dive
Validating Political Position Predictions of Arguments
Jordan Robinson, Angus R. Williams, Katie Atkinson, Anthony G. Cohn
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/20/2026, 10:08:19 PM
Summary
This paper introduces a dual-scale validation framework for political stance prediction in argumentative discourse, addressing the challenge of validating subjective, continuous attributes. Using 22 large language models, the authors constructed a knowledge base of political position predictions for 23,228 arguments from 30 BBC 'Question Time' debates. The study combines pointwise human annotation (for political sentiment classification) and pairwise human annotation (for relative ranking validation). Results show moderate agreement in pointwise evaluation (Krippendorff's α=0.578) but strong alignment in pairwise rankings (α=0.86 for the best model), demonstrating that ordinal structure can be reliably extracted from LLM predictions of subjective discourse.
Entities (13)
Relation Signals (7)
Question Time â sourceof â 23,228 Arguments
confidence 98% · 23,228 arguments drawn from 30 debates that appeared on the UK political television programme Question Time.
Dual-scale validation framework â combines â Pointwise and Pairwise Annotation
confidence 95% · combining pointwise and pairwise human annotation.
22 Language Models â usedfor â Political Position Prediction
confidence 95% · Using 22 language models, we construct a large-scale knowledge base of political position predictions
Best Model â achieves â Alpha 0.86
confidence 90% · pairwise validation reveals substantially stronger alignment between human- and model-derived rankings (α=0.86 for the best model).
Knowledge Base â enables â Graph-based Reasoning
confidence 90% · enabling graph-based reasoning and retrieval-augmented generation in political domains
Neo4j â usedtoinstantiate â Knowledge Base
confidence 90% · Arguments, defeat relations, and AIF dialogue transition structures were instantiated in a Neo4j graph database
ASPICâ â usedtoconvert â Argument Interchange Format
confidence 85% · converted into an ASPIC + argumentation theory
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Real-world knowledge representation often requires capturing subjective, continuous attributes -- such as political positions -- that conflict with pairwise validation, the widely accepted gold standard for human evaluation. We address this challenge through a dual-scale validation framework applied to political stance prediction in argumentative discourse, combining pointwise and pairwise human annotation. Using 22 language models, we construct a large-scale knowledge base of political position predictions for 23,228 arguments drawn from 30 debates that appeared on the UK politicial television programme \textit{Question Time}. Pointwise evaluation shows moderate human-model agreement (Krippendorff's $\alpha=0.578$), reflecting intrinsic subjectivity, while pairwise validation reveals substantially stronger alignment between human- and model-derived rankings ($\alpha=0.86$ for the best model). This work contributes: (i) a practical validation methodology for subjective continuous knowledge that balances scalability with reliability; (ii) a validated structured argumentation knowledge base enabling graph-based reasoning and retrieval-augmented generation in political domains; and (iii) evidence that ordinal structure can be extracted from pointwise language models predictions from inherently subjective real-world discourse, advancing knowledge representation capabilities for domains where traditional symbolic or categorical approaches are insufficient.
Tags
Links
- Source: https://arxiv.org/abs/2602.18351v1
- Canonical: https://arxiv.org/abs/2602.18351v1
Trouble viewing inline? Open PDF directly â
Full Text
59,805 characters extracted from source content.
Expand or collapse full text
Validating Political Position Predictions of Arguments Jordan Robinson 1,2 , Angus R. Williams 2 , Katie Atkinson 1,2 , Anthony G. Cohn 2,3 1 University of Liverpool 2 The Alan Turing Institute 3 University of Leeds Abstract Real-world knowledge representation often requires captur- ing subjective, continuous attributes â such as political po- sitions â that conflict with pairwise validation, the widely accepted gold standard for human evaluation. We address this challenge through a dual-scale validation framework applied to political stance prediction in argumentative dis- course, combining pointwise and pairwise human annotation. Using 22 language models, we construct a large-scale knowl- edge base of political position predictions for 23,228 argu- ments drawn from 30 debates that appeared on the UK politi- cial television programme Question Time. Pointwise eval- uation shows moderate humanâmodel agreement (Krippen- dorffâs α = 0.578), reflecting intrinsic subjectivity, while pairwise validation reveals substantially stronger alignment between human- and model-derived rankings (α = 0.86 for the best model). This work contributes: (i) a practical valida- tion methodology for subjective continuous knowledge that balances scalability with reliability; (i) a validated structured argumentation knowledge base enabling graph-based reason- ing and retrieval-augmented generation in political domains; and (i) evidence that ordinal structure can be extracted from pointwise language models predictions from inherently sub- jective real-world discourse, advancing knowledge represen- tation capabilities for domains where traditional symbolic or categorical approaches are insufficient. 1 Introduction This paper addresses the challenge of validating large-scale, pointwise large language model (LLM) predictions for sub- jective continuous variables, specifically political position scoring in argumentative discourse. While pointwise human annotation is scalable, humans are cognitively ill-equipped for precise, pointwise judgements (Mosteller and Tukey 1977; Likert 1932). Research has long demonstrated that humans excel at comparative judgments but struggle with absolute positioning on continuous scales (Thurstone 1927; Tarlow et al. 2021). Pairwise comparison â where annotators judge which of two items has more of some attribute â aligns better with human cognitive capabilities (Kendall and Smith 1940; Agresti 2018), even though it is significantly more expensive: validating n items requires O(n 2 ) comparisons rather than O(n) pointwise judgments, making full pairwise validation prohibitive at scale. We propose a dual-scale vali- dation framework that combines both approaches: pointwise validation to identify political arguments and pairwise vali- dation to assess the relative ordering of political positions. Grounded in a knowledge base of 23,228 argumentative discourse units (ADUs) (Peldszus and Stede 2013) extracted from 30 BBC Question Time debates, we employ 22 LLMs to predict the political positions of arguments along the leftâ right wing spectrum. Human validation is conducted in two stages using over 1,500 crowdworkers. The resulting knowl- edge graph integrates formal argumentative relations with political position predictions, enabling fine-grained analysis of political discourse and downstream applications, such as graph-based retrieval-augmented generation (RAG). Argument(ation) mining and political science have largely evolved in isolation until recently, with early work on political argumentation mining published by Lippi and Tor- roni (2016a). Argument mining (Lawrence and Reed 2019; Lippi and Torroni 2016b) focusses on identifying argumen- tative structure â such as premises, conclusions, relations like support and attack, and argumentation schemes â pay- ing no attention to an argumentâs political sentiment. This separation leaves a gap: we still lack large-scale, structured resources that jointly represent what is being ar- gued and where those arguments fall on the political spec- trum. Existing approaches struggle to support granular anal- yses of political discourse, such as how ideological positions propagate through argumentative structures or how political bias manifests at the ADU-level. Our work directly addresses this gap. We introduce the first large-scale knowledge base that unifies locutions and their corresponding ADUs with political position predic- tions, enabling political stance analysis at the level of the atoms of individual arguments, rather than entire texts or speakers. Taking 22 LLMs and validating their predictions through a novel dual-scale human annotation methodology, we show that scalable pointwise model predictions can be meaningfully aligned with human comparative judgements. Thus, we extend the literature on political argumentation by combining structured argumentation corpora with vali- dated political stance predictions. The resulting resource 1 opens new research strands in political argumentation, com- 1 Code and the containerised knowledge base are avail- able on GitHub: https://github.com/anonymous-argumentation/ Validating-Political-Position-Predictions-of-Arguments. arXiv:2602.18351v1 [cs.CL] 20 Feb 2026 putational social science and graph-based RAG, supporting analyses that were previously infeasible due to the lack of structured, politically annotated argumentative data. 2 Related Work LLMs as Evaluators of Language Outputs. Prior to re- cent advancements in LLMs, evaluation of natural language generation primarily relied on automated metrics, such as BLEU (Papineni et al. 2002), ROUGE (Lin 2004), and later embedding-based approaches including BERTScore (Zhang et al. 2020), BARTScore (Yuan, Neubig, and Liu 2021) and GPTScore (Fu et al. 2023). While effective for supervised tasks with reference outputs, these metrics are limited in their ability to assess open-ended, normative, or context- dependent language, as they rely on surface-level similarity or static semantic representations. The emergence of powerful LLMs capable of nuanced language understanding and reasoning in unsupervised tasks has led to their adoption as evaluators, or âjudgesâ, of gener- ated text. Recent work demonstrates that LLMs can reliably assess outputs produced by both humans and other models across a range of tasks (Chen et al. 2023; Zhang et al. 2023; Chen et al. 2024b; Wang et al. 2024; Chen et al. 2024a), in- cluding instruction following, question answering, and com- plex domain-specific reasoning (Huang and Chang 2023; Zhao et al. 2025). This shift reflects a broader trend in which generative models are used not only for evaluation but also for prediction and classification tasks that require interpret- ing subtle semantic and pragmatic cues. Political Position Prediction with LLMs. Within politi- cal science and computational social science, language mod- els have been used for political stance prediction at vari- ous granularities. For example, models were tasked with the prediction of politiciansâ political leanings using differ- ent ideological axes, such as gun control (Wu et al. 2023). GPT 3 and 4 were prompted to estimate the probability of a sentence being either conservative or liberal across politi- cal party manifestos, taking the sentence-level average as an analogue for the mean position of each document (Ornstein, Blasingame, and Truscott 2025). Multiple models were used to predict the political stance of sentences, using a scale from 0 (extremely left-wing) to 100 (extremely right-wing), within sets of Tweets, manifestos, and policy speeches across ten different languages (Le Mens and Gallego 2025). Despite differences in scale and domain, this body of work predominately adopts pointwise model evaluation approaches, in which models assign absolute ideological scores or labels that are subsequently compared against ag- gregate human annotations. As we discuss below, such ap- proaches may obscure systematic differences in judgement and place substantial cognitive demands on both human an- notators and models. Pairwise Comparison and Preference-Based Evaluation. A large body of psychological and decision-theoretic re- search has shown that pairwise comparison is cognitively simpler and more reliable than absolute rating for subjec- tive judgement tasks (Thurstone 1927). This insight under- pins preference learning (F Ì urnkranz and H Ì ullermeier 2003), which has become central to recent advances in LLM train- ing. In particular, Reinforcement Learning from Human Feedback (RLHF) relies on pairwise human preferences to train reward models, with BradleyâTerry-style models com- monly used to infer latent strength scores from comparison data (Christiano et al. 2017; Ouyang et al. 2022). In contrast, existing work on LLM-based political posi- tion prediction has largely relied on pointwise scores rather than direct comparative judgements (Le Mens and Gallego 2025). Recent studies have shown that aggregate correla- tion metrics can mask systematic disagreements between model outputs and human annotations, particularly when hu- man uncertainty is high (Elangovan et al. 2025). Accord- ingly, chance-corrected agreement measures, such as Krip- pendorffâs α, have been argued to provide a more appropri- ate basis for LLMâhuman comparison, as they explicitly ac- count for agreement expected by chance (Haldar and Hock- enmaier 2025). These findings motivate the use of pairwise, preference-based evaluation for political stance judgements. Political Argumentation Resources. Existing political argumentation and stance detection datasets predominantly provide categorical annotations. Stance detection resources typically label texts as for or against a target or along a discrete leftâright scale (Sim et al. 2013; Mohammad et al. 2016), while argument mining datasets focus on identifying argumentative components and relations without explicitly modelling ideological position (Lippi and Torroni 2016a; Menini et al. 2018; Haddadan, Cabrio, and Villata 2019; Visser et al. 2020; Mestre et al. 2021; Goffredo et al. 2022). Continuous measures of political stance and compar- ative judgements between arguments remain relatively un- explored, particularly in settings that combine ideological positioning with structured argumentative content. 3 Knowledge Base Construction Figure 1 provides an overview of the pipeline we have de- veloped for knowledge base instantiation, prediction of po- litical stances, and application via graph-based RAG. We ex- plain each part of the pipeline below. 3.1 Data Sources and Preprocessing We constructed a knowledge base from 30 BBC Ques- tion Time debates 2 that were previously annotated for arguments by expert annotators with an inter-annotator agreement (Combined Argument Similarity Score) of 0.56 (Hautli-Janisz et al. 2022). The annotated public corpora, represented in the Argument Interchange Format (AIF) (Ches Ì nevar et al. 2006; Rahwan and Reed 2009), were converted into an ASPIC + argumentation theory follow- ing Bex et al. (2012), restricting the framework to ordinary premises and defeasible inference rules without preferences. 2 https://corpora.aifdb.org/qt30 Figure 1: Overview of the methodology used to instantiate a structured argumentative knowledge base containing political positions predic- tions for all arguments. Green boxes indicate components developed specifically for this study. Arguments, defeat relations, and AIF dialogue tran- sition structures were instantiated in a Neo4j graph database (Neo4j 2012).We adopt the ASPIC + frame- work because it provides a formally rigorous bridge be- tween natural-language argument structure and Dung-style abstract semantics (Prakken 2010; Modgil and Prakken 2014), supporting principled evaluation of argument accept- ability. Granularity is important for graph-based RAG sys- tems emulating political personas, which require ideologi- cally aligned retrieval and coherent argumentative genera- tion. Constructing personas with internally consistent and politically aligned argument sets is left for future work. The resulting knowledge base contains 23,228 arguments, consisting of locution-proposition pairs, and 50,905 rela- tions comprised of support, attack, rephrase and chronolog- ical transition. 3.2 LLMs as Judges of Political Position We used 22 LLMs to predict the political positions of argu- ments in the knowledge base, visible in Table 1. Building on (Le Mens and Gallego 2025), we extended sentence-level political stance predictions to both utterances and their corresponding ADU, interchangeably referred to as locution-proposition pairs or arguments from now on. Each model assigned a political stance score to all locution- proposition pairs on a 0-100 scale (left-right), or âNAâ if an argument lacked political content. An example prompt is shown in Figure 1. Each model scored every node in the graph five times to account for output variability even under near-deterministic settings (i.e., temperature set to 0, top p equal to 0.1, and fixed seed for models that allowed parameter configuration). Prompting was executed using Golem (Blackwell 2024) to ensure reproducible prompts with consistent parameter con- figurations at scale. We derived summary statistics for politi- cal position scores and the probability of an argument being labelled apolitical (or âNAâ) across repetitions, and stored these as properties of nodes in the knowledge base. 3.3 Ensemble Construction and Aggregation We constructed three ensembles of model predictions, each designed to isolate different aspects of model behaviour and robustness and provide additional insights in evaluation. Summary statistics for political position score and probabil- ity of NA were calculated from the set of all individual pre- dictions for all models in an ensemble, and stored as proper- ties for each node in addition to individual model predictions and summary statistics. Ensemble 1 [E 1 ]: All (n = 22). A general-purpose base- line capturing average model behaviour across all models. Ensemble 2 [E 2 ]: Reasoning Models (n = 5). Models with explicit reasoning or chain-of-thought capabilities, to establish whether models designed to show their reasoning steps produce predictions more aligned with human judge- ments. Ensemble 3 [E 3 ]: High-Confidence Models (n = 12). Models with high-confidence judgements, defined by the number of valid political stance predictions between 0-100 exceeding the number of NA predictions. Introduced after initial results indicated that smaller models were incapable of classifying arguments as NA, even for samples which were apolitical. 3 E 3 =mâ E 1 : |Ëy pol (m)| >|Ëy apol (m)|(1) where m is a model, Ëy pol (m) is the number of arguments predicted as political by m, and Ëy apol (m) is the number ar- guments predicted NA by m. ModelOfficial NameE 1 E 2 E 3 Claude 3.5 Haiku claude-3-5-haiku-20241022 (Anthropic 2024)â Claude 3.7 Sonnet claude-3-7-sonnet-20250219 (Anthropic 2025)â DeepSeek-R1 deepseek-r1-0528 (DeepSeek-AI 2025a)â DeepSeek-V3 deepseek-v3-0324 (DeepSeek-AI 2025b)â Gemini 1.5 Pro gemini-1.5-pro-002 (Gemini Team 2024)â Gemini 2.5 Flash gemini-2.5-flash-preview (Gemini Team 2025)â GPT 3.5 Turbo gpt-3.5-turbo-0125 (OpenAI 2024a)â GPT 4 Turbo gpt-4-turbo-2024-04-09 (OpenAI 2024b)â GPT 4o gpt-4o-2024-08-06 (OpenAI 2024c)â GPT 4o Mini gpt-4o-mini-2024-07-18 (OpenAI 2024d)â GPT 4.5 gpt-4.5-preview-2025-02-27 (OpenAI 2025a)â Grok 2 grok-2-1212 (xAI 2024)â Llama 3.1:8b llama3.1:8b (Meta AI 2024a)â Llama 3.1:405b llama-3.1-405b-instruct (Meta AI 2024a)â Llama 3.2:3b llama3.2:3b (Meta AI 2024b)â Llama 3.3:70b llama-3.3-70b-instruct (Meta AI 2024c)â Llama 4 Maverick llama-4-maverick (Meta AI 2025)â Mistral:7b mistral:7b (Mistral AI 2023)â o3 Mini o3-mini-2025-01-31 (OpenAI 2025b)â Phi 4 microsoft/phi-4 (Microsoft 2024)â Qwen 3 qwen3-235b-a22b (Qwen Team 2025)â Qwen QwQ qwq-32b (Qwen Team 2024)â Total22512 Table 1: Models included in each ensemble. 4 Human Annotation and Validation Design In order to validate model predictions of political positions introduced in Section 3, we made use of human crowdwork- ers to annotate arguments from the knowledge base. Crowd- workers were recruited via Prolific. 4 All participants re- cruited resided in the UK, with English as first or primary language, and some form of higher education. Following the structure of prompts presented to LLMs, we divided human annotation into two sequential tasks. 4.1 Pointwise Binary Classification of Political Sentiment This task comprised pointwise binary classification as to whether an argument is political or apolitical, validating model predictions while also providing a pool of high- confidence political arguments to validate model scores on. Annotation. We randomly sampled 1,000 arguments from our knowledge base into three buckets, based on the mean probability of NA for Ensemble 3 ( ÌÏ (E 3 ) ), in order to bias the dataset towards cases of high model confidence while still including some cases of low model confidence, as shown in Table 2. All arguments obtained a majority label from three hu- man annotations (binary labels), using a pool of 600 crowd- workers, each annotating five arguments. Two participants exhibiting invariance and atypical completion speed were flagged and replaced. We used nominal Krippendorffâs α n 3 E.g. some models assigned a score to nodes which contained a locution such as âI agreeâ which possessed no political sentiment. 4 https://w.prolific.com (Accessed on 4th December 2025) Bucket n ÌÏ (E 3 ) Interpretation H pol 400 †0.05high confidence, political L200 â [0.45, 0.55]low confidence H apol 400 â„ 0.95high confidence, apolitical Table 2: Sampling buckets for pointwise binary classification of the presence of political sentiment. (Hayes and Krippendorff 2007; Krippendorff 2004) to mea- sure inter-annotator agreement, accommodating partial an- notator overlap and chance agreement. Confidence-Based Dataset Partitioning. Let D (NA) de- note the full dataset comprised of arguments from the H pol , L, and H apol confidence buckets from Table 2. We partitioned D (NA) into disjoint subsets based on model prediction confi- dence as follows: D (NA) conf =D (NA) LOW âȘD (NA) H apol , D (NA) ambig =D (NA) L .(2) Here, D (NA) conf contained arguments that models predicted as political or apolitical with high confidence, while D (NA) ambig consisted of arguments whose political status was charac- terised by model uncertainty. By construction, these confidence-based subsets formed a strict partition of the dataset.Specifically, the high- confidence and ambiguous subsets were disjoint and their union recovers the full dataset: D (NA) conf â©D (NA) ambig =â and D (NA) conf âȘD (NA) ambig =D.(3) This guaranteed that every argument was assigned to exactly one subset, ensuring complete coverage without overlap and enabling controlled comparisons between high- confidence and ambiguous cases. Evaluation. We evaluated presence of political sentiment using binary human majority labels and thresholded model predictions acrossD (NA) ,D (NA) conf , andD (NA) ambig . We calculated F1 score, precision, recall, and balanced accuracy (treating human labels as ground truth), in addition to α n (framing parties as two equal-weight raters), allowing us to evaluate model performance relative to human reliabil- ity baselines. 4.2 Pairwise Comparison of Political Position To validate the implicit ranking of arguments from pointwise model predictions, we presented human annotators with pairs of arguments, tasking them with annotating which one was more left- or right-leaning. Model-predicted political position scores were converted into pairwise comparisons and evaluated against human pairwise judgements via rank- ings derived from BradleyâTerry (BT) models (Bradley and Terry 1952) and pairwise classification performance metrics across confidence levels. Sampling Pairs. We sampled n = 100 arguments that were unanimously labelled as political by human annotators fromD (NA) H pol , stratified according to position scores predicted by Ensemble 3. Continuous position predictions were dis- cretised into deciles; bins corresponding to the ranges 0â 10 and 90â 100 were empty and therefore left excluded, yield- ing eight non-empty bins B. We denote the binned posi- tion score for argument i, repetition a, under model m as s (m) i,a â B =1,..., 8. In order to do efficient comparison under resource con- straints, we annotated a subset of 934 pairs drawn from the 100 2 = 4, 950 possible argument pairings. Pair selection was guided by the mean predicted position score Ìs (E 3 ) from Ensemble 3. Specifically, we sampled 44 intra-bin pairs from each bin (except for bin 8, where only 10 intra-bin pairs were available) and 22 inter-bin pairs. We denote the result- ing set of annotated pairs asP =(i k ,j k )|k = 1,..., 934 where each element (i k ,j k ) corresponds to an ordered pair of arguments. The resulting set exceeds the n lnn â 460.5 target for pairwise connections (Negahban, Oh, and Shah 2012), where n = 100 for this study. We verified full connectivity of the resulting graph, and confirmed even distributions of connections by computing node connection entropy values, defined as the Shannon entropy (Shannon 1948) H(v i ) of an item v i . This is calculated as the sum of proportions of com- parisons f between v i and each bin bâ B (f i,b ) weighted by their logarithms, shown in Equation (4). With an ideal up- per bound log 2 (8)â 3 (uniform distribution across all eight bins), the median value of the resulting distribution was 2.5, with more than 60% of scores falling within 2.2â2.8, indi- cating balanced connections (lower values relate to bins with availability constraints). H(v i ) =â X bâB f i,b log 2 (f i,b )(4) Win Matrices. We represent political position judgments as a hollow comparison matrix W âR 100Ă100 , where W ij represents aggregated win count of item i over item j. A win indicates that i was judged as more right-wing than j, with draws treated as half-wins, contributing 0.5 to both W ij and W ji (Davidson 1970). Model Comparisons. For each model m (22 LLMs and 3 ensembles), we constructed a dense win matrix W (m) from binned political prediction scores s (m) . Models pre- dicted political position for all possible pairs, a superset of P . Each cell W (m) ij , representing argument i versus argu- ment j, was computed across all combinations of N 5 repe- titions of model predictions: W (m) ij = N X a,b=1 w s (m) i,a ,s (m) j,b w(x,y) = ïŁ± ïŁŽ ïŁČ ïŁŽ ïŁł 1 x > y 0.5 x = y 0 x < y (5) Human Comparisons. We employed 936 crowdworkers across two symmetric tasks. Participants were presented with two arguments and the question âWhich statement 5 For single LLMs, N = 5, whereas for ensembles N = 5Ă ensemble size. E 3 Confident E 3 Unconfident H Confident P 1,1 P 0,1 H Unconfident P 1,0 P 0,0 Table 3: Subsets ofP by model and human confidence. is more left-wing?â or âWhich statement is more right- wing?â, and were tasked with selecting argument i, argu- ment j, or âequalâ. Each pair was annotated 3 times per task (6 annotations per pair overall). We constructed win matri- ces as before, with draws representing a half-win, or +0.5. We transposed left-framed annotations and aggregated across tasks, resulting in 3 human win matrices: inverted left-framed W (H L ) , right-framed W (H R ) , and aggregate hu- man win matrices W (H) = W (H L ) + W (H R ) . Pairwise Modelling. We used BT with the Iterative Luce Spectral Ranking (I-LSR) algorithm (Maystre and Gross- glauser 2015) provided in Python library choix, to model pairwise judgements. 6 This mapped comparisons onto a la- tent scaleΞ, where the probability of an outcome was given by p(i > j) = e Ξ i e Ξ i +e Ξ j . This process yielded a âsmoothedâ 100 Ă 100 probability matrix, imputing missing compar- isons, and a ranking by sorting arguments in descending order of Ξ i , where higher values indicate more right-wing positions. We useR (m) to represent ranking determined by model m. Confidence-based Dataset Partitioning. We assigned pairs fromP into subsets based on the confidence of human and model (using ensemble 3) judgements, to account for uncertainty introduced across model repetitions and through the 6 multi-class human annotations per pair. A given pair i,j in P was assigned two confidence val- ues (|W ij â 0.5| â„ 0.25 â 0, 1), for W (E 3 ) and W (H) , with 1 indicating high-confidence. This provided four sub- setsP E 3 ,H âP across different combinations of model and human confidence (Section 4.2). Evaluation. We quantified agreement between rankings on our n = 100 arguments using Spearmanâs Footrule Dis- tance (Diaconis and Graham 1977) and Kendallâs Ï Distance (Kendall 1938), min-max normalised to [0,1] and inverted to represent similarity metrics, which we represent as d footrule and d Ï respectively. We also calculated ordinal Krippen- dorffâs α o to represent agreement. We calculate these met- rics between each model rankingR (m) and the human rank- ing R (H) , and use distance/agreement between aggregate human rankingR (H) and rankings derived from the left- and right-framed annotation tasks (R (H L ) and R (H R ) ) to pro- vide context to model results. We calculated pairwise performance (macro-f1) of mod- els, using human judgements as ground-truth, for all pairs receiving human annotation (P ). Labels represent a win or loss for the argument i in the pair (i,j) as judged by model m, and are calculated by thresholding W (m) ij at 0.5. 6 https://choix.lum.li/ (Accessed on: 12th January 2026) We present results across high- and low- human and model confidence level subsets P E 3 ,H to breakdown how model performance differs across confidence levels. To further contextualise model performance, we estab- lished two control baselines: 1. Random Baseline: no discriminative ability, i.e. assigns equal strength parameters (zero values) to all items, re- sulting in uniform probability matrices of 0.5. 2. Worst-Case Baseline: systematically inverted predic- tions relative to human judgments, generated by negat- ing the human aggregate model parameters and inverting probability matrices. 5 Experimental Results and Evaluation 5.1 Pointwise Annotation Study Inter-Annotator Agreement. Overall, 48.5% of point- wise annotations were unanimous (3/3). To assess the re- liability of the crowdsourced labels, we computed nominal Krippendorffâs α n . Across the full datasetD (NA) (n=1,000), inter-annotator agreement was low (α n = 0.305), indicating poor agreement amongst annotators. When restricting the analysis to the confident subsetD (NA) conf (n=800), agreement increased slightly to α n = 0.317, whereas when only ambiguous cases were consideredD (NA) ambig (n=200), agreement decreased to α n = 0.259. We note that α n = 0.483 for unanimously-labelled arguments (n=485) and α n = 0.436 for majority labels (n=515). Inter-Model Agreement. Figure 2 reports α n measuring agreement among model predictions under different dataset partitions. When models were evaluated exclusively on the ambiguous subset D (NA) ambig , agreement is at or below chance (α n = â0.048), indicating highly inconsistent predictions in regions of uncertainty. When considering the full datasetD (NA) , there was a sub- stantial increase in agreement between models, such that models exhibited fair agreement (α n = 0.485), with the best model agreement observed in the confident partition D (NA) conf (α n = 0.578), approaching the moderate agreement level. Figure 2: Nominal Krippendorffâs α n measuring inter-model agreement across data partitions. HumanâModel Agreement.Figure 3 shows the distribu- tion of α n between human annotations and model predic- tions across dataset partitions. Humanâmodel agreement is highest onD (NA) conf , with a median of α n = 0.424. In contrast, agreement onD (NA) ambig is consistently negative, indicating sys- tematic divergence between model predictions and human annotations. Figure 3: Distribution of humanâmodel agreement across dataset partitions. Model and Ensemble Performance.Figure 4 illus- trates the relationship between humanâmodel agreement and model performance (macro F1, micro F1 and balanced accu- racy) across all models and ensembles, evaluated onD (NA) , D (NA) conf , andD (NA) ambig . Across all metrics, higher agreement with human judgements strongly predicts better performance. Macro F1 exhibits near-perfect correlations with agreement in both partitions and the full dataset, with similarly strong correlations observed for micro F1 and balanced accuracy. Performance is consistently higher onD (NA) conf , demonstrat- ing that models are most reliable on unambiguous instances. Moreover, D (NA) ambig exhibits substantial class imbalance as demonstrated by comparing macro and micro F1, likely con- tributing to degraded performance across all metrics and weaker agreement in this partition. These results establish humanâmodel agreement as a robust proxy for model qual- ity in pointwise political sentiment classification. Beyond agreement-based analysis, we explicitly evalu- ated predictive performance for the three ensembles intro- duced in Section 3.3, alongside the best performing models (Table 4). All three ensembles demonstrate remarkably con- sistent performance, with near-identical metrics across most conditions. Again, dataset partitioning had a substantial im- pact on performance across all ensembles and models, with results fromD (NA) conf consistently outperforming both ambigu- ous cases and the full dataset. Discussion. Across inter-annotator,inter-model,and humanâmodel analyses, a consistent pattern emerges: agree- ment is systematically higher onD (NA) conf and degrades sharply when ambiguous instances are included. Human annotators Figure 4: Relationship between humanâmodel agreement and model performance metrics. Macro F1Micro F1Bal. Acc. ModelD (NA) D (NA) conf D (NA) ambig D (NA) D (NA) conf D (NA) ambig D (NA) D (NA) conf D (NA) ambig Ensemble 10.6600.7120.3080.6600.7140.4450.6800.7200.500 Ensemble 20.6600.7120.3480.6600.7140.4450.6780.7200.493 Ensemble 30.6750.7120.5280.6770.7140.5300.6810.7200.530 GPT 4o0.6830.7120.5700.6850.7140.5700.6910.7200.575 DeepSeek-V30.6800.7120.5500.6810.7140.5500.6880.7200.559 Grok 20.6730.7120.5120.6740.7140.5150.6820.7200.530 Gemini 2.5 Flash0.6710.7120.4810.6710.7140.5000.6830.7200.527 GPT 3.5 Turbo0.6550.6910.4970.6550.6910.5100.6720.7050.534 GPT 4.50.6730.7120.4480.6810.7140.5500.6740.7200.509 Llama 4 Maverick0.6660.7120.4410.6660.7140.4750.6800.7200.508 Gemini 1.5 Pro0.6600.7120.4440.6630.7140.4600.6630.7200.447 Qwen 30.6630.7120.4350.6630.7140.4600.6760.7200.489 Claude 3.7 Sonnet0.6660.7120.4110.6730.7140.5100.6660.7200.471 Table 4: Ensemble and top ten performing models acrossD (NA) , as well asD (NA) conf andD (NA) ambig partitions. exhibit low overall agreement, reflecting the inherent sub- jectivity of pointwise political sentiment annotation, while models achieve substantially higher consistency on the same confident subset. Both humans and models struggle onD (NA) ambig , indicating a region of genuine semantic uncertainty rather than stochas- tic prediction noise. The convergence between human dis- agreement and model uncertainty provides empirical evi- dence that humanâmodel agreement is a reliable indicator of model prediction quality. Excluding ambiguous instances yields marked improvements in both agreement and down- stream performance, reinforcing the conclusion that point- wise model predictions can approximate human judgements of pointwise binary classification of political sentiment. 5.2 Validating Model Predictions using Pairwise Human Annotations Inter-Annotator Agreement. We measured ordinal Krip- pendorffâs α o between the inverted left-framed and right- framed annotation tasks to assess the reliability of the crowdsourced pairwise comparisons. From a carefully se- lected subset of 100 political arguments, we constructed 934 samples and obtained α o = 0.889, indicating substantial agreement between framing and supporting the robustness of our pairwise annotation protocol. Table 5 reports the distribution of the number of unique labels assigned per pair across annotators for the left- framed and right-framed pairwise annotations tasks.In both conditions, the majority of pairs received two distinct labels (57.3% left-framed; 62.8% right-framed), indicat- ing broader annotator consensus with limited disagreement. Pairs exhibiting complete agreement â characterised by a single unique label â account for a substantial minority of comparisons (27.9% and 22.5%, respectively), while com- plete disagreement (three unique labels) is rare and occurred at comparable rates across framings (14.8% left-framed; 14.7% right-framed). The close correspondence between distributions further confirms that inter-rater agreement pat- terns were stable across annotation framings. Left-framedRight-framed Number of unique labels%n%n 127.926122.5210 257.353562.8587 314.813814.7127 Table 5: Unique labels per pair across annotation tasks. Model and Ensemble Performance, Agreement and Dis- tance Metrics. To evaluate alignment between model- inferred political rankings and human judgements under varying levels of uncertainty, we report agreement and per- formance metrics over the full datasetP (Figure 5) and over four conditional subsets (Figure 6): P 1,1 (high-confidence model predictions and human annotations), P 1,0 (high- confidence model predictions, low-confidence human an- notations), P 0,1 (low-confidence model predictions, high- confidence human annotations), and P 0,0 (low-confidence predictions on both sides). We report Spearmanâs d footrule , Kendallâs Ï , ordinal α o , and macro F1 for all ensembles and the top-performing individual models (Table 6). Across all models, ranking agreement with humans under P remains substantially below human inter-annotator agree- ment. Individual models achieve α o values which ranged from 0.52 to 0.80, corresponding to moderate agreement but falling short of the substantial agreement threshold (α o = 0.81; Figure 5). This gap persists even among the strongest propriety models, indicating that inferring fine-grained po- litical orderings from pointwise judgements remains a chal- lenging task. Figure 5: Macro F1 plotted against ordinal humanâmodel agree- ment for all models and ensembles underP . Conditioning on high-confidence predictions markedly improves performance. Under P 1,1 , several models â in- cluding GPT 4.5, o3 Mini, DeepSeek-R1 and Claude 3.5 Haiku â approached or exceeded α o â 0.85, placing them close to the boundary of agreement between humans across the left- and right-framed tasks (Table 6). As shown in Fig- ure 6, agreement increased and variance decreased in subsets where human confidence is high, whereas low-confidence subsets exhibited substantially weaker and more variable alignment. d footrule d Ï Î± o Macro F1 ModelP P 1,1 P P 1,1 P P 1,1 P P 1,1 Human (agg)1.0001.0001.0001.0001.0001.0001.0001.000 Human (left)0.9040.8670.9340.9070.9740.9560.8271.000 Human (right)0.8860.8730.9220.9140.9660.9600.7901.000 Ensemble 10.7260.7210.8100.8000.8080.7890.6030.492 Ensemble 20.7210.6820.8040.7820.8030.7600.6220.617 Ensemble 30.7280.6570.8060.7560.7990.6920.5830.492 o3 Mini0.6840.7790.7790.8370.7660.8520.6440.723 GPT 4.50.6960.8060.7930.8520.7830.8490.6300.492 DeepSeek-R10.6910.7740.7810.8400.7530.8490.5950.617 Qwen QwQ0.6610.7600.7700.8190.7470.8130.6210.673 GPT 4o Mini0.6610.7800.7670.8290.7300.8270.6450.570 Claude 3.5 Haiku0.6790.7850.7720.8420.7340.8580.6400.483 DeepSeek-V30.6980.7530.7770.8190.7390.8270.5950.484 Gemini 1.5 Pro0.6120.7430.7340.8060.6550.7780.6100.585 Random Baseline0.3330.3330.5000.5000.0000.0000.5000.500 Worst-case Baseline0.0000.0000.0000.000-1.000-1.0000.0000.000 Table 6: Model performance metrics, by metric and dataset. Fea- tures all models that acheived top-5 performance for each metric. Bold cells indicate highest performing LLM for given column. Human Upper Bound. Human annotations established a clear upper bound on achievable performance. Aggregate human inter-annotator agreement reached α o = 1.0 by con- struction, while individual left- and right-framed annotations retained extremely high agreement with the aggregate rank- ing (α o â„ 0.966 forP ; Table 6). Notably, conditioning on high-confidence judgements (P 1,1 ) decreased the agreement slightly but increased annotation performance, yielding per- fect macro F1 for both annotation conditions. Ensembles versus Individual Models. Under the full dis- tribution P , ensemble methods consistently outperformed individual models. All three ensembles achieved α o â 0.80, outperforming most individual models and narrowing (but not closing) the humanâmodel gap (Figure 5 and Table 6). Interestingly, this advantage diminishes under P 1,1 . When attention is restricted to high-confidence predictions, several individual models match or exceed ensemble performance on ranking agreement and macro F1. Ranking Agreement v. Pointwise Model Accuracy. A key insight from from Figure 5 and Table 6 is the partial de- coupling of ranking agreement and macro F1. Some models (e.g., GPT 4o Mini) achieved a strong macro F1 underP de- spite lower ranking agreement, while others prioritise ordi- nal consistency at the expense of pointwise accuracy. Condi- tioning onP 1,1 sharpened this trade-off: macro F1 generally increased for confident subsets but gains in α o were more pronounced, highlighting that ordinal ranking benefits more from uncertainty filtering than categorical accuracy does. Discussion. Our pairwise annotation study clarifies how pointwise model-generated political judgements relate to human comparative reasoning under uncertainty. Rather than framing validation solely as agreement maximisation, our results show that confidence and aggregation play a cen- tral role in shaping the recoverable ordinal structure. The gap between human and model agreement under the full distribution is largely attributable to uncertainty. While models exhibit only moderate alignment overall, con- ditioning on high-confidence judgements reveals substan- tial agreement with the human aggregate ranking for several models. This pattern suggests that disagreement is driven less by systematic ideological bias than by structurally am- biguous regions of the ordering that also challenge human annotators. Accordingly, model outputs encode a partially correct ordinal structure, with uncertainty concentrated in weakly constrained areas of the comparison graph. Our findings further distinguish ordinal consistency from pointwise accuracy. Ordinal agreement and macro F1 are only weakly coupled, and improvements in one do not nec- essarily imply gains in the other. Confidence filtering dispro- portionately benefits ordinal structure, reinforcing the view that global rankings and local classifications capture distinct epistemic aspects of the task. Aggregation effects follow a similar pattern. Ensembles improve ordinal agreement under the full distribution by sta- bilising rankings in ambiguous regions, but offer limited ad- vantage once uncertainty is reduced, where strong individual models perform comparably. Finally, despite being elicited Figure 6: Macro F1 plotted against ordinal humanâmodel agreement for all models and ensembles under four distributions: P 1,1 , which conditioned on high-confidence model predictions and human annotations;P 1,0 , which conditioned on high-confidence model predictions and low-confidence human judgements;P 0,1 , which conditioned on low-confidence model predictions and high-confidence human annotations; andP 0,0 , which conditioned on low-confidence model predictions and human annotations. without explicit comparisons, pointwise model predictions recover a substantial portion of the ordinal information ex- pressed in human pairwise judgements, particularly in high- confidence cases. This supports their use as a scalable proxy for human comparative annotation when exhaustive pairwise judgements are impractical. 6 Limitations We note some limitations. 1) Pointwise political position an- notations is inherently difficult for humans, as seen in sub- stantially lower humanâhuman agreement compared to the pairwise study, reinforcing prior findings that absolute judg- ments on continuous ideological scales are cognitively de- manding. This limits the reliability of pointwise labels as a gold standard and motivates our dual-scale validation de- sign, but does not eliminate the underlying subjectivity of the task. 2) Ensemble construction (esp. Ensemble 3) relies on estimating the probability that an argument is apolitical, requiring multiple model invocations per sample. At run- time, it is not known whether an argument will be unambigu- ous, making the approach more computationally costly and restricting it in low-latency settings. Relatedly, ensembles primarily improve performance in ambiguous regions of the comparison space; confident cases are often handled com- parably well by individual models, though ambiguity itself cannot be identified prior to inference. 3) Our pairwise val- idation relies on discretisation of continuous model outputs and on BT regularisation choices, both of which may ob- scure finer-grained ordinal differences. 4) Pairwise rankings are derived from a relatively small subset of arguments (100) drawn from a single political discourse domain, which may limit generalisability to other corpora or ideological axes. 7 Conclusions and Future Work Our work shows that disagreement in pointwise political position annotation reflects human difficulty with absolute judgements rather than model unreliability.In contrast, pairwise comparison yields substantially higher agreement and reveals that pointwise LLM predictions recover mean- ingful ordinal political structure, particularly under high- confidence conditions. The validated subsets of our data thus delineate where political positioning can be interpreted reliably and where uncertainty is intrinsic. We release a large-scale, structured argumentation knowledge base that integrates formal argumentative structure with political po- sition predictions at the level of individual ADUs. The graph combines attack and support relations, dialogue structure, predictions from 22 LLMs, ensemble aggregates, and uncer- tainty estimates. A subset of the graph has been validated by human annotators, providing a high-confidence core within a larger predictive resource and enabling fine-grained anal- ysis beyond document- or speaker-level stance. The knowl- edge base can support studies on how ideology interacts with argumentative structure, how political distance relates to at- tack and support, and where ideological ambiguity concen- trates. It also has graph-based retrieval-augmented gener- ation, enabling retrieval of ideologically aligned and struc- turally coherent argument sets for downstream applications. Future work will extract political positions from ADUs in new domains, e,g, international political systems that do not have a dichotomous left-right division, and generate politi- cally aligned personas using the shared resource. Ethical Considerations This study received ethical approval from the University of Liverpool (Ref: 17359), University of Leeds (Ref: MEEC 25-001) and Alan Turing Institute (Ref: TR25-17). Human crowdworkers were recruited through the online research platform Prolific. All participants provided informed con- sent prior to undertaking the annotation tasks, in accordance with institutional guidelines. No personally identifiable in- formation was collected by the research team, and partici- pant anonymisation was managed by the Prolific platform. Participants were compensated at an average rate of ÂŁ9.76 per hour across both left- and right-framed annotation tasks. Data and Code Availability We release the structured knowledge base derived from 30 episodes of BBC Question Time, including arguments, their associated political position scores and relations, together with the associated Prolific-based annotation dataset and anonymised demographic metadata. The resource is pro- vided as a database dump and Dockerfile for reproducible deployment, with supporting code available on our GitHub repository:https://github.com/anonymous-argumentation/ Validating-Political-Position-Predictions-of-Arguments. All data and code are released under the MIT License. Credit Assignment Conceptualisation: J.R., K.A., A.G.C.; Knowledge base construction: J.R.; Predicting political positions of ar- guments: J.R.; Human annotation study design: J.R., A.R.W.; Running annotation study on Prolific: A.R.W.; Analysis and data science: J.R., A.R.W.; Tables: J.R., A.R.W.; Visualisation: J.R.; Open-source software: J.R.; Supervision: K.A., A.G.C.; Writing â original draft: J.R.; Writing â review and editing: J.R., A.R.W., K.A., A.G.C.. Acknowledgements The work reported in this paper was supported by fund- ing from the Alan Turing Institute. We also acknowledge support from Microsoft Researchâs Accelerating Founda- tion Models Research programme, which provided Azure resource to run a selection of the LLMs used in the experi- ments that are reported in this paper. References Agresti, A. 2018. Analysis of ordinal paired comparison data. Journal of the Royal Statistical Society Series C: Ap- plied Statistics 41(2):287â297. Anthropic. 2024. Claude 3.5 Haiku. Model identifier: claude-3-5-haiku-20241022. Anthropic. 2025. Claude 3.7 Sonnet. Model identifier: claude-3-7-sonnet-20250219. Bex, F.; Modgil, S.; Prakken, H.; and Reed, C. 2012. On logical specifications of the Argument Interchange Format. Journal of Logic and Computation 23(5):951â989. Blackwell, R. 2024. Golem â robblackwell/golem: v0.0.1- alpha, doi.org/10.5281/zenodo.14035711. Bradley, R. A., and Terry, M. E. 1952. Rank analysis of incomplete block designs: I. the method of paired compar- isons. Biometrika 39:324. Chen, Z.; Jiang, F.; Chen, J.; Wang, T.; Yu, F.; Chen, G.; Zhang, H.; Liang, J.; Zhang, C.; Zhang, Z.; Li, J.; Wan, X.; Wang, B.; and Li, H. 2023. Phoenix: Democratizing Chat- GPT across languages. https://arxiv.org/abs/2304.10453. Chen, G. H.; Chen, S.; Liu, Z.; Jiang, F.; and Wang, B. 2024a. Humans or LLMs as the judge? a study on judge- ment bias. In Al-Onaizan, Y.; Bansal, M.; and Chen, Y.- N., eds., Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 8301â8327. Mi- ami, Florida, USA: Association for Computational Linguis- tics. Chen, J.; Wang, X.; Ji, K.; Gao, A.; Jiang, F.; Chen, S.; Zhang, H.; Song, D.; Xie, W.; Kong, C.; Li, J.; Wan, X.; Li, H.; and Wang, B.2024b.HuatuoGPT- I, one-stage training for medical adaption of LLMs. https://arxiv.org/abs/2311.09774. Ches Ì nevar, C.; Mcginnis, J.; Modgil, S.; Rahwan, I.; Reed, C.; Simari, G.; South, M.; Vreeswijk, G.; and Willmott, S. 2006. Towards an argument interchange format. Knowledge Eng. Review 21:293â316. Christiano, P. F.; Leike, J.; Brown, T. B.; Martic, M.; Legg, S.; and Amodei, D. 2017. Deep reinforcement learning from human preferences. In Proceedings of the 31st Inter- national Conference on Neural Information Processing Sys- tems, NIPSâ17, 4302â4310. Red Hook, NY, USA: Curran Associates Inc. Davidson, R. R. 1970. On extending the Bradley-Terry Model to accommodate ties in paired comparison exper- iments.Journal of the American Statistical Association 65(329):317â328. DeepSeek-AI.2025a.DeepSeek-R1 incentivizes rea- soning in LLMs through reinforcement learning. Nature 645(8081):633â638. DeepSeek-AI.2025b.DeepSeek-V3 Technical Report. https://arxiv.org/abs/2412.19437. Diaconis, P., and Graham, R. L. 1977. Spearmanâs footrule as a measure of disarray. Journal of the Royal Statistical Society. Series B (Methodological) 39(2):262â268. Elangovan, A.; Xu, L.; Ko, J.; Elyasi, M.; Liu, L.; Bodapati, S. B.; and Roth, D. 2025. Beyond correlation: The impact of human uncertainty in measuring the effectiveness of au- tomatic evaluation and LLM-as-a-judge. In The Thirteenth International Conference on Learning Representations. Fu, J.; Ng, S.-K.; Jiang, Z.; and Liu, P. 2023. GPTScore: Evaluate as you desire. https://arxiv.org/abs/2302.04166. F Ì urnkranz, J., and H Ì ullermeier, E. 2003. Pairwise prefer- ence learning and ranking. In Lavra Ë c, N.; Gamberger, D.; Blockeel, H.; and Todorovski, L., eds., Machine Learning: ECML 2003, 145â156. Berlin, Heidelberg: Springer Berlin Heidelberg. Gemini Team.2024.Gemini 1.5: Unlocking multi- modal understanding across millions of tokens of context. https://arxiv.org/abs/2403.05530. Gemini Team.2025.Gemini 2.5: Pushing the fron- tier with advanced reasoning,multimodality,long context,andnextgenerationagenticcapabilities. https://arxiv.org/abs/2507.06261. Goffredo, P.; Haddadan, S.; Vorakitphan, V.; Cabrio, E.; and Villata, S. 2022. Fallacious argument classification in political debates. In Raedt, L. D., ed., Proceedings of the Thirty-First International Joint Conference on Artifi- cial Intelligence, IJCAI-22, 4143â4149. International Joint Conferences on Artificial Intelligence Organization. Main Track. Haddadan, S.; Cabrio, E.; and Villata, S. 2019. Yes, we can! mining arguments in 50 years of US presidential campaign debates. In Korhonen, A.; Traum, D.; and M ` arquez, L., eds., Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 4684â4690. Florence, Italy: Association for Computational Linguistics. Haldar, R., and Hockenmaier, J. 2025. Rating roulette: Self-inconsistency in LLM-as-a-judge frameworks.In Christodoulopoulos, C.; Chakraborty, T.; Rose, C.; and Peng, V., eds., Findings of the Association for Computa- tional Linguistics: EMNLP 2025, 24986â25004. Suzhou, China: Association for Computational Linguistics. Hautli-Janisz, A.; Kikteva, Z.; Siskou, W.; Gorska, K.; Becker, R.; and Reed, C. 2022. Qt30: A corpus of argu- ment and conflict in broadcast debate. In Proceedings of the 13th Language Resources and Evaluation Conference, 3291â3300. European Language Resources Association. Hayes, A. F., and Krippendorff, K. 2007. Answering the call for a standard reliability measure for coding data. Com- munication Methods and Measures 1(1):77â89. Huang, J., and Chang, K. C.-C. 2023. Towards reasoning in large language models: A survey. In Rogers, A.; Boyd- Graber, J.; and Okazaki, N., eds., Findings of the Associa- tion for Computational Linguistics: ACL 2023, 1049â1065. Toronto, Canada: Association for Computational Linguis- tics. Kendall, M. G., and Smith, B. B. 1940. On the method of paired comparisons. Biometrika 31(3-4):324â345. Kendall, M. G. 1938. A new measure of rank correlation. Biometrika 30(1/2):81â93. Krippendorff, K. 2004. Content Analysis: An Introduction to Its Methodology. Business & Economics. Sage. Lawrence, J., and Reed, C. 2019. Argument mining: A survey. Computational Linguistics 45(4):765â818. Le Mens, G., and Gallego, A. 2025. Positioning political texts with large language models by asking and averaging. Political Analysis 1â9. Likert, R. 1932. A Technique for the Measurement of Atti- tudes. Number nos. 136-165 in A Technique for the Mea- surement of Attitudes. Columbia university. Lin, C.-Y. 2004. ROUGE: A package for automatic evalu- ation of summaries. In Text Summarization Branches Out, 74â81. Barcelona, Spain: Association for Computational Linguistics. Lippi, M., and Torroni, P. 2016a. Argument mining from speech: Detecting claims in political debates. Proceedings of the AAAI Conference on Artificial Intelligence 30(1). Lippi, M., and Torroni, P. 2016b. Argumentation mining: State of the art and emerging trends. ACM Trans. Internet Technol. 16(2). Maystre, L., and Grossglauser, M. 2015. Fast and accurate inference of plackettâluce models. In Cortes, C.; Lawrence, N.; Lee, D.; Sugiyama, M.; and Garnett, R., eds., Advances in Neural Information Processing Systems, volume 28. Cur- ran Associates, Inc. Menini, S.; Cabrio, E.; Tonelli, S.; and Villata, S. 2018. Never retreat, never retract: Argumentation analysis for po- litical speeches. Proceedings of the AAAI Conference on Artificial Intelligence 32(1). Mestre, R.; Milicin, R.; Middleton, S. E.; Ryan, M.; Zhu, J.; and Norman, T. J. 2021. M-arg: Multimodal argument min- ing dataset for political debates with audio and transcripts. In Al-Khatib, K.; Hou, Y.; and Stede, M., eds., Proceedings of the 8th Workshop on Argument Mining, 78â88. Punta Cana, Dominican Republic: Association for Computational Linguistics. Meta AI.2024a.The llama 3 herd of models. https://arxiv.org/abs/2407.21783. Meta AI. 2024b. Llama 3.2. https://ai.meta.com/blog/llama- 3-2-connect-2024-vision-edge-mobile-devices/. Large lan- guage model accessed locally; model snapshot llama3.2:3b. MetaAI.2024c.Llama3.3. https://w.llama.com/docs/model-cards-and-prompt- formats/llama3 3/.Large language model accessed via Azureâs API; model snapshot llama3.3:70b. MetaAI.2025.Llama4maverick. https://ai.meta.com/blog/llama-4-multimodal-intelligence/. Large language model accessed via Azureâs API; model snapshot llama-4-maverick. Microsoft.2024.Phi-4technicalreport. https://arxiv.org/abs/2412.08905.Largelanguage model accessed via OpenRouterâs API; model snapshot microsoft/phi-4. MistralAI.2023.Mistral7b. https://arxiv.org/abs/2310.06825.Large language model accessed locally; model snapshot mistral:7b. Modgil, S., and Prakken, H. 2014. The ASPIC+ frame- work for structured argumentation: a tutorial. Argument & Computation 5(1):31â62. Mohammad, S.; Kiritchenko, S.; Sobhani, P.; Zhu, X.; and Cherry, C. 2016. SemEval-2016 task 6: Detecting stance in tweets. In Bethard, S.; Carpuat, M.; Cer, D.; Jurgens, D.; Nakov, P.; and Zesch, T., eds., Proceedings of the 10th International Workshop on Semantic Evaluation (SemEval- 2016), 31â41. San Diego, California: Association for Com- putational Linguistics. Mosteller, F., and Tukey, J. 1977. Data Analysis and Regres- sion: A Second Course in Statistics. Addison-Wesley se- ries in behavioral science. Addison-Wesley Publishing Com- pany. Negahban, S.; Oh, S.; and Shah, D. 2012. Iterative rank- ing from pair-wise comparisons. In Pereira, F.; Burges, C.; Bottou, L.; and Weinberger, K., eds., Advances in Neural Information Processing Systems, volume 25. Curran Asso- ciates, Inc. Neo4j. 2012. Neo4j - the worldâs leading graph database. http://neo4j.org/. OpenAI. 2024a. GPT-3.5 Turbo. Large language model accessed via Azureâs API to OpenAI; model snapshot gpt- 3.5-turbo-0125. OpenAI. 2024b. GPT-4 Turbo. Large language model ac- cessed via Azureâs API to OpenAI; model snapshot gpt-4- turbo-2024-04-09. OpenAI. 2024c. GPT-4o. Large language model accessed via Azureâs API to OpenAI; model snapshot gpt-4o-2024- 08-06. OpenAI. 2024d. GPT-4o Mini. Large language model ac- cessed via Azureâs API to OpenAI; model snapshot gpt-4o- mini-2024-07-18. OpenAI. 2025a. GPT-4.5 Preview. Preview large language model accessed via Azureâs API to OpenAI; model snapshot gpt-4.5-preview-2025-02-27. OpenAI. 2025b. o3-mini. https://openai.com/index/openai- o3-mini/. Large language model accessed via Azureâs API to OpenAI; model snapshot o3-mini-2025-01-31. Ornstein, J. T.; Blasingame, E. N.; and Truscott, J. S. 2025. How to train your stochastic parrot: large language models for political texts. Political Science Research and Methods 13(2):264â281. Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C. L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P.; Leike, J.; and Lowe, R. 2022. Training language models to follow in- structions with human feedback. In Proceedings of the 36th International Conference on Neural Information Processing Systems, NIPS â22. Red Hook, NY, USA: Curran Associates Inc. Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W.-J. 2002. Bleu: a method for automatic evaluation of machine trans- lation. In Isabelle, P.; Charniak, E.; and Lin, D., eds., Pro- ceedings of the 40th Annual Meeting of the Association for Computational Linguistics, 311â318. Philadelphia, Penn- sylvania, USA: Association for Computational Linguistics. Peldszus, A., and Stede, M. 2013. From Argument Dia- grams to Argumentation Mining in Texts: A Survey. Inter- national Journal of Cognitive Informatics and Natural Intel- ligence 7(1):1â31. Prakken, H. 2010. An abstract framework for argumenta- tion with structured arguments. Argument & Computation 1(2):93â124. Qwen Team. 2024. Qwq. https://qwen.ai/blog?id=qwq- 32b-preview. Large language model accessed via Open- Routerâs API; model snapshot qwq-32b. Qwen Team.2025.Qwen3 technical report. https://arxiv.org/abs/2505.09388.Large language model accessed via OpenRouterâs API; model snapshot qwen3- 235b-a22b. Rahwan, I., and Reed, C. 2009. The argument interchange format. In Argumentation in Artificial Intelligence, 383â 402. Shannon, C. E. 1948. A mathematical theory of commu- nication. The Bell System Technical Journal 27:379â423, 623â656. Sim, Y.; Acree, B. D. L.; Gross, J. H.; and Smith, N. A. 2013. Measuring ideological proportions in polit- ical speeches. In Yarowsky, D.; Baldwin, T.; Korhonen, A.; Livescu, K.; and Bethard, S., eds., Proceedings of the 2013 Conference on Empirical Methods in Natural Lan- guage Processing, 91â101. Seattle, Washington, USA: As- sociation for Computational Linguistics. Tarlow, K. R.; Brossart, D. F.; McCammon, A. M.; Gio- vanetti, A. J.; Belle, M. C.; and Philip, J. 2021. Reli- able visual analysis of single-case data: A comparison of rating, ranking, and pairwise methods. Cogent Psychology 8(1):1911076. Thurstone, L. L. 1927. A law of comparative judgment. Psychological Review 34(4):273â286. Visser, J.; Konat, B.; Duthie, R.; Koszowy, M.; Budzynska, K.; and Reed, C. 2020. Argumentation in the 2016 US pres- idential elections: annotated corpora of television debates and social media reaction. Language Resources and Evalu- ation 54(1):123â154. Wang, X.; Chen, G.; Dingjie, S.; Zhiyi, Z.; Chen, Z.; Xiao, Q.; Chen, J.; Jiang, F.; Li, J.; Wan, X.; Wang, B.; and Li, H. 2024. CMB: A comprehensive medical benchmark in Chinese. In Duh, K.; Gomez, H.; and Bethard, S., eds., Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 6184â6205. Mexico City, Mexico: Association for Compu- tational Linguistics. Wu, P. Y.; Nagler, J.; Tucker, J. A.; and Messing, S. 2023. Large language models can be used to estimate the latent positions of politicians. https://arxiv.org/abs/2303.12057. xAI. 2024. Grok 2. https://x.ai/news/grok-1212. Large lan- guage model accessed via Azureâs API to xAI; model snap- shot grok-2-1212. Yuan, W.; Neubig, G.; and Liu, P. 2021. BARTScore: Evaluatinggeneratedtextastextgeneration. https://arxiv.org/abs/2106.11520. Zhang, T.; Kishore, V.; Wu, F.; Weinberger, K. Q.; and Artzi, Y. 2020. Bertscore: Evaluating text generation with BERT. https://arxiv.org/abs/1904.09675. Zhang, H.; Chen, J.; Jiang, F.; Yu, F.; Chen, Z.; Li, J.; Chen, G.; Wu, X.; Zhang, Z.; Xiao, Q.; Wan, X.; Wang, B.; and Li, H. 2023. HuatuoGPT, towards taming language model to be a doctor. https://arxiv.org/abs/2305.15075. Zhao, W. X.; Zhou, K.; Li, J.; Tang, T.; Wang, X.; Hou, Y.; Min, Y.; Zhang, B.; Zhang, J.; Dong, Z.; Du, Y.; Yang, C.; Chen, Y.; Chen, Z.; Jiang, J.; Ren, R.; Li, Y.; Tang, X.; Liu, Z.; Liu, P.; Nie, J.-Y.; and Wen, J.-R. 2025. A survey of large language models. https://arxiv.org/abs/2303.18223.