Paper deep dive
StanceFlip: A Comprehensive Multi-Dimensional Benchmark for Multimodal Conversational Stance Flipping Forecasting
Heyan Chai, Xin Li, Wenjie Wang, Jianyang Qin, Chaoyang Li, Lu Wang, Hao Chen, Qing Liao
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Conversational stance detection has shifted from static text analysis to dynamic multimodal modeling. However, existing benchmarks exhibit three key limitations: failure to capture the dynamic evolution of beliefs, particularly during stance reversals; difficulty in disentangling affective states from logical reasoning; and neglect of the critical role of multimodal cues in resolving pragmatic ambiguities such as sarcasm. To address these limitations, we propose StanceFlip, a benchmark designed for multimodal conversational stance flipping forecasting over multi-turn dialogues across five modalities and multi-scenarios, which includes two novel subtasks: 1) Multimodal Stance Sextuple Extraction, extracting holder, target, emotion, sentiment, stance, and rationale as static state snapshots of dialogue to capture fine-grained cognitive structures. 2) Dynamic Stance Flip Attribution, tracking stance reversals across the conversation and identifying their underlying triggers. Alongside the dataset, we propose a dedicated framework, named ConStaFF, for Multimodal Conversational Stance Flipping Forecasting (MCSFF). Built upon a large language model, ConStaFF performs end-to-end stance reasoning, with a Thought-of-Stance (ToS) reasoning framework and a self-reflective verification mechanism integrated for structured stance modeling and faithful flip attribution. Specifically, ToS decomposes the reasoning process into specialized cognitive personas to formulate target propositions, resolve cross-modal conflicts, and infer historical stance trajectories. Extensive experiments show that our approach achieves state-of-the-art performance on both sextuple extraction and flip-trigger attribution, outperforming strong multimodal large language model baselines by substantial margins.
Tags
Links
- Source: https://arxiv.org/abs/2607.24191v1
- Canonical: https://arxiv.org/abs/2607.24191v1
Trouble viewing inline? Open PDF directly →
Full Text
102,614 characters extracted from source content.
Expand or collapse full text
StanceFlip: A Comprehensive Multi-Dimensional Benchmark for Multimodal Conversational Stance Flipping Forecasting Heyan Chai hychai@szu.edu.cn College of Computer Science and Software Engineering, Shenzhen University Shenzhen, China Xin Li 2024221014@mails.szu.edu.cn College of Computer Science and Software Engineering, Shenzhen University Shenzhen, China Wenjie Wang 2024280356@mails.szu.edu.cn College of Computer Science and Software Engineering, Shenzhen University Shenzhen, China Jianyang Qin hychai@szu.edu.cn Harbin Institute of Technology Shenzhen, China Chaoyang Li hychai@szu.edu.cn Harbin Institute of Technology Shenzhen, China Lu Wang wanglu@szu.edu.cn Shenzhen University Shenzhen, China Hao Chen sundaychenhao@gmail.com City University of Macau Macao SAR, China Qing Liao ∗ liaoqing@hit.edu.cn Harbin Institute of Technology Shenzhen, China Abstract Conversational stance detection has shifted from static text analysis to dynamic multimodal modeling. However, existing benchmarks exhibit three key limitations: failure to capture the dynamic evo- lution of beliefs, particularly during stance reversals; difficulty in disentangling affective states from logical reasoning; and neglect of the critical role of multimodal cues in resolving pragmatic am- biguities such as sarcasm. To address these limitations, we pro- pose StanceFlip, a benchmark designed for multimodal conversa- tional stance flipping forecasting over multi-turn dialogues across five modalities and multi-scenarios, which includes two novel sub- tasks: 1) Multimodal Stance Sextuple Extraction, extracting holder, target, emotion, sentiment, stance, and rationale as static state snapshots of dialogue to capture fine-grained cognitive structures. 2) Dynamic Stance Flip Attribution, tracking stance reversals across the conversation and identifying their underlying triggers. Alongside the dataset, we propose a dedicated framework, named ConStaFF, for Multimodal Conversational Stance Flipping Fore- casting (MCSFF). Built upon a large language model, ConStaFF performs end-to-end stance reasoning, with a Thought-of-Stance (ToS) reasoning framework and a self-reflective verification mech- anism integrated for structured stance modeling and faithful flip attribution. Specifically, ToS decomposes the reasoning process into specialized cognitive personas to formulate target propositions, re- solve cross-modal conflicts, and infer historical stance trajectories, while self-reflective verification iteratively refines generated ratio- nales against multimodal evidence to reduce causal hallucination. Extensive experiments show that our approach achieves state-of- the-art performance on both sextuple extraction and flip-trigger attribution, outperforming strong multimodal large language model baselines by substantial margins. ∗ Corresponding author: Qing Liao. Email: Liaoqing@hit.edu.cn.. CCS Concepts • Computing methodologies→ Artificial intelligence; Keywords Stance Detection, Multimodal Learning, Information Extraction 1 Introduction Stance detection aims to identify user’s attitude toward a specific target in discourse [33,40]. Early studies primarily focused on binary stance classification from isolated utterances [6,26,48]. The task has since evolved to multi-turn conversational settings, where opinions are expressed, negotiated, and sometimes reversed through interaction [11,24]. More recently, the rise of multimedia platforms has pushed stance analysis beyond text-only conversa- tions toward multimodal discourse, in which non-textual signals provide essential context for understanding speaker intent [34,46]. In such scenarios, multimodal signals, including images, audio, stickers, and video, provide essential context that text alone cannot convey, particularly when speakers use sarcasm, irony, or passive aggression to mask their true intent [27,34]. As a result, the field increasingly needs fine-grained structured representations that go beyond simple label assignment and instead support reasoning about why a speaker holds or changes a stance. Despite this progress, existing research still leaves three impor- tant gaps. First, current methods assign a single polarity label to each utterance without disentangling the holder’s affective expres- sion from their target-specific stance. As the dialogue in Figure 1(a) shows, Speaker S2 attaches a sarcastic sticker to express opposition despite the neutral surface text—since emotion and stance are differ- ent dimensions, one can happily oppose or angrily support. Without separate fields for emotion, sentiment, and stance anchored to the same holder-target pair, models inevitably conflate affective signals with logical stance. Second, most benchmarks treat non-textual in- formation as auxiliary features rather than decisive evidence [7,9]. arXiv:2607.24191v1 [cs.CL] 27 Jul 2026 Conference acronym ’X, June 03–05, 2018, Woodstock, NYTrovato et al. Figure 1: Illustration of the StanceFlip benchmark. In the same example, the video showing battery fire hazards is the factual trigger that causes Speaker S1 to reverse a previously firm stance, yet text-only or text-dominant pipelines cannot capture this cross-modal relationship [6,48]. Third, existing methods classify each turn independently and ignore the temporal continuity of stance [11,36]. Stance typically exhibits persistence and tends to remain stable once established unless strong evidence leads to a change. To the best of our knowledge, existing benchmarks neither model the full process of stance evolution nor identify the factors that trigger stance changes. To fill these gaps, we introduce StanceFlip, a large-scale bilingual benchmark for multimodal conversational stance flipping forecast- ing, comprising 7,710 annotated dialogues and 37,836 turns across five modalities and over 100 scenarios, with a flip density exceed- ing 20%. Built from real-world conversational seeds, StanceFlip enriches multi-turn, multi-party dialogues with multimodal cues through GPT-4o-based synthesis and grounded retrieval, followed by dual-consistency filtering and expert verification. To disentangle affective expression from argumentative stance, we define two com- plementary subtasks. Multimodal Stance Sextuple Extraction (subtask-I) extracts a panoptic sextuple(ℎ,푔,푒,푠,푠푡,푟)binding the holder, target, emotion, sentiment, stance, and rationale into a single record, decoupling affective expression from argumentative stance in Figure 1(b).Dynamic Stance Flip Attribution (subtask-I) tracks stance changes across conversations and attributes their underlying causes in Figure 1(c). Together, the two subtasks ad- vance stance detection from static label assignment to cognitively grounded causal reasoning. Compared with conventional stance detection, the proposed task is more challenging, requiring models to track stance dynamics over multi-turn contexts, resolve cross-modal conflicts, and iden- tify stance-change triggers. To address this, we propose ConStaFF, a dedicated framework built upon a multimodal large language model. To compensate for vanilla MLLMs’ lack of structured mech- anisms for target grounding, affective analysis, and temporal stance tracking, we design a novel Thought-of-Stance (ToS) reasoning framework that decomposes the task into four sequential cognitive steps handled by specialized expert personas, augmented with a self-reflective verification mechanism to reduce hallucinations and improve rationale faithfulness. Evaluations on StanceFlip show that ConStaFF consistently outperforms strong LLM baselines across both subtasks and languages. In summary, the contributions of this work are threefold: • We formalize Multimodal Conversational Stance Flip- ping Forecasting via two subtasks—Multimodal Stance Sextuple Extraction and Dynamic Stance Flip Attribution— advancing stance detection from static label assignment to cognitively grounded causal reasoning. •We contribute StanceFlip, a large-scale bilingual benchmark of 7,710 dialogues and 37,836 annotated turns across five modalities, featuring flip density exceeding 20%, expert-level fidelity, and diverse domain coverage. •We propose ConStaFF, an advanced framework equipped with the ToS reasoning framework and self-reflective verifi- cation mechanism, achieving state-of-the-art performance on both sextuple extraction and flip-trigger attribution, es- tablishing a strong baseline for future work. 2 Related Work Stance detection was initially studied in isolated settings, where the goal is to infer a speaker’s attitude toward a target from a single utterance or post [2,33,40,52]. Subsequent work extended this setting to more challenging variants, including multiple targets [39,48], zero-shot or unseen targets [3,4,26], and target extraction in open-world settings [22]. Despite these advances, the dominant formulation remains utterance-centric, treating stance as a local property of individual posts rather than a discourse-level state. Recent work has therefore shifted toward conversational and multi-turn stance analysis, where opinions are expressed, negoti- ated, and revised through interaction [11,24,36]. This setting intro- duces substantially richer interaction patterns, including speaker alternation, topic drift, and target switching, making coarse-grained formulations increasingly inadequate. As a result, stance under- standing requires finer-grained structure that explicitly anchors the holder, target, and supporting evidence across turns [9,30,53]. However, existing formulations still often fail to separate a holder’s affective state from their argumentative stance, which leads to systematic ambiguity in multi-party dialogue and weakens stance attribution across turns [11, 36]. In parallel, stance analysis has moved beyond text-only input to- ward multimodal settings [27,34,46]. This shift is crucial because non-textual signals can determine pragmatic meaning, override the polarity suggested by literal text, and reveal target-relevant evidence that remains implicit in language [15,27,28,34,46]. Mul- timodality is therefore important not only for recognizing stance, Multimodal Conversational Stance Flipping ForecastingConference acronym ’X, June 03–05, 2018, Woodstock, NY but also for explaining stance reversals. Although prior work has in- corporated discourse history and improved interpretability through logic or rationale generation [20,24,31,51,54], existing approaches still largely predict stance turn by turn rather than modeling it as a persistent state whose changes are triggered by new evi- dence. Related advances in adaptive conversational context model- ing, modality-context disentanglement, consequence forecasting, emotion-cause generation, and causal or counterfactual debiasing further suggest that multimodal interaction understanding benefits from explicit modeling of latent causes, confounders, and future consequences [8,18,21,41,42,45]. Against this backdrop, we in- troduce StanceFlip, a novel benchmark, and ConStaFF, reasoning framework that unifies fine-grained stance structure, multimodal grounding, and stance-flip modeling in a single task setting. Table 1 summarizes the key differences between StanceFlip and existing stance benchmarks. 3 StanceFlip Benchmark In this section, we present StanceFlip, a large-scale multimodal benchmark designed to study the dynamic evolution of stance in conversational discourse. We describe the automated construction pipeline and summarize main characteristics of resulting dataset. 3.1 Task Definition We formally define the StanceFlip task as tracking the fine-grained evolution of holder stance as a dynamic, persistent state rather than as a series of isolated classification points. Definition 3.1 (StanceFlip Task). Let퐷=푇 1 ,푇 2 , . . .,푇 푛 be a mul- timodal dialogue of푛turns. Each turn푇 푖 =(푢 푖 ,ℎ 푖 ,푎 푖 )consists of a textual utterance푢 푖 , a speakerℎ 푖 , and an associated multimodal in- formation푎 푖 ∈ 퐼 푖푚푔 ,퐼 푚푒푚푒 ,퐼 푎푢푑 ,퐼 푣푖푑 , N/A. Given a specific debate proposition퐺, which serves as the consistent target푔for each turn, the system tracks the evolution of stances 푠푡 for each participant. Stance Evolution Logic. Unlike traditional stance detection that treats utterances independently, our benchmark assumes a holder’s conviction is a continuous state. We establish the following transition principles for the stance label푠푡 푖 ∈ 푆,푁,푂,푈(where 푆,푁,푂,푈are definitive states, namely Support, Neutral, Oppose, and Unknown): •Stance Establishment: A user forms a clear stance (e.g., Support / Oppose) after previously having no explicit opinion. Note that the transition from Unknown to a definitive state is called "Establishment," not a "Flip." • Stance Persistence: Once a stance is formed, it remains un- changed across subsequent turns unless explicitly updated. Irrelevant or off-topic utterances inherit the previously ex- pressed stance 푠푡 푝푟푒푣 . •Stance Flip: A stance flip occurs when a user changes from one clear stance to another (e.g., from Support to Oppose), which is the main focus of our task. Drawing inspiration from panoptic sentiment analysis [30], we divide the task into two core subtasks. Subtask-I: Multimodal Stance Sextuple Extraction. This subtask extracts a stance sextu- ple푆 푖 = (ℎ,푔,푒,푠,푠푡,푟)for each turn푇 푖 . The elements include: (1) Holderℎand Target푔; (2) Emotion푒(7 classes) and Sentiment푠(3 classes); (3) Stance푠푡; and (4) Rationale푟. These sextuples provide state-aware snapshots of the conversation. Subtask-I: Dynamic Stance Flip Attribution. This subtask identifies turns where a stance flip occurs and attributes the flip to its causal trigger푟. Specif- ically, given a dialogue퐷=푇 1 ,푇 2 , . . .,푇 푛 , the task conducts two steps: (1)Flipping Forecasting. Identify turn푡where푠푡 푡 ≠ 푠푡 푡−1 and 푠푡 푡 ,푠푡 푡−1 ∈ 푆,푁,푂. (2)Trigger Categorization. Classify the trigger mechanism into one of four types: 1) Factual & Logical Appeals, 2) Emotional & Value-based Appeals, 3) Personal Experience & Anecdote, or 4) Social Influence. This task focuses on the temporal dynamics and dynamic reasoning of stance evolution, complement- ing Subtask-I’s element extraction with attribution analysis. 3.2 Automated Construction Pipeline To address the scarcity of multimodal dialogues with complex stance transitions, we propose a multi-stage simulation-retrieval pipeline that turns textual seeds into multimodal discourse. A detailed de- scription is provided in Appendix A of supplementary materials. Data Sourcing and Standardization. We curate textual seeds from three corpora: DailyDialog [23], MELD [37], and ZS-CSD [11]. Dialogues with 3–7 turns are selected to ensure sufficient context for stance evolution modeling, then mapped into a unified schema preserving speaker identities and conversation structure. Directed Multimodal Augmentation. GPT-4o [1] identifies critical turns (e.g., shifts or core arguments) as injection points via stance significance scoring. An online chat decision tree then aligns the medium with the holder’s intent: audio for paralinguistic cues, emojis/memes for emotional feedback, and images/videos for visual evidence. Multimodal Query Synthesis and Grounded Retrieval. GPT- 4o synthesizes queries emphasizing concrete sensory details, and generates cross-modal conflicts (e.g., pairing a compliment with an eye-rolling meme) to simulate irony and passive-aggression. Queries are encoded with SentenceTransformer [38] to retrieve from COCO [29], our proprietary sticker dataset, AudioSet [12], and WebVid [5], selecting candidates with similarity≥ 0.8. Evolution-aware Automated Labeling. GPT-4o performs a panoptic scan based on our evolution logic, prioritizing multimodal evidence to decode latent conviction under irony. This yields per- turn logic-based rationales with keywords, emotions, and sentiment labels. 3.3 Quality Assurance The reliability of StanceFlip rests on rigorous construction protocols and exhaustive manual verification. Construction Protocols and Dual-Consistency Filtering. Structural heuristics enforce contextually justified injection and discourse-level consistency. A cross-modal grounding filter (simi- larity≥0.8) and a stance confidence gate (self-assessed score≥0.9) ensure rationale-label alignment. Expert-level Manual Review and Refinement. To ensure data fidelity, the entire dataset underwent a comprehensive manual review and correction phase. Expert annotators inspected every dialogue thread to verify the coherence of stance trajectories, rele- vance of retrieved information, and accuracy of rationales. Errors in Conference acronym ’X, June 03–05, 2018, Woodstock, NYTrovato et al. Table 1: Systematic comparison of StanceFlip with existing stance detection benchmarks. Context & ModalityGranularity & ScopeCognitive Reasoning BenchmarkInteractionModalityLanguageTarget ScopeStructural Elements Stance State Machine State Rationale Causal Trigger SemEval-16 [33]Single-turnPure TextEN5 Targets(Target, Stance)✗ VAST [3]Single-turnPure TextENMultiple Targets(Target, Stance)✗ CTSDT [24]Multi-turnPure TextENSingle Target(Target, Stance)✗ MT2-CSD [35]Multi-turnPure TextENMultiple Targets(Target, Stance)✗ MmMtCSD [34]Multi-turnText + ImageEN, ZHMultiple Targets(Target, Stance)✗ MultiClimate [46]Multi-turnText + VideoENSingle Target(Target, Stance)✗ StanceFlipMulti-turnT+I+V+A+S*EN, ZHMultiple Targets Sextuple † Stance Attribution ‡ ✓ * Modality: T: Text, I: Image, V: Video, A: Audio, S: Stickers & Memes. † Sextuple: (Holder, Target, Sentiment, Emotion, Stance, Rationale); ‡ Stance-Flip Attribution: (Holder, Target, Stance, Flip, Trigger, Flip-Trig). Figure 2: The overall architecture of our ConStaFF model. Table 2: Detailed statistical profile of the StanceFlip benchmark. Lang. Scale and DepthEvo. LogicMultimodal Coverage (%) Dial. Turn. Av.Sext.Flip.Img. Mem. Aud. Vid.All EN3,898 21,534 5.5221,534 23.99%27.60 37.61 20.52 19.04 62.85 ZH3,812 16,302 4.2816,302 16.53%42.65 42.58 2.68 10.18 72.69 Tot/Avg 7,710 37,836 4.9137,836 20.30%35.04 40.06 11.70 14.66 67.72 automated labeling or cross-modal misalignments were manually corrected, ensuring the benchmark meets expert-level standards. Case Study Validation. A case study on 100 dialogues (∼600 turns) yielded Cohen’s휅=0.84 for stance identification and 0.87 for multimodal relevance [19]. Against human gold standard, auto- mated pipeline achieved 86.4% accuracy for stance detection and 88.5% for multimodal matching. 3.4 Data Highlights The StanceFlip benchmark is partitioned into train/validation/test splits at a ratio of 8:1:1, stratified by dialogue topic and language to ensure balanced target and flip-density distributions across splits. The statistics are shown in Table 2. Below we summarize the key characteristics and highlights of the StanceFlip dataset: Panoptic Granularity and Cognitive Inference. StanceFlip established a high standard for fine-grained analysis by extracting a complete sextupleS=(ℎ,푔,푒,푠,푠푡,푟)for every turn. We introduce natural language rationales that bridge multimodal cues with the holder’s convictions, making the reasoning path easier to trace. Holistic Lifecycle Tracking of Stance States. We formalize the stance lifecycle by tracking the trajectory from initiation to contextual inheritance and substantive inversion. This state-aware modeling captures the cumulative nature of discourse and turns stance detection into a dynamic sequence modeling task. Strategic Multimodal Synergy via Intent Alignment. Guided by an online chat decision tree, multimodal information is inserted based on communicative intents. By explicitly modeling polarity discrepancies (e.g., pairing compliments with ironic memes), we simulate social phenomena like irony and passive-aggression. High-Volatility Dynamics and Anti-Static Bias. To address the static bias in existing datasets, StanceFlip exhibits a flip density above 20%. This forces models to capture subtle turning points and triggers rather than simply fitting majority-class static labels. Bilingual Equilibrium with Expert-level Fidelity. Stance- Flip maintains a near 1:1 balance between English and Chinese. Combined with strict filtering and full manual verification, the benchmark achieves expert-level fidelity at scale. Multimodal Conversational Stance Flipping ForecastingConference acronym ’X, June 03–05, 2018, Woodstock, NY 4 Methodology In this section, we present ConStaFF, a comprehensive frame- work tailored for multimodal stance flipping analysis shown in Figure 2. Our framework is specifically designed to tackle the key challenges of this task, including complex conversational context understanding, multimodal information fusion, and cognitive-level stance flipping reasoning. We then introduce the overall architec- ture, the ToS reasoning framework, the self-reflective verification mechanism, and the multi-stage instruction tuning strategy. 4.1 ToS Reasoning Framework Standard Chain-of-Thought (CoT) reasoning [47] treats inference as a single linear chain, while Tree of Thoughts (ToT) [50] empha- sizes generic branching exploration. Neither directly addresses the core challenges of multimodal conversational stance-flip forecast- ing, namely consistent target grounding in multi-party dialogue, conflict resolution between textual and non-textual cues, discourse- level stance-state tracking, and evidence-grounded explanation. To address these challenges, we propose the Thought-of-Stance (ToS) framework, a stance-centered reasoning process tailored to our task. As shown in Figure 2, rather than expanding arbitrary thought branches, ToS decomposes reasoning into four ordered roles: a Cartographer for target proposition formulation, a Psy- chologist for multimodal affect grounding, a Discourse Analyst for stance-state transition modeling, and a Synthesizer–Critic pair for explanation construction and verification. The hierarchy in ToS is therefore explicit: reasoning proceeds from global dialogue anchor- ing, to turn-level multimodal interpretation, to cross-turn stance updating, and finally to evidence-grounded explanation. Step 1: The Cartographer (Target Proposition Formula- tion). To resolve target ambiguity caused by topic drift and speaker alternation, the model first acts as a “Cartographer” to perform a global scan of the dialogue and multimodal cues to formulate a debate propositionP. This proposition serves as the semantic anchor for all subsequent steps, ensuring that emotion, sentiment, stance, and flip triggers are inferred with respect to the same target rather than different local mentions. This step is formulated as: P ← 푓 cart (퐷,M | I 1 )(1) where퐷represents the dialogue history,Mdenotes multimodal features, andI 1 is the instruction. Step 1: Target Proposition Identification Input Data: Full dialogue text with captions. Instruction: Act as a [Cartographer]. Identify the central ‘Target’. Scan the entire dialogue to identify the main topic of contention. Convert this core issue into a clear, concise debate proposition sentence (e.g., “The debate concerning whether coffee should be tried as a substitute for cigarettes”). This proposition will serve as the consistent reference point. Output: (Target Proposition: [Debate Proposition Sentence]) Step 2: The Psychologist (Multimodal Conflict Resolution). To resolve pragmatic ambiguity, the model acts as a “Psychologist” to ground the holder’s affective state in multimodal evidence, and then to identify the discrepancy when textual polarity conflicts with non-textual cues or target-relevant media evidence, such as supportive wording paired with sarcastic prosody or a mocking sticker. Rather than treating multimodal signals as auxiliary fea- tures, this step uses them to determine the holder’s emotion and sentiment towardP. Step 2: Multimodal Affect Grounding Input Data: Dialogue history, current speaker, utterance with cues, core target. Instruction: Act as a [Psychologist] and [Emotion Ana- lyst]. Infer the holder’s emotion and sentiment toward the target by jointly considering text and multimodal cues. If the literal text conflicts with multimodal evidence, prefer the in- terpretation that best resolves the pragmatic ambiguity. Output: (Holder, Target, Emotion, Sentiment) This step can be formulated as: (ℎ 푖 ,푔 푖 ,푒 푖 ,푠 푖 ) ← 푓 psy (퐷,푢 푖 ,M 푖 ,P | I 2 )(2) where푢 푖 is the current utterance,M 푖 is the multimodal evidence of the current turn, and 푒 푖 and 푠 푖 denote emotion and sentiment. Step 3: Stance-State Transition Modeling Input Data: Dialogue history, current speaker, current emo- tion/sentiment, previous stance, and target proposition. Instruction: Act as a [Discourse Analyst]. Determine the holder’s current stance toward the target by jointly consid- ering dialogue history, current affective evidence, and the previous stance. Mark a flip only when the evidence supports a substantive change in target-specific position. Output: (Holder, Target, Emotion, Sentiment, Stance, Flip) Step 3: The Discourse Analyst (Stance Evolution). The model acts as a “Discourse Analyst” to determine whether the current turn reflects stance persistence or stance change. In ToS, stance is treated as a discourse state rather than a turn-local label. Emotion and sentiment are used as intermediate evidence, but the final stance is inferred by comparing the current turn with the holder’s prior stance state in context. A flip is detected only when the evidence supports a substantive revision of the holder’s position towardP. We formulate this step as: (ℎ 푖 ,푔 푖 ,푒 푖 ,푠 푖 ,푠푡 푖 , 푓 flip ) ← 푓 ana (퐷,ℎ 푖 ,푒 푖 ,푠 푖 ,푠푡 푝푟푒푣 ,P | I 3 )(3) where푠푡 푖 is the current stance,푠푡 푝푟푒푣 is the previous stance, and 푓 flip is the flip indicator. Step 4: Rationale and Trigger Generation Input Data: Dialogue history, multimodal evidence, and in- termediate results. Instruction: Act as a [Synthesizer]. Generate a concise ex- planation of why holder maintains or changes the stance, grounded in the most relevant textual and multimodal evi- dence. Identify the trigger that caused the transition when 푓 푓푙푖푝 = 1. Output: (Rationale, Trigger Type) Step 4: The Synthesizer (Rationale and Trigger Genera- tion). Finally, once the stance state is determined, the model acts a Conference acronym ’X, June 03–05, 2018, Woodstock, NYTrovato et al. “Synthesizer” to generate a rationale explaining why holder main- tains or changes stance, and identifies the trigger휏 푖 when a flip is detected (푓 푓푙푖푝 =1); otherwise (푓 푓푙푖푝 =0), only the rationale is gen- erated. This step summarizes the textual and multimodal evidence most relevant to the inferred stance transition, producing an initial explanation for subsequent verification. We formulate this as: (푟 푖 ,휏 푖 ) ← 푓 sys (퐷,ℎ 푖 ,푔 푖 ,푠 푖 ,푒 푖 ,푠푡 푝푟푒푣 ,푠푡 푖 , 푓 flip ,M 푖 | I 4 )(4) where 푟 푖 denotes the rationale and 휏 푖 denotes the trigger type. 4.2 Self-Reflective Verification Mechanism While ToS provides a structured process for target grounding, affec- tive modeling, stance-state tracking, and initial rationale generation, the resulting explanation for stance transitions are particularly vul- nerable to hallucination, as the model may generate plausible but unsupported reasons that are not fully grounded in the dialogue context or multimodal evidence. Motivation by [32], we introduce a novel self-reflective verification mechanism to improve explanation faithfulness, which verifies and refines the rationale generated by ToS rather than performing a new round of stance inference. Specifically, we adopt a Synthesizer–Critic loop. The Synthesizer first produces a draft rationale based on the predicted stance state and the available dialogue context, while the Critic checks whether this rationale is consistent with the target proposition, aligned with the inferred stance, and explicitly supported by the textual and multimodal evidence. If the rationale is vague, target-inconsistent, or unsupported by the evidence, the Critic feeds corrective feedback back to the Synthesizer for revision. Empirically, we find that the first refinement iteration yields the most significant improvement in rationale quality, while additional iterations provide diminishing returns at the cost of increased inference time. Therefore, we set 푁 max =1 in this work, and the loop terminates either when all three verification criteria inC 푖 are satisfied or upon reaching푁 max . Formally, the verification process is defined as: 푟 ∗ 푖 =Φ refine Φ syn (퐷 푖 ,M 푖 ,푠푡 푖 ) | z Draft Rationale ,C 푖 | I critic (5) whereΦ syn generates an initial rationale draft for the inferred stance transition, andΦ refine iteratively critiques and revises this draft for at most푁 max iterations. The constraint setC 푖 includes three criteria: 1)Target Consistency, requiring the rationale to remain grounded in the debate proposition; 2) Stance Consistency, requiring the explanation to support the predicted stance or stance reversal; and 3) Evidence Alignment, requiring explicit grounding in the dialogue content and, when applicable, the multimodal cues. The final output 푟 ∗ 푖 is a verified rationale that is more specific, evidence-grounded, and faithful to the stance prediction. 4.3 Multi-stage Instruction Tuning A single-stage end-to-end objective tends to entangle multimodal grounding, structured stance reasoning, and explanation verifica- tion. To better instantiate the ToS framework and self-reflective verification, we adopt a progressive three-stage tuning strategy that equips ConStaFF with multimodal perception, structured stance reasoning, and self-correction ability. Stage 1: Multimodal Alignment (Perception). We freeze the pre-trained ImageBind encoder [13] and train a lightweight linear projection layer on multimodal cue-description pairs from Stance- Flip’s training split, optimizing the language modeling loss on cap- tion generation. This maps multimodal representations into the LLM embedding space, providing grounded inputs for subsequent ToS reasoning. Stage 2: ToS Reasoning Tuning (Cognition). We fine-tune the model to follow the ordered ToS reasoning process (Proposition →Sentiment&Emotion→Stance→Rationale) using LoRA [16] while keeping ImageBind fixed: 푊=푊 0 +Δ푊,Δ푊= 퐵퐴,(6) where푊 0 is the frozen pre-trained weight,퐵 ∈ R 푑×푟 and퐴 ∈ R 푟×푘 are trainable low-rank matrices, and푟 ≪ min(푑,푘). This stage instills structured reasoning over target grounding, multimodal evidence interpretation, stance tracking, and rationale generation. Stage 3: Self-Reflection Tuning (Metacognition). To mitigate hallucinations in trigger identification, we introduce a dedicated self-reflection tuning strategy by constructing supervision tuples consisting of a draft rationale, critique feedback, and a revised ra- tionale. These tuples are synthesized by perturbing rationale drafts and training the model to first identify unsupported or inconsistent claims and then rewrite them into evidence-grounded explanations. Rather than introducing a new inference task, this stage directly trains the verification behavior described above and improves the faithfulness of rationale generation, especially for trigger identi- fication. This stage improves the model’s metacognitive ability to detect and revise reasoning errors, particularly benefiting the “Flip-Trigger” metric. 5 Experiments This section presents experiments to demonstrate the effectiveness of ConStaFF for stance flipping forecasting. We aim to answer the following research questions: •RQ1 (Comparative Experiment) Can proposed ConStaFF improve the stance flipping forecasting performance ? •RQ2 (Implicit Stance Experiment) How does ConStaFF per- form in the implicit stance scenario? •RQ3 (Zero-shot Experiment) How does ConStaFF perform in the zero-shot stance flipping prediction scenario? •RQ4 How does the proposed ToS reasoning framework con- tribute to the performance? •RQ5 How significant is the role of multimodal information? •RQ6 Does the self-reflective verification mechanism con- tribute to our model? 5.1 Experimental Settings Evaluations. We evaluate ConStaFF on StanceFlip across both subtasks. For Subtask-I, performance is assessed at three levels: item extraction, pair extraction, and full sextuple prediction. Target, Emotion, Sentiment, and Stance use F1-score; Rationale uses soft matching for semantic equivalence. A pair is correct only when both elements match, and a sextuple only when all components match; we report Micro-F1 on the complete sextuple and Identification F1 on the quadruple (Holder-Target-Stance-Flip). For Subtask-I, Multimodal Conversational Stance Flipping ForecastingConference acronym ’X, June 03–05, 2018, Woodstock, NY Table 3: Main comparative results on subtask-I, Multimodal Stance Sextuple Extraction. "T/E/Se/St/R" represents Target, Emotion, Sentiment, Stance, and Rationale, respectively. All the scores are averaged over five runs under different random seeds. ModelParam. ItemsPairsSext.Quad. HTESeStRT-ET-SeT-StSt-RMicro.Iden. English Vicuna7B96.5145.5545.2238.8448.9132.1226.7522.0532.3422.166.0418.69 Llama27B83.8628.2842.0341.6527.6312.4713.8813.509.3810.280.643.86 Llama38B93.6643.1459.6457.4244.6625.9629.5326.2622.7020.624.9011.28 Qwen2.57B95.5650.9357.9173.9456.1244.9930.4437.5628.9529.846.2422.72 Llama2 (ToS+self-reflection)7B99.9340.9261.3865.9751.7449.9626.0926.2424.1737.958.3020.46 Llama3 (ToS+self-reflection)8B99.3747.8163.5561.4767.5660.4329.9926.9733.8548.7011.4328.36 Flan-T5-XXL (ToS+self-reflection)11B99.6945.4363.9974.6859.6963.5531.0337.5631.6342.4614.8529.84 Qwen2.5 (ToS+self-reflection)7B96.7464.0366.6766.5262.5144.5444.6942.6142.7637.7112.1828.80 Mistral (ToS+self-reflection)7B99.2862.6668.8073.9456.1251.3745.1447.0736.8238.6013.3630.44 ConStaFF (Ours)7B 99.93 64.88 73.05 74.68 75.43 72.01 48.55 50.48 52.41 61.62 25.84 46.77 Chinese qwen2.57B10048.7647.7649.7039.8953.5624.0624.2020.6930.724.3729.29 Llama27B10023.0149.5151.4847.5329.9913.1814.3711.8720.831.328.50 Llama38B10026.9848.4462.6138.2155.3813.7417.8910.7327.982.0816.46 Vicuna (ToS+self-reflection)7B99.1559.4663.5571.9661.5451.7837.9244.2438.7837.2012.4228.87 Llama2 (ToS+self-reflection)7B99.6848.1168.1573.0347.9751.2032.4636.3423.7036.277.9725.35 Llama3 (ToS+self-reflection)8B99.6853.1454.9971.0249.2656.8833.9736.7725.9239.3510.4131.45 Mistral (ToS+self-reflection)7B99.6750.8474.8376.4860.9757.3837.8538.3532.6042.0114.1532.32 ConStaFF(Ours)7B100 64.85 74.90 78.20 61.54 69.08 49.26 50.84 41.01 51.85 21.69 44.31 Table 4: Main comparative results on subtask-I. DatasetModelParam.FlipTrigFlip-Trig English Vicuna7B17.4928.149.89 Llama27B19.235.775.77 Llama38B11.1115.877.94 Qwen2.57B16.4714.904.71 Llama2 (ToS+self-reflection)7B21.3732.8215.27 Llama3 (ToS+self-reflection)8B35.1033.4724.49 Flan-T5-XXL (ToS+self-reflection)11B35.5932.2025.42 Qwen2.5 (ToS+self-reflection)7B39.6625.8625.86 Mistral (ToS+self-reflection)7B40.0025.4525.45 ConStaFF(Ours)7B 44.26 34.04 26.38 Chinese Qwen2.57B13.9516.2811.63 Llama27B17.4617.469.52 Llama38B12.209.769.76 Vicuna (ToS+self-reflection)7B40.0030.3423.45 Llama2(ToS+self-reflection)7B22.8624.7619.05 Llama3(ToS+self-reflection)8B22.2235.5916.30 Mistral(ToS+self-reflection)7B35.9417.1915.62 ConStaFF(Ours)7B 44.30 33.56 26.85 Ours Qwen 7B Llama 3 GPT 3.5 GPT-4o Mini C-3 Haiku 55 60 65 70 F1 (%) (a) Stance Ours Qwen 7B Llama 3 C-3.5 Haiku GPT 3.5 GPT 4o GPT-4o Mini C-3 Haiku 0 20 40 60 80 F1 (%) (b) Flip Ours Qwen 7B Llama 3 C-3.5 Haiku GPT 3.5 GPT 4o GPT-4o Mini C-3 Haiku 20 40 60 F1 (%) (c) Trigger Ours Qwen 7B Llama 3 C-3.5 Haiku GPT 3.5 GPT 4o GPT-4o Mini C-3 Haiku 0 20 40 60 F1 (%) (d) Flip-Trigger Figure 3: Performance evaluation on implicit stance datasets. three F1-level metrics are reported: initial/flipped stance (Flip), trig- ger category (Trig), and flipped stance with trigger (Flip-Trig). All results are averaged over five runs with different random seeds. Baselines and Implementation. We compare ConStaFF against strong MLLM baselines of different scales, including Flan-T5-XXL (11B) [10], Vicuna (7B), Llama-2 (7B) [44], Llama-3 (8B) [43], Mistral (7B) [17], and Qwen2.5 (7B) [49]. For fair comparison, ConStaFF adopts Vicuna-7B for English and Qwen2.5-7B [49] for Chinese, both fine-tuned with LoRA under the ToS and self-reflection set- ting. The experiments were conducted on high-performance device with 8*Nvidia RTX A6000 GPUs. To ensure the reliability and re- producibility of our experiment, all results are averaged over five runs with different random seeds. Due to the space limitation, we provide the detailed description of model details and experimental setting in Appendix C of supplementary materials. 5.2 Main Results (RQ1) Performance on Multimodal Stance Sextuple Extraction task. Table 3 shows that ConStaFF delivers the strongest performance on Multimodal Stance Sextuple Extraction in both English and Chinese. The largest gains appear at the structure level: on English, Con- StaFF improves Micro-F1 and Iden. to 25.84% and 46.77%, clearly above the strongest baselines (14.85% and 30.44%); on Chinese, the corresponding scores reach 21.69% and 44.31%, again establishing a clear margin over prior systems. The improvements are also consis- tent on difficult aspects such as Stance, Rationale, and the St-R pair, suggesting that our method enhances not only element recognition but also cross-element coherence within the full sextuple. More broadly, ToS+self-reflection variants consistently outperform their vanilla backbones, highlighting the effectiveness of our ConStaFF on structured reasoning for fine-grained stance modeling . Performance on Dynamic Stance Flip Attribution task. Table 4 reports the results on Dynamic Stance Flip Attribution. ConStaFF achieves the best Flip and Flip-Trig scores in both languages, reach- ing 44.26%/26.38% on English and 44.30%/26.85% on Chinese. It also obtains the best Trig score on English and remains competitive on Chinese, where the strongest trigger-only baseline still falls short Conference acronym ’X, June 03–05, 2018, Woodstock, NYTrovato et al. GPT 3.5 Deep Seek GPT 4 Join CL Ours KE Pro KPT TTS Dou bao 50 55 60 F1 (%) (a) Bitcoin GPT 3.5 Deep Seek GPT 4 Join CL Ours KE Pro KPT TTS Dou bao 50 55 60 F1 (%) (b) SpaceX GPT 3.5 Deep Seek GPT 4 Join CL Ours KE Pro KPT TTS Dou bao 30 40 50 60 F1 (%) (c) Bar_Trump GPT 3.5 Deep Seek GPT 4 Join CL Ours KE Pro KPT TTS Dou bao 45 50 55 60 F1 (%) (d) Average Figure 4: Evaluation of the generalization to unseen targets. ToSCoSToT 30 40 50 Iden (F1) ToSCoSToT 30 40 Flip (F1) ToSCoSToT 20 25 30 35 Trig (F1) ToSCoSToT 15 20 25 30 Micro (F1) Figure 5: Evaluation of effectiveness of our ToS reasoning strategy. on the joint Flip-Trig metric. This indicates that ConStaFF is more effective at jointly modeling stance reversal and its underlying trig- ger, rather than optimizing the two objectives independently. At the same time, the relatively low joint scores across all systems confirm that flip attribution remains substantially more challenging than static stance prediction. 5.3 Performance on Implicit Stance (RQ2) Figure 3 reports results on implicit stance cases involving sarcasm and visual metaphor, which require both correct stance recognition and evidence-grounded reasoning about why a stance flip occurs. A clear pattern is that strong LLM baselines remain competitive on the easier prediction targets, e.g., GPT-4o achieves 54.55% on Flip, but their performance drops sharply on Trigger and especially on Flip-Trigger, suggesting that they often infer the outcome without accurately identifying its underlying cause. In contrast, ConStaFF achieves the best performance on all four metrics. The gains are especially pronounced on the reasoning-intensive metrics, where ConStaFF surpasses the strongest baseline by 10.18% on Trigger and 19.16% on Flip-Trigger. These results indicate that ConStaFF is more robust to implicit pragmatic cues and better aligns stance prediction with grounded trigger identification. 5.4Zero-Shot Generalization to Unseen Targets To evaluate zero-shot generalization, we test ConStaFF on the MT-CSD dataset [36] across three unseen target domains (Bitcoin, SpaceX, and Trump), where ConStaFF is trained solely on Stance- Flip without any exposure to MT-CSD during training. As shown in Figure 4, ConStaFF consistently ranks first across all three do- mains, achieving 56.26% on Bitcoin, 59.07% on SpaceX, and 49.07% on Trump, with the best average F1 of 54.80%. It outperforms the w/o stick w/o audio w/o video w/o image Ours 42 43 44 45 46 47 48 Iden (F1) (a) Iden (F1) w/o stick w/o audio w/o video w/o image Ours 24.0 24.5 25.0 25.5 26.0 26.5 Micro (F1) (b) Micro (F1) w/o stick w/o audio w/o video w/o image Ours 22 23 24 25 26 27 28 Flip-Trig (F1) (c) Flip-Trig (F1) Figure 6: Evaluation of the contribution of each modality. MicroIdenFlipTrigFlip-Trig 0 10 20 30 40 50 F1 (%) ToS ToS+self-reflection Figure 7: Five-Dimensional Comparison: ToS (w/o Self-Reflection) vs ToS+Self-Reflection. strongest baseline Doubao-Pro [14] by 2.58% and the best transfer- learning method TTS [25] by 4.66%. We attribute this to the ToS reasoning framework, which captures target-invariant stance evolu- tion patterns rather than topic-specific lexical cues, enabling robust generalization to unseen targets. 5.5 Effectiveness of ToS Reasoning (RQ4) Figure 5 compares ToS with CoS and ToT reasoning strategy un- der the same setting. ToS consistently performs best on all four metrics, including Iden. (46.77%), Micro-F1 (25.84%), Flip (44.26%), and Trig (34.04%). The gains are especially clear on the more struc- tured and reasoning-intensive metrics, where ToS substantially outperforms both alternatives. These results show that the staged role-based design of ToS provides a more effective reasoning pro- cess for multimodal stance modeling and stance-flip analysis than sentiment-centered or generic tree-style prompting. 5.6 Impact of Multimodal Information (RQ5) Figure 6 evaluates the contribution of each modality. The full Con- StaFF model achieves the best performance on all metrics, reaching 46.77% on Iden., 25.84% on Micro-F1, and 26.38% on Flip-Trig, con- firming the benefit of joint multimodal modeling. Removing any modality leads to consistent degradation, but the impact is not uniform. Audio removal causes the largest drop on Iden. (42.76%), suggesting that paralinguistic cues are important for fine-grained stance understanding, while removing video yields the lowest Flip- Trig score (22.58%), indicating its particular value for identifying stance-reversal triggers. These results show that different modalities contribute complementary evidence, and that multimodal fusion is critical for both structural stance extraction and flip attribution. Multimodal Conversational Stance Flipping ForecastingConference acronym ’X, June 03–05, 2018, Woodstock, NY 5.7 Impact of Verification Mechanism (RQ6) Figure 7 compares ToS with its full version augmented by self- reflection. Our proposed self-reflective verification mechanism con- sistently improves all five metrics. The largest improvements are observed in Micro and Iden., indicating that verification is par- ticularly beneficial for structured prediction, while the consistent improvements on Flip, Trig, and Flip-Trig further show its value in refining evidence-grounded stance-flip reasoning. These results confirm that self-reflective verification mechanism improves both prediction reliability and explanation faithfulness. 6 Conclusion This paper introduces multimodal conversational stance flipping forecasting, a novel task that advances stance detection by formaliz- ing the cognitive evolution trajectory of human stance. It comprises two subtasks: (1) Multimodal Stance Sextuple Extraction, which provides a state-aware snapshot of each holder’s conviction per turn through a structured record of holder, target, emotion, senti- ment, stance, and rationale; and (2) Dynamic Stance Flip Attribution, which identifies the specific socio-cognitive trigger behind each belief reversal. We benchmark this novel setting with StanceFlip, a large-scale bilingual dataset built via a human-AI collaborative pipeline, covering five modalities, multi-turn multi-party contexts, high flip density, and expert-level annotation across diverse do- mains. We further propose ConStaFF, a novel reasoning framework built on the thought-of-stance architecture and a self-reflective verification mechanism. By decomposing the complex reasoning process into specialized cognitive sub-tasks and mitigating causal hallucination through iterative rationale refinement, ConStaFF pro- vides a strong baseline for future research on StanceFlip. References [1]Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al.2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023). [2] Abeer AlDayel and Walid Magdy. 2020. Stance Detection on Social Media: State of the Art and Trends. CoRR abs/2006.03644 (2020). arXiv:2006.03644 [3] Emily Allaway and Kathleen R. McKeown. 2020. Zero-Shot Stance Detection: A Dataset and Model using Generalized Topic Representations. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics, 8913–8931. [4]Emily Allaway, Malavika Srikanth, and Kathleen McKeown. 2021. Adversarial Learning for Zero-Shot Stance Detection on Social Media. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computa- tional Linguistics: Human Language Technologies. Association for Computational Linguistics, 4756–4767. [5]Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. 2021. Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021, Montreal, QC, Canada, October 10-17, 2021. IEEE, Montreal, QC, Canada, 1708–1718. [6]Heyan Chai, Jinhao Cui, Siyu Tang, Ye Ding, Xinwang Liu, Binxing Fang, and Qing Liao. 2025. MG-SIN: Multigraph Sparse Interaction Network for Multitask Stance Detection. IEEE Trans. Neural Networks Learn. Syst. 36, 2 (2025), 3111–3125. [7]Heyan Chai, Siyu Tang, Jinhao Cui, Ye Ding, Binxing Fang, and Qing Liao. 2022. Improving Multi-task Stance Detection with Multi-task Interaction Network. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP. Association for Computational Linguistics, Abu Dhabi, United Arab Emirates, 2990–3000. [8]Feiyu Chen, Zhengxiao Sun, Deqiang Ouyang, Xueliang Liu, and Jie Shao. 2021. Learning What and When to Drop: Adaptive Multimodal and Contextual Dy- namics for Emotion Recognition in Conversation. In Proceedings of the ACM Multimedia Conference, M. ACM, Virtual Event, China, 1064–1073. [9]Zhanpeng Chen, Zhihong Zhu, Wanshi Xu, Yunyan Zhang, Xian Wu, and Yefeng Zheng. 2024. Aspects are Anchors: Towards Multimodal Aspect-based Sentiment Analysis via Aspect-driven Alignment and Refinement. In Proceedings of the 32nd ACM International Conference on Multimedia, M 2024, Melbourne, VIC, Australia, 28 October 2024 - 1 November 2024. ACM, Melbourne, VIC, Australia, 2292–2300. [10]Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al.2024. Scaling instruction-finetuned language models. Journal of Machine Learning Research 25, 70 (2024), 1–53. [11]Yuzhe Ding, Kang He, Bobo Li, Li Zheng, Haijun He, Fei Li, Chong Teng, and Donghong Ji. 2025. Zero-Shot Conversational Stance Detection: Dataset and Approaches. In Findings of the Association for Computational Linguistics, ACL 2025, Vienna, Austria, July 27 - August 1, 2025 (Findings of ACL, Vol. ACL 2025). Association for Computational Linguistics, Vienna, Austria, 3221–3235. [12] Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter. 2017. Au- dio Set: An ontology and human-labeled dataset for audio events. In 2017 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2017, New Orleans, LA, USA, March 5-9, 2017. IEEE, New Orleans, LA, USA, 776–780. [13] Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra. 2023. ImageBind One Embedding Space to Bind Them All. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023. IEEE, Vancouver, BC, Canada, 15180–15190. [14]Dong Guo, Faming Wu, Feida Zhu, Fuxing Leng, Guang Shi, Haobin Chen, Haoqi Fan, Jian Wang, Jianyu Jiang, Jiawei Wang, et al.2025. Seed1. 5-vl technical report. arXiv preprint arXiv:2505.07062 (2025). [15] Devamanyu Hazarika, Roger Zimmermann, and Soujanya Poria. 2020. MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Anal- ysis. In Proceedings of The 28th ACM International Conference on Multimedia,. ACM, Virtual Event / Seattle, WA, USA, 1122–1131. [16] Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In The Tenth International Conference on Learning Representa- tions, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net, Virtual Event. [17]Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, De- vendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023. Mistral 7B. arXiv preprint arXiv.2310.06825 (2023). [18] Xincheng Ju, Dong Zhang, Suyang Zhu, Junhui Li, Shoushan Li, and Guodong Zhou. 2024. ECFCON: Emotion Consequence Forecasting in Conversations. In Proceedings of the 32nd ACM International Conference on Multimedia, M 2024. ACM, Melbourne, VIC, Australia, 2233–2241. [19]J Richard Landis and Gary G Koch. 1977. The measurement of observer agreement for categorical data. Biometrics (1977), 159–174. [20]Woojin Lee, Jaewook Lee, and Harksoo Kim. 2024. LOGIC: LLM-originated guidance for internal cognitive improvement of small language models in stance detection. PeerJ Comput. Sci. 10 (2024), e2585. [21]Bobo Li, Hao Fei, Lizi Liao, Yu Zhao, Chong Teng, Tat-Seng Chua, Donghong Ji, and Fei Li. 2023. Revisiting Disentanglement and Fusion on Modality and Context in Conversational Multimodal Emotion Recognition. In Proceedings of the 31st ACM International Conference on Multimedia, M. ACM, Ottawa, ON, Canada, 5923–5934. [22]Yingjie Li, Krishna Garg, and Cornelia Caragea. 2023. A New Direction in Stance Detection: Target-Stance Extraction in the Wild. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, July 9-14, 2023. Association for Computational Linguistics, Toronto, Canada, 10071–10085. [23]Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017. DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset. In Proceedings of the Eighth International Joint Conference on Natural Language Processing, IJCNLP 2017, Taipei, Taiwan, November 27 - December 1, 2017 - Volume 1: Long Papers. Asian Federation of Natural Language Processing, Taipei, Taiwan, 986–995. [24]Yupeng Li, Dacheng Wen, Haorui He, Jianxiong Guo, Xuan Ning, and Francis C. M. Lau. 2023. Contextual Target-Specific Stance Detection on Twitter: Dataset and Method. In IEEE International Conference on Data Mining, ICDM, December 1-4, 2023. IEEE, Shanghai, China, 359–367. [25]Yingjie Li, Chenye Zhao, and Cornelia Caragea. 2023. TTS: A Target-based Teacher-Student Framework for Zero-Shot Stance Detection. In Proceedings of the ACM Web Conference 2023, W 2023, Austin, TX, USA, 30 April 2023 - 4 May 2023. ACM, Austin, TX, USA, 1500–1509. [26]Bin Liang, Zixiao Chen, Lin Gui, Yulan He, Min Yang, and Ruifeng Xu. 2022. Zero-Shot Stance Detection via Contrastive Learning. In W ’22: The ACM Web Conference 2022, Virtual Event, Lyon, France, April 25 - 29, 2022. ACM, Virtual Event, Lyon, France, 2738–2747. [27]Bin Liang, Ang Li, Jingqian Zhao, Lin Gui, Min Yang, Yue Yu, Kam-Fai Wong, and Ruifeng Xu. 2024. Multi-modal Stance Detection: New Datasets and Model. In Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11-16, 2024 (Findings of ACL, Vol. ACL Conference acronym ’X, June 03–05, 2018, Woodstock, NYTrovato et al. 2024). Association for Computational Linguistics, Bangkok, Thailand and virtual meeting, 12373–12387. [28]Bin Liang, Chenwei Lou, Xiang Li, Lin Gui, Min Yang, and Ruifeng Xu. 2021. Multi-Modal Sarcasm Detection with Interactive In-Modal and Cross-Modal Graphs. In Proceedings of the ACM Multimedia Conference, October 20 - 24, 2021. ACM, Virtual Event, China, 4707–4715. [29] Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014. Microsoft COCO: Common Objects in Context. In Computer Vision - ECCV 2014 - 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V (Lecture Notes in Computer Science, Vol. 8693). Springer, Zurich, Switzerland, 740–755. [30]Meng Luo, Hao Fei, Bobo Li, Shengqiong Wu, Qian Liu, Soujanya Poria, Erik Cambria, Mong-Li Lee, and Wynne Hsu. 2024. PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment Analysis. In Proceedings of the 32nd ACM International Conference on Multimedia, M 2024, Melbourne, VIC, Australia, 28 October 2024 - 1 November 2024. ACM, Melbourne, VIC, Australia, 7667–7676. [31] Junxia Ma, Changjiang Wang, Hanwen Xing, Dongming Zhao, and Yazhou Zhang. 2024. Chain of Stance: Stance Detection with Large Language Models. In Natural Language Processing and Chinese Computing - 13th National CCF Conference, NLPCC (Lecture Notes in Computer Science, Vol. 15363). Springer, Hangzhou, China, 82–94. [32]Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark. 2023. Self-Refine: Iterative Refinement with Self-Feedback. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS. New Orleans, LA, USA. [33] Saif M. Mohammad, Svetlana Kiritchenko, Parinaz Sobhani, Xiaodan Zhu, and Colin Cherry. 2016. SemEval-2016 Task 6: Detecting Stance in Tweets. In Pro- ceedings of the 10th International Workshop on Semantic Evaluation (SemEval- 2016). Association for Computational Linguistics, San Diego, California, 31–41. doi:10.18653/v1/S16-1003 [34] Fuqiang Niu, Zebang Cheng, Xianghua Fu, Xiaojiang Peng, Genan Dai, Yin Chen, Hu Huang, and Bowen Zhang. 2024. Multimodal Multi-turn Conversation Stance Detection: A Challenge Dataset and Effective Model. In Proceedings of the 32nd ACM International Conference on Multimedia, M 2024, 28 October 2024 - 1 November 2024. ACM, Melbourne, VIC, Australia, 3867–3876. [35]Fuqiang Niu, Genan Dai, Yisha Lu, Jiayu Liao, Xiang Li, Hu Huang, and Bowen Zhang. 2025. MT2-CSD: A New Dataset and Multi-Semantic Knowledge Fu- sion Method for Conversational Stance Detection. CoRR abs/2506.21053 (2025). arXiv:2506.21053 [36] Fuqiang Niu, Min Yang, Ang Li, Baoquan Zhang, Xiaojiang Peng, and Bowen Zhang. 2024. A Challenge Dataset and Effective Models for Conversational Stance Detection. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC/COLING. ELRA and ICCL, Torino, Italy, 122–132. [37] Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik, Erik Cambria, and Rada Mihalcea. 2019. MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers. Association for Computational Linguistics, Florence, Italy, 527–536. [38]Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, No- vember 3-7, 2019. Association for Computational Linguistics, Hong Kong, China, 3980–3990. [39]Parinaz Sobhani, Diana Inkpen, and Xiaodan Zhu. 2017. A Dataset for Multi- Target Stance Detection. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2017, April 3-7, 2017, Volume 2: Short Papers. Association for Computational Linguistics, Valencia, Spain, 551–557. [40]Dhanya Sridhar, James R. Foulds, Bert Huang, Lise Getoor, and Marilyn A. Walker. 2015. Joint Models of Disagreement and Stance in Online Debate. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing of the Asian Federation of Natural Language Processing, ACL 2015, July 26-31, Volume 1: Long Papers. The Association for Computer Linguistics, Beijing, China, 116–125. [41]Teng Sun, Juntong Ni, Wenjie Wang, Liqiang Jing, Yinwei Wei, and Liqiang Nie. 2023. General Debiasing for Multimodal Sentiment Analysis. In Proceedings of the 31st ACM International Conference on Multimedia, M. ACM, Ottawa, ON, Canada, 5861–5869. [42]Teng Sun, Wenjie Wang, Liqiang Jing, Yiran Cui, Xuemeng Song, and Liqiang Nie. 2022. Counterfactual Reasoning for Out-of-distribution Multimodal Sentiment Analysis. In Proceedings of The 30th ACM International Conference on Multimedia, October 10 - 14, 2022. ACM, Lisboa, Portugal, 15–23. [43]Llama Team. 2024. The Llama 3 Herd of Models. CoRR abs/2407.21783 (2024). arXiv:2407.21783 doi:10.48550/ARXIV.2407.21783 [44] Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yas- mine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhos- ale, Dan Bikel, Lukas Blecher, Cristian Canton-Ferrer, Moya Chen, Guillem Cucu- rull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurélien Ro- driguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open Foundation and Fine-Tuned Chat Models. CoRR abs/2307.09288 (2023). arXiv:2307.09288 doi:10.48550/ARXIV.2307.09288 [45]Fanfan Wang, Heqing Ma, Xiangqing Shen, Jianfei Yu, and Rui Xia. 2024. Observe before Generate: Emotion-Cause aware Video Caption for Multimodal Emotion Cause Generation in Conversations. In Proceedings of the 32nd ACM International Conference on Multimedia, M. ACM, Melbourne, VIC, Australia, 5820–5828. [46]Jiawen Wang, Longfei Zuo, Siyao Peng, and Barbara Plank. 2024. MultiClimate: Multimodal Stance Detection on Climate Change Videos. In Proceedings of the Third Workshop on NLP for Positive Impact. Association for Computational Lin- guistics, Miami, Florida, USA, 315–326. [47] Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022. Curran Associates, Inc., New Orleans, LA, USA, 24824–24837. [48]Penghui Wei, Junjie Lin, and Wenji Mao. 2018. Multi-Target Stance Detection via a Dynamic Memory-Augmented Network. In The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval, SIGIR 2018, July 08-12, 2018. ACM, Ann Arbor, MI, USA, 1229–1232. [49]An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li, Tingyu Xia, Xingzhang Ren, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yu Wan, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, and Zihan Qiu. 2024. Qwen2.5 Technical Report. CoRR abs/2412.15115 (2024). arXiv:2412.15115 doi:10.48550/ARXIV.2412.15115 [50] Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023. Curran Associates, Inc., New Orleans, LA, USA, 11809–11822. [51]Jiaqing Yuan, Ruijie Xi, and Munindar P. Singh. 2025. Reasoner Outperforms: Generative Stance Detection with Rationalization for Social Media. In Proceedings of the 36th ACM Conference on Hypertext and Social Media, HT, Chicago, IL, USA, September 15-18. ACM, Chicago, IL, USA, 28–32. [52] Guido Zarrella and Amy Marsh. 2016. MITRE at SemEval-2016 Task 6: Transfer Learning for Stance Detection. In Proceedings of the 10th International Workshop on Semantic Evaluation, SemEval@NAACL-HLT , June 16-17, 2016. The Association for Computer Linguistics, San Diego, CA, USA, 458–463. [53]Changmeng Zheng, Junhao Feng, Ze Fu, Yi Cai, Qing Li, and Tao Wang. 2021. Multimodal Relation Extraction with Efficient Graph Alignment. In Proceedings of the ACM Multimedia Conference, M. ACM, Virtual Event, China, 5298–5306. [54]Arkaitz Zubiaga, Elena Kochkina, Maria Liakata, Rob Procter, Michal Lukasik, Kalina Bontcheva, Trevor Cohn, and Isabelle Augenstein. 2017. Discourse-Aware Rumour Stance Classification in Social Media Using Sequential Classifiers. CoRR abs/1712.02223 (2017). arXiv:1712.02223 A More Details of Dataset A.1 Multi-stage Construction Pipeline Seed Data Selection and Pre-processing To construct a bench- mark that accurately reflects the stance evolution in online dis- course, we implemented a rigorous data curation pipeline targeting high-quality textual seeds. Multimodal Conversational Stance Flipping ForecastingConference acronym ’X, June 03–05, 2018, Woodstock, NY Source Corpus Integration and Rationale. We aggregate raw dia- logues from three sources: DailyDialog, MELD, and ZS-CSD. These datasets were chosen not only because they are conversational, but because they contain emotional conflict and argumentative density, both of which are useful for observable stance dynamics. Unlike generic chitchat corpora, they provide enough friction for opinion shifts to emerge naturally. Heuristic Filtering for Stance Dynamics. We find that stance evo- lution requires a specific temporal window to become visible. Dia- logues shorter than 3 turns often lack enough context for a meaning- ful shift, whereas dialogues longer than 7 turns in existing corpora often suffer from topic drift or diluted context. We therefore apply a strict length filter of 3–7 turns. To keep the dataset rich in argu- mentative content, we also prioritize dialogues containing explicit discourse markers through keyword matching. Speaker and Format Normalization. Raw data from heteroge- neous sources often use inconsistent schemas. We standardize all interlocutor labels into unified identifiers and normalize the conver- sation structure into a canonical JSON format. This pre-processing step reduces parsing errors in the later LLM-based simulation stage. Directed Multimodal Synthesis To transform textual dia- logues into more realistic multimodal online chats, we design a Directed Multimodal Synthesis pipeline. Rather than relying on ran- dom augmentation, the pipeline uses a multi-step reasoning strat- egy with GPT-4o. The detailed instructions that guide this analysis and generation process are shown in Table 5 and Table 6. Step 1: Stance Significance Scoring and Pivot Identification. As shown in Table 5, we use a significance-driven injection strategy rather than a random one. The model first evaluates the full dialogue and assigns a Stance Significance Score (0.0–1.0). Only dialogues above a confidence threshold of 0.7 proceed to injection. To keep multimodal elements relevant to stance detection, the model selects insertion points according to a fixed priority order, giving prefer- ence to stance shifts and stance fortification over generic chatter. This helps ensure that each inserted item carries meaningful infor- mation about the speaker’s conviction. Step 2: Context-Aware Modality Selection (Decision Tree). After identifying a target turn, the model follows the Enhanced Online Chat Decision Tree (Table 5) to select the most appropriate modality. This logic distinguishes between visual evidence, such as screen- shots used in argumentation, paralinguistic sounds such as laughter or sighs, and reactive imagery such as memes. We also include Di- versity Boosters to prevent the model from defaulting to safe and repetitive choices like generic emojis, encouraging broader use of formats such as audio clips for intense emotion and video frames for dynamic reactions. Step 3: Persona-based Query and Description Generation (Irony- Aware). The final step bridges the gap between conversational con- text and multimodal retrieval. As detailed in Table 6, this step is important for simulating linguistic phenomena such as sarcasm. Text-Image Conflict Mechanism. We explicitly instruct the model to analyze the speaker’s latent intention. When sarcasm, irony, or passive aggression is present, the model generates visual descrip- tions that contradict the literal text, such as pairing “Great idea” with a “Facepalm” meme. This design forces downstream models to rely on multimodal reasoning instead of text alone. Rich Description Constraint. To improve grounding quality, we impose a strict length Prompt Box 1: Stance Analysis and Modality Planning Prompt ROLE: You are an expert analyst of online chat conversations. Your task is to identify key moments where users would naturally use multimodal elements. TASK A: Stance Significance Score. Rate from 0.0 to 1.0 how much this dialogue involves opinion, debate, or stance-taking. TASK B: Injection Point Recommendations. If the score is high (>0.7), identify potentialturn_ids using the following Prior- ity Criteria: 1) [Priority 1] Stance Shift Point: A clear change in a speaker’s stance (e.g., realization, concession). 2) [Priority 2] Stance Fortification: A moment of strong, emotional reinforce- ment of a stance. 3) [Priority 3] Core Argumentation: A key fact or argument is presented. 4) [Priority 4] Stance Exploration: A pivotal question that challenges another’s stance. TASK C: Enhanced Online Chat Decision Tree. Q1: Visual Con- tent? Does the intent involve showing proof ? (Sharing photos, screenshots, memes?)→Choose IMAGE or VIDEO_FRAME. Q2: Paralinguistic Audio? Does the intent involve sound? (Laugh- ing "haha", crying, screaming, sound effects like "ding/wow")→ Choose AUDIO. Q3: Emotional Reaction? Does the intent involve emotional expression? (Happy, sad, angry, surprised)→Choose EMOJI_MEME or GIF. DIVERSITY BOOSTERS: 1) Emotional Audio Priority: For strong emotions (laughing, yelling), strongly consider AUDIO first to capture intensity. 2) Visual Variety: Don’t default to memes; con- sider if IMAGE or VIDEO_FRAME provides better evidence. Table 5: The prompts for Step 1 (Significance Analysis) and Step 2 (Modality Selection), emphasizing diverse and logic-driven injection. constraint of 7–30 words and require sensory detail. This prevents vague queries like “happy face” and makes the generated captions more useful for retrieval from large-scale databases such as WebVid and AudioSet. Retrieval and Alignment After synthesizing descriptive queries, we run a retrieval process that maps those semantic descriptions to actual multimodal content. This stage addresses the challenge of aligning short conversational turns with large, weakly curated media repositories. Multimodal Reservoirs and Pre-processing. We build a large can- didate pool by aggregating data from open-source collections. For static imagery, we use the full COCO dataset (∼110k images) and a curated set of more than 15,000 expressive stickers. For temporal modalities, we use AudioSet and WebVid. A practical challenge is that the raw source files are much longer than online chat turns: videos and audio clips often last from 3 to more than 10 minutes, whereas the conversational events we model typically last only a few seconds. Temporal Segmentation and Fine-grained Slicing. To resolve this duration mismatch, we implement a Temporal Segmentation pipeline. Instead of matching against full files, we pre-process raw audio and video into short segments, typically 3–8 seconds long, based on scene changes and sound event boundaries. This produces a large reservoir containing hundreds of thousands of candidate clips. As a Conference acronym ’X, June 03–05, 2018, Woodstock, NYTrovato et al. Prompt Box 2: Persona-based Query Generation Prompt (Irony-Aware) ROLE: You are a creative assistant. Transform a simple reason into a detailed, descriptive query. TASK: Generate aretrieval_querythat describes what the mul- timodal content should be. CRITICAL LOGIC (IRONY & CONFLICT): 1) Direct Match: If the speaker is sincere, the multimodal content should match the text emotion. 2) Conflict Generation (Sarcasm/Irony): If the speaker is being sarcastic, passive-aggressive, or ironic, generate a descrip- tion that contradicts the literal text. Example: Text = "Wow, you are a genius." (Sarcastic context)→Query = "A meme of a person facepalming or rolling their eyes in disbelief." REQUIREMENTS: 1) Latent Intention: Describe the implied mean- ing and communicative goal, not just the literal text. 2) Description Length: All descriptions MUST be 7-30 words long. Do not use short phrases. 3) Specific Details: Include visual/audio details (e.g., lighting, expression intensity, specific sound textures). QUALITY CONTROL EXAMPLES: (Check) GOOD: "An animated sticker of a cartoon character turning red with anger, with steam coming out of its ears." (X) BAD: "Funny meme" (Too short/vague). Table 6: The prompt for Step 3, featuring the specific mechanism for simulating sarcasm through text-image conflict. result, retrieved media are not only semantically relevant but also temporally aligned with the pacing of a single dialogue turn. Semantic Alignment and Diversity-Aware Sampling. We use a dual-encoder framework (SentenceTransformer) to compute co- sine similarity between the synthesized query and the captions of segmented clips. To keep the dataset from becoming repetitive, for example by repeatedly retrieving the same “laughing” clip, we adopt a Diversity-Aware Top-푘Sampling strategy with푘=10. We retrieve the top-10 candidates above a dynamic similarity threshold and sample one at random. This preserves semantic fidelity while maintaining visual and auditory diversity. Automated Labeling The final stage assigns fine-grained panop- tic labels to every turn. We use GPT-4o as the annotator. Importantly, the model does not process text in isolation; it performs Cross-Modal Reasoning to combine textual semantics with visual and auditory cues. The detailed labeling prompt is shown in Table 7. Multimodal Synergy in Annotation. As detailed in the “Cross- Modal Reasoning Rules” of Table 7, the prompt explicitly handles interactions across modalities. We instruct the model to prioritize multimodal evidence when it conflicts with text, as in sarcasm, and to use it to refine emotional intensity, for example by upgrading “displeasure” to “anger” based on audio cues. This helps ensure that the labels reflect the holistic communicative intent rather than the textual surface form alone. Formalized Stance Evolution Model. To ensure rigorous stance tracking, we embed a “Stance State Machine” inside the prompt. This makes a strict distinction between Stance Establishment (initial opinion formation) and True Stance Shift (polarity reversal). By enforcing these definitions, we filter out noise such as temporary Prompt Box 3: Automated Panoptic Labeling and Multi- modal Reasoning Prompt ROLE: You are a master of Discourse Analysis acting as a human conversation observer. Your task is to analyze the dialogue history and the provided multimodal captions to determine the speaker’s true stance and emotion. INPUT DATA: 1) Dialogue History: List of turns (Speaker, Text). 2) Multimodal Description: The specific content of the image/audio attached to the current turn (e.g., "Image: A meme of a rolling eye"). PRIMARY DIRECTIVE: CROSS-MODAL REASONING RULES. You must integrate the text and the multimodal description to form a final judgment. 1) Conflict Resolution (Irony/Sarcasm): If the text is positive (e.g., "Great job") but the multimodal content present negative (e.g., "Image: Facepalm", "Audio: Boo sound"), you MUST prioritize the multimodal evidence. Label the sentiment as Negative and the intent as Sarcasm. 2) Intensity Amplification: If the text is neutral but the multimodal content is high-arousal (e.g., "Audio: Screaming", "Image: Crying face"), update the Emotion label to reflect the high intensity (e.g., from "Neutral" to "Sadness" or "Anger"). 3) Visual Contextualization: Use the image content to resolve ambiguous pronouns (e.g., if text says "Look at this", and Image shows "A broken phone", the Topic/Target is "Phone Quality"). ANALYTICAL FRAMEWORK: THE STANCE STATE MACHINE. 1) Stance Persistence (Default): If the current turn is a filler, ques- tion, or ambiguous remark, the core stance remains UNCHANGED (inherits previous). shift_occurred = "No". 2) Stance Establishment: Moving from ’Unknown’ to the FIRST declared stance is NOT a shift. It is Establishment. shift_occurred = "No". 3) True Stance Shift (Reversal): The ONLY case for "Yes". Requires moving from one ESTABLISHED stance (e.g., Support) to a DIFFERENT ESTAB- LISHED stance (e.g., Oppose). OUTPUT SCHEMA (JSON): - keywords: [Array of 3-5 strings cap- turing essence] - emotion: [Selected from predefined Emotion La- bels] - sentiment: [Positive, Negative, Neutral] - stance_after_shift: [Support, Oppose, Neutral] - stance_reason_detail: "Explanation citing specific Text AND Multimodal evidence." - shift_occurred: "Yes" or "No" - shift_reason_category: [Introduction of New Info, Logical Argument, Emotional Appeal, Social Pressure] Table 7: The detailed system prompt for the labeling phase, explicitly enforcing multimodal integration and stance state logic. hesitation or off-topic diversion, so that positiveshift_occurred labels correspond to genuine cognitive change. Human Verification Protocols The entire dataset was assigned to a panel of 17 expert annotators (postgraduate students in NLP) for comprehensive turn-by-turn verification. Rather than simply re- labeling, the annotators performed targeted refinement and filtering according to the following criteria: •Target and Stance Alignment: Annotators inspected the logi- cal consistency between the textual utterances and the assigned stance labels. If a stance flip point identified by the pipeline was found to be ambiguous or lacked a clear causal trigger after expert deliberation, the sample was flagged and discarded. Multimodal Conversational Stance Flipping ForecastingConference acronym ’X, June 03–05, 2018, Woodstock, NY •Contextual Media Verification: Experts reviewed the multi- modal content associated with each turn. If a media item (image, audio, or video) was identified as semantically misaligned with the dialogue context or the speaker’s intent, it was manually reset to N/A to maintain data fidelity. •Heuristic Correction of Stance Inheritance: Following the Stance Persistence Rule, annotators corrected instances where the pipeline failed to propagate the previous definitive stance (푠푡 푝푟푒푣 ) to filler or off-target turns, ensuring the continuous trajectory of speaker convictions. • Three-Strike Filtering Protocol: During the manual audit, a "Three-Strike" policy was enforced for terminal error handling. A dialogue was permanently removed from the benchmark if: (1) the central proposition퐺remained elusive despite manual review; (2) consensus on the stance transition could not be reached among experts; or (3) the rationale was found to be logically irrecoverable. After the audit and refinement phase, we conducted a validation study to measure the objectivity of the final dataset: •Sampling and Protocol: A random subset of 100 dialogues (∼600 turns) was extracted from the refined pool. Two senior annotators independently re-labeled this subset in a double- blind manner, without access to previous pipeline outputs or each other’s decisions. •Statistical Reliability: Inter-annotator agreement was mea- sured using Cohen’s Kappa (휅). The analysis yielded휅=0.84 for stance identification and휅=0.87 for multimodal relevance. These high scores confirm that the StanceFlip annotation guide- lines are highly objective and the final labels are consistent. A.2 Detailed Summary of Dataset Insights ♦ Panoptic Sextuple: The Structural Decoupling of Affect and Argument. We argue that in multi-party discourse, the binary “Support/Oppose” paradigm is insufficient because Affect (Senti- ment) and Argumentation (Stance) operate on different cognitive dimensions. Treating them as the same signal is one of the main sources of noise in existing benchmarks. (1) Target-Specific Disam- biguation: In complex debates, a holder (ℎ) often directs negative sentiment towards an opponent’s tone or personality (ad hominem attacks) while explicitly supporting their logical premise (푡). A mono- lithic label fails to distinguish "interpersonal hostility" from "topical disagreement." Our Panoptic Sextuple(ℎ,푡,푒,푠,푠푡,푟)resolves this ambiguity by strictly binding the holder’s affective state (푒,푠) and stance (푠푡) to the same specific target (푡). This structural decoupling ensures that the model captures the precise cognitive object of the user’s expression, preventing the misclassification of "hostile agreement" as opposition. (2) Rationale as Structural Evidence: We define the Rationale (푟) not merely as a supplementary text span, but as the grounding evidence that validates the stance. It forces the model to ground its prediction in specific discursive logic rather than spurious lexical correlations (e.g., blindly classifying all turns containing "but" as opposition). Thus, the sextuple serves as the minimal sufficient statistic required to represent a static cognitive state without ambiguity. ♦Lifecycle Dynamics: Modeling the Inertia of Belief. Stance is not a discrete event triggered at every utterance, but a continu- ous cognitive state shaped by Cognitive Inertia. Existing methods often treat each turn as an independent classification task and ig- nore the strong temporal dependencies in dialogue. We introduce a formalized Stance State Machine to capture that continuity: (1) Establishment (푆 ∅ → 푆 퐴 ): The initial formation of an opinion from an unknown state. (2) Persistence (푆 퐴 → 푆 퐴 ): Subsequent turns where the holder reiterates, clarifies, or defends the existing view. Crucially, explicitly modeling persistence allows the system to fil- ter out "false flips" caused by mere conversational elaboration or rhetorical hesitation. (3) True Flipping (푆 퐴 → 푆 퐵 ): A low-frequency, high-energy cognitive rupture where the holder’s internal state is rewritten. This dynamic perspective transforms the task from simple classification to complex state tracking, requiring models to distinguish between the maintenance of an existing belief and the genuine adoption of a new one. ♦ Causal Heptatuple: From Descriptive State to Explana- tory Mechanism. The occurrence of a True Flip marks a qualitative change in the data structure, extending the static Sextuple into a dynamic Heptatuple by adding a seventh element: the Trigger (휏). While the sextuple serves as a Descriptive Snapshot of what the current cognitive state is, it lacks the explanatory power to define the transition mechanism. The Heptatuple functions as a Causal Event Record, where the Trigger (휏) represents the specific External Force—such as Logical Correction, Emotional Resonance, or Social Pressure—that possessed sufficient energy to overcome the holder’s Stance Inertia. This transition effectively models the persuasion process as an "Input-Output" system, where휏is the stimulus and the stance flip is the mandatory response. This structure allows StanceFlip to benchmark not just the detection of opinion shifts, but the attribution of their socio-cognitive causes. ♦Taxonomy: Deep Socio-Cognitive Triggers. We catego- rize persuasion mechanisms (휏) into four cognitive dimensions to capture the “why” behind each flip: (1) Logical Argumentation (Epis- temic Change): The flip is driven by factual correction, reasoning, or the exposure of logical fallacies. This represents a change in knowledge or rationality. (2) Emotional Resonance (Affective Align- ment): The flip is triggered by empathy, anger, or tonal alignment (e.g., "I feel your pain"). This represents a change driven by emo- tional contagion rather than facts. (3) New Information (Contextual Update): The flip occurs because new external evidence (news, data) is introduced that changes the premise of the debate. (4) Social In- teraction (Group Dynamics): The flip is a result of peer pressure, the desire for consensus, or deference to authority. This represents a socially motivated compliance. ♦Authentic Multimodal Context. To reduce the “synthetic bias” common in machine-generated corpora, StanceFlip grounds its textual base in high-quality human conversation datasets. This helps preserve real conversational phenomena such as colloquial language and complex turn-taking. Within this setting, we imple- ment a three-stage causal synthesis pipeline so that non-textual media carry meaningful pragmatic weight. The process begins with strategic trigger localization, which identifies the turn where a stance flip occurs as the logical anchor. It then uses contextual modality selection to choose the most effective medium, such as a cynical audio clip or a counter-evidential image. Finally, semantic alignment Conference acronym ’X, June 03–05, 2018, Woodstock, NYTrovato et al. Attitude & Emotion 20.7% Ordinary Life 18.4% Relationship 14.9% Culture & Education 13.1% Work 10.6% Finance 7.0% Health 6.0% School Life 3.6% Tourism 3.1% Politics 2.5% Figure 8: Distribution of principal domains in the dataset. via captioning produces high-fidelity rationales that connect visual or auditory cues with the dialogue flow, ensuring that multimodal information act as genuine drivers of stance evolution. ♦Multi-domain Diversity. To evaluate model generalizability rigorously, StanceFlip spans ten major domains, from high-stakes areas such as Politics to everyday settings such as Relationship. This breadth matters for benchmarking Stance Inertia, because convictions in social or ethical domains often resist change more strongly than consumer preferences. By covering more than 100 sub-domains with substantial argumentative friction, the dataset makes it harder for models to rely on domain-specific shortcuts. The result is a broader test of adaptability across discursive structures and themes. ♦Human-AI Collaborative Annotation Pipeline. Building a large-scale dataset with complex reasoning labels requires a balance between automated scalability and expert review. Our Human-in- the-loop workflow uses LLMs for labor-intensive preprocessing tasks such as multimodal retrieval and draft rationale generation. To maintain annotation fidelity, every instance then passes through a strict audit protocol carried out by 17 NLP experts. The experts focus on validating the logical entailment between the assigned stance and its rationale, ensuring that each transition is supported by a verifiable cognitive chain. We further apply a “Three-Strike” filtering rule to discard dialogues with unclear goals or irrecoverable logic, resulting in a high Cohen’s Kappa (휅= 0.84). B More Details of Methods B.1 Prompting Design for ToS Reasoning framework To ensure reproducibility and facilitate further research, we provide the detailed prompt templates used in our ToS reasoning framework. The framework utilizes a step-by-step persona-based generation process. B.1.1 Step 1: The Cartographer (Target Proposition Identification). Goal: Perform a global scan of the dialogue to formulate a consis- tent debate proposition. Prompt Template: • Instruction: Act as a Cartographer. Your primary goal is to identify the central ’Target’ of this entire conversation. To do this: –Global Scan: Scan the entire dialogue to identify the main topic of contention, the central proposal, or the core issue being debated. –Formulate Proposition: Convert this identified core issue into a clear, concise debate proposition sentence, using constructions like ’The debate concerning whether...’ or ’The discussion about whether...’. The ’Target’ MUST be phrased as a precise debate proposition, not a simple noun (e.g., instead of ’coffee’, frame it as ’The debate concern- ing whether coffee should be tried as a substitute for cigarettes’). This proposition will serve as the consistent reference point for all stances in this dialogue. Input Data: • Full Dialogue: full_dialogue_text_with_captions Output: • (Target: [Debate Proposition Sentence]) B.1.2 Step 2: The Psychologist (Multimodal Conflict Resolution). Goal: Analyze the sentiment and emotion of a specific turn, resolv- ing conflicts between text and visual cues. Prompt Template: •Instruction: Act as a Psychologist and Emotion Analyst. Analyze holder_name’s current utterance, including its tex- tual content and the provided multimodal cue descriptions, to determine its core Emotion and Sentiment. –Textual Layer: Begin by analyzing what the utterance ex- plicitly states and identifying initial emotion and senti- ment solely based on the text. –Multimodal Integration: Carefully examine the ’multimodal cue descriptions’ to see if these non-verbal cues (e.g., de- scribed facial expressions, tone) reinforce, contradict, or enrich the textual meaning. – Conflict Resolution: If there is a conflict between textual and multimodal cues (e.g., if the text says ’That’s great’ but the multimodal cue describes ’a sarcastic smile’), prior- itize interpretations that resolve ambiguity (like detecting sarcasm) to reveal holder_name’s true feelings. –Final Selection: Choose the Emotion and Sentiment from the allowed lists that best reflect holder_name’s attitude towards gt_target. Input Data: • Dialogue History: history_up_to_current_turn • Current Speaker: holder_name • Utterance & Cues: current_utterance_with_caption • Target: gt_target • Constraints: – Emotions=EMOTION_LABELS – Sentiments=SENTIMENT_LABELS Multimodal Conversational Stance Flipping ForecastingConference acronym ’X, June 03–05, 2018, Woodstock, NY Output: •(Holder: [Name], Target: [Proposition], Sentiment: [Class], Emotion: [Class]) B.1.3 Step 3: The Discourse Analyst (Stance Determination). Goal: Determine the final stance by comparing current sentiment with the historical stance trajectory. Prompt Template: • Instruction: Act as a Discourse Analyst. Based on the complete dialogue history (especially all prior utterances of holder_name and their multimodal cues), the current senti- ment/emotion, and holder_name’s ’Last Stance’, determine holder_name’s final Stance towards gt_target. – Historical Review: Carefully review all prior utterances of holder_name regarding gt_target, focusing on the evolution of their expressed content and emotions. –Temporal Comparison: Evaluate the current utterance’s sentiment (gt_sentiment) and emotion (gt_emotion) by comparing it with holder_name’s previous utterances and ’Last Stance’ (previous_stance_for_holder). Deter- mine if it is consistent, slightly divergent, or a clear shift. –Final Decision: Choose the Stance that best reflects holder_name’s current stance. Select ’Support’ or ’Oppose’ only when the opinion is clear; otherwise, choose ’Neutral’ or ’Unknown’. Input Data: • Dialogue History: history •Current Sentiment/Emotion: (Sentiment: gt_sentiment, Emo- tion: gt_emotion) • Last Stance: previous_stance_for_holder Output: • (Holder: [Name], Target: [Proposition], Stance: [Class]) B.1.4 Step 4: The Synthesizer & Critic (Rationale Generation and Reflexion). Goal: Generate a logic-based rationale and perform self-correction. Prompt Template: • Instruction: Act as a Synthesizer and Critic. –Draft Rationale (Synthesizer): Based on the ’Intermedi- ate Analysis Results’ and ’Dialogue History’, construct a concise ’Rationale’ (1-2 sentences) that logically explains WHY holder_name holds this gt_stance. Reference spe- cific textual evidence. If multimodal cues were crucial, explicitly mention their description (e.g., ’despite positive words, their sarcastic facial expression indicated opposi- tion’). – Self-Reflexion (Critic): Critically evaluate the draft: Does it logically, clearly, and powerfully support the stance? Is it specific enough? If the rationale feels weak or un- convincing, revise it to provide a stronger, more precise explanation. Input Data: •Stance Info: (Speaker: holder_name, Target: gt_target, Stance: gt_stance) • Constraints: Rationale MUST be 1-2 sentences. Output: •(Holder: [Name], ..., Stance: [Class], Rationale: [Analytical Summary]) B.2 Self-Reflection Instruction Tuning Implementation To train the metacognitive capability required by the self-reflective mechanism in Section 4.3, we construct a specialized instruction- tuning dataset. Rather than relying only on inference-time prompt- ing, we explicitly train the model to criticize and correct errors through a “Corrupt-and-Correct” data augmentation strategy. The construction process generates three types of training tasks from the ground-truth dataset: B.2.1 Data Construction Logic. We implemented aSelfRefineDataset class that processes the raw StanceFlip dialogues. For every ground truth stance tuple in the dataset, we apply the following logic: Task 1: Initial Generation (Standard Supervised Fine-Tuning) • Input: The full dialogue history. •Instruction: "Extract all stance information from the dia- logue." • Target: The correct, ground-truth stance sextuples. Task 2: Feedback Generation (Criticism) •Objective: To teach the model to identify errors, we artificially "corrupt" the ground truth tuples. •Corruption Strategy: We verify specific elements of the tuple based on a random selection: –Holder Corruption: Swapping the true speaker with an- other participant in the dialogue. –Stance Corruption: Inverting the stance label (e.g., chang- ing "Support" to "Oppose") or modifying the "Flipped Stance" label in transition samples. –Target Corruption: Replacing the specific proposition with a vague placeholder (e.g., "something"). •Instruction: "Provide specific feedback on the extracted stance tuple based on the dialogue." • Input: The dialogue + The Corrupted Tuple. • Target: A natural language explanation of the error (e.g., "The Stance is incorrect. The speaker’s attitude is ’Support’, not ’Oppose’."). Task 3: Output Refinement (Correction) • Instruction: "Fix the extracted stance tuple based on the feedback." •Input: The dialogue + The Corrupted Tuple + The Feedback from Task 2. • Target: The original, correct Ground Truth tuple. B.2.2 Training Specifications. •Template: All prompts are wrapped in a standard chat tem- plate (e.g., Vicuna format: USER: ... ASSISTANT: ...) to align with the backbone LLM’s instruction format. •Negative Sampling: To maintain data balance and prevent the model from overfitting to correction tasks, we apply a down-sampling rate (probability > 0.4) to the self-reflection samples. • Sequence Construction: The input sequences are tokenized with specific handling for EOS (End of Sentence) tokens to Conference acronym ’X, June 03–05, 2018, Woodstock, NYTrovato et al. ensure the model learns to terminate generation correctly. We utilize masking (setting labels to -100) on the prompt in- structions so the model is only trained on the Target outputs (the extraction, the feedback, and the correction). C Extensions of Settings and Implementations To support reproducibility, we provide detailed specifications of the computational infrastructure, model architecture, and hyperparam- eter settings used in the three-stage curriculum learning pipeline. C.1 Computational Infrastructure All experiments were conducted on a high-performance comput- ing cluster with 10×NVIDIA RTX A6000 GPUs (48GB VRAM per GPU). The software environment used PyTorch 2.3.1, Hug- ging Face Transformers 4.57.1, PEFT 0.18.0, and Accelerate 1.12.0. To improve memory efficiency during distributed training, we used DeepSpeed (v0.18.2) with ZeRO-2 optimization and gradi- ent checkpointing. We also used BitsAndBytes (v0.48.2) for 4-bit and 8-bit quantization, and NCCL for cross-GPU communication. C.2 Model Architecture Configurations Backbone and Visual Encoder. We use Vicuna-7B-v1.5 [? ] as the Large Language Model (LLM) backbone because of its strong instruction-following ability. For visual perception, we use the frozen ImageBind-Huge [13] encoder. ImageBind produces a fixed 1024-dimensional embedding for inputs across modalities, including image, video, and audio. Modality Projector. To map visual embeddings into the LLM in- put space (4096 dimensions for Llama-2/Vicuna-7B), we design a learnable Multi-Layer Perceptron (MLP) projector. As used in Stage 1, the projector is defined as: Projector(푥)=푊 2 (GELU(LayerNorm(푊 1 푥)))(7) where푊 1 ∈ R 1024×4096 and푊 2 ∈ R 4096×4096 . During training, all visual inputs are normalized prior to projection. Parameter-Efficient Fine-Tuning (PEFT).. We use Low-Rank Adap- tation (LoRA) in Stages 2 and 3. To balance performance and mem- ory use on A6000 GPUs, we apply QLoRA-style quantization strate- gies: • Stage 2: 8-bit quantization (load_in_8bit=True). •Stage 3: 4-bit Normal Float (NF4) quantization with double quantization enabled (bnb_4bit_quant_type="nf4"). The LoRA adapters were attached to all linear layers:q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, and down_proj. C.3 Training Protocols by Stage We adopt a progressive three-stage training pipeline. The detailed hyperparameters for each stage are summarized in Table 8. Stage 1: Multimodal Alignment (Perception). In this stage, we freeze the LLM backbone and the ImageBind encoder, and train only the Projector. The objective is to align visual features with the text embedding space using image-caption pairs from our dataset. We optimize this stage with AdamW and a learning rate of 2e-5. Stage 2: ToS Reasoning Tuning (Cognition). We freeze the pre- trained Projector and the ImageBind encoder, and enable gradients only for the LoRA adapters. The model is then trained on the Panop- tic Sextuple Extraction task. We use a larger effective batch size through gradient accumulation (푠푡푒푝푠=32) to stabilize the learning of longer reasoning chains. Stage 3: Self-Reflection Tuning (Metacognition). In the final stage, we continue fine-tuning the LoRA adapters initialized from Stage 2 while keeping the Projector frozen. To reduce the instability often seen in 4-bit training, we use thepaged_adamw_32bitoptimizer and gradient clipping (max norm 0.3). This stage focuses on the “Corrupt-and-Correct” objective so that the model becomes better at revising hallucinated reasoning. Table 8: Hyperparameter settings for the three training stages of ConStaFF. HyperparameterStage 1Stage 2Stage 3 Trainable ParametersProjectorLoRA AdaptersLoRA Adapters LLM Quantization FP16Int8NF4 (4-bit) LoRA Rank (푟 )N/A1616 LoRA Alpha (훼 )N/A3232 LoRA DropoutN/A0.050.05 OptimizerAdamWAdamWPaged AdamW 32bit Learning Rate2e-51e-45e-5 LR SchedulerCosineLinearLinear Warmup Ratio0.00.030.03 Per-Device Batch Size414 Grad. Accumulation13216 Max Grad Norm 1.01.00.3 Epochs322 C.4 Training Objectives Across all three stages, the training objective is to maximize the log-likelihood of target tokens conditioned on the input context and multimodal features. Formally, let푋= 푥 1 ,푥 2 , . . .,푥 푁 be the input sequence (in- cluding the visual embeddings푣mapped by the projector) and 푌=푦 1 ,푦 2 , . . .,푦 푀 be the target response sequence. The standard autoregressive language modeling lossL 퐶퐿푀 is defined as: L 퐶퐿푀 (휃)=− 푀 ∑︁ 푡=1 log푃 휃 (푦 푡 | 푋,푦 <푡 )(8) where휃represents the trainable parameters (Projector in Stage 1, LoRA adapters in Stages 2/3). Stage-Specific Objectives. •Stage 1 (Multimodal Alignment):푋consists of the frozen image features and a prefix prompt, while푌is the ground- truth caption. The loss ensures the projector aligns visual concepts with textual semantics. •Stage 2 (ToS Reasoning):푋is the dialogue history with interleaved visual cues, and푌is the Chain-of-Thought rea- soning path (Cartographer→Psychologist→Analyst). We apply a masking strategy where the loss is calculated only on the reasoning steps and final stance labels, ignoring the instruction prompt tokens (푙푎푏푒푙=−100). Multimodal Conversational Stance Flipping ForecastingConference acronym ’X, June 03–05, 2018, Woodstock, NY •Stage 3 (Self-Reflection): This stage optimizes a multi-task objective. The model learns to generate both the critique (푌 푓푒푑푏푎푐푘 ) and the correction (푌 푟푒푓푖푛푒푑 ). The total loss is the summation over the self-correction samples: L 푆푡푎푔푒3 =L 퐶퐿푀 (푌 푓푒푑푏푎푐푘 | 푋 푑푟푎푓푡 )+L 퐶퐿푀 (푌 푟푒푓푖푛푒푑 | 푋 푑푟푎푓푡 ,푌 푓푒푑푏푎푐푘 ) (9) D Evaluation Specifications Because Large Language Models are generative, exact string match- ing is often too strict for evaluating complex reasoning tasks. We therefore use a multi-layer evaluation protocol that combines rule- based normalization with semantic evaluation through an LLM-as- a-Judge. D.1 Preprocessing and Normalization Prior to evaluation, all predicted (푦 푝푟푒푑 ) and ground-truth (푦 푔표푙푑 ) outputs undergo a standard normalization processN(·), which includes: •Case Folding & Whitespace: Converting all text to lower- case and stripping leading/trailing whitespace. •Holder Alias Resolution: Mapping variations of speaker identifiers to a canonical form (e.g., “Speaker 1”, “s1”, “user 1” → “speaker 1”). D.2 Rule-based Categorical Matching For fields with a closed label set, we employ heuristic keyword matching to handle lexical variations in generation. Stance Mapping. We verify if the predicted stance falls into the correct polarity bucket: • Support: Contains “support”, “agree”, “pro”, “positive”. •Oppose: Contains “oppose”, “disagree”, “con”, “negative”, “refus”. • Neutral: Contains “neutral”. Trigger Categorization. Since the model generates free-text trig- gers, we map them to the four taxonomy categories defined in Section 3.1 using keyword heuristics: •Factual & Logical: “fact”, “logic”, “evidence”, “data”, “stat”. •Emotional & Value: “emotion”, “value”, “moral”, “empa- thy”, “fear”. •Personal Experience: “experi”, “story”, “anecdote”, “life”. •Social Influence: “social”, “pressure”, “group”, “norm”, “peer”. A prediction is a True Positive if its mapped category matches the ground truth. D.3 Model-based Semantic Evaluation For open-ended fields (Target and Rationale), rigid string match- ing yields high false negatives. We adopt an LLM-as-a-Judge ap- proach using GPT-4o-mini. We construct a dynamic promptP 푒푣푎푙 incorporating the full dialogue contextC. The judge is instructed to output “YES” only if semantic equivalence is met. The prompt template used is: System: You are a helpful assistant for evaluating text simi- larity. Respond ONLY with ’YES’ or ’NO’. User: Dialogue Context: [C] Predicted Target: [푦 푝푟푒푑 ] Gold Target: [푦 푔표푙푑 ] Is the Predicted Target semantically equivalent to, contained within, or does it refer to the core entity/proposition of the Gold Target? To ensure reproducibility and efficiency, we implement a persis- tent caching mechanism (saved asapi_match_cache.json). If the API call fails or times out, the system falls back to a lenient string containment check (i.e., match if푦 푝푟푒푑 ⊂ 푦 푔표푙푑 or푦 푔표푙푑 ⊂ 푦 푝푟푒푑 ). D.4 Compound Metric Definitions We adopt the standard Precision (푃), Recall (푅), and F1-score (퐹1) as our primary evaluation metrics. Based on the cumulative count of True Positives (푇푃), False Positives (퐹푃), and False Negatives (퐹푁) across the entire test set, the scores are calculated as follows: 푃= 푇푃 푇푃 + 퐹푃 , 푅= 푇푃 푇푃 + 퐹푁 , 퐹1= 2· 푃 · 푅 푃 + 푅 (10) To capture different granularities of model performance, we define specific tuple structures for matching criteria: •Panoptic Micro F1: Requires the correct extraction of the complete sextuple: (Holder, Target, Emotion, Sentiment, Stance, Rationale). This is the strictest metric. •Identification (Iden) F1: Focuses on the "who" and "why", requiring the tuple: (Holder, Target, Rationale). •Flip-Trig F1: The core metric for our StanceFlip task. It evaluates the dynamic transition logic, requiring the tuple: (푆푡푎푛푐푒 푖푛푖푡푖푎푙 ,푆푡푎푛푐푒 푓푙푖푝푒푑 ,푇푟푖푔푒푟 푐푎푢푠푎푙 )(11) This metric verifies that the model has correctly identified both the stance reversal event and its underlying cause. Received 20 February 2007; revised 12 March 2009; accepted 5 June 2009