Paper deep dive
AffectOmni: RL-Verifiable People-Centric Grounded Affective Reasoning for Social and Art-Related Scenes
Yibo Wang, Rui Yang, Jisheng Dang, Bimei Wang, Yitao Wu, Pengfei Cao, Wencan Zhang, Hong Peng, Bin Hu, Tat-Seng Chua
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 95%
Last extracted: 8/28/2026, 3:09:50 AM
Summary
The paper introduces AffectOmni, a framework for verifiable, people-centric affective reasoning in multimodal large language models (MLLMs). It addresses shortcut learning by employing Group Relative Policy Optimization (GRPO) with specific rewards: People Focus (targeting micro-expressions and body language) and Temporal Order (targeting emotional trajectory coherence). To improve reward discriminability, it uses within-group comparative scoring instead of absolute LLM-as-a-Judge scoring. For external verification, a Thinking Summarizer converts free-form reasoning into structured evidence instructions, which are grounded to pixel-level regions using SAM3, enabling auditable evidence localization.
Entities (11)
Relation Signals (9)
AffectOmni → evaluatedon → IntentBench
confidence 95% · Experiments on IntentBench... show consistent improvements
AffectOmni → evaluatedon → WorldSense
confidence 95% · Experiments on... WorldSense show consistent improvements
AffectOmni → evaluatedon → Daily Omni
confidence 95% · Experiments on... Daily Omni... show consistent improvements
AffectOmni → incorporates → People Focus Reward
confidence 95% · AffectOmni introduces People Focus and Temporal Order rewards
AffectOmni → incorporates → Temporal Order Reward
confidence 95% · AffectOmni introduces People Focus and Temporal Order rewards
Thinking Summarizer → interfaceswith → SAM3
confidence 95% · grounded into pixel level evidence regions via SAM3
AffectOmni → uses → GRPO
confidence 95% · AffectOmni, a GRPO trained framework for verifiable affective reasoning.
AffectOmni → uses → Thinking Summarizer
confidence 95% · a Thinking Summarizer converts free form rationales into executable evidence instructions
AffectOmni → →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Multimodal large language models (MLLMs) achieve strong performance on VQA and scene understanding, yet affective reasoning remains vulnerable to shortcut behavior. Models may predict correct answers while neglecting people-centric cues such as micro expressions and body language, which weakens traceability and external verification. Prior reinforcement learning approaches mainly reward context or logical coherence without explicitly enforcing attention to human evidence. In addition, LLM as a Judge scoring often suffers from score clustering, which reduces reward discriminability. We propose AffectOmni, a GRPO trained framework for verifiable affective reasoning. AffectOmni introduces People Focus and Temporal Order rewards to encourage people-centric evidence selection and temporally structured reasoning, and it adopts within-group comparative scoring to produce more stable and discriminative reward signals. For verification, a Thinking Summarizer converts free form rationales into executable evidence instructions, which are grounded into pixel level evidence regions via SAM3 to provide an externally auditable interface outside the training loop. Experiments on IntentBench, Daily Omni, and WorldSense show consistent improvements over open source 7B scale baselines, including gains of 4.66% on emotion recognition and +14.29% on temporally sensitive tasks. Code is available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2608.26193v1
- Canonical: https://arxiv.org/abs/2608.26193v1
Trouble viewing inline? Open PDF directly →
Full Text
71,527 characters extracted from source content.
Expand or collapse full text
AffectOmni: RL-Verifiable People-Centric Grounded Affective Reasoning for Social and Art-Related Scenes Yibo Wang Rui Yang Jisheng Dang Bimei Wang Yitao Wu Pengfei Cao Wencan Zhang Hong Peng Bin Hu and Tat-Seng Chua Thanks: Yibo Wang, Rui Yang, Jisheng Dang, Bimei Wang, Pengfei Cao, Hong Peng, and Bin Hu are with Lanzhou University, Lanzhou, China. Thanks: Yitao Wu is with Hainan University, Haikou, China. Thanks: Wencan Zhang and Tat-Seng Chua are with the School of Computing, National University of Singapore, Singapore. Thanks: Corresponding authors: Bin Hu, Pengfei Cao, and Wencan Zhang. Abstract Multimodal large language models (MLLMs) achieve strong performance on VQA and scene understanding, yet affective reasoning remains vulnerable to shortcut behavior. Models may predict correct answers while neglecting people-centric cues such as micro expressions and body language, which weakens traceability and external verification. Prior reinforcement learning approaches mainly reward context or logical coherence without explicitly enforcing attention to human evidence. In addition, LLM as a Judge scoring often suffers from score clustering, which reduces reward discriminability. We propose AffectOmni, a GRPO trained framework for verifiable affective reasoning. AffectOmni introduces People Focus and Temporal Order rewards to encourage people-centric evidence selection and temporally structured reasoning, and it adopts within-group comparative scoring to produce more stable and discriminative reward signals. For verification, a Thinking Summarizer converts free form rationales into executable evidence instructions, which are grounded into pixel level evidence regions via SAM3 to provide an externally auditable interface outside the training loop. Experiments on IntentBench, Daily Omni, and WorldSense show consistent improvements over open source 7B scale baselines, including gains of 4.66% on emotion recognition and +14.29% on temporally sensitive tasks. Code is available at https://github.com/eliot127825-rgb/AffectOmni_nobody. Index Terms: Affective Computing, Multimodal Large Language Models, Reinforcement Learning, Visual Grounding, Emotion Recognition, Social Scene Understanding. I Introduction Fig. 1: Shortcut reasoning vs. human-centric reasoning. Left: without People Focus reward, evidence attribution drifts to background cues, leading to an incorrect answer. Right: with People Focus reward, evidence anchors on human cues such as hand placement and posture change, enabling evidence-grounded affective inference. The displayed frame is a representative snapshot from the full video, previous events provide context, while human cues provide discriminative evidence. Multimodal large language models (MLLMs) have made substantial progress [1, 43, 53], and reinforcement learning, especially GRPO [40], further strengthens their capacity for deep reasoning [23, 63]. This capability is widely expected to support high level social cognition tasks involving human emotion and intent [4, 29, 55]. However, Visionary-R1 [50] and HumanOmniV2 [58] report a pronounced shortcut problem in multimodal reasoning. Models often guess answers from global cues while failing to analyze fine-grained evidence. HumanOmniV2 introduces explicit context modeling and an LLM as a Judge reward [68], which partially alleviates limitations in global context understanding. Nevertheless, in affective computing [64, 22], we identify a more fundamental challenge. Even when a model predicts the correct emotion, its reasoning often overlooks people-centric cues such as micro expressions and body language and cannot be verified against external evidence [39, 11]. As shown in Fig. 1, current models [58] may rely on background cues rather than people-centric evidence, yielding an untrustworthy reasoning process and sometimes an incorrect answer. This remains a key barrier to deploying affective AI in real world settings [12, 2]. We distill the trustworthiness challenge in affective reasoning into three interrelated questions. (1) How can we design fine-grained reward signals that guide the model to learn a people-centric reasoning mode, rather than relying on coarse global cues. (2) How can we obtain reward signals with sufficient discriminative power during reinforcement learning to reliably separate evidence grounded reasoning from shortcut reasoning. (3) How can we transform long chain reasoning into an executable structured representation and establish a verification mechanism that links reasoning claims to visual evidence. Existing approaches address these needs only partially. For (1), prior rewards such as global context completeness in HumanOmniV2 [58] mainly target holistic context understanding, but they do not explicitly model or encourage key affective cues, including micro expressions such as furrowed brows and teary eye corners, body language such as gestures and posture changes, and temporal dynamics that reflect emotion evolution over time [29, 32]. For (2), most LLM as a Judge mechanisms adopt an independent scoring strategy. Each candidate response is scored in isolation. However, LLMs can exhibit calibration drift and score clustering in absolute scoring settings [45, 68], causing reasoning paths of different quality to receive similar scores and yielding weakly discriminative rewards for policy optimization. For (3), current models often generate verbose chains of thought of roughly 800 to 1400 characters that contain rich details but remain loosely structured and low density, making them difficult to directly support downstream applications such as counseling assistance and keyframe extraction [31, 36]. Moreover, label only evaluation cannot test whether the model truly attends to the visual regions that are relevant to the emotion judgment [61, 37]. Our central thesis is that trustworthy affective reasoning depends not only on answer correctness, but more importantly on a reasoning process that is traceable, reward signals that are discriminative, and outcomes that are externally verifiable. Building on this view, we propose the AffectOmni framework. Our main contributions are as follows. • We propose the first reasoning to evidence grounding paradigm for affective reasoning tasks. We compress model generated affective reasoning into executable, structured evidence instructions and interface them with SAM3 to produce pixel level evidence regions on video frames. This provides externally traceable and verifiable evidence anchors for the people, actions, and temporal cues referenced in reasoning. • We propose a fine-grained reward mechanism for people-centric reasoning by introducing multi dimensional constraints, including People Focus and Temporal Order, into reinforcement learning. These constraints encourage the model to attend to micro expressions, body language, and human interactions, and to organize reasoning over time. • During GRPO training, we introduce a within-group comparative scoring strategy that mitigates calibration drift and score clustering in LLM as a Judge evaluation, thereby providing a more stable and discriminative optimization signal. • We achieve state of the art performance on multimodal benchmarks including IntentBench, Daily-Omni, and WorldSense, showing consistent improvements over open source 7B scale baselines, including +4.66% on emotion recognition and +14.29% on temporally sensitive tasks. I Related Work Multimodal Affect Understanding. Affective computing aims to enable machines to understand human affect, and the field has progressed from rule based classification to deep reasoning [4, 29]. Early studies focused on unimodal feature extraction and end to end recognition [64], followed by explorations of multimodal fusion strategies. Hazarika et al. [18] propose modality invariant representation learning, and Hu et al. [21] build a unified framework for affect and emotion recognition. With the rise of MLLMs, affective computing has entered a new phase. Lian et al. [32] introduce an explainable multimodal affective reasoning benchmark. Zhou et al. [70] study an audio visual temporal alignment mechanism. Lian et al. [31] propose AffectGPT for open vocabulary affect understanding. Chain of thought prompting has also been explored for implicit affective reasoning [26, 57]. However, existing methods often exhibit shortcut learning. Models may bypass fine-grained cues such as micro expressions and body language and directly predict labels [50, 58]. We introduce a fine-grained reward design centered on people-centric cues to explicitly encourage attention to key affective evidence. Fig. 2: AffectOmni framework. The pipeline consists of three stages. (1) LOCATE applies GRPO training with human-centric rewards (RpplR_ppl, RtmpR_tmp) and comparative scoring to guide evidence-grounded reasoning. (2) REASON generates structured output and compresses the reasoning chain into a Minimal Evidence Package (MEP) via a Thinking Summarizer. (3) VERIFY grounds the MEP to pixel-level masks using SAM3, enabling post-hoc auditing of whether the evidence claims can be localized in the visual input. Stars indicate our key contributions. We further evaluate the verification module in Section V. Reinforcement Learning and Reward Design. Reinforcement learning has become a key approach for eliciting the reasoning capability of MLLMs. Huang et al. [23] and Yuan et al. [63] first apply GRPO to visual reasoning, and subsequent work extends it to video understanding [47, 30], omni modal reasoning [69, 71], and audio visual domains [52, 48]. To mitigate shortcut learning, recent studies design diverse reward functions. HumanOmniV2 [58] proposes context and logical consistency rewards. Chen et al. [8] introduce a step level reward, and Chen et al. [9] design a consistency reward. Other variants include format rewards [66], perceptual rewards [17], and hybrid reward strategies [46, 20]. Most of these methods follow the LLM as a Judge paradigm, where a language model assigns absolute scores to each generated output in isolation. However, calibration drift and score clustering can arise in absolute scoring, so reasoning paths of different quality may receive similar scores, reducing reward discriminability. We propose within-group comparative scoring, which evaluates multiple candidates jointly and enforces relative ranking to obtain more discriminative reward signals. Visual Grounding and Reasoning Verification. Visual grounding aligns language descriptions with regions in images or videos [61, 37]. Reich et al. [39] propose faithful and plausible metrics for grounding, and Das et al. [11] analyze attention discrepancies between humans and deep networks. In medical VQA, studies validate the effectiveness of the localize then answer paradigm [35, 7]. For trustworthy MLLMs, grounding can mitigate hallucinations [13] and support multimodal fact checking [42, 6]. Yu et al. [62] and Sun et al. [41] use RLHF for behavior alignment, and Zhang et al. [65] employ reinforcement learning to improve multi image grounding. The SAM family [38, 5] enables pixel level video segmentation, providing a technical basis for reasoning verification. Prior work explores visual grounding and post-hoc localization, but it typically treats grounding and affective reasoning as separate components. We integrate a reasoning summarizer with SAM3 based pixel level grounding to provide an externally auditable evidence interface for post-hoc verification of affective reasoning. I Method I-A Problem Formulation and Framework Overview Given a multimodal input x that includes a video frame sequence v, an audio signal a, and a textual question q, our goal is to train a policy model πθ _θ to generate a structured reasoning output o=⟨context,think,answer⟩o= context,think,answer . The context segment describes the multimodal context, the think segment provides step by step reasoning, and the answer segment gives the final decision. We train the policy under the Group Relative Policy Optimization (GRPO) framework. For each input x, the model samples G candidate responses oii=1G\o_i\_i=1^G and updates the policy using within-group relative advantages. The training objective is given in Eq. (1) and Eq. (2). (θ)=[1∑i=1G|oi|∑i=1G∑t=1|oi|Li,t(θ)],J(θ)=E [ 1 _i=1^G|o_i| _i=1^G _t=1^|o_i|L_i,t(θ) ], (1) Li,t(θ)=min(ri,tA^i,clip(ri,t,1±ε)A^i),L_i,t(θ)= (r_i,t A_i,\;clip(r_i,t,1± ) A_i ), (2) where ri,t=πθ(oi,t∣,oi,<t)/πθold(oi,t∣,oi,<t)r_i,t= _θ(o_i,t ,o_i,<t)/ _ _old(o_i,t ,o_i,<t) is the importance sampling ratio, and A^i=Ri−mean(Rjj=1G) A_i=R_i-mean(\R_j\_j=1^G) denotes the advantage estimate with a within-group mean baseline. The clipping range is (1±ε)(1± ). The proposed AffectOmni framework is illustrated in Fig. 2. It consists of three complementary components, people-centric reward shaping for task aligned affective reasoning, within-group comparative scoring for more discriminative GRPO rewards, and a post-hoc auditing interface that grounds reasoning claims into visual evidence. The overall reward is defined as R=Rfmt+Racc+λpRppl+λtRtmp,R=R_fmt+R_acc+ _pR_ppl+ _tR_tmp, (3) where RfmtR_fmt and RaccR_acc denote the format reward and the accuracy reward, RpplR_ppl and RtmpR_tmp are our proposed People Focus and Temporal Order rewards, and λp _p, λt _t are weighting coefficients used to balance reward components. I-B People-Centric Fine-Grained Reward Shaping Existing methods such as HumanOmniV2 adopt a context level reward that mainly evaluates global context completeness, but they do not explicitly model key affective cues such as micro expressions and body language. To address this limitation, we design two complementary fine-grained reward functions. These rewards are task-aligned for affective reasoning rather than universal criteria across domains. For new tasks, the framework remains unchanged, while only the evaluation criteria in the judge prompts need to be replaced. People Focus Reward. Let c and h denote the context text and the reasoning text generated by the model. We use an LLM as a Judge model to assess whether the reasoning attends to people-centric cues. As shown in Eq. (4), Rppl=fLLM(c,h,ppl)∈0,1,R_ppl=f_LLM(c,h;P_ppl)∈\0,1\, (4) where pplP_ppl is an evaluation prompt that guides the judge to evaluate three dimensions. The first is facial expression description that covers micro expressions such as furrowed brows and teary eyes. The second is body motion analysis that covers non verbal signals such as gestures and posture changes. The third is interpersonal interaction modeling that captures affective exchange and responses between interlocutors. Temporal Order Reward. Affective states evolve over time, so the model should capture temporal changes along the affective trajectory. We therefore introduce a Temporal Order reward Rtmp=fLLM(c,h,tmp)∈0,1,R_tmp=f_LLM(c,h;P_tmp)∈\0,1\, (5) where tmpP_tmp instructs the judge to evaluate two aspects. The first is the proper use of temporal markers such as initially, then, and finally. The second is temporal coherence of the affective trajectory, such as a gradual transition from surprise to relief. Implementation Details. To reduce API overhead, we adopt a joint evaluation scheme. We merge pplP_ppl and tmpP_tmp into a single prompt jointP_joint and return two scores in one call. In addition, we apply a causal mask so that RpplR_ppl and RtmpR_tmp act only on token positions in the context and think segments, which prevents reward leakage into the answer segment. I-C Within-Group Comparative Scoring The independent absolute scoring paradigm based on LLM as a Judge has inherent drawbacks. Large language models can exhibit calibration drift and score clustering in absolute scoring, causing reasoning paths of different quality to receive similar scores and reducing reward discriminability. Comparative Scoring Mechanism. We propose a within-group comparative scoring strategy that presents the G candidates for the same prompt, o1,…,oG\o_1,…,o_G\, jointly to the judge model and requires relative comparison rather than absolute scoring. Formally, s1,…,sG=fcmp(o1,…,oG,cmp),\s_1,…,s_G\=f_cmp(\o_1,…,o_G\;P_cmp), (6) where si∈[1,10]s_i∈[1,10] is the relative score for the iith candidate. The prompt cmpP_cmp instructs the judge to perform explicit comparison before assigning scores and to produce a differentiated ranking among candidates. Advantage Computation. During training, we compute an overall reward RiR_i for each candidate as a weighted sum of multiple reward components Ri=∑kλkri(k),R_i= _k _k\,r_i^(k), (7) where ri(k)r_i^(k) denotes the kkth reward component and λk _k is the corresponding weight. Comparative scoring is not a separate reward term; it is used to estimate judge-based components over the G candidates. For the G candidates of the same prompt, we compute a within-group mean baseline and obtain the advantage A^i=Ri−R¯g,R¯g=1G∑j=1GRj, A_i=R_i- R_g, R_g= 1G _j=1^GR_j, (8) and optionally apply within-group standardization A^i←A^i/(σg+ϵ) A_i← A_i/( _g+ε) to improve reward scale consistency and training stability. Distributed Implementation. In multi GPU training, comparative scoring requires aggregating candidates across devices. We use a gather and broadcast scheme. Candidates from all GPUs are gathered to the rank 0 process, which calls the judge model to perform comparative scoring. The resulting scores are then broadcast to all ranks to resume backpropagation. I-D Reasoning to Evidence Grounding Framework Long chain reasoning, often around 800 to 1400 characters, can contain rich analytic details, but it is typically low density and loosely structured, which limits its utility for downstream applications and external verification. We propose a reasoning to evidence grounding framework that converts free form reasoning into an executable structured representation and grounds the resulting evidence instructions through visual segmentation. The Thinking Summarizer and SAM3 are used only for post-hoc evidence execution and visualization-based evaluation. The entire verification stage is excluded from GRPO training and is not used as a reward signal. It serves as a decoupled auditing interface rather than an end-to-end optimization loop. Thinking Summarizer. We define a mapping from the reasoning output o to a structured summary z as gϕ:o↦zg_φ o z, where z contains four fields, key points, primary objects, emotional indicators, and SAM3 instruction. We train gϕg_φ with a two stage knowledge distillation pipeline. In the first stage, a large API based model produces high quality summaries to serve as training targets. In the second stage, we apply LoRA fine tuning to a lightweight base model with an autoregressive objective ℒsum=−∑tlogpϕ(zt∣o,z<t).L_sum=- _t p_φ(z_t o,z_<t). (9) SAM3 Based Segmentation Integration. The SAM3 instruction field is parsed into a target entity list e1,…,eK\e_1,…,e_K\ and used to prompt SAM3 to produce pixel level segmentation over video frames k=SAM3(,ek),k=1,…,K,M_k=SAM3(v,e_k), k=1,…,K, (10) where kM_k denotes the segmentation mask for the kkth entity across frames. Verification Framework. Given the SAM3 outputs, we verify whether the reasoning attends to the correct visual regions. For example, if the reasoning states that the man’s expression shifts from surprise to relief, then the SAM3 instruction should include the corresponding person entity, and SAM3 should successfully segment the target. The segmentation success rate provides an external indicator of reasoning trustworthiness by establishing an externally auditable linkage between textual evidence claims and pixel level regions across frames. The full reasoning verification pipeline is illustrated in Fig. 3. Fig. 3: Visualization of free-form reasoning transformed into structured evidence, i.e., key points, focus objects, and SAM3 instructions for SAM3 based pixel level grounding. IV Experiments In this section, we report experimental results and provide further analysis. Implementation details are provided in Appendix. IV-A Experimental Setup Training Data and Configuration. We use Qwen2.5-Omni-7B-Thinker as the base model. To obtain a stable and controllable policy initialization, we follow a standard multimodal reasoning alignment recipe for structured output initialization, enabling the model to consistently produce the three part output ⟨context,think,answer⟩ context,think,answer . This stage serves to establish a strong baseline and control variables. We apply cold start SFT to stabilize long chain reasoning and format consistency, and then conduct two stage GRPO training to suppress reasoning collapse and format drift. This stage does not include the proposed rewards RpplR_ppl and RtmpR_tmp. The training corpus contains 24K video audio samples from Video-R1 [14], Social-IQ 2.0 [49] using the training split, and EMER [32]. During GRPO training, we sample G=4G=4 candidate responses per instance, set ε=0.2 =0.2, use a learning rate of 1×10−61× 10^-6, and cap the maximum output length at 2048 tokens. We set reward weights to λp=0.2 _p=0.2 and λt=0.2 _t=0.2. All experiments run on 4×A100 80GB GPUs with DeepSpeed ZeRO Stage 2 for acceleration and memory optimization. Evaluation Benchmarks. We evaluate on three multimodal benchmarks. IntentBench is our curated benchmark for human intent and emotion understanding. It contains 633 videos and 2,689 questions drawn from the Social-IQ 2.0, EMER, and MDPE subsets. Daily-Omni [70] is a daily scenario audio visual QA benchmark with 684 videos and 1,197 questions covering six task categories. WorldSense [19] is a world knowledge audio visual QA benchmark with 1,662 videos and 3,172 questions spanning eight domains. Baselines. We compare against two groups of models. The first group includes proprietary models, GPT-4o [24], GPT-o1 (think) [25], Gemini-2.5-Pro (think) [16], Gemini 2.0 Flash, Gemini 2.0 Flash Lite, Gemini 1.5 Pro [44], and Claude 3.5 Sonnet [3]. The second group includes open source omni modal models, Qwen2.5-Omni [53], MiniCPM-o [59], Ola [33], HumanOmniV2 (7B) [58], Unified-IO-2 (8B) [34], VideoLLaMA2 (7B) [10], and VITA-1.5 (7B) [15]. IV-B Main Results IntentBench Results. Table I reports results on IntentBench. AffectOmni reaches an average accuracy of 71.89, exceeding the strongest open source baseline HumanOmniV2 (7B) at 69.23. The improvements are moderate but structured, with larger gains on Emotion and How, which are more directly aligned with people-centric evidence. The largest gains occur on Emotion and How, with increases of 4.66 percentage points (p) and 3.64 p, the absolute difference between two accuracy percentages. We attribute these gains to the stronger reliance of these categories on people related fine-grained cues and temporally organized reasoning. Compared with prior models that can answer correctly while bypassing human evidence, AffectOmni more reliably forms people-centric evidence chains. TABLE I: Comparison on IntentBench. We report category-wise intent understanding and social reasoning performance. Models marked with † are proprietary. The last three rows are ablation variants of AffectOmni built incrementally on HumanOmniV2 by adding People-Focus (RpplR_ppl), Temporal-Order (RtmpR_tmp), and both rewards jointly (Full Model). Methods LMM Why How What When Who/Which Other Emotion Deception Avg Proprietary MLLMs GPT-4o† – 61.46 55.69 60.00 35.71 76.00 63.31 60.99 59.00 59.98 GPT-o1† (think) – 68.19 65.82 66.04 57.14 76.00 68.83 67.26 59.50 66.69 Gemini-2.5-Pro† (think) – 68.57 67.41 65.12 57.14 64.00 70.03 68.23 60.00 67.15 Open-Source MLLMs MiniCPM-o 8B 57.87 53.48 57.14 57.14 68.00 61.14 23.85 49.50 54.51 VITA-1.5 7B 53.15 49.36 51.66 71.42 64.00 61.14 53.20 59.50 54.17 Ola 7B 60.60 55.37 56.87 64.28 76.00 62.91 46.66 44.50 57.41 Qwen2.5-Omni 7B 62.60 63.44 63.53 57.14 76.00 69.03 59.74 63.50 64.20 HumanOmniV2 7B 66.05 64.08 68.75 57.14 76.00 74.56 80.78 66.50 69.23 AffectOmni Ablation (built on HumanOmniV2) + People-Focus(RpplR_ppl) 7B 66.48 67.25 68.33 57.14 76.00 75.35 84.23 63.50 69.78 + Temporal-Order(RtmpR_tmp) 7B 64.33 68.20 68.96 57.14 80.00 74.16 83.67 60.50 69.62 Full Model (AffectOmni) 7B 66.91 67.72 70.21 71.43 76.00 74.95 85.44 62.50 71.89 Daily-Omni Results. Table I reports results on Daily-Omni across six task dimensions. Compared with the strongest open source baseline HumanOmniV2 (7B), AffectOmni improves average accuracy from 58.47 to 61.90 and yields consistent gains across several representative dimensions. In particular, Context increases by 4.15 p, Reason increases by 4.01 p, and 60s tasks increase by 5.46 p. On Infer, AffectOmni is slightly lower than HumanOmniV2, suggesting that this dimension relies more on general inference and knowledge transfer than on people-centric evidence constraints. We also note that the proprietary model Gemini 2.0 Flash reaches an average accuracy of 67.84 on this benchmark, indicating that performance under complex daily distributions remains influenced by model scale and general capability. Nevertheless, AffectOmni improves over open source 7B scale baselines on several key dimensions, suggesting transfer beyond the diagnostic set, although the gain is not uniform across all dimensions. TABLE I: Comparison with existing models on Daily-Omni. Models marked with † are proprietary. We report accuracy (%) on representative sub-dimensions; full results are provided in the supplemental material. Method Context Infer. Reason. 60s Avg Gemini 2.0 Flash† 63.73 76.62 75.43 56.57 67.84 Gemini 2.0 Flash Lite† 58.03 74.03 72.00 53.01 61.32 Unified-IO-2 (8B) 26.42 35.06 29.71 30.00 28.24 VideoLLaMA2 (7B) 35.75 40.91 34.29 31.82 35.17 Qwen2.5-Omni (7B) 38.86 57.79 61.71 38.36 47.45 Ola (7B) 39.89 61.03 66.28 48.72 49.87 MiniCPM-o (7B) 49.22 68.83 61.14 52.00 53.13 HumanOmniV2 (7B) 51.81 72.72 74.28 53.09 58.47 + People-Focus (RpplR_ppl) 58.55 75.32 77.14 60.36 62.57 + Temporal-Order (RtmpR_tmp) 55.96 73.38 77.71 58.73 62.41 AffectOmni (Ours) 55.96 72.08 78.29 58.55 61.90 WorldSense Results. Table I reports domain wise performance on WorldSense. AffectOmni achieves an average accuracy of 48.80, exceeding the open source baseline HumanOmniV2 (7B) at 47.70 and slightly surpassing the proprietary baseline Gemini1.5 Pro. The largest domain wise gain appears in Music, with an improvement of 3.9 p. Compared with Gemini1.5 Pro, AffectOmni remains behind on Film, Tech, and Perform., which rely more on external knowledge and cross domain semantics. At the same time, AffectOmni shows a clear advantage in Music, leading to a higher overall average. These results suggest that our gains primarily arise from improved reasoning structures that organize people related cues and temporal evidence, rather than from uniform improvements across knowledge intensive domains. TABLE I: Comparison with existing models on WorldSense. Models marked with † are proprietary. We report accuracy (%) on representative domains; full results are provided in the supplemental material. Method Tech Film Perform. Music Avg Claude 3.5 Sonnet† 43.70 36.50 30.70 33.90 34.80 GPT-4o† 48.00 43.50 41.90 42.70 42.60 Gemini 1.5 Pro† 53.70 50.40 52.40 42.00 48.00 Unified-IO-2 XXL (7B) 27.10 23.70 25.50 27.30 25.90 VideoLLaMA2 (7B) 29.40 24.50 26.20 27.10 25.40 VITA-1.5 (7B) 38.20 39.80 41.20 39.90 36.90 Qwen2.5-Omni (7B) 47.80 43.80 48.30 47.30 45.40 HumanOmniV2 (7B) 49.60 47.50 48.40 43.30 47.70 + People-Focus (RpplR_ppl) 50.90 50.40 46.50 46.70 47.60 + Temporal-Order (RtmpR_tmp) 51.30 48.50 48.80 46.40 48.60 AffectOmni (Ours) 51.50 49.30 50.40 47.20 48.80 IV-C Diagnostic Analysis Effectiveness of Reward Components. To assess how the People-Focus reward RpplR_ppl and the Temporal-Order reward RtmpR_tmp mitigate shortcut reasoning and improve verifiability, we conduct an ablation on IntentBench that varies only reward configurations. Table IV shows that the Baseline with accuracy-only reward reaches 69.23 in Avg. Adding RpplR_ppl increases Avg. to 69.78 and improves Emotion to 84.23, indicating that the People-Focus constraint more reliably encourages reliance on people-related cues. Adding RtmpR_tmp alone yields an Avg. of 69.62 and improves Emotion to 83.67, suggesting that explicit temporal constraints help organize reasoning over time-sensitive evidence. Finally, the full model reaches 71.89 in Avg., improving by 2.66 p over the baseline, and increases the When category by 14.29 p, which confirms the importance of the complete reward design for temporally sensitive questions. We also observe a decrease on Deception. Since the bootstrap confidence interval includes zero, we do not interpret it as a statistically reliable degradation. Instead, it suggests a boundary of the current inductive bias, as deception and sarcasm often require modeling the mismatch between surface behavior and underlying intent. Lastly, although the Baseline, RpplR_ppl, and RtmpR_tmp all obtain 57.14 on When, this does not imply saturation. These settings likely fall into the same granularity bin, whereas the full model moves beyond multiple granularity levels, further suggesting that joint rewards substantially improve temporal reasoning. TABLE IV: We report accuracy (%) on IntentBench across key categories. All variants share the same base model and training setup, and differ only in the enabled reward terms, showing the marginal gain brought by each component. Configuration Emotion Deception When Why Avg Baseline (Acc-only) 80.78 66.50 57.14 66.05 69.23 + People-Focus (RpplR_ppl) 84.23 63.50 57.14 66.48 69.78 + Temporal-Order (RtmpR_tmp) 83.67 60.50 57.14 64.33 69.62 Full Model 85.44 62.50 71.43 66.91 71.89 Comparative Scoring Analysis. We find that when the reward distribution becomes overly concentrated due to score clustering, or drifts across samples due to calibration drift, the advantage gaps between candidates are compressed. This reduces the effectiveness of GRPO [40] updates. To address this issue, we compare independent absolute scoring with comparative scoring that enforces ranking, focusing on reward discriminability and downstream performance. Because the two schemes use different raw score scales, we apply min max normalization to each set of scores and linearly rescale them to the common interval [1,10][1,10]. This enables fair comparison of CV, defined as Std over Mean, within-group CV, and the shape of the score distribution. Table V shows that comparative scoring substantially increases dispersion and within-group separability of reward signals relative to independent scoring. The overall CV increases from 0.153 to 0.325, and the within-group CV increases from 0.049 to 0.138. These results indicate clearer relative differences among candidates within the same group, which yields a more informative and stable update signal for policy optimization. Consistent with these statistics, the unified scale visualization in Fig. 4 reveals pronounced clustering in the high score range under independent scoring. In contrast, comparative scoring reduces perfect score clustering, allocates more probability mass to the 9 to 10 and 8 to 9 ranges, and maintains non zero coverage in the low score tail. For example, the 1 to 2 bin has 7.31. This makes extreme calibration shifts caused by absolute scale drift less likely. Correspondingly, IntentBench average accuracy improves from 69.78 to 71.89, showing that stronger reward discriminability translates into downstream reasoning gains. TABLE V: Comparison of scoring strategies. “Ind.” denotes independent scoring. “Comp.” denotes our comparative scoring. CV is the coefficient of variation, computed as standard deviation over mean. Metric Independent Comparative CV 0.153 0.325 Within-Group CV 0.049 0.138 IntentBench Avg (%) 69.78 71.89 Fig. 4: Reward Signal Discrimination Analysis. Each arc represents a score interval (1–10 scale), with arc length proportional to the percentage of samples falling in that interval. Left (Independent): LLM as a Judge scores each candidate in isolation, causing severe high score clustering, over 73% of scores fall in the 9–10 range. Right (Comparative): within-group comparative scoring forces relative ranking among candidates, yielding a more balanced distribution with meaningful differentiation across the full score spectrum. Fig. 5: Reasoning to evidence grounding framework. The framework connects structured reasoning chains to pixel-level grounding via SAM3: AffectOmni summarizes focus entities from the reasoning, converts them into segmentation instructions, and visualizes the exact regions to enable automatic consistency checks and human inspection. Reasoning Verification Framework. A core challenge in affective reasoning is to verify whether a model truly attends to the visual evidence it cites in its analysis. To address this challenge, we introduce a reasoning verification framework in AffectOmni that connects reasoning with visual localization. As shown in Fig. 5, given a video and the question “Is the man on the left surprised?”, AffectOmni generates a structured reasoning chain. It identifies that the man in a checkered shirt appears calm, with no visual cues of surprise such as widened eyes or a startle posture. The reasoning is then passed to a summarization module that extracts the focus entities in the video, for example the man in a checkered shirt with a calm and neutral expression, and converts them into SAM3 segmentation instructions. Finally, SAM3 produces pixel level masks that visualize the regions referenced by the reasoning. This provides a verifiable linkage between the language description and visual evidence in the video frames, enabling automatic consistency checks and human inspection of whether the claimed content is grounded in the attended regions. Statistical Significance of Performance Gains. To verify that the observed gains reflect systematic effects of the proposed reward design rather than random variation, this study conducts one-sided paired bootstrap tests with 10,000 resamples on the full IntentBench test set (n=2,689n=2,689), with results reported in Table VI. AffectOmni achieves statistically significant improvements in both categories directly targeted by RpplR_ppl. The Emotion category yields the largest absolute gain at +4.66+4.66 p (p=0.008p=0.008, 95% CI [+1.2,+8.3][+1.2,\ +8.3] p), while How achieves the strongest significance at p=0.002p=0.002 with +3.64+3.64 p ([+1.3,+5.9][+1.3,\ +5.9] p). The lower p-value on How, despite its smaller absolute gain, is explained by its much larger sample size (n=632n=632 vs. n=133n=133), which provides greater statistical power. Both confidence intervals exclude zero, confirming that these gains arise from the reward mechanism rather than sampling noise. Beyond the directly targeted categories, Why, What, and Other show positive but non-significant trends (+0.86+0.86, +1.46+1.46, +0.39+0.39 p) with confidence intervals spanning zero, indicating partial but unreliable transfer of people-centric reasoning capabilities. The Deception category stands apart with a decline of −4.00-4.00 p (p=1.000p=1.000, 95% CI [−9.0,+1.0][-9.0,\ +1.0] p), the wide interval reflects high within-category variance, so the drop cannot be distinguished from random variation. Deception requires detecting mismatches between surface behavior and underlying intent, a pragmatic reasoning demand where stronger reliance on observable affective cues may reduce rather than increase sensitivity. TABLE VI: One-sided paired bootstrap significance test (10,000 resamples) on IntentBench (n=2,689n=2,689), testing whether AffectOmni >> HumanOmniV2 per category. Δ denotes the mean accuracy difference (AffectOmni −- HumanOmniV2). Shaded rows are categories directly targeted by RpplR_ppl. p∗∗<0.01^**p<0.01; ∗p<0.05^*p<0.05. Category n Δ (p) 95% CI (p) p Emotion (←Rppl← R_ppl) 133 +4.66∗∗+4.66^** [+1.2,+8.3][+1.2,\ +8.3] 0.0080.008 How (←Rppl← R_ppl) 632 +3.64∗∗+3.64^** [+1.3,+5.9][+1.3,\ +5.9] 0.0020.002 Why 698 +0.86+0.86 [−1.4,+3.2][-1.4,\ +3.2] 0.4840.484 What 480 +1.46+1.46 [−1.2,+4.2][-1.2,\ +4.2] 0.3170.317 Other 507 +0.39+0.39 [−1.8,+2.6][-1.8,\ +2.6] 0.7790.779 Deception 200 −4.00-4.00 [−9.0,+1.0][-9.0,\ +1.0] 1.0001.000 Cross-Domain Generalization on General Video QA. A natural concern is whether affective RL training causes catastrophic forgetting of general video understanding capabilities. To test this, this study evaluates AffectOmni on NExT-QA [51], a general-purpose video QA benchmark covering causal, temporal, and descriptive reasoning, without any domain-specific fine-tuning. As shown in Table VII, AffectOmni achieves 80.52% overall accuracy on the multiple-choice test set, surpassing HumanOmniV2 (79.78%) and several general-purpose video models of comparable scale. Fine-grained results in Table VIII further reveal that the improvements are not uniform. AffectOmni outperforms HumanOmniV2 on the Temporal subset (79.44% vs. 76.50%) and the Descriptive subset (86.48% vs. 83.77%), while HumanOmniV2 retains a slight advantage on Causal (80.50% vs. 79.28%). The gains on Temporal and Descriptive are consistent with the design intent of this study: People-Focus and Temporal-Order rewards train transferable capabilities in entity anchoring, temporal organization, and evidence chain construction, which benefit reasoning tasks beyond affective inference. The modest Causal deficit suggests that causal inference relies more on general semantic priors not directly addressed by the reward design. Together, these results indicate that people-centric RL training does not degrade general video understanding and instead produces structured reasoning capabilities that transfer positively across task domains. Representative qualitative comparisons in Appendix F further illustrate the improvements in micro-expression analysis and temporal affective reasoning. TABLE VII: Results on NExT-QA multiple-choice test set. AffectOmni surpasses HumanOmniV2 without domain-specific fine-tuning, confirming that affective RL training does not cause catastrophic forgetting of general video understanding. Method Size Overall Acc. (%) LLaVA-Video [67] 7B 83.20 PLLaVA [54] 7B 81.00 Magma [56] 8B 80.90 AffectOmni (Ours) 7B 80.52 HumanOmniV2 [58] 7B 79.78 LLaVA-OneVision [27] 7B 79.40 mPLUG-Owl3 [60] 8B 78.60 VideoChat2 [28] 8B 63.20 TABLE VIII: Fine-grained NExT-QA results by question type. AffectOmni shows gains over HumanOmniV2 on Temporal and Descriptive subsets, consistent with the transferable reasoning capabilities trained by People-Focus and Temporal-Order rewards. Method Causal Temporal Descriptive AffectOmni 79.28 79.44 86.48 HumanOmniV2 80.50 76.50 83.77 V Discussion This section evaluates the Minimal Evidence Package (MEP) as an Executable Evidence Interface (EEI). We emphasize that its goal is not segmentation quality itself, but converting reasoning into an executable and checkable evidence carrier for post-hoc auditing. This abstraction is important in high stakes scenarios such as clinical decision support or affective companion systems. In such settings, a system should output not only conclusions but also evidence packages that are auditable and reviewable. This supports explanation and provides a safety fallback. MEP encourages reasoning to produce locatable elements, for example the focus person, key actions, and temporal cues, and maps them into executable verification instructions. MEP interfaces with downstream verifiers through the EEI. The interface layer uses FocusObj, TimeSpan, EvidenceType as a minimal query form, and heterogeneous verifiers can be integrated with lightweight adaptation. This enables reuse of the auditing interface without modifying the upstream reasoning compression core. Our current implementation uses SAM3 pixel level localization as one instance. The same interface can extend to keyframe retrieval, object tracking, speaker segment localization, and textual evidence retrieval. To validate these claims and quantify MEP as an EEI rather than a display component, we design three lightweight diagnostic experiments without retraining. These experiments probe evidence extractability, evidence to verification consistency, and falsifiability under external auditing. V-A Evidence Extractability The first experiment tests whether the reasoning chain provides a stable basis of structured information. We randomly sample 100 think outputs from IntentBench and tally three evidence elements that are directly relevant to generating verification instructions. Visual evidence refers to people related observable cues such as expression, gaze, and posture. Temporal evidence refers to order, change, or stage descriptions over time. Behavioral evidence refers to executable actions or interaction cues such as gestures, approach and avoidance, and conversational interaction. To avoid missed detections and ambiguity from keyword matching, we use Qwen-Max as the judge and apply a strict binary criterion. We also require a minimal supporting span and apply schema constrained parsing to restrict the output format and reduce drift from free form generation. This statistic is used to diagnose interface feasibility rather than as a final evaluation metric. Table IX and Table X show that the occurrence rates of the three evidence types are 84%, 72%, and 89%, and that 93% of samples contain at least two evidence types. These results indicate that the model produces composable evidence primitives for most samples, which provides the information basis for the Summarizer to compress reasoning into an MEP. TABLE IX: Occurrence rates of three types of structured evidence (Visual/Temporal/Behavioral) and coverage rate of ≥ 2 evidence types from 100 randomly sampled think outputs. Evidence Type Count Rate Visual Evidence 84 84% Temporal Evidence 72 72% Behavioral Evidence 89 89% ≥ 2 Types 93 93% TABLE X: Evidence occurrence rates grouped by question type. Type Samples Visual Temporal Behavioral Other 28 75% 75% 89% Why 27 81% 70% 81% How 19 95% 74% 89% What 14 86% 64% 93% Deception 8 88% 62% 100% Emotion 4 100% 100% 100% Overall 100 84% 72% 89% V-B Evidence to Verification Consistency The second experiment evaluates the operationality of the EEI, namely whether the focus object description in think provides sufficient detail for successful localization by downstream visual verifiers. We count elements including person mentions, appearance attributes such as color and clothing, visual features such as expression and posture, and locatable actions. Based on these elements, we define two grounding criteria. Basic Grounding requires an explicitly mentioned and segmentable person. High Quality Grounding requires a person mention and at least one executable localization cue, either a visual feature or a locatable action. Table XI shows that 98% of samples mention a person, 91% include visual feature descriptions, and 94% include locatable actions. Accordingly, 98% of samples satisfy Basic Grounding and 86% satisfy High Quality Grounding. By question type in Table XII, Basic Grounding ranges from 93% to 100% across categories, and High Quality Grounding ranges from 82% to 100%, which indicates consistent executability across question types. Appearance attributes appear in only 39% of cases, which is markedly lower than visual features and action cues. This suggests that current reasoning relies more on dynamic evidence such as expression, posture, and actions. Under occlusion or in crowded multi person scenes, grounding can become more ambiguous without additional discriminative appearance cues or spatial references. TABLE XI: Occurrence rates of grounding-relevant elements and two levels of grounding criteria from 100 randomly sampled reasoning chains. Indicator Count Rate Mentions Person 98 98% Has Appearance (color/clothing) 39 39% Has Visual Feature (expression/posture) 91 91% Has Locatable Action 94 94% Basic Grounding 98 98% High-Quality Grounding 86 86% TABLE XII: Grounding attainment rates grouped by question type. Type Samples Basic High-Quality Other 28 96% 82% Why 27 100% 85% How 19 100% 89% What 14 93% 86% Deception 8 100% 88% Emotion 4 100% 100% Overall 100 98% 86% TABLE XIII: Annotators judge, based only on SAM3 segmentation, whether it is related to the question target, and we report the contingency table between this relevance judgment and answer correctness, together with conditional accuracy (n=100). Seg. Relevant Seg. Irrelevant Total Answer Correct 55 12 67 Answer Wrong 11 22 33 Conditional Acc. 83.3% 35.3% 67% V-C External Auditing and Falsifiability The third experiment evaluates falsifiability through external auditing. It tests whether the auditing interface provides an external signal that helps assess answer trustworthiness, rather than producing superficially plausible visualizations for arbitrary instructions. We randomly sample 100 items from IntentBench and ask annotators to view only the question and the SAM3 segmentation results. Annotators do not have access to the model textual reasoning or the ground truth label. They provide a binary relevance judgment, relevant or irrelevant, on whether the segmentation focuses on objects related to the question target. Relevant means that the segmented subject matches the asked for object and covers the key person. Table XIII shows that when segmentation is judged relevant, conditional accuracy is 83.3%. When segmentation is judged irrelevant, conditional accuracy drops to 35.3%, a difference of 48.0 percentage points. The contingency table yields an overall agreement rate of 77%, computed as (55+22)/100(55+22)/100, and the Phi coefficient is ϕ=0.48φ=0.48, which indicates a moderate positive association between focusing on the key object and answering correctly. Among the 67 correct-answer cases, 12 have irrelevant segmentation, corresponding to a 17.9% false-negative rate. These false negatives reduce diagnostic coverage but do not affect training or final answer accuracy because the verifier is post-hoc and excluded from GRPO optimization. Such errors may come from summarization bias in evidence-query extraction, SAM3 grounding failure, or downstream relevance-judgment noise. Together, these results show that annotators do not need to read the reasoning text. By inspecting the post-hoc visual evidence alone, they can obtain an external signal for answer reliability, which supports the value of the MEP as an evidence carrier that is external, operational, and auditable. When segmentation decouples from the question target, accuracy declines substantially. This demonstrates falsifiability because evidence mismatch exposes potentially unreliable outputs in an observable manner, rather than always producing seemingly reasonable results. Representative failure cases in Appendix G further illustrate pragmatic reasoning failures in sarcasm interpretation and reasoning-selection disconnects. VI Conclusion This paper proposes AffectOmni for verifiable affective reasoning, addressing the trustworthiness bottleneck in multimodal emotion and intent understanding where models produce correct answers while bypassing fine grained people-centric evidence. We introduce People-Focus and Temporal-Order rewards into GRPO training to incorporate people-centered evidence selection and temporal organization into the reinforcement learning objective, and propose within-group comparative scoring to mitigate score clustering and calibration drift in LLM as a Judge evaluation. A post-hoc reasoning-to-evidence auditing interface further compresses chain-of-thought outputs into executable evidence instructions grounded via SAM3, supporting external consistency and falsifiability diagnostics (Section V). Experiments on IntentBench, Daily-Omni, and WorldSense confirm consistent gains over open-source 7B scale baselines on people-centric and temporally sensitive tasks. Temporal interval localization remains coarse and error propagation in the verification module is unquantified, future work will address localization precision, auxiliary verification rewards, and reward generalizability. Acknowledgments This work was supported in part by the National Natural Science Foundation of China (Grant No. 62227807 and Grant No. U24B20186), in part by the Brain Science and Brain-like Intelligence Technology—National Science and Technology Major Project (No. 2021ZD0200600, No. 2021ZD0200408), and in part by the WQ & UCAS Research Academy Intelligent Computing Center (WRA-ICC) and the Supercomputing Center of Lanzhou University. [Supplementary Material] The supplementary material provides additional details and analyses supporting the main text. Appendix A describes the full three-stage training pipeline and Thinking Summarizer configuration. Appendix B reports complete per-subcategory results on Daily-Omni and WorldSense. Appendix C analyzes how the group size G affects reward discriminability and verifies the absence of positional bias. Appendix D examines sensitivity to the reward weights λp _p and λt _t. Appendix E evaluates reasoning chain quality along five diagnostic dimensions using both automatic and human assessment. Appendices F and G present qualitative comparisons of success and failure cases, respectively, illustrating the behavioral patterns induced by the proposed rewards. References [1] J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al. (2023) Gpt-4 technical report. arXiv preprint arXiv:2303.08774. Cited by: §I. [2] M. M. Amin, R. Mao, E. Cambria, and B. W. Schuller (2024) A wide evaluation of chatgpt on affective computing tasks. IEEE Transactions on Affective Computing 15 (4), p. 2204–2212. Cited by: §I. [3] Anthropic (2024) Introducing the next generation of claude. Note: Accessed: 2024-10-22 External Links: Link Cited by: §IV-A. [4] E. Cambria, D. Das, S. Bandyopadhyay, and A. Feraco (2017) Affective computing and sentiment analysis. In A practical guide to sentiment analysis, p. 1–10. Cited by: §I, §I. [5] N. Carion, L. Gustafson, Y. Hu, S. Debnath, R. Hu, D. Suris, C. Ryali, K. V. Alwala, H. Khedr, A. Huang, et al. (2025) Sam 3: segment anything with concepts. arXiv preprint arXiv:2511.16719. Cited by: §I. [6] R. F. Cekinel, P. Karagoz, and Ç. Çöltekin (2025) Multimodal fact-checking with vision language models: a probing classifier based solution with embedding strategies. In Proceedings of the 31st International Conference on Computational Linguistics, p. 4622–4633. Cited by: §I. [7] C. Chen, S. Anjum, and D. Gurari (2022) Grounding answers for visual questions asked by visually impaired people. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 19098–19107. Cited by: §I. [8] H. Chen, X. Lou, X. Feng, K. Huang, and X. Wang (2025) Unveiling chain of step reasoning for vision-language models with fine-grained rewards. arXiv preprint arXiv:2509.19003. Cited by: §I. [9] Y. Chen, Y. Ge, R. Wang, Y. Ge, J. Cheng, Y. Shan, and X. Liu (2026) GRPO-CARE: consistency-aware reinforcement learning for multimodal reasoning. External Links: Link Cited by: §I. [10] Z. Cheng, S. Leng, H. Zhang, Y. Xin, X. Li, G. Chen, Y. Zhu, W. Zhang, Z. Luo, D. Zhao, et al. (2024) Videollama 2: advancing spatial-temporal modeling and audio understanding in video-llms. arXiv preprint arXiv:2406.07476. Cited by: §IV-A. [11] A. Das, H. Agrawal, L. Zitnick, D. Parikh, and D. Batra (2017) Human attention in visual question answering: do humans and deep networks look at the same regions?. Computer Vision and Image Understanding 163, p. 90–100. Cited by: §I, §I. [12] Z. Dong, C. Chen, C. Liao, and X. M. Chen (2025) Integrating large language models and affective computing for human-machine symbiosis in intelligent driving. The Innovation 6 (12). Cited by: §I. [13] A. Favero, L. Zancato, M. Trager, S. Choudhary, P. Perera, A. Achille, A. Swaminathan, and S. Soatto (2024) Multi-modal hallucination control by visual information grounding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 14303–14312. Cited by: §I. [14] K. Feng, K. Gong, B. Li, Z. Guo, Y. Wang, T. Peng, B. Wang, and X. Yue (2025) Video-r1: reinforcing video reasoning in mllms. arXiv preprint arXiv:2503.21776. Cited by: §IV-A. [15] C. Fu, H. Lin, X. Wang, Y. Zhang, Y. Shen, X. Liu, H. Cao, Z. Long, H. Gao, K. Li, et al. (2025) Vita-1.5: towards gpt-4o level real-time vision and speech interaction. arXiv preprint arXiv:2501.01957. Cited by: §IV-A. [16] Google DeepMind (2025) Gemini 2.5 Pro Preview: even better coding performance. Note: Blog post. Available at: https://deepmind.google/blog/gemini-25-pro-preview-even-better-coding-performance/ (accessed 2026-01-21). Cited by: §IV-A. [17] Z. Guo, M. Hong, and T. Jin (2025) Observe-r1: unlocking reasoning abilities of mllms with dynamic progressive reinforcement learning. arXiv preprint arXiv:2505.12432. Cited by: §I. [18] D. Hazarika, R. Zimmermann, and S. Poria (2020) Misa: modality-invariant and-specific representations for multimodal sentiment analysis. In Proceedings of the 28th ACM international conference on multimedia, p. 1122–1131. Cited by: §I. [19] J. Hong, S. Yan, J. Cai, X. Jiang, Y. Hu, and W. Xie (2025) Worldsense: evaluating real-world omnimodal understanding for multimodal llms. arXiv preprint arXiv:2502.04326. Cited by: §IV-A. [20] M. Hong, Z. Guo, Y. Xia, Z. Wang, Z. Zhang, T. Jin, and Z. Zhao (2025) APO: enhancing reasoning ability of mllms via asymmetric policy optimization. arXiv preprint arXiv:2506.21655. Cited by: §I. [21] G. Hu, T. Lin, Y. Zhao, G. Lu, Y. Wu, and Y. Li (2022) UniMSE: towards unified multimodal sentiment analysis and emotion recognition. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Y. Goldberg, Z. Kozareva, and Y. Zhang (Eds.), Abu Dhabi, United Arab Emirates, p. 7837–7851. External Links: Link, Document Cited by: §I. [22] G. Hu, Y. Xin, W. Lyu, H. Huang, C. Sun, Z. Zhu, L. Gui, R. Cai, E. Cambria, and H. Seifi (2024) Recent trends of multimodal affective computing: a survey from nlp perspective. arXiv preprint arXiv:2409.07388. Cited by: §I. [23] W. Huang, B. Jia, Z. Zhai, S. Cao, Z. Ye, F. Zhao, Z. Xu, Y. Hu, and S. Lin (2025) Vision-r1: incentivizing reasoning capability in multimodal large language models. arXiv preprint arXiv:2503.06749. Cited by: §I, §I. [24] A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, et al. (2024) Gpt-4o system card. arXiv preprint arXiv:2410.21276. Cited by: §IV-A. [25] A. Jaech, A. Kalai, A. Lerer, A. Richardson, A. El-Kishky, A. Low, A. Helyar, A. Madry, A. Beutel, A. Carney, et al. (2024) Openai o1 system card. arXiv preprint arXiv:2412.16720. Cited by: §IV-A. [26] W. Lai, H. Xie, G. Xu, and Q. Li (2025) Rvisa: reasoning and verification for implicit sentiment analysis. IEEE Transactions on Affective Computing 16 (3), p. 1760–1771. Cited by: §I. [27] B. Li, Y. Zhang, D. Guo, R. Zhang, F. Li, H. Zhang, K. Zhang, P. Zhang, Y. Li, Z. Liu, et al. (2024) Llava-onevision: easy visual task transfer. arXiv preprint arXiv:2408.03326. Cited by: TABLE VII. [28] K. Li, Y. Wang, Y. He, Y. Li, Y. Wang, Y. Liu, Z. Wang, J. Xu, G. Chen, P. Luo, et al. (2024) Mvbench: a comprehensive multi-modal video understanding benchmark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 22195–22206. Cited by: TABLE VII. [29] S. Li and W. Deng (2020) Deep facial expression recognition: a survey. IEEE transactions on affective computing 13 (3), p. 1195–1215. Cited by: §I, §I, §I. [30] X. Li, Z. Yan, D. Meng, L. Dong, X. Zeng, Y. He, Y. Wang, Y. Qiao, Y. Wang, and L. Wang (2025) Videochat-r1: enhancing spatio-temporal perception via reinforcement fine-tuning. arXiv preprint arXiv:2504.06958. Cited by: §I. [31] Z. Lian, H. Chen, L. Chen, H. Sun, L. Sun, Y. Ren, Z. Cheng, B. Liu, R. Liu, X. Peng, J. Yi, and J. Tao (2025) AffectGPT: a new dataset, model, and benchmark for emotion understanding with multimodal large language models. In Forty-second International Conference on Machine Learning, External Links: Link Cited by: §I, §I. [32] Z. Lian, H. Sun, L. Sun, H. Gu, Z. Wen, S. Zhang, S. Chen, M. Xu, K. Xu, K. Chen, et al. (2023) Explainable multimodal emotion recognition. arXiv preprint arXiv:2306.15401. Cited by: §I, §I, §IV-A. [33] Z. Liu, Y. Dong, J. Wang, Z. Liu, W. Hu, J. Lu, and Y. Rao (2025) Ola: pushing the frontiers of omni-modal language model. arXiv preprint arXiv:2502.04328. Cited by: §IV-A. [34] J. Lu, C. Clark, S. Lee, Z. Zhang, S. Khosla, R. Marten, D. Hoiem, and A. Kembhavi (2024) Unified-io 2: scaling autoregressive multimodal models with vision language audio and action. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 26439–26455. Cited by: §IV-A. [35] D. Nguyen, M. K. Ho, H. Ta, T. T. Nguyen, Q. Chen, K. Rav, Q. D. Dang, S. Ramchandre, S. L. Phung, Z. Liao, et al. (2025) Localizing before answering: a benchmark for grounded medical visual question answering. In Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence, p. 7670–7678. Cited by: §I. [36] M. Niu, Y. El-Tawil, A. Romana, and E. M. Provost (2025) Rethinking emotion annotations in the era of large language models. IEEE Transactions on Affective Computing. Cited by: §I. [37] B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik (2015) Flickr30k entities: collecting region-to-phrase correspondences for richer image-to-sentence models. In Proceedings of the IEEE international conference on computer vision, p. 2641–2649. Cited by: §I, §I. [38] N. Ravi, V. Gabeur, Y. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafson, et al. (2024) Sam 2: segment anything in images and videos. arXiv preprint arXiv:2408.00714. Cited by: §I. [39] D. Reich, F. Putze, and T. Schultz (2023) Measuring faithful and plausible visual grounding in VQA. In Findings of the Association for Computational Linguistics: EMNLP 2023, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, p. 3129–3144. External Links: Link, Document Cited by: §I, §I. [40] Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. Li, Y. Wu, et al. (2024) Deepseekmath: pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300. Cited by: §I, §IV-C. [41] Z. Sun, S. Shen, S. Cao, H. Liu, C. Li, Y. Shen, C. Gan, L. Gui, Y. Wang, Y. Yang, et al. (2024) Aligning large multimodal models with factually augmented rlhf. In Findings of the Association for Computational Linguistics: ACL 2024, p. 13088–13110. Cited by: §I. [42] S. Suryavardan, S. Mishra, P. Patwa, M. Chakraborty, A. Rani, A. Reganti, A. Chadha, A. Das, A. Sheth, M. Chinnakotla, et al. (2023) Factify 2: a multimodal fake news and satire news dataset. arXiv preprint arXiv:2304.03897. Cited by: §I. [43] G. Team, R. Anil, S. Borgeaud, J. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, et al. (2023) Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805. Cited by: §I. [44] G. Team, P. Georgiev, V. I. Lei, R. Burnell, L. Bai, A. Gulati, G. Tanzer, D. Vincent, Z. Pan, S. Wang, et al. (2024) Gemini 1.5: unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530. Cited by: §IV-A. [45] P. Wang, L. Li, L. Chen, Z. Cai, D. Zhu, B. Lin, Y. Cao, L. Kong, Q. Liu, T. Liu, et al. (2024) Large language models are not fair evaluators. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 9440–9450. Cited by: §I. [46] P. Wang, Y. Wei, Y. Peng, X. Wang, W. Qiu, W. Shen, T. Xie, J. Pei, J. Zhang, Y. Hao, et al. (2025) Skywork r1v2: multimodal hybrid reinforcement learning for reasoning. arXiv preprint arXiv:2504.16656. Cited by: §I. [47] Q. Wang, Y. Yu, Y. Yuan, R. Mao, and T. Zhou (2025) VideoRFT: incentivizing video reasoning capability in MLLMs via reinforced fine-tuning. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, External Links: Link Cited by: §I. [48] Q. Wang, X. Jiang, L. He, J. Wu, and N. Mesgarani (2025) SightSound-r1: cross-modal reasoning distillation from vision to audio language models. arXiv preprint arXiv:2509.15661. Cited by: §I. [49] A. Wilf, L. Mathur, S. Mathew, C. Ko, Y. Kebe, P. P. Liang, and L. Morency (2023) Social-iq 2.0 challenge: benchmarking multimodal social understanding. GitHub. Note: https://github.com/abwilf/Social-IQ-2.0-Challenge Cited by: §IV-A. [50] J. Xia, Y. Zang, P. Gao, S. Li, and K. Zhou (2025) Visionary-r1: mitigating shortcuts in visual reasoning with reinforcement learning. arXiv preprint arXiv:2505.14677. Cited by: §I, §I. [51] J. Xiao, X. Shang, A. Yao, and T. Chua (2021) Next-qa: next phase of question-answering to explaining temporal actions. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 9777–9786. Cited by: §IV-C. [52] Z. Xing, X. Hu, C. Fu, W. Wang, J. Dai, and P. Heng (2025) Echoink-r1: exploring audio-visual reasoning in multimodal llms via reinforcement learning. arXiv preprint arXiv:2505.04623. Cited by: §I. [53] J. Xu, Z. Guo, J. He, H. Hu, T. He, S. Bai, K. Chen, J. Wang, Y. Fan, K. Dang, et al. (2025) Qwen2. 5-omni technical report. arXiv preprint arXiv:2503.20215. Cited by: §I, §IV-A. [54] L. Xu, Y. Zhao, D. Zhou, Z. Lin, S. K. Ng, and J. Feng (2024) Pllava: parameter-free llava extension from images to videos for video dense captioning. arXiv preprint arXiv:2404.16994. Cited by: TABLE VII. [55] H. Yang, Y. Zhao, Y. Wu, S. Wang, T. Zheng, H. Zhang, Z. Ma, W. Che, S. Wang, S. Wei, et al. (2025) Large language models meet text-centric multimodal sentiment analysis: a survey. Science China Information Sciences 68 (10), p. 1–29. Cited by: §I. [56] J. Yang, R. Tan, Q. Wu, R. Zheng, B. Peng, Y. Liang, Y. Gu, M. Cai, S. Ye, J. Jang, et al. (2025) Magma: a foundation model for multimodal ai agents. In Proceedings of the computer vision and pattern recognition conference, p. 14203–14214. Cited by: TABLE VII. [57] L. Yang, X. Wang, X. Zhou, Z. Wu, and N. Tan (2025) Application of multiple chain-of-thought in contrastive reasoning for implicit sentiment analysis. arXiv preprint arXiv:2503.07140. Cited by: §I. [58] Q. Yang, S. Yao, W. Chen, S. Fu, D. Bai, J. Zhao, B. Sun, B. Yin, X. Wei, and J. Zhou (2025) HumanOmniV2: from understanding to omni-modal reasoning with context. arXiv preprint arXiv:2506.21277. Cited by: §I, §I, §I, §I, §IV-A, TABLE VII. [59] Y. Yao, T. Yu, A. Zhang, C. Wang, J. Cui, H. Zhu, T. Cai, H. Li, W. Zhao, Z. He, et al. (2024) Minicpm-v: a gpt-4v level mllm on your phone. arXiv preprint arXiv:2408.01800. Cited by: §IV-A. [60] J. Ye, H. Xu, H. Liu, A. Hu, M. Yan, Q. Qian, J. Zhang, F. Huang, and J. Zhou (2024) Mplug-owl3: towards long image-sequence understanding in multi-modal large language models. arXiv preprint arXiv:2408.04840. Cited by: TABLE VII. [61] L. Yu, Z. Lin, X. Shen, J. Yang, X. Lu, M. Bansal, and T. L. Berg (2018) Mattnet: modular attention network for referring expression comprehension. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 1307–1315. Cited by: §I, §I. [62] T. Yu, Y. Yao, H. Zhang, T. He, Y. Han, G. Cui, J. Hu, Z. Liu, H. Zheng, M. Sun, et al. (2024) Rlhf-v: towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 13807–13816. Cited by: §I. [63] R. Yuan, C. Xiao, S. Leng, J. Wang, L. Li, W. Xu, H. P. Chan, D. Zhao, T. Xu, Z. Wei, et al. (2025) Vl-cogito: progressive curriculum reinforcement learning for advanced multimodal reasoning. arXiv preprint arXiv:2507.22607. Cited by: §I, §I. [64] A. Zadeh, R. Zellers, E. Pincus, and L. Morency (2016) Multimodal sentiment intensity analysis in videos: facial gestures and verbal messages. IEEE Intelligent Systems 31 (6), p. 82–88. Cited by: §I, §I. [65] B. Zhang, H. Li, T. Zhang, C. Yan, J. Cai, and Y. Hao (2025) Improving the reasoning of multi-image grounding in mllms via reinforcement learning. arXiv preprint arXiv:2507.00748. Cited by: §I. [66] J. Zhang, J. Huang, H. Yao, S. Liu, X. Zhang, S. Lu, and D. Tao (2025) R1-vl: learning to reason with multimodal large language models via step-wise group relative policy optimization. arXiv preprint arXiv:2503.12937. Cited by: §I. [67] Y. Zhang, J. Wu, W. Li, B. Li, Z. Ma, Z. Liu, and C. Li (2024) Llava-video: video instruction tuning with synthetic data. arXiv preprint arXiv:2410.02713. Cited by: TABLE VII. [68] L. Zheng, W. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al. (2023) Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in neural information processing systems 36, p. 46595–46623. Cited by: §I, §I. [69] H. Zhong, M. Zhu, Z. Du, Z. Huang, C. Zhao, M. Liu, W. Wang, H. Chen, and C. Shen (2025) Omni-r1: reinforcement learning for omnimodal reasoning via two-system collaboration. arXiv preprint arXiv:2505.20256. Cited by: §I. [70] Z. Zhou, R. Wang, and Z. Wu (2025) Daily-omni: towards audio-visual reasoning with temporal alignment across modalities. arXiv preprint arXiv:2505.17862. Cited by: §I, §IV-A. [71] M. Zhu, H. Zhong, C. Zhao, Z. Du, Z. Huang, M. Liu, H. Chen, C. Zou, J. Chen, M. Yang, et al. (2025) Active-o3: empowering multimodal large language models with active perception via grpo. arXiv preprint arXiv:2505.21457. Cited by: §I. Appendix G Biography Section Yibo Wang is currently pursuing the M.S. degree at the Gansu Provincial Key Laboratory of Wearable Computing, School of Information Science and Engineering, Lanzhou University, Lanzhou, China. He received the B.S. degree from Dalian University of Technology, Dalian, China. His research interests include multimodal large language models, affective computing, and reinforcement learning. Rui Yang is currently pursuing the Ph.D. degree with the Gansu Provincial Key Laboratory of Wearable Computing, School of Information Science and Engineering, Lanzhou University. His research interests include biometric authentication, affective computing, large language models (LLMs), and embodied intelligence. Jisheng Dang received the Ph.D. degree from Sun Yat-sen University, China, in 2025, advised by prof. Jianhuang Lai and prof. Huicheng Zheng. He worked as a research fellow at the NExT++ laboratory of the National University of Singapore advised by prof. Tat-Seng Chua. He is now a tenured associate professor at the School of Information Science and Engineering, Lanzhou University. His research interests include multimodal learning, video understanding, and embodied intelligence. He has published several papers as the first author in major journals and conferences including IEEE TIP/TNNLS/TITS/IJCAI/AAAI. He served as a reviewer at some major journals and conferences like IEEE TPAMI, ICML, NIPS, ICLR, IEEE TIP, CVPR, IJCAI, ACM M, AAAI, IEEE TMM, IEEE TCSVT, ACM TOMM. Yitao Wu received the B.S. degree in information system and information management from Hainan University. His research interests include large language models, vision-language-action models, and vision-language models. Hong Peng received the Ph.D. degree from Lanzhou University, Lanzhou, China. From 2010 to 2011, he was a Visiting Scholar with the Institute of Computer System, ETH Zurich, Switzerland. He is currently an Associate Professor with the School of Information Science and Engineering, Lanzhou University. He is also in charge of three projects from the National Natural Science Foundation of China, the Central College Foundation Project of Lanzhou University, and the Youth Cross-Project of Lanzhou University. He has authored or coauthored more than 30 papers in peer-reviewed journals, conferences, and book chapters. His research areas include bioinformation processing and ubiquitous affective computing. Bin Hu (Fellow, IEEE) received the PhD degree in computer science from the Institute of Computing Technology, Chinese Academy of Science, China, in 1998. Since 2008, he has been a professor and the dean of the School of Information Science and Engineering, Lanzhou University, China. He had been also guest professorship in ETH Zurich, Switzerland till 2011. He is a Professor of Lanzhou University and Beijing Institute of Technology. He serves as Editor-in-Chief of IEEE Transactions on Computational Social Systems, Fellow of IET and AAIA, and Chair of Technical Committee on Computational Psychophysiology, IEEE SMC. He has published over 300 papers in domestic and international academic journals and conferences. His research interests include pervasive computing, computational psychophysiology, data modeling, and artificial intelligence. Tat-Seng Chua received the Ph.D. degree from the University of Leeds, U.K. He is the KITHCT chair professor with the School of Computing, National University of Singapore, where he was the acting and founding dean of the School from 1998 to 2000. He is the co-director of NExT, a joint center between NUS and Tsinghua University, to develop technologies for live social media search. He is the 2015 winner of the prestigious ACM SIGMM Award. He is the chair of Steering Committee of the ACM International Conference on Multimedia Retrieval (ICMR) and Multimedia Modeling (M) conference series. He is also the general co-chair of ACM Multimedia 2005, ACM CIVR (now ACM ICMR) 2005, ACM SIGIR 2008, and ACM Web Science 2015. He serves on the editorial boards of four international journals. He is the co-founder of two technology startups in Singapore and a Fellow of the Singapore Academy of Sciences, with 107,618 citations on Google Scholar.