Paper deep dive
CNSL-bench: Benchmarking the Sign Language Understanding Capabilities of MLLMs on Chinese National Sign Language
Rui Zhao, Xuewen Zhong, Xiaoyun Zheng, Jinsong Su, Yidong Chen
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 98%
Last extracted: 4/27/2026, 3:09:57 AM
Summary
CNSL-bench is a new multimodal benchmark designed to evaluate the Chinese National Sign Language understanding capabilities of Multimodal Large Language Models (MLLMs). Grounded in the National Common Sign Language Dictionary, it provides 20,121 instances across text, image, and video modalities, covering diverse manual articulatory forms like air-writing, finger-spelling, and the Chinese manual-alphabet. The study evaluates 21 MLLMs, revealing that current models, including advanced proprietary ones like GPT-5, still significantly underperform compared to humans, particularly in video understanding and complex manual articulations.
Entities (8)
Relation Signals (4)
CNSL-bench → evaluates → MLLM
confidence 100% · designed for evaluating multimodal large language models (MLLMs) in sign language understanding
CNSL-bench → includes → Air-writing
confidence 100% · supporting fine-grained analysis across key manual articulatory forms, including air-writing, finger-spelling, and the Chinese manual-alphabet
GPT-5 → isatypeof → MLLM
confidence 100% · The strongest proprietary model, GPT-5, achieves overall accuracies...
CNSL-bench → isgroundedin → National Common Sign Language Dictionary
confidence 100% · The proposed CNSL-bench is characterized by: 1) Authoritative grounding, as it is anchored to the officially standardized National Common Sign Language Dictionary
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Sign language research has achieved significant progress due to the advances in large language models (LLMs). However, the intrinsic ability of LLMs to understand sign language, especially in multimodal contexts, remains underexplored. To address this limitation, we introduce CNSL-bench, the first comprehensive Chinese em{National Sign Language benchmark designed for evaluating multimodal large language models (MLLMs) in sign language understanding. The proposed CNSL-bench is characterized by: 1) Authoritative grounding, as it is anchored to the officially standardized \textit{National Common Sign Language Dictionary, mitigating ambiguity from regional or non-canonical variants and ensuring consistent semantic definitions; 2) Multimodal coverage, providing aligned textual descriptions, illustrative images, and sign language videos; and 3) Articulatory diversity, supporting fine-grained analysis across key manual articulatory forms, including air-writing, finger-spelling, and the Chinese manual-alphabet. Using CNSL-bench, we extensively evaluate 21 open-source and proprietary up-to-date MLLMs. Our results reveal that, despite recent advances in multimodal modeling, current MLLMs remain substantially inferior to human performance, exhibiting systematic disparities across input modalities and manual articulatory forms. Additional diagnostic analyses suggest that several performance limitations persist beyond improvements in reasoning and that instruction-following robustness varies substantially across models.
Tags
Links
- Source: https://arxiv.org/abs/2604.22367v1
- Canonical: https://arxiv.org/abs/2604.22367v1
Trouble viewing inline? Open PDF directly →
Full Text
101,536 characters extracted from source content.
Expand or collapse full text
CNSL-bench: Benchmarking the Sign Language Understanding Capabilities of MLLMs on Chinese National Sign Language Rui Zhao 1,2,3 , Xuewen Zhong 1,2,3 , Xiaoyun Zheng 1,2,3 , Jinsong Su 1,2 and Yidong Chen 1,2,3 * 1 School of Informatics, Xiamen University, China 2 Key Lab of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian-Taiwan (XMU), Ministry of Culture and Tourism, China 3 National Language Resources Monitoring and Research Center for Education and Teaching Media, Xiamen University, China zhsqzr@stu.xmu.edu.cn ydchen@xmu.edu.cn Abstract Sign language research has achieved significant progress due to the advances in large language models (LLMs). However, the intrinsic ability of LLMs to understand sign language, espe- cially in multimodal contexts, remains underex- plored. To address this limitation, we introduce CNSL-bench, the first comprehensive Chinese National Sign Language benchmark designed for evaluating multimodal large language mod- els (MLLMs) in sign language understanding. The proposed CNSL-bench is characterized by: 1) Authoritative grounding, as it is an- chored to the officially standardized National Common Sign Language Dictionary, mitigat- ing ambiguity from regional or non-canonical variants and ensuring consistent semantic def- initions; 2) Multimodal coverage, providing aligned textual descriptions, illustrative images, and sign language videos; and 3) Articula- tory diversity, supporting fine-grained analysis across key manual articulatory forms, includ- ing air-writing, finger-spelling, and the Chinese manual-alphabet. Using CNSL-bench, we ex- tensively evaluate 21 open-source and propri- etary up-to-date MLLMs. Our results reveal that, despite recent advances in multimodal modeling, current MLLMs remain substantially inferior to human performance, exhibiting sys- tematic disparities across input modalities and manual articulatory forms. Additional diagnos- tic analyses suggest that several performance limitations persist beyond improvements in rea- soning and that instruction-following robust- ness varies substantially across models. 1 Introduction Sign language plays a central role in communi- cation for many people with hearing impairment and has consequently attracted sustained attention from the research community over the past decades. * Corresponding Author. 双手直立,掌心左右相对,前后交替移动几下,表示打手语。 Stand with both hands upright, palms facing each other on the left and right, and move them forward and backward alternately a few times to indicate signing. 手语 sign language 计算 compute 语言 language 一手食指横伸,在嘴前后转动两下。 Extend the index finger of one hand horizontally and rotate it twice in front of the mouth. 双手五指微曲,掌心向上,边交替点动边互碰两下。 Bend the fingers of both hands slightly, with the palms facing up. Alternate tapping them while touching each other twice. Figure 1: Examples from CNSL-bench showing the aligned textual description, illustrative image, and cor- responding sign language video for each sign entry. More recently, advances in large language mod- els (LLMs) have further stimulated progress in automatic sign language understanding, primarily within specific tasks, e.g., sign language translation, where LLMs are incorporated as semantic augmen- tation modules or enhanced text decoders for im- proved performance (Wong et al., 2023; Gong et al., 2024; Chen et al., 2024b; Guo et al., 2025; Kim et al., 2025; Liu et al., 2025; Jang et al., 2025; Asasi et al., 2025; Rao et al., 2025; Hwang et al., 2025). Despite encouraging task-level improvements, existing approaches predominantly embed LLMs into downstream pipelines or datasets, leaving the intrinsic ability of these models to understand sign language largely unexamined. This limitation be- comes even more pronounced in the context of mul- timodal large language models (MLLMs), which have demonstrated strong visual-language capa- arXiv:2604.22367v1 [cs.CL] 24 Apr 2026 bilities over images and videos (Liu et al., 2023, 2024; Zhang et al., 2024; Shen et al., 2025). Cru- cially, sign language possesses the full range of fundamental linguistic properties and is inherently multimodal, with meaning expressed through the coordinated use of linguistically grounded man- ual articulatory cues such as air writing, finger- spelling, and the manual-alphabet (Shi et al., 2021; Yin et al., 2021; Desai et al., 2024; Atwell et al., 2024). While this rich expressiveness poses fun- damental challenges for automatic sign language understanding, it also raises an open question: to what extent can current MLLMs genuinely com- prehend sign language by capturing linguistically grounded structure and semantic meanings rather than relying solely on visual correlations? To answer this question, we introduce CNSL- bench, the first comprehensive Chinese National Sign Language benchmark designed for evaluat- ing MLLMs in sign language understanding (§2.1). CNSL-bench is constructed upon the National Common Sign Language Dictionary, an officially standardized lexical resource for Chinese national sign language. This authoritative grounding pro- vides a canonical semantic reference, reducing am- biguity from regional or variants and enabling con- sistent, controlled evaluation of sign language un- derstanding (Ministry of Education of the People’s Republic of China et al., 2018b; China Disabled Persons’ Federation et al., 2019). To enable multi- modal evaluation, we further align these dictionary entries with a large-scale video dataset covering isolated Chinese national sign language (Jin et al., 2025), resulting in a unified benchmark that pro- vides broad multimodal coverage, where each sign entry is represented by aligned textual description, illustrative image, and sign language video, as ex- emplified in Figure 1. As a consequence, CNSL- bench comprises 20,121 questions spanning text, image, and video. Furthermore, the benchmark explicitly supports manual articulatory diversity, covering three representative categories of signs, including air-writing, finger-spelling, and manual- alphabet, as shown in Figure 2. These categories are treated as dedicated evaluation subsets to enable fine-grained analysis. Supported by high-quality resources, carefully designed evaluation protocols, and human assessment involving the Deaf 1 commu- nity, CNSL-bench serves as a reliable and compre- 1 We follow the recognized convention of using the upper- cased word Deaf to refer to the community of sign language users (Woodward, 1972) hensive benchmark for systematically diagnosing the intrinsic sign language understanding capabili- ties of modern MLLMs. With CNSL-bench, we benchmark a wide range of open- and closed-source up-to-date MLLMs, of- fering a systematic empirical assessment of their sign language understanding capabilities (§ 3.2). Our main results reveal several patterns: (a) De- spite recent advances in multimodal modeling, cur- rent MLLMs remain substantially inferior to hu- man performance in sign language understand- ing. (b) A pronounced modality-dependent per- formance imbalance is observed, with models ex- hibiting substantially weaker performance on vi- sual inputs compared to text. (c) MLLMs demon- strate uneven comprehension across different man- ual articulatory forms, achieving relatively stronger performance on finger-spelling than on air-writing and the specific manual-alphabet. (d) The perfor- mance gap between open-source and proprietary MLLMs is rapidly narrowing, with several open- source models achieving performance compara- ble to lightweight commercial systems. Beyond the main results, additional diagnostic analyses are conducted to further investigate the intrinsic sign language understanding capabilities of current MLLMs (§ 3.3). Test-time scaling via explicit rea- soning mechanisms is examined as a diagnostic tool, revealing heterogeneous gains across mod- els and modalities while failing to fundamentally resolve the observed limitations. Prompt token effects and instruction-following robustness are further considered as complementary diagnostic factors, providing additional insight into the behav- ioral characteristics of current MLLMs. In summary, the contributions of this work are as follows: (1) We introduce CNSL-bench, a multimodal, sign language-centric benchmark for evaluating sign language understanding in MLLMs. (2) We present a comprehensive evaluation of up-to-date MLLMs on the proposed CNSL-bench. (3) Through extensive experiments and analy- ses, we identify persistent challenges in current MLLMs’ sign language understanding capabilities. We hope that CNSL-bench will serve as a di- agnostic foundation and a reference resource for future research toward more robust, reliable, and human-aligned MLLMs in sign language under- standing. Data and code are available athttps: //github.com/rzhao-zhsq/CNSL-bench. 避雷针(LightningRod) (一)双手直立,掌心向外推出。 (二)左手食指直立;右手伸食指,在左手上方书空“ϟ”形,然后点一下左手食指尖。 (1)Standwithbothhandsstraightoutinfrontofyou,palmsfacingoutward. (2)Keeptheindexfingerofthelefthandstraightup;extendtheindexfingeroftherighthandandwritea“ϟ” (lightning)shapeintheairwiththelefthand,thentouchthetipoftheleftindexfinger. 北斗星(ThePlough) (一)双手伸拇、食、中指,手背向外,手腕交叉相搭,仿“北”字形。 (二)一手拇、食指搭成“十”字形,在头前上方做勺形移动,仿北斗七星的形状,眼睛注视手的动作。 (1)Stretchoutthethumb,indexfinger,andmiddlefingerofbothhands,withthebackofthehandsfacingoutward. Crossthewristsandplacethemontopofeachother,imitatingtheshapeofthecharacter"北"(north). (2)Withonehand,forma"十"(plus)shapewiththethumbandindexfinger.Moveitinaspoon-likemotionabove thehead,imitatingtheshapeoftheBigDipper.Keepyoureyesonthemovementofyourhand. 二氧化碳(CO 2 ) 左手打手指字母“C”的指式;右手先打手指字母“O”的指式,再在右下方打数字“2”的手势,表示二氧 化碳的化学分子式。 ThelefthandmakestheChinesemanualalphabet"C";therighthandfirstmakesthesignlanguagehandshape "O",thenmakesthegestureofthenumber"2"atthelowerright,indicatingthechemicalformulaofcarbondioxide. Figure 2: Examples illustrating the three categories of sign articulation: air-writing (top), finger-spelling (middle), and manual-alphabet (bottom). 2 CNSL-bench Sign language understanding requires unified mod- eling of visual, temporal, and fine-grained linguis- tic information across heterogeneous modalities. To enable a systematic and controlled evaluation of such capabilities in MLLMs, we construct CNSL- bench, a Chinese National Sign Language bench- mark grounded in standardized lexical resources and aligned multimodal representations. CNSL- bench is designed to isolate intrinsic sign language understanding from downstream task-specific fac- tors, providing a consistent evaluation framework across text, image, and video inputs. 2.1 Benchmark Construction Benchmark Principle. CNSL-bench is con- structed following three core principles: 1) it is grounded in a standardized lexical foundation. To avoid ambiguity introduced by regional or non- canonical sign variants, the benchmark is anchored to the officially standardized National Common Sign Language Dictionary, ensuring consistent se- mantic definitions across all samples; 2) it em- phasizes aligned multimodal coverage. Each sign entry is systematically represented in text, image, and video formats, enabling controlled evaluation of cross-modal sign language understanding; and 3) the benchmark explicitly incorporates diverse manual articulatory forms, including air-writing, finger-spelling, and specific manual-alphabet signs, supporting fine-grained analysis of linguistically distinct forms in sign language. Data Collection.CNSL-bench is constructed by grounding all entries in officially standardized Chi- nese national sign language resources and aligning them with multimodal representations. The core lexical inventory is derived from the National Com- mon Sign Language Dictionary (China Disabled Persons’ Federation et al., 2019), which is built upon the Lexicon of Common Expressions in Chi- nese National Sign Language jointly issued by the Ministry of Education, the State Language Com- mission, and the China Disabled Persons’ Feder- ation (Ministry of Education of the People’s Re- public of China et al., 2018b). Together, these standards define normative sign realizations that are widely used and stable in education and daily communication. This grounding ensures lexical consistency and reduces regional or informal varia- tion. To ensure that each benchmark item maps to a unique sign entry, we perform systematic sign-level preprocessing to handle dictionary cases where (i) distinct entries share identical hand motions, (i) identical entries are associated with different mean- ings or articulations (e.g., “seat belt” for car and airplane), or (i) the same meaning is realized by distinct hand motions. For each entry, we retain the accompanying textual description and illustra- tive image, which explain the sign entry for educa- tional and communication purposes. Furthermore, each unique sign entry is aligned with an additional sign language video to reflect real-world usage. The video samples are sourced from a large-scale Chinese national sign language dataset (Jin et al., 2025), yielding a unified multimodal representation (a) Overall performance on CNSL-bench.(b) Performance of subsets on CNSL-bench. Figure 3: A concise summary of the MLLMs’ performance on CNSL-bench. (a) provides a comparison detailing the overall performance of the 16 selected MLLMs, including both open-source and closed-source models. (b) illustrates a radar chart that outlines the performance of humans and MLLMs on three subsets (air-writing, finger-spelling, and manual-alphabet) within CNSL-bench. for each sign entry. In addition, CNSL-bench explicitly incorporates specific manual articulatory forms, including air- writing, finger-spelling, and the Chinese manual- alphabet. As exemplified in Figure 2, air-writing refers to tracing graphic forms in the air, and the traced content may correspond to strokes, symbol- like shapes, or a partial character. In our taxonomy, finger-spelling refers to using one or both hands to depict or indicate Chinese character-form structure, prioritizing graphic cues such as outlines, compo- nents, and structural patterns over strictly sequen- tial, letter-by-letter spelling. The manual-alphabet follows the Chinese Manual Alphabet (Ministry of Education of the People’s Republic of China et al., 2018a), which maps conventionalized finger config- urations to individual Chinese Pinyin letters. These letters can be combined to spell Mandarin, form- ing lexical signs, and functioning as morphemic components within signs under the Scheme of the Chinese Phonetic Alphabet (Committee for Lan- guage Reform of China, 1957). The alignment between text, image, and video is detailed in Appendix A.1, and the complete man- ual alphabets in the Chinese Manual Alphabet are listed in Figure 8 (Appendix A.2). Task Definition. For each sign entry, CNSL- bench provides an aligned textual description, an illustrative image, and a sign language video. We formulate the task as a four-way multiple-choice evaluation, where each instance is instantiated with a single input modality. This closed-form design enables controlled and scalable evaluation, as cur- rent state-of-the-art MLLMs remain highly unre- liable in open-ended sign language understanding. We further examine different option construction strategies and observe that semantics-based distrac- tors lead to slightly weaker model performance, while yielding conclusions consistent with random sampling. To simplify the benchmark design and facilitate easy reproduction, we therefore adopt random option sampling for robustness and consis- tency. Detailed task specifications and comparative analyses of option construction strategies, includ- ing an open-ended case and distractor design, are provided in Appendix A.3. 2.2 Benchmark Statistics CNSL-bench is grounded in the National Common Sign Language Dictionary, which contains 8,214 sign glosses. However, it includes cases where dif- ferent glosses share identical hand motions, identi- cal glosses correspond to different meanings and ar- ticulations, or the same meaning is realized through multiple distinct hand motions. After processing at the sign-entry level, CNSL-bench retains 6,707 unique sign entries, yielding 20,121 evaluation in- stances across three modalities (text, image, and video). Among the 6,707 sign entries, we manually identify 407 entries containing air-writing, 77 con- taining finger-spelling, and 592 involving specific Model TextImageVideo ′ 2 f ps Video ′ 10 f ps AWFSMAAllAWFSMAAllAWFSMAAllAWFSMAAll Open&Close -source Image MLLMs LLaVA-NeXT-7B0.742.600.681.6817.2014.2917.5717.5215.7212.9917.7418.67---- Qwen-VL-Plus87.7184.4277.9772.4125.3728.5725.8938.8317.6927.2721.8332.0621.3822.0823.9933.95 Qwen-VL-Max85.5083.1277.8072.6224.8833.7724.8740.1820.6423.3822.0031.0124.0831.1723.9934.43 Open-Source MLLMs Qwen2-VL-2B47.4251.9541.5543.3619.4119.4823.6530.6220.3927.2722.8027.2319.1618.1823.6527.58 Qwen2.5-VL-3B69.0472.7355.4160.0725.8027.2723.8234.2620.6424.6817.9128.3623.5928.5720.6130.34 Intern-VL-3.5-2B68.0679.2260.6458.6819.6628.5724.4933.372.956.494.394.294.426.494.394.50 Qwen3-VL-2B-Instruct73.7183.1258.6162.8321.6238.9621.6234.0119.6618.1821.7927.7821.3819.4823.1430.18 Qwen3-VL-2B66.5874.0352.7057.9723.5927.2720.7832.8020.3916.8817.7426.3822.3623.3818.0728.95 Qwen3-VL-2B✍72.9775.3256.2561.2926.5441.5621.2835.3716.7115.5819.7627.7622.3627.2720.9530.68 LLaVA-NeXT-Video-7B1.722.600.841.348.8510.3912.1612.9414.9918.1815.2015.4313.7614.2915.7115.91 Qwen2-VL-7B68.3075.3256.9358.9826.7831.1725.1732.4422.6029.8721.4529.2423.1027.2723.1431.19 Qwen2.5-VL-7B65.6070.1355.4157.6126.2927.2724.8333.3221.3824.6820.9529.1022.1123.3821.7930.15 GLM-4.1V-9B✍80.3484.4269.59 68.2428.5055.8428.3839.6220.3923.3819.7628.0321.8724.6821.6229.75 Intern-VL-3.5-8B84.7783.1271.7967.5330.4750.6527.3638.3634.1544.1632.2632.2631.7036.3629.3933.59 Qwen3-VL-8B-Instruct84.2879.2270.1067.0627.2735.0622.8038.3921.8722.0823.1430.9425.8024.6828.3833.89 Qwen3-VL-8B81.8280.5267.4064.6526.0428.5725.6836.3422.6015.5819.9328.5120.3920.7822.1330.42 Qwen3-VL-8B✍87.2283.1278.8970.6127.0329.8723.6537.5622.1118.1819.9329.8824.5725.9725.0033.00 Closed-Source MLLMs GPT-4o-mini82.8083.1267.9167.5725.0638.9627.0335.9928.7519.4823.4827.30---- GPT-4o82.0688.3172.1369.0332.4351.9529.7339.0727.7620.7826.0131.2625.3123.3821.9628.43 Qwen3-VL-Plus90.9188.3178.1476.6828.5732.4727.0743.6918.4320.7825.2133.7424.5724.6828.2137.37 Qwen3-VL-Plus✍92.3889.6184.2476.2231.7738.9629.9542.4126.8518.1825.9335.3424.5718.1825.1736.92 Gemini-2.5-Falsh85.0188.3172.6473.0430.4731.9345.5843.5728.2623.3830.0736.6226.0424.6829.3937.44 Gemini-2.5-Falsh✍92.3893.5183.1179.9534.6451.9535.4751.6230.4733.7731.9342.2832.4332.4736.3242.63 Gemini-2.5-Pro✍ 93.3794.8192.2384.7950.3774.0353.2161.1335.3844.1646.1148.3236.6135.0639.0248.35 GPT-5✍96.8197.4097.1389.6459.2181.8259.4666.9644.7242.8646.9653.4246.9353.2553.5556.72 Random25.3527.2724.6925.2324.5724.6724.4924.7325.7923.9824.9225.0324.5724.6724.7625.04 Human98.7797.4096.9696.9398.7798.7098.3197.3999.2698.7097.4797.3999.2698.7097.4797.39 Table 1: Performance of MLLMs on CNSL-bench. AW, FS, and MA indicate air-writing, finger-spelling, and manual-alphabet.✍denotes inference with slow thinking. The best result is bolded, and the second isunderlined. manual-alphabet, enabling dedicated subset evalua- tion for sign-linguistical sensitive analysis. More- over, a sign entry may comprise multiple atomic gestures due to sequential articulation or multi-part realizations, with up to 7 gestures in the most com- plex cases. The detailed statistics and breakdowns are provided in Appendix A.4. Overall, CNSL-bench comprises 20,121 ques- tions spanning text, image, and video modalities, with each question grounded in a standardized lex- ical entry and aligned multimodal evidence. This construction supports systematic and fine-grained evaluation of sign language understanding under a unified benchmark setting. 3 Experiments 3.1 Experimental Settings We evaluate 21 up-to-date MLLMs, spanning open-source families (LLaVA-NeXT, Qwen-VL, InternVL-3.5, GLM-4.1V) and closed-source mod- els (Qwen-Plus/Max, Gemini-2.5, GPT-4/5). De- tailed inference configurations and the human eval- uation protocol are provided in Appendix B.1. 3.2 Main Results The overall performance and subcategory com- parisons (human vs. representative MLLMs) on CNSL-Bench can be quickly glanced at Figure 3. Table 1 details the overall performance of a wide range of MLLMs on CNSL-bench across modali- ties (i.e., text, image, and sign video) and manual articulatory forms (i.e., AW, FS, and MA). Human–MLLMs performance gap.Despite re- cent progress, current MLLMs remain markedly inferior to human-level sign language understand- ing across all modalities. The strongest propri- etary model, GPT-5, achieves overall accuracies of 89.64%, 66.96%, and 56.72% on text, image, and sign language video understanding, respec- tively. Although GPT-5 substantially outperforms other advanced models, a clear and persistent gap remains when compared to human performance, which consistently reaches approximately 97% across all three modalities. This discrepancy high- lights the fundamental difficulty of achieving ro- bust, human-level comprehension of sign language, particularly in visually grounded and temporally Model TextImageVideo ′ 2 f ps Video ′ 10 f ps AWFSMAAllAWFSMAAllAWFSMAAllAWFSMAAll Fast Thinking Qwen3-VL-2B66.5874.0352.7057.9723.5927.2720.7832.8020.3916.8817.7426.3822.3623.3818.0728.95 Qwen3-VL-8B81.8280.5267.4064.6526.0428.5725.6836.3422.6015.5819.9328.5120.3920.7822.1330.42 Qwen3-VL-Plus90.9188.3178.1476.6828.5732.4727.0743.6918.4320.7825.2133.7424.5724.6828.2137.37 Gemini-2.5-Flash85.0188.3172.6473.0430.4731.9345.5843.5728.2623.3830.0736.6226.0424.6829.3937.44 Slow Thinking Qwen3-VL-2B72.9775.3256.2561.2926.5441.5621.2835.3716.7115.5819.7627.7622.3627.2720.9530.68 Qwen3-VL-8B87.2283.1278.8970.6127.0329.8723.6537.5622.1118.1819.9329.8824.5725.9725.0033.00 Qwen3-VL-Plus92.3889.6184.2476.2231.7738.9629.9542.4126.8518.1825.9335.3424.5718.1825.1736.92 Gemini-2.5-Flash92.3893.5183.1179.9534.6451.9535.4751.6230.4733.7731.9342.2832.4332.4736.3242.63 Gemini-2.5-Pro (L)91.1596.1086.1581.3246.6872.7352.2058.0937.5942.8644.7648.8334.6442.8641.7247.96 Gemini-2.5-Pro (M)93.3794.8192.2384.7950.3774.0353.2161.1335.3844.1646.1148.3236.6135.0639.0248.35 Gemini-2.5-Pro (H)92.6394.8190.7184.8449.6374.0354.0561.9236.6138.9642.4048.1738.5736.3645.1048.59 GPT-5 (L)97.0597.4095.6188.9455.5374.0360.6466.7740.5438.9646.4551.8941.0348.0545.1053.13 GPT-5 (M)96.8197.4097.1389.6459.2181.8259.4666.9644.7242.8646.9653.4246.9353.2553.5556.72 GPT-5 (H)97.0598.7096.6289.9563.1483.1259.4668.3442.5137.6648.9953.0944.5838.9646.4554.01 Table 2: Numerical results with test-time scaling on reasoning models. L, M, H: low, medium, and high reasoning effort on the process of thinking before generating an answer. Figure 4: CoT gains across models and multimodal subsets. complex settings. Modality–dependent performance imbalance. A pronounced modality imbalance is consistently observed in MLLMs’ sign language understand- ing. Across nearly all models, performance is high- est for textual descriptions, while accuracy drops substantially for illustrative images and further de- grades for sign language videos. This trend holds for both open-source and closed-source systems and becomes more pronounced in the video setting. These results indicate that, although text-centric language modeling is relatively well developed, robust visual grounding and temporal modeling remain major challenges for current MLLMs. Uneven understanding across manual articu- latory forms. MLLMs exhibit uneven compre- hension across different manual articulatory forms. As shown in Table 1, models consistently achieve stronger performance on finger-spelling than on air- writing and manual-alphabet across all modalities. This disparity suggests that current models handle more discrete and character-like sign components more reliably, while continuous, shape-intensive, or motion-dependent articulations remain difficult. Narrowing gap between open- and closed-source MLLMs. The performance gap between open- source and proprietary MLLMs is rapidly nar- rowing. Several small-scale open-source mod- els, including GLM-4.1V-9B, InternVL-3.5-8B, and Qwen3-VL-8B, attain performance compara- ble to lightweight commercial systems such as GPT-4o-mini and Gemini-2.5-Flash across multi- ple modalities. These models even surpass propri- etary counterparts on specific subsets, underscoring the rapid advancement and increasing competitive- ness of open-source MLLMs in sign language un- derstanding. Nevertheless, despite these encourag- ing trends, open-source models still need to make continued progress on challenging subsets and to match higher-capacity proprietary systems, partic- ularly under multimodal settings. Figure 5: Top-tier reasoning models, i.e., GPT-5 and Gemini-2.5-Pro, exhibit a boundary effect in test-time scaling as reasoning effort increases from low to high, especially in video input settings. Figure 6: KDE plot of reasoning tokens. Videos are at 10 frames per second. Zoom in for better visualization. 3.3 Additional Analysis Test–Time Scaling. To better contextualize em- pirical findings from the main results, we examine test-time scaling (reasoning) via CoT on MLLMs with explicit reasoning capabilities.We com- pare fast-thinking and slow-thinking settings across modalities and manual articulatory forms to assess when additional reasoning is beneficial for sign language understanding. After carefully analyzing these results, we draw several conclusions: (1) CoT gains vary substantially across mod- els. As shown in Table 2 and Figure 4, CoT gains vary substantially across models and sub- sets. Gemini-2.5-Flash consistently benefits most from slow thinking, achieving an average improve- ment of 6.45%, whereas Qwen3-VL-Plus exhibits negative gains on several subsets, including tex- tual descriptions, illustrative images, and sign lan- guage videos (with 10 frames per second). The repeated evaluations in Qwen3-VL-Plus yield con- sistent negative results, suggesting that reasoning does not universally improve performance and may be sensitive to model-specific inference behavior. (2) CoT gains plateau for top-tier models. Fig- ure 5 shows that for top-tier models such as Gemini- 2.5-Pro and GPT-5, increasing reasoning effort from low to high yields marginal or no gains, and can even degrade performance on sign language video inputs. This indicates a boundary effect of test-time scaling, where additional reasoning to- kens no longer translate into improved understand- ing once a strong baseline is reached. (3) Reasoning quality differs by modality. Fig- ure 6 reveals that models exhibit large variance in reasoning length and modality preference. The rea- soning tokens vary by nearly an order of magnitude across models under textual input. The Qwen3-VL and Gemini-2.5-Pro tend to allocate significantly more reasoning effort to text than to image or video inputs. Such modality-dependent behavior indi- cates incomplete multimodal alignment, which is especially problematic given the inherently visual- language nature of sign language. (4) Reasoning length correlates with task diffi- culty. MLLMs tend to generate longer reasoning chains for incorrectly answered instances. The ra- tio between CoT length for incorrect versus correct predictions of GPT-5-M with text input reaches up to 2.89. This pattern is consistent across most mod- els and modalities, suggesting that models spend more effort on harder cases, partially mirroring human problem-solving behavior. Please see Ap- pendix C.1 for detailed analysis. These results indicate that CoT primarily serves as a diagnostic lens rather than a universal perfor- mance enhancer, revealing fundamental challenges in multimodal reasoning and sign language-specific understanding that persist in MLLMs. Prompt Tokens & Instruction Following. Be- yond reasoning-oriented analyses, we further ex- amine prompt token consumption and instruction- following robustness as complementary diagnostic factors for sign language understanding in MLLMs. 笔记本(Laptop) (一)双手五指弯曲,指尖朝下,交替点动几下,如敲击计算机键盘状。 (二)双手横伸,掌心相贴,然后右手做向上打开的动作。 (1)Bendthefingersofbothhands,withthefingertipspointingdown,and alternatelytapafewtimesasifstrikingthekeysofacomputerkeyboard. (2)Stretchbothhandsouthorizontally,withthepalmstouchingeachother, andthenmakeanupwardopeningmotionwiththerighthand. ......选项中有A.笔记本电脑,B.勇士,C.因为(由于),D.传讯。......“因为”可能用手势表示为“手在胸前做某种动作,可能类似表示原因”的动作。观 察视频中的动作:从6秒开始,双手动作,可能是在做“因为”的手势。比如,手在胸前做交叉(Cross)或指向等动作,这可能对应“因为”。再看选项, C是“因为(由于)”。A笔记本电脑通常用手势表示电子设备,可能不匹配;B勇士可能有特定手势,比如手臂伸展;D传讯可能涉及传递信息的动作, 但根据视频中的动作描述(6-11秒的手势),更可能的是“因为”。需要确认标准手语中“因为”的手势。常见的“因为”手语是:一只手放在胸前,另 一只手做类似“原因”的手势,比如手掌朝上,手指弯曲,然后做某种表示因果的动作。视频中第6-10秒的动作可能对应这个。所以正确选项应该是C。 用户现在需要解决的问题是判断视频中手语动作对应的选项。首先得回忆常见手语动作的含义。选项 C“ 因为(由于)”的手语动作特征,或者看动作的对 应。但可能需要分析动作流程,不过通常这种选择题如果有正确答案的话,需要确定。假设正确选项是 C ?不对,可能需要再想。或者可能我之前有记忆, C 选项“因为(由于)”的手语动作表现。所以最终选 C ? Qwen3- VL-2B Qwen3- VL-8B GLM-4.1V Gemini- 2.5-Flash Gemini- 2.5-Pro 视频中,人物先是站立,然后开始做手语动作。......选项里有“因为(由于)”这样的词。......首先分析手语动作:在视频中,手语动作可能是“因为” 或者“由于”的手语。比如,手语中“因为”(由于)的手语是通过手指的动作来表达,比如用双手的特定手势来表示原因。现在看选项,C选项是“因 为(由于)”,所以需要确定是否符合。根据常见手语,比如“因为”或“由于”的手语动作,通常会有类似双手交叉(Cross)或者特定的手势。根据 视频中的动作,人物先做出双手分开,然后交叉(Cross),这可能是“因为”的手语。所以,正确选项应该是C。 ......Thesignerstartswithrelaxedhandsattheirsides.Atthetwo-secondmark,thehandsmoveup,palmsfacingeachother.Then,atthreeseconds,thingsget interesting.Thelefthandformsaflatpalm,facingupwards,andtherighthandformsaflatpalmfacingdownwards.And,importantly,therighthand tapsthelefthandonce.Afterthat,thehandsreturntotherelaxedstartingposition.......Tappingoneflathandontopofanother,orplacingoneflat handonanother,usuallysignifies"because,""reason,"or"dueto."......A: 笔记本电脑 (Laptopcomputer)No,that'susuallyshownbyforminga "laptop"shapewiththehandsorusingakeyboardmotion.Doesn'tmatch.B: 勇士 (Warrior)Definitelynot.Thatwouldinvolvegesturesofstrength, fighting,orholdingaweapon.Doesn'tfit.C: 因为(由于) (Because/Dueto)Bingo!Thisgesturealignsperfectlywiththehandmovementinthevideo.The foundation/reasonconceptisrightontarget.D: 传讯 (Transmitamessage/Message)Doesn'tseemright.Thatwouldinvolvea"sending"gestureor somethingsimilartowriting.Therefore,myconclusionisthatthesigninthevideomostlikelyrepresents**C: 因为(由于) (Because/Dueto). ......First,Ineedtoreallylookatthevideoitself.Handsdowntostart,thentheycomeup.Ah,palmsdown,fingersextended,andtogether.Theright handgoesonthebackofthelefthand,formingasortofbase,right?It'sheldinfrontofthebody,abitlowerthanchestheight.Thisreallyscreams "foundation"tome.......A.LaptopComputer:Nope.Thesignfor"laptop"justdoesn'tinvolvethishandshapeormovement.It'smuchmoredynamic, mimickingopeningandclosingalaptop.That'sout.B.Warrior/Brave:Noway."Brave"usuallyhasafistgesture,nearthechestorshoulder,signifyingstrength. Totallydifferent.C.Because/DueTo:Waitasecond...That'sit!That'sexactlythesign.Thelefthandisthefoundation,andtherighthandisplacedonit, showingthebasisforsomething.Yes,that'sspoton.D.TransmitaMessage/Summon:Nope.Thesignfor"transmit"usuallyinvolvesamovementawayfrom thefaceorhead,showingthesendingofinformation.Thisisn’tthatatall.Soit'sgottabe"Because/DueTo",optionC. A. Laptop A.笔记本电脑 B. Warrior B. 勇士 C. Because/Due to C. 因为(由于) D. Transmit D.传讯 Figure 7: Video-based sign language understanding failure for “laptop”. Despite correctly recognizing the typing motion in a continuous sign language video, models fail to recover the intended meaning, indicating that temporal visual recognition alone is insufficient for reliable sign language understanding. Prompt token usage varies substantially across in- put modalities and model families, with image and video inputs incurring orders-of-magnitude higher token consumption than text, potentially constraining effective context and reasoning bud- gets. In addition, instruction-following robustness differs markedly across models and modalities: while most large-scale models remain stable, sev- eral smaller-capacity models and certain MLLM families exhibit pronounced failures, particularly on sign language videos. Notably, explicit rea- soning mechanisms may also interact non-trivially with instruction adherence. Due to the space limita- tion, the detailed quantitative analyses are provided in Appendix C.2 and Appendix C.3, respectively. 3.4 Case Studies To provide intuitive insights into the sources of observed performance differences and further il- lustrate the limitations revealed by the CNSL- benchmark, we present a case study in Figure 7. A primary observation is the pronounced modal- ity sensitivity. For the concept laptop, the tex- tual description explicitly mentions “striking keys,” allowing models to easily deduce the correct an- swer. However, under video inputs, even advanced models (e.g., Gemini-2.5-Pro, Qwen3-VL) fail to ground the visual motion correctly. Instead of rec- ognizing the “typing” and “opening” gestures, they hallucinate unrelated motions (e.g., “crossing” ges- tures in Qwen3-VL-8B or “foundation” gestures in Gemini-2.5-Pro) and consistently misclassify the sign as “Because” (Option C). This contrast high- lights that while MLLMs possess strong textual rea- soning, their ability to parse complex temporal vi- sual cues in sign language remains fragile. Beyond modality effects, our analysis reveals challenges in implicit semantic association (e.g., “smell” vs “air”) and emerging symbolic understanding (e.g., correctly mapping visual signs to Chinese charac- ters), see cases in Appendix C.4. 4 Related Works 4.1 Sign Language Understanding Over the past few decades, the sign language research community has primarily focused on sign language recognition (SLR), emphasizing the identification of gloss-level units 2 in isolated words (Joze and Koller, 2019; Li et al., 2020) 2 Glosses are spoken-language textual units that approxi- mately capture the meaning of sign language. or in continuous sequences (Koller et al., 2015; Huang et al., 2018). The increasing availability of large-scale sign language translation datasets (Cam- goz et al., 2018; Zhou et al., 2021; Duarte et al., 2021; Tanzer and Zhang, 2024) has further driven progress in end-to-end sign language translation (SLT) (Camgoz et al., 2020; Chen et al., 2022; Fu et al., 2024; Zhao et al., 2024; Zhang et al., 2025; Fu et al., 2025a). In parallel, large language models (LLMs) have become general-NLP purpose backbones for a wide range of NLP tasks (Brown et al., 2020; Touvron et al., 2023), motivating re- cent efforts to incorporate them into sign language research for their strong language modeling and generation capabilities. In particular, prior work has leveraged LLMs either as enhanced text de- coders (Wong et al., 2023; Gong et al., 2024; Liu et al., 2024) or as semantic enhancement modules to improve SLT systems (Guo et al., 2025; Kim et al., 2025; Liu et al., 2025; Jang et al., 2025). Existing work in this line is largely centered on task- or dataset-specific adaptation of LLMs within sign language pipelines. In contrast, system- atic evaluation of models’ intrinsic sign language understanding, particularly for MLLMs operating directly on images and videos, has received compar- atively less attention. Different from focusing on specific downstream tasks or datasets, we propose CNSL-bench, a comprehensive Chinese National Sign Language benchmark designed for evaluating MLLMs in sign language understanding. 4.2 MLLM benchmarks Recent progress in multimodal large language mod- els (MLLMs) has coincided with the establishment of standardized benchmarks designed to evaluate multimodal understanding. Early efforts primar- ily focused on image-based evaluation, including visual question answering, caption-based reason- ing, and broader vision-language understanding benchmarks that assess perception, grounding, and semantic reasoning over static images (Fu et al., 2025b; Yue et al., 2024; Li et al., 2024; Chen et al., 2024a). More recently, the community has extended benchmarking to video-centric settings, introducing datasets and evaluation protocols that emphasize temporal grounding, event understand- ing, long-context reasoning, and multimodal dia- logue over videos (Maaz et al., 2024; Zhou et al., 2025; Fu et al., 2025c). Despite their broad coverage, most existing MLLM benchmarks are designed for general- domain image and video understanding, where the visual content and semantics are dominated by ev- eryday objects, scenes, actions, and events. Bench- marks that explicitly target sign language remain relatively scarce, despite the need for fine-grained modeling of hand articulation, motion trajectories, and linguistically grounded semantics. In contrast, CNSL-bench is constructed as a dedicated eval- uation benchmark for sign language understand- ing, enabling systematic assessment of MLLMs under aligned textual descriptions, illustrative im- ages, and sign language videos. 5 Conclusion In this work, we introduce CNSL-bench, a multi- modal benchmark centered on sign language for evaluating sign language understanding in MLLMs, with authoritative grounding in officially standard- ized sign language resources. Through extensive evaluation of a variety of open- and closed-source models, we systematically uncover persistent chal- lenges in current MLLMs’ ability to comprehend sign language across modalities and manual articu- latory forms. We anticipate that CNSL-bench will serve as both a diagnostic foundation and a ref- erence resource for future research toward more robust, reliable, and human-aligned MLLMs. Limitations CNSL-bench focuses on lexical-level canonical sign understanding and adopts a multiple-choice formulation to enable controlled, scalable, and re- producible evaluation across modalities. While this design does not directly assess open-ended sign language generation, it is motivated by the obser- vation that current MLLMs remain unreliable in interpreting free-form sign language. This limita- tion highlights a fundamental gap between existing model capacities and the demands of open-ended sign interpretation, underscoring the necessity of establishing a diagnostic benchmark target at sign language understanding. In addition, CNSL-bench is centered on Chinese National Sign Language and emphasizes authoritative semantic grounding over linguistic breadth, and thus does not cover cross- linguistic, regional, or dialectal variation present in other sign languages. Extending evaluation to mul- tilingual and multi-regional sign languages, as well as to more open-ended and compositional settings, remains an important direction for future work. Ethical Considerations Data Access. CNSL-bench is constructed by aligning publicly available resources with recently released sign language video datasets and does not involve new data collection or additional human participants. The textual descriptions and illustra- tive images are sourced from officially published materials released by the Ministry of Education of the People’s Republic of China and are publicly accessible for educational and general communi- cation purposes. The sign language videos are derived from an open-source dataset whose data collection and public release were approved by the Ethical Review Board of Leshan Normal Univer- sity, with all participants providing informed con- sent for the use and publication of their identity information and recordings (Ethical Review Num- ber: LSNU-KYLL2025-02-15) (Jin et al., 2025). Human Participant. Our benchmark involves the human assessment, and we invited a profes- sional team consisting of one professor specializing in sign language linguistics and three sign-language students (including one hearing-impaired student). Each student has at least one year of classroom studying experience in sign language; their instruc- tors include the invited professor and Deaf sign language teachers from a local special education institute. Participants are compensated with $30 to complete each task (about two hours of work). Overall, one expert and three students are engaged to fulfill the human assessment tasks. Acknowledgements We are grateful for the efforts and time of the re- viewers and the committee. This work was sup- ported in part by the National Natural Science Foundation of China under Grant 62476232, Grant 62076211, and in part by First Batch of Projects for the 2025 “Intergovernmental International Science, Technology and Innovation Cooperation” of the National Key Research and Development Program of China under Grant 2025YFE0121700. References Sobhan Asasi, Mohamed Ilyas Lakhal, Ozge Mer- canoglu Sincan, and Richard Bowden. 2025. Beyond gloss: A hand-centric framework for gloss-free sign language translation. Preprint, arXiv:2507.23575. Katherine Atwell, Danielle Bragg, and Malihe Alikhani. 2024. Studying and mitigating biases in sign lan- guage understanding models. In Proceedings of the 2024 Conference on Empirical Methods in Natu- ral Language Processing, pages 268–283, Miami, Florida, USA. Association for Computational Lin- guistics. Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023. Qwen-vl: A versatile vision- language model for understanding, localization, text reading, and beyond. Preprint, arXiv:2308.12966. Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhi- fang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, and 45 others. 2025a. Qwen3-vl technical report. Preprint, arXiv:2511.21631. Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wen- bin Ge, Sibo Song, Kai Dang, Peng Wang, Shi- jie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, and 8 oth- ers. 2025b. Qwen2.5-vl technical report. Preprint, arXiv:2502.13923. Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, and 12 others. 2020. Language models are few-shot learners. In Proceedings of the 34th International Conference on Neural Informa- tion Processing Systems, NIPS’20, pages 1877–1901, Red Hook, NY, USA. Curran Associates Inc. Necati Cihan Camgoz, Simon Hadfield, Oscar Koller, Hermann Ney, and Richard Bowden. 2018. Neural sign language translation. In 2018 IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 7784–7793. Necati Cihan Camgoz, Oscar Koller, Simon Hadfield, and Richard Bowden. 2020. Multi-channel trans- formers for multi-articulatory sign language transla- tion. In Computer Vision – ECCV 2020 Workshops, Lecture Notes in Computer Science, pages 301–319, Cham. Springer International Publishing. Qiguang Chen, Libo Qin, Jin Zhang, Zhi Chen, Xiao Xu, and Wanxiang Che. 2024a. M^3cot: A novel benchmark for multi-domain multi-step multi-modal chain-of-thought. In Proceedings of the 62nd Annual Meeting of the Association for Computational Lin- guistics (Volume 1: Long Papers), pages 8199–8221, Bangkok, Thailand. Association for Computational Linguistics. Yutong Chen, Fangyun Wei, Xiao Sun, Zhirong Wu, and Stephen Lin. 2022. A simple multi-modality transfer learning baseline for sign language translation. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5110–5120. Zhigang Chen, Benjia Zhou, Jun Li, Jun Wan, Zhen Lei, Ning Jiang, Quan Lu, and Guoqing Zhao. 2024b. Factorized learning assisted with large language model for gloss-free sign language translation. In Proceedings of the 2024 Joint International Confer- ence on Computational Linguistics, Language Re- sources and Evaluation, pages 7071–7081, Torino, Italia. ELRA and ICCL. China Disabled Persons’ Federation, China Association of Persons with Hearing Disabilities, and National Center for Sign Language and Braille. 2019. Lexicon of Expressions in Chinese National Sign Language. Huaxia Publishing House, Beijing, China. Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Mar- cel Blistein, Ori Ram, Dan Zhang, Evan Rosen, Luke Marris, Sam Petulla, Colin Gaffney, Asaf Aharoni, Nathan Lintz, Tiago Cardal Pais, Henrik Jacobs- son, Idan Szpektor, Nan-Jiang Jiang, and 3416 oth- ers. 2025. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. Preprint, arXiv:2507.06261. Committee for Language Reform of China. 1957. Scheme of the chinese phonetic alphabet. In The 5th Session of The 1st National People’s Congress, Beijing, China. Aashaka Desai, Maartje De Meulder, Julie A. Hochge- sang, Annemarie Kocab, and Alex X. Lu. 2024. Sys- temic biases in sign language ai research: A deaf-led call to reevaluate research agendas. In Proceedings of the LREC-COLING 2024 11th Workshop on the Representation and Processing of Sign Languages: Evaluation of Sign Language Resources, pages 54– 65, Torino, Italia. ELRA and ICCL. Amanda Duarte, Shruti Palaskar, Lucas Ventura, Deepti Ghadiyaram, Kenneth DeHaan, Florian Metze, Jordi Torres, and Xavier Giro-i-Nieto. 2021. How2sign: A large-scale multimodal dataset for continuous ameri- can sign language. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog- nition, pages 2735–2744. Biao Fu, Liang Zhang, Peigen Ye, Pei Yu, Cong Hu, Xiaodong Shi, and Yidong Chen. 2025a. Improving end-to-end sign language translation via multi-level contrastive learning. IEEE Transactions on Audio, Speech and Language Processing, 33:1230–1242. Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, Yunsheng Wu, Rongrong Ji, Caifeng Shan, and Ran He. 2025b. Mme: A compre- hensive evaluation benchmark for multimodal large language models. Preprint, arXiv:2306.13394. Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, Peixian Chen, Yanwei Li, Shaohui Lin, Sirui Zhao, Ke Li, Tong Xu, Xiawu Zheng, Enhong Chen, Caifeng Shan, and 2 others. 2025c. Video-mme: The first-ever compre- hensive evaluation benchmark of multi-modal llms in video analysis. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 24108–24118. Honghao Fu, Liang Zhang, Biao Fu, Rui Zhao, Jin- song Su, Xiaodong Shi, and Yidong Chen. 2024. Signer diversity-driven data augmentation for signer- independent sign language translation. In Findings of the Association for Computational Linguistics: NAACL 2024, pages 2182–2193, Mexico City, Mex- ico. Association for Computational Linguistics. Team GLM-V., Wenyi Hong, Wenmeng Yu, Xiaotao Gu, Guo Wang, Guobing Gan, Haomiao Tang, Jiale Cheng, Ji Qi, Junhui Ji, Lihang Pan, Shuaiqi Duan, Weihan Wang, Yan Wang, Yean Cheng, Zehai He, Zhe Su, Zhen Yang, Ziyang Pan, and 69 others. 2025. Glm-4.5v and glm-4.1v-thinking: Towards versatile multimodal reasoning with scalable reinforcement learning. Preprint, arXiv:2507.01006. Jia Gong, Lin Geng Foo, Yixuan He, Hossein Rahmani, and Jun Liu. 2024. Llms are good sign language translators. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, pages 18362–18372. Jianyuan Guo, Peike Li, and Trevor Cohn. 2025. Bridg- ing sign and spoken languages: Pseudo gloss gen- eration for sign language translation. In The Thirty- Ninth Annual Conference on Neural Information Pro- cessing Systems. Jie Huang, Wengang Zhou, Qilin Zhang, Houqiang Li, and Weiping Li. 2018. Video-based sign language recognition without temporal segmentation. In Pro- ceedings of the Thirty-Second AAAI Conference on Artificial Intelligence and Thirtieth Innovative Ap- plications of Artificial Intelligence Conference and Eighth AAAI Symposium on Educational Advances in Artificial Intelligence, AAAI’18/IAAI’18/EAAI’18, pages 2257–2264, New Orleans, Louisiana, USA. AAAI Press. Eui Jun Hwang, Sukmin Cho, Junmyeong Lee, and Jong C. Park. 2025. An efficient gloss-free sign lan- guage translation using spatial configurations and mo- tion dynamics with llms. In Proceedings of the 2025 Conference of the Nations of the Americas Chap- ter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Pa- pers), pages 3901–3920, Albuquerque, New Mexico. Association for Computational Linguistics. Youngjoon Jang, Haran Raajesh, Liliane Momeni, Gül Varol, and Andrew Zisserman. 2025. Lost in transla- tion, found in context: Sign language translation with contextual cues. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog- nition, pages 8742–8752. Peng Jin, Hongkai Li, Jun Yang, Yazhou Ren, Yuhao Li, Lilan Zhou, Jin Liu, Mei Zhang, Xiaorong Pu, and Siyuan Jing. 2025. A large dataset covering the chinese national sign language for dual-view iso- lated sign language recognition. Scientific Data, 12(1):660. Hamid Reza Vaezi Joze and Oscar Koller. 2019. Ms- asl: A large-scale data set and benchmark for understanding american sign language. Preprint, arXiv:1812.01053. Jungeun Kim, Hyeongwoo Jeon, Jongseong Bae, and Ha Young Kim. 2025. Leveraging the power of mllms for gloss-free sign language translation. In Proceedings of the IEEE/CVF International Confer- ence on Computer Vision, pages 21048–21058. Oscar Koller, Jens Forster, and Hermann Ney. 2015. Continuous sign language recognition: Towards large vocabulary statistical recognition systems handling multiple signers. Computer Vision and Image Under- standing, 141:108–125. Bohao Li, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang, Ruimao Zhang, and Ying Shan. 2024. Seed- bench: Benchmarking multimodal large language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13299–13308. Dongxu Li, Cristian Rodriguez Opazo, Xin Yu, and Hongdong Li. 2020. Word-level deep sign language recognition from video: A new large-scale dataset and methods comparison. In 2020 IEEE Winter Con- ference on Applications of Computer Vision, pages 1448–1458. Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. 2024. Llava- next: Improved reasoning, ocr, and world knowledge. Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual instruction tuning. In Proceed- ings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, pages 34892–34916, Red Hook, NY, USA. Curran Asso- ciates Inc. Yuqi Liu, Wenqian Zhang, Sihan Ren, Chengyu Huang, Jingyi Yu, and Lan Xu. 2025. Scope: Sign language contextual processing with embedding from llms. Proceedings of the AAAI Conference on Artificial Intelligence, 39(6):5739–5747. Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Khan. 2024. Video-chatgpt: Towards detailed video understanding via large vision and language models. In Proceedings of the 62nd An- nual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12585– 12602, Bangkok, Thailand. Association for Compu- tational Linguistics. Ministry of Education of the People’s Republic of China, State Language Commission, and China Disabled Persons’ Federation. 2018a. Chinese manual alpha- bet. Huaxia Publishing House, Beijing, China. Ministry of Education of the People’s Republic of China, State Language Commission, and China Disabled Persons’ Federation. 2018b. Lexicon of Common Expressions in Chinese National Sign Language. Huaxia Publishing House, Beijing, China. OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Ale- man, Diogo Almeida, Janko Altenschmidt, Sam Alt- man, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haim- ing Bao, Mohammad Bavarian, Jeff Belgum, and 262 others. 2024. Gpt-4 technical report. Preprint, arXiv:2303.08774. Zhi Rao, Yucheng Zhou, Benjia Zhou, Yiqing Huang, Sergio Escalera, and Jun Wan. 2025. Rvlf: A reinforcing vision-language framework for gloss-free sign language translation. Preprint, arXiv:2512.07273. Haozhan Shen, Peng Liu, Jingcheng Li, Chunxin Fang, Yibo Ma, Jiajia Liao, Qiaoli Shen, Zilun Zhang, Kangjia Zhao, Qianqian Zhang, Ruochen Xu, and Tiancheng Zhao. 2025. Vlm-r1: A stable and general- izable r1-style large vision-language model. Preprint, arXiv:2504.07615. Bowen Shi, Diane Brentari, Greg Shakhnarovich, and Karen Livescu. 2021. Fingerspelling detection in american sign language.In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pages 4166–4175. Garrett Tanzer and Biao Zhang. 2024. Youtube-sl-25: A large-scale, open-domain multilingual sign language parallel corpus. In The Thirteenth International Con- ference on Learning Representations. Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. Llama: Open and efficient foundation language models. Preprint, arXiv:2302.13971. Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhi- hao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. 2024. Qwen2-vl: Enhancing vision-language model’s per- ception of the world at any resolution. Preprint, arXiv:2409.12191. Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu, Long Cui, Xingguang Wei, Zhaoyang Liu, Linglin Jing, Shenglong Ye, Jie Shao, Zhaokai Wang, Zhe Chen, Hongjie Zhang, Ganlin Yang, Haomin Wang, Qi Wei, Jinhui Yin, Wenhao Li, Erfei Cui, and 56 others. 2025. Internvl3.5: Advancing open-source multimodal models in versatility, reasoning, and effi- ciency. Preprint, arXiv:2508.18265. Ryan Wong, Necati Cihan Camgoz, and Richard Bow- den. 2023. Sign2gpt: Leveraging large language models for gloss-free sign language translation. In The Twelfth International Conference on Learning Representations. James C. Woodward. 1972. Implications for sociolin- guistic research among the deaf. Sign Language Studies, 1(1):1–7. Kayo Yin, Amit Moryossef, Julie Hochgesang, Yoav Goldberg, and Malihe Alikhani. 2021. Including signed languages in natural language processing. In Proceedings of the 59th Annual Meeting of the Asso- ciation for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 7347– 7360, Online. Association for Computational Lin- guistics. Xiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, and 3 others. 2024. Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9556–9567. Ruiquan Zhang, Rui Zhao, Zhicong Wu, Liang Zhang, Haoqi Zhang, and Yidong Chen. 2025. Dynamic feature fusion for sign language translation using hy- pernetworks. In Findings of the Association for Com- putational Linguistics: NAACL 2025, pages 6227– 6239, Albuquerque, New Mexico. Association for Computational Linguistics. Zhuosheng Zhang, Aston Zhang, Mu Li, Hai Zhao, George Karypis, and Alex Smola. 2024. Multi- modal chain-of-thought reasoning in language mod- els. Transactions on Machine Learning Research. Rui Zhao, Liang Zhang, Biao Fu, Cong Hu, Jinsong Su, and Yidong Chen. 2024. Conditional variational autoencoder for sign language translation with cross- modal alignment. In Proceedings of the AAAI Con- ference on Artificial Intelligence, volume 38, pages 19643–19651. Hao Zhou, Wengang Zhou, Weizhen Qi, Junfu Pu, and Houqiang Li. 2021. Improving sign language transla- tion with monolingual data by sign back-translation. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1316–1325. Junjie Zhou, Yan Shu, Bo Zhao, Boya Wu, Zhengyang Liang, Shitao Xiao, Minghao Qin, Xi Yang, Yong- ping Xiong, Bo Zhang, Tiejun Huang, and Zheng Liu. 2025. Mlvu: Benchmarking multi-task long video understanding. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13691–13701. A CNSL-bench A.1 Dataset Alignment This section details the data processing and align- ment procedures underlying the construction of CNSL-bench. The textual descriptions and illustra- tive images are sourced from the National Common Sign Language Dictionary, which contains 8,214 commonly used sign entries and serves as an author- itative reference for standardized CSL (Ministry of Education of the People’s Republic of China et al., 2018b; China Disabled Persons’ Federation et al., 2019). During data preparation, we identify sev- eral forms of redundancy in the original dictionary. First, some entries share identical meanings and identical sign realizations, which are merged into a single sign entry. Second, certain lexical items correspond to multiple meanings, which are dis- tinguished in the dictionary using auxiliary index markers. Third, some entries share the same lexical form and meaning but differ in their sign realiza- tions. To ensure a consistent representation, we remove these auxiliary markers and retain all valid lexical variants associated with each sign realiza- tion, followed by sign-level processing. As a result, a set of 6,707 unique sign entries is obtained. In addition, the dictionary includes explanatory and instructional content intended for human readers, and we remove such descriptive text and preserve only the core lexical information. Moreover, the dictionary does not explicitly annotate linguisti- cally manual articulatory forms such as air-writing, finger-spelling, or manual-alphabet, we manually identify and annotate these categories to support fine-grained analysis. The sign language videos are sourced from the CNSL-DP dataset (Jin et al., 2025), which was collected under institutional ethical approval and provides synchronized video recordings for indi- vidual sign entries. For each sign, multiple video instances from different signers are available. To ensure both consistency and representativeness, we select one representative recording per sign entry for inclusion in the benchmark. The original videos are recorded at a resolution of 1920×1080 and 50 frames per second, with the signer centered in the frame. We uniformly downsample the videos to 24 frames per second and apply center cropping followed by resizing to 512×512 to standard- ize visual inputs. In cases where multiple synony- mous lexical entries correspond to the same sign realization, the original CNSL-DP dataset retains Figure 8: All manual alphabets in the Chinese Manual Alphabet, including 26 single-letter alphabets, 4 double-letter alphabets, and 2 alphabets with symbols. only a single representative form. To construct a lexically complete benchmark aligned with the official dictionary, we explicitly recover the omit- ted synonymous entries and re-associate them with the corresponding video instances during dataset alignment, thereby restoring the full set of lexical variants for each sign. As a result, CNSL-bench establishes a unified mapping among textual de- scriptions, illustrative images, and sign language videos, thereby constructing a coherent and well- aligned benchmark for evaluating sign language understanding. A.2 The China Manual Alphabet Figure 8 shows all manual alphabets in the Chinese Manual Alphabet, including 26 single-letter man- ual alphabets, 4 double-letter manual alphabets, and 2 manual alphabets with symbols. A.3 Detailed Task Definition This section provides additional description sup- porting the task formulation and option construc- tion strategies adopted in CNSL-bench. We first ex- amine the feasibility of open-ended sign language understanding and find that current state-of-the-art MLLMs remain highly unreliable in this setting. As illustrated in Figure 9, we draw a case from CSL-Daily (Zhou et al., 2021), a Chinese sign lan- guage dataset targeted at sign language recognition and translation. The results show that even for a short video expressing a simple and common sentence (e.g., “The weather is very nice today, neither cold nor hot”), flagship models such as Gemini-3-Pro and GPT-5.1 fail to recover the in- tended meaning, instead misinterpreting isolated handshapes or hallucinating unrelated lexical con- cepts. The observed failures suggest that existing models struggle to robustly integrate temporal dy- namics, sequential articulation, and lexical compo- sition in free-form generation, leading to unstable and difficult-to-interpret outputs. DistractorTextImageVideo ′ 2 Video ′ 10 Random72.4138.8332.0633.95 Semantic-Based64.3336.9532.6032.47 Table 3: An analysis of distractor on Qwen-VL-Plus. In addition, we further analyze different option construction strategies for the multiple-choice for- mulation. As shown in Table 3, semantics-based distractors yield slightly lower absolute accuracy than random sampling across modalities on a rep- resentative model (i.e., Qwen-VL-Plus), yet they lead to qualitatively consistent conclusions about model performance. This indicates that the over- all ranking and modality-dependent trends are sta- ble under alternative distractor designs, and that benchmark outcomes are not driven by overly ob- vious negatives. In other words, random sampling already provides sufficiently challenging and re- liable evaluation signals for CNSL-bench. Since semantics-based distractors introduce additional en- gineering overhead and potential sensitivity to simi- larity heuristics without changing the main findings, we adopt random option sampling to streamline the 今天气很好,不冷不热。(The weather is very nice today. It’s neither cold nor hot.) Gemini-3-Pro GPT-5.1 根据视频中的动作,这个中国手语( CSL )手势表达的意思是:忘记 (Forget) 以下是该手势的分解动作含义:指头 (Memory/Brain): 视频开始时,演示者用食指着自己的太阳穴 / 头部。这在手语中通常代表“大脑”、“思维”或 “记忆”。挥手 (Gone/Clear): 紧接着,她的手从头部向外挥动(有时候是一个类似“抹去”或“扔掉”的动作)。这象征着把脑子里的东西“丢掉”或 “抹去”了。总结:这一连串动作形象地表达了“脑子里的记忆被抹去了”,即“忘记”的意思。 Accordingtotheactionsinthevideo,thisChineseSignLanguage(CSL)gesturemeans:Forget Thefollowingisthebreakdownofthegesture'smovementsandtheirmeanings:Finger(Memory/Brain):Atthebeginningofthevideo,thedemonstratorpointsto hertempleorheadwithherindexfinger.Thisisusuallyusedinsignlanguagetorepresent"brain","thought",or"memory".Waving(Gone/Clear):Then,herhand wavesoutwardfromherhead(sometimesit'samotionsimilarto"erasing"or"throwingaway").Thissymbolizes"throwingaway"or"erasing"thethingsinthe mind.Summary:Thisseriesofmovementsvividlyexpresses"thememoryinthemindhasbeenerased",whichmeans"forget". 根据视频中提取的若干帧可以看到:仅右手参与。手形为食指单指上伸,其余手指弯曲。手臂从下方向上抬起,但没有明显的左右、前后移动轨迹。手指 方向是竖直向上。在中国手语中,这种单手食指向上竖起、指尖朝上的手形,最常见的含义是:表示数字“ 1” 。这与中国手语中最典型且标准的“ 1” 的表 达一致:右手食指伸出,其余手指屈曲,无特定空间移动。因此,视频中表达的含义是:数字: 1 。如果你希望,我可以进一步对比其与“上、第一、你” 等手势的区别,便于你核对。 Basedonseveralframesextractedfromthevideo,itcanbeseenthatonlytherighthandisinvolved.Thehandshapeiswiththeindexfingerextendedupwardsand theotherfingersbent.Thearmisliftedfromthebottomupwards,butthereisnoobviousleft-rightorfront-backmovementtrajectory.Thedirectionofthefingers isstraightupwards.InChineseSignLanguage,thishandshapeofasinglehandwiththeindexfingerextendedupwardsandthefingertipfacing upwardsmostcommonlymeansthenumber"1".Thisisconsistentwiththemosttypicalandstandardexpressionof"1"inChineseSignLanguage:theright indexfingerisextended,andtheotherfingersarebent,withoutspecificspatialmovement.Therefore,themeaningexpressedinthevideois:Number:1.Ifyouwish, Icanfurthercompareitwiththegesturesfor"up,first,you",etc.,tofacilitateyourverification. 今天(today) 天气(weather) 好(nice) 热(hot) 不(no) 不(no)冷(cold) Figure 9: An example of open-ended sign language understanding from the CSL-Daily dataset (Zhou et al., 2021). The input video expresses a simple sentence (“The weather is very nice today, neither cold nor hot”), yet flagship models (e.g., Gemini-3-Pro and GPT-5.1) fail to generate the correct meaning. Note that the realization of concepts such as “weather” in this example differs from the canonical forms used in CNSL-bench, reflecting natural variation in sign expression, while remaining readily interpretable to human signers. benchmark design and support robust, reproducible evaluation. A.4 Detailed Dataset Statistics The detailed statistics of CNSL-bench are illus- trated in Table 4, and the frame count is calculated at a rate of 24 frames per second. Subset#Sign Entry#Frames Air-Writing40799.8 Finger-Spelling77109.2 Manual-Alphabet59298.0 w/ 1 gesture2,97789.2 w/ 2 gestures3,287100.2 w/ 3 gestures369120.9 w/ 4 gestures62136.4 w/ 5 gestures8148.4 w/ 6 gestures2161.0 w/ 7 gestures2201.5 All670796.89 Table 4: The detailed statistics of CNSL-bench. B Experiments B.1 Experimental Settings MLLMs Participants. A total of 21 MLLMs (13 Open-source and 8 closed-source MLLMs) are included for validation, which includes a) 3 open&closed-source Image MLLMs: LLaVA- NeXT (Mistral-7B) (Liu et al., 2024), Qwen-VL- Plus/Max (Bai et al., 2023), b) 12 open-source MLLMs: Qwen2-VL-2B/7B (Wang et al., 2024), Qwen2.5-VL-3B/7B (Bai et al., 2025b), Intern3.5- VL-2B/8B (Wang et al., 2025), Qwen3-VL-2B/8B- Instruct, Qwen3-VL-2B/8B-Thinking (Bai et al., 2025a), LLaVA-NeXT-Video-7B (Liu et al., 2024), and GLM-4.1V-9B-Thinking (GLM-V. et al., 2025), and c) 6 close-source MLLMs: Qwen3- VL-Plus (Bai et al., 2025a), Gemini-2.5-Flash, Gemini-2.5-Pro (Comanici et al., 2025), GPT-4o- mini, GPT-4o, and GPT-5 (OpenAI et al., 2024). Evaluation Details. For open source MLLMs, we primarily use the Hugging Face transformers library 3 for model inference. To accelerate decod- ing for thinking models under the slow-thinking setting, we adopt vLLM 4 as the inference backend, and set the maximum generation length to 8,192 tokens per response. For closed-source MLLMs, 3 https://hugging-face.cn/docs/transformers 4 https://github.com/vllm-project/vllm any samples that fail due to API errors, timeouts, or malformed generations are excluded from scoring, so that the reported metric scores are computed only over valid outputs. As for video inputs, we evaluate two frame sampling rates: 2 fps (the de- fault in many video-understanding benchmarks) and a denser 10 fps setting, which better captures the high-speed and fine-grained spatiotemporal mo- tions characteristic of sign language. To mitigate input length constraints, when a dense sampling rate (e.g., FPS=10) would exceed a model’s maximum context window, we adaptively resample the video and include as many frames as possible without surpassing the input length limit (e.g., for LLaVA-NeXT-Video, we reduce the sam- pling rate but maximize the number of frames al- lowed within its context budget). For the GPT- series models, we uniformly cap the visual input to at most 50 frames to comply with the API re- striction. Unless stated otherwise, all hyperparame- ters follow the official recommendations for each model, as summarized in Table 6. Accuracy is computed by an exact match between the model prediction and the ground-truth answer. For think- ing models, we manually extract the final answer from the generated response to avoid conflating intermediate reasoning with the predicted label. Human Assessment. Deaf community involve- ment is essential for developing sign language un- derstanding systems (Yin et al., 2021; Atwell et al., 2024). To establish a human reference for CNSL- bench, we invited a professional team consisting of one professor specializing in sign language linguis- tics and three sign-language students (including one hearing-impaired student). Each student has at least one year of classroom studying experience in sign language; their instructors include the invited professor and Deaf sign language teachers from a local special education institute. To make the evaluation feasible while preserving articulation diversity, we constructed a 1,500-entry subset from the 6,707 sign entries. Specifically, we retained all entries involving air-writing, finger-spelling, and manual-alphabet articulations, yielding 1,018 en- tries, and then randomly sampled an additional 482 entries from the remaining gesture-only en- tries. For each entry, we generated three multiple- choice questions corresponding to the aligned tex- tual description, illustrative image, and sign lan- guage video, resulting in 4,500 questions in total. Each evaluator completed all questions, and we report the average accuracy over the three student evaluators as the human performance. C Additional Analysis C.1 Detailed Reasoning Tokens Table 5 reports the average reasoning tokens gen- erated by each model, stratified by correctness and modality. Across all evaluated MLLMs, incorrect predictions are consistently associated with sub- stantially longer reasoning traces than correct ones, reinforcing the observation that models tend to “think longer” when facing harder or ambiguous inputs, partially mirroring human problem-solving behavior. This effect is particularly pronounced for stronger models. For instance, GPT-5 (M) exhibits a ratio of 2.89 between incorrect and correct cases under text input, indicating that failed attempts of- ten trigger nearly three times as many reasoning tokens. Similar trends are observed in Gemini-2.5- Flash, whose ratios exceed 2.0 in the text modality. This phenomenon generalizes to image and video inputs, albeit with a noticeably attenuated magnitude. In image settings, the ratios between incorrect and correct reasoning length typically fall within 1.2–1.7, while with video input, the ratios further decrease to around 1.0–1.3. This compres- sion suggests that the tendency to engage in longer reasoning on more difficult cases, which is clearly observed in the text-only setting, becomes less pro- nounced once multimodal perception is introduced. Rather than reflecting increased confidence, the re- duced gap more plausibly indicates that multimodal perception and alignment imperfections constrain the model’s ability to adaptively allocate reason- ing effort: when visual evidence is noisy, under- specified, or imperfectly aligned with the language space, the model may fail to trigger longer, ex- ploratory reasoning even on genuinely difficult in- stances. A further supporting signal is that increas- ing the temporal resolution of video inputs does not yield a systematic restoration of the gap. The ratios observed at 2 FPS and 10 FPS remain highly simi- lar across models, including GPT-5 and Gemini-2.5 variants, despite the substantially increased number of frames. This suggests that the attenuation is not primarily driven by insufficient temporal evidence, but rather by broader limitations in multimodal un- derstanding, such as imperfect robustness in visual feature extraction, temporal integration, or cross- modal grounding, which prevent additional frames from translating into more accurately calibrated Model TextImageVideo ′ 2 f ps Video ′ 10 f ps CorrectFaultAllRatioCorrectFaultAllRatioCorrectFaultAllRatioCorrectFaultAllRatio Qwen3-VL-8B1,6882,9052,0461.722553032851.193904144071.062522752671.09 GLM-4.1V-9B1793042181.702783793391.363354003821.193183973741.25 Qwen3-VL-Plus1,2122,1241,4291.753604314011.201,2641,4111,3591.123223173190.98 Gemini-2.5-Flash8181,7551,0062.151,0901,8271,4471.677199388451.307209348431.30 Gemini-2.5-Pro (L)3473493471.001321351331.027673740.958284831.02 Gemini-2.5-Pro (M)1,1331,7551,2281.559411,2811,0731.367007887461.137088037571.14 Gemini-2.5-Pro (H)3655323901.469151,1771,0151.296757577181.126997977491.14 GPT-5 (L)2486372912.573394943911.453253953591.223534423941.25 GPT-5 (M)6351,8377592.891,0171,6141,2141.591,0931,4201,2451.301,3121,7011,4801.30 GPT-5 (H)1,3923,7371,6282.682,2483,2112,5531.432,3052,7772,5271.202,5973,1502,8521.21 Table 5: Detailed reasoning tokens across models and modalities. L, M, H: low, medium, and high reasoning effort on the process of thinking before generating an answer. Model TextV-L top_ptop_kTtop_ptop_kT LLaVA-NeXT-7B0.95501.00.95501.0 LLaVA-NeXT-Video-7B0.95501.00.95501.0 Qwen2/2.5-VL-Instruct0.95501.00.95501.0 Intern-VL-3.50.95501.00.95501.0 GLM-4.1V-9B0.95501.00.95501.0 Qwen3-VL-Instruct1.00401.00.80200.7 Qwen3-VL0.95201.00.95201.0 Gemini-2.5-Series0.95-1.00.95-1.0 GPT-Series1.0-1.01.0-1.0 Table 6: Detailed hyperparameters. T means tempera- ture. V-L denotes multimodal input settings. reasoning effort. Conclusively, the results point to a modality- dependent divergence in test-time behavior specific to sign language understanding. Although MLLMs display difficulty-sensitive processing in text-based settings, this characteristic is notably attenuated for visual inputs. Such attenuation suggests that current models may struggle to effectively utilize visual linguistic cues, indicating that limitations in multimodal perception and alignment contribute to the reduced adaptability observed in sign language understanding tasks. C.2 Detailed Prompt Tokens As shown in Table 8, the consumption of prompt tokens varies substantially across both input modal- ities and models. For a fixed model, image and video inputs consistently generate orders of magni- tude more prompt tokens than text, resulting in an extreme length gap in multimodal processing. This disparity in multimodal processing may partially account for the observed performance discrepan- cies across modalities. Beyond modality effects, the token usage also differs markedly under iden- tical text inputs, which can be largely attributed to heterogeneous tokenization and visual encoding strategies (e.g., LLaVa-NeXT vs. Qwen). Such tokenizer-induced prompt length variations may further affect effective context allocation and rea- soning budget, introducing an additional source of performance variability in cross-model compar- isons. Finally, closed-source models exhibit dis- tinct budget characteristics shaped by their pricing- oriented design choices. Although GPT-4o-mini of- fers a lower per-token cost, its substantially higher token consumption for multimodal inputs results in significantly increased overall usage, leading us to exclude it from further evaluation due to prohibitive cost considerations. C.3 Instruction Following As shown in Table 7, instruction adherence varies substantially across models and input modalities. While most large-scale open-source and proprietary MLLMs achieve near-perfect instruction-following accuracy across settings, several smaller-capacity models and certain MLLM families exhibit pro- nounced failures. Specifically, InternVL-3.5-2B maintains high accuracy on text and image inputs but collapses to around 15% accuracy on sign lan- guage videos, indicating severe difficulty in jointly satisfying visual, temporal, and task-level con- straints. In contrast, an opposite pattern is observed in Qwen2-VL-2B and models from the LLaVA- Next family, where instruction-following perfor- mance is already unstable across different modali- ties. Beyond model scale, we further observe that explicit CoT mechanisms may interact negatively with instruction adherence, likely due to verbosity and drifting constraints. For example, compared to Qwen3-VL-8B-Instruct, Qwen3-VL-8B-Thinking (marked with✍in Table 7) exhibits a slight but consistent degradation in instruction-following ac- curacy. A similar trend is observed in Gemini-2.5- Flash, whose instruction-following performance Model TextImageVideo ′ 2 f ps Video ′ 10 f ps AWFSMAAllAWFSMAAllAWFSMAAllAWFSMAAll Open& Close -source Image MLLMs 0% 50% 100% LLaVA-NeXT-7B1.26.52.54.961.254.662.362.868.871.473.771.2---- Qwen-VL-Plus100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.099.999.8100.0100.099.9 Qwen-VL-Max100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0 Open-Source MLLMs Qwen2-VL-2B85.583.185.083.498.597.499.098.3100.0100.0100.099.997.190.998.397.7 Qwen2.5-VL-3B100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0 Intern-VL-3.5-2B99.8100.099.799.499.397.497.897.515.215.616.415.317.414.316.415.4 Qwen3-VL-2B99.8100.0100.099.9100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0 LLaVA-NeXT-Video-7B3.23.93.74.845.244.247.149.863.658.463.764.060.058.463.362.6 Qwen2-VL-7B100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0 Qwen2.5-VL-7B100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0 GLM-4.1V-9B✍99.5100.099.799.898.898.799.399.199.5100.099.299.599.398.799.299.4 Intern-VL-3.5-8B100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0 Qwen3-VL-8B-Instruct100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0 Qwen3-VL-8B100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0 Qwen3-VL-8B✍99.096.194.999.1100.0100.0100.0100.0100.0100.099.8100.0100.0100.0100.0100.0 Closed-Source MLLMs Qwen3-VL-Plus✍100.0100.0100.099.9100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0 Gemini-2.5-Flash100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0 Gemini-2.5-Flash✍99.8100.099.399.892.992.292.996.398.396.197.598.497.598.799.098.4 Gemini-2.5-Pro✍100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0 GPT-4o-mini100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0---- GPT-4o99.898.799.899.898.8100.099.298.999.5100.099.399.299.098.798.398.6 GPT-5✍100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0 Table 7: Instruction Following.✍ denotes inference with slow thinking. ModelTextImageVideo ′ 2 Video ′ 10 Open-Source MLLMs LLaVA-NeXT-7B2722,4538,913- LLaVA-NeXT-Video✂3072,4801,3063,698 Qwen2/2.5-VL1697421,4136,669 Intern-VL-3.51954152,10210,659 Qwen3-VL15859011635,445 GLM-4.1V-9B✍1537261,4096,713 Closed-Source MLLMs Qwen-VL-Plus2181,5193,55215,586 Qwen-VL-Max2181,6213,67316,892 Qwen3-VL-Plus2061,3503,37520,842 Qwen3-VL-Plus✍2101,3963,22814,334 Gemini-2.5-Flash2254,3883,1243,343 Gemini-2.5-Flash✍2062,0253,0433,022 Gemini-2.5-Pro✍1982,5762,5012,451 GPT-4o-mini27071,274230,971- GPT-4o2641,9935,5078,383 GPT-5✍2021,1613,2198,382 Table 8: Average prompt token consumption across different input modalities.✍denotes inference with slow thinking.✂indicates that videos are sampled at a maximum of 6 FPS due to the context window limitation. decreases when slow thinking is enabled. These findings suggest that instruction-following robust- ness in sign language understanding is influenced by modalities, architectural and training choices, and may be further affected by the introduction of explicit reasoning. C.4 Case Studies The qualitative cases in Figure 10 to Figure 14 pro- vide intuitive insights into the sources of observed performance differences and further illustrate the challenges revealed by the CNSL-benchmark. A primary observation concerns modality sen- sitivity. For the lexical concept laptop (from Fig- ure 10 to Figure 12), models consistently succeed under textual descriptions but frequently fail un- der sign language video inputs, despite correctly recognizing the underlying typing motion. In con- trast, image-based inputs sometimes enable correct predictions by exposing complementary cues (e.g., typing combined with an “open” gesture). This contrast indicates that sign language understand- ing is strongly affected not only by the presence of visual information, but also by how spatial and temporal cues are presented across modalities. Be- yond modality effects, the cases suggest that sign language comprehension often requires integrat- ing perceptual recognition with implicit semantic associations. As shown in Figure 13, Gemini-2.5- Flash may correctly identify individual sign com- ponents (e.g., a nose-related gesture corresponding to “smell”) while failing to connect them to the intended meaning in context. Such failures reflect the fact that sign language is not a simple compo- sition of isolated gestures, but relies on culturally and linguistically grounded conventions that are not always explicitly observable. At the same time, 笔记本(Laptop) (一)双手五指弯曲,指尖朝下,交替点动几下,如敲击计算机键盘状。 (二)双手横伸,掌心相贴,然后右手做向上打开的动作。 (1)Bendthefingersofbothhands,withthefingertipspointingdown,and alternatelytapafewtimesasifstrikingthekeysofacomputerkeyboard. (2)Stretchbothhandsouthorizontally,withthepalmstouchingeachother, andthenmakeanupwardopeningmotionwiththerighthand. Qwen3- VL-2B Qwen3- VL-8B GLM-4.1V Gemini- 2.5-Flash Gemini- 2.5-Pro A. Laptop A.笔记本电脑 B. Warrior B. 勇士 C. Because/Due to C. 因为(由于) D. Transmit D.传讯 ......Let'sstartwithpartone:( 一 ) 双手五指弯曲,指尖朝下,交替点动几下,如敲击计算机键盘状。 "Thisistheclassic"typing"gesture,clearasday.The fingersarecurved,pointingdown,andtappinglikethey'rehittingkeys.Prettymuchscreaming"keyboard."Next,parttwo:( 二 ) 双手横伸,掌心相贴,然后右 手做向上打开的动作。 "Okay,we'vegotthehandsextended,palmstogether,andthenarighthandopeningupwards.That’saclassic"opening" motion.......Thiscouldbeopeningabook,alaptop,anythingthatopens.So,combiningthosetwoparts...TypingandOpening...Hmm...We'redefinitely lookingatsomethingthatinvolvesakeyboardandhasanopeningmechanism......Now,let'slookattheoptions.A. 笔记本电脑 (bǐjìběndiànǎo-laptop computer):Bingo!Thisfitsperfectly.Typing(keyboard)+opening(screen).Thefirstpartisthetyping,andthesecondpartisthelaptopscreenopening.This isincrediblystraightforward,frommyknowledgeandexperience.B.......C.......D.......Therefore,thebestanswerisA.Itleveragesmyunderstandingof CSLstructuresandthemostcommonconceptualinterpretationsused. ......第一部分:......这听起来像在模仿敲击键盘的动作。在手语中,敲击键盘的表示可能与“电脑”或“输入”相关。第二部分:......这可能表示 “打开”或“传递”,但需要看上下文。在中文手语中,常见的手势:“传讯”通常表示“传递信息”或“传递”,可能涉及手势如双手交叉或特定动 作。“因为”或“由于”可能有特定的手势,但我不确定。“勇士”可能不是标准手语词,但可能表示“英雄”或类似的东西。“笔记本电脑”是具体 的设备,但手语可能用动作表示。我应该考虑手语中常见的词汇。......第二部分:双手横伸,掌心相贴,然后右手向上打开。这听起来像“打开”或 “传递”。在中文手语中,有一个手势表示“传讯”或“传递”。我搜索记忆:......(太长省略)...... A“ 笔记本电脑”可能,但第一部分是敲击键盘,可 能表示“电脑”,但“传讯”更通用。或许问题是要表达的意思是“传讯”,因为第二部分匹配。但第一部分也相关。让我们看问题 :“ 这个手语动作表达 的是什么意思?”动作是两个部分。第一部分:敲击键盘动作,可能表示“电脑”或“输入”。第二部分:双手动作,可能表示“传递”。但“传讯”是 “传递消息”,所以可能 D 。 C“ 因为”可能用其他手势。我认为 D 是正确答案。...... A“ 笔记本电脑”可能,但第一部分是敲击键盘,可能表示“电脑”, 但“传讯”更具体。在上下文中,选项 D“ 传讯”可能就是答案。所以,我选择 D 。 ...... 分析动作:(一)双手五指弯曲,指尖朝下,交替点动几下,如敲击计算机键盘状。这听起来像在模拟敲击键盘,可能表示“计算机”或“键盘”。 (二)双手横伸,掌心相贴,然后右手做向上打开的动作。双手横伸掌心相贴可能表示“屏幕”或“显示器”,然后右手向上打开可能表示“打开”或 “展开”。结合两个部分,整体可能表示“笔记本电脑”,因为笔记本电脑有键盘和屏幕。动作(一)模拟敲击键盘,对应笔记本电脑的键盘。动作 (二)双手横伸掌心相贴可能表示屏幕,然后右手向上打开表示打开屏幕。在手语中,笔记本电脑通常用手势表示键盘和屏幕。现在,检查其他选项: B.......C.......D. 传讯:意思是传递信息或消息,可能涉及手势表示发送或接收消息。但动作(一)是敲键盘,动作(二)是打开屏幕,这更直接地指向笔 记本电脑。在标准中国手语中,笔记本电脑的表示通常是:用手势模拟键盘敲击,然后模拟打开屏幕。是的,我认为 A. 笔记本电脑是正确的。所以, ...... 最合理的应该是 A. 笔记本电脑。要求是只输出选项,不要输出内容。因此,我应该只输出 "A" 。在输出中,它说 " 只输出 A 、 B 、 C 、 D 的选项 " ,所以输出 "A" 。 确保没有多余输出。最后,答案应该只是 A 。 首先分析每个动作:第一部分“双手五指弯曲,指尖朝下,交替点动几下,如敲击计算机键盘状”—这是模拟敲键盘,和电脑设备相关;第二部分 “双手横伸,掌心相贴,然后右手做向上打开的动作”—结合这两个动作,对应选项里的A是笔记本电脑?不对,再想,或者看选项D?不对,再仔细 想。......?或者再查常见手语。......?不,C的常见手语是双手比“因为”的符号?不对,传讯(D)也不是。哦,可能我之前错了,再想。第一动作是 敲键盘(输入),第二动作可能和设备操作,那A是笔记本电脑?或者再确认。另一种思路:第一个动作是模拟敲击键盘(输入操作),第二个动作可能 是打开设备?那组合起来是笔记本电脑?或者可能正确答案是A?不对,再查资料。哦,可能正确选项是A?不对,再想。或者看选项D?......?不对,重 新分析:第一个动作是敲键盘(计算机键盘操作),第二个动作是打开设备,所以组合起来是笔记本电脑?那选A?不对,可能我错了。或者看选项 D?......可能正确答案是A?不对,可能我之前理解错了。或者正确选项是A?最终判断选A?不对,再想。哦,可能正确答案是A。所以输出A? ......Thefirstpartdescribesa"typing"action.Thatjumpsoutasveryiconic.Thesecondpartdescribessomethingopening.This,combinedwiththe "typing,"shouldimmediatelynarrowdownthepossibilities.Now,Ilookattheoptions."Laptopcomputer,""warrior,""because,"and"tosendamessage."Let's eliminatetheobvious."Warrior"......"Because"......And"Tosendamessage"......NowIfocusontheconnectionbetweenthedescribedactionsand "laptopcomputer."The"typing"isspot-on.That'sadirectreferencetousingacomputer.The"opening"gesture...thatsealsthedeal.Itclearly mimicsopeningaclamshell-styledevice.Thehandstogetheratthestartrepresenttheclosedlaptop,andtheopeningisthescreen.Thisisadirect representationofalaptopinCSL.Let'squicklycheckthe"sendmessage"option.There'sanassociationbutusing"sendmessage"isn'tadirectconnection liketheotherone.It'smoreabouttheactionandnottheobject.TheCSLsignisoftenexpressedthroughhandmovementsfromthemouthorother gesturesrepresentingsending.Basedonthis,optionAistheonlylogicallycorrectanswer.Itmapsperfectlytothetwopartsofthegesturesequence:typing (computer)andopening(laptop).So,IneedtomakesureIformatmyanswercorrectly.Theinstructionssayonlytheletter,noexplanations.It'sjust"A." Figure 10: Text-based sign language understanding for the lexical concept “laptop”. Under textual descriptions, models successfully associate the described typing action with the intended concept, demonstrating robust perfor- mance when sign language information is abstracted into natural language. we observe emerging sign-specific capabilities in stronger models. In Figure 14, Gemini-2.5-Pro correctly maps a visual sign to the Chinese charac- ter “neng” (meaning “can” or “ability”), indicating partial acquisition of sign knowledge. Conclusively, these examples emphasize that the difficulties faced by current MLLMs stem primarily from the intrinsic linguistic and multimodal com- plexity of sign language, rather than from isolated perceptual errors alone. 笔记本(Laptop) (一)双手五指弯曲,指尖朝下,交替点动几下,如敲击计算机键盘状。 (二)双手横伸,掌心相贴,然后右手做向上打开的动作。 (1)Bendthefingersofbothhands,withthefingertipspointingdown,and alternatelytapafewtimesasifstrikingthekeysofacomputerkeyboard. (2)Stretchbothhandsouthorizontally,withthepalmstouchingeachother, andthenmakeanupwardopeningmotionwiththerighthand. 首先看选项,C是“因为(由于)”。手语中“因为”通常用手势表示原因,比如一只手表示原因,另一只手表示结果,或者有特定的手势动作。图中可 能有表示原因的手势动作,比如(一)可能是手部动作,(二)是另一部分,结合常见手语,“因为”是常用的手语表达。而其他选项:A笔记本电脑、B 勇士、D传讯,这些对应的手语手势应该不太符合图中的动作。比如“传讯”可能有传递信息的手势,但图中动作更像表示原因。所以正确选项是C。 首先得分析图示内容。图示有两部分, ...... 。选项 D 是“传讯”,传讯在手语中通常有特定手势。比如,传讯可能 ...... 。 C 是“因为”,“因为”在手语 中 ...... 。选项 A 是“笔记本电脑”,但图示里没有出现电脑,所以排除。选项 B 是“勇,勇士的手语手势可能有特定的姿势 ...... 。选项 D 是“传讯”, ...... 。 再仔细看图示, ( 一 ) 是两组手部动作, ( 二 ) 是另一组。可能 ( 二 ) 的动作是手指向某个方向,比如传递信息。再检查其他选项, A 是笔记本电脑,但图示没 有出现电脑,排除。 B 是勇士,勇士的字可能需要特定手势,比如“勇”字的手语,但图示中可能不是。所以综合考虑,正确选项应该是 D 。 首先得回忆中国手语中各手势的含义。先看选项 C “因为(由于)”,需要判断手势是否匹配。不过可能需要更准确的分析,不过根据常见手语,这里可 能对应“因为”?不对,再想。或者看图示动作,不过可能我之前记忆有误?或者再仔细想,可能正确选项是 C ?不对,再检查。哦,可能正确答案是 C ? 不,再想。或者看选项中的“因为(由于)”的手势动作?不对,可能我之前错了。重新想:中国手语中,“因为”的手势动作?或者看选项 C 是“因为 (由于)”,而图示如果是这个手势,那选 C ?或者可能我之前记错了?现在要确定唯一正确选项。假设正确选项是 C ,所以输出 C ?不对,再想。或者看 其他选项, A 是笔记本电脑,手势不太像; B 勇士,手势也不像; D 传讯,也不对。那剩下的 C 可能正确。所以最终选 C ? ......Part(一)showstwohands,fingersslightlycurled,likethey'regrabbingsomethingormaybevibratingoverawavyline.Thewavylineisa commonindicatorofrepetitiveactionorsomethinghappening.......There'sanarrowindicatingthisdescendingmotion.Laptopsignsusuallyhave clearhandshapesresemblinganopendevice........"Because/DueTo"isthestrongestcontender.Part(二),withthedominanthanddescending...... Puttingitalltogether,thetwopartsalignperfectlywiththestructureof"because"(因为)inCSL:1.Part(一)setsthecontext,presentingthesituation/reason. 2.Part(二)showsthedominanthandsignifyingthereasonplacingintothepalmoftheotherhandrepresentingtheresult.Consideringthis,optionC,"因为 (because/dueto)",looksliketheclearwinner. ......First,Iseetwodistinctmovements.Movement( 一 )involvesatypingmotion,palmsdown,fingersbent,handsmovingupanddownslightly.That's clearlyrepresentingtyping,almostuniversally.Next,Movement( 二 )depictsopeningandclosing,ahandmovingdowntocoveranother,shapedto mimicalaptopclosing.It’sveryclearlythelaptopitself.Combiningthose,ithastobea"laptop".IknowhowCSL,andindeedmanysignlanguages, oftenuseiconicgestures......D. 传讯 (chuánxùn)-Tosendamessage/dispatch:Couldinvolveinformation,butthespecificityofthegesturestoalaptopare tooprecise.Messageusesindexfingers.Basedonthat,I'mconfidentin"A".Therefore,thefinaloutputneedstobe"A". Qwen3- VL-2B Qwen3- VL-8B GLM-4.1V Gemini- 2.5-Flash Gemini- 2.5-Pro A. Laptop A.笔记本电脑 B. Warrior B. 勇士 C. Because/Due to C. 因为(由于) D. Transmit D.传讯 Figure 11: Image-based sign language understanding for “laptop”. With illustrative images, some models correctly infer the target meaning by jointly recognizing multiple visual cues (e.g., “type” and “open” gestures), highlighting the role of complementary spatial information in sign language interpretation. 笔记本(Laptop) (一)双手五指弯曲,指尖朝下,交替点动几下,如敲击计算机键盘状。 (二)双手横伸,掌心相贴,然后右手做向上打开的动作。 (1)Bendthefingersofbothhands,withthefingertipspointingdown,and alternatelytapafewtimesasifstrikingthekeysofacomputerkeyboard. (2)Stretchbothhandsouthorizontally,withthepalmstouchingeachother, andthenmakeanupwardopeningmotionwiththerighthand. ......选项中有A.笔记本电脑,B.勇士,C.因为(由于),D.传讯。......“因为”可能用手势表示为“手在胸前做某种动作,可能类似表示原因”的动作。 观察视频中的动作:从6秒开始,双手动作,可能是在做“因为”的手势。比如,手在胸前做交叉或指向等动作,这可能对应“因为”。再看选项,C是 “因为(由于)”。A笔记本电脑通常用手势表示电子设备,可能不匹配;B勇士可能有特定手势,比如手臂伸展;D传讯可能涉及传递信息的动作,但根 据视频中的动作描述(6-11秒的手势),更可能的是“因为”。需要确认标准手语中“因为”的手势。常见的“因为”手语是:一只手放在胸前,另一只 手做类似“原因”的手势,比如手掌朝上,手指弯曲,然后做某种表示因果的动作。视频中第6-10秒的动作可能对应这个。所以正确选项应该是C。 用户现在需要解决的问题是判断视频中手语动作对应的选项。首先得回忆常见手语动作的含义。选项 C“ 因为(由于)”的手语动作特征,或者看动作的对 应。但可能需要分析动作流程,不过通常这种选择题如果有正确答案的话,需要确定。假设正确选项是 C ?不对,可能需要再想。或者可能我之前有记忆, C 选项“因为(由于)”的手语动作表现。所以最终选 C ? Qwen3- VL-2B Qwen3- VL-8B GLM-4.1V Gemini- 2.5-Flash Gemini- 2.5-Pro 视频中,人物先是站立,然后开始做手语动作。......选项里有“因为(由于)”这样的词。......首先分析手语动作:在视频中,手语动作可能是“因为” 或者“由于”的手语。比如,手语中“因为”(由于)的手语是通过手指的动作来表达,比如用双手的特定手势来表示原因。现在看选项,C选项是“因 为(由于)”,所以需要确定是否符合。根据常见手语,比如“因为”或“由于”的手语动作,通常会有类似双手交叉或者特定的手势。根据视频中的动 作,人物先做出双手分开,然后交叉,这可能是“因为”的手语。所以,正确选项应该是C。 ......Thesignerstartswithrelaxedhandsattheirsides.Atthetwo-secondmark,thehandsmoveup,palmsfacingeachother.Then,atthreeseconds,thingsget interesting.Thelefthandformsaflatpalm,facingupwards,andtherighthandformsaflatpalmfacingdownwards.And,importantly,therighthand tapsthelefthandonce.Afterthat,thehandsreturntotherelaxedstartingposition.......Tappingoneflathandontopofanother,orplacingoneflat handonanother,usuallysignifies"because,""reason,"or"dueto."......A: 笔记本电脑 (Laptopcomputer)No,that'susuallyshownbyforminga "laptop"shapewiththehandsorusingakeyboardmotion.Doesn'tmatch.B: 勇士 (Warrior)Definitelynot.Thatwouldinvolvegesturesofstrength, fighting,orholdingaweapon.Doesn'tfit.C: 因为(由于) (Because/Dueto)Bingo!Thisgesturealignsperfectlywiththehandmovementinthevideo. Thefoundation/reasonconceptisrightontarget.D: 传讯 (Transmitamessage/Message)Doesn'tseemright.Thatwouldinvolvea"sending"gestureor somethingsimilartowriting.Therefore,myconclusionisthatthesigninthevideomostlikelyrepresents**C: 因为(由于) (Because/Dueto). ......First,Ineedtoreallylookatthevideoitself.Handsdowntostart,thentheycomeup.Ah,palmsdown,fingersextended,andtogether.Theright handgoesonthebackofthelefthand,formingasortofbase,right?It'sheldinfrontofthebody,abitlowerthanchestheight.Thisreallyscreams "foundation"tome.......A.LaptopComputer:Nope.Thesignfor"laptop"justdoesn'tinvolvethishandshapeormovement.It'smuchmoredynamic, mimickingopeningandclosingalaptop.That'sout.B.Warrior/Brave:Noway."Brave"usuallyhasafistgesture,nearthechestorshoulder,signifyingstrength. Totallydifferent.C.Because/DueTo:Waitasecond...That'sit!That'sexactlythesign.Thelefthandisthefoundation,andtherighthandisplacedonit, showingthebasisforsomething.Yes,that'sspoton.D.TransmitaMessage/Summon:Nope.Thesignfor"transmit"usuallyinvolvesamovementawayfrom thefaceorhead,showingthesendingofinformation.Thisisn’tthatatall.Soit'sgottabe"Because/DueTo",optionC. A. Laptop A.笔记本电脑 B. Warrior B. 勇士 C. Because/Due to C. 因为(由于) D. Transmit D.传讯 Figure 12: Video-based sign language understanding failure for “laptop”. Despite correctly recognizing the typing motion in a continuous sign language video, models fail to recover the intended meaning, indicating that temporal visual recognition alone is insufficient for reliable sign language understanding. A. Server A.服务器 B. Weather/Air B.天气(空气) C. Sect C. 宗派 D. Mango D.芒果 天气/空气(Function) (一)一手食指直立,在头一侧上方转动一圈。 (二)一手打手指字母“Q”的指式,指尖朝内,置于鼻孔处。 (1)Raisetheindexfingerofonehandandextenditvertically,thenperform acircularmotionaboveonesideofthehead. (2)Formthehandshapecorrespondingtothemanualalphabetletter"Q" withthefingertipsorientedinwardandpositionedatthelevelofthenostrils. 用户现在需要分析手语动作对应的选项。首先看视频里的动作:从画面看,手语动作可能是“天气(空气)”。因为手语中“天气”相关的动作可能涉及 手势指向天空或表示空气。选项B是天气(空气),所以选B。 用户现在需要解决手语动作对应的含义问题。首先得回忆手语中各选项对应的动作, B 选项“天气(空气)”相关的手语动作,可能需要分析动作手势, 但这里可能更直接的是识别常见手语。通常天气或空气类手语动作可能与手的位置、动作表达空气相关,而其他选项如服务器、宗派、芒果的手语动作不 太符合。所以判断为 B 。 Qwen3- VL-2B Qwen3- VL-8B GLM-4.1V Gemini- 2.5-Flash Gemini- 2.5-Pro 用户现在需要分析手语动作表达的意思。首先看视频中的动作:第2到7秒之间,手语动作包括抬手食指向上(可能表示“点”或“开始”),然后食 指靠近嘴边(可能表示“小心”或“注意”),接着挥手动作。这些动作组合起来,手语中常见的是表达“天气(空气)”吗?不,再想,可能对应的 是“服务器”(A)?但更常见的是“天气(空气)”可能对应B?不对,再看手语的常见含义。不过用户需要明确,根据中国手语中“天气(空气)” 这个说法是否对应。可能手语动作中,先指向上(比如点),然后靠近嘴边(可能表示“注意”),挥手可能表示“空气”相关,但不太对。再想,可能 “天气”是B选项,但需要确认。不,正确的分析应该是:手语动作中,指向上(可能表示“点”),然后食指靠近嘴边(可能表示“注意”),挥手可 能表示“空气”?不对。可能正确的选项是B.天气(空气)?不,再仔细看选项,B是天气(空气),可能对应手语动作中表示天气或空气的含义。或者 可能,手语动作是“天气”对应的,比如用手指动作表示天气。但可能正确的是B。不过,我需要确认。其实,常见手语中,“天气”可能对应的是某个 动作,比如“点”(手指向上)表示“天气”,然后其他动作。但具体来说,可能视频中的动作是“天气(空气)”,即B选项。 AnalysisoftheCSLSign:Okay,sotheuserwantsmetofigureoutwhatthissignmeans.Let'sbreakitdown:righthandup,fistclosed,thendowntothe nose,withtheindexfingerandthumbpinching.Classic"smell"or"nose"inChineseSignLanguage(CSL).Easyenough.Now,theoptions.Server?Nope. That'sa"serve"and"machine"kindofdeal.Sect?"Religion"and"group,"definitelynot.Mango?They'reshowingmecuttingandpeelingimaginaryfruit, totallyoff.ButWeather/Air?Hmm..."Air"usuallyhasanopenhand,likeabreeze,ormovingaroundtheface.Thepinching'snotquiterightfor"air" onitsown.However...whatifit'saboutthesmelloftheair?Pollution,maybe?That'sapossibility,butlet'ssee.LookingatsomesignsforWeatherand AirinCSL,Iseethisgestureisn'tadirectsignforthose.Okay,let'skeepdigging.Searchingforsignsforeachconceptshowsthatit'sjustthesignfornose,or smell,ortheactofsmelling.IfIambeingspecifictothemovementtheoptionsaren'tright,yetitdoesappearthatthere'saconnection.Mangoisdefinitelyout. ServerandSectarenogo.Nowfor"Weather/Air".ThisistheonlyoneIcaneventryandconnectthesignto.IntermsofCSL,isthereanyassociationbetween itandthenoseortheactofsmelling?"Smellingtheair,"couldbelinked.Giventhatthisisamultiplechoicequiz,andthesignmeanssmellandthebest optiontopickwouldbe"Weather/Air",let'sreviewthesignagain.Thissignmeansnoseorsmell,andintheCSL.Sooutoftheoptionsgiven,picking "Weather/Air"istheclosestandonlyreasonableinterpretationinthegivencontext.Thesignisdefinitively"smell,"andthatstronglysuggests" 空气 (air)"whenthinkingaboutairqualityorthesensoryexperienceoftheair(doesitsmellbad?).Evenifit'snottheprimarysign,thisisthebestfit.I'm goingwith"B. 天气(空气) ". Alright,here'swhatI'mthinking.I'mbeingaskedtoidentifythesignlanguagedisplayedinthevideo.Timetobreakitdownsystematically,asIalwaysdo.First, Ineedtoreallyanalyzethevideo.Okay,sothewomanstartswithafistnearhershoulder,thumbup.Immediately,thatshouts"sky"or"heaven"( 天 )to me.Classic.Then,shebringsthatclosedhandtohernose.That'saclear"air"or"breath"( 气 )gesture,potentiallyeven"air"( 空气 )onitsown dependingonthecontext.Nowforthecrucialpart:puttingitalltogether.It'salmostcertainthatcombiningthosetwosignsgivesus"weather"( 天气 ).Next,I needtoevaluatetheoptionsprovided,startingwithA," 服务器 (fúwùqì)-Server."Nope,thesignsfor"serve"or"server"(asinthemachine)aretotally different;probablyinvolveflathandsormotionsrelatedtocomputing.Doesn'tfit.B," 天气(空气) -Weather(Air)"...Bingo!Thecombinationof"sky"and "air"directlytranslatesto"weather"( 天气 ),oreven"air"( 空气 )onitsown.It'saperfectmatch.C," 宗派 (zōngpài)-Sect/Denomination."That'sanabstract concept.I'dexpectsignsrelatedto"group"or"belief "forthis,nothingclosetowhat'sinthevideo.Wrong.Andfinally,D," 芒果 (mángguǒ)-Mango."Noway. Thesignformangowouldinvolvehandshapesdepictingthefruitortheactionofeatingit.Completelyoffbase.Therefore,thevideodemonstrablydisplays thesignfor"weather"or"air".Theanswer,withoutadoubt,isB. Figure 13: Partial semantic recognition without correct contextual integration. Models correctly identify individual sign components (e.g., a nose-related gesture), but fail to integrate them into the intended semantic concept, reflecting the implicit and convention-driven nature of sign language semantics. A. Blue A.蓝 B. The Salar nationality B.撒拉族 C. phrase C. 短语(短句) D. Function D.功能 功能(Function) (一)左手食、中指与右手食指先搭成“工”字形,然后右手食指在左手旁书空“力”字, 仿“功”字形。 (二)一手直立,掌心向外,然后食、中、无名、小指弯动一下。 (1)First,theindexandmiddlefingersofthelefthandandtheindexfingeroftherighthandform theshapeofthecharacter“gōng”(工).Then,theindexfingeroftherighthandwritesthecharacter “lì”(力)intheairbesidethelefthand,imitatingtheshapeofthecharacter"gōng"(功). (2)Holdonehanduprightwiththepalmfacingoutward,andthenbendtheindex,middle,ringand littlefingersonce. 用户现在需要解决手语图示的问题。图(一)是左手食指和右手食指的动作,可能表示某种手势;图(二)是手掌张开,手指弯曲,有箭头指示动作。然 后看选项:A蓝,B撒拉族,C短语(短句),D功能。首先,手语中“短语”或“短句”可能对应特定手势。比如,可能用手指的动作表示分句,或者重 复手势。另外,“功能”是更抽象的概念,可能手势不同。撒拉族是民族,蓝是颜色,通常颜色用手语中“蓝”对应的手势是蓝色的表示,但需要具体手 势。再想,图(一)可能是手势的组合,图(二)是手掌动作。可能“短语(短句)”的手语是表示短句的结构,比如用手指划动表示短句。或者查常见 手语词汇,“短语”可能对应某种手势。另外,选项C是“短语(短句)”,D是“功能”。需要回忆中国手语中“短语”或“短句”的表达。比如,短语 可能用手势表示分隔,短句可能重复或特定动作。而“功能”可能手势不同。可能图(一)是手势示意,图(二)是手掌动作,组合起来表示“短语”或 “短句”。例如,在手语中,“短语”可能用手指在空中划动表示,而“短句”可能类似。或者“短语”对应特定手势。再想,可能这个图示是“短语 (短句)”的手语表达,因为其他选项不太符合。比如A蓝:颜色蓝的手语通常是用蓝色的表示,比如“蓝”字的手势,而图中动作不是;B撒拉族:民族 的手语可能用特定手势,比如“撒拉族”可能用指代的动作,但图中不太像;D功能:功能的手语可能用“功能”这个词的表达,可能不同。所以综合判 断,选项C更可能正确。 用户现在需要解决手语图示对应的意思,选项是 A 到 D 。首先得回忆中国手语中各个手语的表示。看选项 C 是短语(短句),需要判断图示是否符合。不过 可能需要结合手语知识,通常这类手势如果是表示短语(短句)的手势?或者再仔细想,可能正确选项是 C ?不对,再检查。或者可能我之前记错了?不 对,重新分析。用户提供的图示,如果是表示“短语(短句)”的手语,那选 C 。但需要确定每个选项对应的手语。假设正确答案是 C ,那输出 C ?不对, 可能我之前有误?或者再想,可能正确选项是 C ?哦,不对,再仔细看问题,选项 C 是短语(短句),所以判断后选 C ?不对,可能我记错了,再想。或者 可能正确选项是 C ?最终确定选 C ?不对,再检查。哦,可能正确答案是 C ?所以最终输出 C ? Qwen3- VL-2B Qwen3- VL-8B GLM-4.1V Gemini- 2.5-Flash Gemini- 2.5-Pro 用户现在需要解决手语图示的意思。首先看选项,选项A是“蓝”,B是“撒拉族”,C是“短语(短句)”,D是“功能”。手语图示(一)和(二)的 描述需要理解。手语通常是表达特定含义,可能涉及字词或短语。但看选项中的C是“短语(短句)”,而手语里的“短语”可能是表达句子结构。不过 需要分析图示。通常,手语图示可能用于表示某种短句或词汇。不过需要看选项。可能这个手语图示是表示“短语”或者“短句”的意思,因为C选项是 “短语(短句)”。其他选项:A“蓝”是颜色,手语可能有颜色词,但图示看起来是手势动作。B“撒拉族”是民族,一般用拼音或文字,手语可能涉及民 族,但图示不是这个。D“功能”是功能,比如工具功能,但图示是手语手势。再仔细看,手语图示可能表达的是“短语”或者“短句”。因为选项C是 “短语(短句)”,而正确答案可能是C。因为短语(短句)在手语中常用手语表达。所以选C。 Alright,let'sbreakthisdown.ThetaskistodeciphertheseChineseSignLanguagegestures.I'mseeingtwodistinctsignsintheimage,labeled( 一 )and( 二 ). Analyzing( 一 ),thefirstgesture,itlookslikea"scissors"shapewiththeindexandmiddlefingersextendedononehand.Theotherhand'sindex fingerismovingtowardsthose"scissors."There'sevena" 力 "symbol,hintingatforceoremphasisduringthemovement.Myinitialthoughtisthat thiscouldrepresentakindof"cutting"orarticulation,mayberelatingtoseparatingsomething.Thiscouldpotentiallyrepresenttheconceptof"word"or "phrase"inCSL.Nowfor( 二 ).Thisoneshowsahandwithfingersextended,butthentheycurldownintoafist,oratleastarelaxedpositionwiththe fingertipspointingdown.Thedownwardcurlingindicatessomethingisshortening.Thisisaveryclearrepresentationof"short"inCSL.Now,puttingitall together.Consideringtheoptions,let'seliminatetheoutliersfirst." 蓝 "(Blue)isdefinitelyout,asitwouldlikelyinvolvegesturesrelatedtotheskyoreyes." 撒 拉族 "(Salarethnicgroup)isunlikely,asit'stoospecific.And" 功能 "(Function)hasnothingtodowithshorteningorwords.Thatleavesuswith" 短语 (短句) "-"Phrase(ShortSentence)".Thisiswhereitgetsinteresting." 短 "(duǎn,short)definitelymapstogesture( 二 )." 语 "(yǔ,language/speech/phrase) isthekey.Thinkingaboutit,themovementingesture( 一 )withthe"scissors"andtheindexfingercouldbeinterpretedas"articulating"or "separating"words.The" 力 "symbolwouldthenreinforcetheconceptofforceoremphasis.Itrepresentswordscomingoutinsomeway.I'm prettyconfidentthattheimagedepictsthesignsfor" 短 "and" 语 "-or" 句 ",inthecontextofsentences.Letmedouble-checkmyinternalknowledgeof CSL.RememberingCSLresources," 短 "(short)isoftenadownwardmovementofthefingers,whichmatches( 二 )perfectly." 语 "(word/phrase)ofteninvolves atappingormovementbetweenthefingers.Gesture( 一 )representsaformofarticulationbyusingtheindexfingeronthe"scissors"likethemouthorthe soundofwordscomingout.I'mprettyconfidentinthefinalanswer.Therefore,I'mgoingwithC," 短语(短句) ". Alright,here'swhatI'mthinking.I'mbeingaskedtodecipheraChineseSignLanguagegesture.Okay,let'sbreakitdownmethodically.First,Ineedtosee what'sgoingonintheimage.( 一 )showsahandusingthe"scissors"handshape,withanotherhandtracingthecharacter" 力 "(lì),meaning"power" or"function".Irecognizethat.Then,in( 二 ),thehandshapechangestothreefingersextended,andthentheybenddownwards.Thatsignalsthe character" 能 "(néng),meaning"can"or"ability".Now,timetocombinethose.“ 力” and“ 能” togetherimmediatelybringtomind" 功能 "(gōngnéng), whichtranslatesto"function"or"capability.“So,I'llgothroughtheoptions,justtobethorough.A. 蓝 (lán-blue):Nope,that'snotrelated.Thesignfor"blue" wouldlookentirelydifferent.B. 撒拉族 (Sālāzú-Salarpeople):No,theSalarsignwouldn’tusethosehandshapesormovements.C. 短语(短句) (duǎnyǔ (duǎnjù)-phrase/shortsentence):Nope.Thesignfor“phrase”or“sentence”isadifferentgesturealtogether.D. 功能 (gōngnéng-function):Bingo!That's thecombinationof" 力 "and" 能 ".Therefore,theanswerisD. Figure 14: Emerging sign-specific symbolic understanding in advanced MLLMs. A stronger model successfully maps a visual sign to the Chinese character “neng” (meaning “can” or “ability”), suggesting partial acquisition of sign language knowledge, while still falling short of comprehensive understanding.