Paper deep dive
Overview of the NLPCC 2026 Shared Task 1: Difficulty-Aware Multilingual and Multimodal Medical Instructional Video Understanding Evaluation
Shenxi Liu, Kan Li, Mingyang Zhao, Yuhang Tian, Bin Li
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 7/9/2026, 4:23:14 AM
Summary
The paper introduces DA-MIVQA, a difficulty-aware shared task for NLPCC 2026 that evaluates multilingual and multimodal medical instructional video question answering. It categorizes questions into simple (subtitle-based) and complex (visual/procedural) subsets across three tracks: temporal grounding in single videos, video corpus retrieval, and joint retrieval-grounding. The task provides a benchmark, evaluation metrics, baselines, and competition results to assess system robustness across varying evidence complexity.
Entities (17)
Relation Signals (18)
DA-MIVQA → partof → NLPCC 2026
confidence 98% · we introduce the Difficulty-Aware Medical Instructional Video Question Answering (DA-MIVQA) shared task for NLPCC 2026.
DA-MIVQA → containstrack → DA-TAGSV
confidence 97% · The challenge contains three tracks: Difficulty-Aware Temporal Answer Grounding in Single Video (DA-TAGSV)...
DA-MIVQA → containstrack → DA-VCR
confidence 97% · ...Difficulty-Aware Video Corpus Retrieval (DA-VCR)...
DA-MIVQA → containstrack → DA-TAGVC
confidence 97% · ...Difficulty-Aware Temporal Answer Grounding in Video Corpus (DA-TAGVC).
DA-MIVQA → categorizesquestionsas → Simple
confidence 95% · It divides questions into simple and complex subsets: simple questions are typically answerable from subtitle-aligned textual cues...
DA-MIVQA → categorizesquestionsas → Complex
confidence 95% · ...complex questions require visual grounding and procedural understanding.
Simple → requiresevidence → Subtitle-based textual cues
confidence 94% · simple questions can often be answered from subtitle-based textual cues
Complex → requiresevidence → Visual grounding and procedural understanding
confidence 94% · complex questions require visual grounding, procedural understanding, and cross-modal evidence integration.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Following the CMIVQA, MMI-VQA, and M4IVQA challenges in NLPCC 2023--2025, we introduce the Difficulty-Aware Medical Instructional Video Question Answering (DA-MIVQA) shared task for NLPCC 2026. DA-MIVQA extends previous multilingual and multimodal medical video benchmarks by explicitly distinguishing questions according to the type and complexity of evidence required for answering. Specifically, simple questions can often be answered from subtitle-based textual cues, whereas complex questions require visual grounding, procedural understanding, and cross-modal evidence integration. The challenge contains three tracks: Difficulty-Aware Temporal Answer Grounding in Single Video (DA-TAGSV), Difficulty-Aware Video Corpus Retrieval (DA-VCR), and Difficulty-Aware Temporal Answer Grounding in Video Corpus (DA-TAGVC). The dataset is collected from public medical instructional channels, covers diverse scenarios such as first aid, emergency response, rehabilitation, nursing, and general medical education, and is manually verified with difficulty annotations. This paper presents the task motivation, dataset construction, evaluation protocol, participation overview, competition results, and representative systems of DA-MIVQA. DA-MIVQA provides a practical benchmark for evaluating medical instructional video question answering systems under varying textual, visual, temporal, and procedural reasoning requirements.
Tags
Links
- Source: https://arxiv.org/abs/2607.06618v1
- Canonical: https://arxiv.org/abs/2607.06618v1
Trouble viewing inline? Open PDF directly →
Full Text
33,271 characters extracted from source content.
Expand or collapse full text
11institutetext: School of Computer Science and Technology, Beijing Institute of Technology 11email: liushenxi,likan,tianyuhang@bit.edu.cn 22institutetext: Department of Computing, The Hong Kong Polytechnic University 22email: 25019897r@connect.polyu.hk 33institutetext: Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences 33email: b.li2@siat.ac.cn Overview of the NLPCC 2026 Shared Task 1: Difficulty-Aware Multilingual and Multimodal Medical Instructional Video Understanding Evaluation Shenxi Liu Kan Li Mingyang Zhao Yuhang Tian Bin Li Corresponding author Abstract Following the CMIVQA, MMI-VQA, and M4IVQA challenges in NLPCC 2023–2025, we introduce the Difficulty-Aware Medical Instructional Video Question Answering (DA-MIVQA) shared task for NLPCC 2026. DA-MIVQA extends previous multilingual and multimodal medical video benchmarks by explicitly distinguishing questions according to the type and complexity of evidence required for answering. Specifically, simple questions can often be answered from subtitle-based textual cues, whereas complex questions require visual grounding, procedural understanding, and cross-modal evidence integration. The challenge contains three tracks: Difficulty-Aware Temporal Answer Grounding in Single Video (DA-TAGSV), Difficulty-Aware Video Corpus Retrieval (DA-VCR), and Difficulty-Aware Temporal Answer Grounding in Video Corpus (DA-TAGVC). The dataset is collected from public medical instructional channels, covers diverse scenarios such as first aid, emergency response, rehabilitation, nursing, and general medical education, and is manually verified with difficulty annotations. This paper presents the task motivation, dataset construction, evaluation protocol, participation overview, competition results, and representative systems of DA-MIVQA. DA-MIVQA provides a practical benchmark for evaluating medical instructional video question answering systems under varying textual, visual, temporal, and procedural reasoning requirements. 1 Introduction Recent advances in AI-assisted healthcare have shown promise in clinical and educational scenarios such as medical imaging, decision support, and procedural training [7, 24, 25]. Medical instructional videos are increasingly used for acquiring practical skills in first aid, emergency response, nursing, rehabilitation, and general medical education [27, 26, 9]. Compared with text alone, such videos demonstrate medical actions, tool usage, posture transitions, and temporal procedure flows, providing richer evidence for medical learning and question answering [16, 30, 17]. This value has motivated a series of NLPCC shared tasks on medical instructional video question answering, including CMIVQA 2023 [13], MMIVQA 2024 [15], and M4IVQA 2025 [12]. These tasks have expanded from Chinese medical video QA to multilingual, multimodal, and multi-hop reasoning settings [19], with two core directions: temporal answer grounding, which localizes answer-relevant spans in untrimmed videos, and video corpus retrieval, which identifies relevant videos from large collections [4, 33]. However, existing benchmarks do not explicitly distinguish questions by the evidence required for answering in realistic medical scenarios [23, 35]. While some questions can be answered from subtitles or basic semantic matching, others require visual grounding, action understanding, object-state recognition, and procedural context modeling. Thus, difficulty in medical instructional video understanding should not be defined only by reasoning depth or inference hops, but also by whether explicit visual evidence is required beyond textual cues. To address this gap, NLPCC 2026 introduces the Difficulty-Aware Medical Instructional Video Question Answering challenge, or DA-MIVQA111https://med-m3-dataset.github.io/for more informaton.. Unlike previous tasks emphasizing language scope, modality coverage, or multi-hop reasoning, DA-MIVQA evaluates systems according to the evidence structure required by each question. It divides questions into simple and complex subsets: simple questions are typically answerable from subtitle-aligned textual cues or explicit single-source evidence, whereas complex questions require visual grounding and procedural understanding. This design better reflects real-world medical use, where users need not only textual hints but also visual evidence about what to do, where to act, and when a procedure step occurs [2, 5, 18]. DA-MIVQA retains the three-track formulation of previous NLPCC shared tasks while applying difficulty-aware evaluation to each track. DA-TAGSV requires systems to localize the answer span in a single video; DA-VCR retrieves the most relevant video from a corpus; and DA-TAGVC first retrieves the relevant video and then grounds the answer temporally within it. Evaluating all tracks under simple, complex, and mixed settings enables fine-grained analysis of robustness across procedural and cross-modal difficulty levels. The DA-MIVQA dataset is collected from public medical instructional channels on YouTube and covers scenarios such as first aid, emergency management, rehabilitation, nursing, and general medical education [6]. Following previous shared tasks, questions and temporal answers are manually verified by annotators with medical backgrounds, and each question is labeled as simple or complex according to its required evidence. This annotation helps reveal whether systems understand visual and procedural content or mainly rely on subtitle matching. Table 1: A conceptual comparison of the NLPCC medical instructional video QA shared tasks. Task Languages Multi-modal Chain-of-thought Difficulty-aware CMIVQA (2023) [13] Chinese ✓ ✗ ✗ MMIVQA (2024) [15] Chinese/English ✓ ✗ ✗ M4IVQA (2025) [12] Chinese/English ✓ ✓ ✗ DA-MIVQA (2026) Chinese/English ✓ ✓ ✓ DA-MIVQA supports more reliable medical video QA for education, skill acquisition, emergency guidance, and multilingual knowledge access. It also provides a benchmark for distinguishing systems that mainly match subtitles from those that integrate visual, textual, and procedural evidence in medical scenarios. This paper presents an overview of DA-MIVQA, covering its motivation, task definition, dataset construction, evaluation protocol, participation summary, and representative systems. 2 Task Introduction The DA-MIVQA challenge aims to promote medical instructional video question answering under a difficulty-aware evaluation setting. Different from previous NLPCC shared tasks that mainly emphasized multilingual understanding, multimodal fusion, or multi-hop reasoning [13, 15, 12, 19], DA-MIVQA explicitly evaluates system performance across questions with different evidence requirements [22, 10]. All questions are categorized into simple and complex subsets. Simple questions are mainly answerable through subtitle-aligned textual matching or basic semantic understanding, whereas complex questions require visual grounding, procedural interpretation, and integration of textual and visual evidence from medical instructional videos. Following the design paradigm of previous NLPCC medical video shared tasks, DA-MIVQA contains three tracks: Difficulty-Aware Temporal Answer Grounding in Single Video, Difficulty-Aware Video Corpus Retrieval, and Difficulty-Aware Temporal Answer Grounding in Video Corpus. 2.1 Definition of Each Track The challenge contains three complementary tracks covering temporal localization, corpus retrieval, and their joint modeling in medical instructional video understanding. Figure 1: The multi-modal features of complex questions in DA-MIVQA. Track 1: Difficulty-Aware Temporal Answer Grounding in Single Video (DA-TAGSV). Given a medical or health-related question and a single untrimmed medical instructional video, DA-TAGSV requires systems to localize the start and end timestamps of the answer span. It evaluates fine-grained temporal localization ability under both subtitle-dominant simple questions and visually grounded complex questions. Track 2: Difficulty-Aware Video Corpus Retrieval (DA-VCR). Given a question and a large collection of untrimmed medical instructional videos, DA-VCR requires systems to retrieve the most relevant video from the corpus. It evaluates semantic matching between questions and candidate videos under multilingual and multimodal conditions, while further comparing retrieval robustness across simple and complex questions. Track 3: Difficulty-Aware Temporal Answer Grounding in Video Corpus (DA-TAGVC). Given a question and a video corpus, DA-TAGVC requires systems to first retrieve the relevant video and then localize the answer span within it. This track combines the challenges of DA-TAGSV and DA-VCR, making it more demanding because retrieval and localization errors may accumulate, especially for complex questions. 2.2 Evaluation Metrics Following previous NLPCC medical instructional video shared tasks, DA-MIVQA adopts task-specific metrics for the three tracks and reports results on simple-only, complex-only, and mixed subsets [13, 15, 12]. This protocol measures not only overall effectiveness but also robustness under different levels of evidence complexity. For DA-TAGSV, evaluation is based on temporal overlap between predicted and gold answer spans. We adopt Intersection over Union (IoU), mean Intersection over Union (mIoU), and R@1, IoU = μ, where μ∈0.3,0.5,0.7μ∈\0.3,0.5,0.7\. The main ranking metric is mIoU. For DA-VCR, evaluation is based on whether the relevant video is ranked highly in the returned list. We use recall-based metrics R@n with n∈1,10,100n∈\1,10,100\ and Mean Reciprocal Rank (MRR). The main ranking metric is the Overall score, computed as the average of R@1, R@10, R@100, and MRR. For DA-TAGVC, evaluation combines retrieval and localization. We report R@1|mIoU, R@10|mIoU, and R@100|mIoU, which reflect localization quality under different retrieval depths. The main ranking metric is the Average score, computed as the mean of these three values. 2.3 Dataset The DA-MIVQA dataset is collected from public medical instructional channels on YouTube and covers diverse medical and health-related scenarios, including first aid, medical emergency management, rehabilitation guidance, nursing practice, and general medical education [1]. Following previous shared tasks, questions and temporal answers are manually verified by annotators with medical backgrounds to ensure annotation quality and medical validity [8]. Each video may contain multiple question-answer pairs, and each question is associated with a unique temporal answer span marked by start and end timestamps. Compared with previous datasets, DA-MIVQA further introduces difficulty labels, namely simple and complex, according to the type of evidence required for answering each question. This re-annotation strategy distinguishes subtitle-dominant matching questions from genuinely multimodal procedural understanding questions [20]. The dataset is divided into training, validation, and test subsets, each containing Chinese and English questions with simple and complex labels. Overall, DA-MIVQA provides a more fine-grained benchmark for multilingual and multimodal medical instructional video understanding. Table 2: Statistics of sample counts for the three tracks. Track Language Complexity Training Set Validation Set Test Set Total Track 1 Chinese Simple 2101 362 351 2814 Complex 1842 285 339 2466 English Simple 1273 199 202 1674 Complex 1341 223 219 1783 Track 2 Chinese Simple 2101 362 394 2857 Complex 1842 285 280 2407 English Simple 1273 199 210 1682 Complex 1341 223 234 1798 Track 3 Chinese Simple 2101 362 394 2857 Complex 1842 285 280 2407 English Simple 1273 199 210 1682 Complex 1341 223 234 1798 3 Baseline To provide reference points for the proposed DA-MIVQA evaluation, we design two complementary baselines: a text-only oracle baseline and a multimodal encoder–decoder baseline. The former estimates how far a system can go using subtitle information alone, while the latter provides a practical reference for modeling visual evidence, temporal context, and cross-modal interactions. Together, these baselines allow us to examine the performance gap between subtitle-based reasoning and multimodal video understanding under different evidence requirements. 3.1 Text-only Oracle Baseline For the text-only oracle baseline, we use a frozen Qwen3.5-9B [28] model as the backbone large language model. Only timestamped SRT subtitles are provided as input, and the model is adapted to DA-MIVQA through task-specific prompts without updating the backbone parameters. Each instance consists of a question, the corresponding subtitle sequence, and the expected task-specific output. For DA-TAGSV, the model predicts the temporal span in the given video that best supports the answer. For DA-VCR, it identifies the most relevant video from the candidate corpus based on subtitle-level textual evidence. For DA-TAGVC, it first infers the relevant video and then predicts the answer-supporting span within that video. This baseline is not intended as a deployable multimodal solution, but as an estimate of the upper-bound performance achievable from subtitles alone. 3.2 Multimodal Encoder–Decoder Baseline The second baseline adopts a multimodal encoder–decoder framework that jointly models questions, timestamped subtitles, and visual clips from medical instructional videos. It is conceptually inspired by cross-modal mutual knowledge transfer for visual answer localization [31], but is adapted to the evidence-aware setting of DA-MIVQA rather than directly reproducing the original model. The framework first extracts subtitle, question, and visual clip representations. It then performs multimodal evidence mining to select question-relevant subtitle spans and video clips, followed by iterative cross-modal reasoning between textual and visual representations. The resulting multimodal representation is fed into task-specific decoders for the three tracks: timestamp prediction for DA-TAGSV, video relevance scoring for DA-VCR, and joint retrieval plus temporal localization for DA-TAGVC. The model is optimized with localization, retrieval, and cross-modal consistency objectives. Overall, this baseline provides a practical multimodal reference by combining subtitle cues, visual demonstrations, and temporal evidence within a unified architecture. Compared with the text-only oracle, it is expected to better handle complex questions that require visual grounding and procedural understanding. 4 Evaluation Results In this section, we summarize the participation status and evaluation results of the NLPCC 2026 shared task on Difficulty-Aware Medical Instructional Video Question Answering (DA-MIVQA). Following the reporting style of previous NLPCC shared task overview papers, we present the overall participation overview, the official rankings of each track, and brief comparative observations across different difficulty settings [13, 15, 12]. 4.1 Participation Overview The DA-MIVQA challenge attracted teams from universities, research institutes, and industry participants interested in multilingual and multimodal medical video understanding. For NLPCC 2025 Shared Task 4, a total of 22, 14, 16 teams registered for Track 1, 2, 3, respectively. During the test phase, 17, 10, and 11 teams submitted valid results, respectively. In the end, Amazon Inc., Team_WuKong, and BIGC achieved the SOTA for Tracks 1, 2, and 3. In general, the participation distribution across the three tracks can reflect the relative technical difficulty of each task setting, especially because DA-TAGVC requires systems to jointly solve retrieval and temporal grounding under difficulty-aware evaluation. 4.2 Results of Track 1: DA-TAGSV Track 1, Difficulty-Aware Temporal Answer Grounding in Single Video, evaluates whether a system can accurately localize the temporal answer span in an untrimmed medical instructional video for a given question. The official ranking of this track is based on the mIoU score, while R@1, IoU = 0.3, R@1, IoU = 0.5, and R@1, IoU = 0.7 are reported as complementary metrics for detailed analysis. Table 3 presents the final ranking results for Track 1. Table 3: Results of Track 1: Difficulty-Aware Temporal Answer Grounding in Single Video (DA-TAGSV). Rank Team ID R@1,IoU=0.3 R@1,IoU=0.5 R@1,IoU=0.7 mIoU 1 Amazon Inc. 0.5512 0.4007 0.2217 0.3912 2 Karamay 0.5334 0.4110 0.2269 0.3904 3 Ouc_AI [34] 0.5101 0.3579 0.2146 0.3608 4 Dvoe protiv vetra 0.4946 0.3688 0.2183 0.3606 5 SETAG[37] 0.4116 0.3363 0.1790 0.3089 6 HIIT 0.4191 0.3287 0.1760 0.3079 7 M-Baseline [31] 0.4265 0.3211 0.1729 0.3068 8 Text-Only Baseline [28] 0.4206 0.3127 0.1441 0.2925 9 BFSU 0.4147 0.3043 0.1153 0.2782 10 Random Pick Method [14] 0.0774 0.0818 0.0403 0.0665 Under the difficulty-aware setting, Track 1 is expected to reveal a clear performance gap between simple and complex questions, because the latter require more reliable visual grounding and procedural understanding beyond subtitle-based matching. Compared with simple questions, complex questions in this track typically place higher demands on identifying medically relevant actions, body positions, tool usage, and temporal transitions in the video. 4.3 Results of Track 2: DA-VCR Track 2, Difficulty-Aware Video Corpus Retrieval, evaluates whether a system can retrieve the most relevant medical instructional video from a large corpus for an input question [33]. The official ranking of this track is based on the Overall score, which summarizes the retrieval quality measured by R@1, R@10, R@100, and MRR. Table 4 shows the final ranking results for Track 2. Table 4: Results of Track 2: Difficulty-Aware Video Corpus Retrieval (DA-VCR). Rank Team ID R@1 R@10 R@100 MRR Overall 1 Team_Wukong 0.3751 0.4150 0.5189 0.3931 0.4255 2 Chaldeas 0.3559 0.4160 0.5492 0.3537 0.4187 3 DIMA[29] 0.3268 0.4355 0.5260 0.3325 0.4052 4 sun [32] 0.3086 0.3721 0.4442 0.3000 0.3562 5 HDU_Team2 0.3139 0.3647 0.4447 0.2982 0.3554 6 DSG-1 [11] 0.3192 0.3572 0.4452 0.2964 0.3545 7 M-Baseline [31] 0.1570 0.1953 0.2046 0.1546 0.1779 8 AC Automation 0.1518 0.1712 0.1704 0.1461 0.1599 9 Text-Only Baseline [28] 0.1465 0.1470 0.1361 0.1375 0.1418 10 Random Pick Method [14] 0.0361 0.0822 0.0962 0.0636 0.0695 Compared with Track 1, Track 2 places greater emphasis on identifying question-video relevance at the corpus level, and therefore tests the semantic matching ability of systems under multilingual and multimodal conditions. Under the simple setting, retrieval performance may benefit more from lexical overlap and subtitle-level semantic clues, whereas the complex setting is expected to require stronger modeling of visually grounded procedures and action semantics. 4.4 Results of Track 3: DA-TAGVC Track 3, Difficulty-Aware Temporal Answer Grounding in Video Corpus, is the most comprehensive setting in DA-MIVQA because it requires systems to retrieve the relevant video and localize the answer span within that video at the same time. The official ranking of this track is based on the Average score, which is computed from R@1|mIoU, R@10|mIoU, and R@100|mIoU. Table 5 provides the final ranking results for Track 3. Table 5: Results of Track 3: Difficulty-Aware Temporal Answer Grounding in Video Corpus (DA-TAGVC). Rank Team ID R@10,mIoU R@100,mIoU R@1,mIoU Average 1 BIGC 0.1646 0.2879 0.3660 0.2728 2 UWM 0.1415 0.2727 0.3338 0.2493 3 IIEleven[21] 0.1238 0.2513 0.3566 0.2439 4 UESC 0.1364 0.2551 0.3359 0.2425 5 MedEcho[36] 0.1490 0.2588 0.3152 0.2410 6 Nsddd[3] 0.1436 0.2129 0.3235 0.2267 7 DesiWen 0.1232 0.2351 0.3159 0.2248 8 M-Baseline [31] 0.1028 0.2572 0.3082 0.2228 9 Text-Only Baseline [28] 0.0885 0.2415 0.2853 0.2051 10 Random Pick Method [14] 0.0454 0.0872 0.0673 0.0666 Because retrieval and localization errors may accumulate in this track, its performance is expected to be lower than that of the first two tracks, especially on complex questions requiring both correct video selection and accurate visual-temporal grounding. Therefore, the results of this track are particularly useful for examining the end-to-end capability of systems in realistic medical instructional video question answering scenarios. 4.5 Comparative Observations Across all three tracks, the difficulty-aware setting makes it possible to compare system behavior on text-dominant and visually grounded questions in a more fine-grained manner. The expected comparison between Simple Only and Complex Only results can help reveal whether a model mainly relies on subtitle matching or can truly integrate textual, visual, and procedural evidence. In particular, if a system maintains relatively stable performance across the two difficulty subsets, it may indicate stronger robustness in handling medically grounded multimodal reasoning. By contrast, a large performance drop from simple to complex questions would suggest that current methods still have limitations in visual grounding, action understanding, and temporal procedural modeling for medical videos. 5 Conclusion In this paper, we presented an overview of the NLPCC 2026 shared task on Difficulty-Aware Medical Instructional Video Question Answering (DA-MIVQA), which extends the CMIVQA, MMIVQA, and M4IVQA research line toward a more fine-grained and practically grounded evaluation setting. Unlike previous benchmarks that mainly focused on language coverage, modality expansion, or multi-hop reasoning, DA-MIVQA distinguishes questions according to the type of evidence required for answering. Simple questions are mainly answerable from subtitle-aligned textual cues, whereas complex questions require visual grounding, procedural understanding, and cross-modal evidence integration. The challenge consists of three tracks: Difficulty-Aware Temporal Answer Grounding in Single Video, Difficulty-Aware Video Corpus Retrieval, and Difficulty-Aware Temporal Answer Grounding in Video Corpus. Together, they evaluate temporal localization, corpus retrieval, and end-to-end retrieval-grounding ability in multilingual and multimodal medical instructional video scenarios. Built from public medical instructional videos and enriched with difficulty annotations, DA-MIVQA provides a practical benchmark for distinguishing systems that mainly rely on subtitle matching from those that can integrate textual, visual, and procedural evidence. We hope this challenge will encourage future research on difficulty-aware medical video understanding and support more robust medical question answering systems for education, emergency guidance, rehabilitation training, and cross-lingual knowledge access. Acknowledgement This work was supported by National Natural Science Foundation of China (Nos. 4222037 and L181010), and Sanming Project of Medicine in Shenzhen (No. SZZYSM202311002). References [1] C. J. Brame (2016) Effective educational videos: principles and guidelines for maximizing student learning from video content. CBE Life Sciences Education 15. External Links: Link Cited by: §2.3. [2] A. Burgess, C. van Diggele, C. Roberts, and C. Mellis (2020-12-03) Tips for teaching procedural skills. BMC Medical Education 20 (2), p. 458. External Links: ISSN 1472-6920, Document, Link Cited by: §1. [3] S. Cheng, Z. Zhou, J. Liu, J. Ye, H. Luo, and Y. Gu (2023) A unified framework for optimizing video corpus retrieval and temporal answer grounding: fine-grained modality alignment and local-global optimization. In CCF International Conference on Natural Language Processing and Chinese Computing, p. 199–210. Cited by: Table 5. [4] S. Cheng, Z. Zhou, J. Liu, J. Ye, H. Luo, and Y. Gu (2023) A unified framework forăoptimizing video corpus retrieval andătemporal answer grounding: fine-grained modality alignment andălocal-global optimization. In Natural Language Processing and Chinese Computing, F. Liu, N. Duan, Q. Xu, and Y. Hong (Eds.), Cham, p. 199–210. External Links: ISBN 978-3-031-44699-3 Cited by: §1. [5] J. Colgan, S. Kourouche, G. Tofler, and T. Buckley (2023) Use of videos by health care professionals for procedure support in acute cardiac care: a scoping review. Heart, Lung and Circulation 32 (2), p. 143–155. External Links: ISSN 1443-9506, Document, Link Cited by: §1. [6] V. Curran, K. Simmons, L. Matthews, L. Fleet, D. L. Gustafson, N. A. Fairbridge, and X. Xu (2020-12-01) YouTube as an educational resource in medical education: a scoping review. Medical Science Educator 30 (4), p. 1775–1782. External Links: ISSN 2156-8650, Document, Link Cited by: §1. [7] D. Escobar-Castillejos, A. Y. Barrera-Animas, J. Noguez, A. J. Magana, and B. Benes (2025) Transforming surgical training with ai techniques for training, assessment, and evaluation: scoping review. Journal of Medical Internet Research 27. External Links: ISSN 1438-8871, Document, Link Cited by: §1. [8] D. Gupta, K. Attal, and D. Demner-Fushman (2023-03-22) A dataset for medical instructional video classification and question answering. Scientific Data 10 (1), p. 158. External Links: ISSN 2052-4463, Document, Link Cited by: §2.3. [9] I. R. Krumm, M. C. Miles, A. Clay, W. G. Carlos I, and R. Adamson (2022) Making effective educational videos for clinical teaching. Chest 161 (3), p. 764–772. External Links: ISSN 0012-3692, Document, Link Cited by: §1. [10] J. Lei, L. Yu, T. Berg, and M. Bansal (2020-07) TVQA+: spatio-temporal grounding for video question answering. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault (Eds.), Online, p. 8211–8225. External Links: Link, Document Cited by: §2. [11] N. Lei, J. Cai, Y. Qian, Z. Zheng, C. Han, Z. Liu, and Q. Huang (2023) A two-stage chinese medical video retrieval framework with llm. In CCF International Conference on Natural Language Processing and Chinese Computing, p. 211–220. Cited by: Table 4. [12] B. Li, S. Liu, Y. Weng, Y. Du, Y. Tian, and S. Zhou (2026) Overview of the nlpcc 2025 shared task 4: multi-modal, multilingual, and multi-hop medical instructional video question answering challenge. In Natural Language Processing and Chinese Computing, X. Mao, Z. Ren, and M. Yang (Eds.), Singapore, p. 367–379. External Links: ISBN 978-981-95-3352-7 Cited by: Table 1, §1, §2.2, §2, §4. [13] B. Li, Y. Weng, H. Guo, B. Sun, S. Li, Y. Luo, M. Qi, X. Liu, Y. Han, H. Liang, S. Gao, and C. Chen (2023) Overview of the nlpcc 2023 shared task: chinese medical instructional video question answering. In Natural Language Processing and Chinese Computing, F. Liu, N. Duan, Q. Xu, and Y. Hong (Eds.), Cham, p. 233–242. External Links: ISBN 978-3-031-44699-3 Cited by: Table 1, §1, §2.2, §2, §4. [14] B. Li, Y. Weng, H. Guo, B. Sun, S. Li, Y. Luo, M. Qi, X. Liu, Y. Han, H. Liang, et al. (2023) Overview of the nlpcc 2023 shared task: chinese medical instructional video question answering. In CCF International Conference on Natural Language Processing and Chinese Computing, p. 233–242. Cited by: Table 3, Table 4, Table 5. [15] B. Li, Y. Weng, Q. Song, L. Liang, X. Min, and S. Zhou (2025) Overview of the nlpcc 2024 shared task 7: multi-lingual medical instructional video question answering. In Natural Language Processing and Chinese Computing, D. F. Wong, Z. Wei, and M. Yang (Eds.), Singapore, p. 429–439. External Links: ISBN 978-981-97-9443-0 Cited by: Table 1, §1, §2.2, §2, §4. [16] B. Li, Y. Weng, B. Sun, and S. Li (2022) Learning to locate visual answer in video corpus using question. ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 1–5. External Links: Link Cited by: §1. [17] S. Li, B. Li, B. Sun, and Y. Weng (2024) Towards visual-prompt temporal answer grounding in instructional video. IEEE Transactions on Pattern Analysis and Machine Intelligence 46 (12), p. 8836–8853. External Links: Document Cited by: §1. [18] H. Liu, X. Ma, C. Zhong, Y. Zhang, and W. Lin (2024) TimeCraft: navigate weakly-supervised temporal grounded video question answering via bi-directional reasoning. In Computer Vision – ECCV 2024: 18th European Conference, Milan, Italy, September 29–October 4, 2024, Proceedings, Part V, Berlin, Heidelberg, p. 92–107. External Links: ISBN 978-3-031-72651-4, Link, Document Cited by: §1. [19] S. Liu, K. Li, M. Zhao, Y. Tian, B. Li, S. Zhou, H. Li, and F. Yang (2025) M3-med: a benchmark for multi-lingual, multi-modal, and multi-hop reasoning in medical instructional video understanding. External Links: 2507.04289, Link Cited by: §1, §2. [20] S. Liu, K. Li, M. Zhao, Y. Tian, S. Zhou, and B. Li (2025) Med-craft: automated construction of interpretable and multi-hop video workloads via knowledge graph traversal. External Links: 2512.01045, Link Cited by: §2.3. [21] T. Ma, Y. Hu, S. Jiang, Z. Yin, and T. Zang (2024) Multilingual temporal answer grounding in video corpus with enhanced visual-textual integration. In CCF International Conference on Natural Language Processing and Chinese Computing, p. 471–483. Cited by: Table 5. [22] J. Park, K. J. Jang, B. Alasaly, S. Mopidevi, A. Zolensky, E. Eaton, I. Lee, and K. Johnson (2024) Assessing modality bias in video question answering benchmarks with multimodal large language models. ArXiv abs/2408.12763. External Links: Link Cited by: §2. [23] J. Park, K. J. Jang, B. Alasaly, S. Mopidevi, A. Zolensky, E. Eaton, I. Lee, and K. Johnson (2025) Assessing modality bias in video question answering benchmarks with multimodal large language models. AAAI’25/IAAI’25/EAAI’25, AAAI Press. External Links: ISBN 978-1-57735-897-8, Link, Document Cited by: §1. [24] E. W. Riddle, D. Kewalramani, M. Narayan, and D. B. Jones (2024) Surgical simulation: virtual reality to artificial intelligence. Current Problems in Surgery 61 (11), p. 101625. External Links: ISSN 0011-3840, Document, Link Cited by: §1. [25] A. Shahrezaei, M. Sohani, S. Taherkhani, and S. Y. Zarghami (2024-11-13) The impact of surgical simulation and training technologies on general surgery education. BMC Medical Education 24 (1), p. 1297. External Links: ISSN 1472-6920, Document, Link Cited by: §1. [26] K. Srinivasa, A. Charlton, F. Moir, and F. Goodyear-Smith (2023) How to develop an online video for teaching health procedural skills: tutorial for health educators new to video production. JMIR Medical Education 10. External Links: Link Cited by: §1. [27] K. Srinivasa, F. Moir, and F. Goodyear-Smith (2022) The role of online videos in teaching procedural skills in postgraduate medical education: a scoping review. Journal of Surgical Education 79 (5), p. 1295–1307. External Links: ISSN 1931-7204, Document, Link Cited by: §1. [28] Q. Team (2026-02) Qwen3.5: accelerating productivity with native multimodal agents. External Links: Link Cited by: §3.1, Table 3, Table 4, Table 5. [29] Y. Wang, T. Tan, and Y. Wang (2026) Hierarchical indexing with knowledge enrichment for multilingual video corpus retrieval. In Natural Language Processing and Chinese Computing, X. Mao, Z. Ren, and M. Yang (Eds.), Singapore, p. 393–404. External Links: ISBN 978-981-95-3352-7 Cited by: Table 4. [30] Y. Weng and B. Li (2022) Visual answer localization with cross-modal mutual knowledge transfer. External Links: 2210.14823, Link Cited by: §1. [31] Y. Weng and B. Li (2023) Visual answer localization with cross-modal mutual knowledge transfer. In ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Vol. , p. 1–5. External Links: Document Cited by: §3.2, Table 3, Table 4, Table 5. [32] G. Yu, X. Bi, J. Tang, M. Gu, T. Chen, Z. Li, and M. Zhu (2024) MQuA: multi-level query-video augmentation for multilingual video corpus retrieval. In CCF International Conference on Natural Language Processing and Chinese Computing, p. 353–364. Cited by: Table 4. [33] H. Zhang, A. Sun, W. Jing, G. Nan, L. Zhen, J. T. Zhou, and R. S. M. Goh (2021-07) Video corpus moment retrieval with contrastive learning. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’21, p. 685–695. External Links: Link, Document Cited by: §1, §4.3. [34] H. Zhang, C. Zheng, Y. He, Y. Zhao, and Y. Lai (2024) Improving multilingual temporal answering grounding in single video via llm-based translation and ocr enhancement. In CCF International Conference on Natural Language Processing and Chinese Computing, p. 145–156. Cited by: Table 3. [35] Y. Zhong, J. Xiao, W. Ji, Y. Li, W. Deng, and T. Chua (2022) Video question answering: datasets, algorithms and challenges. External Links: 2203.01225, Link Cited by: §1. [36] Y. Zhou, J. Wu, and Y. Li (2026) Multi-hop knowledge-enhanced query reasoning for multi-modal medical video qa. In Natural Language Processing and Chinese Computing, X. Mao, Z. Ren, and M. Yang (Eds.), Singapore, p. 380–392. External Links: ISBN 978-981-95-3352-7 Cited by: Table 5. [37] Z. Zhou, J. Liu, S. Cheng, H. Luo, Y. Gu, and J. Ye (2023) Improving cross-modal visual answer localization in chinese medical instructional video using language prompts. In CCF International Conference on Natural Language Processing and Chinese Computing, p. 221–232. Cited by: Table 3.