Paper deep dive
MMCOMET: A Large-Scale Multimodal Commonsense Knowledge Graph for Contextual Reasoning
Eileen Wang, Hiba Arnaout, Dhita Pratama, Shuo Yang, Dangyang Liu, Jie Yang, Josiah Poon, Jeff Pan, Caren Han
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 89%
Last extracted: 7/20/2026, 2:57:50 AM
Summary
The paper introduces MMCOMET, a large-scale multimodal commonsense knowledge graph (MMKG) that extends the ATOMIC2020 text-based knowledge graph with visual grounding. It integrates physical, social, and eventive knowledge into nearly 1 million multimodal triples. The authors propose a hybrid visual alignment pipeline combining embedding-based similarity matching and concreteness-aware web retrieval to associate images with commonsense phrases. Experimental results demonstrate that MMCOMET improves performance in visual storytelling, visual question answering, and image captioning tasks compared to text-only or baseline multimodal approaches.
Entities (16)
Relation Signals (16)
MMCOMET → contains → 989K multimodal triples
confidence 95% · The resulting graph contains nearly 1M multimodal commonsense tuples spanning 19 relation types.
MMCOMET → extends → ATOMIC2020
confidence 95% · MMCOMET extends the ATOMIC2020 knowledge graph to include a visual dimension
MMCOMET → integrates → physical, social, and eventive knowledge
confidence 92% · MMCOMET, the first multimodal commonsense knowledge graph (MMKG) that integrates physical, social, and eventive knowledge.
MMCOMET → supports → Image Captioning
confidence 92% · We integrate MMCOMET into: ... (3) Image Captioning.
MMCOMET → supports → Visual Storytelling
confidence 92% · We integrate MMCOMET into: (1)Visual Storytelling
MMCOMET → supports → Visual Question Answering
confidence 92% · We integrate MMCOMET into: ... (2) Visual Question Answering
MMCOMET → integrates → social knowledge
confidence 90% · MMCOMET ... integrates physical, social, and eventive knowledge.
MMCOMET → integrates → eventive knowledge
confidence 90% · MMCOMET ... integrates physical, social, and eventive knowledge.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We present MMCOMET, the first multimodal commonsense knowledge graph (MMKG) that integrates physical, social, and eventive knowledge. MMCOMET extends the ATOMIC2020 knowledge graph to include a visual dimension, through an efficient image retrieval process, resulting in over 900K multimodal triples. This new resource addresses a major limitation of existing MMKGs in supporting complex reasoning tasks like image captioning and storytelling. Through a standard visual storytelling experiment, we show that our holistic approach enables the generation of richer, coherent, and contextually grounded stories than those produced using text-only knowledge. This resource establishes a new foundation for multimodal commonsense reasoning and narrative generation.
Tags
Links
- Source: https://arxiv.org/abs/2603.01055v1
- Canonical: https://arxiv.org/abs/2603.01055v1
Trouble viewing inline? Open PDF directly →
Full Text
45,265 characters extracted from source content.
Expand or collapse full text
MMCOMET: A Large-Scale Multimodal Commonsense Knowledge Graph for Contextual Reasoning Eileen Wang ∗ ewan9058@uni.sydney.edu.au University of Sydney Sydney, Australia Hiba Arnaout ∗ hiba.arnaout@tu-darmstadt.de Technische Universität Darmstadt Darmstadt, Germany Dhita Putri Pratama ∗ dhita.pratama@unimelb.edu.au University of Melbourne Melbourne, Australia Shuo Yang shuo.yang.3@unimelb.edu.au University of Melbourne Melbourne, Australia Dangyang Liu danyang.liu@ed.ac.uk University of Edinburgh Edinburgh, United Kingdom Jie Yang jyan4704@uni.sydney.edu.au The University of Sydney Sydney, Australia Josiah Poon † josiah.poon@sydney.edu.au The University of Sydney Sydney, Australia Jeff Pan † j.z.pan@ed.ac.uk University of Edinburgh Edinburgh, United Kingdom Soyeon Caren Han † caren.han@unimelb.edu.au University of Melbourne Melbourne, Australia Abstract We present MMCOMET, the first multimodal commonsense knowl- edge graph (MMKG) that integrates physical, social, and even- tive knowledge. MMCOMET extends the ATOMIC2020 knowledge graph to include a visual dimension, through an efficient image re- trieval process, resulting in over 900K multimodal triples. This new resource addresses a major limitation of existing MMKGs in support- ing complex reasoning tasks like image captioning and storytelling. Through a standard visual storytelling experiment, we show that our holistic approach enables the generation of richer, coherent, and contextually grounded stories than those produced using text- only knowledge. This resource establishes a new foundation for multimodal commonsense reasoning and narrative generation. Keywords Commonsense Knowledge Graph, Multimodal Language Model, Visual Storytelling, Visual Question Answer, Image Captioning ACM Reference Format: Eileen Wang, Hiba Arnaout, Dhita Putri Pratama, Shuo Yang, Dangyang Liu, Jie Yang, Josiah Poon, Jeff Pan, and Soyeon Caren Han. 2026. MMCOMET: A Large-Scale Multimodal Commonsense Knowledge Graph for Contextual Reasoning. In Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR’26). ACM, Mel- bourne, VIC, Australia, 8 pages. https://doi.org/X.X ∗ Co-first Authors. † Corresponding Authors Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org. SIGIR’26, Naarm, Melbourne © 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-1-4503-X-X/2026/06 https://doi.org/X.X Figure 1: An example of automated visual storytelling: Baseline: Family members enjoyed leisurely moments together. Grandpa shared memories during the trip; Ours: Family spent the day relaxing in the boat, enjoying beer. Grandpa took the grandson on his lap as he drove, with the grandson’s parents watching nearby. 1 Introduction Commonsense Knowledge Graphs (KGs) encode everyday knowl- edge in the form of relational triples< head, relation, tail>, and have become important external resources for information re- trieval and generation systems [11]. In multimodal settings, KGs are frequently queried to provide auxiliary knowledge for tasks such as image captioning(IC) [48], visual question answering (VQA) [9], and visual storytelling (VST) [39]. In such pipelines, the quality of the retrieved knowledge directly affects downstream reasoning and generation quality [27]. However, most existing KGs are text-only, limiting their ability to support visually grounded reasoning. Recent work on multimodal knowledge graphs (MMKGs) has incorporated images into encyclopedic KGs [1,8]. While these re- sources are valuable for factual and entity-centric queries, they largely focus on encyclopedic relations (e.g., capital-of, located-in) and lack coverage of everyday commonsense knowledge involving social interactions, event dynamics, and human intentions. On the other hand, large-scale commonsense KGs such as ATOMIC2020 [13] provide rich social and eventive reasoning signals, but remain purely textual and are not visually grounded. This gap is partic- ularly evident in visual-language tasks. As illustrated in Figure 1, a baseline storytelling model produces a generic description of a family outing. In contrast, a system augmented with multimodal arXiv:2603.01055v1 [cs.AI] 1 Mar 2026 SIGIR’26, July, Naarm, MelbourneEileen Wang, Hiba Arnaout, Dhita Putri Pratama, Shuo Yang, Dangyang Liu, Jie Yang, Josiah Poon, Jeff Pan, and Soyeon Caren Han commonsense can infer that people are relaxing in a boat, that a grandfather is driving with his grandson on his lap, and that parents are watching nearby. These inferences require integrating (i) physi- cal knowledge about objects and settings, (i) eventive knowledge about actions and temporal dynamics, and (i) social knowledge about family roles and interactions. Existing MMKGs do not jointly model these dimensions. To address this limitation, we introduce MMCOMET, the first large-scale multimodal commonsense knowl- edge graph that integrates physical, social, and eventive knowledge into a unified resource. MMCOMET extends ATOMIC2020 with a visual dimension by associating commonsense triples with rep- resentative images retrieved via a hybrid pipeline that combines similarity-based matching and web search. The resulting graph contains nearly 1M multimodal commonsense tuples spanning 19 relation types. Beyond constructing the resource, we evaluate its utility across multiple visual-language tasks to demonstrate its general applicability. We integrate MMCOMET into: (1)Visual Sto- rytelling, (2) Visual Question Answering, and (3) Image Cap- tioning. Across all three tasks, MMCOMET consistently improves visual grounding, contextual coherence, and answer accuracy, in- dicating that multimodal commonsense provides complementary signals beyond purely visual or textual features. Main contributions: •We present the first multimodal commonsense knowledge graph that covers physical, social, and event relations in a unified frame- work, comprising 989K multimodal triples. •We propose an efficient visual alignment pipeline that reduces similarity-matching computations by approximately 60×com- pared to brute-force retrieval. • We provide statistics and human evaluation analyses validating image-text alignment quality across relation types. •We demonstrate the effectiveness of MMCOMET across three representative visual-language tasks (VST, VQA, and IC), showing consistent performance gains. MMCOMET establishes a new foundation for multimodal common- sense retrieval and reasoning, and serves as a reusable resource for future multimodal IR and generation research. 2 Related Work We situate MMCOMET within prior work on commonsense knowl- edge and multimodal knowledge graphs, highlighting the gap be- tween textual commonsense and visually grounded knowledge. 2.1Text-only Commonsense Knowledge Graphs Commonsense knowledge graphs (CSKGs) encode everyday knowl- edge beyond factual encyclopedic information. Early resources such as ConceptNet [17,33,34] collect structured commonsense relations (e.g.,IsA,UsedFor,CapableOf) via crowdsourcing. Other projects including WebChild [35,36], Quasimodo [28,29], TupleKB [20], and Ascent [21] extract commonsense statements automatically from large-scale web corpora using pattern-based and open infor- mation extraction methods [23]. TransOMCS [45] further refines plausibility using statistical and neural scoring. ATOMIC [30] and ATOMIC2020 [13] extend commonsense reasoning to event-centric and social interactions. Using COMET [5], ATOMIC2020 scales inferential commonsense generation across social, physical, and MMKGDomain or Focus Size Multimodal Image Source IMGpedia [8]Ency442ME, RWMC ImageGraph [24]Ency560KEWSE MMKG [18]Ency814KEWSE Richpedia [40]Ency, Geo172ME, RWSE, WP VisualSem [1]Ency, Multi-L1.5MEWP, ImageNet MMEKG [19]CS (Ev)934ME, TimSitu TIVA-KG [41]CS (Phys)1.4ME, TWSE MMpedia [42]Ency19MEWSE, WP AspectMMKG [46]Ency645KEWSE, WP MMCOMET (ours)CS (Soc, Phys, Ev)989KE, RWSE, OID Table 1: Comparison of existing MMKGs with MMCOMET. Ours uniquely covers domains (social, physical, eventive) with nearly 1M tuples (Ency: Encyclopedic; Geo: Geographic; Multi-L: Multilingual; Soc: Social; Phys: Physical; Ev: Even- tive; E: Entity; R: Relation; T: Tuples; CS: Commonsense; WSE: Web Search Engine; WP: Wikipedia; WMC: Wikimedia Commons; OID: open-source image datasets.) eventive relations. These resources have proven effective for dia- logue generation and narrative reasoning [7,39]. However, they remain purely textual and lack direct visual grounding, limiting their applicability to multimodal retrieval and generation settings. Several works explore specialized commonsense dimensions, in- cluding hyper-relational CSKG integration [14], cultural common- sense [22], and negative commonsense statements [3,15]. Despite their diversity, these resources operate exclusively in text. 2.2 Multimodal Knowledge Graphs Multimodal knowledge graphs (MMKGs) augment textual entities with images or other modalities to support cross-modal reasoning and retrieval. Most MMKGs focus on encyclopedic knowledge. Encyclopedic MMKGs. Resources such as IMGpedia [8], Image- Graph [24], MMKG [18], Richpedia [40], VisualSem [1], MMpe- dia [42], and AspectMMKG [46] link structured entities from DB- pedia, Freebase, or Wikidata with images. These graphs primarily encode factual relations (e.g., geographic, biographical, organiza- tional) and support entity-level visual queries. While valuable for structured retrieval, they do not capture implicit social and eventive knowledge required for narrative and contextual reasoning. Commonsense-oriented MMKGs. A smaller body of work ex- plores multimodal commonsense. MMEKG [19] integrates event knowledge with visual signals derived from situation recognition [43]. TIVA-KG [41] incorporates text, image, audio, and video modalities, with a strong emphasis on physical commonsense. However, these resources typically focus on either event-centric or physical aspects and do not jointly model physical, social, and eventive dimensions in a unified framework. 2.3 Positioning of MMCOMET MMCOMET bridges the gap between textual commonsense graphs and encyclopedic multimodal KGs. Unlike text-only CSKGs, M- COMET provides explicit visual grounding for social, physical, and eventive relations. Unlike existing MMKGs, it focuses on everyday inferential commonsense rather than factual entity relations. In Ta- ble 1, MMCOMET is the first large-scale multimodal commonsense knowledge graph that unifies social interaction, event dynamics, and physical attributes within nearly 1M multimodal triples. This positioning enables MMCOMET to function as a reusable resource for multimodal retrieval, reasoning, and generation tasks. MMCOMET: A Large-Scale Multimodal Commonsense Knowledge Graph for Contextual ReasoningSIGIR’26, July, Naarm, Melbourne Figure 2: A small subset of MMCOMET (top) and the Con- struction Pipeline (bottom). 3 Method 3.1 Base Commonsense Topology MMCOMET is constructed by extending the textual commonsense graph ATOMIC2020 with a visually grounded layer. For each com- monsense triple< head, relation, tail>, we associate represen- tative images for the head and tail phrases through a hybrid retrieval pipeline. Figure 2 illustrates a subset of MMCOMET (top) and the overall construction workflow (bottom). We adopt ATOMIC2020 [13] as the textual backbone. ATOMIC2020 contains∼1.33M inferential commonsense triples spanning social, physical, and eventive rela- tions. After filtering incomplete and non-visualizable relation types, MMCOMET retains 19 relation categories and 989K valid triples. The goal of MMCOMET is not to alter the relational structure, but to ground each textual node with representative visual evidence, enabling multimodal retrieval and reasoning. 3.2 Hybrid Visual Alignment Pipeline Associating images with commonsense phrases presents two chal- lenges: (i) physical concepts are often visually concrete and (i) social and eventive concepts are frequently abstract. To address this, we design a two-stage retrieval strategy: 3.2.1Embedding-based Similarity Matching. For visually concrete phrases (e.g.,ObjectUse,AtLocation), we adopt embedding-based retrieval. Head and tail phrases are encoded using CLIP [26] text encoders and matched against pre-encoded image embeddings from open-source captioning and storytelling datasets via cosine simi- larity. To ensure scalability, we introduce a noun-indexed filtering mechanism. Image captions are POS-tagged to extract noun to- kens, which are organized into an inverted index mapping nouns to candidate images. Similarity is first computed at the noun level to identify a reduced candidate subset, after which fine-grained phrase-to-image similarity is applied. This hierarchical filtering reduces similarity computations by approximately 60×compared to brute-force search while preserving retrieval quality. 3.2.2 Concreteness-aware Web Retrieval. For abstract social and eventive phrases (e.g.,React,Intent,Want), caption-based similar- ity matching often fails to produce adequate visual grounding. We therefore employ a concreteness-aware routing strategy. We com- pute phrase-level concreteness scores using the 40K Concreteness Ratings [6]. Phrases with an average concreteness score below 4.0, a threshold aligned with the empirical distribution of social and even- tive relations, are routed to web-based image retrieval. Structured web queries are issued with entity normalization and text-filtering heuristics to reduce noise. Approximately 72% of unique phrases are handled through this abstract-concept routing process. 3.3 Post-processing and Image Selection All candidate images (from similarity matching and web retrieval) are re-ranked using CLIP similarity with the original phrase. Low- alignment samples are filtered using a threshold of 0.15, and the top 10–15 images per phrase are retained. The final MMCOMET graph, therefore, consists of 989K multimodal commonsense triples, 578K unique textual nodes, and∼3.9M associated images in social, phys- ical, and eventive relations. This hybrid pipeline enables scalable multimodal grounding while maintaining high alignment quality, forming a retrieval-ready multimodal commonsense resource. 4 Experimental Setup We evaluate MMCOMET along two complementary axes: (1) In- trinsic Quality Evaluation, measuring the semantic alignment between commonsense phrases and retrieved images; and (2) Ex- trinsic Utility Evaluation, assessing whether MMCOMET im- proves downstream vision-language reasoning tasks. 4.1 Datasets Commonsense Source. We adopt ATOMIC2020 [13] as the textual backbone. After removing incomplete tuples and theisFilledBy relation (which contains blank placeholders unsuitable for visual grounding), MMCOMET contains 989K multimodal tuples spanning 19 relations, with 578K unique phrases. Image Corpora. Similarity-based grounding utilizes images from Conceptual Captions [32], COCO [16], Flickr30K [44], and the Vi- sual Storytelling dataset [12]. Abstract phrases are grounded via web image retrieval and filtered using CLIP similarity. Auxiliary Resources. Concreteness ratings from [6] determine routing between similarity-based and web-based retrieval. 4.2 Implementation Details All textual and visual embeddings are computed using CLIP (ViT- B/32) [26]. For similarity-based retrieval, cosine similarity is used. For web retrieval, we collect the top 10 images and filter them using a similarity threshold of 0.15 before retaining the top 15 final images per phrase. We use NLTK [4] for POS tagging and lemmatization when constructing the noun-indexed image dictionary. 4.3 Intrinsic Evaluation: Knowledge Quality We conduct a human evaluation to assess the semantic alignment between commonsense phrases and their retrieved images. Sampling. For each of the 19 relations, 60 phrases (20 heads and 40 tails) are randomly sampled, resulting in 1140 phrases. To balance SIGIR’26, July, Naarm, MelbourneEileen Wang, Hiba Arnaout, Dhita Putri Pratama, Shuo Yang, Dangyang Liu, Jie Yang, Josiah Poon, Jeff Pan, and Soyeon Caren Han retrieval strategies, 50% of the samples are drawn from phrases that required web search. Annotation Protocol. For each phrase, annotators are shown the top 7 retrieved images. Each image is labeled as: 1) Full Match: All key concepts expressed in the phrase are visually present, 2) Partial Match: The image is thematically relevant but missing important elements, 3) No Match: The image does not correspond to the phrase. Additionally, annotators rate phrase concreteness as Low, Medium, High. In total, 7,980 phrase–image pairs are evaluated. We report the proportion of phrases where at least 5 out of 7 images are Full or Partial matches, indicating reliable grounding. 4.4 Extrinsic Evaluation: Downstream Utility We evaluate whether MMCOMET improves reasoning performance in three downstream tasks. 4.4.1Visual Storytelling (VST). We augment a standard visual sto- rytelling framework [39] by injecting multimodal commonsense retrieved from MMCOMET into intermediate reasoning steps. Per- formance is evaluated using RoViST-VG, RoViST-C, and RoViST- NR [38], along with SPICE [2], BLEURT [31], and MoverScore [47]. 4.4.2 Visual Question Answering (VQA). We evaluate on VQAv2 [10]. Given an image-question pair, we retrieve relevant MMCOMET phrases using combined image– and question–phrase similarity: 푆= sim(퐸 푖푚푎푔푒 ,퐸 푝ℎ푟푎푠푒 )+ sim(퐸 푞푢푒푠푡푖표푛 ,퐸 푝ℎ푟푎푠푒 ) Top-ranked phrases are injected as auxiliary knowledge in a few- shot prompting setup. We compare performance with and without MMCOMET knowledge injection using accuracy as the metric. 4.4.3Image Captioning. For caption generation, we retrieve M- COMET phrases whose associated images are most similar to the input image from Flickr30k [44]. The highest-scoring head/tail phrases are appended as auxiliary contexts to the captioning model. Performance is evaluated using standard captioning metrics: BLEU [25], CIDEr [37], and BLEURT [31]. 5 Results and Analysis We present results via two dimensions: (1) intrinsic evaluation of multimodal alignment quality, and (2) extrinsic validation of MMCOMET as a reusable knowledge resource for downstreams. 5.1 Intrinsic Evaluation: Human Study We conduct a human evaluation to assess the semantic alignment between commonsense phrases and their associated images in M- COMET. Table 2 reports the percentage of phrases for which at least 5 out of the top 7 retrieved images were judged as either Full Match (FM) or Partial Match (PM). Across all relations and image sources, MMCOMET achieves 90.7% FM+PM alignment and 77.5% FM alignment, indicating that the majority of commonsense phrases are reliably grounded with visually relevant evidence. Although Full Match scores are naturally lower, partially matched images re- main semantically relevant and preserve core contextual elements, making them suitable for downstream retrieval and reasoning. Web- based retrieval (WSE) consistently outperforms similarity matching over open-source caption datasets (OID). Aggregated across all re- lations, WSE achieves 94.5% (FM+PM) and 85.2% (FM), compared Web (WSE)Open-source Data (OID)Both PhysicalFM+PM FMFM+PMFMFM+PM FM ObjectUse1007396769874.5 AtLocation10092.910010010096.4 MadeUpOf 10010089.389.394.194.1 HasProperty92.692.688.584.690.688.7 CapableOf95.887.595.870.895.879.2 Desires 10095.71007210083.3 NotDesires10092.6887294.282.7 EventFM+PM FMFM+PMFMFM+PM FM isAfter91.369.677.345.584.457.8 HasSubEvent 96.392.692.971.494.581.8 isBefore87.562.588.96388.262.7 HinderedBy 86.773.376.242.980.655.6 Causes96.296.287.579.29288 Reason1009283.37591.883.7 SocialFM+PM FMFM+PMFMFM+PM FM Need90.58185.771.488.176.2 Attr928089.578.990.979.5 Effect 959061.942.97865.9 React90.986.461.155.677.572.5 Want80.873.186.45083.362.5 Intent958088.258.991.970.3 All Relations94.585.28769.790.777.5 Table 2: Human evaluation results showing the percentage of commonsense phrases whose retrieved images were judged as full (FM) or partial matches (PM). Web search (WSE) yields higher matching accuracy than open-source image datasets (OID), particularly for abstract social and eventive relations. to 87.0% and 69.7% for OID. This difference reflects the advantage of query-sensitive search for abstract and compositional phrases, while caption datasets often lack explicit coverage of eventive or so- cial expressions. Approximately 72% of phrases in MMCOMET are grounded via web retrieval, underscoring the importance of the hy- brid routing strategy. Alignment quality also varies across relation types. Physical relations achieve the highest grounding rates (96.1% FM+PM), followed by eventive (88.4%) and social (85.0%) relations. This trend aligns with phrase concreteness: tangible entities are easier to visually ground, whereas abstract emotional or intentional states exhibit greater ambiguity. Despite this inherent difficulty, even the most challenging relation types maintain alignment above 77% (FM+PM). Inter-annotator agreement is strong, with intraclass correlation coefficients of 0.76 (FM+PM) and 0.82 (FM), indicating consistent annotation quality. Overall, the human evaluation con- firms that MMCOMET provides high-quality multimodal grounding across diverse commonsense relations, supporting its reliability as a retrieval-oriented multimodal resource. 5.2 Concreteness Analysis by Relation Type We analyze phrase-level concreteness scores across relation types to better understand grounding difficulty and the necessity of hybrid retrieval. The left sub-table of Table 3 reports average concreteness scores per relation. Physical relations exhibit consistently higher concreteness scores than social and eventive relations, with the ex- ception of temporal relations such asisBeforeandisAfter. Social- interaction relations show the lowest average concreteness, reflect- ing their reliance on abstract emotional and behavioral descriptions. For example,Reactfrequently involves affective expressions, while Attroften consists of adjectival characterizations, both of which are inherently difficult to ground visually. The middle and right MMCOMET: A Large-Scale Multimodal Commonsense Knowledge Graph for Contextual ReasoningSIGIR’26, July, Naarm, Melbourne RelationAvg. ConcretenessRelation% Heads SearchedRelation% Tails Searched React (94228)2.95NotDesires2.82AtLocation25.52 Attr (115815)2.98Desires2.85MadeUpOf36.47 Causes (376)3.29AtLocation18.78isAfter48.98 Intent (49312)3.40MadeUpOf20.03isBefore53.50 Effect (113494)3.53HasProperty22.77CapableOf59.55 Reason (334)3.55CapableOf23.98HinderedBy64.72 HasSubEvent (12845)3.58ObjectUse25.26ObjectUse67.65 Want (110831)3.58Reason50.90HasProperty71.46 Need (89391)3.65Intent51.76HasSubEvent73.03 HinderedBy (106647)3.80Need54.11NotDesires74.03 HasProperty (5617)3.80React54.85Need75.94 Desires (2737)3.88Effect54.85Reason76.35 ObjectUse (165590)3.90isAfter55.96Effect76.54 isBefore (23208)3.91isBefore56.07Causes77.93 NotDesires (2838)3.91Want56.15Desires79.65 isAfter (22453)3.94HinderedBy56.19Want80.92 CapableOf (7968)3.96Attr57.13Intent90.35 MadeUpOf (3345)4.08HasSubEvent60.19Attr98.82 AtLocation (20221)4.15Causes65.96React99.48 Table 3: The left sub-table shows the average concreteness scores for each commonsense relation (sample size in brackets) in ascending order. The second and third sub-tables show the proportion of heads and tail phrases across each relation that required web search. Green, orange and blue highlighting indicate social, event and physical relations, respectively. ModelR-VG R-C R-NRRoViST (R)SPICE (S) BLEURT (B) MoverScore (M) UNION (U) PerplexityR+S+B+M+UStory L. SRL-pmi70.472.891.678.311.534.756.075.914.7256.351.2 TGCN-SRL-cos70.372.390.577.710.934.956.084.014.9263.452.3 VILT-pmi71.376.492.179.910.235.154.482.213.2261.851.2 VILT-cos70.276.394.280.29.334.352.183.312.1259.250.1 LLaVa-cos72.378.196.182.29.534.153.984.210.5263.651.9 LLaVA-cos-MMCOMET 푆 73.378.495.882.59.935.254.282.610.8264.450.8 LLaVA-cos-MMCOMET 푆+퐼 75.2 79.1 96.283.510.236.757.484.211.2272.953.4 Table 4: Automatic evaluation results of visual storytelling. MMCOMET-enhanced models (bottom rows) outperform baselines (SRL-pmi and TGCN-SRL-cos) in visual grounding, coherence, and overall story quality metrics, demonstrating the benefit of multimodal commonsense knowledge (for Perplexity, lower score is better.) Model VQAImage Captioning Acc.BLEUCIDErBLEURT Qwen2.5-VL0.6300.4810.2580.357 Qwen2.5-VL+MMCOMET0.6420.4980.2960.363 Table 5: Automatic evaluation results of VQA, using∼189k validation data from VQAv2, and Image Captioning tasks us- ing Flickr30k. MMCOMET-enhanced model (bottom row) out- performs the baseline in both tasks and all metrics, demon- strating the benefit of multimodal commonsense knowledge. sub-tables further report the proportion of head and tail phrases routed to web-based retrieval under the concreteness threshold. Consistent with their lower concreteness scores, social and eventive relations require substantially more web searches than physical re- lations. An interesting asymmetry emerges within certain physical relations such asHasPropertyandDesires/NotDesires. While head phrases typically contain tangible noun entities and therefore exhibit high concreteness, tail phrases often describe abstract prop- erties or mental states. As a result, tail phrases in these relations show a disproportionately high reliance on web-based grounding. The analysis confirms a strong correlation between relation-level concreteness and retrieval routing behavior. These findings empiri- cally justify the concreteness-aware hybrid alignment strategy, as a uniform similarity-based approach would inadequately ground abstract social and eventive commonsense expressions. 5.3 Extrinsic Valication of MMCOMET: Downstream Task Evaluation To assess the practical utility of MMCOMET as a reusable mul- timodal commonsense resource, we evaluate its impact on three representative vision-language tasks: VST, VQA, and IC. 5.3.1 Visual Storytelling (VST). Table 4 reports automatic evalua- tion results on the VIST benchmark. We compare baseline story- telling systems with variants augmented using MMCOMET-based multimodal commonsense retrieval. Incorporating MMCOMET consistently improves visual grounding (R-VG), coherence (R-C), and non-redundancy (R-NR). The best-performing configuration (LLaVA-cos-MMCOMET 푆+퐼 ) achieves the highest overall compos- ite score (R+S+B+M+U = 272.9), substantially outperforming non- augmented vision-language baselines. Notably, improvements are most pronounced in grounding-sensitive metrics (R-VG) and se- mantic similarity measures (BLEURT, MoverScore), indicating that multimodal commonsense contributes additional contextual cues beyond raw visual features. Lower perplexity further suggests im- proved narrative fluency. Qualitative examples (Tables 6 show that MMCOMET-enhanced stories exhibit richer event descriptions, clearer social interactions, and more specific contextual details compared to generic baseline outputs. SIGIR’26, July, Naarm, MelbourneEileen Wang, Hiba Arnaout, Dhita Putri Pratama, Shuo Yang, Dangyang Liu, Jie Yang, Josiah Poon, Jeff Pan, and Soyeon Caren Han ImageBaseline (SRL-pmi) Ours (LLaVA-cos-MMCOMET 푆+퐼 ) Merchants set up their booths. Merchants set up vibrant booths filled with handmade goods. It was a family gathering. It was a family gathering in a restau- rant where everyone enjoyed various delicacies. Celebrated with a birthday cake. Blowing out the candles of her birth- day cake, she made a wish for hap- piness and adventure. Table 6: Selected story frames and their associated text to showcase the effect of MMCOMET in visual story telling. ImageQuestionGTBaselineOurs How many cars on truck?0over 20+0 Is the bus in a parking space?YesNoYes Are the people getting on the bus or off the bus?onoff buson Table 7: Selected visual question-answer exemplars to show- case the effect of MMCOMET as additional contexts for VQA tasks. (GT: Ground Truth, Baseline: Qwen2.5-VL Only, Ours: Qwen2.5-VL-MMCOMET-augmented) ImageBaselineOurs A woman cradles a sleeping child, both appearing calm and serene. A woman tenderly holds a sleeping child, symbolizing the comfort and care pro- vided by a mother. A person walks down a wet brick side- walk under a yellow checkered umbrella, wearing a black coat and tall boots. A person walks down a wet sidewalk un- der a yellow umbrella, protecting them- selves from the rain. A young boy wearing a blue cap looks through a telescope, focusing intently on the view. A young boy observes through a tele- scope, learning about the wonders of the cosmos. Table 8: Comparison of image captions from baseline and MMCOMET-enhanced models to provide the impact of M- COMET in image captioning tasks. (GT: Ground Truth, Baseline: Qwen2.5-VL Only, Ours: Qwen2.5-VL-MMCOMET- augmented) 5.3.2 Visual Question Answering (VQA). Table 5 presents results on the VQAv2 validation set. Augmenting Qwen2.5-VL with M- COMET improves accuracy from 0.630 to 0.642 (+1.1%), demonstrat- ing that retrieved multimodal commonsense provides complemen- tary reasoning signals for question answering. Qualitative examples (Table 7) show that MMCOMET particularly benefits questions re- quiring implicit commonsense reasoning, such as understanding intent, spatial context, or event interpretation. 5.3.3 Image Captioning. We evaluate MMCOMET on Flickr30k caption generation. Knowledge-augmented models consistently outperform baselines across BLEU, CIDEr, and BLEURT metrics (Table 5). The improvement in CIDEr (+0.038) indicates better align- ment with reference descriptions, while BLEURT gains suggest enhanced semantic richness. Example captions (Table 8) demon- strate that MMCOMET introduces socially and eventively grounded interpretations (e.g., intent, purpose, or emotional context), leading to more informative and context-aware captions. Overall, across all three tasks, MMCOMET consistently improves performance without model retraining, confirming its effectiveness as a plug-and-play multimodal commonsense resource. 5.4 Qualitative Analysis We analyze representative examples from VSR (Table 6), VQA (Table 7), and IC (Table 8) to examine how MMCOMET influ- ences model behavior. Across all tasks, improvements extend be- yond verbosity; MMCOMET enables models to incorporate socially grounded, event-structured, and contextually coherent information that extends beyond literal visual content. 5.4.1Visual Storytelling. Baseline models typically generate literal, scene-level descriptions. In contrast, MMCOMET-augmented out- puts consistently introduce additional contextual structure that can be grouped into three categories: (1) Physical specificity. Descrip- tions become more concrete by grounding objects and attributes beyond minimal recognition. For example, instead of stating that “merchants set up their booths,” the MMCOMET-enhanced model specifies “vibrant booths filled with handmade goods,” enriching the scene with object-type information inferred from common- sense associations. (2) Social interaction modeling. Enhanced outputs incorporate implicit interpersonal dynamics and collec- tive reactions. Rather than simply noting that a father told a joke, the augmented model describes shared emotional responses (“had everyone in stitches”), reflecting socially grounded reasoning. (3) Eventive structure and intent. In scenarios such as birthday celebrations, MMCOMET enables the model to describe culturally structured actions (“blowing out candles” and “making a wish”), capturing event schemas and associated intentions that are not directly observable from pixels alone. 5.4.2Visual Question Answering. Qualitative VQA examples show that MMCOMET particularly benefits questions requiring contex- tual inference rather than direct visual counting or object detection. For instance, the baseline misinterprets bus boarding directions and parking contexts, whereas the MMCOMET-augmented model correctly resolves them using commonsense constraints on trans- portation events and spatial usage. 5.4.3 Image Captioning. In captioning, baseline outputs focus on visible entities and actions. MMCOMET-augmented captions in- troduce inferred purpose (“protecting themselves from the rain”), relational interpretation (“comfort and care provided by a mother”), and learning intent (“learning about the wonders of the cosmos”). These additions reflect eventive and social commonsense integra- tion rather than mere lexical expansion. 6 Conclusion We introduce MMCOMET, a large-scale multimodal commonsense knowledge graph that integrates physical, social, and eventive re- lations with visually grounded evidence. Intrinsic evaluation con- firms high-quality multimodal alignment, while experiments on visual storytelling, visual question answering, and image caption- ing demonstrate consistent improvements when MMCOMET is used as an external resource. These gains are achieved without task-specific retraining, highlighting MMCOMET’s plug-and-play applicability. By focusing on everyday inferential commonsense rather than encyclopedic facts, MMCOMET complements existing multimodal knowledge graphs and supports retrieval-augmented vision-language reasoning. Future work will expand modality cov- erage and further improve scalable grounding strategies. MMCOMET: A Large-Scale Multimodal Commonsense Knowledge Graph for Contextual ReasoningSIGIR’26, July, Naarm, Melbourne References [1]Houda Alberts, Ningyuan Huang, Yash Deshpande, Yibo Liu, Kyunghyun Cho, Clara Vania, and Iacer Calixto. 2021. VisualSem: a high-quality knowledge graph for vision and language. In Proceedings of the 1st Workshop on Multilingual Representation Learning. [2]Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016. Spice: Semantic propositional image caption evaluation. In European conference on computer vision. Springer, 382–398. [3]Hiba Arnaout, Simon Razniewski, Gerhard Weikum, and Jeff Z Pan. 2022. Un- commonsense: Informative negative knowledge about everyday concepts. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management. [4]Steven Bird, Ewan Klein, and Edward Loper. 2009. Natural language processing with Python: analyzing text with the natural language toolkit. " O’Reilly Media, Inc.". [5]Antoine Bosselut, Hannah Rashkin, Maarten Sap, Chaitanya Malaviya, Asli Ce- likyilmaz, and Yejin Choi. 2019. COMET: Commonsense transformers for auto- matic knowledge graph construction. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. [6]Marc Brysbaert, Amy Beth Warriner, and Victor Kuperman. 2014. Concreteness ratings for 40 thousand generally known English word lemmas. Behavior research methods (2014). [7]Hua Cai, Xuli Shen, Qing Xu, Weilin Shen, Xiaomei Wang, Weifeng Ge, Xiaoqing Zheng, and Xiangyang Xue. 2023. Improving Empathetic Dialogue Generation by Dynamically Infusing Commonsense Knowledge. In Findings of the Association for Computational Linguistics. [8] Sebastián Ferrada, Benjamin Bustos, and Aidan Hogan. 2017. IMGpedia: a linked dataset with content-based analysis of Wikimedia images. In The Semantic Web– ISWC 2017: 16th International Semantic Web Conference, Vienna, Austria, October 21-25, 2017, Proceedings, Part I 16. [9]François Gardères, Maryam Ziaeefard, Baptiste Abeloos, and Freddy Lecue. 2020. Conceptbert: Concept-aware representation for visual question answering. In Findings of the Association for Computational Linguistics. [10] Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017. Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering. In Conference on Computer Vision and Pattern Recognition (CVPR). [11]Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020. Retrieval augmented language model pre-training. In The 37th International Conference on Machine Learning. [12] Ting-Hao K. Huang, Francis Ferraro, Nasrin Mostafazadeh, Ishan Misra, Jacob Devlin, Aishwarya Agrawal, Ross Girshick, Xiaodong He, Pushmeet Kohli, Dhruv Batra, et al.2016. Visual Storytelling. In 15th Annual Conference of the North American Chapter of the Association for Computational Linguistics. [13] Jena D Hwang, Chandra Bhagavatula, Ronan Le Bras, Jeff Da, Keisuke Sakaguchi, Antoine Bosselut, and Yejin Choi. 2021. (comet-) atomic 2020: On symbolic and neural commonsense knowledge graphs. In Proceedings of the AAAI Conference on Artificial Intelligence. [14]Filip Ilievski, Pedro Szekely, and Bin Zhang. 2021. Cskg: The commonsense knowledge graph. In The 18th International Semantic Web Conference. [15]Liwei Jiang, Antoine Bosselut, Chandra Bhagavatula, and Yejin Choi. 2021. "I’m Not Mad": Commonsense Implications of Negation and Contradiction. In Pro- ceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. [16] Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. In Computer Vision–ECCV: 13th European Conference. [17]Hugo Liu and Push Singh. 2004. ConceptNet—a practical commonsense reasoning tool-kit. BT technology journal (2004). [18]Ye Liu, Hui Li, Alberto Garcia-Duran, Mathias Niepert, Daniel Onoro-Rubio, and David S Rosenblum. 2019. MMKG: multi-modal knowledge graphs. In The 16th European Semantic Web Conference. [19] Yubo Ma, Zehao Wang, Mukai Li, Yixin Cao, Meiqi Chen, Xinze Li, Wenqi Sun, Kunquan Deng, Kun Wang, Aixin Sun, et al.2022. MMEKG: Multi-modal event knowledge graph towards universal representation across modalities. In Proceed- ings of the 60th Annual Meeting of the Association for Computational Linguistics: System Demonstrations. [20]Bhavana Dalvi Mishra, Niket Tandon, and Peter Clark. 2017. Domain-targeted, high precision knowledge extraction. Transactions of the Association for Compu- tational Linguistics (2017). [21]Tuan-Phong Nguyen, Simon Razniewski, Julien Romero, and Gerhard Weikum. 2022. Refined commonsense knowledge from large-scale web contents. IEEE Transactions on Knowledge and Data Engineering (2022). [22]Tuan-Phong Nguyen, Simon Razniewski, Aparna Varde, and Gerhard Weikum. 2023. Extracting cultural commonsense knowledge at scale. In Proceedings of the ACM Web Conference. [23]Christina Niklaus, Matthias Cetto, André Freitas, and Siegfried Handschuh. 2018. A survey on open information extraction. In Proceedings of the 27th International Conference on Computational Linguistics. [24] Daniel Oñoro-Rubio, Mathias Niepert, Alberto García-Durán, Roberto González, and Roberto J López-Sastre. 2017. Answering visual-relational queries in web- extracted knowledge graphs. arXiv preprint (2017). [25] Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, Pierre Isabelle, Eugene Charniak, and Dekang Lin (Eds.). Association for Computational Linguistics, Philadelphia, Pennsylvania, USA, 311–318. doi:10.3115/1073083. 1073135 [26] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al.2021. Learning transferable visual models from natural language supervision. In International conference on machine learning. [27]Simon Razniewski, Hiba Arnaout, Shrestha Ghosh, and Fabian Suchanek. 2024. Completeness, Recall, and Negation in Open-world Knowledge Bases: A Survey. ACM Comput. Surv. (2024). [28]Julien Romero and Simon Razniewski. 2020. Inside quasimodo: Exploring con- struction and usage of commonsense knowledge. In Proceedings of the 29th ACM International Conference on Information & Knowledge Management. [29]Julien Romero, Simon Razniewski, Koninika Pal, Jeff Z. Pan, Archit Sakhadeo, and Gerhard Weikum. 2019. Commonsense properties from query logs and question answering forums. In Proceedings of the 28th ACM International Conference on Information and Knowledge Management. [30]Maarten Sap, Ronan Le Bras, Emily Allaway, Chandra Bhagavatula, Nicholas Lourie, Hannah Rashkin, Brendan Roof, Noah A Smith, and Yejin Choi. 2019. Atomic: An atlas of machine commonsense for if-then reasoning. In Proceedings of the AAAI conference on artificial intelligence. [31]Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020. BLEURT: Learning Robust Metrics for Text Generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 7881–7892. [32]Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018. Con- ceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics. [33]Robyn Speer, Joshua Chin, and Catherine Havasi. 2017. Conceptnet 5.5: An open multilingual graph of general knowledge. In Proceedings of the AAAI conference on artificial intelligence. [34] Robert Speer and Catherine Havasi. 2013. ConceptNet 5: A large semantic network for relational knowledge. The People’s Web Meets NLP: Collaboratively Constructed Language Resources (2013). [35] Niket Tandon, Gerard De Melo, Fabian Suchanek, and Gerhard Weikum. 2014. Webchild: Harvesting and organizing commonsense knowledge from the web. In Proceedings of the 7th ACM international conference on Web search and data mining. [36]Niket Tandon, Gerard De Melo, and Gerhard Weikum. 2017. Webchild 2.0: Fine- grained commonsense knowledge distillation. In Proceedings of ACL 2017, System Demonstrations. [37] Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015. CIDEr: Consensus-based image description evaluation. In 2015 IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR). 4566–4575. doi:10.1109/CVPR.2015. 7299087 [38]Eileen Wang, Caren Han, and Josiah Poon. 2022. RoViST: Learning Robust Metrics for Visual Storytelling. In Findings of the Association for Computational Linguistics. [39]Eileen Wang, Soyeon Caren Han, and Josiah Poon. 2024. SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics. [40]Meng Wang, Haofen Wang, Guilin Qi, and Qiushuo Zheng. 2020. Richpedia: a large-scale, comprehensive multi-modal knowledge graph. Big Data Research (2020). [41]Xin Wang, Benyuan Meng, Hong Chen, Yuan Meng, Ke Lv, and Wenwu Zhu. 2023. TIVA-KG: A multimodal knowledge graph with text, image, video and audio. In Proceedings of the 31st ACM International Conference on Multimedia. [42]Yinan Wu, Xiaowei Wu, Junwen Li, Yue Zhang, Haofen Wang, Wen Du, Zhidong He, Jingping Liu, and Tong Ruan. 2023. Mmpedia: A large-scale multi-modal knowledge graph. In The 22nd International Semantic Web Conference. [43]Mark Yatskar, Luke Zettlemoyer, and Ali Farhadi. 2016. Situation recognition: Visual semantic role labeling for image understanding. In Proceedings of the IEEE conference on computer vision and pattern recognition. [44]Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014. From image descriptions to visual denotations: New similarity metrics for semantic infer- ence over event descriptions. Transactions of the Association for Computational Linguistics (2014). SIGIR’26, July, Naarm, MelbourneEileen Wang, Hiba Arnaout, Dhita Putri Pratama, Shuo Yang, Dangyang Liu, Jie Yang, Josiah Poon, Jeff Pan, and Soyeon Caren Han [45]Hongming Zhang, Daniel Khashabi, Yangqiu Song, and Dan Roth. 2020. Tran- somcs: From linguistic graphs to commonsense knowledge. In Proceedings of the 29th International Conference on International Joint Conferences on Artificial Intelligence. [46]Jingdan Zhang, Jiaan Wang, Xiaodan Wang, Zhixu Li, and Yanghua Xiao. 2023. Aspectmmkg: A multi-modal knowledge graph with aspect-aware entities. In Pro- ceedings of the 32nd ACM International Conference on Information and Knowledge Management. [47] Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M Meyer, and Steffen Eger. 2019. MoverScore: Text Generation Evaluating with Contextualized Em- beddings and Earth Mover Distance. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing. [48]Yimin Zhou, Yiwei Sun, and Vasant Honavar. 2019. Improving image captioning by leveraging knowledge graphs. In IEEE winter conference on applications of computer vision.