Paper deep dive
Beyond Description: A Multimodal Agent Framework for Insightful Chart Summarization
Yuhang Bai, Yujuan Ding, Shanru Lin, Wenqi Fan
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 7/20/2026, 9:28:54 PM
Summary
The paper introduces Chart Insight Agent Flow (CIAF), a plan-and-execute multi-agent framework that leverages Multimodal Large Language Models (MLLMs) to generate insightful chart summaries. It addresses the limitation of existing methods that focus on low-level data description by decomposing the task into planning, insight extraction (data and domain), and summarization. The authors also introduce ChartSummInsights, a new dataset of real-world charts with expert-authored insightful summaries, demonstrating significant performance improvements over baselines.
Entities (8)
Relation Signals (7)
Chart Insight Agent Flow → consistsof → Summarizer Agent
confidence 95% · CIAF decomposes the complex task... into three specialized stages... Planner Agent, an Insight Extraction Agent, and a Summarizer Agent
Chart Insight Agent Flow → consistsof → Planner Agent
confidence 95% · CIAF decomposes the complex task... into three specialized stages... Planner Agent, an Insight Extraction Agent, and a Summarizer Agent
Chart Insight Agent Flow → consistsof → Insight Extraction Agent
confidence 95% · CIAF decomposes the complex task... into three specialized stages... Planner Agent, an Insight Extraction Agent, and a Summarizer Agent
Chart Insight Agent Flow → uses → Multimodal Large Language Models
confidence 95% · CIAF is a plan-and-execute multi-agent framework effectively leveraging the perceptual and reasoning capabilities of MLLMs
ChartSummInsights → contains → real-world charts
confidence 92% · ChartSummInsights, a new dataset featuring a diverse collection of real-world charts paired with high-quality, insightful summaries
ChartSummInsights → usedfor → Chart Summarization
confidence 90% · we introduce ChartSummInsights... to overcome the lack of suitable benchmarks... for chart summarization
ChartSummInsights → sourcedfrom → Our World in Data
confidence 88% · We collected a chart insight dataset of 240 images from Our World in Data
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Chart summarization is crucial for enhancing data accessibility and the efficient consumption of information. However, existing methods, including those with Multimodal Large Language Models (MLLMs), primarily focus on low-level data descriptions and often fail to capture the deeper insights which are the fundamental purpose of data visualization. To address this challenge, we propose Chart Insight Agent Flow, a plan-and-execute multi-agent framework effectively leveraging the perceptual and reasoning capabilities of MLLMs to uncover profound insights directly from chart images. Furthermore, to overcome the lack of suitable benchmarks, we introduce ChartSummInsights, a new dataset featuring a diverse collection of real-world charts paired with high-quality, insightful summaries authored by human data analysis experts. Experimental results demonstrate that our method significantly improves the performance of MLLMs on the chart summarization task, producing summaries with deep and diverse insights.
Tags
Links
- Source: https://arxiv.org/abs/2602.18731v1
- Canonical: https://arxiv.org/abs/2602.18731v1
Trouble viewing inline? Open PDF directly →
Full Text
26,360 characters extracted from source content.
Expand or collapse full text
Beyond Description: A Multimodal Agent Framework for Insightful Chart Summarization Yuhang Bai The Hong Kong Polytechnic University Yujuan Ding The Hong Kong Polytechnic University Shanru Lin City University of Hong Kong Wenqi Fan The Hong Kong Polytechnic University Abstract—Chart summarization is crucial for enhancing data accessibility and the efficient consumption of information. How- ever, existing methods, including those with Multimodal Large Language Models (MLLMs), primarily focus on low-level data descriptions and often fail to capture the deeper insights which are the fundamental purpose of data visualization. To address this challenge, we propose Chart Insight Agent Flow, a plan- and-execute multi-agent framework effectively leveraging the perceptual and reasoning capabilities of MLLMs to uncover profound insights directly from chart images. Furthermore, to overcome the lack of suitable benchmarks, we introduce Chart- SummInsights, a new dataset featuring a diverse collection of real-world charts paired with high-quality, insightful summaries authored by human data analysis experts. Experimental results demonstrate that our method significantly improves the perfor- mance of MLLMs on the chart summarization task, producing summaries with deep and diverse insights. Index Terms—Chart Summarization, Insight, LLM agents I. INTRODUCTION Data visualization, particularly through charts, is an indis- pensable tool for efficiently communicating complex informa- tion, making chart analysis a significant research topic [1]. A key task in this field, chart summarization, aims to automat- ically generate a natural language summary that describes a chart’s core content and findings [2], [3]. This capability is essential for enhancing data accessibility, enabling the rapid consumption of large volumes of data, and facilitating data- driven storytelling [4]. Ultimately, it helps users quickly grasp key information, greatly enhancing efficiency and addressing information overload [5]–[7]. To achieve these goals, chart summarization methods must effectively bridge the gap be- tween visual data and natural language, which is a challenging yet significant research area [8]. The emergence of powerful Multimodal Large Language Models (MLLMs) has introduced new paradigms for this task by demonstrating remarkable abilities in visual under- standing and reasoning. While many existing methods have made progress in describing basic visual elements and factual data, they remain limited to low-level analysis [9]–[11] such as semantic understanding [12], often overlooking the most critical aspect of visualization: insight [6], [13]. The true purpose of a visualization is not merely to display data but to convey the deeper insights hidden within it [14]. These insights include high-level patterns, trends, and domain-related impacts [8], [14]. Effectively summarizing a chart, therefore, requires uncovering these underlying insights with accurate language. However, extracting such insights remains a con- siderable challenge. It necessitates deep analysis and a solid understanding of both the presented data and the underly- ing domain knowledge [15], [16]. Despite these challenges, several recent studies have made attempts. Some methods leverage Large Language Models (LLMs) to extract insights from raw structured data, but they fail to interpret visual information [17]–[19]. Other works like ChartInsights [20] and ChartInsighter [21] focus on multimodal understanding but are limited to low-level ChartQA tasks or specific chart types, and often require access to the raw data. To fill this research gap, we propose a novel plan-and- execute multi-agent framework, the Chart Insight Agent Flow (CIAF), which is a training-free pipeline based on MLLMs. As illustrated in Fig. 1, the CIAF decomposes the complex task of generating insightful chart summaries into three specialized stages. These stages are processed sequentially by a Planner Agent, an Insight Extraction Agent, and a Summarizer Agent, with each agent responsible for a distinct part of the process, from preliminary planning to final summary generation. A significant obstacle to progress in this field has been the lack of a suitable benchmark. Existing datasets for insight generation are typically derived from raw data tables [18], lacking the visual modality of charts. Conversely, chart-specific datasets often focus on low-level tasks visual question answering and summarization [2], [7], [22]. We contribute the ChartSum- mInsights dataset, which is uniquely constructed from real- world chart images and corresponding summaries annotated by human experts to capture profound, domain-specific in- sights. This expert-level annotation provides a reliable ground truth and aligns with the rich prior knowledge of MLLMs, enhancing their ability to generate coherent and contextually accurate analyses. Our contributions are threefold: 1) We contribute the Chart- SummInsights dataset, a unique collection of multi-type charts with expert-crafted, insightful summaries; 2) We propose the Chart Insight Agent Flow (CIAF), a method that leverages MLLMs’ capabilities to generate insightful summaries for charts; and 3) We design a new evaluation mechanism to arXiv:2602.18731v1 [cs.AI] 21 Feb 2026 Insight Plan Domain Context The chart reveals distinct trends in fertility across age groups from 1950 to 2023. 50-59 age groups had consistently low birth rates, reflecting biological limitations and lifestyle factors that delay or limit childbearing at older ages. Collectively, these trends... Task: Propose potential analytical angles as insight plans, and identify professional domain of the chart. Data Insights Domain Insights Task: Review the chart image , execute the analysis based on the insight plan extract data- level insight from corresponding perspective. Task:Interpret the data insights according to chart image and your domain knowledge specializing in the field of domain . Task:Synthesize polished paragraph with insight according to data insights and domain insights. Input Chart Planner Insight Extractor Summarizer Output Text Fig. 1. Chart Insight Agent Flow (CIAF) Framework, which contains three core components: Planner, Insight Extractor and Summarizer. assess models based on the quality and diversity of insights, providing a valid measure for this task. I. METHOD In this paper, we introduce a straightforward yet effective agent flow for generating insightful chart summaries, which we call the Chart Insight Agent Flow (CIAF). We leverage the capabilities of state-of-the-art MLLMs to understand visual data and interpret semantic information, as well as reasoning. Rather than building a new MLLM from scratch with complex structures or optimization techniques, our approach focuses on a novel agent-based methodology that effectively utilizes existing models to accomplish a specific task: interpreting and summarizing key insights from a given chart in clear natural language. Therefore, our method is feasible to be applied on different MLLM backbones. We do not aim to improve the chart interpretation capabilities of visual models or the sentence generation quality of language models. Instead, our primary goal is to produce chart summaries with higher-quality insights. To achieve this, our framework, as shown in Fig. 1, consists of three core agent components—Planner, Insight Ex- tractor, and Summarizer—each responsible for distinct tasks. Planner. The Planner serves as the initial stage for interpret- ing the input chart image, which accepts chart images and performs two key functions: Insight Plan Generation: The Planner analyzes the visual and semantic elements of the chart image and proposes a set of potential analytical perspectives as an insight plan. This plan outlines a sequence of analytical actions to guide subsequent insight extraction process. We use examples of possible insight perspectives paired with according chart type and specific professional domain as an In Context Learning (ICL) prompt, which facilitate MLLM learns and imitates the structured approach to insight planning. Domain Identification: The Planner identifies the profes- sional domain most relevant to the chart, which is used to prompt the following Insight Extraction Module, enabling more specialized and knowledge-grounded analysis. The output of the Planner is a structured plan containing both a list of insight plan and a domain label, which are passed to the Insight Extraction Module for execution. Insight Extractor. Based on theoretical research on In- sights [15], [23], two key sources to produce insights from a chart are data and relevant domain behind it. Therefore, we device two independent agents in this stage to collaboratively extract more comprehensive insights, serving as Data Analyst and Domain Analyst respectively: Data Analyst: It follows the insight plan generated by the Planner and extracts data insights from the chart. An ICL prompt containing multiple examples of data insight are provided. Insights extracted in this module are used for further in-depth analysis of subsequent modules. Domain Analyst: It is prompted with domain context pro- vided by the Planner and plays the role of a domain expert. It takes the data insights produced by the Data Analyst and en- hances them with relevant background knowledge, contextual implications, and real-world significance. The result is a set of domain insights that are both mean- ingful and actionable within the identified domain. The two agent analysts complement each other, thereby pro- ducing two types of insights: data and domain, which together form a comprehensive analytical view, ensuring diverse and complete insights to be extracted in this part. Summarizer. The Summarizer is implemented with an LLM, which functions to integrate the data and domain insights into a fluent, logically structured summary. The model is prompted to synthesize all extracted information, and produce a coherent narrative that captures both statistical and domain- OnMarch13,2020,anationalclosureof schoolsinVenezuelawasimplentedinthe contextoftheCOVID-19outbreak.Asof August1,2020,closetosevenmillionchildren andteenagershadbeenaffectedbythe contingencymeasuresintheLatinAmerican country.Nearly5.7millionofthestudents wereenrolledinprimaryandsecondary schoolsatthetime,whilecloseto1.2million wereenrolledinpre-primaryinstitutions. WhenitcomestotheimpacttheirDNAtest resultshavehadonhowtheyviewthemselves, 15%ofmail-intestuserssaytheirresults changedthewaytheythinkabouttheirracial orethnicidentity.Nonwhitesaretwiceas likelyaswhitestosaythis(24%vs.12%). TheozoneholeoverAntarcticawasgrowing rapidlythroughoutthe1980sandearly1990s,as thedatainthechartshows.Atitslargest,the ozoneholewasmorethan25millionsquare kilometers-slightlybiggerthanthesizeofSub- SaharanAfrica.Theearth'sozonelayeris importantastheozoneabsorbsmostofthesun's ultravioletradiation,andhelpstokeepEarth habitable.Humanemissionsofozone-depleting substances-mostlychlorofluorocarbons-were breakingdownozonehighintheatmosphere.But in1987,theworldagreedtophaseoutthese ozone-depletingsubstancesbysigningthe MontrealProtocol.Sincethen,emissionshave fallenclosetozero.Asaconsequence,theozone holestoppedgrowinginthelate1990s.Itwill takedecadestorecoverfully,butit'sslowly startingtorebuild. ChartSumm: Chart-to-Text: ChartSummInsights(Ours): 0 10 20 30 40 50 60 0% 20% 40% 60% 80% 100% Health Environment Demographic Economic Technology Democracy Agriculture Community Violence Education Total Number of Charts Percentage of Chart Types (%) areabarlinemapothersscatterTotal Images (a) (b) Fig. 2. ChartSummInsight dataset. (a) Chart Type and Domain Distribution; (b) Sample comparison with existing chart summarization datasets [2], [7]. Contains Insight? The analysis of cherry blossom peak bloom dates in Kyoto, Japan, reveals a clear long-term trend of earlier flowering, shifting from around April 20 to as early as March 21, ... This shift aligns with rising spring temperatures, offering compelling evidence of climate change's impact on biological phenology. ... * Trend in cherry blossom peak bloom dates in Kyoto. * Correlation between earlier blooming and rising spring temperatures. ... N Y Analyze Perspective Merge Perspective Input Text Output Perspectives Fig. 3. Insight Perspective Analysis Process specific aspects of the chart. The final output is a well-rounded summarization that is not only factually accurate but also contextually relevant and rich in insight. I. EXPERIMENTS Dataset. We collected a chart insight dataset of 240 images from Our World in Data 1 , which is a platform producing charts regarding the problems faced by the world, as well as the in- sightful elaboration produced by researchers at the University of Oxford and the non-profit organization Global Change Data Lab. In our collected dataset, each data point consists of a data visualization chart paired with a corresponding, expertly written summarization with insight. 1 https://ourworldindata.org/ TABLE I CHART INSIGHT SUMMARIZATION PERFORMANCE BaselinesID-RCID-SpanIQ Score Unichart [3]0.270.271.22 Matcha [24]0.110.111.16 ChartLlama [25]0.320.321.91 ChartInstruct [1]0.450.442.25 ChartGemma [26]0.380.372.71 QwenVL-3b [27]0.540.533.68 QwenVL-7b [27]0.540.524.09 QwenVL-plus [27]0.620.604.14 LlaVa-7b [28]0.650.632.29 InternVL-2b [29]0.540.533.25 InternVL-8b [29]0.640.623.97 QwenVL-plus-ours [27]0.650.634.67 The dataset spans multiple professional domains and a variety of common chart types. As shown in Fig. 2, Compared with other chart summarization datasets that simply describe the data information in the chart, our dataset contains more in-depth and actionable insights, which can better evaluate the model’s visual perception capabilities across different chart types and its knowledge reasoning abilities within various professional domains. Baselines. After surveying previous chart-to-text generation approaches, we selected the following models from 3 cat- egories for comparison, including Early-stage end-to-end chart-to-text models (Unichart [3], Matcha [24]); Fine-tuned MLLM-based moedls (ChartLlama [25], ChartInstruct [1] and ChartGemma [26]); Off-the-shelf MLLMs (Qwen-VL [27], LLaVa [28] and Intern-VL [29] series). TABLE I PERFORMANCE ON DIFFERENT BACKBONE LLMS QwenVL QwenVL QwenVLLlaVaInternVL InternVL -3b-ours -7b-ours -plus-ours -7b-ours -2b-ours -8b-ours ID-RC0.560.610.650.690.640.69 (+)0.020.070.030.040.100.05 (+%)3.7012.964.846.1518.517.81 ID-Span0.540.600.630.660.610.67 (+)0.010.080.030.030.080.05 (+%)1.8915.385.004.7615.108.06 IQ Score4.144.334.672.513.984.56 (+)0.460.240.530.220.730.59 (+%)12.505.8712.809.6122.4614.86 Summarized Insight Evaluation. To assess the performance of various baseline models, we designed a two-pronged eval- uation framework focused specifically on the insight level, rather than just the overall summary. This approach allows us to measure both the quality and diversity of the generated content. For Insight Quality (IQ) Score, we apply an LLM (GPT) to measure the depth and factual correctness of the generated insights. Each output is given a score from 1 to 5 based on how well it aligns with the ground-truth insights and 0.54 0.54 0.62 0.65 0.54 0.64 0.54 0.57 0.62 0.56 0.55 0.65 0.49 0.56 0.56 0.58 0.6 0.65 0.56 0.61 0.65 0.69 0.64 0.69 0 0.2 0.4 0.6 0.8 QwenVL- 3b QwenVL- 7b QwenVL- plus LlaVa-7b InternVL- 2b InternVL- 8b ID-RC 3.68 4.09 4.14 2.29 3.25 3.97 4.02 4.29 4.63 2.21 3.86 4.41 3.56 3.89 4.04 2.27 3.36 3.74 4.14 4.33 4.67 2.51 3.98 4.56 0 1 2 3 4 5 QwenVL- 3b QwenVL- 7b QwenVL- plus LlaVa-7b InternVL- 2b InternVL- 8b IQ-Score original+extraction+plannerCIAF Fig. 4. Performance of models with different agent components on various backbone models. facts. This method provides a nuanced measure of the model’s ability to produce meaningful and accurate observations. For Insight Diversity (ID), we employ multiple BERT-based metrics [9] including remote-clique (RC) and Span, which measure different aspects of the variety within the generated insights. To facilitate this evaluation, we first use a separate LLM agent (GPT) to post-process the generated summary. As shown in Fig. 3, this agent analyzes each sentence in the summary to extract distinct insight perspectives. These extracted perspectives are then used to calculate the diversity metrics, ensuring our evaluation is based on the core ideas, not just the surface-level wording. Experimental Results. Experimental results are shown in Table I. Our approach effectively enhances both the GPT score and the SBERT-based diversity score of the generated insight summarization across the tested backbone models. Furthermore, our method performance surpasses all fine- tuned MLLM-based methods and Early end-to-end models compared. This indicates that our method not only produces summaries with greater depth and accuracy but also achieves more comprehensive analysis, leading to more insightful sum- marization with rich content. To further analyze the performance of each module, we conducted ablation study on each core component. As shown in Fig. 4, without insight planning, the extraction module can generate in-depth insights, while lacking diversity. Thus, insight planner has been proved to successfully enrich the insight generation perspective. The overall score significantly decreased after the removal of the insight extraction module, demonstrating its effectiveness of executing the insight plan. Case Study. Fig. 5 shows an output case generated by our model, compared with output of off-the-shelf MLLM backbones. The text colored in black indicates sentences that lack insight and merely state facts, while the text colored in blue indicates insightful chart analysis. The text colored in red indicates incorrect facts. These results indicate that our Chart: QwenVL-3b: This chart shows the share of deaths from underlying causes in Denmark, Australia, the United Kingdom, and Canada from 1950 to 2022. The data indicates that cardiovascular diseases have been the leading cause of death in all four countries over the years, with a slight decline in recent decades. This trend suggests an increasing focus on cancer prevention and treatment in these regions. QwenVL-3b-ours: Across Denmark, Australia, the United Kingdom, and Canada, trends in mortality from cancer and neoplasms and cardiovascular diseases reveal significant shifts over the past seven decades. While cardiovascular disease mortality has consistently declined in all four countries—likely due to effective public health interventions, lifestyle changes, and improved chronic disease management—cancer and neoplasm mortality has shown more variation. In Denmark, cancer mortality peaked in 1960 before declining, whereas in Australia and Canada, it nearly doubled from the 1950s to 2023, suggesting rising incidence linked to environmental and demographic factors. The UK has seen relative stability in cancer mortality, indicating possible regional or cancer-type disparities. Fig. 5.A comparison case showing summary generation results of our proposed method and compared QwenVL-3b. approach can generate chart summaries with domain-relevant insights more effectively, surpassing the baseline model in analytical depth. At the same time, our method significantly reduces the factual errors caused by LLMs’ hallucinations. IV. CONCLUSION In this work, we propose a plan-and-execute multi-agent framework Chart Insight Agent Flow based on MLLMs, which can facilitate the chart summarization task with insightful findings. Additionally, a new dataset ChartSummInsights is collected with real-world charts annotated by human data analysis experts as a benchmark. Our method shows improved performance, demonstrating its ability to improve insight in chart summarization tasks. REFERENCES [1] Ahmed Masry, Mehrad Shahmohammadi, Md Rizwan Parvez, Enamul Hoque, and Shafiq Joty, “Chartinstruct: Instruction tuning for chart comprehension and reasoning,” arXiv preprint arXiv:2403.09028, 2024. [2] Raian Rahman, Rizvi Hasan, Abdullah Al Farhad, Md Tahmid Rahman Laskar, Md Hamjajul Ashmafee, and Abu Raihan Mostofa Kamal, “Chartsumm: A comprehensive benchmark for automatic chart summa- rization of long and short summaries,” arXiv preprint arXiv:2304.13620, 2023. [3] Ahmed Masry, Parsa Kavehzadeh, Xuan Long Do, Enamul Hoque, and Shafiq Joty, “Unichart: A universal vision-language pretrained model for chart comprehension and reasoning,” arXiv preprint arXiv:2305.14761, 2023. [4] Qianwen Wang, Zhutian Chen, Yong Wang, and Huamin Qu, “A survey on ml4vis: Applying machine learning advances to data visualization,” IEEE transactions on visualization and computer graphics, vol. 28, no. 12, p. 5134–5153, 2021. [5] Remco Chang, Caroline Ziemkiewicz, Tera Marie Green, and William Ribarsky,“Defining insight for visual analytics,”IEEE Computer Graphics and Applications, vol. 29, no. 2, p. 14–17, 2009. [6] Stuart K Card, Jock Mackinlay, and Ben Shneiderman, Readings in information visualization: using vision to think, Morgan Kaufmann, 1999. [7] Shankar Kantharaj, Rixie Tiffany Ko Leong, Xiang Lin, Ahmed Masry, Megh Thakkar, Enamul Hoque, and Shafiq Joty,“Chart-to-text: A large-scale benchmark for chart summarization,”arXiv preprint arXiv:2203.06486, 2022. [8] Yi He, Shixiong Cao, Yang Shi, Qing Chen, Ke Xu, and Nan Cao, “Leveraging foundation models for crafting narrative visualization: A survey,” arXiv preprint arXiv:2401.14010, 2024. [9] Hyung-Kwon Ko, Hyeon Jeon, Gwanmo Park, Dae Hyun Kim, Nam Wook Kim, Juho Kim, and Jinwook Seo, “Natural language dataset generation framework for visualizations powered by large language models,” in Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, 2024, p. 1–22. [10] Nicole Sultanum and Arjun Srinivasan, “Datatales: Investigating the use of large language models for authoring data-driven articles,” in 2023 IEEE Visualization and Visual Analytics (VIS). IEEE, 2023, p. 231– 235. [11] Benny J Tang, Angie Boggust, and Arvind Satyanarayan, “Vistext: A benchmark for semantically rich chart captioning,” arXiv preprint arXiv:2307.05356, 2023. [12] Xiao Zhang, Dongyuan Li, Liuyu Xiang, Yao Zhang, Cheng Zhong, and Zhaofeng He, “Do mllms really understand the charts?,” arXiv preprint arXiv:2509.04457, 2025. [13] Zachary Pousman, John Stasko, and Michael Mateas,“Casual in- formation visualization: Depictions of data in everyday life,” IEEE transactions on visualization and computer graphics, vol. 13, no. 6, p. 1145–1152, 2007. [14] Eun Kyoung Choe, Bongshin Lee, et al., “Characterizing visualization insights from quantified selfers’ personal data presentations,” IEEE computer graphics and applications, vol. 35, no. 4, p. 28–37, 2015. [15] Leilani Battle and Alvitta Ottley, “What exactly is an insight? a literature review,” 2023 IEEE Visualization and Visual Analytics (VIS), p. 91–95, 2023. [16] Zhicheng Liu and Jeffrey Heer, “The effects of interactive latency on exploratory visual analysis,” IEEE transactions on visualization and computer graphics, vol. 20, no. 12, p. 2122–2131, 2014. [17] R Ding, S Han, and D Zhang,“Insightpilot: An llm-empowered automated data exploration system,” in EMNLP 2023. ACL special interest group on linguistic data (SIGDAT), 2023. [18] Gaurav Sahu, Abhay Puri, Juan Rodriguez, Amirhossein Abaskohi, Mohammad Chegini, Alexandre Drouin, Perouz Taslakian, Valentina Zantedeschi, Alexandre Lacoste, David Vazquez, et al., “Insightbench: Evaluating business analytics agents through multi-step insight genera- tion,” arXiv preprint arXiv:2407.06423, 2024. [19] Luoxuan Weng, Xingbo Wang, Junyu Lu, Yingchaojie Feng, Yihan Liu, Haozhe Feng, Danqing Huang, and Wei Chen, “Insightlens: Augmenting llm-powered data analysis with interactive insight management and nav- igation,” IEEE Transactions on Visualization and Computer Graphics, 2025. [20] Yifan Wu, Lutao Yan, Leixian Shen, Yunhai Wang, Nan Tang, and Yuyu Luo, “Chartinsights: Evaluating multimodal large language models for low-level chart question answering,” arXiv preprint arXiv:2405.07001, 2024. [21] Fen Wang, Bomiao Wang, Xueli Shu, Zhen Liu, Zekai Shao, Chao Liu, and Siming Chen, “Chartinsighter: An approach for mitigating hallucination in time-series chart summary generation with a benchmark dataset,” IEEE Transactions on Visualization and Computer Graphics, 2025. [22] Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque, “Chartqa: A benchmark for question answering about charts with visual and logical reasoning,” arXiv preprint arXiv:2203.10244, 2022. [23] Yang Chen, Jing Yang, and William Ribarsky,“Toward effective insight management in visual analytics systems,” in 2009 IEEE Pacific Visualization Symposium. IEEE, 2009, p. 49–56. [24] Fangyu Liu, Francesco Piccinno, Syrine Krichene, Chenxi Pang, Kenton Lee, Mandar Joshi, Yasemin Altun, Nigel Collier, and Julian Martin Eisenschlos, “Matcha: Enhancing visual language pretraining with math reasoning and chart derendering,” arXiv preprint arXiv:2212.09662, 2022. [25] Yucheng Han, Chi Zhang, Xin Chen, Xu Yang, Zhibin Wang, Gang Yu, Bin Fu, and Hanwang Zhang, “Chartllama: A multimodal llm for chart understanding and generation,” arXiv preprint arXiv:2311.16483, 2023. [26] Ahmed Masry, Megh Thakkar, Aayush Bajaj, Aaryaman Kartha, Enamul Hoque, and Shafiq Joty, “Chartgemma: Visual instruction-tuning for chart reasoning in the wild,” arXiv preprint arXiv:2407.04172, 2024. [27] Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al., “Qwen2- vl: Enhancing vision-language model’s perception of the world at any resolution,” arXiv preprint arXiv:2409.12191, 2024. [28] Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee, “Visual instruction tuning,” Advances in neural information processing systems, vol. 36, p. 34892–34916, 2023. [29] Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al., “Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, p. 24185–24198.