Paper deep dive
A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images
Jennifer D'Souza, Fahad Ahmed, Cecilia Andrea Bustamante Andrade, Lina Frolova, Poorani Gnanasambandan, Dilshad Hussain, Muhammad Uzair Khan, Nkembeng Kevin Nkengfoa, Paul Praveen J., Fabio Priante, Sjoerd Franciscus van der Werf, Thomas Frederik Jan van Roeden
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/17/2026, 4:58:37 AM
Summary
This paper introduces the ALD/E-ImageMiner benchmark and the ICDAR 2026 Competition, which provide 1,951 expert-annotated scientific figures from Atomic Layer Deposition/Etching literature. The benchmark evaluates multimodal AI capabilities including figure classification, data table extraction, summarization, and visual question answering (VQA). The authors propose using Bloom's taxonomy to design VQA tasks that probe deeper scientific reasoning, aiming for a long-term goal of 'scientific conceptual understanding from images' to advance machine-actionable scientific visual knowledge.
Entities (10)
Relation Signals (9)
ALD/E-ImageMiner â contains â 1,951 figures
confidence 98% ¡ The ALD/E-ImageMiner benchmark ... provide 1,951 figures from 205 publications
Jennifer DâSouza â affiliatedwith â TIB Leibniz Information Centre for Science and Technology
confidence 97% ¡ Jennifer DâSouza 1 ... Data Science and Digital Libraries, TIB Leibniz Information Centre for Science and Technology
ALD/E-ImageMiner â supportstask â data table extraction
confidence 95% ¡ evaluates four complementary capabilities: figure classification, data table extraction, summarization, and VQA
ALD/E-ImageMiner â supportstask â summarization
confidence 95% ¡ evaluates four complementary capabilities: figure classification, data table extraction, summarization, and VQA
ALD/E-ImageMiner â supportstask â Visual Question Answering
confidence 95% ¡ evaluates four complementary capabilities: figure classification, data table extraction, summarization, and VQA
ALD/E-ImageMiner â supportstask â figure classification
confidence 95% ¡ evaluates four complementary capabilities: figure classification, data table extraction, summarization, and VQA
ALD/E-ImageMiner â usesframework â Bloom's Taxonomy
confidence 94% ¡ The VQA component was deliberately designed using Bloomâs revised taxonomy as a scaffold
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and multimodal AI systems to retrieve and interpret. The ALD/E-ImageMiner benchmark and ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching Scientific Figures provide 1,951 figures from 205 publications, expert-annotated for classification, data table extraction, summarization, and visual question answering. In these companion proceedings, we present a forward-looking perspective on how the benchmark can guide future scientific-image challenges. We examine how its tasks probe capabilities from visual and quantitative reading to domain-grounded reasoning and evidential justification, and how Bloom-informed question design can support deeper scientific understanding. We propose "scientific conceptual understanding from images" as a long-term benchmark objective, with future directions including broader domains and figure types, contextual and cross-document synthesis, hypothesis evaluation, provenance, uncertainty, counterfactual grounding, and open-ended multimodal research. This perspective connects the ICDAR 2026 challenge to a broader agenda for machine-actionable scientific visual knowledge and verifiable multimodal scientific AI.
Tags
Links
- Source: https://arxiv.org/abs/2608.14075v1
- Canonical: https://arxiv.org/abs/2608.14075v1
Trouble viewing inline? Open PDF directly â
Full Text
55,063 characters extracted from source content.
Expand or collapse full text
ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figures Sci-ImageMiner 2026 Competition Proceedings TIB-OP will set DOI with Š Authors. This work is licensed under a Creative Commons Attribution 4.0 International License Submitted: 2026-08-10 A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images Jennifer DâSouza 1 , Fahad Ahmed 1 , Cecilia Andrea Bustamante Andrade 2 , Lina Frolova 3 , Poorani Gnanasambandan 2 , Dilshad Hussain 4 , Muhammad Uzair Khan 5 , Nkembeng Kevin Nkengfoa 6 , Paul Praveen J. 7 , Fabio Priante 8 , Sjoerd Franciscus van der Werf 2 , and Thomas Frederik Jan van Roeden 1 Data Science and Digital Libraries, TIB Leibniz Information Centre for Science and Technology, Hannover, Germany 2 Department of Applied Physics and Science Education, Eindhoven University of Technology, Eindhoven, Netherlands 3 Department of Biology, Chemistry, and Pharmacy, Freie Universit Ě at Berlin, Berlin, Germany 4 HEJ Research Institute of Chemistry, International Center for Chemical and Biological Sciences, University of Karachi, Karachi, Pakistan 5 School of Interdisciplinary Engineering and Sciences, National University of Sciences and Technology, Islamabad, Pakistan 6 Department of Chemistry, University of Warwick, Coventry, United Kingdom 7 Department of Physics, PSG College of Technology, Coimbatore, Tamil Nadu, India 8 Department of Chemistry and Materials Science, Aalto University, Helsinki, Finland *Correspondence: jennifer.dsouza@tib.eu Abstract: Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and multimodal AI systems to retrieve and inter- pret. The ALD/E-ImageMiner benchmark and ICDAR 2026 Competition on Informa- tion Extraction from Atomic Layer Deposition/Etching Scientific Figures provide 1,951 figures from 205 publications, expert-annotated for classification, data table extrac- tion, summarization, and visual question answering. In these companion proceed- ings, we present a forward-looking perspective on how the benchmark can guide future scientific-image challenges. We examine how its tasks probe capabilities from visual and quantitative reading to domain-grounded reasoning and evidential justification, and how Bloom-informed question design can support deeper scientific understanding. We propose âscientific conceptual understanding from imagesâ as a long-term benchmark objective, with future directions including broader domains and figure types, contextual and cross-document synthesis, hypothesis evaluation, provenance, uncertainty, coun- terfactual grounding, and open-ended multimodal research. This perspective connects the ICDAR 2026 challenge to a broader agenda for machine-actionable scientific visual knowledge and verifiable multimodal scientific AI. Keywords: scientific images, multimodal AI, vision-language models, digital libraries, benchmarks, scientific reasoning arXiv:2608.14075v1 [cs.AI] 14 Aug 2026 DâSouza & Ahmed |Open Conf Proc X (2025) âICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figuresâ ⢠URL: https://sciknoworg.github.io/ALD-E-ImageMiner/ ⢠Huggingface: https://huggingface.co/datasets/SciKnowOrg/ALD-E-ImageMiner ⢠License: Mixed right non-commercial 1 Introduction Scientific communication is intrinsically multimodal. Researchers convey experimental observations, quantitative results, structural relationships, and methodological details through charts, spectra, microscopy images, molecular and reaction diagrams, process flows, and apparatus schematics. These visual artifacts are not merely illustrations of the surrounding text; they often constitute the primary record of the evidence on which scientific claims are based. For digital libraries, this creates a fundamental challenge: scholarly knowledge cannot be comprehensively indexed, retrieved, or analyzed when figures and tables remain accessible only as undifferentiated image objects. Future scientific AI systems must therefore be able to extract and reason over information represented jointly in visual and textual forms, from interpreting spectroscopic plots and microstructural features to understanding experimental configurations [1]. Without this capability, both AI-assisted scientific analysis and the transformation of digital libraries into multimodal knowledge infrastructures remain incomplete. Recent visionâlanguage models have broadened natural-language access to visual content, but their performance on scientific figures remains uneven [1], [2]. Scientific interpretation often requires numerical fidelity, domain-specific notation, spatial and re- lational reasoning, cross-panel evidence integration, and experimental context. Models that perform well on general captioning or visual question answering (VQA) may still misread axes and legends, overlook visual relations, or produce conclusions unsup- ported by the figure. Materials-science chart-to-table studies make this gap especially clear: advanced multimodal models may identify axes, legends, and overall trends while hallucinating or omitting data points in dense plots with overlapping markers or fit- ted curves [3]. Reliable scientific-image understanding therefore requires both seman- tic interpretation and numerically faithful evidence extraction. These limitations con- strain applications such as structured experimental-data extraction, evidence-grounded search, visual literature synthesis, and scientific question answering. Although sys- tems such as ORKGEx already incorporate visionâlanguage methods into scholarly knowledge-curation workflows [4], domain-grounded benchmarks remain necessary for evaluating their reliability. The ALD/E-ImageMiner benchmark and the ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching Scientific Figures provide a concrete response to this need [2]. The benchmark brings together expert-annotated figures from experimental and simulation-based atomic layer deposition and etching (ALD/E) literature and evaluates four complementary capabilities: figure classification, data ta- ble extraction, summarization, and VQA. The competition report describes the dataset construction, annotation process, evaluation protocols, participating systems, and re- sults. In the present vision paper, we use the benchmark and the experience gained from organizing the competition as a foundation for considering what future scientific- image challenges should measure and how their scope could develop. Within these companion proceedings, this contribution complements the competition report [2] and the participating-team studies by connecting the current challenge to a longer-term research agenda. We first examine ALD/E-ImageMiner as a testbed for ca- pabilities ranging from visual localization and quantitative reading to domain-grounded DâSouza & Ahmed |Open Conf Proc X (2025) âICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figuresâ interpretation, reasoning, and evidential justification. We then discuss how Bloom- informed question design can support different forms of engagement with the same scientific evidence and introduce scientific conceptual understanding from images as a broader benchmark objective. Finally, we outline future extensions across materi- als science and related engineering domains, richer visual representations, contextual and cross-document synthesis, hypothesis evaluation, provenance, uncertainty, coun- terfactual grounding, and open-ended multimodal research. 2 Related Work: Gaps in Scientific Image Understanding 2.1 Benchmarks for Chart and Scientific Figure Understanding Chart-understanding benchmarks have established important tasks for extracting and reasoning over visually represented data, but much of this work has focused on syn- thetic or general-purpose visualizations. FigureQA [5] and DVQA [6] use programmati- cally generated charts paired with template-based questions, while PlotQA [7] provides a large collection of automatically generated plots and associated question-answer pairs. StructChart extends this line of work through SimChart9K, a language-model- driven collection of synthetic chartâtable pairs intended to support structured chart un- derstanding [8]. ChartQA [9] introduced 20,882 real-world charts with human-authored questions, thereby broadening the visual and linguistic diversity of chart question an- swering; however, its web-sourced economic, social, and survey visualizations differ substantially from the specialized plots, diagrams, spectra, and composite figures found in scientific publications. More directly related to materials science, Circi et al. introduce PolyCompChartIE and MetalThermoChartIE, two human-annotated benchmarks for chart-to-table extrac- tion from real scatter plots concerning polymer composites and the thermophysical properties of metal systems [3]. Their figures contain heterogeneous axis ranges, dense or overlapping data points, fitted curves, and categorical information repre- sented through marker shapes, colors, and legends. The work also proposes Rel- ative Coordinate-Label Similarity (RCLS), which evaluates whether extracted tables preserve the correspondence among numerical coordinates, axis labels, and series labels. This demonstrates that domain-specific chart understanding requires both rep- resentative scientific data and evaluation measures aligned with the structure of the visual evidence. Together, these benchmarks have advanced chart classification, chart-to-table con- version, and visual question answering, while also revealing the additional demands introduced by scientific figures. Their meaning may depend on domain-specific nota- tion, experimental conditions, measurement techniques, spatial organization, and re- lationships among several panels or modalities. A spectrum, microscopy image, reac- tion scheme, process diagram, or apparatus schematic cannot always be interpreted through the representations and question templates used for conventional statistical charts. Scientific-image benchmarks must therefore combine visual diversity with do- main knowledge, numerically appropriate evaluation, and task definitions that reflect how evidence is communicated and used within scientific practice. 2.2 Limitations of VisionâLanguage Models on Scientific Visual Evidence Recent evaluations indicate that strong performance on general visual benchmarks does not reliably transfer to scientific figures. MaCBench [1], which evaluates visionâ DâSouza & Ahmed |Open Conf Proc X (2025) âICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figuresâ language models on chemistry and materials-science problems, reports strong perfor- mance on tasks such as laboratory-equipment recognition and direct numerical extrac- tion, but substantially weaker results when spatial reasoning, cross-modal synthesis, or multi-step inference is required. The MAC benchmark [10], constructed from scientific journal covers and their associated stories, similarly shows that models can recognize salient entities while relying heavily on textual cues; their performance declines when those cues are removed and interpretation must depend more directly on the visual content. More general investigations of the visual binding problem also show that mod- els such as GPT-4V can fail to reliably associate objects, attributes, and spatial relations even in comparatively simple scenes [11]. Together, these findings suggest that visual recognition and textual fluency do not by themselves ensure reliable interpretation of scientific evidence. Evaluations focused on particular visual forms reinforce this conclusion. Compar- isons of GPT-4V and Gemini identify difficulties with table reasoning and formula recog- nition despite strong general captioning and question answering [12]. ChartInsights [13] reports considerable weaknesses in fine-grained chart analysis, while FlowLearn [14] shows that model performance varies substantially across flowchart tasks such as OCR, node identification, and structural interpretation. Evaluations on real materials- science scatter plots further show that models may recover axes and legend infor- mation while producing incomplete or numerically inaccurate tables, with performance deteriorating on denser and more visually complex figures [3]. Domain adaptation can improve results: visual instruction tuning supports stronger general multimodal perfor- mance [15], and LLaVA-Med [16] demonstrates the value of training on biomedical fig- ures. Nevertheless, domain-adapted systems may still hallucinate visual details or fail on questions requiring several inferential steps. These results motivate benchmarks that evaluate complementary capabilities on the same domain-specific figures rather than relying on a single aggregate measure of visual understanding. 2.3 Multimodal AI for Scholarly Knowledge Infrastructures Scientific-image understanding is also becoming relevant to digital-library systems that seek to transform publications into structured and searchable knowledge. ORKGEx [4], for example, combines language and vision models with the Open Research Knowl- edge Graph [17] to support the annotation of scholarly contributions. Its workflow rec- ognizes that conventional OCR is insufficient for figures in which text, graphical marks, spatial relations, and domain-specific symbols jointly express meaning, and it incorpo- rates visionâlanguage methods for tasks including figure classification, chart-to-table conversion, and summarization. This illustrates how multimodal models can contribute to scholarly knowledge curation rather than serving only as standalone VQA systems. However, the integration of scientific visual evidence into digital infrastructures re- mains at an early stage. Existing systems generally provide limited support for domain- specific interpretation, panel-level grounding, numerically faithful extraction, and the joint evaluation of several capabilities on the same figures. The literature therefore re- veals a gap between general chart benchmarks, evaluations that expose weaknesses in scientific visual reasoning, and digital-library tools intended to produce structured scholarly knowledge. ALD/E-ImageMiner addresses this intersection through an expert- annotated, domain-grounded benchmark covering heterogeneous scientific figures and four complementary tasks: figure classification, data table extraction, summarization, and visual question answering [2]. Next we examine how these tasks provide a testbed for scientific visual intelligence and a foundation for future challenge design. DâSouza & Ahmed |Open Conf Proc X (2025) âICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figuresâ 3 ALD/E-ImageMiner as a Testbed for Scientific Visual Intelligence ALD/E-ImageMiner is not merely a collection of scientific figures, but a testbed for evaluating increasingly demanding vision-language capabilities in scientific settings. It contains 1,951 figures from 205 experimental and simulation-based ALD/E publica- tions, organized into 49 figure categories [2]. These include line and multi-line charts, spectra, scatter plots, heatmaps, molecular-structure and reaction diagrams, band and apparatus diagrams, process flows, and composite image panels. The Hugging Face release enables systematic browsing by figure type: https://huggingface.co/ datasets/SciKnowOrg/ALD-E-ImageMiner. This diversity reflects the many ways scien- tific evidence is encoded, including through axes and legends, spectral peaks, chemical notation, spatial relationships, and domain-specific visual conventions. Further details on the benchmark, annotations, evaluation, competition, systems, and results are pro- vided in the ICDAR competition report [2]. In this paper, we use the benchmark as a basis for articulating desiderata for future vision-language model research, including the design of scientific benchmarks. 3.1 From Seeing to Scientific Reasoning Scientific figure comprehension can be understood as a progression from seeing to reading, understanding, reasoning, and ultimately justifying. First, a model must see the relevant visual structure by distinguishing subfigures, axes, legends, labels, sym- bols, curves, regions, and other graphical elements. It must then read the figure with sufficient quantitative fidelity to recover values, units, categories, and relations without losing their structural organization. To understand the figure, the model must connect these observations to the represented scientific variables, experimental conditions, pro- cesses, and domain terminology. Scientific reasoning further requires deriving rela- tionships and implications that cannot be obtained through direct transcription alone, including causal mechanisms, comparative trends, structureâproperty relations, and application-oriented conclusions. A trustworthy system should additionally be able to justify its response by locating the relevant visual evidence, distinguishing direct ob- servations from domain-informed interpretations, and expressing uncertainty when the available evidence is insufficient. The four ALD/E-ImageMiner tasks provide complementary operationalizations of this progression. TASK 1. Figure classification probes whether a model can identify the representational form through which information is communicated. TASK 2. Data table extraction tests whether quantitative values and their relations can be recovered with structural and numerical fidelity. TASK 3. Summarization examines whether a model can move beyond transcription to identify the central scientific message conveyed by a figure. TASK 4. Visual question answering (VQA) then targets the most explicit transition from visual perception to domain-grounded scientific reasoning. 3.2 Bloom-Informed Visual Question Answering The VQA component was deliberately designed using Bloomâs revised taxonomy as a scaffold for eliciting progressively deeper levels of reasoning [18]. The questions are grounded in the scientific characteristics of atomic layer deposition and etching (ALD/E), where material growth or removal is controlled through cyclic, surface-mediated processes. Figures in this literature therefore communicate more than isolated numer- ical values: they describe precursor and co-reactant exposures, purge and reaction steps, process windows, growth-per-cycle behavior, changes in composition and struc- DâSouza & Ahmed |Open Conf Proc X (2025) âICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figuresâ ture, and the resulting material or device performance [2]. Understanding such figures may require a model to trace a sequence of surface reactions, relate processing condi- tions to measured trends, connect film structure or composition to physical properties, or interpret the implications of those properties for a particular application. In contrast to generic VQA datasets dominated by object recognition or direct lookup, the questions in ALD/E-ImageMiner are organized around four domain-grounded fami- lies: process-oriented, comparative/trend, structureâproperty, and application/performance questions. Together, these families span cognitive operations ranging from remem- bering and understanding visible information to applying domain concepts, analyzing relationships, and evaluating scientific or technological implications. Bloomâs taxonomy is used here as a question-design framework without implying that model computation is equivalent to human cognition. Moreover, the question families do not correspond rigidly to individual Bloom levels. The same figure can support a direct factoid question, an analytical comparison, a causal explanation, or an evaluative judgment. This makes it possible to vary not only the visual content presented to a model, but also the depth at which that content must be interpreted. Process-oriented questions. These questions concern the cyclic and surface-mediated nature of ALD/E, including precursor exposures, purge steps, surface reactions, modification and removal stages, and complete deposition or etching workflows. A sequencing question may primarily involve Remembering and Understanding; explaining why a purge step is necessary requires Understanding and Analyzing; and predicting the consequence of changing a pulse or purge condition requires Applying scientific knowledge to a new situation. Comparative/trend questions. These questions compare experimental variables and determine how changes in temperature, pulse length, cycle count, precursor ra- tio, or composition affect outcomes such as growth per cycle, etch per cycle, film thickness, or spectral intensity. Directly reporting a trend may involve Understand- ing, whereas explaining correlations, identifying competing effects, or predicting behavior under an unobserved condition requires Applying and Analyzing. Com- parisons among alternative process routes can additionally elicit Evaluation. Structureâproperty questions. These questions ask models to connect precursor chemistry, film composition, molecular or layer structure, doping arrangement, or interface organization with resulting material properties. They therefore move beyond identifying visible entities and require the model to Apply domain knowl- edge and Analyze how changes in structure or composition influence electronic, optical, chemical, mechanical, or surface behavior. Application/performance questions. These questions relate experimentally observed material behavior to device-level performance or practical use. They may re- quire a model to determine whether a spectral change indicates improved emis- sion, whether a chromaticity coordinate is relevant to a lighting application, or how a measured material response could support sensing, energy conversion, or semiconductor-device fabrication. Such questions primarily target Analysis and Evaluation, because the answer must connect visual evidence to a broader sci- entific or technological objective. Questions are paired with four answer formats: yes/no, factoid, list, and paragraph. These formats permit different levels of granularity while avoiding the assumption that all scientific reasoning must be expressed as long-form text. Factoid and list answers can test precise recognition, quantitative recovery, or process sequencing, whereas paragraph answers require models to make their reasoning and interpretation more DâSouza & Ahmed |Open Conf Proc X (2025) âICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figuresâ explicit. The annotation design includes explanatory paragraph answers so that the benchmark can evaluate not only whether a conclusion is correct, but also whether it is scientifically coherent and grounded in the figure. The following multi-panel example illustrates how this Bloom-informed VQA design is instantiated through panel-level annotation and machine-readable visual grounding in ALD/E-ImageMiner. 3.3 A Multi-Panel Example from ALD/E-ImageMiner Figure 1 provides a concrete example of how the benchmark combines heteroge- neous scientific visual content with panel-specific task annotations. During dataset preparation, MinerU [19] was first applied to extract the structured textual content and high-resolution figures from each publication. The document text and figure captions were stored in a top-level content.json file, while the extracted figures and their cor- responding annotation files were organized separately. To facilitate machine reading of composite figures, each subfigure was manually associated with its panel letter and localized using machine-readable bounding-box coordinates (x,y, width, height). This enabled annotations associated with a particular panel to be paired with the corre- sponding image region when presented to a vision-language model, rather than requir- ing the model to infer the target region from the complete multi-panel figure. Figure 1. Reproduced from Figure 3 in Mattinen et al., Van der Waals epitaxy of continuous thin films of 2D materials using atomic layer deposition in low temperature and low vacuum conditions [20]. The ALD/E-ImageMiner classification annotations identify (a) a molecular structure diagram, (b) a polar chart (rose chart), (câd) spectra charts, and (e) a stacked spectra chart. The figure also supports domain-grounded questions concerning epitaxial growth, crystallographic orientation, and disorder in PbI 2 films grown on sapphire. For this example, the five panels were localized using the coordinates shown in Ta- ble 1, expressed in pixels relative to the complete figure. The panel identifiers link each task annotation to its intended subfigure, while the coordinates allow the corresponding visual region to be deterministically isolated. For example, a question associated with panel (b) can be supplied together with the crop defined by (623, 40, 508, 436). The bounding boxes are not assumed to improve model performance by themselves; rather, they provide explicit machine-readable grounding DâSouza & Ahmed |Open Conf Proc X (2025) âICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figuresâ Table 1. Panel-level bounding boxes for Figure 1. PanelxyWidthHeight (a)1555597392 (b)62340508436 (c)17493374286 (d)393488363290 (e)778492360286 so that the annotation, visual input, and evaluation consistently refer to the intended panel. This figure was annotated with four domain-grounded questions covering three ques- tion families and multiple answer formats: Panel (a): Process-Oriented; Yes/No. Question: Can PbI 2 be grown epitaxially over sapphire? Answer: Yes. Panel (b): StructureâProperty; Paragraph. Question: Why was the pole figure mea- sured at 25.95 ⌠? Answer: The pole figure measured at 25.95 ⌠2θ corresponds to the asymmetric (1 0 Ě 1 1) PbI 2 reflection, but is also sufficiently close to the asymmetric (0 1 Ě 1 2) sapphire reflection to provide information about the filmâsubstrate registration. Three sharp peaks over a 360 ⌠rotation were observed for the sapphire reflection, whereas six broader peaks were observed for PbI 2 ; considering the symmetry of the reflections, both patterns are consistent with a single in-plane orientation. The alignment of the (1 0 Ě 1 1) PbI 2 and (0 1 Ě 1 2) sapphire reflections corresponds to the epitaxial relation (0 0 1)[2 Ě 1 1 0] PbI 2 ⼠(0 0 1)[2 Ě 1 1 0] Al 2 O 3 , indicating parallel alignment of the unit-cell axes of PbI 2 and sapphire, as expected from the relatively small lattice mismatch between the film and substrate. Panel (d): Comparative/Trend; Factoid. Question: What suggests disorder in the PbI 2 films? Answer: The relatively broad peaks in the β scan (FWHM = 13 ⌠, panel (c)) and the Îą scan (FWHM = 9 ⌠, panel (d)), extracted from the in-plane pole figure, sug- gest disorder in the in-plane and out-of-plane orientations, respectively. Panel (e): StructureâProperty; Factoid. Question: Are there 30 ⌠/90 ⌠domains in PbI 2 grown on sapphire? Answer: In-plane XRD measurements confirmed the in-plane alignment, the ab- sence of 30 ⌠/90 ⌠domains, and the presence of some non-epitaxial domains. This example illustrates how the benchmark combines panel-level visual ground- ing with different forms of scientific interpretation. A model may need to recognize the representational form of an individual panel, relate measurements across pan- els, and interpret the evidence using concepts such as epitaxy, filmâsubstrate reg- istration, crystallographic orientation, and structural disorder. The complete annota- tion record, including the summarization and data-table extraction targets, is avail- able at https://github.com/sciknoworg/ALD-E-ImageMiner/blob/main/icdar2026- competition-data/train/atomic-layer-deposition/experimental-usecase/37/images/ figure_3.json [2]. DâSouza & Ahmed |Open Conf Proc X (2025) âICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figuresâ 3.4 Toward Create-Level Scientific Visual Reasoning The preceding example shows how the current ALD/E-ImageMiner annotations progress from visual recognition to the analysis and evaluation of scientific evidence represented in a figure. A natural direction for future challenge editions is to extend this progres- sion toward Create, the highest level in Bloomâs revised taxonomy, by asking models to formulate scientifically plausible experimental workflows or material designs under explicit constraints. Such tasks would require models to synthesize distributed visual in- formation into an ordered and coherent plan while distinguishing supported steps from parameters not specified in the available evidence, such as pulse duration, deposition temperature, target thickness, or donorâacceptor distance. Their evaluation could con- sider the inclusion and ordering of essential stages, consistency with the visual and textual evidence, scientific feasibility, treatment of missing information, and avoidance of unsupported procedural details. The following hypothetical question, based on Fig- ure 1 of Ghazy et al. [21], illustrates this prospective extension and is not part of the current ALD/E-ImageMiner annotations. Question: Based on the illustrated process and optical configuration, outline a minimal experimental workflow for preparing and testing a EuâHQA hybrid layer for a FRET-based application. Answer: 1. Repeat the following ALD/MLD cycle until the target thickness is reached: Eu(thd) 3 pulseâ N 2 purgeâ HQA pulseâ N 2 purge. 2. Deposit the hybrid layer either on a plasmonic nanostructure for emission enhance- ment or on flat Si as a reference. 3. Introduce AF647 under conditions that control the donorâacceptor spacing required for measurable FRET. More broadly, future challenge editions could introduce paired questions of increas- ing cognitive depth for the same figure, complemented by cross-panel reasoning, provenance- aware answers, and uncertainty assessment. These additions would enable a richer evaluation of scientific visual reasoning. 4 Toward âScientific Conceptual Understandingâ from Images: A Challenge Roadmap We propose scientific conceptual understanding from images as a long-term objec- tive for multimodal AI and digital libraries. This capability extends beyond recognizing visual forms, extracting values, or producing plausible descriptions. It requires a sys- tem to use visual evidence within a scientific process: to identify what was observed, relate it to experimental conditions and domain knowledge, compare it with alterna- tive evidence, and determine what conclusions are warranted. The current ALD/E- ImageMiner tasks provide complementary and independently evaluable probes of this objective. Classification identifies how evidence is represented, data extraction recov- ers machine-readable quantitative content, summarization captures the principal mes- sage of a figure, and VQA evaluates domain-grounded interpretation and reasoning. Building on this foundation, we outline two directions for the future development of Sci-ImageMiner. The first concerns the incremental expansion of its figure types and scientific domains, initially across materials science and subsequently into related en- gineering disciplines. The second concerns new task objectives that connect scientific DâSouza & Ahmed |Open Conf Proc X (2025) âICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figuresâ images with contextual interpretation, cross-figure synthesis, hypothesis evaluation, provenance, uncertainty, and broader research workflows. 4.1 Figure-Type and Domain Extensions The first direction is to expand ALD/E-ImageMiner incrementally while preserving its expert-annotated and domain-grounded character. Within materials science, the near- est extensions include thin-film synthesis and characterization, semiconductor process- ing, two-dimensional materials, catalysis, electrochemistry and battery materials, pho- tovoltaics, polymers, ceramics, and composite materials. These areas share many of the experimental techniques and processingâstructureâproperty relations already encountered in ALD/E, while introducing different materials, mechanisms, and perfor- mance criteria. Such a progression would make it possible to evaluate whether a model transfers its understanding of familiar visual conventions to related scientific settings. It would also support concrete applications such as retrieving experiments conducted under comparable conditions, comparing fabrication routes across material classes, constructing structured materials-property records, and identifying earlier studies that report similar structural or performance trends. The expansion should be guided by the scientific functions of visual representa- tions rather than by visual diversity alone. Microscopy images from scanning elec- tron microscopy (SEM), transmission electron microscopy (TEM), and atomic force mi- croscopy (AFM), and related techniques could support the characterization of morphol- ogy, interfaces, defects, particle distributions, and film coverage. Diffraction patterns, reciprocal-space maps, and spectroscopic figures could support phase identification, crystallographic-orientation analysis, peak assignment, and comparison of chemical or structural states. Phase diagrams and multidimensional property maps could be con- verted into searchable process windows, stability regions, and compositionâproperty relations, while atomistic models, band structures, density-of-states plots, device cross- sections, reaction schemes, process diagrams, and apparatus schematics could rep- resent mechanisms, configurations, and experimental workflows that cannot be ade- quately reduced to ordinary tables. The resulting benchmark would test whether mod- els can transform different visual languages into scientifically useful representations while preserving quantities, relations, constraints, and uncertainty. Related work on chemical structure recognition provides a concrete example: molecular and Markush structure images are translated into machine-readable SMILES or CXSMILES records encoding atoms, bonds, variable sites, and substituent constraints [22]. Such annota- tions would extend the benchmark toward chemical diagrams while enabling program- matic structure search, database indexing, and the construction of training data for downstream chemistry models. After deepening its materials-science coverage, the dataset could expand into related engineering sciences that communicate knowledge through similarly specialized visual forms. Chemical and process engineering would introduce reactor schematics, pro- cess flowsheets, transport plots, and kinetic models; electrical and electronic engineer- ing would contribute circuits, device architectures, semiconductor cross-sections, and timing diagrams; mechanical and aerospace engineering would add stress fields, frac- ture images, simulation outputs, and CAD-like representations; and energy or biomed- ical engineering would provide system configurations, diagnostic images, and device- performance figures. A modular annotation framework could preserve a shared core of panel localization, classification, structured extraction, summarization, VQA, and source provenance while adding schemas and question families specific to each field. DâSouza & Ahmed |Open Conf Proc X (2025) âICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figuresâ Metadata describing the domain, subdomain, visual form, experimental or computa- tional setting, and required background knowledge would permit controlled compar- isons between in-domain performance, transfer to neighboring fields, and generaliza- tion across substantially different scientific conventions. 4.2 Task-Objective Extensions The scientific research lifecycle spans literature discovery, evidence synthesis, hypoth- esis formulation, experimentation, content production, and evaluation [23]. Future Sci- ImageMiner challenges could connect figure understanding to these wider research activities while retaining clear task-wise evaluation and extending the benchmark to- ward increasingly contextual, comparative, and open-ended uses of scientific visual evidence. Context-grounded and machine-actionable interpretation. A contextual figure-understanding task could provide a figure together with selected captions, methods, and surrounding passages and require the model to distinguish among information directly visible in the figure, information reported in the text, and conclusions requiring domain knowledge. The output could include both a natural- language interpretation and a structured record of variables, values, units, conditions, and relations. For example, a system might identify a spectral shift from the image, retrieve the corresponding deposition condition from the methods, and mark a pro- posed reaction mechanism as an interpretation rather than an observation. Such out- puts would support figure-aware paper chat, scientific database population, evidence- grounded search, and the integration of visual observations into knowledge graphs. Hypothesis formation, discriminating tests, and revision. A more demanding track could present visual evidence and ask a system to generate competing explanations, identify which additional measurement would best distinguish them, and revise its assessment when new evidence is introduced. For example, broad diffraction peaks might be explained by small crystallite size, strain, or orientational dis- order; a model could be asked to identify which microscopy, diffraction, or complemen- tary spectroscopy result would discriminate among these possibilities. The benchmark could then add a new panel and test whether the system updates its conclusion rather than rationalizing its initial answer. This would support experiment planning, measure- ment selection, materials troubleshooting, and the design of follow-up studies. Such evaluation could use a structured epistemic record rather than an unrestricted chain of thought. A record might contain the selected visual evidence, candidate hy- potheses, predicted observations, chosen test, resulting evidence, and revised conclu- sion. These fields expose claims that can be checked against the figure and domain annotations without assuming that the record reproduces the modelâs internal compu- tation. Bloom-informed question families could define the intended cognitive demand, while explicit evidenceâhypothesisâtestârevision links would measure whether the ob- servable response follows a scientifically disciplined structure. Cross-panel and cross-document synthesis. Many scientific conclusions depend on several complementary measurements. Cross- panel tasks could require models to connect a structural image with a spectrum, re- DâSouza & Ahmed |Open Conf Proc X (2025) âICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figuresâ late a control sample to a treated sample, or explain how changes in process condi- tions correspond to changes in measured properties. At the document level, a sys- tem could trace a material from synthesis and characterization to device evaluation. Cross-document tasks could then compare related figures from multiple publications, normalize experimental variables, and determine whether reported results agree, dif- fer because of their conditions, or remain genuinely contradictory. These capabilities would support visual literature reviews, evidence synthesis, reproducibility analysis, and the construction of searchable processingâstructureâproperty networks. Counterfactual grounding, provenance, and uncertainty. Because plausible explanations may be unfaithful [24] and relevant computation need not be fully visible in generated text [25], explanation quality alone should not be used as evidence that a model reasoned correctly. Future tasks could instead introduce controlled counterfactuals: remove a relevant panel, alter a plotted value, exchange two labels, add contradictory evidence, or modify irrelevant formatting. A grounded model should change its answer when the scientific evidence changes, remain stable under irrelevant visual perturbations, and abstain when the necessary evidence is absent. Requiring each answer to identify its supporting panel, bounding box, plotted value, or document passage would further enable claim verification, figure-aware peer review, and transparent scholarly retrieval. Open-ended multimodal research challenges. The eventual objective should extend beyond answering predetermined questions to- ward evaluating whether multimodal agents can use visual literature in an open-ended research process. Recent case studies suggest that agents can complete substantial research engineering while still struggling with experimental judgment, evidence se- lection, backtracking, and determining whether a research question has actually been resolved [26], [27]. A future âshadowâ challenge could give an agent a research ques- tion derived from an unpublished or carefully held-out study, together with a literature collection containing figures, tables, methods, and data. The agent could be asked to synthesize prior evidence, formulate hypotheses, select analyses or experiments, and produce a justified research proposal or result for assessment by domain experts. To preserve its multimodal focus, the task would center on reasoning over figures and tables, with textual passages, and associated data serving as supporting context. Such open-ended evaluations should complement rather than replace automatically scored tasks. Classification, extraction, summarization, and VQA provide reproducible diagnostics of specific capabilities, while structured epistemic tasks evaluate evidence use and revision, and expert-assessed challenges probe scientific judgment under less constrained conditions. Together, these evaluation layers could distinguish a system that merely produces a plausible final answer from one that reliably connects visual observations, scientific claims, tests, and decisions. The resulting vision is a continuously expanding multimodal scientific environment rather than a static collection of images. A future Sci-ImageMiner system could retrieve related visual evidence, transform it into machine-actionable knowledge, compare find- ings across publications, propose competing explanations, request the next informative measurement, revise its conclusions, and report what remains uncertain. Its progress would remain measurable through independently scored tasks, but its broader pur- DâSouza & Ahmed |Open Conf Proc X (2025) âICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figuresâ pose would be to evaluate whether multimodal AI can participate responsibly in the evidence-driven and self-correcting practices of science. 5 Conclusion Scientific figures and tables are primary carriers of experimental evidence, yet much of their content remains difficult to search, compare, and integrate within digital knowl- edge infrastructures. ALD/E-ImageMiner provides a domain-grounded testbed for ad- dressing this gap through complementary tasks covering classification, structured data extraction, summarization, and visual question answering. Building on this founda- tion, we have proposed scientific conceptual understanding from images as a broader benchmark objective: the ability to interpret visual evidence in context, relate it to sci- entific claims and experimental conditions, and determine which conclusions are sup- ported. Rather than treating this objective as a single measure of âunderstanding,â the proposed roadmap decomposes it into independently evaluable capabilities that can be extended across figure types, scientific domains, and research activities. The future development of Sci-ImageMiner should therefore proceed incrementally, first broadening its coverage within materials science and related engineering disci- plines, and then introducing tasks involving contextual interpretation, cross-figure syn- thesis, hypothesis evaluation, provenance, uncertainty, and open-ended multimodal re- search. Such benchmarks could enable digital libraries to move beyond caption-based indexing toward machine-actionable representations of visual evidence, evidence-grounded retrieval, visual literature synthesis, and support for experimental planning and scien- tific evaluation. Achieving this vision will require sustained domain-expert annotation, transparent task design, and evaluation methods that reward numerical fidelity, eviden- tial grounding, appropriate revision, and abstention when the available information is insufficient. In this way, scientific image understanding can become a concrete and progressively measurable foundation for general-purpose scientific AI that learns from the full multimodal scholarly record. Data availability statement The ALD/E-ImageMiner benchmark data discussed in this contribution are publicly available at https://github.com/sciknoworg/ALD-E-ImageMiner/tree/main/icdar2026- competition-data and are described in the ICDAR 2026 competition report [2]. The release includes scientific figure images, machine-readable annotations and metadata, structured content files, evaluation scripts, and submission guidelines. Source article PDFs are intentionally excluded from the GitHub distribution. ALD/E-ImageMiner is released as a mixed-rights, non-commercial benchmark re- source. The annotations and generated metadata are available under C BY 4.0, whereas extracted figure images and source-derived content remain subject to the rights and reuse terms of their respective source articles. There is therefore no blan- ket open license covering the complete corpus. Users must consult the corresponding source publication and rights-holder terms before redistributing, reusing, or commer- cially exploiting individual images. Further details are provided in the repository-level li- cense notice: https://github.com/sciknoworg/ALD-E-ImageMiner/blob/main/LICENSE. DâSouza & Ahmed |Open Conf Proc X (2025) âICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figuresâ Underlying and related material The ALD/E-ImageMiner competition project page provides access to the benchmark dataset, documentation, and related resources: https://sciknoworg.github.io/ALD-E-ImageMiner/ Codabench submission-format guidelines and the official evaluation scripts used on the Codabench competition pages are available at: ⢠Submission-format guidelines: https://github.com/sciknoworg/ALD-E-ImageMiner/ tree/main/icdar2026-competition-data/test/submission_guidelines ⢠Official competition evaluation scripts: https://github.com/sciknoworg/ALD-E- ImageMiner/tree/main/icdar2026-competition-data/evaluation_scripts The live Codabench leaderboards remain open indefinitely for submissions, subject to the continued availability of the Codabench platform. The competition pages for the respective tasks are listed below. ⢠Classification task: https://w.codabench.org/competitions/12901/ ⢠Data table extraction task: https://w.codabench.org/competitions/12902/ ⢠Summarization task: https://w.codabench.org/competitions/12909/ ⢠Visual question answering task: https://w.codabench.org/competitions/12908/ Author contributions J.D.: Conceptualization, Funding acquisition, Methodology, Project administration, Re- sources, Supervision, Writing â original draft. F.A.: Conceptualization, Methodology, Project administration, Resources, Software. C.A.B.A., L.F., P.G., D.H., M.U.K., N.K.N., P.P.J., F.P., S.F.v.d.W., and T.F.J.v.R.: Data curation, Investigation, Validation. Authors 3â12 contributed as domain-expert dataset annotators and are listed alpha- betically by family name. Competing interests The authors declare that they have no competing interests. Funding The creation of the ALD/E-ImageMiner benchmark dataset was supported by the NFDI4 DataScience initiative, funded by the German Research Foundation (DFG, Grant ID: 460234259). Acknowledgements We would like to thank all participants in the Sci-ImageMiner competition for their valu- able contributions, thoughtful engagement, and the scientific exchange that made the competition a particularly enriching experience. We also thank the ICDAR 2026 com- petition chairs for the opportunity to host the competition as part of ICDAR 2026. We also thank Eleni Poupaki (TUE, NL), Alex Watkins (UOW, UK), Bora Karasulu (UOW, UK), Adrie Mackus (TUE, NL), Erwin Kessels (TUE, NL), and Vijay K. Narasimhan (M Ventures, USA) for extended discussions. These scientists, together with the first au- DâSouza & Ahmed |Open Conf Proc X (2025) âICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figuresâ thor, participated in the âAI-Aware Pathways to Sustainable Semiconductor Process and Manufacturing Technologies (AWASES)â program, funded by Merck and Intel. References [1] N. Alampara et al., âProbing the limitations of multimodal language models for chemistry and materials research,â Nature computational science, p. 1â10, 2025. [2] F. Ahmed, S. Auer, and J. DâSouza, âIcdar 2026 competition on information extraction from atomic layer deposition/etching (ald/e) scientific figures,â arXiv preprint arXiv:2607.26848, 2026. [Online]. Available: https://arxiv.org/abs/2607.26848 [3] D. Circi et al., âInformation extraction from diverse charts in materials science,â in LLM for Scientific Discovery: Reasoning, Assistance, and Collaboration. [4] H. Hussein, F. Ahmed, A. Oelen, R. Ewerth, and S. Auer, âORKGEx: Leveraging lan- guage and vision models with knowledge graphs for research contribution annotation,â in Proceedings of the Thirteenth International Conference on Building and Exploring Web Based Environments (WEB 2025), P. Dini, Ed., IARIA Conference, March 9â13, 2025, Lisbon, Portugal: IARIA, Mar. 2025, ISBN: 978-1-68558-243-2. [5] S. E. Kahou, V. Michalski, A. Atkinson, Ě A. K Ě ad Ě ar, A. Trischler, and Y. Bengio, âFigureqa: An annotated figure dataset for visual reasoning,â arXiv preprint arXiv:1710.07300, 2017. [6] K. Kafle, B. Price, S. Cohen, and C. Kanan, âDvqa: Understanding data visualizations via question answering,â in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, p. 5648â5656. [7] N. Methani, P. Ganguly, M. M. Khapra, and P. Kumar, âPlotqa: Reasoning over scien- tific plots,â in Proceedings of the ieee/cvf winter conference on applications of computer vision, 2020, p. 1527â1536. [8] R. Xia et al., âStructchart: On the schema, metric, and augmentation for visual chart understanding,â arXiv e-prints, arXivâ2309, 2023. [9] A. Masry, X. L. Do, J. Q. Tan, S. Joty, and E. Hoque, âChartqa: A benchmark for question answering about charts with visual and logical reasoning,â in Findings of the Association for Computational Linguistics: ACL 2022, 2022, p. 2263â2279. [10] M. Jiang, J. Gao, J. Zhan, and D. Wang, âMac: A live benchmark for multimodal large language models in scientific understanding,â in Proceedings of the Conference on Lan- guage Modeling (COLM), 2025. [Online]. Available: https://arxiv.org/abs/2508. 15802 [11] D. Campbell et al., âUnderstanding the limits of vision language models through the lens of the binding problem,â Advances in Neural Information Processing Systems, vol. 37, p. 113 436â113 460, 2024. [12] Z. Qi et al., âGemini vs gpt-4v: A preliminary comparison and combination of vision- language models through qualitative cases,â CoRR, 2023. [13] Y. Wu, L. Yan, L. Shen, Y. Wang, N. Tang, and Y. Luo, âChartinsights: Evaluating mul- timodal large language models for low-level chart question answering,â arXiv preprint arXiv:2405.07001, 2024. [14] H. Pan, Q. Zhang, C. Caragea, E. Dragut, and L. Jan Latecki, âFlowlearn: Evaluating large vision-language models on flowchart understanding,â in ECAI 2024, IOS Press, 2024, p. 73â80. [15] H. Liu, C. Li, Q. Wu, and Y. J. Lee, âVisual instruction tuning,â Advances in neural infor- mation processing systems, vol. 36, p. 34 892â34 916, 2023. DâSouza & Ahmed |Open Conf Proc X (2025) âICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figuresâ [16] C. Li et al., âLlava-med: Training a large language-and-vision assistant for biomedicine in one day,â Advances in Neural Information Processing Systems, vol. 36, p. 28 541â 28 564, 2023. [17] S. Auer et al., âImproving access to scientific literature with knowledge graphs,â Bibliothek Forschung und Praxis, vol. 44, no. 3, p. 516â529, 2020. [18] D. R. Krathwohl, âA revision of bloomâs taxonomy: An overview,â Theory into practice, vol. 41, no. 4, p. 212â218, 2002. [19] B. Wang et al., âMineru: An open-source solution for precise document content extraction,â arXiv preprint arXiv:2409.18839, 2024. [20] M. Mattinen et al., âVan der waals epitaxy of continuous thin films of 2d materials using atomic layer deposition in low temperature and low vacuum conditions,â 2D Materials, vol. 7, no. 1, p. 011 003, 2020. [21] A. Ghazy, J. Yl Ě onen, N. Subramaniyam, and M. Karppinen, âAtomic/molecular layer de- position of europiumâorganic thin films on nanoplasmonic structures towards fret-based applications,â Nanoscale, vol. 15, no. 38, p. 15 865â15 870, 2023. [22] A. Andonian, S. G. Rodriques, A. D. White, and S. M. Narayanan, âMarkushglyph and oc- srglyph: Improved chemical structure recognition,â arXiv preprint arXiv:2607.28532, 2026. [23] S. Eger et al., âTransforming science with large language models: A survey on ai-assisted scientific discovery, experimentation, content generation, and evaluation,â arXiv preprint arXiv:2502.05151, 2025. [24] M. Turpin, J. Michael, E. Perez, and S. Bowman, âLanguage models donât always say what they think: Unfaithful explanations in chain-of-thought prompting,â Advances in Neural Information Processing Systems, vol. 36, p. 74 952â74 965, 2023. [25] V. Baherwani, T. Goldstein, and A. Panda, âNot all llm reasoning is visible in the chain-of- thought,â arXiv preprint arXiv:2607.22925, 2026. [26] M. R Ě Äąos-Garc Ě Äąa et al., âAi scientists produce results without reasoning scientifically,â arXiv preprint arXiv:2604.18805, 2026. [27] P. Kirgis et al., âCan ai agents conduct open-ended ai research? early evidence from two case studies,â arXiv preprint arXiv:2607.27191, 2026.