Paper deep dive
From Materials Database to Materials Bank: Assetizing Data for AI Driven Materials Innovation
Chenyao Ma, Di Zhang, Weibo Gong, Wei Du, Rui Su, Yuhang Chen, Kan Xu, Huan Gu, Limin Li, Piao Ma, Zhenghao Li, Hao Li
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 7/7/2026, 2:19:03 AM
Summary
The paper introduces the Materials Bank, an industrialization-oriented framework that transforms passive materials databases into active, assetized repositories for AI-driven materials innovation. It outlines a closed-loop ecosystem integrating databases, AI large models, automated high-throughput experimentation, and a multi-dimensional assessment framework. Through the BankCard framework, qualified material candidates are systematically elevated into standardized, upgradable material assets, establishing a clear value progression from raw data to commercial products and bridging the gap between academic discovery and industrial translation.
Entities (9)
Relation Signals (9)
Material Assets → createdvia → BankCard Framework
confidence 95% · The “BankCard” framework establishes standardized, multi-dimensional admission and classification criteria for materials entering the Materials Bank.
Materials Bank → manages → Material Assets
confidence 95% · It does not merely curate high-quality data but systematically elevates qualified candidates into standardized, upgradable materials assets
Materials Bank → buildsupon → Materials Database
confidence 92% · Built upon conventional databases, this AI-powered professional management framework centers on material assetization and industrial translation.
Materials Bank → integrates → Large Models
confidence 90% · It integrates databases, AI predictive models, autonomous experimentation and multi-criteria evaluation into a unified innovation ecosystem
Materials Bank → utilizes → Assessment Framework
confidence 90% · Standardized data generated must pass a four-tier asset qualification system spanning from scientific discovery to industrial translation before being incorporated into the Materials Bank as formal assets.
Materials Bank → integrates → Modular High-Throughput Experimental Platforms
confidence 88% · The system operates as a complete closed-loop workflow: raw database data feeds initial AI screening, prospective candidates are validated through automated experiments
Large Models → screens → Material Assets
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Driven by high-throughput experimentation, computational modeling, and artificial intelligence (AI), materials data has expanded at an unprecedented rate. Conventional materials databases function only as passive repositories, archiving raw experimental records indiscriminately including both successful and failed data, without systematic value filtering or asset management. This creates a critical gap between massive data accumulation and actionable innovation, hindering the identification of high-potential materials and industrial translation. To address this bottleneck, we propose an industrialization-oriented Materials Bank, a dedicated valuefiltering and assetization layer that operates beyond traditional databases. It does not merely curate high-quality data but systematically elevates qualified candidates into standardized, upgradable materials assets via a multi-dimensional BankCard framework covering scientific validity, synthesis feasibility, application readiness, and industrial value. By unifying databases, AI models, automated experimentation, and multi-criteria assessment into a cohesive closed-loop ecosystem, the Materials Bank establishes a clear trajectory from data to knowledge, candidate, asset, and product. It serves not as an enhanced database or screening tool, but as a decision infrastructure bridging academic discovery and industrial demand, offering a scalable paradigm to accelerate AI-driven materials innovation and deliver tangible real-world impact.
Tags
Links
- Source: https://arxiv.org/abs/2606.31366v2
- Canonical: https://arxiv.org/abs/2606.31366v2
Trouble viewing inline? Open PDF directly →
Full Text
22,803 characters extracted from source content.
Expand or collapse full text
Commentary From Materials Database to Materials Bank: Assetizing Data for AI- Driven Materials Innovation Chenyao Ma a , Di Zhang c d , Weibo Gong b , Wei Du a b , Rui Su a , Yuhang Chen a , Kan Xu a , Huan Gu a , Limin Li a b *, Piao Ma a b *, Zheng-Hao Li e *, and Hao Li c * a. Suzhou MatSource Technology Co., Ltd., Suzhou 215000, Jiangsu, China. b. Gusu Laboratory of Materials, Suzhou 215000, Jiangsu, China c. Advanced Institute for Materials Research (WPI-AIMR), Tohoku University, Sendai 980-8577, Japan d. Frontier Research Institute for Interdisciplinary Sciences (FRIS), Tohoku University, Sendai, 980-8577, Japan e. State Key Laboratory of Advanced Environmental Technology, Department of Environmental Science and Engineering, University of Science and Technology of China, 230026, China Abstract Driven by high-throughput experimentation, computational modeling, and artificial intelligence (AI), materials data has expanded at an unprecedented rate. Conventional materials databases function only as passive repositories, archiving raw experimental records indiscriminately including both successful and failed data, without systematic value filtering or asset management. This creates a critical gap between massive data accumulation and actionable innovation, hindering the identification of high-potential materials and industrial translation. To address this bottleneck, we propose an industrialization-oriented Materials Bank, a dedicated value- filtering and assetization layer that operates beyond traditional databases. It does not merely curate high-quality data but systematically elevates qualified candidates into standardized, upgradable materials assets via a multi-dimensional BankCard framework covering scientific validity, synthesis feasibility, application readiness, and industrial value. By unifying databases, AI models, automated experimentation, and multi-criteria assessment into a cohesive closed-loop ecosystem, the Materials Bank establishes a clear trajectory from data to knowledge, candidate, asset, and product. It serves not as an enhanced database or screening tool, but as a decision infrastructure bridging academic discovery and industrial demand, offering a scalable paradigm to accelerate AI-driven materials innovation and deliver tangible real-world impact. Introduction Figure 1. The closed-loop innovation ecosystem of the Materials Bank for AI-driven materials discovery. The Materials Bank acts as the core industrialization-oriented knowledge hub, underpinned by high-quality databases, domain-specific large models, modular high-throughput experimental platforms, and full-chain feedback mechanisms. This integrated ecosystem enables accelerated discovery, multi-dimensional evaluation, and iterative optimization of advanced energy materials Rapid advances in high-throughput experimentation, computation, and artificial intelligence (AI) have driven explosive growth in materials data. Traditional databases serve as essential repositories for raw data, including both successful and failed experimental records 1,2 . However, they remain confined to passive storage, lacking systematic evaluation and value-screening mechanisms. Consequently, databases cannot identify high-potential candidates or guide industrial translation, creating a persistent gap between data abundance and actionable materials innovation 3,4 . In response to this mismatch between massive raw data and industrial demand, a range of industrial material data platforms have been developed alongside the emerging concept of material assetization. Representative computational repositories such as Materials Project 5 and AFLOW 6 establish unified property screening criteria for computational bulk material candidates, while industrial informatics platforms including Granta EduPack introduce fixed multi-indicator scoring to assess commercial material performance. Nevertheless, most existing frameworks only adopt static evaluation metrics locked to preset weights and thresholds 7 . They lack full-lifecycle asset management coupled with automated experimental closed loops: for example, their scoring rules cannot shift priorities between fundamental computational screening of solid electrolytes and pilot-scale performance verification of electrocatalysts, failing to realize flexible multi-dimensional screening for diverse research scenarios. To tackle this issue, we propose an industrialized Materials Bank tailored for AI- driven materials innovation. Built upon conventional databases, this AI-powered professional management framework centers on material assetization and industrial translation. Far from being a mere combination of curated databases, AI screening tools and TRL assessments, the Materials Bank relies on three core pillars: validated datasets, industrial maturity evaluation and closed-loop feedback. It integrates databases, AI predictive models, autonomous experimentation and multi-criteria evaluation into a unified innovation ecosystem focused on full-lifecycle asset governance. The system operates as a complete closed-loop workflow: raw database data feeds initial AI screening, prospective candidates are validated through automated experiments, and evaluation results are fed back to update both databases and the candidate inventory of the Materials Bank. High-value material assets are further advanced for industrial scaling. This workflow bridges fundamental research and industrial demands, forming a complete value progression chain: data → knowledge → material candidates → material assets → end products. Data: The raw bottom-layer information collected by high-throughput experiments and computational simulations, including unfiltered successful and failed experimental records, crystal structural parameters, and primitive electrochemical signals. Stored passively in conventional materials databases without value screening or logical refinement, data serves as the fundamental raw input for the entire innovation pipeline. Knowledge: Scientific laws, structure-property relationships and screening guidelines generalized from massive raw data via domain-specific large models. Knowledge extracts interpretable correlations out of fragmented datasets, completing the first value upgrade from scattered records to reusable scientific insights, which support rapid preliminary screening of promising material candidates. Materials Assets: Standardized, iterable and upgradable material candidates that pass four-dimensional assessments (scientific validity, synthesis feasibility, application readiness, and industrial potential) and are formally managed via the Materials Bank’s BankCard framework. Unlike theoretical knowledge alone, materials assets are experimentally validated, fully indexed, and treated as manageable long-term R&D resources with clear commercialization prospects, representing the transformation from theoretical insights to operable industrial R&D assets. Products: Commercially viable materials or devices derived from high-grade materials assets after pilot-scale verification, mass-production optimization and market adaptation. While assets only embody industrialization potential at the laboratory stage, finished products realize tangible commercial value through full engineering scaling, marking the highest-value output of the whole closed-loop system. The four tiers feature an inherent progressive dependency: data yields knowledge, which screens candidates standardized into formal material assets via multi- dimensional assessment; mature assets then undergo engineering iteration to create marketable products. Every upper tier refines and elevates the value of lower layers, outlining the full value chain spanning basic academic research to industrial implementation, and this tiered framework equips the Materials Bank with a scalable paradigm to expedite AI-led materials innovation and deliver tangible practical value. Database The rapid development of high-throughput experiments, computational simulations and AI has triggered explosive growth in materials data. As core data carriers, materials databases comprehensively collect various raw data covering both successful and failed experiments. They lay a solid foundation for intelligent research and experimental verification, provide sufficient data support for energy materials research, and consolidate the underlying basis of the entire materials innovation system. 8,9 Crucially, databases are repositories where only data accumulation and archiving are performed, while no value filtering, asset definition, or transformation management is conducted. At present, such data resources have been widely applied in the fields of solid electrolytes and electrocatalysis. Specifically, they facilitate the development of solid electrolytes with high ionic conductivity and the screening of high-performance catalysts, and effectively support mechanism analysis, performance optimization and rapid screening of candidate materials in relevant research directions. Large Models Supported by massive raw data stored in materials databases, domain-specific large models enable preliminary intelligent screening and prediction of materials. These models efficiently mine potential correlations within datasets and screen out promising candidate materials, greatly boosting screening efficiency. They build an effective link between basic data and candidate materials, delivering accurate guidance for subsequent experimental verification. Researchers established data-driven intelligent research frameworks based on large language models (LLMs) tailored for solid electrolytes, accelerating the development of innovative hydride-based solid electrolytes. 10-12 Advanced LLMs also serve as powerful tools for electrocatalytic materials research, clarifying inherent relationships among catalytic activity, structural stability, electronic configuration, defect morphology and reaction intermediates from abundant experimental data. Furthermore, multimodal and vision language models are capable of processing multi-source information including texts, graphs and crystal structures, realizing comprehensive extraction and integrated analysis of research data in materials science. Modular Labs Figure 2 Architecture of the Modular Automated High-Throughput Energy Materials R&D Platform Schematic of the high-throughput integrated research platform for advanced energy materials. The core hub (High-Throughput Integrated Workstation) coordinates three parallel synthesis stations (ball milling, sintering, solvothermal), feeds materials into a unified post- processing station, and enables automated characterization via high-throughput XRD and multi- channel electrochemical testing stations, forming a closed-loop workflow for accelerated materials discovery and optimization. Modular high-throughput experimental platforms conduct standardized verification on candidate materials screened by large models. They efficiently generate experimental data to confirm the scientific rationality and synthesis feasibility of target materials, laying a solid experimental foundation for further screening. 13 One typical example is one of the self-driving labs developed by Suzhou MatSource Technology. Equipped with high-energy ball milling, high-temperature sintering and solvothermal/hydrothermal reaction units, the self-developed platform enables fully automatic material preparation. After standardized post-treatment, samples are characterized in parallel via high-throughput XRD for structural analysis and multi-channel electrochemical tests for performance evaluation. All experimental data are stored in the Materials Bank as evidence for asset grading and status tracking, enabling traceable management and in-depth analysis. Featuring flexible modular configuration, the platform adapts well to cathodes, anodes, solid electrolytes, and other energy material systems without adjusting the overall structure. Centralized intelligent control reduces manual work, shortens research cycles and minimizes human errors, enabling high-throughput parallel screening of diverse material formulas. Integrating automatic synthesis, efficient characterization and data management based on the Materials Bank, it builds a complete data-driven research loop and greatly improves research efficiency and experimental reliability. Assessment Modular automated high-throughput platforms enable parallel synthesis and preliminary characterization of candidate materials. Standardized data generated must pass a four-tier asset qualification system spanning from scientific discovery to industrial translation before being incorporated into the Materials Bank as formal assets. This system adopts progressive criteria: scientific validity, materials feasibility, application readiness, and industrial translation potential, with dynamically adjustable thresholds and weights to screen materials at each stage and avoid invalid data accumulation. As a concrete illustration, two representative energy material research scenarios demonstrate the dynamic tuning logic of thresholds and weights. For fundamental screening of novel solid electrolyte candidates at the early laboratory stage, the framework assigns a high weight to scientific validity and raises the threshold for structural stability, while reducing the proportion of industrial translation potential; the primary goal is to eliminate materials with unreasonable crystal structures and poor theoretical ionic conductivity. In contrast, for electrocatalytic catalysts undergoing pilot trial development close to industrialization, the weight of synthesis scalability, element abundance and production cost is greatly elevated, together with tighter thresholds for cycling durability and batch reproducibility, so as to sift out candidates compatible with mass manufacturing. Such differentiated parameter settings explicitly reflect the flexible adjustment strategy of the four-tier assessment system across distinct research contexts. Tiered data flows back dynamically: early-stage data optimizes large model prediction accuracy, while mature data iteratively refines evaluation criteria, forming a self-reinforcing loop. Notably, this multi-dimensional assessment is not merely an extended TRL evaluation; it serves as the core grading mechanism for material assets. 14 The Materials Bank thus functions as a tiered asset library covering the full R&D pipeline rather than a mere industrial database or assessment tool, effectively bridging lab research and translation, and improving the reliability, efficiency and translation potential of energy materials R&D. This four-dimensional scoring standard forms the essential dividing line between curated raw datasets and formal material assets archived in the Materials Bank. Curated data only undergoes basic cleaning and formatting without multi-dimensional industrial evaluation, while material assets are curated data that pass all four evaluation metrics and are managed as standardized, lifecycle-tracked R&D resources. Materials Bank Table 1. The “BankCard” Framework: Admission Criteria for the Materials Bank. The “BankCard” framework establishes standardized, multi-dimensional admission and classification criteria for materials entering the Materials Bank. It includes modules for material identity, evidence reliability, performance metrics, synthesis feasibility, technological readiness, application value, and development status, ensuring only verified, high-potential materials are archived and tracked throughout the innovation pipeline. Module Content Identity composition, structure, phase, synthesis route Evidence literature source, experiment, computation, uncertainty Performance activity, conductivity, stability, selectivity, etc. Feasibility synthesis difficulty, scalability, element abundance Readiness TRL, device compatibility, reproducibility Value cost, IP space, sustainability, market/application fit Status archived, candidate, validated, scalable, deployable Table 1 lists seven core functional modules of the BankCard framework, which jointly enable full-lifecycle evaluation of materials from basic research to industrialization. The Identity module defines unique material fundamental information; the Evidence module records data sources and their reliability uncertainty; the Performance module collects key functional indicators of energy materials; the Feasibility module assesses synthesis and scale-up constraints; the Readiness module quantifies industrial maturity based on TRL, compatibility and reproducibility; the Value module measures economic, patent and market potential; and the Status module marks the current R&D progress of each material asset. The multi-angle indicators covered by all seven modules complement each other without omission, completely matching the multi-stage screening demands of solid electrolytes, electrocatalysts and other energy materials. This modular design forms the core asset classification logic of BankCard, supporting flexible dynamic evaluation for diverse research scenarios. As the core assetization and transformation hub of AI-driven material innovation systems, the Materials Bank does not merely collect data but governs materials as upgradable assets. It sorts out scattered original data, candidate materials and experimental achievements generated in previous research stages, and converts them into standardized, reusable and iterable high-quality material assets. It formulates unified BankCard asset admission rules to standardize the warehousing process and set practical operational specifications. Only materials that pass multi-dimensional evaluation and meet all specified access indicators can be officially included to realize standardized management of material assets 15 . In addition, to safeguard asset integrity and long-term data reliability, the Materials Bank implements full traceability and multi-version archiving for all stored assets. All raw datasets, initial four-dimensional assessment records and subsequent experimental/computational logs are permanently retained without overwriting or removal. If new measurements or simulations contradict an asset’s documented properties, the platform creates a labeled new asset version and launches a secondary re-evaluation via the four-dimensional scoring framework. Conflicting outcomes are retained to retrain domain-specific large models instead of being discarded as defective data; these discrepancies act as informative training samples to improve screening precision, aligning with the insight that experimental inconsistencies facilitate iterative progress for both researchers and AI models. Discussion and Summary The proposed Materials Bank is not merely a curated database combined with AI screening and TRL assessment. Unlike conventional databases that passively store raw data, it functions as a dedicated assetization layer built on value filtering and systematic evaluation. Its core innovation lies in the BankCard framework, which standardizes multi-dimensional asset grading and lifecycle management, transforming fragmented data into reusable materials assets. By integrating data curation, AI prediction, automated experimentation, and closed-loop feedback, it bridges academic discovery and industrial demand. In the future, when industry requires materials with specific new properties, it can turn to the Materials Bank just as people seek funds from banks, enabling fast and precise matching of qualified candidates. Further optimization of evaluation criteria and AI-experimental integration will accelerate advanced energy materials development and industrialization, driving next-generation energy technologies. Acknowledgement The authors acknowledge financial support from Suzhou MatSource Technology Co., Ltd. and Gusu Laboratory of Materials under Grant No. Y2501. Reference 1 Zhuang, Y. et al. Materials Databases: Foundations of Modern Digital Materials. Precision Chemistry (2026). 2 Wang, Y., Wang, Q., Jang, S.-H., Cheng, E. J. & Li, H. Discovering new materials knowledge from “old data”. Chemical Communications 62, 9536–9549 (2026). https://doi.org/10.1039/D6C01716A 3 Zhang, D. et al. Digital materials ecosystem: from databases to AI agents for autonomous discovery. Chemical Science (2026). https://doi.org/10.1039/D5SC09229A 4 Smit, B. & Garcia, S. The data-only illusion in materials discovery. Nature Materials (2026). https://doi.org/10.1038/s41563-026-02578-7 5 Jain, A. et al. Commentary: The Materials Project: A materials genome approach to accelerating materials innovation. APL Materials 1, 011002 (2013). https://doi.org/10.1063/1.4812323 6 Curtarolo, S. et al. AFLOW: An automatic framework for high-throughput materials discovery. Computational Materials Science 58, 218–226 (2012). https://doi.org/https://doi.org/10.1016/j.commatsci.2012.02.005 7 Hegde, V. I. et al. Quantifying uncertainty in high-throughput density functional theory: A comparison of AFLOW, Materials Project, and OQMD. Physical Review Materials 7, 053805 (2023). https://doi.org/10.1103/PhysRevMaterials.7.053805 8 Horton, M. K. et al. Accelerated data-driven materials science with the Materials Project. Nature Materials 24, 1522–1532 (2025). https://doi.org/10.1038/s41563-025-02272-0 9 Shen, L. et al. Harnessing database-supported high-throughput screening for the design of stable interlayers in halide-based all-solid-state batteries. Nature Communications 16, 3687 (2025). https://doi.org/10.1038/s41467-025-58522-x 10 Zhang, D. et al. “DIVE” into hydrogen storage materials discovery with AI agents. Chemical Science 17, 3031–3042 (2026). https://doi.org/10.1039/D5SC09921H 11 Zhang, D. et al. Accelerating Catalyst Materials Discovery With Large Artificial Intelligence Models. Angewandte Chemie International Edition n/a, e26150 (2026). https://doi.org/https://doi.org/10.1002/anie.202526150 12 Wang, Q. et al. Unraveling the Complexity of Divalent Hydride Electrolytes in Solid-State Batteries via a Data-Driven Framework with Large Language Model. Angewandte Chemie International Edition 64, e202506573 (2025). https://doi.org/https://doi.org/10.1002/anie.202506573 13 Szymanski, N. J. et al. An autonomous laboratory for the accelerated synthesis of inorganic materials. Nature 624, 86–91 (2023). https://doi.org/10.1038/s41586-023-06734-w 14 Xin, H. et al. Transparent Reporting for Agentic Catalysis Enabled by Artificial Intelligence (TRACE-AI): Community Guidelines and A Publication Checklist. ChemRxiv 2026 https://doi.org/10.26434/chemrxiv.15001239/v1 15 Li, L. et al. AI as a catalyst for transforming scientific research: a perspective. AI Agent 1, 8 (2025). https://doi.org/10.20517/aiagent.2025.08