Machine learning (ML) in materials science refers to computational methods that learn statistical, structural, or physics-informed relationships from experimental, computational, and literature-derived materials data. These methods are used to predict materials properties, identify structure–property and process–structure–property–performance relationships, discover candidate materials, optimize synthesis and processing routes, and guide functional applications. ML is narrower than artificial intelligence (AI), which also includes broader reasoning, planning, search, and automation capabilities. It is also distinct from materials informatics, which is the wider data-centered framework that includes databases, descriptors, metadata, workflows, visualization, and knowledge management. ML can complement high-throughput computation by building surrogate models from density functional theory, finite-element simulation, molecular dynamics, or experimental data, but it is not identical to high-throughput screening itself. Unlike conventional physics-based modeling, which begins with explicit governing equations or mechanistic assumptions, ML usually infers predictive relationships from data; modern approaches increasingly combine both perspectives through physics-informed features, uncertainty quantification, and human expertise.
1.1. From Empirical Materials Development to Data-Driven Materials Science
Materials development has historically combined empirical observation, thermodynamics, processing experience, and mechanistic modeling. Data-driven materials science extends this tradition by using curated experimental and computational data to infer patterns that are difficult to capture by trial-and-error experimentation alone. Materials informatics and big-data approaches have been described as part of a broader fourth-paradigm mode of scientific inquiry, in which data, models, and domain knowledge are used together to accelerate materials discovery
[1,2][1][2].
1.2. Emergence of Materials Informatics
Materials informatics emerged from the need to organize, connect, and reuse increasingly heterogeneous materials datasets generated by computation, experimentation, characterization, and literature-derived sources. Its development was stimulated by high-throughput computation, combinatorial experimentation, and the need to connect composition, process, structure, property, and performance information in reusable ways
[3,4][3][4].
1.3. Growth of Machine Learning in Materials Research
ML became prominent in materials science as public databases, automated calculations, imaging pipelines, and laboratory records increased the availability of machine-readable data. Solid-state materials studies have used ML for property prediction, descriptor construction, stability screening, and ranking of candidate compounds
[5,6][5][6]. In continuum materials mechanics, ML and data mining have also been adopted for constitutive behavior, damage, fatigue, and microstructure–property modeling
[7].
1.4. Scope of the Entry
This entry defines core concepts, summarizes data infrastructure and algorithm families, and describes representative applications in functional and structural materials. It is not a systematic literature review and does not rank publications bibliometrically. Its emphasis is on established terminology, technical principles, common workflows, validation practices, limitations, and future directions relevant to ML-enabled materials discovery and functional applications.
The purpose of this entry is to provide a concise but analytical overview of how ML is used in materials science, with attention to both its opportunities and its methodological constraints. The intended audience includes materials scientists, engineers, graduate students, computational researchers, and interdisciplinary readers who need a structured introduction to data-driven materials research. Relative to specialized reviews focused on a single algorithm family, database, or materials class, the distinctive contribution of this entry is its integrated treatment of definitions, materials data infrastructure, ML method families, validation and reproducibility issues, representative functional applications, industrial deployment considerations, and future perspectives within one encyclopedic framework.