The Data Gap: Why Materials Science’s Machine Learning Era Hinges on Better Mechanical Testing

Published on 

July 14, 2026

Why metallurgy has trailed other sciences in the shift to data-driven discovery, and why the pace of change from here on will depend on what experimental testing can deliver.

Materials science is entering its big-data era. Machine learning, high-throughput simulation, and large computational datasets are reshaping how new materials are discovered, screened, and developed. Initiatives such as the Materials Genome Initiative [1], and companies including Altrove [2] and Kebotix [3], are demonstrating what becomes possible when computational tools are combined with fast, systematic experimental validation. In scientific fields such as genomics and astrophysics, this shift has already produced fundamental advances. Up to this point, materials science has made substantially less progress.

The obstacle is more specific than a shortage of interest or capability. Machine learning has been applied to alloy design across virtually every major class of structural metal, and simulation platforms of remarkable sophistication are in routine use. What has been harder to solve is a challenge that sets metallurgy apart from many other disciplines: the difficulty of generating enough experimental data to make those models reliable. Alloy performance is not controlled by composition alone. Processing, microstructure, defects, and mechanical behaviour all matter, and each is expensive to characterise at scale.

The central challenge for the field has shifted. Where the question was once whether machine learning could be applied to metallurgy, it is now whether the experimental data required to make those models reliable can be generated at the speed and scale that a data-driven approach requires.

The rise of materials informatics

Some of the most striking progress in computational materials science has come from open, large-scale databases. Platforms such as the Materials Project [4], AFLOW [5], OQMD [6], and NOMAD [7] now contain hundreds of thousands of first-principles calculations covering crystal structures, formation energies, phase stability, and other properties that can be predicted consistently from computational methods. These resources have transformed screening in areas including catalysis, energy storage, and functional materials, where the property of interest is calculable from atomic-scale structure alone.

For structural metallurgy, the picture is more complicated. Machine learning is being actively applied to alloy design in systems ranging from advanced steels and aluminium alloys to nickel superalloys, titanium alloys, and high-entropy alloys. High-entropy alloys are a useful illustration. The compositional space is so large that traditional trial-and-error development is impractical, and computational screening is the only realistic way to identify optimal candidates. However, engineering performance in these materials, as in most structural alloys, is not determined by composition alone. Factors such as processing route, thermal history, microstructure, and mechanical response all shape whether a candidate that looks attractive on the screen will perform in service. Computational predictions still need experimental validation, particularly where strength, fatigue life, and toughness are the properties that matter.

Machine learning has not caught up to increasing computational capacity.

Case in point: additive manufacturing

If any application area crystallises why metallurgy needs a better data infrastructure, it is additive manufacturing. Process development for AM involves many interacting variables. Laser power, scan speed, hatch spacing, layer thickness, powder condition, build orientation, scan strategy, and post-processing route each influence the microstructure and properties of the final part, often in ways that are difficult to predict from first principles. The relationships between them shift with alloy system, machine, and application.

Efforts such as the NIST AM-Bench series [8] have begun to establish benchmark datasets against which AM simulations can be tested and refined. Machine learning is increasingly being applied to derive process maps, predict defects, optimise print parameters, and support the digital-twin approaches that regulated industries are beginning to adopt for AM qualification. Each of these applications depends on the same underlying requirement: high-quality, traceable, structured experimental data collected across the space the model is trying to explain. Without it, the model reflects the gaps in the training set rather than the physics of the process

Simulation is only as good as the data behind it

The same principle holds across the wider simulation toolkit that metallurgists rely on. CALPHAD [9], phase-field modelling, crystal plasticity, finite-element analysis, and AM melt-pool simulation are each becoming more capable year on year. Commercial platforms such as Thermo-Calc [10] and Pandat [11], alongside open-source alternatives including OpenCalphad [12], are widely used. Integrated Computational Materials Engineering (ICME) [13] workflows now let engineers connect composition, processing, and properties in ways that would have been out of reach even a decade ago. Machine learning is also being used to build surrogate models that run orders of magnitude faster than the underlying physics-based simulations they approximate.

Each of these tools requires calibration and validation against real measurements. For structural metals, this means reliable mechanical property data, not just chemistry, phase predictions, or micrographs. The familiar principle of “good data in, good results out” is especially binding here. Sparse, biased, or inconsistent mechanical property datasets constrain the accuracy of even the most sophisticated model, and the field is now generating computational predictions faster than experimental measurements can validate them.

The closed-loop vision

The most ambitious vision for data-driven metallurgy is a closed-loop workflow:

  • Machine learning proposes the next composition, or set, of process parameters.
  • The material is produced.
  • Its properties are measured.
  • The data is fed back into the model, which updates its predictions and proposes the next experiment.

Composition-gradient samples and additively manufactured parameter matrices allow many conditions to be screened within a smaller number of physical samples than would previously have been possible, and the same logic applies as much to alloy screening and heat-treatment optimisation as it does to AM process development.

In this loop, the bottleneck is increasingly the rate at which useful experimental data can be generated, rather than modelling. For applications where mechanical properties matter, and structural metallurgy is chiefly interested in mechanical properties, the testing method chosen sets the ceiling on how quickly the loop can turn.

Automated hardness and nanoindentation mapping are already being used in some laboratories to generate dense mechanical-property maps across microstructures and composition gradients. The direction of travel is clear: faster measurements, larger datasets, and less manual intervention. What is less clear is how to get transferable, engineering-grade mechanical property data at the speed the loop requires.

In a closed-loop workflow, the bottleneck is increasingly the rate at which useful experimental data can be generated, rather than modelling.

The tensile & hardness trade-off

For most of the last century, the mechanical property data that structural metallurgy has acted on has come from tensile testing. Yield strength, ultimate tensile strength, elongation, and work-hardening behaviour are all obtained cleanly, and the method is defensible enough to underwrite design and qualification decisions across regulated industries. The cost of that defensibility is that tensile testing is slow, requires machined specimens, and consumes significant material. Scaling it across an AM parameter study or a heat-treatment optimisation trial is prohibitive at the volumes those studies increasingly require.

Hardness testing is easier to automate and in theory can be applied quickly across many samples. But hardness measures the response of the material to a specific indenter at a specific load, and its link to engineering strength properties relies on empirical correlations that break down across different alloys, heat treatments, and microstructures. As our previous newsletter explored in more detail, the same material tested per the same standard can return readings that differ by 20 to 30% depending on the load applied. For data-driven work that depends on consistent, transferable inputs, hardness has significant limitations.

Both tensile testing and hardness testing have known limitations which are only exacerbated when applied to high-throughput, high-volume workflows that depend on defensible data.

What data-driven metallurgy needs

The method that data-driven metallurgy needs combines the best of both of these test methods. It should be faster and more scalable than tensile testing, but more mechanically informative than hardness. It should return the same kind of stress-strain data that has anchored design decisions for a century, and it should do so from samples small enough to fit within alloy development coupons, AM witness cubes, small heat-treatment trials, and local regions of interest within a component. And for machine learning and simulation workflows, testing quickly is only part of the value. The rest lies in generating consistent, structured mechanical property data across composition, process route, heat treatment, and microstructure.

Profilometry-based Indentation Plastometry (PIP testing) sits in that space. PIP uses indentation to extract mechanical property information from small material volumes, returning yield strength, ultimate tensile strength, and stress-strain behaviour rather than a hardness value. It requires less material and simpler preparation than tensile testing, which makes it more scalable across the kind of parameter studies and screening campaigns that data-driven metallurgy relies on. In an AM process-parameter study, PIP can compare many build conditions without a full tensile specimen for every parameter set. In alloy screening, it can generate early mechanical property data before committing to larger-scale manufacture and full qualification testing.

As datasets grow, manual testing workflows become the limiting factor. Automation is essential for high-throughput mechanical testing at scale, including automated positioning, repeatable test execution, and structured data capture compatible with modern informatics pipelines. Plastometrex’s PLX-AutoStage has been developed with this in mind, supporting automated mechanical testing across samples, builds, or material libraries

The stakes

Big data and machine learning will play an increasingly important role in metallurgy over the coming decade. The transformation, when it comes, will come from combining advanced modelling with high-quality, high-throughput experimental data. Algorithms are one half of that equation. For structural applications, the other half is mechanical property data generated at a speed and scale that the current tools have struggled to reach.

For teams building models that need to say anything meaningful about strength, ductility, or stress-strain behaviour, the bottleneck now sits on the experimental side. Closing that gap is what will decide whether metallurgy follows other sciences into an optimised, data-driven era, or whether it stays a step behind.

Learn more about PIP Testing and subscribe to the Material Reality Newsletter to get new issues straight to your inbox.

References

1. Materials Genome Initiative (US National Programme). https://www.mgi.gov/

2. Altrove — AI-driven materials discovery. https://www.altrove.ai/

3. Kebotix — Self-driving lab for materials discovery. https://www.kebotix.com/

4. The Materials Project. https://materialsproject.org/

5. AFLOW — Automatic FLOW for materials discovery. https://aflow.org/

6. OQMD — Open Quantum Materials Database. https://oqmd.org/

7. NOMAD Laboratory — Novel Materials Discovery. https://nomad-lab.eu/

8. NIST Additive Manufacturing Benchmark Test Series (AM-Bench). https://www.nist.gov/ambench

9. CALPHAD (community journal and portal). https://calphad.org/

10. Thermo-Calc Software. https://thermocalc.com/

11. Pandat Software (CompuTherm). https://computherm.com/software

12. OpenCalphad — Open-source thermodynamic software. https://opencalphad.org/

13. ICME — Integrated Computational Materials Engineering (TMS). https://www.tms.org/icme