An AI turns 176 thin film coating papers (atomic layer deposition) into a queryable database, but its F1 (precision recall) score drops to 0.344 out of 1.0 on indium gallium zinc oxide (IGZO), the display backplane chemistry that still needs a human.
A materials scientist hunting for every published recipe to grow a specific thin-film coating used to spend an afternoon chasing citations. Atomic-layer-deposition (ALD) papers, especially for display-grade metal oxides, bury the useful parameters in paragraphs of process narrative: the substrate temperature, the precursor pulse times, and the cycle counts that decide whether a run works. A new framework, SciKGExtract, shows that an AI can now do most of that reading, and that a small but well-chosen set of controls is what lets it decide when to stop trusting itself.
SciKGExtract, posted to arXiv on 5 October 2026 by Jennifer D'Souza and declared for submission to Nature Communications Materials, is a pipeline of three checks rather than a single model. A hand-built schema of 65 experimental properties and 155 quantitative measurement slots tells the extractor what to look for. Every chemical name it pulls is normalized against PubChem, the public chemistry database, so the same precursor is not recorded under three spellings. An agent loop then scores the model's own extractions, flags the low-quality ones, and reruns them before they enter the knowledge graph, a database that stores entities and the relationships between them, designed to be queried by software.
The system was tested on 176 ALD papers, with an expert-annotated subset for the deepest schema. On zinc oxide, a single-component ALD target that is the easier case, the framework's exact-match F1 score (a 0-to-1 accuracy measure combining precision and recall) rose from 0.591 on direct normalized extraction to 0.805 once the agent refinement loop ran. On indium-gallium-zinc oxide (IGZO), a multicomponent target that is the harder case, the same system stalled at F1 0.344. The 0.805 is the number the paper leads with. The 0.344 is the number that has to travel with it.
The 0.344 is also the most informative number in the paper. IGZO is a metal-oxide semiconductor used in advanced display backplanes, and the chemistry runs as a "supercycle": alternating pulses of three different precursors in a precise order. SciKGExtract's extraction breaks down at exactly the steps where the order and timing of those pulses are the actual result the paper is reporting. The framework can tell you the substrate temperature. It cannot reliably tell you how many cycles of indium came before the first gallium pulse, and that is the recipe.
The same 65-property schema that lifts the F1 on zinc oxide is what exposes the place where the system fails on indium-gallium-zinc oxide. The depth of the schema forces the extractor to assign a numeric value to a specific slot, and the failure mode the paper itself flags, the segmentation of a multi-step process into the right numbered slots, becomes diagnostic instead of silent. A wrong temperature is a confident number. A missing supercycle step is a structural gap, and a structural gap is something a human reader can catch on the first pass.
The paper is a preprint, submitted to Nature Communications Materials but not peer-reviewed, and the 176-paper corpus is small by machine-learning standards. The expert-annotated subset that drives the 0.591-to-0.805 lift is smaller still. Anyone building on this result should wait for review, and should treat the agent loop's gain as a proof of concept rather than a production benchmark.
For a working ALD lab, that is a useful place to be. The chemist no longer has to skim 176 papers to find a recipe; the framework pulls the structured fields, the chemical names are already normalized, and the cases the extractor is not sure about are flagged. The table-of-parameters work is what the AI handles, and the supercycle-order work is what the chemist still owns: the cycle counts, the pulse-timing differences, the small choices that decide whether the run produces uniform films.
The framework's shape, a schema plus a canonical reference plus a verification loop, is starting to look like the structure of other reliable AI infrastructure. Three controls, layered together, do most of the reading work the chemist used to do by hand. The parts that still fail are exactly the parts where the human reader is cheapest: the process steps that are the point of the paper.
The next gate is peer review. If the schema and the agent loop survive it, the question is whether working ALD groups will actually load their lab's process notes into a knowledge graph on top of this pipeline, and whether publisher agreements around corpora like these will let them.