Data quality used to be something data scientists handled on their own. The model needed clean inputs, they cleaned them, and the organization moved on. AI agents changed that. They consume data at machine speed, with no human in the loop to fix what’s broken, and 37% of manufacturers polled at a December 2025 IIoT World session named siloed OT data as their top barrier to trusted AI. Cognitive data readiness, a framework panelists from Michelin, ATELIC, Körber Pharma, and 4D1 described, goes beyond clean or contextualized data to encode cause and effect so AI agents can reason about decisions.
What Is Cognitive Data Readiness?
Cognitive data readiness sits above two layers most organizations are already working on. Physical data from sensors and machines is the first. Contextual data that bridges IT and OT silos is the second. The cognitive layer adds governance rules, compliance requirements, and operational tradeoffs so AI agents can make autonomous decisions. Legacy systems with inconsistent data formats ranked second in the same poll at 30%, and insufficient governance policies came third at 22%. All three barriers feed each other: silos create format problems, and both resist governance.
Passing a data quality scorecard of completeness, accuracy, consistency, and visibility is necessary but not sufficient. The data also needs stability: low drift between training and production samples and good coverage of edge cases. If training data is biased or incomplete, the model will fail the moment it hits conditions outside its training set. And the features used during training must be reproducible in production. When offline and online pipelines diverge, model behavior becomes unpredictable. Data readiness is the difference between a model that works in a demo and one that works on a real production line.
How Should Manufacturers Govern AI Agents?
Governance has a reputation problem because it relies on committees, manual approvals, and PDF files, all of which operate at human speed. AI agents that react in milliseconds will outrun it.
Minimum viable governance is one answer: five to seven metrics instead of forty. Automated checks for drift, validation, and anomalies run inside the data pipeline, so governance moves at the same speed as the data. About 80% of the process can be automated. The remaining 20% stays with humans for exceptions, ethical decisions, and situations that require judgment.
For autonomous AI agents, governance itself needs to become code. Policy-as-code embedded in the agentic platform replaces PDF files and Excel spreadsheets. Continuous decision and risk analysis runs alongside the agents rather than as a periodic review. The governance layer can use AI agents to monitor and control other AI agents, similar to a security operations center but at machine speed.
Where Should Data Readiness Start?
The first step is honest: look at existing master data and data flows. Most manufacturing organizations have spent 10 to 20 years optimizing processes individually, creating pockets of harmonization but not a connected whole. Comparing that reality against the end goal, whether a digital shadow of a single process or a full digital twin, shows where the work needs to happen.
One counter-intuitive approach is to start with processes still in development rather than the most mature ones. New processes introduce variation, and that variation generates the characterization data that builds real understanding of how a process behaves. It also gives AI models the kind of data diversity they need for training.
On the factory floor, reconciling logical MES data with physical reality is where a stronger foundation comes from. Real-time location systems with millimeter-level accuracy capture how work actually gets done, giving AI a ground truth for every task, motion, and resource.
Operators should not carry the burden of data enrichment. If a solution requires closing one system and opening another to add information, adoption will stall. The tools need to work where operators already work, adding value in the moment rather than creating extra steps.
FAQ
1. What is cognitive data readiness for manufacturing AI?
Cognitive data readiness goes beyond clean or contextualized data. It means data that encodes cause and effect, governance and compliance rules, and operational tradeoffs so AI agents can reason about decisions rather than find correlations. At a December 2025 IIoT World panel, an adviser described three readiness layers: physical data from sensors and machines, contextual data that bridges IT and OT silos, and cognitive data that supports autonomous AI agent reasoning.
2. How should manufacturers govern autonomous AI agents?
Traditional governance based on committees and manual approvals operates at human speed and cannot keep pace with AI agents making decisions in milliseconds. Panelists at a December 2025 IIoT World panel recommended minimum viable governance with five to seven automated checks embedded in data pipelines, where 80% of governance runs automatically and the remaining 20% is reserved for exceptions and ethical decisions. Policy-as-code replaces PDF-based compliance for agentic platforms.
3. Why does data integrity cause clinical trials to fail?
A Körber Pharma adviser cited an FDA statement that the majority of clinical trials fail because of data integrity problems. In pharma and biotech, poor data quality affects patient outcomes, delays trials for diseases without current cures, and creates significant financial costs. The same principle applies to manufacturing: AI models trained on flawed data produce flawed outputs regardless of how sophisticated the algorithm is.
4. Where should manufacturers start improving data readiness for AI?
Start with a clear assessment of existing master data and data flows, then compare that baseline against the end goal. Rather than starting with the most mature process, one approach is to begin with processes still in development, because the variation they introduce generates characterization data that improves process understanding. Reconciling logical MES data with physical spatial data from real-time location systems gives AI a verified ground truth about what is actually happening on the factory floor.
Related from IIoT World
- How UNS Prepares Manufacturing Data for AI
- Build a Factory Digital Twin in 14 Weeks
- Agentic AI in Manufacturing: ROI vs. Reality
Sources:
- IIoT World Manufacturing & Supply Chain Day, December 2025, panel: “Garbage In, Garbage Out: Fixing Industrial Data Before It Hits Your AI Models.” Session sponsored by ATELIC.
This article is based on a panel discussion, “Garbage In, Garbage Out: Fixing Industrial Data Before It Hits Your AI Models,” with Romina Guevara, formerly Chief Product and Digital Officer at Michelin; Dr. Omid Givehchi of ATELIC; Douglas Langen, CEO of 4D1; and Dr. Judith Koliwer, Senior Industry Advisor at Körber Pharma; moderated by Sebastian Trolli of Frost & Sullivan at IIoT World Manufacturing & Supply Chain Day, December 2025. AI tools were used to help summarize and organize the content. Reviewed and edited by the IIoT World editorial team.