Skip to content

Why Factory Data Never Reaches AI

5 min read Sponsored by HiveMQ, Cybus, and Barbara | Editorially Independent

manufacturing data contextualization

Manufacturers have spent ten to twenty years installing sensors and connecting machines. The data flows, but it rarely reaches AI. At the Industrial AI Summit 2026, Jonathan Alexander of Albemarle Corporation, Peter Sorowka of Cybus, and Kudzai Manditereza of HiveMQ explained where the data gets stuck: in historian-centric architectures that store without context, in application-centric integrations that cannot scale, and in a missing contextualization layer where raw readings would acquire the process knowledge that AI needs.

Where Are Manufacturers Stuck in the Data Value Chain?

Most manufacturers sit between two solved stages and one missing stage. They have mastered collection, where sensor feeds and machine outputs flow into the facility’s systems, and historization, where that data gets forwarded into historians for long-term storage. The stage they have not reached is contextualization, where raw readings acquire the process knowledge, equipment relationships, and operating conditions that make the data usable for AI. That contextual knowledge still lives inside the heads of experienced operators and engineers rather than inside any system.

Value Chain Stage Current State What Blocks AI
Collection Mature: sensors, PLCs, and connectivity installed over 10-20 years Largely solved
Historization Widespread: data forwarded to historians across plants Data stored without structure or governance
Contextualization Missing: process knowledge trapped in people Biggest gap preventing AI from scaling
Intelligence Emerging: AI models available but lack domain-specific data Blocked by upstream gaps

 

Historian-centric architectures make this worse because they treat data as something to archive rather than something to use. When manufacturers forward sensor data into a historian, the storage system captures values and timestamps without preserving relationships, operating context, or access governance. AI needs more than time-series records to make decisions, and historians were never designed to serve that function.

Why Does Application-Centric Data Collection Block AI Scaling?

The prevailing approach in most manufacturing environments connects each application directly to its data sources through dedicated point-to-point integrations. This application-centric model works for a single use case. Every subsequent project requires building a new set of connections from scratch, creating a tangle of redundant integrations that grows more fragile with each addition.

A manufacturer running one predictive maintenance project might connect a vibration sensor feed to one analytics application through a dedicated pipeline. When a second project needs energy optimization data from the same production line, the team builds an entirely new connection, and a third project for quality monitoring adds another on top of that. Each integration carries its own maintenance burden, and the resulting spaghetti architecture becomes a primary cause of the lack of scalability that prevents organizations from moving beyond isolated AI pilots.

The alternative is an infrastructure-centric approach where a shared data platform connects once to all operational sources and makes that data available to any application through standardized interfaces. New AI use cases draw from existing connections rather than justifying and building their own, which reduces both the time and cost of each additional deployment. Organizations that adopt this model can add use cases incrementally instead of treating every new AI project as a greenfield infrastructure effort.

Do LLMs Need Clean Data to Work in Manufacturing?

The panel split on whether public LLMs can handle manufacturing data. Each panelist identified a different barrier, and the answer shapes how manufacturers should approach AI deployment.

Jonathan Alexander of Albemarle Corporation pointed to a fundamental gap in training data. Manufacturing process data, engineering drawings, proprietary process parameters, and equipment-specific PIDs simply do not exist inside the training sets of frontier AI models. Companies like Anthropic and OpenAI built their models on publicly available information, and decades of operational knowledge from individual plants were never made public. Jonathan tested this firsthand: a general-purpose AI can troubleshoot a home HVAC system because residential equipment documentation is widely available online, but the same model failed when asked to interpret plant-specific process instrumentation because that data has never left the facility’s internal systems.

Peter Sorowka of Cybus challenged that position. Traditional machine learning required ninety percent of the work to go into data labeling and data cleansing, but that bottleneck does not apply to LLMs. LLMs handle messy, inconsistent data far more effectively and do not need private factory data trained into the model to be useful. Manufacturers can supply context at inference time, even through a badly scanned PDF or a photograph, and the model will produce useful results. Waiting for perfectly structured data is, in Peter Sorowka’s view, not a valid reason to delay AI in manufacturing.

Kudzai Manditereza of HiveMQ talked, in this context, about meaning. A knowledge graph connects work orders, equipment relationships, and downstream impacts in ways that a language model processing isolated documents cannot replicate. An LLM can structure and answer a homework problem, but it cannot tell you which downstream processes break when a specific work order fails, because that semantic context lives in the relationships between systems, not in any single document.

All three agree that manufacturers need to expose their operational data to AI systems with proper data infrastructure and governance. They disagree on whether LLMs can work with that data as-is or whether it first needs the structured meaning that knowledge graphs provide.

What Governance Does AI Need on the Plant Floor?

Connecting data to AI systems without a governance framework creates a different category of risk than leaving the data disconnected. AI systems operating on plant-floor data need an identity within the organization’s security architecture, with access control rules that define which data they can read, which systems they can influence, and which decisions they can trigger. Audit trails must record what the AI accessed, what it recommended, and what actions followed, for both compliance and troubleshooting when outputs go wrong.

The governance requirement extends beyond traditional data quality frameworks because AI introduces agency into data systems. A historian stores values passively, while an AI system interprets data and may initiate actions. The organization needs to govern behavior, permissions, and accountability alongside data format and storage. Manufacturers implementing governance as part of their data infrastructure build avoid the scenario where an AI system reaches production readiness and then stalls because security, compliance, and access control have not kept pace.

Frequently Asked Questions

1. Why is contextualization the biggest gap in manufacturing data for AI?

Manufacturers have spent ten to twenty years installing sensors and connecting machines, so data collection is largely solved. The missing layer is contextualization, where raw sensor readings acquire the process knowledge, relationships, and operating context that AI needs to make decisions. That knowledge still lives inside experienced operators rather than inside systems, which blocks AI from scaling across facilities.

2. What is the difference between application-centric and infrastructure-centric data collection?

Application-centric collection connects each application directly to its data sources through dedicated integrations, forcing every new AI project to rebuild connectivity from scratch. Infrastructure-centric collection uses a shared data platform that connects once to all operational sources and makes data available to any application through standardized interfaces, so new projects can deploy without rebuilding connections.

3. Can LLMs work with messy manufacturing data?

Large language models handle inconsistent and unstructured data far better than traditional machine learning, which required extensive labeling and cleansing. The bigger limitation is access: manufacturing process data, engineering drawings, and equipment-specific parameters were never included in any foundation model’s training set. Organizations need to expose their operational data to AI systems, even when that data is imperfect.

4. What governance framework does industrial AI need before scaling?

AI systems operating on plant-floor data require their own identity within the organization’s security architecture, with access control rules defining which data they can read and which systems they can influence. Audit trails recording what the AI accessed, recommended, and triggered are necessary for both compliance and troubleshooting, governing behavior and permissions rather than just data format.

This article is based on a panel discussion with Kudzai Manditereza of HiveMQ, Peter Sorowka of Cybus, David Purón of Barbara, and Jonathan Alexander of Albemarle Corporation, moderated by Hamish Mackenzie of IIoT World, recorded at the Industrial AI Summit 2026 hosted by IIoT World. Editorially Independent, Sponsored by HiveMQ, Cybus, and Barbara. AI tools were used to help summarize and organize the content. Reviewed and edited by the IIoT World editorial team.

Related from IIoT World

Share with your network