Factories Threw Away the Data AI Needs Now

Manufacturers have been disconnecting hard drives full of production data for years because their SCADA databases ran out of capacity. AI models now need exactly that data, full-resolution sensor readings and process histories, for training and real-time inference. Doug Pagnutti of Tiger Data, a former automation engineer who lived this cycle firsthand, explains how the manufacturing data infrastructure problem has flipped.

Why Has the Manufacturing Data Problem Changed?

Ten years ago, manufacturers collected large volumes of data and could not figure out how to use it. GenAI and machine learning have flipped that problem. The tools to extract value from operational data exist now, and the constraint has shifted to whether the data infrastructure can supply what those tools demand.

AI models training on historical production data need precise, full-resolution records. An hourly average of pressure tells a machine learning model almost nothing. The model needs to see how pressure fluctuated within each interval, because the fluctuation pattern is what the AI uses to recognize what will happen next. At the same time, AI inference requires “now data,” the trend from the last five seconds of operation. The AI compares real-time behavior against historical patterns and projects the likely outcome. Both requirements, high-resolution historical archives and low-latency real-time feeds, must run simultaneously on the same infrastructure.

AI Requirement What the Model Needs Why Averages Fail
Historical training data Full-resolution sensor readings over extended periods Hourly or daily averages erase the fluctuation patterns that AI uses for prediction
Real-time inference data Last five seconds of trend behavior Delayed or aggregated feeds prevent the AI from comparing current conditions against historical patterns

What Happens When a SCADA Database Runs Out of Capacity?

The database stops accepting new data within twelve to eighteen months, engineers sacrifice existing monitoring tags to make room for new ones, dashboards slow until operators outpace the screen refresh, and hard drives full of production history get disconnected and lost permanently.

Before joining Tiger Data, Doug Pagnutti lived through it. The system started with dashboards that loaded instantly. Within a year to a year and a half, the underlying database stopped accepting additional data at the rate production generated it. The team purchased a new server and upgraded hardware. The problem persisted.

Adding monitoring for a new piece of equipment meant removing tags from an existing system. Engineers had to evaluate which data streams they could sacrifice to make room for new ones, a zero-sum tradeoff where expanding visibility in one area required going blind in another. The dashboards that operators relied on slowed so badly that floor workers started racing the data refresh, competing to finish tasks before the screen updated.

Hard drives filled up and were disconnected. The data on those drives, potentially years of production history at full resolution, became permanently inaccessible. “We were throwing away some super valuable resources.” That historical data is exactly what AI models now require for training, and it is gone.

What Can AI Do With Full-Resolution Factory Data?

Three categories of AI application open up when the data infrastructure can sustain both high-resolution storage and real-time access: machine learning optimization that correlates variables engineers never connected manually, AI-assisted troubleshooting that reaches root causes faster than manual diagnosis, and energy management that shifts production schedules to match grid pricing.

A dosing system used machine learning to determine optimal parameters by correlating variables that engineers had not previously connected. Humidity levels, conditions from other processes in the same room, and ambient factors all influenced the dosing outcome. “All this data that we didn’t know was so related to the tuning.” The machine learning model identified relationships across data streams that manual analysis had missed.

AI-assisted troubleshooting changes how automation engineers diagnose equipment problems. Where diagnosis was previously “a very manual process,” AI helps engineers reach the root cause quickly. The result is “solutions that last” rather than “the band-aid” fixes, reducing downtime.

Energy management represents a third category where production scheduling adapts to grid pricing. Manufacturers in energy-intensive industries shift when they run certain processes based on electricity costs that vary throughout the day. Optimizing that scheduling requires historical consumption data at high resolution combined with real-time production monitoring, exactly the combination that conventional data architectures struggle to deliver simultaneously.

This article is based on a video interview with Doug Pagnutti, Industrial Developer Advocate at Tiger Data, and Lucian Fogoros of IIoT World, recorded ahead of IMTS 2026. Editorially Independent, Sponsored by Tiger Data. AI tools were used to help summarize and organize the content. Reviewed and edited by the IIoT World editorial team.