Your OT Data Isn't Ready for AI — And That's Why Pilots Fail
ERP & Business Systems

Your OT Data Isn't Ready for AI — And That's Why Pilots Fail

Most mid-market manufacturers have years of shop-floor data but it was collected for compliance, not AI model training. Here's what to audit before committing pilot budget.

7 min readJuly 5, 2026
Back to News
TL;DR
  • -Compliance-grade OT data — collected for monitoring and traceability — is structurally different from analytics-grade data required for AI model training.
  • -AI projects in industrial environments fail because OT data is incomplete, inconsistently structured, or poorly contextualized — not because of algorithmic limits.
  • -Data lineage from sensors and PLCs through MES and ERP to any analytics platform is rarely documented at mid-market manufacturers.
  • -An OT data governance audit should precede any AI vendor selection or pilot budget commitment.
  • -Cybersecurity controls must be embedded in OT data pipelines — edge protection, transit security, and access logging — not treated as a separate initiative.

The Signal

In February 2026, Schneider Electric published a post titled "OT data: The foundation behind industrial analytics and AI" that draws a hard line between two categories most manufacturers treat as interchangeable: compliance-grade OT data and analytics-grade OT data. The argument is direct. The constraint blocking industrial AI is not algorithmic capability or vendor availability. It is the OT data management foundation underneath it.

That framing comes from a vendor with a product interest in OT infrastructure, so treat the editorial positioning accordingly. The underlying structural claim holds up against operational reality: manufacturers who have been collecting sensor, PLC, MES, and WMS data for years to satisfy compliance and traceability requirements are sitting on records that were never designed to train AI models or drive continuous analytics.

The post cites a projection — attributed broadly to industrial segment forecasts, with no originating research firm named — that organizations will collect 4.4 zettabytes of OT data globally by 2030, more than double the 2023 total. Treat that figure as a vendor-cited statistic rather than independently verified data. The directional point stands regardless: volume is not the gap. Structure, context, and governance are.

Why This Matters for Mid-Market Manufacturers

There is a specific failure mode this signal describes, and it is expensive. A manufacturer budgets for a predictive maintenance or quality analytics pilot. The vendor is selected. The engagement begins. Then the data preparation work reveals that years of historian records have inconsistent tag naming across production lines, missing contextual metadata — which line, which shift, which material — sampling rates optimized for alarm thresholds rather than trend detection, and no documented ownership for who is responsible when sensor calibration drifts.

The pilot stalls. The vendor extends the engagement. The budget is consumed in data cleanup that was never scoped. In the worst cases, the model is trained on the available data anyway and produces recommendations that reflect the noise and gaps in the underlying records rather than actual equipment behavior.

The Schneider Electric post argues that AI projects in industrial environments fail or underperform not because of algorithmic limitations, but because OT data is incomplete, inconsistently structured, poorly contextualized, or difficult to access. The post also cites a widely repeated rule of thumb: approximately 80% of effort in an AI project goes to data preparation, with AI deployment positioned as "the finale, not the opening act." That figure circulates broadly across AI implementation literature; the post does not independently source it, so use it as a directional benchmark, not a precision claim.

The compliance-grade versus analytics-grade distinction is the most operationally useful concept in the post. Compliance-grade data was designed to prove that a process ran within spec or that a batch met traceability requirements. Analytics-grade data must support continuous model training, anomaly detection, and optimization — which requires different sampling rates, richer contextual fields, consistent schema across sites and lines, and enough labeled historical data to establish normal and abnormal patterns.

Where the Exposure Shows Up

The gap between what manufacturers have and what AI needs tends to cluster in predictable places:

Sensor and PLC records captured at intervals optimized for alarm detection, not for the sub-second or rolling-average patterns that predictive models require. Changing the sampling rate retroactively does not fix historical data.

MES and WMS records that exist in separate systems with no consistent cross-reference to the OT layer. A quality event in the MES may have no machine-level timestamp that aligns it to the corresponding PLC alarm in the historian.

ERP production and inventory records that represent business-layer events — work orders opened and closed, materials consumed — but carry no operational context about what actually happened on the line during that interval.

Manual workarounds — spreadsheets, emailed reports, printed-and-scanned shift logs — that feed into production decisions and sometimes into BI dashboards but are invisible to any AI pipeline and unauditable as training data.

The Schneider Electric post notes that OT data is continuous, heterogeneous, and often produced by assets that were never designed to broadly share information. That distinguishes it structurally from the transactional IT records that ERP systems manage. The integration path from the OT layer through the IT layer is rarely documented, and the contextual translation required at each handoff is rarely automated.

The post makes a pointed observation: organizations do not actually "connect data" — they connect assets. Only once assets are properly connected can data be retrieved, contextualized, and reused across use cases. For mid-market manufacturers, the instinct is often to solve the problem at the analytics layer — buying a BI platform or an AI tool — when the real work is upstream at the asset and integration level.

What to Audit Before Any AI Budget Commitment

Before committing any budget to an AI vendor, predictive analytics platform, or consulting engagement, operations and IT leaders should work through five specific areas:

OT data collection points: Map every active data source — sensors, PLCs, MES systems, WMS — and document the current schema, sampling rate, and context metadata for each. Identify whether those specifications meet analytics-grade requirements or compliance-grade requirements. They are not the same standard.

Data lineage and ownership: Trace the path each OT data stream takes from collection point through any intermediate system (historian, edge gateway, MES) into ERP and then into any analytics or BI platform. Identify every manual handoff, every spreadsheet transfer, and every point where lineage breaks. Document which team or role is accountable for each stream's accuracy — sensor calibration, MES record integrity, WMS inventory sync. Gaps in ownership are gaps in data quality.

Data quality controls upstream of AI inputs: Identify which OT data undergoes validation, deduplication, or cleaning before it reaches any system an AI model or dashboard will consume. A model trained on unvalidated records reflects the errors in those records. This is a pipeline problem, not a model problem, and it must be solved before model selection.

Integration paths: Document every integration connecting OT systems to analytics and AI platforms. Confirm whether each path is versioned and maintained, or whether it is a one-time export someone built to meet a deadline. Undocumented integrations break silently.

Cybersecurity controls in OT data pipelines: OT data now flows across edge devices, cloud connectors, and analytics platforms that were not built with open network access in mind. Audit edge protection, transit security, and access logging for every path that carries OT data off the shop floor. The Schneider Electric post treats cybersecurity as a foundational requirement embedded in OT data collection and movement — not a parallel initiative to address after deployment.

What to Watch

The more consequential risk is not that manufacturers will ignore this signal. It is that they will act in the wrong order. AI vendor selection is a visible procurement exercise with demos, timelines, and executive sponsorship. Data governance is not. Organizational pressure to show AI progress pushes teams toward pilot contracts before infrastructure questions are answered.

The Schneider Electric post warns that ad hoc, use-case-by-use-case approaches to OT data preparation do not scale as digital ambitions grow from local analytics to enterprise-wide optimization. A manufacturer that cleans up data for one predictive maintenance pilot without establishing a repeatable governance structure will face the same preparation cost again for every subsequent use case.

Watch for AI vendor contracts that scope data preparation as a separate, variable-cost line item. That structure transfers the data readiness risk to the manufacturer after the contract is signed.

Bottom Line

The sequencing decision is the decision. Selecting an AI platform or predictive analytics vendor before auditing OT data readiness does not accelerate the path to working AI — it delays it, at greater cost. The audit described above is not a technology project. It is a documentation and ownership exercise that operations leaders and IT teams can conduct without purchasing anything. The output — a clear map of what data exists, what shape it is in, who owns it, and how it flows — is the prerequisite for every AI investment that follows.

Sources and supporting resources
Previous
Predictive Maintenance AI Readiness: The Data Pipeline Decision Manufacturers Face Now
Next
AWS AI Customer-Service Agents: The Data Governance Audit Manufacturers Need Before Deployment

Get Business Technology Updates

News, insights, and practical guidance across ERP, Cloud, Data, AI, digital transformation, and technology projects.

No spam. Unsubscribe anytime.