Blog Unlocking OT Data for AI & Analytics: Insight’s Data Rescue

engineer inspecting with a tablet

Key takeaways

  • You don't have a data problem. You have an access problem. OT data sits trapped in historians, fleet systems, and vendor controllers that were never built to talk to each other.
  • Blindness has a price: unplanned failures, lost production, and safety and regulatory exposure you could have seen coming.
  • Predictive maintenance, autonomous haulage, digital twins — every AI use case worth having sits on the same foundation of unified, trusted, accessible OT data.
  • OT Data Rescue moves that data into Microsoft Fabric, AWS, or Databricks, staged so your operational network is never directly exposed to the cloud.
  • Start small. Where the operations team is engaged, first workloads run in four to six weeks — and one proven use case funds the next.

One haul truck. Hundreds of sensors. Add the drills, conveyors, and processing units, and a single site generates more data in a shift than any team could ever review. OT data analytics is how you stop losing it, and it's closer, easier, and lower risk than most operators assume.

We called our recent webinar “OT Data Rescue” because that's exactly what the work is: recovering stranded operational technology data and turning it into decisions, efficiency, and AI use cases. Watch the full session on demand.

engineer working on 3d model

What industrial data silos cost you

Most operations run partially blind, and it isn't for lack of data. The fixed plant runs on a historian — AVEVA PI, Rockwell, or something similar. The mobile fleet reports somewhere else entirely. Drills, conveyors, and processing units each sit behind their own vendor's controllers. Every system holds a piece of the truth. None of them share it.

So the haul truck engine starts trending toward failure, and the warning signs sit in the onboard telematics while the maintenance team works from a spreadsheet. The truck dies mid-shift. A planned repair becomes lost production.

That's the pattern, and it compounds. Margin walks out the door because you can't see where you're losing it. Risk builds because you're reacting instead of anticipating. And the competitors who've already unified their operational data move while you're still assembling the picture.

Unified OT data changes what's possible

Unify the data and the same foundation pays off differently depending on where you sit: predictive maintenance and autonomous operations in mining; digital twins, computer vision quality control, and dynamic line balancing in manufacturing; throughput optimization and consolidated remote operations centers in oil, gas, and energy.

None of this is a science project. It's running in production today. And notice what all of it has in common: clean, accessible, real-time OT data. The AI part is the straightforward part. The data foundation is what's missing. The leaders didn't start with AI, they started by getting their data out of the silos.

In mining, six opportunities are drawing attention right now. Every one of them needs that foundation:

  1. Predictive maintenance on haul trucks, SAG mills, crushers, and conveyors, using vibration, temperature, and oil analysis
  2. Autonomous operations across trucks, drills, and even rail, with safer sensor-driven control
  3. Grade and recovery optimization through ore-aware process control, including computer vision that reads flotation froth
  4. Energy and decarbonization, turning emissions data into ESG reporting and fuel savings as fleets shift to battery-electric
  5. Safety and tailings, a board-level issue since the Brumadinho disaster, pairing piezometers with satellite radar and InSAR
  6. Mine-to-mill digital twins, letting you simulate pit to port and test a plan before you commit to it

Pick one or two opportunities where the payback is clearest. That's your beachhead. You don't need a full digital transformation to start.

Opportunity in action

New EU water reporting rules mean a miner that can't report may lose their permit to work. Install the sensors, pipe the data out of the historian, and you've satisfied the requirement. Then use that same data in a digital twin to optimize the heap leach process. In one opportunity we're exploring with a gold miner, that points to higher output, lower water use, better safety, and a payback measured in months. See how Insight approaches business analytics and data modernization.

engineer using tablet

How OT Data Rescue gets the data out

We've worked with mining, manufacturing, oil and gas, and utility clients long enough to know what was missing: something secure, scalable, and multicloud. That's OT Data Rescue — capture at the field site, pipelines into an analytical platform in a hyperscaler cloud, integration with the systems you already run, and an MLOps foundation so AI value gets built repeatably instead of one project at a time.

Capture at the edge, safely.

We use Cogent DataHub for field-side acquisition, though we've also delivered this with Ignition from Inductive Automation, with AVEVA directly, and with custom software written at the edge.

So why Cogent DataHub? It's multi-source and multi-cloud, so it fits wherever your data starts and wherever you want it to land. It stages that data through ISA-95 or Purdue boundaries — out of the OT network, through the DMZ, into IT, then the cloud — with store-and-forward, queuing, and batching so nothing drops on the way. It integrates natively with the major clouds and other downstream consumers. And it's sold under a perpetual license: You own it, rather than renting it forever.

Microsoft Fabric, real time and historical

In Fabric, DataHub connects natively to Eventstream and Event Hub, and data lands in Eventhouse, the real-time time series database. Query it live, put it on real-time dashboards within seconds of the source event, and let thresholds fire alerts by SMS, email, or Teams out of the box. The same data then flows into a lakehouse and combines with master data from your ERP for aggregated reporting — one favorite in flight is a shift report that leaderboards throughput and quality shift over shift.

AWS and Databricks

On AWS, the same output pushes into Kinesis Data Streams, through Firehose into S3, then Redshift and Athena for modeling — feeding AI outputs or visualizations in Amazon Quick Suite. And because the capture layer streams through Kafka or structured streaming, we've built the same pattern into Databricks. One cloud or both, the pattern holds.

Operators decide whether this works.

A portfolio of hundreds or thousands of tags means nothing to us out of the box. We need the operators and site supervisors who own those devices to tell us what the tags represent and which ones matter. Then, once we've captured the data, we show it back and let them tell us how well we're reading it.

That loop is what makes the numbers trustworthy enough to send up to corporate. Site superintendents know the numbers and they know the math. These projects succeed as a joint effort or they don't succeed.

Start small, scale from the beachhead.

It starts with an executive briefing and an operational co-design session. We bring starting-point dashboards to your business and operating teams, and ask whether this is what you're after. The answer is almost always no — which is the point, because it tells us what you actually want.

From there, supported by AWS or Microsoft partner funding, first workloads run in a few weeks. Then we expand: more datasets and tags, operational data combined with master data, the reports and data products that drive value. Then we scale that beachhead into production.

The OT data analytics questions we get

We already have a historian, why do we need this?

Historians store data well. They aren't reporting platforms, they won't scale to an enterprise, and it's hard to combine what's inside them with fleet systems, ERP platforms, or mine planning tools. We're not replacing your historian. We're unlocking what's already in it and combining it with everything else in a common data model.

Is it safe to connect OT systems to the cloud?

Directly? Probably not, and you shouldn't want to. That's why the design uses multi-stage exfiltration, with protections at every boundary between OT, the DMZ, and the IT and cloud layer. There should be no direct path between your operating environment and the far broader attack surface the cloud represents, and messaging should only ever run in the intended direction.

Is four to six weeks realistic, or is that a sales number?

Fair question. It depends almost entirely on engagement from the operations team. Where they benefit directly — say we're taking manual daily reporting off their plate — they're invested, and they'll fact-check numbers that are heading up the chain. Four to six weeks is realistic there. Where we're waiting on responses, it takes longer.

Every site differs, too: a different historian, different fleet and plant systems, no common data model, different tags. A good naming convention helps enormously, but it has to be learned site by site.

How do we justify the investment?

Proof beats projection. Build something small, quickly realizable and demonstrably valuable, with a roadmap that lets you build forward from it, and the next phase gets much easier to fund.

About the authors:

Headshot of Stream Author

Dan Kronstal

Principal Architect, Insight Canada

Dan Kronstal is a principal architect with Insight Canada who helps organizations modernize their data and AI capabilities to drive business outcomes. He specializes in cloud-native data platforms, advanced analytics, and AI-enabled solutions using technologies such as AWS, Azure, Microsoft Fabric, Power BI, Amazon Bedrock, Amazon SageMaker, and Azure AI Foundry. Dan has led advisory and implementation engagements across energy, mining, transportation, healthcare, and logistics, helping clients turn complex data into actionable insights and scalable solutions.

Headshot of Stream Author

Scott Yaworski

Data & AI Lead, Insight Canada

Scott Yaworski is the manager of data and AI at Insight Canada, where he helps organizations maximize the value of modern data, analytics, and AI platforms. With expertise in solution strategy, technical leadership, and client engagement, he works closely with clients to align technology investments with business goals. Scott collaborates with consulting, architecture, and delivery teams to design and implement cloud, data, and AI solutions that accelerate digital initiatives, improve decision-making, and deliver measurable business outcomes.