Industrial Data Infrastructure for Digital Twin Readiness in Oil & Gas

Industrial Data Infrastructure for Digital Twin Readiness in Oil & Gas | Crunch-IS Case Study
We established a secure, repeatable data pipeline from the isolated plant network to AWS, continuously collecting selected sensor and SCADA data, structuring it around physical assets, and landing it into an AWS data lake — giving a US oil & gas operator the data foundation required for future digital twin use cases.
Industry:

Oil & Gas

Location:

USA

Team Size:

6

Duration:

5 months

Technologies
AWS IoT Core
AWS IoT SiteWise
Amazon S3
AWS Glue
AWS Glue Data Catalog
Amazon Athena
Grafana
Amazon CloudWatch
Litmus Edge
01

About the Client

The client operates a natural gas processing and compression facility in the United States. Their plant handles inlet separation, gas compression, dehydration, amine treatment, condensate storage, export metering, and flare management — a dense operational environment generating continuous sensor data across dozens of assets.

The facility ran SCADA systems, PLCs, a local historian, and engineering spreadsheets. All of it was confined to the plant network.

Industrial Data Infrastructure for Digital Twin Readiness in Oil & Gas | Crunch-IS Case Study
02

Challenge

The client had years of valuable plant data. None of it was ready for digital twin use.

The core problems were:

  1. Data locked inside the plant network. Sensor readings, historian tags, and SCADA outputs were accessible only on-site. No cloud infrastructure existed. Digital twin vendors had no path to the data.
  2. Tags without context. SCADA and historian tags were unstandardized and not mapped to physical assets. A tag value alone told you nothing about which compressor, separator, or valve it described.
  3. No asset hierarchy. Engineering teams had no clean, structured model connecting sensor data to the physical equipment it monitored. Data existed — but it wasn’t organized in a way that supported modeling or analysis.
  4. Fragmented historical data. Historical records were scattered across local systems, lacked a unified structure, had inconsistent timestamps, and lacked quality flags. Pulling a coherent picture of past plant behavior required manual effort that the team couldn’t sustain.
  5. No path to cloud analytics. Without a data lake, there was no foundation for digital twin scenarios, predictive maintenance, or operational dashboards — regardless of what tools the client might want to use later.
03

Solution We Delivered

We built a secure, structured data infrastructure that took plant data from SCADA and historian systems and delivered it to a cloud-based AWS data lake — organized by asset, standardized for quality, and ready for digital twin consumption.

OT-to-Cloud Data Ingestion

We established a secure pathway from the plant network to AWS. Litmus Edge was deployed at the OT/IT boundary to connect in read-only mode to the plant SCADA system and local historian. It collected approved operational tags, performed initial validation and data conditioning, buffered data when needed, and prepared telemetry for cloud ingestion.

From Litmus Edge, selected telemetry was securely forwarded to AWS IoT Core, using MQTT over TLS. AWS IoT Core handled device authentication, encrypted communication, and telemetry routing from the industrial edge layer into AWS. AWS IoT SiteWise organized incoming time-series data around physical assets — compressors, separators, pumps, valves — turning raw SCADA tags into asset-contextualized data that made sense to downstream systems.

Data Lake Architecture on Amazon S3

Raw, cleansed, and curated data were loaded into a structured Amazon S3 data lake with three distinct zones. Raw ingestion preserved the source data as-is. The cleansed zone applied standardized timestamps, units, tag names, and quality flags. The curated zone held processed datasets shaped specifically for digital twin and analytics use.

Tag Mapping and Asset Hierarchy

We reviewed and exported SCADA and historian tag lists, then mapped each sensor tag to its corresponding equipment and process unit. The result was a clean asset hierarchy for the facility — connecting every data point to the physical asset it described.

Data Transformation with AWS Glue

AWS Glue handled cleaning, transformation, and preparation at scale. AWS Glue Data Catalog stores metadata, schemas, and table definitions so datasets are discoverable and queryable. Amazon Athena provided engineers with direct SQL access to data in S3.

Data Quality Monitoring

We ran initial data validation to confirm completeness, accuracy, and expected sensor ranges across the tag population. Ongoing monitoring — built using Amazon CloudWatch — detected missing values, stale tags, abnormal readings, and ingestion gaps in real time.

Grafana Dashboards

Grafana was used as the engineering dashboard layer on top of curated AWS data. It allowed operations and engineering users to visualize selected SCADA trends, monitor data quality metrics, identify ingestion gaps, and review key asset signals before the data was consumed by future digital twin applications.

Crunch-IS image case study
04

Client’s Results

The client received a secure data foundation for digital twin implementation. Plant sensor and SCADA data were extracted, standardized, mapped to physical assets, and landed in AWS. This allowed the client to move from isolated plant data toward cloud-based analytics, digital twin scenarios, and future predictive maintenance use cases.

Have a Question? Let’s Get in Touch!

Tell us what you’re building or where you’re stuck. We work with engineering and product teams on custom software, AI & ML, cloud infrastructure, DevOps, and UI/UX — from early scoping to long-term delivery. One conversation is usually enough to know whether we’re the right fit.

Email: [email protected]

    I have read and accepted the Terms of Use and Privacy Policy *

    We and our partners use technology such as cookies on our site to personalize content and ads, provide social media features, and analyze our traffic. Click “Accept” to consent to the use of this technology across the web.

    Decline