Skip to content

Why energy AI pilots fail in production and how to fix it

by Anthony OlazabalOCT 6, 20268 min read

Your last AI pilot worked because a person hand-built the data foundation. Getting to production means building a foundation that does that work without them.

TL;DR

Energy AI pilots usually work because a data scientist quietly supplies the context the model needs, and nobody can do that across a whole fleet. A continuous data foundation is what gets AI into production. A bigger model won't do it.

  • Most energy AI pilots succeed on hand-mapped tags, filled gaps and converted units, and that manual work disappears when the model moves to the next site.
  • A foundation that scales has five layers: real-time streaming, a governed Unified Namespace, context and metadata, validation close to the source and reusable KPI features.
  • The fastest route to a production backlog is to list everything the data scientist did by hand on your last successful pilot.

Who this blog is for: CDOs, CTOs, VPs of Digital and transformation leaders at energy companies who are evaluating AI or agentic tooling and deciding where to invest first.

Here's a pattern I've seen at energy companies many times. A team picks a sharp use case, such as turbine fault prediction, inverter underperformance or compressor maintenance. They export history from the historian, clean it by hand and build a model. The pilot works. Then they try to roll it out across the fleet, and it stalls.

The pilot worked because a person created the context the model needed. The next site has different tags and the next OEM uses different units. There's no real-time data path, and no one is available to hand-assemble all of that at fleet scale. Energy AI pilots fail in production because the data foundation the data scientist built by hand can't be reproduced at scale.

Fixing that is a foundation problem, and it's the subject of this post.

This is the sixth post in the From Fragmented Data to Agentic Operations series and the start of the "Act" step. Earlier posts covered why a Unified Namespace matters now, real-time streaming from edge to enterprise and making KPIs comparable across your fleet.

Why do energy AI pilots work but fail in production?

Pilots succeed because the hardest part of the work is done by a person and never written down.

A data scientist who knows one site, one OEM and one historian can fill every gap from memory. That knowledge stays invisible until the model has to run somewhere else.

A data foundation for AI is the set of systems that continuously supplies a model with live, trusted, contextualized operational data, so no individual has to assemble that context by hand.

In a pilot, the data scientist is the foundation. In production, software has to do the job.

The outcome at stake is practical: A fault prediction that reaches the control room after the turbine has tripped doesn't help anyone, nor does a methane alert that arrives a day late.

Production AI is only worth funding if it reaches the people who can act while there's still time to change what happens.

What does energy AI need beyond operational data to work in production?

AI needs to know what the data means and more raw data doesn't supply that. Energy companies already collect huge volumes from turbines, inverters, batteries, substations, wells, compressors, SCADA systems and OEM portals. It’s clear that volume isn't the constraint.

A model needs answers to questions like these:

  • Which asset produced this signal, and which unit applies?
  • Is the value valid, or was it interpolated or flagged as bad quality?
  • Is the asset curtailed, faulted or running normally?
  • What metadata describes the equipment, and which operating constraints apply?
  • Can this signal be trusted for this specific decision?

In a pilot, the data scientist answers all of these on the fly. Once data streaming is in place you have data but not shared understanding, and intelligence doesn't scale without shared understanding. A pilot builds that understanding by hand, and production has to build it into the architecture.

This is the role of industrial data contextualization: adding the meaning, relationships and governance that turn raw signals into information people and AI can use. Read our blog, Contextualize: Turning Data into Understandable Operational Information.

What breaks when a pilot moves to the fleet?

The list of failures is almost always the same:

  • Tags differ by site. The mapping the data scientist built covers one location only.
  • Units differ by OEM. Conversions were applied by hand and never written into the data path.
  • Historical gaps were quietly filled. Production data arrives with the same gaps and nobody fills them.
  • There's no real-time path. The pilot ran on a static export, while production needs live operational state.
  • KPI definitions differ across sites. The model learns one definition and gets fed another.
  • Asset metadata is incomplete. It often lives in a spreadsheet on one person's laptop.
  • Outputs can't reach operators in time. A recommendation that arrives after the event has no value.
  • Approval boundaries were never defined. No one agreed what the model is allowed to recommend or trigger.

This list is the gap between a promising pilot and AI that operations teams rely on every day.

What are the five layers of an AI-ready data foundation that scales?

If a human supplying context is the bottleneck, the fix is the five layers that supply it continuously instead.

  1. Real-time streaming. Many energy decisions live in the operational window, fault response, curtailment, dispatch, leak triage, grid events. Batch data trains models; it can't run them in production. You need live operational state, which in practice means MQTT-based event streaming from edge to enterprise.

  2. A governed Unified Namespace. Instead of each model wiring into each system, the AI layer subscribes to governed topics representing the current state of assets, sites, KPIs, events, and metadata. This is the structure that replaces the data scientist's per-site tag knowledge with something every model can read the same way, ISA-95 as scaffolding, extended with KKS, IEC 61850, and oil and gas identifiers.

  3. Context and metadata. A turbine power value means little without nameplate capacity, wind speed, curtailment state, and maintenance history. A BESS state-of-charge value means little without state of health, cycle count, and operating mode. A methane reading means little without pipeline segment, pressure drop, location, and severity threshold. A definitional _meta namespace, asset registry, models, nameplate, commissioning, warranty, geo, is what turns a stream into knowledge. This is the layer pilots fake most heavily and production needs most.

  4. Governance and validation. A model fed bad data produces a bad recommendation; an agent acting on bad data produces operational risk. Validation has to happen close to the source, catching a malformed payload as a deviation the moment live data stops matching the expected model, before it reaches the model, not after a wrong call has been made.

  5. Reusable intelligence products. Not raw telemetry, normalized availability, capacity factor, performance ratio, BESS state-of-health score, methane severity, compressor health index, grid-event classification, operating state, maintenance risk score. These are the features a model actually consumes, computed and validated once and reused, rather than re-derived by hand for every project.

A pilot needs one data scientist. A fleet needs these five layers working without one.

How does the foundation support real energy use cases?

The same five layers support every major energy AI use case. What changes is the context each model needs and the outcome it delivers.

Use caseWhat the model needsWhat the foundation suppliesOperational outcome
Predictive maintenanceVibration, temperature, operating mode, maintenance history, fault codes, behavior of similar assetsConsistent namespace access, validated and enriched inputsA developing fault reaches maintenance crews before it becomes unplanned downtime
Dispatch and tradingCurrent availability, constraints, curtailment, market and grid signalsOne governed capacity factor and curtailment value shared by operations and tradingTeams decide from the same number, so decisions don't stall over definitions
Emissions and methane monitoringLive sensor data, pressure context, flow anomalies, location, severity rulesStructured, validated events routed to the right respondersA leak is triaged and routed to a crew while it can still be contained
Grid intelligenceFast correlation across substations, feeders, meters and protection relaysIEC 61850 logical nodes built into the namespace designGrid events are interpreted fast enough to support a timely response
Oil and gas productionPressure, flow, vibration, pump, compressor and leak-detection dataUpstream, midstream and downstream namespace patternsOne set of signals serves monitoring, maintenance, safety and compliance

Can energy companies skip straight to agentic AI?

No. Most energy companies are at the analytics stage today, which is the right place to be. A responsible path moves from dashboards to recommendations, then to workflow automation, then to constrained agentic action with a human in the loop, and finally to narrow autonomy inside hard limits.

Each step depends on the same thing: trusted live data. An agent with authority over operations on an ungoverned foundation will fail the way a pilot fails, except it will be taking actions when it does. In energy, uptime and safety come first, so any increase in AI authority has to be earned through reliable data and clear approval boundaries.

You don't pass through the foundation on the way to AI. The AI depends on it permanently.

How do you turn your last pilot into a production backlog?

Start with your last successful AI pilot and write down everything the data scientist did by hand:

  1. List every tag they mapped and every site-specific naming fix.
  2. List every data gap they filled and how they filled it.
  3. List every unit conversion and scaling factor they applied.
  4. List every KPI definition they chose and where it differs from other sites.
  5. Note every metadata source they relied on, including spreadsheets.
  6. Note how the model's output reached people, and how long that took.

That list is your real production backlog, and it's a data foundation backlog. Until streaming, a governed namespace, metadata, validation and reusable features can reproduce the list automatically, the next deployment will stall the same way the last one did. If the proposed investment is a bigger model, you're working on the wrong layer.

This is the foundation the HiveMQ Platform is built around. It connects industrial data, contextualizes it so it can be trusted, analyzes it where it's created and supports governed action from the edge to the cloud. The outcome is AI recommendations that operators can rely on at the first site and the 40th.

A trustworthy foundation makes AI recommendations reliable. It also raises a harder question as soon as anyone suggests letting AI act: what is software allowed to do to a live energy operation, and how do you limit it? The next post in this series covers that.

Talk to a HiveMQ expert about building the data foundation your AI pilots depend on.

Frequently asked questions

Energy AI pilots usually succeed because a data scientist manually maps tags, fills data gaps, converts units and chooses KPI definitions for one site. At fleet scale, nobody can repeat that manual work site by site. Without a continuous data foundation that supplies context automatically, the model gets inconsistent inputs and its results stop being reliable.

A data foundation for AI is the set of systems that continuously supplies a model with live, validated and contextualized operational data. In energy it typically includes real-time streaming, a governed Unified Namespace, asset metadata, validation close to the source and reusable KPI features. It replaces the manual context a data scientist provides during a pilot.

Usually not. When a model works in a pilot and fails at the next site, the cause is usually the data feeding it. Differences in tags, units, KPI definitions and metadata across sites break models that worked well in isolation. Investing in the data foundation fixes the cause, while a larger model trained on the same inconsistent inputs inherits the same problem.

A Unified Namespace gives every AI model one governed, consistent view of assets, sites, KPIs, events and metadata. Models subscribe to standard topics and don't need custom integrations with each system. In energy, structuring the namespace with ISA-95 and extending it with KKS and IEC 61850 lets models interpret signals the same way across every site and asset class.

Move in stages: dashboards, then recommendations, then workflow automation, then constrained agentic action with a human in the loop, then narrow autonomy within hard limits. Each stage depends on trusted, validated live data and defined approval boundaries. Giving an agent operational authority before the foundation is governed adds risk to live operations.

Share this on social media

Anthony Olazabal

Anthony Olazabal

Anthony is the Director of Solution Engineering at HiveMQ, leading the Solutions Engineering organization and helping customers and partners design and adopt scalable, enterprise-grade solutions around MQTT, IoT, and cloud technologies. With extensive experience across software development, cloud architecture, and infrastructure, he combines deep technical expertise with a strong focus on business and customer outcomes. His experience spans IaaS, PaaS, and SaaS, with a particular focus on cloud-native architectures, distributed systems, and connected technologies. Anthony is passionate about helping organizations turn complex technology challenges into practical, scalable solutions and about driving the adoption of MQTT and IoT architectures across the enterprise. He regularly shares his insights and experience on MQTT, IoT, cloud technologies, and the evolution of connected systems.

Related Content