Troubleshooting Plant Issues with Agentic AI
I started using LLMs heavily in November 2022, right after ChatGPT launched. Within a few weeks I started wondering whether they could help troubleshoot chemical plant issues, because troubleshooting was a large (and time-consuming) part of my day job. Just like it is for thousands of plant engineers and operators around the world.
There was an obvious obstacle: LLMs aren’t trained on your company’s internal data. They don’t know Reactor C’s current temperature or how it’s controlled, that your boiler’s economizer hasn’t been cleaned in 3 years, or how Line 6 performed last quarter relative to its target. Companies go to great lengths to protect this data to ensure LLMs aren’t trained on their IP.
But what if you could link an LLM to your plant’s data sources, allowing relevant plant data to flow into the LLM’s context window alongside the engineer’s prompt? Then take it one step further by providing deterministic tools that query the data and perform calculations rather than making the LLM do those itself. Then, the LLM application could reason across the data the way an engineer does and potentially solve the problem.

This troubleshooting question sat with me for over 2 years, until Anthropic released Claude Code which gave me a promising answer.
What I built
I used Claude Code to build an agentic AI manufacturing troubleshooting assistant that reasons across a plant’s data to find a root cause and recommend corrective actions.
All of it written by Claude. Keep in mind that I had practically no prior coding experience when I built this. All I did was hand Claude a few datasets, along with a process description, and explained in plain English that I wanted an agentic application that linked to all of them and could troubleshoot plant issues. In under 2 hours, Claude had built me a working application that was correctly diagnosing process issues.
The application runs on a public time-series dataset from a real coal-fired industrial boiler in Zhejiang, China. I open sourced the project on GitHub, including install steps, if you’d like to try it on your own.
The architecture
The assistant connects to the following 4 plant data sources through an MCP server.

A time-series database acting as a mock historian: 30 tags, 5 days of data sampled every 5 seconds.
A database of DCS alarm and event logs.
Relationships between equipment, instrumentation, control loops and process areas. Claude Code built this on its own.
A RAG vector database holding SOPs, troubleshooting guides and equipment datasheets. Also built by Claude Code.
The goal was to link Claude to all data sources a plant employee would need to solve a process issue. Leaving out any data sources would deprive the LLM of the necessary context it needs to successfully troubleshoot the issue.
Working a real problem
Let’s say a plant engineer notices the boiler’s steam temperature dipped to 517°C, 13°C below its target range of 530 to 545°C, and asks:
What caused the steam temperature dip on March 28?
The agentic assistant works through a troubleshooting sequence on its own, making use of the tools in the MCP server as well as the LLM for reasoning:

-
Claude Code bundles the engineer's question with the conversation history and the list of available tools: historian query, alarm log search, knowledge graph traversal, document search, etc. That package gets sent to the LLM.
-
The LLM processes the context window and decides it needs more data, so it emits a tool request: pull the steam temperature tag
TE_8332Afrom the historian across March 28. -
The tool queries the historian and returns a daily low of 517.1°C at 11:13am, and that the temperature rose to 543.1°C just before it crashed. A rise before a significant drop could mean something overcorrected.
-
The LLM emits a tool request against the alarm log. It returns the following alarms: high induced draft fan bearing vibration at 10:28am, low primary air flow at 10:48am, low flue gas O2 at 10:58am, and low steam temperature at 11:09am.
-
Looking for troubleshooting guidance, the LLM runs a semantic search of the vector database, finding the low steam temperature troubleshooting guide. 6 potential root causes are listed: a stuck spray water valve, insufficient firing, poor combustion, a load increase, fouled superheater tubes, and a faulty thermocouple. The LLM then emits a tool call to traverse the knowledge graph, mapping each cause to the tags that measure it.
-
The LLM emits a historian tool request to pull every one of those tags across the event window. Five of the six root causes are eliminated using the historian data, but the data points towards a stuck desuperheater spray valve.
-
The engineer can issue a work order to stroke test the spray valve and repair if needed, along with a work order to inspect the induced draft fan due to high vibration.
Every step is logged
To ensure an SME can review the assistant’s thought process, the troubleshooting sequence above is shown to the user along with the answer. The assistant logs all steps as it executes them: which tool it called, what input arguments it provided, what came back, and what it concluded before deciding on the next call.
When it finishes, the assistant outputs a detailed report containing that full trail alongside its answer. If the reasoning is wrong, it’s somewhere visible in the report, and a reviewer can point at the exact step where it went sideways. This is the difference between an application an SME can sign off on and one they have to trust on faith.
Where this goes
I believe an AI troubleshooting assistant like this will exist in most chemical plants in 5 to 10 years. Some companies are ahead of the curve and are already building something similar.
If implemented well, annual troubleshooting hours and downtime drop, and engineers spend more time improving their processes instead of firefighting. Most process engineers I know have a backlog of improvements they never get to because they are stuck chasing the same recurring problems.
Keep in mind that this is a small prototype, and implementing in a real chemical plant will be far more complex for a variety of reasons:
- Tag volume and tag quality. Most historians contain hundreds or thousands of tags, not 30. And many of these tags are often poorly or incorrectly labeled, making it difficult for the LLM to select the correct one to query. Most plants will need to put in a lot of up-front work to resolve these inconsistencies and build an accurate knowledge graph.
- Data availability. Good manufacturing data is hard to come by, especially for older plants and processes. Proper instrumentation doesn’t always exist, and documentation is often missing or outdated. Without this data, the agentic application won’t have the necessary context to solve process issues.
- Security and IP. Manufacturing sites are hesitant to expose their data to LLMs. Process data is intellectual property, and historians sit behind an OT/IT boundary that security teams are slow to open. Expect to need a private or self-hosted deployment, read-only access, and contractual guarantees that nothing is retained or trained on.
All of these problems are solvable though. I believe the end goal is worth the time and money needed to resolve them.