Best Chemical Manufacturing Public Datasets
If you work in chemical manufacturing and want to experiment with using AI to solve plant problems, the first thing you need is a good dataset. Public chemical manufacturing data is hard to come by, but I’ve managed to find 9 high-quality sources – 4 time-series datasets, 5 P&ID datasets – you can use for your AI projects. Each section also includes a handful of project ideas you can get started on with Claude Code today.
You can visit my GitHub repository for more detailed information on each dataset.
Time-Series datasets
Below is a collection of both real and simulated time-series datasets for a variety of chemical processes. Each dataset has unique strengths and weaknesses, so the best one for you will depend on your particular use case.
Tennessee Eastman Process (TEP)
The standard benchmark for fault detection. It’s a simulation of a chemical process with a reactor, separator, and stripper. 52 variables, normal operation plus 20 fault scenarios, with hundreds of labeled runs per scenario. If you’re testing a fault detection method, this is what everyone compares against. The main downside of this dataset is that it’s simulated and isn’t grounded in real plant data. I made a YouTube video building machine learning models with Claude Code on this dataset.
NOBOOM
The first public collection of real chemical process data labeled for anomaly detection. It’s six distillation datasets covering batch and continuous operation, five from controlled lab and pilot plants and one from a real BASF industrial process. Most public process data is simulated, making NOBOOM’s real data a standout feature. Good for building and benchmarking anomaly detection across different process types and plant scales. Keep in mind there’s a bit of a learning curve to this dataset, but worth the investment if you want a good dataset.
Coal-Fired Industrial Boiler Dataset
Operating data from a real coal-fired boiler at a chemical plant in Zhejiang, China. 30 process variables like pressure, temperature, flow, and oxygen, logged every 5 seconds over 5 days. This is a smaller dataset than the others but excellent for proof of concepts. Great for anomaly detection on real plant data. I built an agentic AI troubleshooting assistant with Claude Code on top of this dataset.
Industrial Penicillin Simulation (IndPenSim)
A model of a 100,000-liter penicillin fermentation. The simulation is validated against data from a real industrial process, so it behaves realistically. 100 batches with 39 process variables plus Raman spectroscopy measurements. The first 90 batches run under three different control strategies and the last 10 contain faults. Good for batch process modeling, soft sensors, and fault detection on bioprocess data.
Project ideas · time-series
Below are a few projects you can build with time-series data. Feel free to copy and paste these ideas directly into Claude Code, and it will start building immediately.
- Fault detection classifier. Train a machine learning model on labeled process data so it learns what each fault looks like across dozens of variables at once. Once trained, it reads live sensor values and tells you which fault is developing before an operator would catch it. TEP is the standard starting point because the faults are already labeled for you.
- Anomaly detection without labels. Train a machine learning model what normal operation looks like on its own, then flags anything that deviates from it. No fault labels required, which matters because your plant does not have them either. NOBOOM is built for this and includes real industrial data rather than a simulation.
- Soft sensor. Use a model to predict a lab result or analyzer reading from cheap process variables you already measure continuously. It gives you a value every few seconds instead of waiting hours for a sample, and IndPenSim is the right dataset because it pairs 39 process variables with the measurements you would want to predict.
- A historian an LLM can query. Load a dataset into a time-series database and expose query tools through an MCP server, so an LLM can pull real numbers instead of guessing at them. Then you can ask a question in plain English and get an answer grounded in the data. This is the foundation for anything agentic, and it is where I would start if that is where you want to end up. The Chinese Coal-Fired Industrial Boiler dataset is perfect for this.
P&ID Datasets
Public sets of P&IDs are even harder to come by, but I recently found a workaround specifically for water treatment plants. If you discharge treated water to a river or the ground, you need a permit, and the permit application goes into a public docket along with the supporting engineering drawings (including P&IDs). Water treatment plants contain a lot of the same unit operations as chemical manufacturing plants, making them perfect substitutes for your AI projects.
Haile Gold Mine Contact Water Treatment Plant
Industrial wastewater treatment at an operating gold mine in South Carolina. P&IDs can be found on pages 47-60 and cover staged lime and ferric chloride reaction tanks, clarifiers, microfiltration skids, backwash and pH neutralization, sludge transfer, and six separate chemical dosing sheets.
Pahala Wastewater Treatment Plant
Design-build documents for a replacement municipal WWTP on the island of Hawaii. P&IDs can be found on pages 9-25 and cover anoxic and pre-aeration basins, a membrane bioreactor with permeate collection and acid cleaning, pH adjustment, grit removal, and standby generator fuel systems.
Virden Wastewater Treatment Facility
Phase 2 drawings for a replacement municipal secondary treatment plant in Manitoba, Canada. P&IDs can be found on pages 5-14 and cover the mechanical treatment train for the new facility, replacing an aging plant whose deep shaft reactor and flotation clarifiers had failed.
Salina Wastewater Treatment Plant Basis of Design Report
Preliminary design for an upgrade to a 7.25 MGD Kansas plant originally built in 1926. P&IDs can be found on pages 114, 117, 119, 121, 126, 127, 132, and 134 and cover the intermediate pump station, aeration blowers, BNR basins, scum and final clarifiers, digester heating water loop, and digester gas handling. Same file contains basis of design report and electrical drawings.
Ruttan Mine Water Treatment Plant
Acid mine drainage treatment at an orphaned mine site in northern Manitoba, using a low density sludge lime process to raise pH and precipitate dissolved metals. P&IDs can be found on pages 50-52, 62-63 and cover lime slaker and grit removal, slurry mixing and storage tanks, and two reactor trains. The same file carries the plant operating manual.
Project ideas · P&ID
Below are a few projects you can build using P&ID datasets.
- Structured extraction. Have an LLM read a P&ID and pull every equipment item, instrument, valve, line number and control loop into a table you can sort and search. It turns a drawing you have to squint at into data you can query, and it is the right first project because the output is easy to check. Count the instruments on one sheet by hand and see how many the model missed.
- Knowledge graph creation. Go a step beyond the table and capture process connectivity. The model builds a maps out the relationships between equipment, instruments and control loops rather providing than a flat list. That lets you ask questions a table cannot answer, like which instruments sit upstream of a given vessel. This is the same graph an agentic troubleshooting assistant needs to trace a problem back through a process.
- Cross-checking drawings against documents. Give an LLM both a P&ID and the operating manual or basis of design for the same plant, then ask it to find where they disagree. Tags that appear in one and not the other, equipment described differently in each. The Ruttan Mine set is good for this because the drawings and the operating manual are in the same file.