The Shift Test measures whether an AI system can identify a developing industrial hazard from what a crew can actually see — using situations generated in a simulator rather than incidents drawn from the record, so that no contestant can succeed by remembering the answer.
IRVINE, CA – August 10, 2026 – EON AI Ventures today published the design of the Shift Test, an open benchmark for evaluating artificial-intelligence systems on industrial safety judgment, and invited the industry to run it.
The claim the benchmark exists to settle is a narrow one. It is not that any system is generally more intelligent than another. It is that a model trained on the industrial incident record should outperform a model trained on the open internet, at the specific task of understanding industrial incidents. That proposition is widely assumed and has never been measured.
It has never been measured because the obvious way of measuring it does not work. Read more in the EON Delta Shift Test white paper.
Why the obvious test fails
The natural design is to take real incidents, hide the outcomes, and ask each system what was developing. It is intuitive, it is what most people propose first, and it is unsound.
Large general-purpose models are retrained continuously against corpora whose contents are not disclosed. It is therefore not possible for an evaluator to establish that any documented incident was absent from a given system’s training material, nor to verify a stated cutoff date. A well-known incident is in the training data of every serious system, and a test built on well-known incidents measures recall alongside reasoning in a proportion nobody can determine.
Choosing more recent incidents does not fix it. It narrows an unverifiable margin that continues to shrink with every retraining cycle. The defect is structural, and no amount of care in selecting historical cases removes it.
Click on the image below to access the EON Delta Shift Test presentation.
The answer is the one aviation reached fifty years ago
| We test AI the way airlines test pilots.
Real physics. New situations. No answer sheet to memorise. |
Airlines do not certify pilots with a written examination about historical crashes. They put them in a simulator and give them a situation they have not seen before. Nobody has ever argued that this is unfair. It is the most trusted evaluation method in safety-critical work, and it has been for decades.
The Shift Test applies the same principle. Its situations are generated in a simulation of industrial equipment rather than selected from the incident record. Three properties make that legitimate:
- Real physics. The equipment behaves as real equipment behaves. Nothing is invented about how the world works.
- New situations. The particular combination of conditions has not occurred before, so there is no historical case for any system to recall.
- No answer sheet. Because the situation has never happened, nothing has ever been written about it. The answer has to be worked out from the physics and the readings.
| Why EON
EON AI Ventures has built simulators of industrial work for twenty-five years. Testing an AI system inside one is not a novel departure. It is the obvious use of the thing the company already has. |
What is being asked
A shift is the unit of industrial work — the stretch of time one crew is responsible for a plant. The benchmark asks precisely the question a crew asks at handover:
| Given only what can be seen right now, what is most likely to hurt us before the next shift ends — and what check would confirm it? |
Each contestant answers the same question on the same situations. Responses are stripped of any identifying characteristic and graded blind. Scoring is reported across separate dimensions rather than collapsed into a single number, because the dimensions fail differently and a single number would conceal that.
| Dimension | What it measures |
|---|---|
| Identification | Whether the developing mechanism is correctly named |
| Evidence | Whether a real, checkable comparable case is cited — and whether it exists |
| Actionability | Whether the recommended check would actually confirm or rule out the hypothesis |
| Calibration | Whether the stated confidence matches the observed accuracy |
| Abstention | Whether the system declines to answer when the evidence genuinely does not support one |
| False alarms | Performance on a separate track of situations in which nothing is wrong |
The false-alarm track carries more weight than it appears to. Serious industrial events are rare. A system with excellent detection and an ordinary false-alarm rate still produces far more false warnings than real ones, and a warning system that cries wolf is switched off by the people it was built to protect. Any benchmark that omits this measures the wrong thing.
The obvious objection, and the answer
A benchmark whose questions are generated by machine, taken by machines and graded by machines would measure only the agreement of correlated systems. The Shift Test is validated by three checks, none of which is derived from any model.
1. Working engineers cannot tell the situations apart from real ones
Generated situations are interleaved with situations taken from documented events and shown to practising process safety engineers, who are asked to identify which is which. The set is admitted only where their accuracy is statistically indistinguishable from guessing. No model participates in this procedure at any point. If engineers can spot the generated ones, the set is rejected and the generation method is revised.
2. Human experts write the answers
The reference answer against which contestants are scored is authored by domain experts from the situation alone. It is not produced by the generator and not read off the simulation’s subsequent behaviour. Where the experts and the simulation disagree, the rate of disagreement is published rather than resolved in either direction.
3. The ranking holds across generated and historical situations alike
Every contestant is also scored on situations drawn from the documented record. If the ordering of contestants is preserved across both sets and the magnitudes correspond, the two are measuring the same capability. This check has a property worth stating plainly: EON cannot arrange to pass it, because most of the contestants are systems EON did not build and cannot configure.
| The one criticism that is simply true
There is no evaluation whose author cannot influence it. Anyone who claims otherwise is selling something. What can be done is to make every point at which a hand could be placed visible to everybody else. Publishing the construction method, publishing every per-dimension score including EON’s own losses, using answer keys written by people who do not work for EON, releasing a public sample anyone can run, and having the full scoring set held and administered by a neutral party — each of those removes one such point. That is the whole of the defence, and it is offered as sufficient rather than as perfect. |
Published openly, including the losses
The construction method, the scoring rubric, every per-dimension result and a public sample of situations will be published. The complete scoring set will be held and administered by a neutral party rather than by EON, for a reason that has nothing to do with fairness: a published test set is absorbed into the next round of training and stops measuring anything.
EON expects to lose on some dimensions. A general-purpose model that reasons more fluently in the abstract, but cannot cite a real comparable case, cannot state honest uncertainty and cannot decline to answer, would demonstrate exactly the point this benchmark exists to establish. Results will be published in full, including those.
| There is no score yet
This release announces a method, not a result. EON has not run the Shift Test and is making no claim about how any system performs on it, including its own. No performance figure will be published by EON until the benchmark has been run under the conditions described here, and any figure that appears before then has not come from EON. |
Invitation
EON invites process safety engineers, operators and AI research groups to take part — as members of the expert panel that authors the answer keys, as participants in the discrimination check, or by entering a system. Contact details are below.
Find more information in the EON Delta Shift Test white paper.
Learn more by tuning to our podcast.
About EON AI Ventures
EON AI Ventures is the company behind Work Intelligence — the captured, verified, and compounding knowledge of how expert work is actually done. Its Intelligence Flywheel platform (Genesis, Field IQ, Assess IQ) enables industrial enterprises to encode expert procedures into AI-guided simulations, deliver them to any worker on any device, and verify competency in the field. EON AI Ventures builds on a 25-year foundation of immersive learning technology deployed across more than 80 countries. For more information, visit www.eonaiventures.com