enzwa lab concept

Augur

An agentic predictive-maintenance system for the plant that keeps a data centre running.

in plain terms

A tool for deciding when to service a machine, rather than waiting for it to break.

The problem

Every pump, chiller and switchboard fails eventually. Service everything early and you throw away parts with years left in them; wait for the failure and you pay for the breakage and all the downtime around it. Most operations cannot tell which of the two they are doing.

What it does

Augur ages a set of equipment along its real failure curve, lets you choose a maintenance strategy and how much sensor coverage to pay for, and prices the outages you avoided against what the monitoring cost.

where the outages start
45%

of impactful data-centre outages start in the power train. Most often the UPS.

Uptime Institute, Annual Outage Analysis 2025.

what one costs
1 in 5

of operators say their most recent significant outage cost more than a million dollars. More than half say it cost over a hundred thousand.

Uptime Institute, Annual Outage Analysis 2025.

why it keeps happening

Critical plant is
maintained to a calendar.

The calendar decides when a machine needs attention months before the machine does. Sometimes that is early and a good asset is stripped for nothing. Sometimes it is a week late.

part one

The window

Machines give warning before they stop. The warning is physical and measurable, and it arrives long before anything trips.

the interval the whole system rests on

Between the first detectable sign and the moment the machine stops, there is time.

P · first detectable sign F · the machine stops health the p–f interval

Reliability engineers have called this the P–F interval for fifty years. Its length is a property of the failure mode: bearing wear gives weeks, an arc fault gives seconds.

the whole proposition

Every hour of that window you can see is an hour you can act in.

Augur finds the window early, works out what it is worth, and spends it.

part two

The system

Four layers over plant that is already instrumented for most of it. Most of the sensors are on the floor. Little reads them continuously.

what gets built

Four layers, one register, an agent for every subsystem.

01

Instrument

Vibration, motor current, temperature, battery impedance and differential pressure, read continuously off the BMS and the sensors already fitted to the plant.

02

Model

Every asset carries its failure modes, its P–F intervals, its duty history and what an hour of its downtime costs, in one register the agents read from.

03

Judge

An agent per subsystem watches its own signals, checks them against the neighbouring plant, and estimates how long the asset has against the failure mode it thinks it has found.

04

Act

It raises the work order, reserves the part, books the window against the load profile, and tells the engineer what it saw and how sure it is.

how it knows

Each subsystem gives a different amount of warning.

UPS battery string
Internal impedance climbing across a string, cell by cell, against its own baseline.
warning
30days
Chiller compressor
Vibration drifting off baseline at running speed, with motor current following it.
warning
35wks
CRAC unit
Fan vibration, coil fouling read as rising differential pressure, supply air off set point.
warning
26wks
Standby generator
Coolant temperature drift and start-time creep across successive test runs.
warning
days
PDU
Load imbalance across phases and thermal rise at the busway joints.
warning
4872hrs

Lead times are published predictive-maintenance benchmark ranges for this class of plant. They are what the design assumes, not measurements from one site.

part three

How it is run

You decide how much of it runs without you. That is a setting, and it moves as trust does.

the boundary is a setting you own

Three ways to run it.

The system is identical in all three. What changes is how much of its own decision it carries out before a person sees it.

In the loop

It proposes. An engineer approves every action before anything moves. Useful while the floor is deciding whether it believes the model.

On the loop

It acts inside limits you set: routine work orders, parts, scheduling. Anything touching live plant or spend above a threshold comes back to a person.

Out of the loop

It runs the routine work and calls you for the exceptions. The engineer's day becomes the hard calls and the escalations.

what arrives on the shift

A ranked list, and the reason for the ranking.

UPS-2 · string 4
Impedance up 18% over 40 days on cells 9 to 14. Consistent with dry-out. String carries the A feed to hall 1.
26 daysact within
CH-01 · compressor
Vibration 2.1× baseline at running speed, motor current following. Consistent with early bearing wear.
5 weeksact within
CRAC-07
Supply air 3.4°C above set point with coil differential pressure rising. Consistent with fouling.
6 weeksact within

Ranked on criticality multiplied by imminence, so the work that protects the most load and has the least time left sits at the top. Illustrative readings.

part four

What changes

Faults are found while the machine is still running. The work moves into planned windows, and the part is on site before the engineer is.

what the case is built on

The benchmark ranges the model is anchored to.

4060%
less unplanned downtime
2540%
lower maintenance cost
612mo
to payback on a mid-size hall

Published ranges for condition-based programmes on this class of plant. Your own figure depends on your outage cost per hour, your sensor coverage and how much of the estate is already past its design life.

the model is live

You can run this one yourself.

Set the maintenance strategy, choose how much sensor coverage to pay for, then watch five subsystems age along the curve. It prices the downtime you avoided against running the plant until it breaks.

Open Augur
enzwa
ENZWA

A studio for documents that have to work the first time they are read.

hello@enzwa.com · enzwa.com · Hong Kong

→ or scroll