A health score for a transformer fleet sounds like a simple idea: take the relevant condition indicators, weight them, combine them into a number between 0 and 100, and now you have a basis for prioritizing maintenance. In practice, the number is only useful if the weights are right, the inputs are credible, and the score is interpretable by the people who have to act on it. Get any of those wrong and you have a number that correlates poorly with actual failure risk, generates maintenance decisions that are either too conservative or too late, and erodes trust with the engineers who need to use it.
This piece is about the technical decisions involved in building a health score that actually works, drawn from the methodology we use in the Magnefy platform.
What the Score Needs to Represent
A transformer health score should represent the probability that the asset will experience a fault event requiring forced outage within a defined forward window. In practice we use a 30-day and 90-day window as the primary operating points, because those timeframes map to the planning cycle most operations teams work with: 90 days covers the next maintenance planning period, 30 days is the threshold for elevated attention.
This framing matters because it defines what the score is measuring. A score that is high because the transformer is old does not mean the same thing as a score that is high because there is an active EM anomaly present. Age is a risk factor, not a fault indicator. An old transformer with clean DGA, no EM anomaly, and a recent bushing inspection has a lower near-term failure probability than a younger transformer with elevated hydrogen trend and a developing EM signature change. A score that treats age as the dominant input will systematically overestimate the risk on well-maintained old transformers and underestimate it on younger units with active fault development.
The Input Categories
The inputs we treat as credible health indicators fall into four categories, with different data quality and temporal resolution characteristics:
Continuous electromagnetic monitoring. This is the real-time layer. The normalized harmonic deviation from the transformer's own statistical baseline is the primary continuous input to the score. It has the highest temporal resolution and the earliest detection of winding fault onset, but lower specificity than DGA or PD testing for fault type characterization. We weight this input heavily for the 30-day window, where near-term winding fault detection matters most, and less heavily for long-range condition assessment where age and historical test results are more informative.
Dissolved gas analysis history. DGA results are high-specificity diagnostic data when they show anomalies, but they are point-in-time and typically infrequent. We encode DGA as two sub-inputs: the absolute concentration levels relative to IEEE C57.104 thresholds, and the rate of change between successive samples. A transformer with hydrogen at 80 ppm and flat trend is in a different risk category than one with hydrogen at 80 ppm and rising by 20 ppm per month. The rate-of-change input requires at least two data points separated in time; for transformers with only a single historical DGA result, we treat the DGA contribution as lower weight with wider confidence bounds.
Load and thermal history. The cumulative thermal aging index from IEEE C57.91, computed from the actual load profile and ambient temperature history, gives us an estimate of the remaining paper insulation life. Transformers that have absorbed sustained overload events in their history have a meaningfully higher brittle-winding risk than those that have run near nameplate continuously. We also track the number of through-fault events in the protection relay history, where available, because each significant through-fault imposes mechanical stress on the winding structure and accelerates the brittle aging progression.
Asset age and last inspection record. Transformer age relative to the design life, combined with the vintage of the last comprehensive inspection, provides a prior on risk. A 30-year-old transformer that has not had an FRA or internal inspection since installation has an inherently higher uncertainty about its winding condition than one of the same age with a recent inspection result. We encode this as an information quality factor: missing or aged inspection data widens the confidence interval on the score, it does not necessarily push the score itself higher.
Weighting and Combination Logic
The combination of these inputs into a single score requires decisions about relative weighting and about how to handle conflicting signals. The weighting framework we use is time-window dependent: for the 30-day window, the continuous EM input has dominant weight; for the 90-day window, DGA trend and thermal aging model contribute substantially.
The most important decision in the combination logic is how to handle a high-confidence anomaly in any single input category. If the EM monitoring shows a significant deviation from baseline that has persisted for more than 3 weeks and is consistent with a winding fault pattern, that input should drive the score into the elevated range regardless of what age or DGA says. We implement this as a floor function: any confirmed anomaly in a high-confidence category raises the score floor to at least the medium-risk band, which means the transformer gets maintenance attention even if other inputs would score it as healthy.
The opposite case, conflicting signals where one input says high risk and others say low risk, is more common than the straightforward cases. A 28-year-old transformer with a modest EM anomaly and clean recent DGA: the age-based prior says elevated risk, the DGA says currently clean, and the EM says something is developing. In these cases, the score should reflect the uncertainty honestly. Our approach is to compute a confidence interval on the score, not just a point estimate, and to surface the conflicting inputs to the user rather than hiding them behind a single number.
The Confidence Dimension
A health score without a confidence measure is of limited use. Two transformers with health scores of 62 can be in very different situations: one with a high-confidence score based on 18 months of continuous monitoring plus two recent DGA results and a recent inspection, and one with a score based primarily on age and a single 4-year-old DGA result. Both scoring 62 does not mean both have the same maintenance priority. The confidence dimension disambiguates them: the first is a moderately elevated risk with good certainty, the second is an uncertain risk that should get additional diagnostic attention to improve the score confidence.
This is why Magnefy surfaces both a score and a confidence indicator in the fleet dashboard. The maintenance decision for a high-score, high-confidence unit is straightforward: schedule intervention. The decision for a high-score, low-confidence unit is: schedule a targeted diagnostic (DGA sample or field inspection) to either confirm the risk or resolve the uncertainty. The decision for a low-score, high-confidence unit is: maintain current monitoring cadence. A low-score, low-confidence unit is the information gap: it needs a diagnostic to establish a credible baseline.
Score Drift and Recalibration
Health scores that only move in one direction, toward higher risk as assets age, are not useful for maintenance prioritization after the first year. If a transformer receives a successful refurbishment or inspection with no findings, the score should reflect the improved condition, not remain locked at its pre-intervention level. We implement score recalibration as a first-class operation: after a DGA result shows clean chemistry following an alert, the score rolls back to reflect the current diagnostic state. After an FRA confirms winding geometry unchanged, the winding-related components of the score reset.
The recalibration principle also applies to the EM baseline. After a confirmed maintenance intervention, the post-intervention period establishes a new EM baseline rather than comparing against the pre-fault baseline, which would immediately flag the transformer as anomalous relative to a baseline that no longer reflects the current asset state.
What a Useful Score Produces in Practice
The test of a health score is whether it produces better maintenance decisions than a rules-of-thumb approach. Rules of thumb like "replace everything over 30 years" or "DGA sample annually" are consistent and auditable but they treat all assets uniformly. A well-calibrated health score tells you which transformer in your fleet needs attention this month and which 30-year-old can safely wait until the next planned window because its condition data is actually good.
We are careful about the claims we make here. Our approach has been developed with a relatively small number of monitored units, and the calibration of the weighting parameters benefits from more data. The fundamental architecture reflects the logic we believe is sound: near-term risk dominated by continuous EM; medium-term risk incorporating DGA trend; long-term risk incorporating thermal aging model; confidence tracked separately from score level. Whether the specific weights we have chosen are optimal will take more fleet-years of data to determine.
What we are confident of is the structural point: a score that conflates age-based risk with active fault development, or that fails to track confidence independently from the point estimate, produces maintenance decisions that are either systematically too conservative or systematically miss specific developing faults. Getting the structure right is prior to getting the calibration right.