These notes come from my 15 years in substation protection and automation engineering before co-founding Magnefy. They are not a literature review. They are an account of what I observed working alongside maintenance engineers and operations staff on distribution and substation transformers, what approaches earned operational trust and why, and what specific failure patterns these approaches consistently missed.
I want to be careful about scope here. I am describing my direct experience and observations from those years, which span a particular set of utility types, equipment vintages, and operational contexts. Other engineers working in different contexts may have had different experiences. These are notes, not a definitive survey.
Thermal Monitoring: Useful for What It Measures, Blind to Most Fault Families
Top-oil temperature monitoring is the most widely deployed continuous monitoring technology for oil-filled power transformers. It is cheap, reliable, and well-integrated with SCADA systems. For what it measures, it works well: it tracks thermal loading and provides a direct input to the IEEE C57.91 thermal life calculation. When a transformer is being thermally abused, top-oil temperature monitoring will flag it.
The limitation is that most of the fault families that lead to in-service failures do not produce elevated top-oil temperature until they have progressed well beyond the stage where intervention would change the outcome. A developing turn-to-turn winding fault, in its early stages, involves a small fraction of the total winding turns. The heat generated by the fault current in the shorted turns is dissipated locally into the oil and does not produce a measurable increase in bulk oil temperature until the fault has progressed to a large enough fraction of the winding to represent a meaningful thermal load. By that point, the transformer is often already in a condition where a trip is imminent.
I have been in post-mortems on winding failures where the top-oil temperature records showed normal values right up to the protection trip. That is not a failure of the measurement. It is a reflection of the fact that top-oil temperature is not a sensitive indicator of early winding fault development. The engineers who monitor top-oil temperature know this and do not expect it to catch winding faults. But when top-oil monitoring is the only continuous monitoring method in use, the implication is that winding faults are effectively invisible until they are catastrophic.
Load-Based Thresholds: Where They Work and Where They Do Not
A common approach in SCADA-connected transformer management is to set alarm thresholds on measured quantities: load current relative to nameplate, power factor deviation, voltage unbalance, neutral current on wye-connected windings. These thresholds are useful for catching overload conditions, system-level imbalances, and gross electrical anomalies.
For the winding fault family, load-based thresholds have a structural limitation: the electrical quantities measured at the meter point are dominated by the transformer's normal operating behavior. A developing turn-to-turn fault creates a low-impedance parallel current path within the winding, but the primary effect of this path is on the internal current distribution within the winding, not on the bulk electrical quantities measured at the terminals. From the terminal perspective, a transformer with a small developing turn-to-turn fault looks nearly identical to a healthy transformer at the same load.
The neutral current deviation alarm, which is sometimes used as an indicator of winding faults in grounded wye transformers, can be useful for detecting significant winding faults but is generally not sensitive to early fault onset. The sensitivity depends on the fault geometry, its location within the winding, and the grounding configuration. Faults in the middle third of a winding in a solidly grounded transformer produce the largest neutral current contribution; faults near the neutral end of the winding or in high-impedance grounded systems may produce minimal neutral current anomaly even with significant winding involvement.
Load-based threshold monitoring is not monitoring the fault families that produce the highest-consequence in-service failures. It is monitoring load and system-level electrical health. Both are important, but they are different things.
DGA: Where It Is Well Positioned and Its Operating Constraints
Dissolved gas analysis is the most effective diagnostic tool in the standard toolkit for oil-filled transformer fault detection, when applied correctly. The chemistry is sound: specific fault types produce specific dissolved gas signatures, and the key gas method and Duval triangle provide a principled interpretation framework. For detecting active high-energy arcing, sustained partial discharge, and significant thermal faults, DGA is reliable and well-validated.
The operating constraint is sampling frequency. Standard practice for unmonitored distribution transformers is quarterly or semi-annual DGA sampling. For a fault that progresses over 6 to 12 months, quarterly sampling may catch it in an intermediate state. For a fault that develops rapidly over 4 to 8 weeks, a quarterly sampling cycle may entirely miss the window between detectable gas elevation and failure.
Online DGA systems close this gap for the transformers where they are installed. But online DGA is not economically justified for every distribution transformer in a typical fleet. Utilities typically instrument their highest-risk and highest-consequence assets with online DGA and accept the detection gap for lower-tier assets.
The other constraint is that DGA, like top-oil temperature, responds to the thermal and chemical consequences of fault energy dissipation. For the winding fault family, there is a lag between fault onset and detectable gas evolution. Partial discharge from deteriorating turn-to-turn insulation produces hydrogen and methane, but the quantities at early fault stages are small relative to background gas levels and the gas transport time from the fault location to the oil sampling point adds additional delay. Early-stage winding fault detection by DGA requires either very good background data on each transformer's normal gas evolution profile or the fault to be producing gas at a rate well above background levels.
What the Monitoring Gap Looks Like in Practice
In the years I worked in substation protection, the most common failure scenario I saw was a transformer that had gone through all its scheduled maintenance cycles without anomalous findings, and then failed in service. Post-mortem inspection would typically reveal significant winding insulation degradation that had been developing for a year or more, with no surface indication during routine inspection and no anomalous DGA at the most recent scheduled sampling.
The failure was not a surprise from a physics standpoint: the insulation degradation was visible in the post-mortem, and the signs were there to be read if you knew where to look and had a way to look continuously. But the monitoring architecture, which relied on periodic DGA sampling and thermal monitoring, had no mechanism for continuous observation of the internal winding condition between samples.
The gap is not a failure of the monitoring tools. It is a coverage gap between what those tools measure and what needs to be measured to catch early winding fault onset. Thermal sensors and load monitoring cover the thermal loading and electrical performance envelope. Periodic DGA covers the chemical evolution of the oil. What is missing is a continuous physical measurement of the winding current distribution condition, which changes when a fault begins to develop regardless of whether a DGA sample is due or whether the thermal loading is elevated.
What EM Signature Monitoring Adds to the Stack
The EM signature approach that Magnefy uses addresses the winding current distribution gap directly. The harmonic composition of the terminal EM field changes when the winding current distribution changes, and it changes continuously as the fault develops rather than only when a periodic sample is taken or when a thermal threshold is exceeded.
I want to be precise about what this adds and what it does not replace. It adds a continuous, always-on measurement of winding condition that responds to the internal current distribution change from early winding fault onset. It does not replace DGA for detecting active arcing faults or significant thermal faults, for which DGA is the better-established and more directly interpretable signal. It does not replace thermal monitoring for managing thermal life and detecting overload conditions.
The value is in fault family coverage. The combination of thermal monitoring, DGA, and EM signature monitoring covers the major fault families with at least one continuous signal. Without continuous EM monitoring, the winding fault family has essentially no continuous early-stage indicator in the standard monitoring stack. That is the gap this approach closes, and in the context of what I observed in field operations, it is the gap that most often corresponded to unexpected in-service failures.