Technical

Early Fault Detection in Oil-Filled Transformers: A Signal Processing Perspective

Nadia Osei 9 min read
Early Fault Detection in Oil-Filled Transformers: A Signal Processing Perspective
Back to blog

I came to transformer fault detection from a background in time-series signal analysis, which gave me a particular angle on the problem. The question I kept returning to in those first months at Magnefy was: where does the fault signal appear first, and how much earlier does it appear than existing monitoring methods can resolve it?

That question matters because it determines the intervention window. A fault that produces a detectable signal 30 days before catastrophic failure is qualitatively different from one that produces no signal until 48 hours before failure. The former gives an operations team time to schedule a planned outage and order parts. The latter does not. Most of the decisions in transformer maintenance come down to that distinction.

Where Faults Begin and Why They Take Time to Manifest

Oil-filled transformers fail through a progression that typically spans weeks to months from the onset of the initiating condition to catastrophic failure. This is not accidental. The oil medium provides both electrical insulation and thermal dissipation, and its properties change gradually as fault conditions develop rather than abruptly.

For the winding fault family, which accounts for a substantial share of mid-life and end-of-life failures in distribution and substation transformers, the progression typically starts with the degradation of turn-to-turn insulation. The degrading insulation increases local electric field stress in the affected region, which accelerates further degradation in a positive feedback loop. At some point in this progression, the degraded insulation begins to allow small partial discharge events: brief, low-energy ionization events in the insulation voids and along the degraded insulation surface.

These partial discharge events produce detectable chemical byproducts, particularly hydrogen and methane in the dissolved gas profile, but the quantities are initially small relative to normal background gas evolution. They also produce acoustic emissions and high-frequency electromagnetic transients. The electromagnetic transients are particularly interesting from a signal processing standpoint because they produce a characteristic distortion in the winding current harmonic composition that can be extracted from measurements at the terminal.

As the degradation progresses, partial discharge transitions to partial winding turn-to-turn fault: a low-impedance path develops between adjacent turns in the affected winding region. This changes the winding current distribution more substantially. The turns in the shorted region carry fault current, which is not reflected in the terminal measurements of the total winding current but which produces a specific perturbation in the spatial distribution of the magnetic flux in the transformer core and hence in the external electromagnetic field.

The Detection Window: What Different Signals Can Resolve

The key signal processing question is: at what stage in this progression does each monitoring method produce a signal with enough signal-to-noise ratio to be actionable?

DGA sampling at standard quarterly or semi-annual intervals means the detection window is limited by the sample frequency. Even if elevated dissolved gas is present between samples, it is not observed. If DGA is the only monitoring method, the practical detection window is at most equal to the sampling interval from when the gas generation rate exceeds background. For slowly evolving faults, that can be adequate. For faults that progress quickly once partial discharge begins, it may not be.

Online DGA systems improve this by providing continuous or near-continuous oil sampling. The detection window shifts to: how long before failure does the dissolved gas signature rise above the detection threshold? For hydrogen evolution from partial discharge, the gas concentration in the oil typically rises over weeks to months for a slowly progressing fault, but the background gas level from normal operation varies and creates a noise floor that limits early detection sensitivity. A transformer with an already elevated hydrogen background from normal hot-spot activity requires a larger incremental rise before a gas-driven alert can be distinguished from baseline variation.

Acoustic partial discharge monitoring has a different noise profile. The substation acoustic environment is dominated by load-related noise, mechanical vibration from auxiliary equipment, and ambient acoustic sources. Early-stage partial discharge produces signals in the 100 to 300 kHz range that are attenuated by the oil and tank wall before reaching an external acoustic sensor. The detection range for early-stage PD from an external sensor is typically limited to the interior of the transformer tank, and the minimum detectable discharge magnitude depends on the attenuation path between the discharge location and the sensor.

EM Signature: What the Signal Captures and When

The electromagnetic measurement approach works differently because it does not require the fault to produce a specific chemical or acoustic byproduct above a detection threshold. What it measures is the harmonic composition of the current-related EM field at the transformer terminals, and any change in the winding current distribution from its baseline state produces a corresponding change in that harmonic composition.

The winding current distribution changes when any fault geometry develops that creates a parallel current path or a region of reduced impedance within the winding. This includes partial turn-to-turn fault onset before the fault has progressed far enough to produce elevated DGA at background-distinguishable levels. In the signal processing framing: the EM signature change is a direct byproduct of the changed current distribution, not a secondary byproduct (chemical decomposition of oil, acoustic radiation) that requires additional physical processes to produce a detectable signal.

In practice, this means the EM signature method can detect winding fault onset at an earlier stage in the fault progression than methods that depend on detecting the byproducts of fault energy dissipation. The relevant comparison is not "EM vs. DGA": both have value and detect different things. The relevant observation is that for the winding fault family, EM signature changes can precede detectable DGA elevation by weeks in slowly progressing cases, and that earlier detection directly extends the intervention window.

Signal Separation: The Core Engineering Problem

None of this is straightforward in practice, because the substation electromagnetic environment is not quiet. The signals that indicate winding fault onset are small relative to the operating signal, and they coexist with noise sources that produce harmonic content in the same frequency bands: variable-frequency drives, power factor correction capacitors, nonlinear industrial loads, and switching transients from breaker operations.

The signal processing approach at Magnefy uses three layers to extract the fault-related signal from this environment. First, load normalization: the harmonic composition of the EM field varies with load level and load type even for a healthy transformer, so raw harmonic measurements are normalized against concurrent load measurements before any change detection is applied. A raw increase in third harmonic amplitude during a period of high load may be normal; the same increase normalized against load is anomalous.

Second, supply-side reference: harmonics that originate on the supply side of the transformer appear in both the primary and secondary terminal measurements in a predictable ratio. By tracking the supply-side harmonic reference, it is possible to separate source-side harmonic variation from variation that is specific to the transformer's own harmonic generation. This removes a major category of false positive sources.

Third, transformer-specific baseline: the harmonic signature baseline is fit to each individual transformer's historical measurements, not derived from a fleet average or a generic transformer model. This accounts for the substantial unit-to-unit variation in harmonic behavior that exists within nominally identical transformer types due to manufacturing variation and differences in operating history.

Even with these three layers, the detection is not certain at very early fault stages. The change in harmonic signature from one or two shorted turns is small, and the separation from residual noise depends on the signal-to-noise ratio in the specific substation environment. We do not claim the method detects every fault at the earliest possible stage. What the architecture provides is a detection capability that is consistent across installations, not dependent on specific chemical or acoustic thresholds that vary with transformer condition and environment.

What Early Detection Actually Changes

The operational value of earlier detection depends on what you can do with the additional lead time. A 4-week detection lead time that is not acted upon produces no value. The signal has to trigger a response that takes advantage of the window.

When an EM signature anomaly is detected, the appropriate first action is confirmation with DGA. If the EM signature indicates early winding fault onset, the DGA sample taken at that point may or may not show elevated dissolved gases depending on how early in the fault progression the detection occurred. If DGA is already elevated, the combined signal from both methods significantly increases the maintenance team's confidence. If DGA is not yet elevated, the EM alert is the leading indicator and the DGA follow-up schedule is increased to catch any gas evolution as the situation develops.

We are not saying that EM monitoring replaces DGA in the monitoring stack for any transformer where DGA is already being performed. The two methods complement each other: DGA provides confirmation and additional fault-type discrimination; EM signature provides continuous coverage and earlier winding-fault detection. The combination is better than either method alone, and the value is concentrated in the additional lead time that earlier detection provides for the highest-consequence failure scenarios.

See fault detection in action

Request a short live session where we run your real transformer data against the detection model and show you what surfaces.

Request a demo