OmniScientist
Seismology · signal 750 noise-labelled traces 3-component broadband verdict: mixed

One noise trace in five is not noise

A catalogue says these 750 seismograms contain nothing. The system looked at the raw waveforms and found real arrivals sitting in a fifth of them.

21.7%
163 of 750 traces that STEAD labels noise carry coherent, polarised, cross-component transient bursts that are statistically indistinguishable from genuine body-wave arrivals. 95% CI 18.8–24.9%; a station-cluster bootstrap, which allows for several traces coming from the same station, gives 17.9–25.8%. The system had guessed 5–15% before it looked.

1The detector never sees a label

The score is built from four ingredients measured on the raw three-component waveform: STA/LTA amplitude (is there a burst?), rectilinearity and planarity (does the ground move along a line or in a plane, the way a body wave makes it move, rather than wandering?), and cross-channel onset coincidence (do the three components start moving at the same instant?). No labels enter anywhere.

The threshold is not chosen by eye either. It is calibrated to a 1% false-alarm rate against IAAFT surrogates, synthetic traces built to match each real trace's amplitude distribution and power spectrum while destroying its phase structure. Measured on those nulls the realised false-alarm rate is 1.07%, so the calibration holds.

Anomaly-score histograms and ECDFs for noise traces, surrogate nulls and labelled earthquakes
Score distributions for noise-labelled traces, the matched surrogate null, and labelled earthquakes. The noise population has a tail the null does not.

2It is the coincidence term doing the work

Amplitude alone finds almost nothing. A naive amplitude-only STA/LTA detector at the same 1% false-alarm rate flags just 2.0% of the same traces, an order of magnitude less. And when the coincidence term alone is ablated out of the full score, prevalence collapses straight back to 2.0%.

That single ablation is the load-bearing result: what separates these traces from background is not that they are loud, it is that three components start moving together. Contrasts between flagged and unflagged traces on onset spread, rectilinearity and planarity all land at p < 10−16.

Take one ingredient out at a time. Only the coincidence term matters, and without it the detector also stops finding real earthquakes.
Where flagged traces sit inside the labelled-earthquake score distribution: median percentile 77.
Median rectilinearity and planarity, flagged against unflagged. Both are dimensionless and share this axis; onset spread is kept off it because it is measured in samples — there the medians are 21 against 2,304, a hundredfold gap (p = 1.3×10−84).

3Its own hypothesis was wrong, and it said so

Before running, the system pre-registered a guess: that short-period and accelerometer channels (HN, EH, SH) would be the contaminated ones. That was refuted. Prevalence is lowest on HN accelerometers at 2.4% and highest on the broadband channels, 29.1% for HH and 32.8% for BH.

Channel code is not the operative variable at all. Prevalence runs from 0% to 65% across networks (χ² p = 3.5×10−14), a stronger split than channel type gives. The right reading is that this is a site-specific noise environment effect, not an instrument-class effect. This is why the run's own verdict is recorded as mixed rather than supported: the headline prevalence held, the mechanism guess did not.

By instrument type. The pre-registered prediction is inverted.
By network. Site identity separates the data far more sharply.

4Does it survive being poked?

The prevalence is stable across false-alarm rates from 0.1% to 5%, across three independent null models, and across four STA/LTA window settings. The station-cluster bootstrap confirms it is not an artefact of a handful of chatty stations.

Prevalence against the false-alarm rate you demand, for three independent nulls.

5Every number on this page

Noise-labelled traces flagged as coherent arrivals21.7% (163/750)
95% confidence interval18.8 – 24.9%
Station-cluster bootstrap CI17.9 – 25.8%
Realised false-alarm rate on the null1.07%
Amplitude-only baseline, same false-alarm rate2.0%
Full score with the coincidence term removed2.0%
Prevalence on HN accelerometer channels2.4%
Prevalence on HH / BH broadband channels29.1% / 32.8%
Flagged traces' median percentile in the earthquake distribution77.1

What this does not show

It does not show that the STEAD labels are wrong. A trace labelled noise means no catalogued source, and a coherent local arrival too small to be catalogued is exactly what you would expect to find in that category. The claim is about what is present in the waveform, not about mislabelling. The system's own pre-registered hypothesis about instrument type was also refuted, which is why the run is recorded as mixed; the site-level explanation that replaced it was arrived at after seeing the breakdown, so it should be treated as a hypothesis for a fresh dataset rather than as a tested result.

Dataset
STEAD, 60-second three-component broadband seismograms at 100 Hz; 750 noise-labelled traces of a balanced 1,500-trace sample
Run
engine/examples/stead_seismic
Paper title
Coherent Polarized Signals in a Substantial Fraction of Sampled Noise-Labeled STEAD Traces
Figures
Produced by the run itself, reproduced here unmodified