Supercell tracking & mesocyclone detection

An assisted-annotation pipeline on MRMS radar. Every storm below was found, tracked and annotated with no human input; the white stars are the manual archive, overlaid for comparison only — they feed into nothing.

overview scan→SAM3 storm mask→ detector inside mask→linked meso_id segments

—
Raw annotation file · download
rows where in = out — one value entered twice, so vrot is unreliable · the row outlined in blue is the scan the frame on screen is aligned to · # is the row number used in the walkthrough below
Most frames have no click, and that is the data, not a gap in the viewer. MRMS updates every 2 minutes but the radar volume-scans every 6–10 minutes, so the annotator could not mark more often than that: 3–5 frames pass between consecutive rows. On two of the three cases the viewer is already at that ceiling. The faint hollow dots fill the gaps with positions interpolated onto the 2-minute grid — they are not clicks, and they are drawn differently so they cannot be mistaken for one.
A frame and its row differ by about four minutes on the clock. That is deliberate: the archive's scan time is when the radar swept that elevation, and it lines up best with the MRMS frame roughly four minutes later — a calibrated offset measured across six events, not a mismatch. The Δ shown beside the frame is the residual after it.

Scroll to zoom · drag to pan · hover to read the value under the cursor · ←→ step frames · space play/pause

How to read it. Azimuthal shear is the rotation field: red is cyclonic. A mesocyclone is a compact red core sitting on or beside a strong reflectivity core, often with blue (anticyclonic) alongside it — that pair is the velocity couplet. A mesocyclone is about 5.6 km across, so zoom in to see its structure. The detector probability field is what the network outputs; it is only allowed to fire inside the dilated search region.

What the three cases show

CaseRegimeVolume scans Auto segmentsMedian errorWithin 5 km

Error is the distance from an automatic detection to the manual click, compared only within the same storm (storms are matched first). Without that step, a day like Hurricane Ian — 33 annotated storms packed into one area — pairs one storm's annotation with another storm's detection and reports a meaningless 40 km. For scale: a manual click is itself a median 1.67 km from the true couplet centre, measured against NEXRAD Level II.

Finding the storms in the first place

Overview pane, 19 April 2023
The top 24 candidates the overview layer ranked for 19 April 2023, with no human input. Green border = a manually annotated storm really is there: 24 of 24. Across 21 events, leave-one-event-out: the top 20 rows cover 86.6% of the day's real supercells, the top 30 cover 95.0%.

Two layers that are easy to confuse

Overview seed versus annotation product
The overview pane draws one circle per storm — that storm's highest-probability moment, used as the seed point for SAM3. The mesocyclone annotation product is the bottom row: the detector run inside the mask at every volume scan, linked into meso_id segments.

A real defect in the archive, and how it was found

The archive defines vrot = (|v_out| + |v_in|) / 2 — half the velocity difference across the couplet. That holds for all 7,816 rows, so the arithmetic is consistent. But in 28.2% of scans the annotator put the same number in both columns (44.9% of the test events), and those rows do not mean what they appear to.

MRMS AzShear settles it. It measures the same rotation from the same radars, NOAA computes it independently, and the annotator never saw it — they worked in Level II velocity displays. So it can be used as a referee that neither party can have tuned to.

Three scans of one storm at the same measured rotation
Three scans of one storm, all at essentially the same independently measured rotation (AzShear 0.010–0.011 s−1, which is "strong"). Reading both velocity extremes gives 12.25 m/s. Entering one number twice gives 1.00 m/s in one scan and 10.00 m/s in another — for the same true rotation.
row in 20220412-10-track.csvtimetiltinout vrot recordedAzShear measured
row 8 — same number twice23:49Z3.118° 1.01.01.000.011
row 40 — same number twice01:23Z6.374° 10.010.010.000.010
row 11 — both extremes read00:03Z3.119° 8.016.512.250.010

What is happening. A velocity couplet has two sides: one flowing toward the radar, one away. vrot is meant to average their magnitudes. When both columns hold the same number the annotator read one side and entered it twice, so vrot collapses onto that single side. If the couplet happened to be symmetric the answer is still about right (row 40). If it was asymmetric — and a real couplet usually is — the answer can be almost anything (row 8, recording 1.0 m/s while the storm was strongly rotating).

This is lost information, not a bias you can correct. Pooled across events with an event fixed effect, these rows report 0.74× the rotation at equal measured AzShear (95% CI 0.67–0.82, p = 5×10−9), and 8 of 11 events lean the same way. But a single multiplier would be the wrong fix: the row itself carries no record of which case it is. The pipeline therefore flags them (vrot_suspect) and keeps using their position, which is unaffected — the annotator still clicked the right place, they just wrote down one number instead of two.

On the choice of example: event 20220412 is one of the three (of eleven) where the pooled effect runs the other way — visible as the grey bar in the chart. That is deliberate. Picking an event that agrees with the archive-wide claim would be choosing the evidence to fit it, and the per-row unreliability, which is the actual finding, shows up either way.

The corroborating check. Among rows the archive records as weak (vrot < 10 m/s), the same-number-twice group has higher measured AzShear than the others (0.008 against 0.006 s−1). Rows labelled as the weakest rotation are, by independent measurement, rotating harder. That inversion is difficult to explain any other way, and it is why filtering the archive on vrot would quietly throw away real mesocyclones.

What is automated, and what is not

#Annotation stepStatusEvidence
1Choose which storms to annotateAutomated Top 20 covers 86.6%, top 30 covers 95.0% (21 events, leave-one-event-out)
2Track the storm through its lifeMostly automated 306/308 tracked; oversized masks cut from 25.8% to 17.0%
3Locate the couplet each scanAutomated 83.3% of manual segments recovered on tracked storms, 1.4 km median
4Read the out/in velocitiesNot possible Automatic Level II reading correlates r = −0.03 with the human value
5Split into meso_id segmentsPartial Linking produces segments; fragmentation still too high
7QC: units, conventions, anomaliesAutomated Quantisation-grid test found all 83 files recorded in knots
Step 4 is a hard limit, and it is where this archive's value sits. out and in have to be read from a single radar's raw velocity field, with aliasing resolved: 58% of the sampled volumes are aliased and 26% of the annotated values exceed the volume's Nyquist velocity — the display shows a folded number and the annotator unfolds it mentally. An automatic reader finds the couplet correctly (median diameter 5.62 km against a known 5.6 km; centre 1.67 km from the human click) but cannot read its speed: automatic dealiasing fails hardest on steep gradients, and a strong couplet is a steep gradient. The annotators are not doing mechanical work here.

Held-out test set

ModelTraining eventsRecall Over-productionPosition error
3-model ensemble, single split1859.2% 3.23×3.1 km
5-fold ensemble, cross-validated21 65.8%2.66×3.1 km

Split by event: Southeast cold-season 68.6% / Texas supercells 65.2% / Hurricane Ian 64.3% — the three regimes are within four points of each other. The test set was looked at twice and is now closed; everything since is measured on cross-validation.