An assisted-annotation pipeline on MRMS radar. Every storm below was found, tracked and annotated with no human input; the white stars are the manual archive, overlaid for comparison only — they feed into nothing.
in =
out — one value entered twice, so vrot is unreliable
· the row outlined in blue is the scan the frame on screen is aligned to
· # is the row number used in the walkthrough below
Scroll to zoom · drag to pan · hover to read the value under the cursor · ←→ step frames · space play/pause
| Case | Regime | Volume scans | Auto segments | Median error | Within 5 km |
|---|
Error is the distance from an automatic detection to the manual click, compared only within the same storm (storms are matched first). Without that step, a day like Hurricane Ian — 33 annotated storms packed into one area — pairs one storm's annotation with another storm's detection and reports a meaningless 40 km. For scale: a manual click is itself a median 1.67 km from the true couplet centre, measured against NEXRAD Level II.
meso_id segments.The archive defines
vrot = (|v_out| + |v_in|) / 2 — half the velocity difference across the
couplet. That holds for all 7,816 rows, so the arithmetic is consistent. But in
28.2% of scans the annotator put the same number in both columns (44.9% of the test
events), and those rows do not mean what they appear to.
MRMS AzShear settles it. It measures the same rotation from the same radars, NOAA computes it independently, and the annotator never saw it — they worked in Level II velocity displays. So it can be used as a referee that neither party can have tuned to.
| row in 20220412-10-track.csv | time | tilt | in | out | vrot recorded | AzShear measured |
|---|---|---|---|---|---|---|
| row 8 — same number twice | 23:49Z | 3.118° | 1.0 | 1.0 | 1.00 | 0.011 |
| row 40 — same number twice | 01:23Z | 6.374° | 10.0 | 10.0 | 10.00 | 0.010 |
| row 11 — both extremes read | 00:03Z | 3.119° | 8.0 | 16.5 | 12.25 | 0.010 |
What is happening. A velocity couplet has two
sides: one flowing toward the radar, one away. vrot is meant to average their
magnitudes. When both columns hold the same number the annotator read one side and
entered it twice, so vrot collapses onto that single side. If the couplet happened
to be symmetric the answer is still about right (row 40). If it was asymmetric — and a
real couplet usually is — the answer can be almost anything (row 8, recording
1.0 m/s while the storm was strongly rotating).
vrot_suspect) and keeps
using their position, which is unaffected — the annotator still clicked the right
place, they just wrote down one number instead of two.
On the choice of example: event 20220412 is one of the three (of eleven) where the pooled effect runs the other way — visible as the grey bar in the chart. That is deliberate. Picking an event that agrees with the archive-wide claim would be choosing the evidence to fit it, and the per-row unreliability, which is the actual finding, shows up either way.
The corroborating check. Among rows the archive
records as weak (vrot < 10 m/s), the same-number-twice group has
higher measured AzShear than the others (0.008 against
0.006 s−1). Rows labelled as the weakest rotation are, by independent
measurement, rotating harder. That inversion is difficult to explain any other way, and it is
why filtering the archive on vrot would quietly throw away real mesocyclones.
| # | Annotation step | Status | Evidence |
|---|---|---|---|
| 1 | Choose which storms to annotate | Automated | Top 20 covers 86.6%, top 30 covers 95.0% (21 events, leave-one-event-out) |
| 2 | Track the storm through its life | Mostly automated | 306/308 tracked; oversized masks cut from 25.8% to 17.0% |
| 3 | Locate the couplet each scan | Automated | 83.3% of manual segments recovered on tracked storms, 1.4 km median |
| 4 | Read the out/in velocities | Not possible | Automatic Level II reading correlates r = −0.03 with the human value |
| 5 | Split into meso_id segments | Partial | Linking produces segments; fragmentation still too high |
| 7 | QC: units, conventions, anomalies | Automated | Quantisation-grid test found all 83 files recorded in knots |
out and
in have to be read from a single radar's raw velocity field, with aliasing
resolved: 58% of the sampled volumes are aliased and 26% of the annotated values exceed the
volume's Nyquist velocity — the display shows a folded number and the annotator
unfolds it mentally. An automatic reader finds the couplet correctly (median diameter
5.62 km against a known 5.6 km; centre 1.67 km from the human click) but cannot read its
speed: automatic dealiasing fails hardest on steep gradients, and a strong couplet is a
steep gradient. The annotators are not doing mechanical work here.
| Model | Training events | Recall | Over-production | Position error |
|---|---|---|---|---|
| 3-model ensemble, single split | 18 | 59.2% | 3.23× | 3.1 km |
| 5-fold ensemble, cross-validated | 21 | 65.8% | 2.66× | 3.1 km |
Split by event: Southeast cold-season 68.6% / Texas supercells 65.2% / Hurricane Ian 64.3% — the three regimes are within four points of each other. The test set was looked at twice and is now closed; everything since is measured on cross-validation.