Validation

Prediction vs reality.

The question every coordinator, agency and reviewer asks of drift-prediction software is simple: when it predicted where someone would be, how close was it? This page reports how SAROS has been tested against real recorded outcomes — real ocean drifter tracks, twenty-one published lost-person incidents, and a fully documented accident investigation — under acceptance thresholds fixed in writing before the tests were first run. Every result is published, including the exclusions.

Five layers of assurance

Each layer makes one claim, and makes it cleanly.

1

Automated engineering tests

4,250+ tests on every code change: the software behaves as specified.

2

Golden regression baselines

29 fixed scenarios re-run on every release: model outputs cannot change unintentionally, and any deliberate model change appears as an auditable diff.

3

Closure against published science

Leeway physics verified against the US Coast Guard field-trial literature (Allen & Plourde 1999): drift speed reproduced exactly, direction error 0.0° under controlled wind.

4

Real-outcome golden set

Predictions tested against where real drifting objects — and, in one documented accident case, real people — actually were. The results below. This set runs on every release.

5

Live field trials

Instrumented drift targets with the NZ Police Marine Unit: both targets recovered inside the SAROS search area, 0.27 and 0.7 miles from prediction after 8-hour drifts.

One hundred real-world drift tracks.

One hundred satellite-tracked ocean drift tracks — one hundred different buoys, selected by a script under rules committed in writing before any data was downloaded, spanning some twenty ocean regions in both hemispheres — plus one fully documented person-in-the-water case. Ninety-nine of one hundred fell inside their pre-declared thresholds; the one exception missed by 70 metres and is published with its explanation.

0.18 NM Median prediction error, 48-hour drift
0.48 NM Median across all durations, to five days
0.97 Predicted uncertainty vs actual error (ideal 1.00)
99/100 Within pre-declared threshold
Drift durationTracksMedian error 9 in 10 withinWorst case
24 hours30.39 NM0.60 NM0.65 NM
48 hours150.18 NM1.28 NM1.41 NM
72 hours640.48 NM1.56 NM2.04 NM
120 hours181.23 NM2.35 NM2.77 NM
All durations1000.48 NM1.74 NM2.77 NM
Region groupTracks
UK & European waters39
North American waters31
Southern Hemisphere (NZ, Australia, S Atlantic, S Pacific)20
Arabian Sea / Gulf region7
Open ocean3

Honestly sized search areas

Accuracy asks how close the prediction was. Calibration asks whether the stated uncertainty was honest. SAROS measures and reports both.

Every SAROS prediction is a cloud of simulated drift outcomes, and the size of that cloud is a promise: the real position should sit about this far from our best estimate. That promise decides the size of the search area — so it has to be honest. A cloud that is too confident produces a search area that is too small and can miss the casualty; one that is under-confident spreads aircraft hours across empty water. SAR practice tests this with the spread-to-error ratio: the uncertainty the model predicted, divided by the error it actually made. A perfectly honest model scores 1.0.

0.97 Spread-to-error ratio (ideal 1.00)
0.50 NM Uncertainty the cloud predicted, typical case
0.48 NM Error actually measured, typical case

Across the one hundred real drift tracks, the typical SAROS case scores 0.97 — the cloud promised about half a nautical mile of uncertainty and the measured error was 0.48 NM. Practitioners regard anything within about ±0.1 of the ideal as well calibrated.

This number is measured, not assumed — and it is maintained by a closed loop. When analysis of the longest drifts showed the cloud slightly understating uncertainty in the tail, the cause was traced, an uncertainty term was engineered, calibrated on half the population, and then verified blind on the other half against acceptance bands declared in advance. The default SAROS envelope remains the tightest credible one — the fastest route to a find on the first search — and wider, fully characterised envelopes for expanded-search planning are documented in the validation report.

A real case, treated with respect

The loss of the yacht Ouzo, English Channel, August 2006.

The MAIB investigation into the loss of the yacht Ouzo (Report 7/2007) is one of the most thoroughly documented person-in-the-water drift records in UK public literature: a fixed encounter position and time, recorded weather, and a recorded recovery position some 35 hours later. Reconstructing the case from the report's own figures — nothing tuned, thresholds declared first — SAROS's predicted search datum fell 3.55 nautical miles from the actual recovery position, and every one of the simulated drift particles finished within 12 nautical miles. A search area drawn from the SAROS prediction 35 hours after the incident would have contained the casualty. The case is used because its documentation honours the men involved with accuracy; it is reported here with the same intent.

Twenty-one real lost-person incidents.

The land search engine, scored against the published Koester casebook — every case re-run from the original planning point, every case on the record.

86% Find locations inside the P95 search area (18 of 21)
1.0 km Mean error against the central estimate, all 21 cases
80% Pass target, fixed before the run — met on both scoring bases
3 Finds beyond P95 — published by name, not removed

Twenty-one real search and rescue incidents from Robert J. Koester's Lost Person Behavior (2010) — the standard reference casebook for land SAR, selected and distance-measured by the field's reference author, not by us — were re-run through SAROS from each incident's original planning point: dementia, despondent, children, hikers, hunters, autism and gatherer behaviour across urban, agricultural, forest and mountain terrain. Eighteen of the twenty-one subjects, including the five found at the planning point itself, lay inside the P95 search area a planner would have drawn; counting only the sixteen who travelled, the result is 13 of 16 (81%) — both bases clear the 80% target declared before the run. The three finds beyond P95 are published with their overshoot distances, and where the model errs it errs generous: the search area contains the subject rather than excludes them. Every run is seeded, executes offline in seconds, and repeats on every SAROS release.

Why these numbers can be trusted

The discipline matters more than any single result.

Thresholds first, results second

Every acceptance threshold was fixed in writing before the test first ran, sized from each case's documented uncertainties — and never adjusted afterwards. A failure at the declared threshold is a published finding, not something to tune away.

Selection rules first, data second

For the global drifter sample, the ocean regions, seasons, durations and mechanical first-match selection rules were committed to a version-controlled audit trail before any data was downloaded. A script selected every track; no human chose one. The ordering — criteria, then data, then results — is provable from the record.

Everything published, reproducible by anyone

All sources are public (Crown copyright/OGL and US Government open data); all environmental inputs are embedded in the test fixtures; all runs are seeded. The full set executes offline in seconds, identically, on every SAROS release — and the three excluded segments are published with the same prominence as the passes.

SAROS has additionally been benchmarked head-to-head against the leading open-source implementation of the same international leeway science, on identical scenarios, identical environmental forcing and an identical statistical recipe — including calibration analysis (how often the find truly falls inside each stated probability region) and search-area efficiency at equal containment. The full comparative study, the validation test report and the one-page accuracy datasheet are available to agencies and professional reviewers on request.

How good is good enough?

No SAR authority publishes a pass mark for drift-prediction accuracy. The defensible answer is to define “good enough” operationally — and meet it three ways.

There is precedent for this. CASP, the US Coast Guard's computerised search planning system, served for over three decades without an accuracy threshold ever existing. When SAROPS replaced it in 2007, the case for fielding rested on demonstrations, not a number: validated components, performance at least equal to the incumbent, and characterised uncertainty. SAROS makes the same case — explicitly, and with the evidence published.

1 — Calibrated

Honest about its uncertainty

A drift prediction can never contain more information than the wind and current data feeding it — but it can be honest about how uncertain it is. A calibrated tool is, by definition, as good as the physics and the data allow. SAROS measures this: across one hundred real tracks the uncertainty it stated matched the error it made almost exactly — spread-to-error ratio 0.97 against an ideal of 1.0.

2 — Non-inferior

At least as good as the state of practice

Benchmarked head-to-head against the leading open implementation of the same international leeway science — identical scenarios, identical environmental forcing, identical statistical recipe. SAROS matched its prediction skill case-for-case, and was markedly stronger in the worst case.

3 — Superior where it counts

Search area at equal containment

The number that changes decisions is the size of the search area at the same probability of containing the casualty: a smaller area means more probability of detection per sortie — found sooner. This is the operational currency on which SAROPS itself was justified over CASP, and the measure the SAROS comparative study reports directly. The full study is available to reviewers on request.

Limitations, stated plainly

A validation page that lists no limitations should not be believed.

  • One hundred valid tracks across some twenty ocean regions support the percentile statements above — but not all duration bands are equally deep: the 24-hour group holds only three tracks. Read the table with the track counts beside it.
  • The Ouzo reconstruction is wind-driven per the investigation record; the Channel's tidal streams are absorbed in the declared threshold. Embedding a tidal stream series is the planned refinement and is expected to tighten the result.
  • Drifter cases isolate the drift computation from forecast-data quality, which varies by provider and region. Operational accuracy also depends on the environmental data available on the day — as it does for every drift tool.
  • SAROS is a decision-support tool for trained SAR personnel. It informs, and never replaces, coordinator judgement.

Request the validation reports