Validation
The question every coordinator, agency and reviewer asks of drift-prediction software is simple: when it predicted where someone would be, how close was it? This page reports how SAROS has been tested against real recorded outcomes — real ocean drifter tracks, twenty-one published lost-person incidents, and a fully documented accident investigation — under acceptance thresholds fixed in writing before the tests were first run. Every result is published, including the exclusions.
Each layer makes one claim, and makes it cleanly.
4,250+ tests on every code change: the software behaves as specified.
29 fixed scenarios re-run on every release: model outputs cannot change unintentionally, and any deliberate model change appears as an auditable diff.
Leeway physics verified against the US Coast Guard field-trial literature (Allen & Plourde 1999): drift speed reproduced exactly, direction error 0.0° under controlled wind.
Predictions tested against where real drifting objects — and, in one documented accident case, real people — actually were. The results below. This set runs on every release.
Instrumented drift targets with the NZ Police Marine Unit: both targets recovered inside the SAROS search area, 0.27 and 0.7 miles from prediction after 8-hour drifts.
One hundred satellite-tracked ocean drift tracks — one hundred different buoys, selected by a script under rules committed in writing before any data was downloaded, spanning some twenty ocean regions in both hemispheres — plus one fully documented person-in-the-water case. Ninety-nine of one hundred fell inside their pre-declared thresholds; the one exception missed by 70 metres and is published with its explanation.
| Drift duration | Tracks | Median error | 9 in 10 within | Worst case |
|---|---|---|---|---|
| 24 hours | 3 | 0.39 NM | 0.60 NM | 0.65 NM |
| 48 hours | 15 | 0.18 NM | 1.28 NM | 1.41 NM |
| 72 hours | 64 | 0.48 NM | 1.56 NM | 2.04 NM |
| 120 hours | 18 | 1.23 NM | 2.35 NM | 2.77 NM |
| All durations | 100 | 0.48 NM | 1.74 NM | 2.77 NM |
| Region group | Tracks |
|---|---|
| UK & European waters | 39 |
| North American waters | 31 |
| Southern Hemisphere (NZ, Australia, S Atlantic, S Pacific) | 20 |
| Arabian Sea / Gulf region | 7 |
| Open ocean | 3 |
Accuracy asks how close the prediction was. Calibration asks whether the stated uncertainty was honest. SAROS measures and reports both.
Every SAROS prediction is a cloud of simulated drift outcomes, and the size of that cloud is a promise: the real position should sit about this far from our best estimate. That promise decides the size of the search area — so it has to be honest. A cloud that is too confident produces a search area that is too small and can miss the casualty; one that is under-confident spreads aircraft hours across empty water. SAR practice tests this with the spread-to-error ratio: the uncertainty the model predicted, divided by the error it actually made. A perfectly honest model scores 1.0.
Across the one hundred real drift tracks, the typical SAROS case scores 0.97 — the cloud promised about half a nautical mile of uncertainty and the measured error was 0.48 NM. Practitioners regard anything within about ±0.1 of the ideal as well calibrated.
This number is measured, not assumed — and it is maintained by a closed loop. When analysis of the longest drifts showed the cloud slightly understating uncertainty in the tail, the cause was traced, an uncertainty term was engineered, calibrated on half the population, and then verified blind on the other half against acceptance bands declared in advance. The default SAROS envelope remains the tightest credible one — the fastest route to a find on the first search — and wider, fully characterised envelopes for expanded-search planning are documented in the validation report.
The loss of the yacht Ouzo, English Channel, August 2006.
The MAIB investigation into the loss of the yacht Ouzo (Report 7/2007) is one of the most thoroughly documented person-in-the-water drift records in UK public literature: a fixed encounter position and time, recorded weather, and a recorded recovery position some 35 hours later. Reconstructing the case from the report's own figures — nothing tuned, thresholds declared first — SAROS's predicted search datum fell 3.55 nautical miles from the actual recovery position, and every one of the simulated drift particles finished within 12 nautical miles. A search area drawn from the SAROS prediction 35 hours after the incident would have contained the casualty. The case is used because its documentation honours the men involved with accuracy; it is reported here with the same intent.
The land search engine, scored against the published Koester casebook — every case re-run from the original planning point, every case on the record.
Twenty-one real search and rescue incidents from Robert J. Koester's Lost Person Behavior (2010) — the standard reference casebook for land SAR, selected and distance-measured by the field's reference author, not by us — were re-run through SAROS from each incident's original planning point: dementia, despondent, children, hikers, hunters, autism and gatherer behaviour across urban, agricultural, forest and mountain terrain. Eighteen of the twenty-one subjects, including the five found at the planning point itself, lay inside the P95 search area a planner would have drawn; counting only the sixteen who travelled, the result is 13 of 16 (81%) — both bases clear the 80% target declared before the run. The three finds beyond P95 are published with their overshoot distances, and where the model errs it errs generous: the search area contains the subject rather than excludes them. Every run is seeded, executes offline in seconds, and repeats on every SAROS release.
The discipline matters more than any single result.
Every acceptance threshold was fixed in writing before the test first ran, sized from each case's documented uncertainties — and never adjusted afterwards. A failure at the declared threshold is a published finding, not something to tune away.
For the global drifter sample, the ocean regions, seasons, durations and mechanical first-match selection rules were committed to a version-controlled audit trail before any data was downloaded. A script selected every track; no human chose one. The ordering — criteria, then data, then results — is provable from the record.
All sources are public (Crown copyright/OGL and US Government open data); all environmental inputs are embedded in the test fixtures; all runs are seeded. The full set executes offline in seconds, identically, on every SAROS release — and the three excluded segments are published with the same prominence as the passes.
SAROS has additionally been benchmarked head-to-head against the leading open-source implementation of the same international leeway science, on identical scenarios, identical environmental forcing and an identical statistical recipe — including calibration analysis (how often the find truly falls inside each stated probability region) and search-area efficiency at equal containment. The full comparative study, the validation test report and the one-page accuracy datasheet are available to agencies and professional reviewers on request.
No SAR authority publishes a pass mark for drift-prediction accuracy. The defensible answer is to define “good enough” operationally — and meet it three ways.
There is precedent for this. CASP, the US Coast Guard's computerised search planning system, served for over three decades without an accuracy threshold ever existing. When SAROPS replaced it in 2007, the case for fielding rested on demonstrations, not a number: validated components, performance at least equal to the incumbent, and characterised uncertainty. SAROS makes the same case — explicitly, and with the evidence published.
A drift prediction can never contain more information than the wind and current data feeding it — but it can be honest about how uncertain it is. A calibrated tool is, by definition, as good as the physics and the data allow. SAROS measures this: across one hundred real tracks the uncertainty it stated matched the error it made almost exactly — spread-to-error ratio 0.97 against an ideal of 1.0.
Benchmarked head-to-head against the leading open implementation of the same international leeway science — identical scenarios, identical environmental forcing, identical statistical recipe. SAROS matched its prediction skill case-for-case, and was markedly stronger in the worst case.
The number that changes decisions is the size of the search area at the same probability of containing the casualty: a smaller area means more probability of detection per sortie — found sooner. This is the operational currency on which SAROPS itself was justified over CASP, and the measure the SAROS comparative study reports directly. The full study is available to reviewers on request.
A validation page that lists no limitations should not be believed.