Fleet distribution
Each point is one engineer-calibrated basin with a held-out validation split. The dashed line is 1:1; the dashed crosshairs mark Moriasi monthly NSE = 0.5. The cloud shows how validation tracks calibration across the fleet.
Source: calibrationFleetSummary.json ← fleet_summary.py / fleet_summary.json. Paired median ΔNSE -0.039 (display -0.04).
High-skill exemplars
Four basins where monthly calibration and validation skill both land in the high range. Display values are rounded; exact scores are in the fleet JSON.
California · USGS 11481200
monthly NSE · calibration / validation
Kentucky · USGS 03262001
monthly NSE · calibration / validation
Oregon · USGS 14161500
monthly NSE · calibration / validation
Pennsylvania · USGS 01451800
monthly NSE · calibration / validation
Regional MMPSO basins
Three west-central Florida basins finished the same information-cascade path — pooled prior → diagnosis → short regional MMPSO — documented in the calibration methods write-up.
Peace River (57,998 HRUs)
Terminal calibration stage COMPLETE — regional MMPSO, converged at iteration 8.
Fleet calibration flagship: regionalized MMPSO seeded from the pooled prior, converged at iteration 8.
| Field | Value |
|---|---|
| Target model | 57,998-HRU Peace River basin model (HUC8 03100101, NHDPlus HR region 0310) |
| Gauge roster | 38 evaluated → 11 structural exclusions (drainage-area mismatch, lake-outlet orphaning, regulated canals); the terminal objective couples the retained gauges as per-gauge sub-objectives — 21 returned computable NSE over the scored window |
| Regionalization | cn3_swf, perco, esco, epco expanded over 9 tributary regions — 73 optimizer dimensions (insensitive + basin-uniform parameters stay global) |
| Optimizer | MMPSO with per-gauge sub-objectives (by-gauge grouping); pooled-prior vector injected as a protected initial particle |
| Objective hardening | Per-gauge NSE floor at −1 inside the optimizer sum (reported metrics unclipped) — pathological reaches are protected in-objective and reported honestly, not hand-excluded |
| Convergence | Global best reached at iteration 8 and held byte-identical through iteration 23 (spot instance reclaimed) — 15 iterations of zero improvement; ~100 model evaluations at 12 particles |
The search converged fast: the global-best objective (−19.25 in penalized units) was reached at iteration 8 and held unchanged for the remaining 15 iterations until the AWS spot instance was reclaimed — convergence, not a truncated run. A good pooled prior plus a fixed search space places the satisficing target within ~8 iterations, at roughly 100 model evaluations against the several hundred a from-scratch campaign would spend.
Per-gauge skill (2016–2022 calibration window (daily and monthly Moriasi NSE))
| Channel | Group | Daily NSE | Monthly NSE |
|---|---|---|---|
| 427 | Mainstem / Charlie core | 0.79 | 0.88 |
| 422 | Mainstem / Charlie core | 0.78 | 0.86 |
| 439 | Mainstem / Charlie core | 0.78 | 0.87 |
| 332 | Mainstem / Charlie core | 0.77 | 0.92 |
| 399 | Mainstem / Charlie core | 0.77 | 0.86 |
| 2718 | Mainstem / Charlie core | 0.68 | 0.77 |
| 269 | Mainstem / Charlie core | 0.64 | 0.91 |
| 451 | Mainstem / Charlie core | 0.63 | 0.76 |
| 220 | Interior tributary | 0.39 | 0.80 |
| 650 | Interior tributary | 0.24 | 0.47 |
| 968 | Lower-basin (floor-protected) | 0.19 | 0.12 |
| 603 | Lower-basin, regulated (floor-protected) | 0.12 | −0.27 |
| 933 | Tidal reach (floor-protected) | −0.27 | −0.49 |
| 947 | Tidal — daily degenerate, monthly recovers (floor-protected) | −2.36 | 0.53 |
Regionalization is visible in the solution: percolation (perco) ranges from 1.98 in Horse Creek to 2.83 in Shell Creek across tributary groups (Prairie 2.79, Charlie 2.70, Peace mainstem 2.52, other 2.63) — spatially-varying calibrated parameters a single global scalar could not produce, reconciling the opposite per-tributary biases.
Tampa Bay (77,636 HRUs)
Terminal calibration stage COMPLETE — regional MMPSO, converged at iteration 6.
Cross-basin case: pooled prior carried no Tampa Bay information; terminal search converged at iteration 6.
| Field | Value |
|---|---|
| Target model | 77,636-HRU Tampa Bay integrated model (70 HUC12s, HUC12-outlet 031002060700, region 0310) — a separate drainage feeding a nitrogen-limited estuary |
| Gauge roster | 32 gauges retained after coverage + drainage-area vetting; 27 returned computable NSE over the scored window |
| Regionalization | runoff-controlling parameters (cn3_swf, perco, esco, epco) distributed by tributary sub-watershed; snow / storage / channel / baseflow stay global |
| Optimizer | MMPSO with per-gauge NSE sub-objectives (floor −1); pooled-prior vector (built entirely from the Peace-area pool, no Tampa Bay information) injected as a protected initial particle |
| Convergence | Global best (−15.48 in penalized units) reached at iteration 6 and held byte-identical through the final synchronized iteration — a converged plateau, not a truncated run |
The well-gauged core is Satisfactory-to-Very-Good on the reaches that carry the basin (core daily NSE 0.52–0.70, monthly 0.62–0.85, with several interior gauges reaching monthly NSE 0.80–0.85). A cluster of seven small urban canals stays weak — their monthly PBIAS pins at roughly −100% (near-zero measured baseflow the coarse 500 m routing cannot reproduce) — and the per-gauge floor is why they neither dominated the search nor were quietly dropped. The same cascade, the same honesty rail: seeded from the pooled prior, converged in six iterations on a 77.6k-HRU model, weak tail reported at face value.
Per-gauge skill (2016–2022 calibration window (daily and monthly Moriasi NSE))
| Channel | Group | Daily NSE | Monthly NSE |
|---|---|---|---|
| 817 | Well-gauged core | 0.70 | 0.85 |
| 1106 | Well-gauged core | 0.61 | 0.74 |
| 837 | Well-gauged core | 0.60 | 0.68 |
| 1671 | Well-gauged core | 0.55 | 0.62 |
| 1979 | Well-gauged core | 0.54 | 0.85 |
| 2184 | Well-gauged core | 0.53 | 0.66 |
| 844 | Well-gauged core | 0.52 | 0.64 |
| 1868 | Interior (strong monthly) | 0.42 | 0.82 |
| 1013 | Interior (strong monthly) | 0.42 | 0.80 |
| 1495 | Interior (strong monthly) | 0.35 | 0.82 |
| 1265 | Small urban canal (floor-protected) | −0.19 | −0.78 |
| 1689 | Small urban canal (floor-protected) | −0.28 | −0.71 |
| 1267 | Small urban canal (floor-protected) | −0.29 | −0.42 |
Myakka River (29,069 HRUs)
Terminal calibration stage COMPLETE — regional MMPSO, converged at iteration 5.
Flat wet-prairie / flatwoods basin; coverage-vetted roster; converged at iteration 5.
| Field | Value |
|---|---|
| Target model | 29,069-HRU Myakka River model (HUC8 03100102, region 0310) — a flat wet-prairie / flatwoods basin draining to phosphorus-driven Charlotte Harbor |
| Gauge roster | 11-gauge coverage-vetted roster — the corrected roster after a −99 sentinel value was discovered and scrubbed; 10 returned computable NSE |
| Regionalization | runoff-controlling parameters distributed by tributary sub-watershed, insensitive + basin-uniform parameters global (same dispatch as Peace / Tampa Bay) |
| Optimizer | MMPSO with per-gauge NSE sub-objectives (floor −1); pooled-prior vector injected as a protected initial particle |
| Convergence | Global best (−7.03 in penalized units) reached at iteration 5 and held byte-identical through the final synchronized iteration |
Myakka is the flattest, most diffuse basin in the fleet, and its daily skill is inherently lower — but the whole coverage-vetted roster stays positive at both time scales. The core mainstem/wet-prairie gauges reach monthly NSE 0.56–0.65 (gauges 3592, 148, 191, 297), with daily NSE a modest 0.17–0.40 reflecting the flatwoods hydrology. No pathological floored tail: a fully-positive but honestly modest result from the same pooled-prior seed and the same short terminal search.
Per-gauge skill (2016–2022 calibration window (daily and monthly Moriasi NSE))
| Channel | Group | Daily NSE | Monthly NSE |
|---|---|---|---|
| 3592 | Mainstem core | 0.40 | 0.56 |
| 148 | Mainstem core | 0.33 | 0.64 |
| 191 | Mainstem core | 0.31 | 0.65 |
| 297 | Interior tributary | 0.21 | 0.64 |
| 676 | Interior tributary | 0.25 | 0.51 |
| 490 | Interior tributary | 0.25 | 0.42 |
| 220 | Interior tributary | 0.23 | 0.31 |
| 506 | Interior tributary | 0.18 | 0.46 |
| 950 | Flatwoods / low-yield | 0.17 | 0.21 |
| 433 | Flatwoods / low-yield | 0.09 | 0.19 |
Cost context for this class of run: about $14 AWS compute for a full regional calibration of a ~57k-HRU model — see AWS calibration page. Fleet home: Florida nonpoint-source page.
Transferable priors
Why the 44th calibration is cheaper than the 1st
A parameter distribution pooled from previously calibrated Peace-area models was applied to the 77.6k-HRU Tampa Bay catalog model in one simulation, with no optimizer. That zero-shot transfer is why later basins start closer to a usable solution — the 44th calibration inherits what the first forty-three already learned.
Largest |Δ monthly NSE| among well-assigned gauges (top 12). Green = improvement, amber = decline. Source: tampaPooledPriorDeltaNse.json. Roster move: Satisfactory+ 19→25, Very Good 4→11; correctly-assigned median monthly NSE 0.05→0.50.
| Gauge set | Satisfactory+ | Good+ | Very good | Median monthly NSE |
|---|---|---|---|---|
| All 58 evaluated gauges | 19 → 25 | 9 → 17 | 4 → 11 | −0.11 → +0.26 |
| Correctly-assigned subset (n=32, DA ratio 0.7–1.4) | 11 → 16 | 7 → 12 | 4 → 8 | 0.05 → 0.50 |
Manuscripts
Preprint · Under peer review (2026)
Platform paper · SSRN
In peer review (2026)
Engine acceleration · method page
Calibration methods write-up: calibration methods.
Appendix — Florida & Illinois worked examples
Compact single-gauge appendices: initialization → calibration → verification metrics, Morris screening, and web hydrographs. Florida reaches strong split-sample skill; Illinois is the harder snowmelt/baseflow case shown unedited.
Florida controlled basin (USGS 02297600)
Calibration station: 02297600 (gage channel 2).
Interactive figures: hydrographs and Morris tornado are web-rendered from committed JSON (same source as the metric chips / table).
Calibration global best (daily)
| Field | Value |
|---|---|
| Calibration station | USGS streamgage 02297600 (NHDPlus HR region 0310) |
| Scored calibration window | 2013-01-01 to 2018-12-31 |
| Independent verification window | 2019-01-01 to 2024-12-31 |
| Simulation & optimizer | Simulated 2010–2018 with a 3-year warm-up; particle-swarm optimization with 36 particles over 70 iterations, 6 run concurrently. |
Sensitivity analysis (Morris)
| Rank | Parameter | μ* (mean effect) | σ (interaction) |
|---|---|---|---|
| 1 | surq_lag | 0.6435 | 0.4685 |
| 2 | perco | 0.288 | 0.4113 |
| 3 | cn3_swf | 0.1413 | 0.1699 |
| 4 | alpha_bf | 0.0921 | 0.0987 |
| 5 | dep_wt | 0.0882 | 0.1565 |
| 6 | spec_yld | 0.0545 | 0.0945 |
| 7 | flo_min | 0.0519 | 0.0774 |
| 8 | dp_es | 0.05 | 0.0282 |
Calibration and verification results
| Stage | Period | Daily NSE | Monthly NSE | Daily KGE | Monthly KGE | PBIAS (%) |
|---|---|---|---|---|---|---|
| Initialization pool best | 2013–2018 | 0.197 | 0.266 | 0.088 | 0.21 | 46.746 |
| Calibration global best | 2013–2018 | 0.762 | 0.746 | 0.561 | 0.546 | 42.9 |
| Verification global best | 2019–2024 | 0.69 | 0.73 | 0.679 | 0.725 | 22.87 |
Illinois controlled basin (USGS 05536265)
Calibration station: 05536265 (gage channel 25).
Interactive figures: hydrographs and Morris tornado are web-rendered from committed JSON (same source as the metric chips / table).
Calibration global best (daily)
| Field | Value |
|---|---|
| Calibration station | USGS streamgage 05536265 (NHDPlus HR region 0712) |
| Scored calibration window | 2020-01-01 to 2024-12-31 |
| Independent verification window | 2012-01-01 to 2015-12-31 |
| Simulation & optimizer | Simulated 2018–2024 with a 2-year warm-up; particle-swarm optimization with 48 particles over 50 iterations, 6 run concurrently. |
Sensitivity analysis (Morris)
| Rank | Parameter | μ* (mean effect) | σ (interaction) |
|---|---|---|---|
| 1 | melt_min | 3.82 | 9.8425 |
| 2 | k | 3.6832 | 3.7334 |
| 3 | cn3_swf | 3.3198 | 2.0242 |
| 4 | perco | 3.1507 | 4.9258 |
| 5 | melt_max | 2.9509 | 9.8071 |
| 6 | urban_cn_c | 2.8681 | 2.9149 |
| 7 | mann | 1.782 | 1.3773 |
| 8 | surq_lag | 1.7715 | 2.685 |
Calibration and verification results
| Stage | Period | Daily NSE | Monthly NSE | Daily KGE | Monthly KGE | PBIAS (%) |
|---|---|---|---|---|---|---|
| Initialization pool best | 2020–2024 | -0.581 | -1.288 | 0.248 | -0.083 | 42.488 |
| Calibration global best | 2020–2024 | 0.105 | -0.139 | 0.445 | 0.369 | 36.459 |
| Verification global best | 2012–2015 | 0.233 | 0.427 | 0.42 | 0.647 | -18.841 |
Notes
- Fleet metrics come from the published engineer-run digest (paired calibration vs held-out validation).
- Verification uses an independent window with no overlap with the scored calibration period (split-sample).
- If you order a calibration run on your own model, the wizard shows runtime/cost estimates before launch.
