The Duck Chugger
← The Duck Chugger

The Duck Chugger · Field Data · The Model in Production · Part 2 of 2

From White Paper to Duck Forecast

The white paper earned an honest +6.3%. The production engine that powers a private club's daily hunt outlooks scores +20% at club blinds and +27% at the refuges — and both numbers are true. This is the story of the gap: the statistical bridge, the trust dials, the walk-forward tuning loop, and the graveyard of good ideas that earned exactly zero. Told three ways — the TL;DR, the plain-English walkthrough, and the computer-science deep dive.

Want to grade the model yourself — every refuge, every season, every one of 42,819 blind predictions, day by day, with the factor badges that explain each call?

Open the production explorer →

The article · three depths, pick yours

The figures · jump to a chart

01 The gap 02 The training ladder 03 34 seasons, graded 04 The trust dials 05 When to trust it

The code, data & docs — all public

Part ISixty seconds, no math

The TL;DR

Last week's white paper proved a careful statistical model can predict California refuge duck hunting +6.3% better than the smartest naive guess. This article is about what happened when that finding went to work — inside Filthy Spoon, a private duck-club reservation app whose members see a hunt outlook for every blind, every day. The production version of the model scores +20.3% on club-style predictions and +26.7% at refuge level, on the same public data, judged by the same walk-forward rules.

+43%

Forecasting a whole season

On season averages the engine misses by just 0.17 birds per hunter — climatology misses by 0.30. This dividend is earned by tracking recent form: the paper's static model loses to the naive guess at this scale.

+32%

Forecasting your week

Zoom out to whole refuge-weeks and the skill grows: a typical miss of 0.44 birds per hunter vs climatology's 0.65. Aggregation smooths the day noise faster than it dulls the edge.

~0.6

The typical miss, in birds per hunter

At day scale the engine's forecast lands within 0.59 birds per hunter of what actually happens, on days that average about 2 — where guessing from the refuge's own history misses by 0.80.

+6.3→20.3%

The gap is not cheating

The paper answers "which effects are real?" with full uncertainty accounting. The engine answers "what's tomorrow's number?" and is free to use anything that survives a walk-forward — including signals a clean inference model must exclude.

1 + λ(m−1)

One formula runs everything

Every factor is a multiplier m pulled toward neutral by a learned trust dial λ between 0 and 1. Thirty dials, each tuned only on seasons the model hadn't seen. λ=0 means "measured, and ignored."

+6.8pp

The biggest win isn't weather

"How did this refuge shoot its last three hunt-days?" adds more skill than every weather factor combined — but only for refuges. A club blind has no such history, which is why the two headline numbers differ.

λ=0

The graveyard is the proof

Bird counts, satellite water, snowpack, refuge water deliveries, neighbor spillover: all measured, all swept, all earned zero trust. A tuning loop you can't lose to is a tuning loop you can't trust.

34/34

No losing seasons

The refuge-profile model beat climatology in every one of the 34 held-out seasons it predicted blind — the worst year it ever posted was still +17%. Consistency, not a hot streak.

The fine print, up front: production skill is measured exactly like the paper's — walk-forward, out of sample, hunter-weighted, against a day-of-season climatology baseline — across 42,819 held-out hunt-days over 34 seasons. The difference is what the two systems are allowed to use and what they're for. Everything here replicates from one public repo, offline, including the weather cache.

Part IIThe plain-English tour · no stats background needed

The walkthrough

What a hunt outlook actually is

Open the app on a Tuesday night and every blind on the club's calendar carries a number: a 0–100 outlook, where 50 is a typical day for that spot at that point of the season. Behind the number is a predicted birds-per-hunter rate — never a claim about how many ducks exist, always a claim about what a hunter in that blind should expect to bag. The number comes with its reasons in plain English: "on opening weekend", "under a bright moon", "a migration push off the Klamath Basin".

The engine that produces it is almost insultingly simple to state. Start with the obvious guess — what this spot typically does around this date. Then multiply by a stack of small corrections, one per condition that matters:

rate = base × timing × weather × front × moon × day-of-week × opener × water × …

Each correction is just history: on days like this — same weather bucket, same phase of the season — this refuge shot 1.3× its normal. The intelligence is not in the multipliers. It's in how much each one is trusted.

The trust dials

Raw historical multipliers lie constantly. A refuge with nine snowy days in its record will show some wild snow multiplier that's pure luck. So every factor passes through one gate:

shrink(m, λ) = 1 + λ × (m − 1)

λ is a dial from 0 to 1. At λ=1 the multiplier is used at face value. At λ=0 the factor is completely ignored — measured, computed, and then deliberately not believed. At λ=0.3, a raw "1.5× better" claim becomes a cautious 1.15×.

And the dials are not set by opinion. Each one is tuned by walk-forward backtest — the same honesty protocol as the paper. Stand in 1992 knowing only 1987–91, predict every hunt-day, step forward, repeat through 2025. A factor's λ rises only if trusting it more makes those blind predictions better. Overconfident factors get dialed down automatically; useless ones get dialed to zero. The result is a model that is exactly as opinionated as its track record earns — 42,819 held-out predictions' worth.

01

Three honest numbers, one data set

Walk-forward skill vs day-of-season climatology. Hover a bar for what each system is allowed to use.

The paper (+6.3%) is an inference machine: pre-registered covariates, uncertainty on everything, nothing the science can't defend. The club engine (+20.3%) may use any signal that survives the walk-forward and transfers to a blind with no harvest history of its own. The refuge engine (+26.7%) additionally knows each refuge's own recent form and rest schedule. Same data, same test, different rules — that's the whole gap.

The paper asks "is this effect real?" The engine asks "does this number help tomorrow?" Those are different jobs, and they license different tools.

How the skill was actually earned

The engine didn't start at +20%. It started at +1.3% — barely better than the calendar — and climbed through fourteen training rounds over a summer, each one a specific idea, tested blind, adopted only if the walk-forward paid it. The ladder below is the honest changelog. Three rungs are worth telling.

The day-of-week discovery. The single strongest per-day signal in the record isn't weather at all. California's public areas run on a weekly rhythm — shoot days on Wednesday and Saturday at most refuges — so a rested Wednesday shoots nearly a bird per hunter better than a pressured Sunday morning, the day after the whole state emptied its shell boxes. The backtest pinned that factor at λ=1.0, full trust — the only dial in the model at maximum. Weather is nice; rest is destiny.

The paper's rain hangover, shipped. The white paper's one exploratory win — success peaks on days that are dry but recently wet — went into production as three factors (position in a wet spell, prior-week rainfall, days since rain). The dials landed at 0.2/0.3/0.1: real, modest, and in the model today. That's the pipeline working as designed: reviewer idea (thank you, J. Coslovich and C. Overton) → frozen protocol → paper addendum → production factor, in about two weeks.

Recent form, the monster. One question — "how did this refuge shoot its last three hunt-days, versus its own normal?" — added +6.8 points of skill, more than every weather factor combined. Birds are either around or they aren't, and the most direct evidence is what happened Wednesday. But a club blind can't use it (no published harvest history), so the engine runs two profiles off one factor vector: the club-transferable prior at +20.3%, and the refuge model with recent form and rest at +26.7%. Neither number borrows the other's evidence.

02

The training ladder: fourteen rounds from +1.3% to +20.3%

Club-transferable walk-forward skill after each shipped training round. Hover a rung for the story.

Every rung is a re-run of the full 34-season walk-forward with one new idea in. The big steps: the day-of-season baseline (the migration moves in days, not half-months), the water-year and opener factors, and the day-of-week / wind-direction / barometer round that nearly doubled skill in one go. The gold extension is the refuge profile: + rest cycle, + recent form.

Where it wins and where it doesn't

Averages hide things, so the backtest grades every slice of itself: every season, every region, every refuge, every weather condition. The per-season scoreboard below is the production twin of the paper's Figure 01 — and it should look familiar, just taller. The refuge-profile model is positive in every one of the 34 held-out seasons; the worst year it ever posted was still +17% over climatology. Where the paper's leaner model lost whole seasons (the data-poor early folds, the genuinely weird years), the production stack's recent-form factor reads the weirdness off the refuge's own last three hunt-days and rides it instead of fighting it.

03

34 seasons predicted blind, production edition

Walk-forward skill by target season, refuge profile. Hover for detail; letter grades are the app's own.

Each bar is a season the model had never seen, scored with the refuge profile (the one the in-app report cards use). Every season grades A — the same grading scale (A ≥ 10%, B ≥ 6%, C ≥ 3%) that club admins see, and the drill-down in the explorer goes all the way to individual days.

Zoom out and it gets better

Everything above scores the engine day by day — the hardest possible grain, where a single hot blind or a busted flight can swing a refuge's number wildly. Most real questions are coarser: how will my week at the refuge go? What kind of season is this shaping up to be? So I re-scored the same out-of-sample predictions at refuge-week and refuge-season scale, aggregating the forecast and the climatology baseline identically so neither side gets an unfair average (the reproduction script is scripts/experiment_aggregation.py in the repo).

The engine's edge doesn't just survive the zoom-out — it compounds: +26.7% at day scale becomes +31.9% over refuge-weeks and +43.1% on season averages, where the typical miss falls to 0.17 birds per hunter against climatology's 0.30. And this is not automatic — it has to be earned. The white paper's static model, put through the identical exercise, loses to climatology at these scales (−3.7% weekly, −9.6% seasonal). The difference is the recent-form factor: average away the daily weather and what's left of the season question is "is this refuge running hot or cold this year?" — drift that only a model tracking the refuge's own current state can see. The part 1 deep dive has the full experiment.

The graveyard, and why it matters

Half the credibility of a tuned model is the list of things it refused to believe. The backtest swept a dial for every idea I could compute, and these all landed at λ=0: the mid-winter aerial bird count, the pre-season breeding flight index, satellite open-water acreage, Sierra snowpack, federal refuge water deliveries, Yolo Bypass inundation, and neighbor-refuge spillover. Some of those are genuinely informative about ducks — just not about your odds of bagging one, once the model already knows the water situation. The paper's abundance null, re-confirmed by a completely different machine, with money on the line.

The dials that did earn trust, and how hard each is believed — for ducks, and separately for geese, which turn out to be a different animal entirely — are all in Figure 04.

04

Every trust dial in the engine

Backtest-tuned λ per factor. Toggle the target; hover a bar for the factor's story.

λ=1 means full face-value trust; λ=0 means measured and ignored (the graveyard, bottom block). Note the geese toggle: the goose model leans ~10× harder on the day's weather (conditions λ 0.1→1.0) and managed water, and zeroes the moon, opener, ENSO, and migration pushes that drive ducks. Ducks ride the calendar and the flyway; geese ride the day.

When to trust the forecast

The last thing a member deserves to know is when not to listen. The backtest grades the model inside every weather bucket, and the honest answer is: trust it most on clear, cold, rested, cold-front days — the model's A-grades — and hold it loosely in snow and gales, where the record is thin and the grade drops toward a shrug. Those per-condition grades ship in the app as hover-explained report cards, and they're all in Figure 05.

05

The model's report card, by condition

Out-of-sample skill inside each bucket of each factor. Pick a factor; hover for sample sizes.

Skill within a bucket means: on just those days, how much better than climatology was the model? Cold fronts, clear skies, drought days and short-rest days are the model's best calls. Snow and gale days are near-guesses — printed here with the same prominence as the wins, because a report card you only publish when it's good isn't a report card.
Part IIIThe compsci deep dive · architecture, bridge, and tuning loop

The deep dive

The bridge: why a GLMM and a multiplier stack are the same model

The white paper's inferential engine is a Bayesian negative-binomial GLMM: log-rate = intercept + covariate effects + refuge and season random intercepts. The production engine is a base rate times shrunk multipliers. The technical bridge (also in the repo as markdown) works the equivalence: exponentiate the GLMM and it is a multiplicative model — exp(Σ effects) = Π multipliers. The random intercepts are the base rates. The regularizing priors are the λ dials: a prior that pulls a coefficient toward zero is exactly a λ that pulls a multiplier toward 1, applied on the log scale versus the linear scale. Empirical-Bayes shrinkage of data-poor cells appears in both systems under different names.

The two implementations differ where their jobs differ. The GLMM buys calibrated uncertainty — posteriors, credible intervals, joint hypothesis tests — at the cost of an MCMC fit that takes hours and a covariate set that must stay small and pre-registered. The engine buys speed and coverage — 30 factors, retuned nightly if needed, serving predictions in microseconds from a dictionary lookup — at the cost of point estimates only. The bridge's punchline: they agree where they overlap. Rain hurts, wind helps, the opener towers, the moon taxes, abundance is null — in both machines, at compatible magnitudes.

Architecture: batch-calibrate, live-serve

The production constraint that shapes everything: the app must never train at request time, never call an external service on the hot path, and never break when a data source disappears. The answer is a strict two-phase design — every expensive thing happens offline and lands in a committed JSON artifact; the running app only reads.

1 · The white paper

Bayesian NB GLMM on 39 seasons of public data. Establishes which effects are real, with uncertainty. Frozen, pre-registered, replicable.

gitlab.com/duckchugger/refuge-harvest-model
↓the technical bridge: GLMM ≡ regularized log-additive engine

2 · ETL — build the record

51k+ CDFW hunt-day reports × Open-Meteo ERA5 weather × nine habitat/climate factor datasets (USGS flow, CDEC reservoirs, PDSI, ENSO, rice acreage, midwinter counts, snowpack, surface water, eBird), all cached and committed.

scripts/build_*_dataset.py → apps/club/data/*.json + scripts/.wxcache/
↓one committed conditions table + factor tables

3 · Calibrate — the walk-forward tuning loop

Replay 34 seasons blind; coordinate-descent every λ (plus matching scheme, baseline window, per-region overrides) against hunter-weighted out-of-sample MAE. Emits the calibration artifact, every scorecard slice, and all 42,819 day-level predictions.

scripts/backtest_refuge_model.py → refuge_backtest*.json (~20 MB)
↓committed artifacts; the app never touches the ducks repo or any API at runtime

4 · Serve — the engine in the app

Pure-Python module reads the artifacts and answers predicted_rate(refuge, date, conditions) in microseconds: base × timing × Π shrink(mf, λf). Two profiles: club-transferable (blends into each blind's own history) and refuge (adds rest + recent form). Weather comes from the app's cached forecast layer; every factor degrades to neutral when its data is missing.

apps/club/refuges.py · moon.py · (private app: outlook blend, report-card views)
↓report cards grade the model in public

5 · Report cards & the explorer

Admin dashboards render every scorecard slice with drill-down to individual predicted-vs-actual days, factor badges explaining what the model saw. The public explorer on this page is a static rebuild of that surface from the same artifacts.

refuge_backtest_days.json → day_rows() / score_rows()

The serving engine, precisely

# apps/club/refuges.py - the entire prediction, conceptually
rate(refuge, day) = base                       # recency-weighted birds/hunter for this refuge
                  × timing                    # day-of-season shape, ±14d pooled, New-Year-seam aware
                  × Πf (1 + λf·(mf − 1))       # ~30 factors, each an empirical multiplier

mf  = rate in this factor's bucket / base   # skipped (neutral) if under ~150 hunter-days
λf  = backtest-tuned trust, per factor      # per-region overrides where CV says they earn it

Implementation notes that matter in practice:

The tuning loop as a systems problem

The backtest is the heart of the repo: ~1,400 lines that replay history honestly. Per candidate configuration it must rebuild training tables per fold (rates from strictly prior seasons — the recent-form factor reads only prior days within the fold), predict every held-out day, and score hunter-weighted MAE against a climatology that is itself refit per fold. The λ search is coordinate descent over a ~30-dimensional grid — sweep one dial holding others fixed, iterate to a fixed point — because the response surface is cheap to evaluate per-dial but the joint grid is astronomically large. Full run: minutes, not hours, because everything is preaggregated into per-bucket count pairs before the sweep.

Leakage is the failure mode that would silently inflate everything, so the invariants are explicit: factor rates from training seasons only; the precipitation z-score normalizer may use the full weather record (exogenous to harvest — weather statistics can't leak outcomes) while the harvest rates in the water table stay training-only; opener anchors come from the published schedule; recent form reads prior days only. The unit tests pin the bucket edges between the standalone engine and the private app's weather module, so the two can never drift apart silently — the drift test is in the public repo.

PRD and EDD, in one paragraph each

The PRD (model sections extracted) specifies the feature from the member's side: every future day carries a 0–100 outlook with plain-language reasons; club-blind predictions blend the public-refuge prior with the club's own logged history, weighted by how much comparable history exists (a new club leans on 37 years of public data, a well-logged club trusts itself); every outlook carries an evidence-based confidence; past days show the model's graded accuracy, never a self-graded number. The PRD's change log — 25 model rows of it — is the training ladder in prose, and reads like a lab notebook.

The EDD (§16 extracted) specifies the engineering: the two-phase batch/serve split, the artifact contracts, the factor registry (EXTRA_ARTIFACTS/EXTRA_FACTORS — one tuple registers a new region×season factor end-to-end: ETL → sweep → serving), failure tolerance, and the rule that retraining is a data commit, not a deploy. Two standing plan documents — the refuge model plan and the season pipeline — govern the search discipline and the in-season operational loop (weekly CDFW results → shrunk level recalibration → sharper remaining-season forecast; a 2025–26 replay lifted held-out season-level skill from +9.8% to +13.3%).

Ducks vs geese: one engine, two personalities

Re-running the identical walk-forward on the species-split record produces the article's cleanest systems lesson. Ducks (~95% of the bag) reproduce the shipped model almost exactly: +20.3% club, +27.1% refuge. Geese come in at +10.4%/+16.9% — and with a different tuned personality: the goose vector cranks the day's weather to full trust (conditions λ 0.1→1.0), leans harder on managed water, and zeroes the moon, the opener bump, ENSO, and every migration push that drives ducks. None of that was designed; all of it fell out of the same coordinate descent pointed at different birds. Flip the toggle on Figure 04 to see it.

Threats to validity, production edition

The paper trail

This project generated an unusual amount of writing, and all of the model-relevant parts are public in the replication repo's docs/: the technical bridge (the GLMM≡engine equivalence, markdown + PDF), the PRD model sections with the full 25-row model change log, the EDD prediction-engine design, the refuge model plan (the standing search discipline: season-CV gates, coarsest-grain-that-wins, the interaction matrix), the persistence factor plan (the paper-finding-to-production case study), and the season pipeline (the living-forecast operations manual). The private app around the engine — bookings, members, the Django views — stays private; nothing in it affects a number on this page.

Sources & reproducibility

Everything replicates from gitlab.com/duckchugger/spoon-model: the serving engine (apps/club/refuges.py, pure Python, no framework), the full backtest and every ETL and experiment script, all committed model artifacts (~20 MB of JSON), the cached ERA5 weather record (~76 MB, so the backtest reproduces the published numbers fully offline), the unit tests, and the design docs. pip install -r requirements.txt && pytest && python scripts/backtest_refuge_model.py rebuilds the headline numbers from scratch. Harvest data is CDFW public records via the ducks dashboard snapshot; weather is Open-Meteo ERA5; hydrology USGS/CDEC; climate NOAA.

Every figure on this page is generated from those artifacts by this article's own build script, so the numbers here cannot drift from the shipped model's. The white-paper companion is The Duck Day Model.