The Duck Chugger
← The Duck Chugger

The Duck Chugger · Field Data · The White Paper · Part 1 of 2

The Duck Day Model

The production version of this model predicts refuge duck hunting 20–27% better than the smartest naive guess — typically within about six-tenths of a bird per hunter of what actually happens. This article is part 1 of 2: the white paper underneath that engine — what I can prove, held to a scientific standard, and why that provable number is +6.3%. Told three ways: the TL;DR, the plain-English walkthrough, and the full nerd deep dive. Part 2 is the engineering story of the production engine itself.

This is not the paper — this is the overview. The full peer-review-grade analysis report is linked below, alongside all data and code.

Want to poke the model yourself — watch it predict any refuge's season, spin the weather dials, see where it wins and loses?

Open the model explorer →

The article · three depths, pick yours

The figures · jump to a chart

01 The scoreboard 02 What moves the needle 03 The water paradox 04 The rain hangover 05 Calibration

The paper, data & code — all public

Part ISixty seconds, no math

The TL;DR

My first article toured the raw record: 39 seasons, 41 public refuges, 6.9 million birds, and a pile of averages that couldn't tell you why. This is the follow-up I promised — a real statistical model, held to the only standard that matters for a forecast: predict seasons you've never seen, and beat the obvious guess.

One definition before anything else, because every number below uses it. Skill is how much closer to reality the model gets than the smartest naive guess — the climatology: "what does this refuge usually do around this date?" Skill of 0% means the model knows nothing the calendar doesn't; +20% means its errors are a fifth smaller. The goal was never to predict the exact number of birds in the bag — single hunt-days are far too noisy for that, for anyone. The goal is to beat climatology, reliably, on seasons the model has never seen; whatever margin survives that test is real knowledge about ducks.

+43%

Forecasting a whole season

On season averages the engine misses by just 0.17 birds per hunter — climatology misses by 0.30. But this dividend is earned, not free: the paper's static model actually loses to the naive guess at this scale.

+32%

Forecasting your week

Zoom out to whole refuge-weeks and the engine's skill grows: a typical miss of 0.44 birds per hunter vs climatology's 0.65. Aggregation smooths the day noise faster than it dulls the edge.

+20–27%

Where this ends up (part 2)

The production engine built on this paper scores +20.3% on club-style predictions and +26.7% at refuge level, on the same walk-forward test (in bird terms: a typical miss of 0.59 birds per hunter). This paper is the first step toward that number.

~0.6

The typical miss, in birds per hunter

The production model's typical miss: 0.59 birds per hunter overall; by refuge, a median of 0.67 (best 0.39, hardest 1.09); 0.45–0.71 across regions; 0.49–0.71 across 34 blind seasons — on days that average about 2 birds per hunter.

+6.3%

The scientifically defensible first step

Held to scientific standards — pre-registered covariates, uncertainty on everything — timing plus daily weather beats climatology by 6.3% (95% CI +3.4 to +9.2) on 47,028 held-out hunt-days — a typical miss of 0.71 birds per hunter, vs 0.75 guessing from history alone. Modest, and provable.

+12.6%

The post-paper upgrade

Letting each refuge keep its own seasonal curve doubles the paper's skill on the same blind test (typical miss: 0.67 birds per hunter) — but that figure stays "exploratory" until it survives a season the model has never seen.

water?

The water paradox

Drought, river flow, reservoirs, rice acreage: clearly associated with success in the data it saw — and it degrades prediction on seasons it didn't. Both halves are true. That's the paper's most interesting result.

null

Bird counts don't help

Mid-winter census and pre-season flight indices add nothing to per-hunter prediction. How many ducks exist and how many end up in your bag are different questions.

+0.7pp

The rain hangover is real

A "recent rain history" block — how wet the last week was, how deep into a wet spell you are — was the one reviewer suggestion that added out-of-sample skill. Best day: dry, but recently wet.

~2×

Nothing beats the opener

Opening weekend nearly doubles the per-hunter rate all by itself, with every condition held fixed — the largest single effect in the model, bigger than any weather dial.

~28%

The ceiling

Oracle tests that cheat with hindsight put the practical limit near +28% skill at day scale. The engine already sits within about a point of it; what remains is mostly day noise no honest forecast can see.

×34

Every claim tested blind

Every number here is scored walk-forward: the model predicts 34 straight seasons it has never seen — 47,028 hunt-days — refit at each step using only the past. No hindsight, anywhere.

Why is the paper's number so much smaller than the engine's? Because they play by different rules, and the paper's rules are the point. This article only claims what survives pre-registration, full uncertainty accounting, and a frozen covariate list — the standard that makes a result scientifically defensible rather than merely profitable. The production engine (part 2) is free to use anything that earns its keep in backtests — each refuge's own recent form, the weekly hunting-pressure rhythm, thirty tuned trust dials — none of which belongs in a hypothesis test. Practically, +6.3% means the science is decent and honest at the single-day level. I originally assumed the edge would compound as you aggregate — predict a week or a season average and the daily noise washes out — so I measured it. Half true: everyone's errors shrink with aggregation, including the naive guess's, and whether the edge survives depends on tracking slow season-scale drift. The engine's does — spectacularly, reaching +43% on season averages — while the paper's static model hands its edge back. That contrast turned out to be one of the most instructive results in the whole project.

And the honest fine print, up front: every effect is associational, not causal; the response is birds bagged per hunter-day, never bird abundance; and every claim below carries a bootstrap confidence interval computed by resampling whole seasons. If a number doesn't have an interval, I don't trust it, and neither should you.

Part IIThe plain-English tour · no stats background needed

The walkthrough

The question, stated carefully

Every duck camp has a theory: hunt the fronts, skip the full moon, pray for wind. The raw averages in my first article couldn't referee those theories, because everything is tangled — rainy days cluster mid-season, cold fronts are also windy, and the opener falls on a warm clear weekend every single year. A model's job is to untangle: hold everything else fixed, and ask what each condition is worth on its own.

So I ask two separate questions, and the distinction is the whole paper. One: which conditions are associated with success in the historical record? Two: does knowing them let you predict a season the model has never seen? Those sound like the same question. They are not — and the gap between them is where most confident hunting-forecast claims go to die.

How to test a crystal ball honestly

Here's the trap most analyses fall into: fit a model on all your data, notice it fits well, and declare victory. Of course it fits — it already saw the answers. The honest test is called a walk-forward: stand in 1992 knowing only 1987–1991, predict every hunt-day of 1992, then step forward and repeat, season by season, through 2025. The model predicts 34 straight seasons blind, 47,028 hunt-days in all.

A detail worth being precise about: after each season is scored, that season's actual results join the training data and the entire model is refit from scratch — coefficients, refuge intercepts, even the climatology baseline it's graded against. So the latest season always informs the next prediction. But there is no explicit error correction: the model never inspects its own misses and adjusts for them, and it never sees a single day of the season it's currently predicting — even if it's badly wrong on opening weekend, it can't course-correct by January. It learns only in the statistical sense: more truth in, better fit out.

And "predict" has to beat something. The baseline is the smartest naive guess available — the day-of-season climatology: for this refuge, around this date, what has the average been in prior years? That baseline already knows about openers and Slowvember. Whatever the model earns above it is genuine skill, credited only for knowing about conditions: the weather, the moon, the calendar's fine structure. Remember the aim: not the exact bag — beat climatology, blind, season after season.

01

The scoreboard: 34 seasons predicted blind

Walk-forward skill vs the day-of-season climatology, by target season. Hover a bar for detail.

Each bar is one season the model had never seen, scored as improvement in hunter-weighted error over the climatology baseline. Overall: +6.3% (95% CI +3.4, +9.2). Positive in most seasons; the losses cluster in data-poor early folds and genuinely weird years. That is what real, modest predictive skill looks like — not a hot streak.

Plus six percent doesn't sound like much — until you remember the baseline already knows the calendar, and that most published hunting-weather claims have never survived a single blind season.

What moves the needle

Inside the model, every condition gets a coefficient: how much success shifts when that dial moves, with everything else held fixed. This is where the raw record's flat, confusing weather averages finally sharpen into shape. Wind helps — genuinely, not just because windy days fall in the good weeks. Rain on the day hurts. Cold days beat warm days once you stop letting the calendar take credit for them. A bright moon quietly taxes the morning flight. And the opener towers over everything: nearly a doubling of the per-hunter rate, all by itself.

02

Every dial in the model, with its uncertainty

Posterior effect on the per-hunter rate, with 95% credible intervals. Hover for the plain-English read.

Effects are per standard deviation of each condition (or vs its reference category), on the log rate scale, converted here to % change in birds per hunter. Green: helps. Red: hurts. The water block (amber) is the paradox — real in-sample association, no out-of-sample value (Figure 03). Intervals crossing zero mean the data can't pin the direction down.

Two raw-record puzzles from the first article now resolve. The raw averages said a cold-front day was indistinguishable from a steady one; the model agrees the front itself is nearly worthless once its wind and cold are counted separately — the front was never magic, its ingredients were. And the raw record's "clear beats rain" finding survives: rain on the hunt day is a real negative, duck-camp lore notwithstanding. The lore isn't entirely wrong, though — it was just pointing at the wrong day. Hold that thought for the rain hangover.

The water paradox

Now the result that makes this a paper instead of a blog post. California's water year — drought indices, river flow, reservoir storage, flooded rice acreage — shows a clear statistical association with hunter success in the seasons the model was fit on. The mechanism reads beautifully: water moves birds around the landscape, concentrating them on or off the refuges. H1a, supported, p < 0.0001. Champagne?

No. Hand those same water indices to the walk-forward — ask them to help predict a season the model never saw — and they don't just fail to help. They make prediction slightly worse.

03

Association is not prediction

Walk-forward skill with and without each covariate block, on identical rows. Whiskers: 95% bootstrap CIs.

Adding the water block drops out-of-sample skill by 5.2 points (95% CI −13.2 to −0.4 — reliably negative). Adding abundance indices changes nothing (CI spans zero: the pre-registered null). The honest headline: neither block beats timing + weather where it counts.

How can both be true? Resolution. The water indices exist at region-by-season grain — one number for the whole Sacramento Valley's year. When a held-out season's water sits outside anything in training, the model extrapolates that one number across every hunt-day in the region and eats a season-sized loss (1996–97 was the bloodbath). Within seasons it saw, the association is real; projected onto seasons it didn't, the coarseness is fatal. The fix isn't to abandon the mechanism — it's finer data: refuge-level flooding timing instead of valley-level annual indices. That's the collaboration pitch at the end of the paper.

More ducks ≠ more ducks in your bag

The same test, run on bird-count indices — the mid-winter aerial survey and the pre-season breeding flight — comes back empty, exactly as the analysis plan predicted before anyone fit anything. A big flight year does not raise the per-hunter take. Your bag is set by whether birds are workable — where they sit, how they move, whether weather shuffles them — not by how many exist in the flyway. It's the statistical version of something every refuge regular already knows: the marsh can be covered up in birds and still shoot two ducks a gun.

The rain hangover

After two expert reviewers — J. Coslovich and C. Overton — read the paper, they pushed on one thing: weather isn't just the day's snapshot — duration matters. Prolonged weather should act differently from fresh weather. I froze a protocol and tested five re-encodings of the weather, blind, on the same 47,028 held-out days. One cleared the bar, and it was theirs, not mine.

04

Five ideas, one survivor

Change in walk-forward skill vs the base model, paired on identical rows. Whiskers: 95% bootstrap CIs.

The precipitation-persistence block (V3) — prior-week rainfall, days since rain, position in a wet spell — adds +0.7 points of skill (CI +0.4 to +1.0), lifting the model to +7.0%. Re-encoding temperature as departure-from-normal (V2) actively hurts: raw cold carries real between-season signal that anomalies throw away. Bundling everything (V4) shows why you test blocks separately — V2's damage cancels V3's gain.

The persistence terms sketch a shape no single-day variable could: success peaks on days that are dry but recently wet, sags through a prolonged soak, and sags again deep into a drought spell. A wet week floods the off-refuge landscape and scatters birds; the day the storm quits, they're reshuffled and workable; weeks of nothing lets everything go stale. The raw record hinted at this — "clear beats rain" — but only the model can see that the best clear day is the one right after the rain. That's the answer to the duck-camp lore: the storm does matter. Just hunt the day it leaves, not the day it arrives.

Part IIIThe full spec · for people who read appendices first

The deep dive

The model, precisely

One row per refuge × hunt-day (50,154 rows; 41 areas, 5 regions, seasons 1987–2025). The response is the bagged count with hunters as exposure, so coefficients live on the log per-hunter-rate scale. Over-dispersion is severe (α = 4.34, HDI 4.26–4.42), so negative binomial, not Poisson:

harvest ~ NegBinomial(mu, alpha)
log(mu) = log(hunters)            # exposure offset — everything is a *rate*
        + Xβ                      # standardized covariates (timing, weather, water, ENSO)
        + u_refuge + u_season     # random intercepts: 41 areas, 39 seasons

u_refuge ~ Normal(0, σ_r)         σ_r = 0.23   # between-area spread
u_season ~ Normal(0, σ_s)         σ_s = 0.16   # between-season spread
β ~ Normal(0, 1)   α, σ ~ HalfNormal          # weakly-informative priors

Fit in PyMC via bambi on the 28,249 water-complete hunt-days: 4 chains, 1,000 tune + 1,000 draws, max R̂ = 1.010. Covariates: a natural cubic day-of-season spline (4 df), opener weekend, day-of-week, max temperature, 3-day cold-front drop, max wind, wind quadrant, barometric tendency, sky category, moon illumination, ENSO ONI, and the four-term water block (PDSI, flow, reservoir, rice). The random intercepts are not decoration — hunt-days nest inside refuges and seasons, and ignoring that is pseudoreplication that would shrink every interval dishonestly.

Two engines, one structure

The Bayesian GLMM is the inferential engine: posteriors, credible intervals, the H1a joint test. The walk-forward is a different computational animal — 34 folds × model variants × 2,000 bootstrap draws — so the out-of-sample loop runs the pre-registered frequentist cross-check: an NB GLM with identical fixed-effect structure, handling the refuge level with empirical-Bayes shrunk offsets (each refuge's log-rate deviation shrunk toward the pool by ~500 hunter-equivalents — the frequentist analog of the random intercept, which also prevents separation blowups on data-poor early folds). A fold's fit is rejected if it predicts implausible rates, stepping down NB → Poisson → pooled-mean; data-poor folds predict conservatively instead of exploding.

The baseline and the metric

The climatology baseline is per-refuge, hunter-weighted, pooled over ±14 days of day-of-season, trained only on prior seasons — the strongest naive forecast I could construct. Skill = 1 − MAEmodel/MAEclimatology, hunter-weighted; I also track NB log predictive density, which scores the full predictive distribution rather than the point. Every skill number and every contrast carries a season-block bootstrap 95% CI: resample whole seasons, 2,000 draws, because hunt-days within a season are correlated and row-level bootstraps would fake precision.

Collinearity, handled before the fact

The water covariates co-move, so the plan pre-committed to interpreting them jointly. The pre-fit diagnostic came back milder than feared — max VIF 2.36, and rice acreage nearly orthogonal to the hydrology terms (r = 0.10–0.34). The PCA sensitivity check told the same story from the other side: the full block is decisively significant in-sample (LRT p ≈ 2×10−10) while a 2-component reduction is not (p ≈ 0.79) — rice loads on the components the reduction throws away. Keeping the terms individual was the right pre-registered call.

Calibration — where the model is honest and where it isn't

05

Out-of-sample calibration by predicted decile

Held-out hunt-days binned by predicted rate; observed vs predicted, hunter-weighted.

Point calibration is excellent: when the model says 0.7, reality says 0.7; when it says 3.2, reality says 3.3 — across all ten deciles of 47,028 blind predictions. The known weakness is interval width: the NB's 90% predictive intervals cover ~100% of outcomes because the estimated over-dispersion is huge. The mean is trustworthy; the error bars are conservative. Flagged in the paper as the next refinement.

Exploratory discipline

Everything confirmatory was frozen in ANALYSIS_PLAN.md before any model was fit — hypotheses, covariates, metrics, folds, even the bootstrap seed protocol. When reviewer suggestions arrived post-hoc, they went through a second frozen protocol (EXPLORATORY_PROTOCOL.md): five variants specified before fitting, all five reported regardless of outcome, judged only by paired out-of-sample contrast. That's why the V2 failure is printed as prominently as the V3 win — the graveyard is part of the result. The persistence block's clean promotion path is to pre-register it as confirmatory and test on the next unseen season; meanwhile it has already shipped to the production forecast engine (that story is the next article).

Threats to validity, without flinching

Where this sits in the literature

The full literature memo (17 sources) maps the field on two axes: what scale the response lives at (annual flyway populations vs a hunter's day), and how claims are validated (in-sample association vs out-of-sample prediction). The map below is the lit review's own figure, redrawn live — hover any study for what it found and why it's placed where it is.

06

A map of the waterfowl-harvest literature

Response scale vs validation standard, 17 sources. Hover a dot for the study's story.

Most refereed work lives in the lower-left: population-scale and associational, built for regulation rather than forecasting. The shaded upper-right quadrant — per-hunter, daily, validated out-of-sample — is nearly empty; its only neighbors are two unrefereed commercial analyses. That corner is where this paper sits, which is both the opportunity and the reason there was no yardstick to borrow.

Three landmarks orient everything:

Next steps: the dynamic refuge intercept — planned, then tested

One upgrade seemed obvious enough that I ran it before someone could suggest it in review. Right now each refuge gets a single random intercept — a fixed "how good is this refuge, on average" number, shrunk toward the pool and re-estimated after every season. That refit drags all of history along equally: a refuge whose habitat changed in 2018 is still being predicted partly by its 1990s self. But refuge quality isn't fixed. Water allocations shift, ponds get reconfigured, vegetation cycles turn over, management priorities change — and a static intercept can only chase that drift slowly, one full-history refit at a time.

The idea: let each refuge's baseline move. I implemented it as exponential forgetting — each training season's contribution to a refuge's offset is discounted by a retention factor ρ per season of age (the steady-state form of a random-walk intercept), with the same shrinkage guarding against overreacting to one loud year. Two things made me expect it to pay. First, the walk-forward misses show persistent, refuge-specific runs of over- or under-prediction that no weather covariate explains. Second, the production engine already ran the experiment by accident: its recent-form factor (how a refuge shot lately versus its own norm) earned the single largest skill gain of any factor it tested — and much of that gain looks like baseline drift in disguise.

The discipline was the same as everything else here: a grid of three forgetting speeds (ρ = 0.95, 0.90, 0.80 — half-lives of roughly 13.5, 6.6, and 3.1 seasons) frozen in the exploratory protocol before any variant was fit, scored by the identical walk-forward on identical rows against the identical climatology baseline, paired season-block bootstrap, all rows reported win or lose. And the honest result is: it didn't win.

Where the remaining skill lives — three quick experiments

With the drift question settled, I ran three cheap experiments to map what's left. All three use the committed out-of-sample artifacts, so nothing was refit with hindsight.

1. The ceiling. I scored oracle predictors that cheat. A season oracle that knows every refuge-season's final average comes in at −1.3% — worse than climatology, because a season's level without its day texture is worthless. A local oracle that knows the neighboring ±2 hunt-days' actual results reaches only +14.9% — barely half the engine's +26.7%, because it averages across the weekly pressure rhythm instead of riding it. And granting the engine a hindsight correction from its own neighbors' misses adds just +1.3 points. Translation: season-level foresight is worthless, state-style signals (form, drift, local habitat level) are nearly tapped out, and the engine's edge was always the conditions-and-calendar texture. What remains beyond ~28% is mostly day noise no ex-ante forecast can see.

2. The ensemble. This paper's GLMM and the production engine err differently — one weighs all history equally, one is recency-tilted — so I blended their predictions on the 42,813 hunt-days they both scored. A fixed 85/15 engine/paper mix lifts skill from +26.7% to +27.3%, and unlike everything else this week it passed the season-split cross-validation gate in both directions. Small, cheap, real: the humble white-paper model earns its 15% by remembering what the recency-tilted engine has deliberately half-forgotten.

3. The structure test. The paper's model is additive on the log scale — interactions exist only where hand-built. A gradient-boosted tree challenger, fit on the same walk-forward folds with the same information and scored on identical rows against the identical baseline, reaches +14.0% skill against the GLM's +6.3% — a paired gain of +7.7 points (95% CI +4.5 to +10.7). That's the largest signal found since this paper shipped, and it's a finding about structure, not data: the covariates already carry roughly twice the skill the linear model extracts, in interactions the GLM can't represent. It's also, in hindsight, a big part of why the production engine wins — its per-refuge condition tables are effectively hand-built interactions. The obvious next chapter: understand which interactions the trees found, and either port them into an interpretable model or admit the challenger into production through a pre-registered test on the next unseen season.

Coda: does the challenger help the app? Measured directly, yes. On the engine's own 42,813 backtest days the GBM alone scores +18.0% — no match for the engine's +26.7% — but blended 75/25 engine/GBM the backtest reaches +28.0%, a +1.3-point gain that clears the season-split cross-validation gate in both directions and displaces the 85/15 paper blend entirely (in a three-way mix, the GLMM's weight goes to zero — the trees subsume it). And +28.0% is almost exactly the hindsight ceiling from experiment 1: the blend reaches by conditions-texture what the oracle could only reach by cheating. The two experiments triangulate the same boundary from opposite sides, which is what a real limit looks like.

The autopsy — which interactions did the trees find?

I then took the challenger apart, three ways, on the same walk-forward folds. First, structure: an additive gradient-boosted model — stumps only, so no tree can represent an interaction — scores +8.9%. So of the challenger's +7.7-point gap over the GLM, roughly a third is honest non-linearity and two-thirds is pure interaction (+5.1 points, 95% CI +2.7 to +7.3, paired). Second, ablation: refit the full model with one feature family removed and see what dies. Removing refuge identity costs −12.9 points; removing the day-of-week column costs −4.3; removing the entire weather block costs only −2.5 (CI just barely excluding zero). Third, measurement: Friedman's H statistic on the fitted model puts the interaction share of joint effect variance at 0.156 for refuge × day-of-season — five times any other pair tested (refuge × weekday 0.027, day-of-season × weekday 0.026; every weather pairing ≤ 0.010).

Three methods, one verdict: the signal the paper's model leaves on the table is overwhelmingly refuge-specific seasonal shape — Delevan's December is not Wister's December — with a side of refuge-specific weekday texture, and only a garnish of weather interaction. The GLM can't represent it: one global day-of-season spline plus a scalar per-refuge offset forces every refuge onto the same curve at a different height. And this, in hindsight, is the cleanest explanation of the paper-to-production gap yet: the engine's per-refuge day-of-season anchor is the slug × day-of-season interaction, hand-built — the single structure the trees rediscovered as their biggest find. The paper's next methodological step writes itself: per-refuge seasonal splines (partially pooled so thin refuges borrow strength), pre-registered, tested on the next unseen season.

Executing the verdict — the per-refuge seasonal shape

So I executed it (exploratory addendum 3, spec frozen before fitting). The change is one line of model structure: the empirical-Bayes refuge offset generalizes from a scalar to a shape. Instead of "Delevan runs 40% above the pool, every day," the offset becomes "Delevan on day 45 runs at its own training-history rate for days 31–59, shrunk toward its scalar level when the window is thin" — partial pooling by hunter-equivalents, so data-rich refuges get their own curve and thin ones fall back gracefully to the confirmatory model. The covariates then fit only what conditions add on top. Nothing else moves: same engine, same folds, same 47,028 out-of-sample hunt-days, same baseline, paired contrasts.

Three prespecified variants (window ±14 or ±21 days, shrinkage 100 or 300 hunter-equivalents), and all three win decisively. The best, ±14 days at K=100, doubles the paper's skill: +6.3% → +12.6% (paired gain +6.3 points, 95% CI +4.7 to +8.0) — in bird terms, the typical miss falls from 0.71 to 0.67 birds per hunter against the baseline's 0.75 — with the probability forecasts improving too. And the punchline: the paired contrast against the gradient-boosted challenger is −1.4 points with a CI spanning zero — the interpretable model is now statistically indistinguishable from the black box. The trees' +7.7-point edge was never really about a thousand trees; it was one interaction, and once the GLM is allowed to express it, the gap closes to noise.

That's the week's cleanest result, and it lands the whole arc in one sentence: the production engine wins because its per-refuge day-of-season anchor was this structure all along; the black-box challenger won by rediscovering it; and the paper's model, granted the same single idea, catches the challenger while keeping every coefficient readable. What remains before this graduates from exploratory to headline is the discipline this paper is built on — a fresh pre-registration, tested on the 2026–27 season the model has never seen.

The aggregation dividend — measured, and only half true

One more claim needed testing, because I'd been repeating it without a number attached: "errors shrink as you aggregate, so a modest daily edge compounds when you predict a week or a season average." Plausible, intuitive — and checkable in an afternoon, since it needs no new model, just re-scoring the committed out-of-sample predictions at coarser grains (exploratory addendum 4, spec frozen before computing). I summed model, actual, and climatology counts within each refuge-week and refuge-season — aggregating the forecast and its baseline identically, so neither side gets an unfair average — and re-scored skill against the identically aggregated climatology, bootstrap CIs and all.

The first half of the claim is true for everyone: the typical miss falls from 0.71 birds per hunter at day scale to 0.35 at season scale for the paper's model, and from 0.59 to 0.17 for the engine. Aggregation absolutely makes forecasts more accurate. The second half — that the edge over climatology compounds — is where it gets interesting. The paper's confirmatory model hands its edge back as you zoom out: +5.8% at day scale (scored on the per-hunter rate, hence the whisker of difference from the headline), −3.7% at week scale, −9.6% at season scale — worse than the naive guess, CIs excluding zero. The shape-offset upgrade keeps a real week-scale edge (+4.6%) but ties at season scale (−4.6%, CI spanning zero). The engine inverts the story entirely: +26.7% → +31.9% → +43.1%, day to week to season, with a season-average miss of just 0.17 birds per hunter against climatology's 0.30.

The mechanism, once you see it, is obvious. A daily conditions edge — wind, fronts, moon — is exactly that: daily. Average over a season and the weather cancels out of both the forecast and the truth; what's left of the season-average question is "is this refuge running hot or cold this year?" — season-scale drift, the one thing a static intercept structurally cannot know and the very signal the engine's recent-form factor exists to track. So the aggregation dividend is real, but it isn't free: it's earned by whoever tracks slow state. The engine earns it; the paper, by design, doesn't. Which retroactively explains the dynamic-intercept experiment above — the drift signal it hunted at day scale, where it's a rounding error, is the entire game at season scale.

Every data source, and whether it earned skill

Across the paper and the production engine, I tested every dataset I could defensibly attach to a refuge hunt-day. This is the complete list — the wins, the nulls, and the graveyard — because a forecast you should trust tells you what it threw away. "Paper" verdicts follow the pre-registered rules of this article; "engine" verdicts are the production model's tuned trust weight λ (0 = ignored entirely, 1 = fully trusted), the part 2 story. Each source name links to the upstream provider; the exact SHA-256-manifested snapshots used by the paper are committed in the replication repo.

Tested in the white paper

SourceWhat it isAdded skill?
CDFW hunt-day reports50,154 refuge×day records of birds and hunters, 1987–2025, from check stationsThe response itself — what everything else tries to predict
The calendar (day-of-season spline, opener, day-of-week)Pure date structure, no external dataYes — the backbone; engine trusts day-of-week at λ 1.0, opener at 0.5
ERA5 daily weather (Open-Meteo)Max temp, wind, sky category, precipitation per refuge-dayYes — the core of the +6.3%; rain and sky strongest
Cold-front drop (derived from ERA5)Day's high vs the prior 3 days' meanSmall but real; engine λ 0.2
Wind quadrant (derived from ERA5)Dominant direction N/E/S/WSmall but real (west winds hurt); engine λ 0.2
Barometric tendency (derived from ERA5)Sea-level pressure vs the prior 3 daysSmall; falling pressure helps; engine λ 0.2
Moon illumination (computed; cross-checkable at USNO)Astronomical fraction litYes — bright moon reliably hurts; engine λ 0.6
ENSO (NOAA ONI)El Niño / La Niña index, season-constantMarginal; engine λ 0.1
Palmer drought index (NOAA nClimDiv)Climate-division PDSIThe water paradox: associated in-sample, degrades out-of-sample; engine λ 0.1
River flow (USGS NWIS)Oct–Jan discharge index near each refugeSame paradox in the paper; engine salvages λ 0.2
Reservoir storage (CDEC)Statewide managed-water indexSame paradox; the engine's best water term at λ 0.3
Flooded rice acreage (USDA NASS)Harvested rice by district — the flooded food baseDecisively significant in-sample, no OOS skill; engine λ 0.1
Mid-winter aerial survey (CDFW)January duck counts by regionNull — abundance ≠ success; engine λ 0
Breeding-flight index (USFWS WBPHS, via the Pacific Flyway Data Book)Pre-season Pacific-source breeding totalNull; engine λ 0 (a small geese-only 0.2)
Precipitation persistence (derived from ERA5)Prior-7-day rain, wet-spell position, days since rainYes — the +0.7pp exploratory win; engine λ 0.3 / 0.2 / 0.1

Explored in the production engine (part 2 territory)

SourceWhat it isAdded skill?
Recent form (derived from the CDFW record)The refuge's last 3 hunt-days vs its own normThe single biggest production gain (+19.9% → +26.7% at refuge level); λ 0.5–0.6
Migration pushes (up-flyway ERA5)Cold drops at Klamath Basin, Great Salt Lake; Y-K Delta freeze-upKlamath λ 0.1, Great Salt Lake λ 0.2; the Alaska freeze signal earned nothing (λ 0)
Rest-day cycle (derived from the CDFW record)Days since the refuge last huntedReal; refuge-only λ 0.2
Weather cell & wind×speed buckets (derived from ERA5)Exact sky×wind×temp match; joint wind direction×speedMarginal; λ 0.1 each (the cell earns full trust for geese)
Water-year anomaly (derived from ERA5)Cumulative precip since Aug 1 vs the refuge's historyλ 0 — redundant with the daily rain block
Satellite open water (JRC Global Surface Water)Monthly flooded-area fractionλ 0 — fog-obscured valley flooding undercounts; flow/reservoir proxies beat it
Refuge water deliveries (USBR)Level-4 refuge water, 1993–2008λ 0 — short record, redundant with drought/reservoir
Sierra snowpack (CDEC Apr-1 SWE)Snow water equivalent feeding summer storageλ 0 — a real lead signal, but redundant with the reservoirs it fills
Yolo Bypass inundation (DWR/CDEC)Floodplain activationλ 0 — only ~11 well-covered seasons; not enough record to earn trust
Neighbor spillover (derived from the CDFW record)Nearby refuges running hot in the last 4 daysλ 0 — the refuge's own recent form already carries it
eBird Status & Trends (Cornell)Weekly relative-abundance rasters, 21 species, 9 kmTested as the seasonal anchor: −10.3% — bird presence peaks weeks before harvest rate. Kept as a spatial prior for new properties
Dawn-hour weather (hourly ERA5, 6–10am)Shooting-window temp/wind/sky/precipBest factor +0.09pp held-out — redundant with the daily aggregates; parked

The pattern across 27 sources is the paper's thesis in miniature: what the birds are doing right now (weather, rain history, recent form, pressure rhythms) predicts; how many birds exist somewhere (counts, surveys, satellite water, snowpack) does not — either null or redundant with a more direct read.

Glossary

The load-bearing terms, in the order the article leans on them.

Skill
The fraction by which the model's forecast error beats the climatology baseline: skill = 1 − MAEmodel / MAEclimatology. 0% = you know nothing the calendar doesn't; +6.3% = errors are 6.3% smaller; +26.7% = errors are about a quarter smaller. Borrowed from weather forecasting, where beating climatology is the definition of a forecast being worth anything. Note the goal it encodes: not "predict the exact bag," but beat the smartest naive guess.
Climatology (the baseline)
The smartest guess that uses no conditions at all: for this refuge, within ±14 days of this date, what has the hunter-weighted average been in prior seasons? It already knows about openers and mid-season slumps, which is exactly why beating it is hard — and meaningful.
Walk-forward validation
The honest test. Train only on seasons before year Y, predict every hunt-day of year Y blind, score it, step forward, repeat — 34 consecutive seasons, 47,028 held-out hunt-days. The model never sees the answer before it guesses.
Out-of-sample (OOS)
Data the model has never seen during fitting. A pattern that holds in-sample but dies out-of-sample (the water block here) is a description of the past, not knowledge about the future.
MAE (hunter-weighted)
Mean absolute error: the average size of the miss, in birds. Hunter-weighted means a 500-hunter day counts 500 times more than a one-hunter day, so the score reflects the experience of actual hunters rather than of rows in a table.
Confidence / credible intervals
The "give or take" on every number. The paper's skill CI comes from a season-block bootstrap: resample whole seasons 2,000 times, because hunt-days within a season move together and pretending they're independent would fake precision. Bayesian effect estimates carry credible intervals — the range holding 95% of the posterior's belief.
Pre-registration
Freezing hypotheses, covariates, metrics, and folds in a written plan before fitting anything, so you can't quietly shop for the analysis that flatters you. Post-hoc ideas (the reviewer variants) went through a second frozen protocol and are labeled exploratory. This discipline is why the paper's headline stays at +6.3% while the unconstrained engine reaches +20–27%.
Negative binomial GLMM
The paper's model: a generalized linear mixed model where the bagged count follows a negative binomial distribution (a Poisson that admits real-world mess — birds arrive in flocks, days blow up or die) with hunters as exposure, so effects live on the per-hunter rate scale.
Random intercepts
Each refuge and each season gets its own baseline level, estimated with partial pooling. Without them, 47k correlated hunt-days would masquerade as 47k independent observations (pseudoreplication) and every interval would be dishonestly narrow.
Empirical-Bayes shrinkage
The walk-forward's frequentist stand-in for random intercepts: each refuge's deviation is pulled toward the pooled average in proportion to how little data supports it (~500 hunter-equivalents), so data-poor early folds predict conservatively instead of wildly.
Calibration
Whether predicted numbers mean what they say: when the model says 0.7 birds/hunter, does reality average 0.7? (Here: yes, across all ten deciles.) Distinct from skill — a model can be well-calibrated and useless, or skillful and biased.
VIF (variance inflation factor)
A collinearity diagnostic: how much a covariate's uncertainty inflates because other covariates carry the same information. Max here was 2.36 — mild; below the conventional worry threshold of 5.
λ (lambda, the trust dial)
Part 2 vocabulary: the production engine multiplies climatology by per-factor adjustments, each raised to a tuned exponent λ between 0 (ignore the factor entirely) and 1 (trust its historical rate fully). Tuned by walk-forward backtest, per region and per target. The data-source table above reports each factor's λ.

Sources & reproducibility

Everything is public and replicable from one repository: gitlab.com/duckchugger/refuge-harvest-model — the frozen analysis plan, the SHA-256-manifested raw data snapshot (CDFW harvest records, ERA5 weather, NOAA/USGS/CDEC/NASS water and climate indices), the tidy 50,154-row analysis table, all model and validation code, and the built PDFs of the analysis report, literature review, exploratory addendum, technical bridge, and pre-registered plan. make all reruns snapshot → dataset → models → figures end to end with pinned dependencies and fixed seeds.

Every figure on this page is generated from the repo's committed artifacts by this article's own build script, so the numbers here cannot drift from the paper's. The raw-record companion piece is Thirty-Nine Seasons of Duck Days, and the full dark-mode data explorer lives at theduckchugger.com/duckdata.

Acknowledgments

J. Coslovich and C. Overton — thank you for taking a look and for your suggestions!