feat!: N-team outcome prediction with draw mass, replacing the 2-team panic

`predict_outcome` asserted `teams.len() == 2` and returned `[p, 1 - p]`,
allocating no probability to a draw even with `p_draw > 0`. For a
draw-enabled model the numbers were simply wrong, at any team count.

It now returns `Result<Prediction, InferenceError>` and supports N teams.

Two algorithms, both deterministic:

- Who finishes first. Performances are independent Gaussians, so this
  separates into a one-dimensional integral per team rather than a
  multivariate orthant probability. Adaptive Gauss-Kronrod evaluates it
  to ~1e-15, matching the exact two-team closed form.
- A specific finishing order. The factor graph only constrains
  rank-adjacent teams, so a full order is a chain of local constraints,
  not a general orthant integral. That chain collapses into a sequential
  recursion over cumulative integrals: O(teams * grid) per order.

Fixed-node Gauss-Hermite is the obvious tool for the first and is a trap:
when a rival's sigma is small the CDF product becomes a step narrower
than the node spacing, and the nodes step over it. Measured 4.4e-4 off
the closed form on a mildly skewed matchup and 1.7e-2 on a small-sigma
one, while still returning something that looks like a probability.
Adaptive refinement is what makes that case safe, and
`win_probabilities_survive_a_rival_with_a_tiny_sigma` pins it down.

The acceptance test is an identity rather than a golden: the outcome
space is exhaustive and disjoint, so the probabilities sum to one. Any
drift is integration error and nothing else. Gauss-Hermite failed it at
4.4e-4; this holds to ~1e-9.

Also from #21: unknown keys are now reported rather than dropped, so a
team of strangers can no longer produce a confident-looking prediction.
`predict_quality` returns `Result` for the same reason.

BREAKING CHANGE: `predict_outcome` returns `Result<Prediction, _>`
instead of `Vec<f64>`; `predict_quality` returns `Result<f64, _>`.

Refs #21, #39

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
This commit is contained in:
2026-09-07 14:55:11 +02:00
co-authored by Claude Opus 5
parent 87fca8dcca
commit bb2a845882
8 changed files with 1513 additions and 50 deletions
+37
View File
@@ -40,6 +40,24 @@ pub enum InferenceError {
},
/// Negative precision: a Gaussian with `pi < 0` slipped into an API call.
NegativePrecision { pi: f64 },
/// A prediction referenced a key the history has no skill for.
///
/// Reported rather than skipped: dropping unknown keys turns a team of
/// strangers into a confident-looking probability about nobody.
UnknownKey { team: usize, member: usize },
/// A prediction was given a team with no members.
EmptyTeam { team: usize },
/// Fewer than two teams were supplied to a prediction.
NotEnoughTeams { got: usize },
/// The full outcome distribution was requested for too many teams.
///
/// Each realisation sorts into exactly one (order, tie-pattern) event, so
/// the space holds `n! * 2^(n-1)` members — 1_920 at five teams, 23_040 at
/// six, 322_560 at seven. Past `max` this stops being something to
/// enumerate on a caller's behalf; ask for individual rankings with
/// `predict_ranking`, or for `predict_win_probabilities`, both of which
/// stay cheap at any team count.
TooManyTeams { got: usize, max: usize },
}
impl fmt::Display for InferenceError {
@@ -90,6 +108,25 @@ impl fmt::Display for InferenceError {
Self::NegativePrecision { pi } => {
write!(f, "precision must be non-negative; got {pi}")
}
Self::UnknownKey { team, member } => {
write!(
f,
"team {team}, member {member}: no skill recorded for this key"
)
}
Self::EmptyTeam { team } => {
write!(f, "team {team} has no members")
}
Self::NotEnoughTeams { got } => {
write!(f, "prediction needs at least 2 teams, got {got}")
}
Self::TooManyTeams { got, max } => {
write!(
f,
"the outcome distribution over {got} teams is too large to enumerate (limit {max}); \
use predict_ranking or predict_win_probabilities instead"
)
}
}
}
}