Six of fifteen variants carried a `&'static str` discriminator, about
thirty magic strings between them, and the only thing a caller could do
with one was print it. Four new enums replace them:
Parameter 13 variants, replacing 9 strings in InvalidParameter
Shape 4 variants, replacing 10 in MismatchedShape
OutcomeKind 2 variants, replacing WrongOutcomeKind's three fields
CompetitorField 2 variants, replacing ConflictingCompetitorConfig's
`InvalidProbability` folds into `InvalidParameter` as
`Parameter::PDraw`. It was a bespoke variant for one scalar while every
other scalar shared `InvalidParameter`, and it omitted the parameter
name — so the same parameter had two mechanisms.
`JointUnavailable { reason: &'static str }` splits into `EmptyHistory`,
`JointRequiresScoredEvents` and `NotPositiveDefinite`. The three are
conditions a caller branches on differently — add events, use
`predict_win_probabilities`, or reconsider the priors — and telling them
apart used to mean string-matching English. One test already proved the
distinction was load-bearing: the blanket conversion mapped the
empty-history case onto the ranked one and `an_empty_history_has_no_joint`
caught it immediately.
`NonFiniteResult` splits into `NonFiniteStep { context, step }` and
`NonFiniteSkill { mu, sigma }`. One `step: (f64, f64)` field was
carrying a sweep step from `converge` and a skill's own moments from a
prediction — two situations in one variant, and a field name that could
only be right for one of them.
`InvalidParameter { name: "beta with point-mass skills" }` becomes
`NoPerformanceVariance`. It was never a parameter out of range: both
values are individually valid and it is their combination that leaves
nothing varying.
Three `Display` impls did not meet the standard the others set, and the
typed data is what makes fixing them possible:
before drift variance is invalid: NaN
after drift variance must be finite and non-negative (got NaN)
before kinds: expected length 3, got 2
after the outcome describes a different number of teams than the
event has: expected 3, got 2
before Game::ranked: expected Outcome::Ranked, got Outcome::Scored
after expected Outcome::Ranked, got Outcome::Scored; call
Game::scored for a scored outcome
`Parameter::range()` states each parameter's actual bounds, which no
`&'static str` name could have. `error::message_tests` renders every one
and asserts each is a sentence rather than a label, and that the three
above now carry a range or a next step.
The four internal `MismatchedShape` kinds — `results`, `times`, `kinds`,
and the weights array — collapse to `Shape::Internal`, whose `Display`
says plainly that reaching it is a bug in this crate. They are checks on
`add_events_with_prior`'s own parallel arrays and are unreachable
through the public API; they stay checked rather than becoming
`debug_assert!`s, because release is where this crate's defects hide.
Closes #74.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
218 lines
8.0 KiB
Rust
218 lines
8.0 KiB
Rust
//! Inference must report numerical breakdown rather than call it convergence.
|
|
//!
|
|
//! The boundary rejects inputs that are *not numbers*, but finite inputs can
|
|
//! still overflow during inference — `beta.powi(2)` at 1e300 is infinite, and
|
|
//! infinity minus infinity is NaN. `NonFiniteResult` is the guard for that, and
|
|
//! it matters because the alternative is silent: NaN fails every comparison, so
|
|
//! a naive `step < epsilon` check reads a NaN step as *converged*.
|
|
//!
|
|
//! That is why the crate has `step_converged` / `step_is_finite` rather than
|
|
//! `!tuple_gt(..)`. These tests pin the guard from outside.
|
|
|
|
use smallvec::smallvec;
|
|
use trueskill_tt::{
|
|
ConstantDrift, Event, Gaussian, History, InferenceError, Member, Outcome, Team,
|
|
};
|
|
|
|
fn scored_fit(
|
|
sigma: f64,
|
|
beta: f64,
|
|
score_sigma: f64,
|
|
scores: [f64; 2],
|
|
) -> Result<bool, InferenceError> {
|
|
let mut h = History::builder()
|
|
.mu(0.0)
|
|
.sigma(sigma)
|
|
.beta(beta)
|
|
.score_sigma(score_sigma)
|
|
.build();
|
|
h.add_events(vec![Event {
|
|
time: 1i64,
|
|
teams: smallvec![
|
|
Team::with_members([Member::new("a")]),
|
|
Team::with_members([Member::new("b")]),
|
|
],
|
|
outcome: Outcome::scores(scores),
|
|
}])?;
|
|
h.converge().map(|r| r.converged)
|
|
}
|
|
|
|
/// Every one of these is built from finite, individually legal parameters. The
|
|
/// overflow happens inside inference, which is exactly the case the boundary
|
|
/// checks cannot catch.
|
|
///
|
|
/// Matched rather than merely `is_err()`: an assertion that only checks "some
|
|
/// error" would keep passing if these started failing at the boundary for an
|
|
/// unrelated reason, and would then be testing nothing.
|
|
#[test]
|
|
fn overflow_during_inference_is_reported_not_hidden() {
|
|
let cases: [(&str, f64, f64, f64, [f64; 2]); 5] = [
|
|
("huge sigma", 1e300, 1.0, 1.0, [3.0, 1.0]),
|
|
("huge beta", 6.0, 1e300, 1.0, [3.0, 1.0]),
|
|
("tiny sigma", 1e-300, 1.0, 1.0, [3.0, 1.0]),
|
|
("tiny score_sigma", 6.0, 1.0, 1e-300, [3.0, 1.0]),
|
|
("huge scores", 6.0, 1.0, 1.0, [1e308, -1e308]),
|
|
];
|
|
|
|
for (name, sigma, beta, score_sigma, scores) in cases {
|
|
match scored_fit(sigma, beta, score_sigma, scores) {
|
|
Err(InferenceError::NonFiniteStep { context, step, .. }) => {
|
|
assert_eq!(context, "History::converge", "{name}");
|
|
assert!(
|
|
!step.0.is_finite() || !step.1.is_finite(),
|
|
"{name}: reported NonFiniteResult with a finite step {step:?}"
|
|
);
|
|
}
|
|
other => panic!("{name}: expected NonFiniteResult, got {other:?}"),
|
|
}
|
|
}
|
|
}
|
|
|
|
/// The trap the invariant exists for: NaN fails every comparison, so a naive
|
|
/// `step < epsilon` test reads a NaN step as converged. A breakdown must never
|
|
/// come back as a successful fit.
|
|
#[test]
|
|
fn a_broken_fit_is_never_reported_as_converged() {
|
|
let mut h = History::builder().build();
|
|
h.add_events(vec![Event {
|
|
time: 1i64,
|
|
teams: smallvec![
|
|
Team::with_members([Member::new("a").with_prior(Gaussian::from_ms(1e300, 1e-300))]),
|
|
Team::with_members([Member::new("b")]),
|
|
],
|
|
outcome: Outcome::winner(0, 2),
|
|
}])
|
|
.unwrap();
|
|
|
|
let err = h.converge().unwrap_err();
|
|
assert!(
|
|
matches!(err, InferenceError::NonFiniteStep { .. }),
|
|
"a breakdown must not be reported as convergence: {err:?}"
|
|
);
|
|
|
|
// `converge_partial` must not launder it into an `Ok` either — the
|
|
// permissive path is permissive about *stopping short*, not about NaN.
|
|
let mut h2 = History::builder().build();
|
|
h2.add_events(vec![Event {
|
|
time: 1i64,
|
|
teams: smallvec![
|
|
Team::with_members([Member::new("a").with_prior(Gaussian::from_ms(1e300, 1e-300))]),
|
|
Team::with_members([Member::new("b")]),
|
|
],
|
|
outcome: Outcome::winner(0, 2),
|
|
}])
|
|
.unwrap();
|
|
assert!(matches!(
|
|
h2.converge_partial().unwrap_err(),
|
|
InferenceError::NonFiniteStep { .. }
|
|
));
|
|
}
|
|
|
|
/// The neighbouring case, so the tests above cannot pass by the fit simply
|
|
/// always failing: ordinary extreme-but-workable parameters still converge.
|
|
#[test]
|
|
fn merely_extreme_parameters_still_converge() {
|
|
assert!(scored_fit(1e6, 1.0, 1.0, [3.0, 1.0]).unwrap());
|
|
assert!(scored_fit(1e-6, 1.0, 1.0, [3.0, 1.0]).unwrap());
|
|
assert!(scored_fit(6.0, 1.0, 1e6, [3.0, 1.0]).unwrap());
|
|
assert!(scored_fit(6.0, 1.0, 1.0, [1e150, -1e150]).unwrap());
|
|
}
|
|
|
|
/// A NaN in one competitor must not be masked by a healthy competitor reduced
|
|
/// after it.
|
|
///
|
|
/// The convergence step is a fold over a `HashMap`, so which competitor is
|
|
/// reduced last is per-process hash order. Before the fix, `tuple_max` dropped
|
|
/// a NaN accumulator in favour of the next finite delta and this returned
|
|
/// `Ok(converged: true)` with a NaN posterior in **16 of 30 runs** on identical
|
|
/// input. Deterministic now, but note this test can only ever sample one hash
|
|
/// order per run — the ordering guarantee itself is pinned by
|
|
/// `tuple_max_propagates_a_nan_from_any_position` in the crate's unit tests.
|
|
#[test]
|
|
fn a_nan_competitor_is_not_masked_by_a_healthy_one() {
|
|
let mut h = History::builder()
|
|
.mu(0.0)
|
|
.sigma(6.0)
|
|
.beta(1.0)
|
|
.p_draw(0.1)
|
|
.build();
|
|
h.add_events(vec![
|
|
Event {
|
|
time: 1i64,
|
|
teams: smallvec![
|
|
Team::with_members([Member::new("a").with_prior(Gaussian::from_ms(0.0, 1e-200))]),
|
|
Team::with_members([Member::new("b")]),
|
|
],
|
|
outcome: Outcome::winner(0, 2),
|
|
},
|
|
// A healthy pair in the same slice, to be reduced alongside the NaN.
|
|
Event {
|
|
time: 1i64,
|
|
teams: smallvec![
|
|
Team::with_members([Member::new("c")]),
|
|
Team::with_members([Member::new("d")]),
|
|
],
|
|
outcome: Outcome::winner(0, 2),
|
|
},
|
|
])
|
|
.unwrap();
|
|
|
|
let err = h
|
|
.converge()
|
|
.expect_err("a NaN fit must never be reported as converged");
|
|
assert!(
|
|
matches!(err, InferenceError::NonFiniteStep { .. }),
|
|
"{err:?}"
|
|
);
|
|
}
|
|
|
|
/// A tie observed with a narrow draw margin between far-apart competitors must
|
|
/// produce a fit, not NaN skills.
|
|
///
|
|
/// The tie branch forms the truncated variance from `v^2 - u`, and both grow as
|
|
/// `alpha^2` while their difference stays `O(1)`. Deep enough into the tail
|
|
/// that subtraction had four digits left: measured, it returned `1 - w`
|
|
/// negative and `sqrt` of it was NaN. The half-line escape hatch did not cover
|
|
/// it, because that keys on how many window-widths from the mean the window
|
|
/// sits and a narrow window fails that however deep it is.
|
|
///
|
|
/// These parameters are ordinary for a precise-scoring domain, and the
|
|
/// neighbouring wider-margin case always worked — so this was a cliff, not
|
|
/// "extreme inputs break".
|
|
#[test]
|
|
fn a_narrow_draw_margin_far_into_the_tail_still_fits() {
|
|
for (beta, p_draw, sd, gap) in [
|
|
(1e-2, 1e-8, 1e-2, 10.0),
|
|
(1e-3, 1e-9, 1e-3, 1.0),
|
|
(1e-4, 1e-12, 1e-4, 1.0),
|
|
] {
|
|
let mut h = History::builder()
|
|
.mu(0.0)
|
|
.sigma(sd)
|
|
.beta(beta)
|
|
.p_draw(p_draw)
|
|
.drift(ConstantDrift::new(0.0))
|
|
.build();
|
|
h.add_events(vec![Event {
|
|
time: 1i64,
|
|
teams: smallvec![
|
|
Team::with_members([Member::new("a").with_prior(Gaussian::from_ms(0.0, sd))]),
|
|
Team::with_members([Member::new("b").with_prior(Gaussian::from_ms(gap, sd))]),
|
|
],
|
|
outcome: Outcome::draw(2),
|
|
}])
|
|
.unwrap();
|
|
|
|
let report = h
|
|
.converge()
|
|
.unwrap_or_else(|e| panic!("beta {beta:e}, p_draw {p_draw:e}: {e:?}"));
|
|
assert!(report.converged);
|
|
|
|
let skill = h.current_skill(&"a").unwrap();
|
|
assert!(
|
|
skill.mu().is_finite() && skill.sigma().is_finite() && skill.sigma() > 0.0,
|
|
"beta {beta:e}, p_draw {p_draw:e}: {skill:?}"
|
|
);
|
|
}
|
|
}
|