Six of fifteen variants carried a `&'static str` discriminator, about
thirty magic strings between them, and the only thing a caller could do
with one was print it. Four new enums replace them:
Parameter 13 variants, replacing 9 strings in InvalidParameter
Shape 4 variants, replacing 10 in MismatchedShape
OutcomeKind 2 variants, replacing WrongOutcomeKind's three fields
CompetitorField 2 variants, replacing ConflictingCompetitorConfig's
`InvalidProbability` folds into `InvalidParameter` as
`Parameter::PDraw`. It was a bespoke variant for one scalar while every
other scalar shared `InvalidParameter`, and it omitted the parameter
name — so the same parameter had two mechanisms.
`JointUnavailable { reason: &'static str }` splits into `EmptyHistory`,
`JointRequiresScoredEvents` and `NotPositiveDefinite`. The three are
conditions a caller branches on differently — add events, use
`predict_win_probabilities`, or reconsider the priors — and telling them
apart used to mean string-matching English. One test already proved the
distinction was load-bearing: the blanket conversion mapped the
empty-history case onto the ranked one and `an_empty_history_has_no_joint`
caught it immediately.
`NonFiniteResult` splits into `NonFiniteStep { context, step }` and
`NonFiniteSkill { mu, sigma }`. One `step: (f64, f64)` field was
carrying a sweep step from `converge` and a skill's own moments from a
prediction — two situations in one variant, and a field name that could
only be right for one of them.
`InvalidParameter { name: "beta with point-mass skills" }` becomes
`NoPerformanceVariance`. It was never a parameter out of range: both
values are individually valid and it is their combination that leaves
nothing varying.
Three `Display` impls did not meet the standard the others set, and the
typed data is what makes fixing them possible:
before drift variance is invalid: NaN
after drift variance must be finite and non-negative (got NaN)
before kinds: expected length 3, got 2
after the outcome describes a different number of teams than the
event has: expected 3, got 2
before Game::ranked: expected Outcome::Ranked, got Outcome::Scored
after expected Outcome::Ranked, got Outcome::Scored; call
Game::scored for a scored outcome
`Parameter::range()` states each parameter's actual bounds, which no
`&'static str` name could have. `error::message_tests` renders every one
and asserts each is a sentence rather than a label, and that the three
above now carry a range or a next step.
The four internal `MismatchedShape` kinds — `results`, `times`, `kinds`,
and the weights array — collapse to `Shape::Internal`, whose `Display`
says plainly that reaching it is a bug in this crate. They are checks on
`add_events_with_prior`'s own parallel arrays and are unreachable
through the public API; they stay checked rather than becoming
`debug_assert!`s, because release is where this crate's defects hide.
Closes #74.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
153 lines
4.6 KiB
Rust
153 lines
4.6 KiB
Rust
//! No prediction path may answer from a fit it cannot answer from.
|
|
//!
|
|
//! `converge` grew a `NonFiniteResult` guard; nothing stopped a caller from
|
|
//! ignoring that error and predicting anyway. The three failures that produced
|
|
//! were each differently wrong: `Ok(NaN)`, a panic out of a `Result`-returning
|
|
//! method, and `Ok([0.0, 0.0])` — finite, plausible, summing to zero against a
|
|
//! doc that promises one.
|
|
//!
|
|
//! Every test here has a healthy control, so none can pass by everything
|
|
//! returning `Err`.
|
|
|
|
use trueskill_tt::{
|
|
ConstantDrift, Event, Gaussian, History, InferenceError, Member, Outcome, Team,
|
|
};
|
|
|
|
type H = History;
|
|
|
|
fn build(beta: f64, prior: Option<Gaussian>, outcome: Outcome) -> H {
|
|
let mut h: H = History::builder()
|
|
.beta(beta)
|
|
.drift(ConstantDrift::new(0.0))
|
|
.build();
|
|
|
|
let member = |k: &'static str| match prior {
|
|
Some(p) => Member::new(k).with_prior(p),
|
|
None => Member::new(k),
|
|
};
|
|
|
|
let _ = h.add_events(vec![Event {
|
|
time: 1,
|
|
teams: [
|
|
Team::with_members([member("a")]),
|
|
Team::with_members([member("b")]),
|
|
]
|
|
.into_iter()
|
|
.collect(),
|
|
outcome,
|
|
}]);
|
|
h
|
|
}
|
|
|
|
/// Point-mass priors with `beta(0.0)` on a *ranked* event: `converge` reports
|
|
/// `NonFiniteResult` and the stored posteriors are `pi: NaN, tau: NaN`.
|
|
fn nan_poisoned() -> H {
|
|
let mut h = build(
|
|
0.0,
|
|
Some(Gaussian::from_ms(0.0, 0.0)),
|
|
Outcome::winner(0, 2),
|
|
);
|
|
let err = h.converge().expect_err("this fixture must not converge");
|
|
assert!(
|
|
matches!(err, InferenceError::NonFiniteStep { .. }),
|
|
"{err:?}"
|
|
);
|
|
h
|
|
}
|
|
|
|
/// The same degenerate parameters on a *scored* event, where inference
|
|
/// converges cleanly and leaves legitimate point-mass posteriors behind. The
|
|
/// fit is fine; it is prediction that has nothing to work with.
|
|
fn degenerate_but_converged() -> H {
|
|
let mut h = build(
|
|
0.0,
|
|
Some(Gaussian::from_ms(0.0, 0.0)),
|
|
Outcome::scores([1.0, 0.0]),
|
|
);
|
|
h.converge().expect("this fixture converges");
|
|
h
|
|
}
|
|
|
|
fn healthy() -> H {
|
|
let mut h = build(1.0, None, Outcome::winner(0, 2));
|
|
h.converge().expect("control converges");
|
|
h
|
|
}
|
|
|
|
macro_rules! all_predictions {
|
|
($h:ident, $f:expr) => {{
|
|
let teams: &[&[&&'static str]] = &[&[&"a"], &[&"b"]];
|
|
let f = $f;
|
|
f("quality", $h.quality(teams).map(|_| ()));
|
|
f(
|
|
"predict_win_probabilities",
|
|
$h.predict_win_probabilities(teams).map(|_| ()),
|
|
);
|
|
f("predict_outcome", $h.predict_outcome(teams).map(|_| ()));
|
|
f(
|
|
"predict_ranking",
|
|
$h.predict_ranking(teams, &[0, 1]).map(|_| ()),
|
|
);
|
|
f(
|
|
"expected_information_gain",
|
|
$h.expected_information_gain(teams).map(|_| ()),
|
|
);
|
|
}};
|
|
}
|
|
|
|
#[test]
|
|
fn a_nan_poisoned_fit_is_refused_by_every_prediction_path() {
|
|
let h = nan_poisoned();
|
|
|
|
all_predictions!(h, |name: &str, r: Result<(), InferenceError>| {
|
|
match r {
|
|
Err(InferenceError::NonFiniteSkill { .. }) => {}
|
|
other => panic!("{name} answered from a NaN fit: {other:?}"),
|
|
}
|
|
});
|
|
}
|
|
|
|
#[test]
|
|
fn degenerate_performances_are_refused_rather_than_answered_wrongly() {
|
|
let h = degenerate_but_converged();
|
|
|
|
// The fit itself is sound — the posteriors are point masses, not NaN.
|
|
let skill = h.current_skill("a").expect("a played");
|
|
assert_eq!(skill.sigma(), 0.0);
|
|
assert!(skill.mu().is_finite());
|
|
|
|
// `quality` previously PANICKED here, out of a method that returns
|
|
// `Result`: the contrast covariance is exactly singular when beta is zero
|
|
// and every skill is a point mass.
|
|
all_predictions!(h, |name: &str, r: Result<(), InferenceError>| {
|
|
match r {
|
|
Err(InferenceError::NoPerformanceVariance) => {}
|
|
other => panic!("{name} predicted from a degenerate fit: {other:?}"),
|
|
}
|
|
});
|
|
}
|
|
|
|
#[test]
|
|
fn the_control_history_answers_every_prediction() {
|
|
let h = healthy();
|
|
|
|
all_predictions!(h, |name: &str, r: Result<(), InferenceError>| {
|
|
assert!(r.is_ok(), "{name} failed on a healthy history: {r:?}");
|
|
});
|
|
}
|
|
|
|
#[test]
|
|
fn win_probabilities_sum_to_one_on_the_control() {
|
|
// The promise the `Ok([0.0, 0.0])` case broke. Asserted on the control so
|
|
// the guard above cannot be "fixed" by making every path error.
|
|
let h = healthy();
|
|
let p = h
|
|
.predict_win_probabilities(&[&[&"a"], &[&"b"]])
|
|
.expect("control predicts");
|
|
let total: f64 = p.iter().sum();
|
|
assert!(
|
|
(total - 1.0).abs() < 1e-6,
|
|
"win probabilities sum to {total}"
|
|
);
|
|
}
|