Three issues from two downstream consumers, all small, all sharing a theme: the crate had the information and would not hand it over. #44 — `UnknownKey { team: 0, member: 0 }` did not say which key. A consumer upgrading 0.1.2 -> 0.4.1 had every one of 5591 predictions return this error, fell back to a neutral 0.5, and lost its entire metadata model for a day. Nothing crashed and nothing logged; it was found by sweeping an unrelated parameter and noticing the output did not move. The 0.4.0 change that made unknown keys an error was right — the error was just too anonymous to act on. It now carries the key's `Debug` rendering, and its `Display` says what to do about it. The precondition is documented on every prediction entry point, which the reporter said would alone have saved the day. #43 — `cdf` was `pub(crate)`, so a consumer asking "is this competitor below the cutoff" approximated it with a `mu + z * sigma` band and had no way to say what confidence any `z` bought. Adds `Gaussian::probability_below` / `probability_above`. The second is separate on purpose: `1 - cdf` collapses to exactly zero past ~8.3 sigma, and a stopping rule is evaluated precisely there. Both route through the survival function added in 0.4.1, so this is visibility rather than new numerics. #50 — `ConvergenceReport` was not `#[must_use]`, so the one signal that a fit stopped short was trivially discarded. It now is, and that immediately found 78 sites doing exactly that — including this crate's own ATP example, which was capped at 10 sweeps when the history needs 30. The example now reads the report and says so. `ITERATIONS = 30` is documented as the floor it is, with the three measurements to hand: 400 events over 100 competitors already stops there at ~7e-3 against a 1e-6 tolerance, the ATP example needs 30 at a much looser one, and a consumer's 2000-node model needs 76 to 161. BREAKING CHANGE: `InferenceError::UnknownKey` gains a `key` field, and the prediction methods now require `K: Debug` in order to fill it. Closes #43, #50. Refs #44 — its third ask, an opt-in `UnknownKeys::Skip` mode, is a live API question and deliberately not answered here. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
168 lines
5.7 KiB
Rust
168 lines
5.7 KiB
Rust
//! Property-based tests over generated histories.
|
|
//!
|
|
//! The golden suite pins exact values against the Python/Julia reference on a
|
|
//! handful of fixtures. These pin *invariants* over inputs nobody wrote by
|
|
//! hand, which is where the defects this crate has actually shipped were
|
|
//! hiding: a linear evidence product that underflowed only past ~1000 teams,
|
|
//! and a batching path no golden exercised because every golden ingests in one
|
|
//! call.
|
|
|
|
mod common;
|
|
|
|
use common::assert_finite;
|
|
use proptest::prelude::*;
|
|
use smallvec::smallvec;
|
|
use trueskill_tt::{ConvergenceOptions, Event, History, Member, Outcome, Team};
|
|
|
|
/// Distinct competitors, so no event pits someone against themselves.
|
|
fn pairs() -> impl Strategy<Value = Vec<(usize, usize)>> {
|
|
prop::collection::vec((0usize..8, 0usize..8), 1..24)
|
|
.prop_map(|v| v.into_iter().filter(|(a, b)| a != b).collect::<Vec<_>>())
|
|
.prop_filter("needs at least one valid pair", |v| !v.is_empty())
|
|
}
|
|
|
|
const KEYS: [&str; 8] = ["a", "b", "c", "d", "e", "f", "g", "h"];
|
|
|
|
fn history_from(games: &[(usize, usize)]) -> History {
|
|
let mut h = History::builder()
|
|
.convergence(ConvergenceOptions {
|
|
max_iter: 200,
|
|
epsilon: 1e-10,
|
|
..ConvergenceOptions::default()
|
|
})
|
|
.build();
|
|
|
|
let events: Vec<Event<i64, &'static str>> = games
|
|
.iter()
|
|
.enumerate()
|
|
.map(|(i, &(a, b))| Event {
|
|
time: i as i64 + 1,
|
|
teams: smallvec![
|
|
Team::with_members([Member::new(KEYS[a])]),
|
|
Team::with_members([Member::new(KEYS[b])]),
|
|
],
|
|
outcome: Outcome::winner(0, 2),
|
|
})
|
|
.collect();
|
|
|
|
h.add_events(events).unwrap();
|
|
|
|
h
|
|
}
|
|
|
|
proptest! {
|
|
#![proptest_config(ProptestConfig::with_cases(48))]
|
|
|
|
/// Whatever the schedule of games, convergence must not produce NaN or an
|
|
/// improper posterior. `converge` returns `NonFiniteResult` rather than
|
|
/// silently reporting a NaN step as converged, so a break shows up here as
|
|
/// either an Err or a non-finite curve point.
|
|
#[test]
|
|
fn converged_posteriors_are_always_finite(games in pairs()) {
|
|
let mut h = history_from(&games);
|
|
|
|
let _ = h.converge().unwrap();
|
|
|
|
for key in KEYS {
|
|
for (time, g) in h.learning_curve(key) {
|
|
assert_finite(g, &format!("{key} at t={time}"));
|
|
}
|
|
}
|
|
}
|
|
|
|
/// Log-evidence is a log probability: finite, and never above zero.
|
|
///
|
|
/// The linear-product implementation this replaced underflowed to zero on
|
|
/// long chains, making `ln(0)` = -inf — finite-ness is the property that
|
|
/// would have caught it.
|
|
#[test]
|
|
fn log_evidence_is_a_finite_log_probability(games in pairs()) {
|
|
let mut h = history_from(&games);
|
|
|
|
let _ = h.converge().unwrap();
|
|
|
|
let batch = h.log_evidence();
|
|
let filtered = h.filtered_log_evidence();
|
|
|
|
prop_assert!(batch.is_finite(), "batch log-evidence {batch} is not finite");
|
|
prop_assert!(batch <= 0.0, "batch log-evidence {batch} exceeds zero");
|
|
prop_assert!(filtered.is_finite(), "filtered log-evidence {filtered} is not finite");
|
|
prop_assert!(filtered <= 0.0, "filtered log-evidence {filtered} exceeds zero");
|
|
}
|
|
|
|
/// Filtered estimates must not depend on whether `converge` has run — the
|
|
/// property the whole forward-only design rests on.
|
|
#[test]
|
|
fn filtered_evidence_is_invariant_to_convergence(games in pairs()) {
|
|
let mut h = history_from(&games);
|
|
|
|
let before = h.filtered_log_evidence();
|
|
|
|
let _ = h.converge().unwrap();
|
|
|
|
let after = h.filtered_log_evidence();
|
|
|
|
prop_assert!(
|
|
(before - after).abs() < 1e-8,
|
|
"filtered evidence moved across converge(): {before} -> {after}"
|
|
);
|
|
}
|
|
|
|
/// Ingesting the same games one at a time must reach the same fixed point
|
|
/// as ingesting them in one call.
|
|
#[test]
|
|
fn ingestion_order_does_not_change_the_answer(games in pairs()) {
|
|
let batched = {
|
|
let mut h = history_from(&games);
|
|
let _ = h.converge().unwrap();
|
|
h
|
|
};
|
|
|
|
let incremental = {
|
|
let mut h = History::builder()
|
|
.convergence(ConvergenceOptions {
|
|
max_iter: 200,
|
|
epsilon: 1e-10,
|
|
..ConvergenceOptions::default()
|
|
})
|
|
.build();
|
|
|
|
for (i, &(a, b)) in games.iter().enumerate() {
|
|
h.add_events([Event {
|
|
time: i as i64 + 1,
|
|
teams: smallvec![
|
|
Team::with_members([Member::new(KEYS[a])]),
|
|
Team::with_members([Member::new(KEYS[b])]),
|
|
],
|
|
outcome: Outcome::winner(0, 2),
|
|
}])
|
|
.unwrap();
|
|
}
|
|
|
|
let _ = h.converge().unwrap();
|
|
h
|
|
};
|
|
|
|
for key in KEYS {
|
|
let one = batched.current_skill(key);
|
|
let other = incremental.current_skill(key);
|
|
|
|
match (one, other) {
|
|
(Some(one), Some(other)) => {
|
|
prop_assert!(
|
|
(one.mu() - other.mu()).abs() < 1e-6
|
|
&& (one.sigma() - other.sigma()).abs() < 1e-6,
|
|
"{key}: batched mu={} sigma={}, incremental mu={} sigma={}",
|
|
one.mu(),
|
|
one.sigma(),
|
|
other.mu(),
|
|
other.sigma()
|
|
);
|
|
}
|
|
(None, None) => {}
|
|
_ => prop_assert!(false, "{key} present in only one history"),
|
|
}
|
|
}
|
|
}
|
|
}
|