Three issues from two downstream consumers, all small, all sharing a theme: the crate had the information and would not hand it over. #44 — `UnknownKey { team: 0, member: 0 }` did not say which key. A consumer upgrading 0.1.2 -> 0.4.1 had every one of 5591 predictions return this error, fell back to a neutral 0.5, and lost its entire metadata model for a day. Nothing crashed and nothing logged; it was found by sweeping an unrelated parameter and noticing the output did not move. The 0.4.0 change that made unknown keys an error was right — the error was just too anonymous to act on. It now carries the key's `Debug` rendering, and its `Display` says what to do about it. The precondition is documented on every prediction entry point, which the reporter said would alone have saved the day. #43 — `cdf` was `pub(crate)`, so a consumer asking "is this competitor below the cutoff" approximated it with a `mu + z * sigma` band and had no way to say what confidence any `z` bought. Adds `Gaussian::probability_below` / `probability_above`. The second is separate on purpose: `1 - cdf` collapses to exactly zero past ~8.3 sigma, and a stopping rule is evaluated precisely there. Both route through the survival function added in 0.4.1, so this is visibility rather than new numerics. #50 — `ConvergenceReport` was not `#[must_use]`, so the one signal that a fit stopped short was trivially discarded. It now is, and that immediately found 78 sites doing exactly that — including this crate's own ATP example, which was capped at 10 sweeps when the history needs 30. The example now reads the report and says so. `ITERATIONS = 30` is documented as the floor it is, with the three measurements to hand: 400 events over 100 competitors already stops there at ~7e-3 against a 1e-6 tolerance, the ATP example needs 30 at a much looser one, and a consumer's 2000-node model needs 76 to 161. BREAKING CHANGE: `InferenceError::UnknownKey` gains a `key` field, and the prediction methods now require `K: Debug` in order to fill it. Closes #43, #50. Refs #44 — its third ask, an opt-in `UnknownKeys::Skip` mode, is a live API question and deliberately not answered here. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
72 lines
2.4 KiB
Rust
72 lines
2.4 KiB
Rust
//! Regression: a single time slice with many distinct competitors must converge to finite
|
|
//! skills. Before the `pi <= 0` guard in `Gaussian::mu()/sigma()`, EP message cancellation
|
|
//! produced a tiny-negative precision whose `sigma() = 1/sqrt(pi)` was NaN, which the
|
|
//! moment-space `Sub` in the game chain propagated into every skill once the slice grew past
|
|
//! ~75 competitors (e.g. a real ranking dataset with hundreds of players).
|
|
use trueskill_tt::{ConstantDrift, ConvergenceOptions, EPSILON, History, ITERATIONS, NullObserver};
|
|
|
|
/// Tiny deterministic LCG — avoids a dev-dependency on `rand`.
|
|
struct Lcg(u64);
|
|
impl Lcg {
|
|
fn next(&mut self) -> u64 {
|
|
self.0 = self
|
|
.0
|
|
.wrapping_mul(6364136223846793005)
|
|
.wrapping_add(1442695040888963407);
|
|
self.0
|
|
}
|
|
fn below(&mut self, n: usize) -> usize {
|
|
(self.next() >> 33) as usize % n
|
|
}
|
|
fn coin(&mut self) -> bool {
|
|
self.next() & 1 == 0
|
|
}
|
|
}
|
|
|
|
fn nan_after_fit(players: usize) -> usize {
|
|
let mut h: History<i64, ConstantDrift, NullObserver, String> = History::builder_with_key()
|
|
.beta(1.0)
|
|
.sigma(6.0)
|
|
.drift(ConstantDrift(0.1))
|
|
.convergence(ConvergenceOptions {
|
|
max_iter: ITERATIONS,
|
|
epsilon: EPSILON,
|
|
..Default::default()
|
|
})
|
|
.build();
|
|
|
|
let ids: Vec<String> = (0..players).map(|i| format!("p{i:04}")).collect();
|
|
let mut rng = Lcg(1);
|
|
for _ in 0..(players * 4) {
|
|
let a = rng.below(players);
|
|
let mut b = rng.below(players - 1);
|
|
if b >= a {
|
|
b += 1;
|
|
}
|
|
let (w, l) = if rng.coin() { (a, b) } else { (b, a) };
|
|
h.record_winner(&ids[w], &ids[l], 0).unwrap();
|
|
}
|
|
let _ = h.converge().unwrap();
|
|
|
|
ids.iter()
|
|
.filter(|id| {
|
|
h.current_skill(id.as_str())
|
|
.map(|g| !g.mu().is_finite() || !g.sigma().is_finite())
|
|
.unwrap_or(true)
|
|
})
|
|
.count()
|
|
}
|
|
|
|
#[test]
|
|
fn many_competitors_converge_to_finite_skills() {
|
|
// The NaN regression onset was between 70 and 80 competitors; 250 is comfortably past it
|
|
// and in the range of a real ranking dataset.
|
|
for players in [12usize, 75, 150, 250] {
|
|
assert_eq!(
|
|
nan_after_fit(players),
|
|
0,
|
|
"{players}-competitor history produced NaN skills"
|
|
);
|
|
}
|
|
}
|