`Add`, `Sub`, `exclude` and `forget` combined variances by way of standard
deviations: `sigma()` takes a square root, `.powi(2)` squares it away,
`var.sqrt()` takes another, and `from_ms` squares that one back. Three roots
to compute a value that is `1/pi` all along.
They now go through `variance()` and a new `from_mv(mu, var)`, which skip
both conversions. `Sub` is the hot one — `RankDiffFactor::propagate` is
`a - b`, run for every adjacent team pair on every forward and backward
sweep of every EP iteration.
`run_chain` also stopped recomputing each team's weighted performance in the
likelihood loop; the fold is already in `arena.team_prior`, indexed by the
sorted position the loop has in hand. Each `performance()` is itself a
`forget`, so the duplicate cost scaled with players per team.
Measured on this machine, before and after, same fixtures:
Batch::iteration 23.57us -> 19.31us (-18%)
scored_history_60_events 1.071ms -> 983us (-8%)
The `Gaussian::add`/`sub` microbenchmarks cannot resolve the change: they
sit at ~234ps against a ~218ps floor that `mul`/`div` also hit, so the
harness overhead dominates a single operation.
One golden moved. Two identical competitors drawing must land on their
shared prior mean exactly, by symmetry; the root-free path now returns
25.0 where the reference transcription recorded 24.999999 — that value
rounded to six decimals. Asserting a six-decimal transcription at
epsilon 1e-6 left no headroom, so the expectation is now the exact value.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DnsaJg74eNSva3PJjK2eej
57 lines
2.1 KiB
Rust
57 lines
2.1 KiB
Rust
//! Equivalence tests: every historical golden from the pre-T2 tests is
|
|
//! reproduced here at the integration level via the new public API.
|
|
//!
|
|
//! The in-crate tests in `src/history.rs::tests` and
|
|
//! `src/time_slice.rs::tests` are the primary regression net for numerical
|
|
//! behavior. This file provides Game-level goldens that stand alone and are
|
|
//! more naturally expressed as integration tests.
|
|
|
|
use approx::assert_ulps_eq;
|
|
use trueskill_tt::{ConstantDrift, Game, GameOptions, Gaussian, Outcome, Rating};
|
|
|
|
type R = Rating<i64, ConstantDrift>;
|
|
|
|
fn ts_rating(mu: f64, sigma: f64, beta: f64, gamma: f64) -> R {
|
|
R::new(Gaussian::from_ms(mu, sigma), beta, ConstantDrift(gamma))
|
|
}
|
|
|
|
#[test]
|
|
fn game_1v1_golden_matches_historical() {
|
|
let a = ts_rating(25.0, 25.0 / 3.0, 25.0 / 6.0, 25.0 / 300.0);
|
|
let b = ts_rating(25.0, 25.0 / 3.0, 25.0 / 6.0, 25.0 / 300.0);
|
|
let (a_post, b_post) = Game::<i64, _>::one_v_one(&a, &b, Outcome::winner(0, 2)).unwrap();
|
|
// Historical golden from pre-T2 test_1vs1 (team 0 wins):
|
|
assert_ulps_eq!(
|
|
a_post,
|
|
Gaussian::from_ms(29.205220, 7.194481),
|
|
epsilon = 1e-6
|
|
);
|
|
assert_ulps_eq!(
|
|
b_post,
|
|
Gaussian::from_ms(20.794779, 7.194481),
|
|
epsilon = 1e-6
|
|
);
|
|
}
|
|
|
|
#[test]
|
|
fn game_1v1_draw_golden() {
|
|
let a = ts_rating(25.0, 25.0 / 3.0, 25.0 / 6.0, 25.0 / 300.0);
|
|
let b = ts_rating(25.0, 25.0 / 3.0, 25.0 / 6.0, 25.0 / 300.0);
|
|
let g = Game::<i64, _>::ranked(
|
|
&[&[a], &[b]],
|
|
Outcome::draw(2),
|
|
&GameOptions {
|
|
p_draw: 0.25,
|
|
score_sigma: 1.0,
|
|
convergence: Default::default(),
|
|
},
|
|
)
|
|
.unwrap();
|
|
let p = g.posteriors();
|
|
// Historical golden from pre-T2 test_1vs1_draw. The mean is 25.0 exactly
|
|
// by symmetry — two identical competitors drawing cannot move apart — and
|
|
// the reference's 24.999999 is that value transcribed to six decimals.
|
|
assert_ulps_eq!(p[0][0], Gaussian::from_ms(25.0, 6.469480), epsilon = 1e-6);
|
|
assert_ulps_eq!(p[1][0], Gaussian::from_ms(25.0, 6.469480), epsilon = 1e-6);
|
|
}
|