`grid_shape` asked for 12 nodes across the narrowest feature and then clamped to MAX_GRID_POINTS with no detection that the request was not met. Past `step/sigma ~ 1.7` the trapezoid rule stops resolving the density, and the result is unbounded: sigma_a step/sig_a P(a first) exact total 2.0e-3 0.86 0.515953 0.515953 1.000000 1.0e-3 1.72 0.517185 0.515953 1.002388 1.0e-4 17.17 2.791336 0.515953 5.410065 A probability of 2.79. Reachable through `predict_outcome` with a pinned reference competitor — a documented pattern — where `predict_outcome` and `predict_win_probabilities` disagreed 44x and `predict_outcome` was the wrong one. There is no useful answer on the far side of that cliff, so this reports `GridTooCoarse` rather than guessing, and the message points at `predict_win_probabilities`, which answers the same matchup through adaptive quadrature and is accurate there to 1e-13. The floor is 4 nodes per feature rather than the 12 requested, because the request carries margin: measured accurate to 2.2e-12 at 1.4 nodes per sigma and wrong by 1.2e-3 at 0.7. This also fixes the `ln k` ceiling violation. `expected_information_gain` weights `probability * divergence`, so probabilities of 3.97 and 2.62 made it return 3.237828 nats against `ln 2 = 0.693147` — 4.67x over. The crate's docs call that ceiling its sharpest test and record a prototype once returning 4.77 nats; it was live again by a different route. The new sweep then caught a second, independent defect: `kl_divergence` returned NEGATIVE values, worst -5.55e-17, exactly one ULP of its `- 1.0`. Rewritten as `0.5*(u - ln1p(u)) + gap^2/(2*var_p)` with `u = var_q/var_p - 1`, so both terms are non-negative by construction. It is also more accurate where it matters: at `u = 1e-9` the old form returned 0.0 where the true value is 2.5e-19, and well-conditioned cases are unchanged. tests/prediction_bounds.rs sweeps rather than spot-checks, because a single fixture cannot defend a bound like this — the previous check passed throughout. It asserts the sweep still reaches the coarse-grid regime, so it cannot quietly stop testing the case it was written for. BREAKING CHANGE: `predict_outcome`, `predict_ranking` and `expected_information_gain` return `GridTooCoarse` for matchups whose performance sigmas are too far apart to integrate on one grid. They previously returned wrong answers, including probabilities above 1. Closes #55, closes #56 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
137 lines
5.0 KiB
Rust
137 lines
5.0 KiB
Rust
//! Bounds that any correct implementation must satisfy, swept rather than
|
|
//! spot-checked.
|
|
//!
|
|
//! The crate's docs call the `ln k` ceiling "the sharpest available test of an
|
|
//! implementation", and record that an early prototype returned 4.77 nats. It
|
|
//! was violated again — 3.237828 nats against `ln 2` — because the existing
|
|
//! check sampled one fixture and the violation lives in a specific regime: a
|
|
//! large ratio between the widest and narrowest performance sigma, where the
|
|
//! shared prediction grid could not resolve the narrow density and returned
|
|
//! probabilities greater than one.
|
|
//!
|
|
//! A single fixture cannot defend a bound like this. A sweep can.
|
|
|
|
use trueskill_tt::{
|
|
ConstantDrift, GameOptions, Gaussian, InferenceError, Rating, expected_information_gain,
|
|
};
|
|
|
|
type R = Rating<i64, ConstantDrift>;
|
|
|
|
/// Deterministic LCG, so a failure is reproducible from the printed seed.
|
|
struct Lcg(u64);
|
|
|
|
impl Lcg {
|
|
fn next_f64(&mut self) -> f64 {
|
|
self.0 = self
|
|
.0
|
|
.wrapping_mul(6_364_136_223_846_793_005)
|
|
.wrapping_add(1_442_695_040_888_963_407);
|
|
// Top 53 bits to [0, 1).
|
|
((self.0 >> 11) as f64) / ((1u64 << 53) as f64)
|
|
}
|
|
|
|
fn in_range(&mut self, lo: f64, hi: f64) -> f64 {
|
|
lo + (hi - lo) * self.next_f64()
|
|
}
|
|
|
|
/// Log-uniform, so the sweep spends its samples across magnitudes rather
|
|
/// than crowding the top of the range — the violations live at small sigma.
|
|
fn log_uniform(&mut self, lo: f64, hi: f64) -> f64 {
|
|
let t = self.next_f64();
|
|
(lo.ln() + t * (hi.ln() - lo.ln())).exp()
|
|
}
|
|
}
|
|
|
|
#[test]
|
|
fn information_gain_never_exceeds_the_entropy_of_the_outcome() {
|
|
let mut rng = Lcg(0x5eed_1234_abcd_ef01);
|
|
let ceiling = 2.0_f64.ln();
|
|
let mut evaluated = 0usize;
|
|
let mut refused = 0usize;
|
|
|
|
for i in 0..2_000 {
|
|
let mu_a = rng.in_range(-100.0, 100.0);
|
|
let mu_b = rng.in_range(-100.0, 100.0);
|
|
let sigma_a = rng.log_uniform(1e-4, 1e2);
|
|
let sigma_b = rng.log_uniform(1e-4, 1e2);
|
|
let beta = rng.log_uniform(1e-4, 1e1);
|
|
|
|
let a = R::new(Gaussian::from_ms(mu_a, sigma_a), beta, ConstantDrift(0.0));
|
|
let b = R::new(Gaussian::from_ms(mu_b, sigma_b), beta, ConstantDrift(0.0));
|
|
let options = GameOptions {
|
|
p_draw: 0.0,
|
|
..GameOptions::default()
|
|
};
|
|
|
|
match expected_information_gain(&[&[a], &[b]], &options) {
|
|
Ok(gain) => {
|
|
evaluated += 1;
|
|
assert!(
|
|
gain.is_finite(),
|
|
"sample {i}: non-finite gain {gain} \
|
|
(mu {mu_a}, {mu_b}; sigma {sigma_a:e}, {sigma_b:e}; beta {beta:e})"
|
|
);
|
|
assert!(
|
|
gain >= 0.0,
|
|
"sample {i}: negative gain {gain} \
|
|
(mu {mu_a}, {mu_b}; sigma {sigma_a:e}, {sigma_b:e}; beta {beta:e})"
|
|
);
|
|
assert!(
|
|
gain <= ceiling + 1e-9,
|
|
"sample {i}: gain {gain} exceeds ln 2 = {ceiling} \
|
|
(mu {mu_a}, {mu_b}; sigma {sigma_a:e}, {sigma_b:e}; beta {beta:e})"
|
|
);
|
|
}
|
|
// Refusing to answer is acceptable; answering wrongly is not.
|
|
Err(InferenceError::GridTooCoarse { .. }) => refused += 1,
|
|
Err(e) => panic!("sample {i}: unexpected error {e:?}"),
|
|
}
|
|
}
|
|
|
|
// The sweep must actually exercise the function, not pass by refusing
|
|
// everything.
|
|
assert!(
|
|
evaluated > 1_000,
|
|
"only {evaluated} of 2000 samples were evaluated ({refused} refused); \
|
|
the sweep is no longer testing anything"
|
|
);
|
|
// And it must still reach the regime where the ceiling was violated —
|
|
// large sigma ratios, which is exactly where the grid now refuses. Without
|
|
// this the sweep could drift into only-easy inputs and stop being a guard.
|
|
assert!(
|
|
refused > 0,
|
|
"no sample reached the coarse-grid regime; the sweep no longer covers \
|
|
the case that produced 3.24 nats"
|
|
);
|
|
}
|
|
|
|
/// The regime that produced 3.237828 nats, pinned exactly.
|
|
#[test]
|
|
fn the_known_ceiling_violation_no_longer_answers_wrongly() {
|
|
let a = R::new(
|
|
Gaussian::from_ms(9.577_887_112_129_012, 0.000_132_507_526_585_134_38),
|
|
0.000_307_235_559_013_096_2,
|
|
ConstantDrift(0.0),
|
|
);
|
|
let b = R::new(
|
|
Gaussian::from_ms(-14.114_932_828_525_696, 91.586_690_140_921_16),
|
|
0.000_307_235_559_013_096_2,
|
|
ConstantDrift(0.0),
|
|
);
|
|
let options = GameOptions {
|
|
p_draw: 0.0,
|
|
..GameOptions::default()
|
|
};
|
|
|
|
match expected_information_gain(&[&[a], &[b]], &options) {
|
|
Ok(gain) => assert!(
|
|
gain <= 2.0_f64.ln() + 1e-9,
|
|
"returned {gain}, over the ln 2 ceiling"
|
|
),
|
|
Err(InferenceError::GridTooCoarse { needed, max }) => {
|
|
assert!(needed > max, "needed {needed} should exceed max {max}");
|
|
}
|
|
Err(e) => panic!("unexpected error {e:?}"),
|
|
}
|
|
}
|