feat: add expected information gain for active matchup selection

`quality()` answers "is this matchup fair". Callers picking which
comparison to run next need "is this matchup informative", and the two
coincide only for two evenly matched competitors. Without a principled
alternative, downstream code was reaching for hand-rolled heuristics
like `quality * sigma_a^2 * sigma_b^2`, which double-counts uncertainty:
the two factors are not independent.

Adds `expected_information_gain`, the outcome-weighted divergence
between current beliefs and the beliefs each result would produce:

    EIG = SUM P(outcome) * KL(posterior_after(outcome) || prior)

Available standalone over `Rating`s, and as
`History::expected_information_gain` using current skills and the
history's own beta, drift and p_draw — so the outcomes it weighs are the
ones that would actually be fitted.

This is the mutual information between the outcome and the skills, which
gives an analytic ceiling: gain cannot exceed the entropy of the thing
being observed, so at most `ln k` nats for k outcomes. That bound is the
sharpest test available, because an acquisition function is unusually
exposed to returning finite, plausible, monotone numbers while being
wrong — it would simply select slightly worse matchups forever. A
prototype of this returned 4.77 nats from a sign error while passing
every monotonicity check; `never_exceeds_the_entropy_of_the_outcome`
catches that class unconditionally.

Measured against the ceiling the values are meaningful rather than
vacuous: 0.382 nats for an even matchup between diffuse priors against
an 0.693 ceiling, falling to 0.013 for a lopsided one and 0.000 for a
hopeless one.

`disagrees_with_the_quality_times_variance_heuristic` pins down that
this is not a monotone transform of the heuristic it replaces — the two
rank a lopsided matchup and a confident even one in opposite orders — so
a later "simplification" cannot quietly revert to it.

Cost is one inference pass per possible outcome, documented on the
public API alongside the shortlist-then-score pattern, so callers do not
discover it in production.

Also folds the duplicated key-gathering in `predict_quality` and
`performances` into one validated `member_skills`.

Refs #39

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
This commit is contained in:
2026-09-07 15:08:12 +02:00
co-authored by Claude Opus 5
parent 507894dae7
commit 3c2f9ac64c
5 changed files with 581 additions and 45 deletions
+77
View File
@@ -212,3 +212,80 @@ fn team_size_affects_the_prediction() {
assert!((p.total() - 1.0).abs() < 1e-6, "total = {}", p.total());
assert!(p.probability_of(&[0, 0]) > 0.0);
}
// ---------------------------------------------------------------------------
// Expected information gain
// ---------------------------------------------------------------------------
/// The whole point of #39: "which comparison should I run next?" is a
/// different question from "who will win?" or "is this fair?".
#[test]
fn information_gain_prefers_the_uncertain_pairing() {
let mut h = History::builder().build();
// "known" and "rival" have played a lot; "newcomer" has played once.
for t in 1..=15 {
h.record_winner(&"known", &"rival", t).unwrap();
h.record_winner(&"rival", &"known", t + 100).unwrap();
}
h.record_winner(&"known", &"newcomer", 500).unwrap();
h.converge().unwrap();
let settled = h
.expected_information_gain(&[&[&"known"], &[&"rival"]])
.unwrap();
let unknown = h
.expected_information_gain(&[&[&"known"], &[&"newcomer"]])
.unwrap();
assert!(
unknown > settled,
"pairing against the newcomer should teach more: {unknown} vs {settled}"
);
}
/// The analytic ceiling, through the `History` entry point rather than the
/// standalone one.
#[test]
fn information_gain_respects_the_entropy_ceiling() {
let h = history_with(&["a", "b", "c"], 0.0);
let two = h.expected_information_gain(&[&[&"a"], &[&"b"]]).unwrap();
assert!(
(0.0..=std::f64::consts::LN_2).contains(&two),
"two-team EIG {two} outside [0, ln 2]"
);
let three = h
.expected_information_gain(&[&[&"a"], &[&"b"], &[&"c"]])
.unwrap();
assert!(
(0.0..=6.0f64.ln()).contains(&three),
"three-team EIG {three} outside [0, ln 6]"
);
}
#[test]
fn information_gain_reports_unknown_keys() {
let h = history_with(&["a", "b"], 0.0);
assert_eq!(
h.expected_information_gain(&[&[&"a"], &[&"ghost"]])
.unwrap_err(),
InferenceError::UnknownKey { team: 1, member: 0 }
);
}
/// A draw-enabled history has three outcomes to weigh rather than two, so the
/// draw branch must actually be reachable through this path.
#[test]
fn information_gain_accounts_for_draws() {
let with_draws = history_with(&["a", "b"], 0.25);
let g = with_draws
.expected_information_gain(&[&[&"a"], &[&"b"]])
.unwrap();
assert!(g > 0.0 && g <= 3.0f64.ln(), "{g}");
// The draw outcome carries mass, so it is genuinely being weighed.
let dist = with_draws.predict_outcome(&[&[&"a"], &[&"b"]]).unwrap();
assert!(dist.probability_of(&[0, 0]) > 0.0);
}