feat: add expected information gain for active matchup selection
`quality()` answers "is this matchup fair". Callers picking which
comparison to run next need "is this matchup informative", and the two
coincide only for two evenly matched competitors. Without a principled
alternative, downstream code was reaching for hand-rolled heuristics
like `quality * sigma_a^2 * sigma_b^2`, which double-counts uncertainty:
the two factors are not independent.
Adds `expected_information_gain`, the outcome-weighted divergence
between current beliefs and the beliefs each result would produce:
EIG = SUM P(outcome) * KL(posterior_after(outcome) || prior)
Available standalone over `Rating`s, and as
`History::expected_information_gain` using current skills and the
history's own beta, drift and p_draw — so the outcomes it weighs are the
ones that would actually be fitted.
This is the mutual information between the outcome and the skills, which
gives an analytic ceiling: gain cannot exceed the entropy of the thing
being observed, so at most `ln k` nats for k outcomes. That bound is the
sharpest test available, because an acquisition function is unusually
exposed to returning finite, plausible, monotone numbers while being
wrong — it would simply select slightly worse matchups forever. A
prototype of this returned 4.77 nats from a sign error while passing
every monotonicity check; `never_exceeds_the_entropy_of_the_outcome`
catches that class unconditionally.
Measured against the ceiling the values are meaningful rather than
vacuous: 0.382 nats for an even matchup between diffuse priors against
an 0.693 ceiling, falling to 0.013 for a lopsided one and 0.000 for a
hopeless one.
`disagrees_with_the_quality_times_variance_heuristic` pins down that
this is not a monotone transform of the heuristic it replaces — the two
rank a lopsided matchup and a confident even one in opposite orders — so
a later "simplification" cannot quietly revert to it.
Cost is one inference pass per possible outcome, documented on the
public API alongside the shortlist-then-score pattern, so callers do not
discover it in production.
Also folds the duplicated key-gathering in `predict_quality` and
`performances` into one validated `member_skills`.
Refs #39
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
This commit is contained in:
@@ -212,3 +212,80 @@ fn team_size_affects_the_prediction() {
|
||||
assert!((p.total() - 1.0).abs() < 1e-6, "total = {}", p.total());
|
||||
assert!(p.probability_of(&[0, 0]) > 0.0);
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Expected information gain
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/// The whole point of #39: "which comparison should I run next?" is a
|
||||
/// different question from "who will win?" or "is this fair?".
|
||||
#[test]
|
||||
fn information_gain_prefers_the_uncertain_pairing() {
|
||||
let mut h = History::builder().build();
|
||||
|
||||
// "known" and "rival" have played a lot; "newcomer" has played once.
|
||||
for t in 1..=15 {
|
||||
h.record_winner(&"known", &"rival", t).unwrap();
|
||||
h.record_winner(&"rival", &"known", t + 100).unwrap();
|
||||
}
|
||||
h.record_winner(&"known", &"newcomer", 500).unwrap();
|
||||
h.converge().unwrap();
|
||||
|
||||
let settled = h
|
||||
.expected_information_gain(&[&[&"known"], &[&"rival"]])
|
||||
.unwrap();
|
||||
let unknown = h
|
||||
.expected_information_gain(&[&[&"known"], &[&"newcomer"]])
|
||||
.unwrap();
|
||||
|
||||
assert!(
|
||||
unknown > settled,
|
||||
"pairing against the newcomer should teach more: {unknown} vs {settled}"
|
||||
);
|
||||
}
|
||||
|
||||
/// The analytic ceiling, through the `History` entry point rather than the
|
||||
/// standalone one.
|
||||
#[test]
|
||||
fn information_gain_respects_the_entropy_ceiling() {
|
||||
let h = history_with(&["a", "b", "c"], 0.0);
|
||||
|
||||
let two = h.expected_information_gain(&[&[&"a"], &[&"b"]]).unwrap();
|
||||
assert!(
|
||||
(0.0..=std::f64::consts::LN_2).contains(&two),
|
||||
"two-team EIG {two} outside [0, ln 2]"
|
||||
);
|
||||
|
||||
let three = h
|
||||
.expected_information_gain(&[&[&"a"], &[&"b"], &[&"c"]])
|
||||
.unwrap();
|
||||
assert!(
|
||||
(0.0..=6.0f64.ln()).contains(&three),
|
||||
"three-team EIG {three} outside [0, ln 6]"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn information_gain_reports_unknown_keys() {
|
||||
let h = history_with(&["a", "b"], 0.0);
|
||||
assert_eq!(
|
||||
h.expected_information_gain(&[&[&"a"], &[&"ghost"]])
|
||||
.unwrap_err(),
|
||||
InferenceError::UnknownKey { team: 1, member: 0 }
|
||||
);
|
||||
}
|
||||
|
||||
/// A draw-enabled history has three outcomes to weigh rather than two, so the
|
||||
/// draw branch must actually be reachable through this path.
|
||||
#[test]
|
||||
fn information_gain_accounts_for_draws() {
|
||||
let with_draws = history_with(&["a", "b"], 0.25);
|
||||
let g = with_draws
|
||||
.expected_information_gain(&[&[&"a"], &[&"b"]])
|
||||
.unwrap();
|
||||
assert!(g > 0.0 && g <= 3.0f64.ln(), "{g}");
|
||||
|
||||
// The draw outcome carries mass, so it is genuinely being weighed.
|
||||
let dist = with_draws.predict_outcome(&[&[&"a"], &[&"b"]]).unwrap();
|
||||
assert!(dist.probability_of(&[0, 0]) > 0.0);
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user