feat: add expected information gain for active matchup selection

`quality()` answers "is this matchup fair". Callers picking which
comparison to run next need "is this matchup informative", and the two
coincide only for two evenly matched competitors. Without a principled
alternative, downstream code was reaching for hand-rolled heuristics
like `quality * sigma_a^2 * sigma_b^2`, which double-counts uncertainty:
the two factors are not independent.

Adds `expected_information_gain`, the outcome-weighted divergence
between current beliefs and the beliefs each result would produce:

    EIG = SUM P(outcome) * KL(posterior_after(outcome) || prior)

Available standalone over `Rating`s, and as
`History::expected_information_gain` using current skills and the
history's own beta, drift and p_draw — so the outcomes it weighs are the
ones that would actually be fitted.

This is the mutual information between the outcome and the skills, which
gives an analytic ceiling: gain cannot exceed the entropy of the thing
being observed, so at most `ln k` nats for k outcomes. That bound is the
sharpest test available, because an acquisition function is unusually
exposed to returning finite, plausible, monotone numbers while being
wrong — it would simply select slightly worse matchups forever. A
prototype of this returned 4.77 nats from a sign error while passing
every monotonicity check; `never_exceeds_the_entropy_of_the_outcome`
catches that class unconditionally.

Measured against the ceiling the values are meaningful rather than
vacuous: 0.382 nats for an even matchup between diffuse priors against
an 0.693 ceiling, falling to 0.013 for a lopsided one and 0.000 for a
hopeless one.

`disagrees_with_the_quality_times_variance_heuristic` pins down that
this is not a monotone transform of the heuristic it replaces — the two
rank a lopsided matchup and a confident even one in opposite orders — so
a later "simplification" cannot quietly revert to it.

Cost is one inference pass per possible outcome, documented on the
public API alongside the shortlist-then-score pattern, so callers do not
discover it in production.

Also folds the duplicated key-gathering in `predict_quality` and
`performances` into one validated `member_skills`.

Refs #39

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
This commit is contained in:
2026-09-07 15:08:12 +02:00
co-authored by Claude Opus 5
parent 507894dae7
commit 3c2f9ac64c
5 changed files with 581 additions and 45 deletions
+70
View File
@@ -164,6 +164,75 @@ h.event(1)
h.converge().unwrap();
```
## Prediction
`predict_outcome` gives the full distribution over finishing orders. Each entry
is a rank vector in the same shape `Outcome::ranking` takes — equal ranks mean a
tie — so an outcome feeds straight back into inference.
```rust
use trueskill_tt::History;
let mut h = History::builder().p_draw(0.1).build();
h.record_winner(&"alice", &"bob", 1).unwrap();
h.converge().unwrap();
let p = h.predict_outcome(&[&[&"alice"], &[&"bob"]]).unwrap();
// Probabilities are exhaustive and disjoint, so they sum to one.
assert!((p.total() - 1.0).abs() < 1e-6);
let (best, likelihood) = p.most_likely().unwrap();
println!("most likely: {best:?} at {likelihood:.3}");
println!("draw: {:.3}", p.probability_of(&[0, 0]));
```
Supports any number of teams. Because the outcome space grows factorially, the
full distribution is capped at `MAX_PREDICTED_TEAMS`; two cheaper entry points
stay available at any size:
- `predict_win_probabilities(teams)``P(team i finishes strictly first)`,
quadratic in team count.
- `predict_ranking(teams, ranks)` — one specific finishing order.
Unknown keys are an error, not a silent omission: a team the history has never
seen cannot produce a confident-looking probability.
## Which match to play next
`quality()` measures whether a matchup is *fair*. That is not the same as
whether it is *informative*, and the two only coincide for two evenly matched
competitors. When each observation costs something, ask
`expected_information_gain` instead — the outcome-weighted divergence between
what you believe now and what you would believe afterwards.
```rust
use trueskill_tt::History;
let mut h = History::builder().build();
for t in 1..=10 {
h.record_winner(&"veteran", &"regular", t).unwrap();
h.record_winner(&"regular", &"veteran", t + 100).unwrap();
}
h.record_winner(&"veteran", &"newcomer", 500).unwrap();
h.converge().unwrap();
let settled = h.expected_information_gain(&[&[&"veteran"], &[&"regular"]]).unwrap();
let unknown = h.expected_information_gain(&[&[&"veteran"], &[&"newcomer"]]).unwrap();
// Playing the newcomer teaches you more than replaying a settled rivalry.
assert!(unknown > settled);
```
The result is in nats, and is bounded by the entropy of the outcome: at most
`ln 2 ≈ 0.693` for a two-way result, `ln 3` once draws are possible, `ln k` for
`k` outcomes. A value near zero means you already know how it ends.
This costs one full inference pass **per possible outcome**, so it is far more
expensive than `quality()`. Scoring every pairing among `n` competitors is
`O(n² × outcomes)` passes — shortlist with `quality()` or
`predict_win_probabilities` first, then score only the shortlist.
## Todo
- [x] Implement approx for Gaussian
@@ -172,6 +241,7 @@ h.converge().unwrap();
- [x] Add examples (`examples/atp.rs`, `examples/scored.rs`)
- [x] Add Observer (`Observer` / `NullObserver`)
- [x] Benchmark the inference loop (`benches/batch.rs`, `benches/history_converge.rs`, `benches/ingest.rs`)
- [x] N-team `predict_outcome` with draw mass, and `expected_information_gain`
- [ ] Cross-check `quality()` against [sublee/trueskill](https://github.com/sublee/trueskill/tree/master) — N-group support works and is covered by invariants, but no reference values are asserted
## License