feat: add expected information gain for active matchup selection
`quality()` answers "is this matchup fair". Callers picking which
comparison to run next need "is this matchup informative", and the two
coincide only for two evenly matched competitors. Without a principled
alternative, downstream code was reaching for hand-rolled heuristics
like `quality * sigma_a^2 * sigma_b^2`, which double-counts uncertainty:
the two factors are not independent.
Adds `expected_information_gain`, the outcome-weighted divergence
between current beliefs and the beliefs each result would produce:
EIG = SUM P(outcome) * KL(posterior_after(outcome) || prior)
Available standalone over `Rating`s, and as
`History::expected_information_gain` using current skills and the
history's own beta, drift and p_draw — so the outcomes it weighs are the
ones that would actually be fitted.
This is the mutual information between the outcome and the skills, which
gives an analytic ceiling: gain cannot exceed the entropy of the thing
being observed, so at most `ln k` nats for k outcomes. That bound is the
sharpest test available, because an acquisition function is unusually
exposed to returning finite, plausible, monotone numbers while being
wrong — it would simply select slightly worse matchups forever. A
prototype of this returned 4.77 nats from a sign error while passing
every monotonicity check; `never_exceeds_the_entropy_of_the_outcome`
catches that class unconditionally.
Measured against the ceiling the values are meaningful rather than
vacuous: 0.382 nats for an even matchup between diffuse priors against
an 0.693 ceiling, falling to 0.013 for a lopsided one and 0.000 for a
hopeless one.
`disagrees_with_the_quality_times_variance_heuristic` pins down that
this is not a monotone transform of the heuristic it replaces — the two
rank a lopsided matchup and a confident even one in opposite orders — so
a later "simplification" cannot quietly revert to it.
Cost is one inference pass per possible outcome, documented on the
public API alongside the shortlist-then-score pattern, so callers do not
discover it in production.
Also folds the duplicated key-gathering in `predict_quality` and
`performances` into one validated `member_skills`.
Refs #39
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
This commit is contained in:
@@ -164,6 +164,75 @@ h.event(1)
|
||||
h.converge().unwrap();
|
||||
```
|
||||
|
||||
## Prediction
|
||||
|
||||
`predict_outcome` gives the full distribution over finishing orders. Each entry
|
||||
is a rank vector in the same shape `Outcome::ranking` takes — equal ranks mean a
|
||||
tie — so an outcome feeds straight back into inference.
|
||||
|
||||
```rust
|
||||
use trueskill_tt::History;
|
||||
|
||||
let mut h = History::builder().p_draw(0.1).build();
|
||||
h.record_winner(&"alice", &"bob", 1).unwrap();
|
||||
h.converge().unwrap();
|
||||
|
||||
let p = h.predict_outcome(&[&[&"alice"], &[&"bob"]]).unwrap();
|
||||
|
||||
// Probabilities are exhaustive and disjoint, so they sum to one.
|
||||
assert!((p.total() - 1.0).abs() < 1e-6);
|
||||
|
||||
let (best, likelihood) = p.most_likely().unwrap();
|
||||
println!("most likely: {best:?} at {likelihood:.3}");
|
||||
println!("draw: {:.3}", p.probability_of(&[0, 0]));
|
||||
```
|
||||
|
||||
Supports any number of teams. Because the outcome space grows factorially, the
|
||||
full distribution is capped at `MAX_PREDICTED_TEAMS`; two cheaper entry points
|
||||
stay available at any size:
|
||||
|
||||
- `predict_win_probabilities(teams)` — `P(team i finishes strictly first)`,
|
||||
quadratic in team count.
|
||||
- `predict_ranking(teams, ranks)` — one specific finishing order.
|
||||
|
||||
Unknown keys are an error, not a silent omission: a team the history has never
|
||||
seen cannot produce a confident-looking probability.
|
||||
|
||||
## Which match to play next
|
||||
|
||||
`quality()` measures whether a matchup is *fair*. That is not the same as
|
||||
whether it is *informative*, and the two only coincide for two evenly matched
|
||||
competitors. When each observation costs something, ask
|
||||
`expected_information_gain` instead — the outcome-weighted divergence between
|
||||
what you believe now and what you would believe afterwards.
|
||||
|
||||
```rust
|
||||
use trueskill_tt::History;
|
||||
|
||||
let mut h = History::builder().build();
|
||||
for t in 1..=10 {
|
||||
h.record_winner(&"veteran", &"regular", t).unwrap();
|
||||
h.record_winner(&"regular", &"veteran", t + 100).unwrap();
|
||||
}
|
||||
h.record_winner(&"veteran", &"newcomer", 500).unwrap();
|
||||
h.converge().unwrap();
|
||||
|
||||
let settled = h.expected_information_gain(&[&[&"veteran"], &[&"regular"]]).unwrap();
|
||||
let unknown = h.expected_information_gain(&[&[&"veteran"], &[&"newcomer"]]).unwrap();
|
||||
|
||||
// Playing the newcomer teaches you more than replaying a settled rivalry.
|
||||
assert!(unknown > settled);
|
||||
```
|
||||
|
||||
The result is in nats, and is bounded by the entropy of the outcome: at most
|
||||
`ln 2 ≈ 0.693` for a two-way result, `ln 3` once draws are possible, `ln k` for
|
||||
`k` outcomes. A value near zero means you already know how it ends.
|
||||
|
||||
This costs one full inference pass **per possible outcome**, so it is far more
|
||||
expensive than `quality()`. Scoring every pairing among `n` competitors is
|
||||
`O(n² × outcomes)` passes — shortlist with `quality()` or
|
||||
`predict_win_probabilities` first, then score only the shortlist.
|
||||
|
||||
## Todo
|
||||
|
||||
- [x] Implement approx for Gaussian
|
||||
@@ -172,6 +241,7 @@ h.converge().unwrap();
|
||||
- [x] Add examples (`examples/atp.rs`, `examples/scored.rs`)
|
||||
- [x] Add Observer (`Observer` / `NullObserver`)
|
||||
- [x] Benchmark the inference loop (`benches/batch.rs`, `benches/history_converge.rs`, `benches/ingest.rs`)
|
||||
- [x] N-team `predict_outcome` with draw mass, and `expected_information_gain`
|
||||
- [ ] Cross-check `quality()` against [sublee/trueskill](https://github.com/sublee/trueskill/tree/master) — N-group support works and is covered by invariants, but no reference values are asserted
|
||||
|
||||
## License
|
||||
|
||||
Reference in New Issue
Block a user