`quality()` answers "is this matchup fair". Callers picking which
comparison to run next need "is this matchup informative", and the two
coincide only for two evenly matched competitors. Without a principled
alternative, downstream code was reaching for hand-rolled heuristics
like `quality * sigma_a^2 * sigma_b^2`, which double-counts uncertainty:
the two factors are not independent.
Adds `expected_information_gain`, the outcome-weighted divergence
between current beliefs and the beliefs each result would produce:
EIG = SUM P(outcome) * KL(posterior_after(outcome) || prior)
Available standalone over `Rating`s, and as
`History::expected_information_gain` using current skills and the
history's own beta, drift and p_draw — so the outcomes it weighs are the
ones that would actually be fitted.
This is the mutual information between the outcome and the skills, which
gives an analytic ceiling: gain cannot exceed the entropy of the thing
being observed, so at most `ln k` nats for k outcomes. That bound is the
sharpest test available, because an acquisition function is unusually
exposed to returning finite, plausible, monotone numbers while being
wrong — it would simply select slightly worse matchups forever. A
prototype of this returned 4.77 nats from a sign error while passing
every monotonicity check; `never_exceeds_the_entropy_of_the_outcome`
catches that class unconditionally.
Measured against the ceiling the values are meaningful rather than
vacuous: 0.382 nats for an even matchup between diffuse priors against
an 0.693 ceiling, falling to 0.013 for a lopsided one and 0.000 for a
hopeless one.
`disagrees_with_the_quality_times_variance_heuristic` pins down that
this is not a monotone transform of the heuristic it replaces — the two
rank a lopsided matchup and a confident even one in opposite orders — so
a later "simplification" cannot quietly revert to it.
Cost is one inference pass per possible outcome, documented on the
public API alongside the shortlist-then-score pattern, so callers do not
discover it in production.
Also folds the duplicated key-gathering in `predict_quality` and
`performances` into one validated `member_skills`.
Refs #39
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
263 lines
9.4 KiB
Markdown
263 lines
9.4 KiB
Markdown
# TrueSkill - Through Time
|
||
|
||
Rust port of [TrueSkillThroughTime.py](https://github.com/glandfried/TrueSkillThroughTime.py).
|
||
|
||
## Other implementations
|
||
|
||
- [ttt-scala](https://github.com/ankurdave/ttt-scala)
|
||
- [ChessAnalysis #F](https://github.com/lucasmaystre/ChessAnalysis)
|
||
- [TrueSkillThroughTime.jl](https://github.com/glandfried/TrueSkillThroughTime.jl)
|
||
- [TrueSkillThroughTime.R](https://github.com/glandfried/TrueSkillThroughTime.R)
|
||
- [TrueSkill Through Time: Revisiting the History of Chess](https://www.microsoft.com/en-us/research/wp-content/uploads/2008/01/NIPS2007_0931.pdf)
|
||
- [TrueSkill Through Time. The full scientific documentation](https://glandfried.github.io/publication/landfried2021-learning/)
|
||
|
||
## Drift
|
||
|
||
Skill drift models how a competitor's true skill can change between appearances.
|
||
Each time they reappear after a gap, their skill uncertainty is widened by the
|
||
drift model before the new evidence is incorporated.
|
||
|
||
Drift is represented by the `Drift` trait (`src/drift.rs`), generic over the
|
||
history's time type:
|
||
|
||
```text
|
||
pub trait Drift<T: Time>: Copy + Debug + Send + Sync {
|
||
fn variance_delta(&self, from: &T, to: &T) -> f64;
|
||
fn variance_for_elapsed(&self, elapsed: i64) -> f64;
|
||
}
|
||
```
|
||
|
||
Both methods return the amount to add to `σ²`, not to `σ`. `variance_delta`
|
||
works from two timestamps; `variance_for_elapsed` takes an already-computed
|
||
elapsed count, and is used on the paths that cache it. `Gaussian::forget`
|
||
applies the result entirely in variance space — `from_mv(mu, variance() +
|
||
variance_delta)` — taking no square root.
|
||
|
||
That block is a quotation rather than a doctest. The custom-drift example below
|
||
is compiled by CI, so it is what actually pins the signature.
|
||
|
||
### ConstantDrift
|
||
|
||
The built-in `ConstantDrift` implements a linear random walk — skill uncertainty
|
||
grows proportionally to time:
|
||
|
||
```text
|
||
variance_delta = elapsed * γ²
|
||
```
|
||
|
||
This is the standard TrueSkill Through Time model. Pass a `ConstantDrift(gamma)`
|
||
when constructing a `Rating`:
|
||
|
||
```rust
|
||
use trueskill_tt::{ConstantDrift, Gaussian, Rating};
|
||
|
||
// gamma = 0.1 means skill can shift ~0.1 per time unit.
|
||
let rating: Rating<i64, ConstantDrift> =
|
||
Rating::new(Gaussian::from_ms(0.0, 6.0), 1.0, ConstantDrift(0.1));
|
||
|
||
assert_eq!(rating.drift().0, 0.1);
|
||
```
|
||
|
||
The type annotation is load-bearing: `ConstantDrift` implements `Drift<T>` for
|
||
every `T: Time`, so without it `T` is ambiguous.
|
||
|
||
### Custom drift
|
||
|
||
Implement `Drift<T>` to express any other model. For example, a drift that
|
||
saturates after a long absence, with uncertainty growing as the square root of
|
||
elapsed time instead of linearly:
|
||
|
||
```rust
|
||
use trueskill_tt::{Drift, Gaussian, History, Rating, Time};
|
||
|
||
#[derive(Clone, Copy, Debug)]
|
||
struct SqrtDrift {
|
||
gamma: f64,
|
||
}
|
||
|
||
impl<T: Time> Drift<T> for SqrtDrift {
|
||
fn variance_delta(&self, from: &T, to: &T) -> f64 {
|
||
let elapsed = from.elapsed_to(to).max(0) as f64;
|
||
elapsed.sqrt() * self.gamma * self.gamma
|
||
}
|
||
|
||
fn variance_for_elapsed(&self, elapsed: i64) -> f64 {
|
||
(elapsed.max(0) as f64).sqrt() * self.gamma * self.gamma
|
||
}
|
||
}
|
||
|
||
// On a single Rating:
|
||
let rating: Rating<i64, SqrtDrift> =
|
||
Rating::new(Gaussian::from_ms(0.0, 6.0), 1.0, SqrtDrift { gamma: 0.5 });
|
||
|
||
// Or for a whole History, via the builder:
|
||
let history = History::builder().drift(SqrtDrift { gamma: 0.5 }).build();
|
||
|
||
assert_eq!(rating.beta(), 1.0);
|
||
assert_eq!(history.log_evidence(), 0.0);
|
||
```
|
||
|
||
`HistoryBuilder::drift` is the only way to set a history's drift model; there is
|
||
no `gamma()` shorthand. The default is `ConstantDrift(GAMMA)`.
|
||
|
||
### Per-competitor drift
|
||
|
||
A `History` has one drift model, but individual competitors can scale it.
|
||
`Member::with_drift_scale(s)` multiplies the drift *variance* that competitor
|
||
accumulates, so `s` is in the same units as `gamma`: `ConstantDrift(g)` at
|
||
scale `s` behaves exactly as `ConstantDrift(g * s)` would, for that competitor
|
||
alone.
|
||
|
||
`0.0` pins a competitor still. That is what makes a **fixed reference point**
|
||
expressible in the same graph as moving competitors — a bot at a known
|
||
strength, a rating floor, a course difficulty:
|
||
|
||
```rust
|
||
use trueskill_tt::{ConstantDrift, Event, History, Member, Outcome, Team};
|
||
|
||
let mut h = History::builder().drift(ConstantDrift(0.1)).build();
|
||
|
||
h.add_events(vec![Event {
|
||
time: 0,
|
||
teams: [
|
||
Team::with_members([Member::new("player")]),
|
||
// A course does not improve. Pin it, and the round's evidence
|
||
// lands on the player instead of being split between the two.
|
||
Team::with_members([Member::new("layout_7").with_drift_scale(0.0)]),
|
||
]
|
||
.into_iter()
|
||
.collect(),
|
||
outcome: Outcome::winner(0, 2),
|
||
}])
|
||
.unwrap();
|
||
|
||
h.converge().unwrap();
|
||
```
|
||
|
||
Like `with_prior`, the scale is **competitor configuration captured at first
|
||
appearance** — setting it on a key the history already knows has no effect. It
|
||
must be finite and non-negative; ingestion otherwise fails with
|
||
`InferenceError::InvalidParameter`.
|
||
|
||
Note that the fluent `EventBuilder` (`h.event(t).team([...])`) sets weights but
|
||
not `drift_scale` or `prior`; those need the typed `Event` / `Team` / `Member`
|
||
shape shown above.
|
||
|
||
## Scored outcomes
|
||
|
||
Use `Outcome::scores([...])` when you have continuous per-team scores rather
|
||
than just ranks. Adjacent score margins flow into a `MarginFactor` that adds
|
||
soft Gaussian evidence about the latent performance diff. Configure
|
||
`HistoryBuilder::score_sigma(σ)` to control how much you trust the margins
|
||
(smaller σ = more trust).
|
||
|
||
```rust
|
||
use trueskill_tt::History;
|
||
|
||
let mut h = History::builder().score_sigma(2.0).build();
|
||
h.event(1)
|
||
.team(["alice"])
|
||
.team(["bob"])
|
||
.scores([21.0, 9.0])
|
||
.commit()
|
||
.unwrap();
|
||
h.converge().unwrap();
|
||
```
|
||
|
||
## Prediction
|
||
|
||
`predict_outcome` gives the full distribution over finishing orders. Each entry
|
||
is a rank vector in the same shape `Outcome::ranking` takes — equal ranks mean a
|
||
tie — so an outcome feeds straight back into inference.
|
||
|
||
```rust
|
||
use trueskill_tt::History;
|
||
|
||
let mut h = History::builder().p_draw(0.1).build();
|
||
h.record_winner(&"alice", &"bob", 1).unwrap();
|
||
h.converge().unwrap();
|
||
|
||
let p = h.predict_outcome(&[&[&"alice"], &[&"bob"]]).unwrap();
|
||
|
||
// Probabilities are exhaustive and disjoint, so they sum to one.
|
||
assert!((p.total() - 1.0).abs() < 1e-6);
|
||
|
||
let (best, likelihood) = p.most_likely().unwrap();
|
||
println!("most likely: {best:?} at {likelihood:.3}");
|
||
println!("draw: {:.3}", p.probability_of(&[0, 0]));
|
||
```
|
||
|
||
Supports any number of teams. Because the outcome space grows factorially, the
|
||
full distribution is capped at `MAX_PREDICTED_TEAMS`; two cheaper entry points
|
||
stay available at any size:
|
||
|
||
- `predict_win_probabilities(teams)` — `P(team i finishes strictly first)`,
|
||
quadratic in team count.
|
||
- `predict_ranking(teams, ranks)` — one specific finishing order.
|
||
|
||
Unknown keys are an error, not a silent omission: a team the history has never
|
||
seen cannot produce a confident-looking probability.
|
||
|
||
## Which match to play next
|
||
|
||
`quality()` measures whether a matchup is *fair*. That is not the same as
|
||
whether it is *informative*, and the two only coincide for two evenly matched
|
||
competitors. When each observation costs something, ask
|
||
`expected_information_gain` instead — the outcome-weighted divergence between
|
||
what you believe now and what you would believe afterwards.
|
||
|
||
```rust
|
||
use trueskill_tt::History;
|
||
|
||
let mut h = History::builder().build();
|
||
for t in 1..=10 {
|
||
h.record_winner(&"veteran", &"regular", t).unwrap();
|
||
h.record_winner(&"regular", &"veteran", t + 100).unwrap();
|
||
}
|
||
h.record_winner(&"veteran", &"newcomer", 500).unwrap();
|
||
h.converge().unwrap();
|
||
|
||
let settled = h.expected_information_gain(&[&[&"veteran"], &[&"regular"]]).unwrap();
|
||
let unknown = h.expected_information_gain(&[&[&"veteran"], &[&"newcomer"]]).unwrap();
|
||
|
||
// Playing the newcomer teaches you more than replaying a settled rivalry.
|
||
assert!(unknown > settled);
|
||
```
|
||
|
||
The result is in nats, and is bounded by the entropy of the outcome: at most
|
||
`ln 2 ≈ 0.693` for a two-way result, `ln 3` once draws are possible, `ln k` for
|
||
`k` outcomes. A value near zero means you already know how it ends.
|
||
|
||
This costs one full inference pass **per possible outcome**, so it is far more
|
||
expensive than `quality()`. Scoring every pairing among `n` competitors is
|
||
`O(n² × outcomes)` passes — shortlist with `quality()` or
|
||
`predict_win_probabilities` first, then score only the shortlist.
|
||
|
||
## Todo
|
||
|
||
- [x] Implement approx for Gaussian
|
||
- [x] Add more tests from `TrueSkillThroughTime.jl`
|
||
- [x] Generalise a time axis — `Time` is now a trait (`Untimed`, `i64`), not an enum
|
||
- [x] Add examples (`examples/atp.rs`, `examples/scored.rs`)
|
||
- [x] Add Observer (`Observer` / `NullObserver`)
|
||
- [x] Benchmark the inference loop (`benches/batch.rs`, `benches/history_converge.rs`, `benches/ingest.rs`)
|
||
- [x] N-team `predict_outcome` with draw mass, and `expected_information_gain`
|
||
- [ ] Cross-check `quality()` against [sublee/trueskill](https://github.com/sublee/trueskill/tree/master) — N-group support works and is covered by invariants, but no reference values are asserted
|
||
|
||
## License
|
||
|
||
Licensed under either of
|
||
|
||
- Apache License, Version 2.0 ([LICENSE-APACHE](LICENSE-APACHE) or
|
||
<http://www.apache.org/licenses/LICENSE-2.0>)
|
||
- MIT license ([LICENSE-MIT](LICENSE-MIT) or
|
||
<http://opensource.org/licenses/MIT>)
|
||
|
||
at your option.
|
||
|
||
### Contribution
|
||
|
||
Unless you explicitly state otherwise, any contribution intentionally submitted
|
||
for inclusion in the work by you, as defined in the Apache-2.0 license, shall be
|
||
dual licensed as above, without any additional terms or conditions.
|