`EventBuilder` could set weights and nothing else, so `prior` and
`drift_scale` were reachable only through the typed
`Event`/`Team`/`Member` shape plus `add_events`. Which ingestion route a
competitor arrived through decided whether it could be configured.
`members(...)` takes `Member` values directly, so `Member`'s own builder
expresses everything. `team(...)` stays the common case.
One escape hatch rather than `priors` and `drift_scales` setters beside
`weights`, as the issue suggested and then argued against itself: a
parallel array per field means a parallel length check per field, and
each one is a new way to get the lengths wrong. `Member` already has a
builder; this just lets the fluent path reach it.
`record_winner`/`record_draw` are deliberately left alone. They are the
two-argument convenience path, and extending them would be a breaking
signature change. The issue's reason for wanting them extended has also
weakened: it said a competitor arriving through them was "permanently
stuck on the history defaults", and since 8c087ad that is no longer true
— a later `add_events` carrying the `Member` refits the whole history.
Measured, late configuration through that route reaches mu 40.000000000,
identical to configuring from the start.
Refs #37
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
308 lines
11 KiB
Markdown
308 lines
11 KiB
Markdown
# TrueSkill - Through Time
|
||
|
||
Rust port of [TrueSkillThroughTime.py](https://github.com/glandfried/TrueSkillThroughTime.py).
|
||
|
||
## Other implementations
|
||
|
||
- [ttt-scala](https://github.com/ankurdave/ttt-scala)
|
||
- [ChessAnalysis #F](https://github.com/lucasmaystre/ChessAnalysis)
|
||
- [TrueSkillThroughTime.jl](https://github.com/glandfried/TrueSkillThroughTime.jl)
|
||
- [TrueSkillThroughTime.R](https://github.com/glandfried/TrueSkillThroughTime.R)
|
||
- [TrueSkill Through Time: Revisiting the History of Chess](https://www.microsoft.com/en-us/research/wp-content/uploads/2008/01/NIPS2007_0931.pdf)
|
||
- [TrueSkill Through Time. The full scientific documentation](https://glandfried.github.io/publication/landfried2021-learning/)
|
||
|
||
## Drift
|
||
|
||
Skill drift models how a competitor's true skill can change between appearances.
|
||
Each time they reappear after a gap, their skill uncertainty is widened by the
|
||
drift model before the new evidence is incorporated.
|
||
|
||
Drift is represented by the `Drift` trait (`src/drift.rs`), generic over the
|
||
history's time type:
|
||
|
||
```text
|
||
pub trait Drift<T: Time>: Copy + Debug + Send + Sync {
|
||
fn variance_delta(&self, from: &T, to: &T) -> f64;
|
||
fn variance_for_elapsed(&self, elapsed: i64) -> f64;
|
||
}
|
||
```
|
||
|
||
Both methods return the amount to add to `σ²`, not to `σ`. `variance_delta`
|
||
works from two timestamps; `variance_for_elapsed` takes an already-computed
|
||
elapsed count, and is used on the paths that cache it. `Gaussian::forget`
|
||
applies the result entirely in variance space — `from_mv(mu, variance() +
|
||
variance_delta)` — taking no square root.
|
||
|
||
That block is a quotation rather than a doctest. The custom-drift example below
|
||
is compiled by CI, so it is what actually pins the signature.
|
||
|
||
### ConstantDrift
|
||
|
||
The built-in `ConstantDrift` implements a linear random walk — skill uncertainty
|
||
grows proportionally to time:
|
||
|
||
```text
|
||
variance_delta = elapsed * γ²
|
||
```
|
||
|
||
This is the standard TrueSkill Through Time model. Pass a `ConstantDrift(gamma)`
|
||
when constructing a `Rating`:
|
||
|
||
```rust
|
||
use trueskill_tt::{ConstantDrift, Gaussian, Rating};
|
||
|
||
// gamma = 0.1 means skill can shift ~0.1 per time unit.
|
||
let rating: Rating<i64, ConstantDrift> =
|
||
Rating::new(Gaussian::from_ms(0.0, 6.0), 1.0, ConstantDrift(0.1));
|
||
|
||
assert_eq!(rating.drift().0, 0.1);
|
||
```
|
||
|
||
The type annotation is load-bearing: `ConstantDrift` implements `Drift<T>` for
|
||
every `T: Time`, so without it `T` is ambiguous.
|
||
|
||
### Custom drift
|
||
|
||
Implement `Drift<T>` to express any other model. For example, a drift that
|
||
saturates after a long absence, with uncertainty growing as the square root of
|
||
elapsed time instead of linearly:
|
||
|
||
```rust
|
||
use trueskill_tt::{Drift, Gaussian, History, Rating, Time};
|
||
|
||
#[derive(Clone, Copy, Debug)]
|
||
struct SqrtDrift {
|
||
gamma: f64,
|
||
}
|
||
|
||
impl<T: Time> Drift<T> for SqrtDrift {
|
||
fn variance_delta(&self, from: &T, to: &T) -> f64 {
|
||
let elapsed = from.elapsed_to(to).max(0) as f64;
|
||
elapsed.sqrt() * self.gamma * self.gamma
|
||
}
|
||
|
||
fn variance_for_elapsed(&self, elapsed: i64) -> f64 {
|
||
(elapsed.max(0) as f64).sqrt() * self.gamma * self.gamma
|
||
}
|
||
}
|
||
|
||
// On a single Rating:
|
||
let rating: Rating<i64, SqrtDrift> =
|
||
Rating::new(Gaussian::from_ms(0.0, 6.0), 1.0, SqrtDrift { gamma: 0.5 });
|
||
|
||
// Or for a whole History, via the builder:
|
||
let history = History::builder().drift(SqrtDrift { gamma: 0.5 }).build();
|
||
|
||
assert_eq!(rating.beta(), 1.0);
|
||
assert_eq!(history.log_evidence(), 0.0);
|
||
```
|
||
|
||
`HistoryBuilder::drift` is the only way to set a history's drift model; there is
|
||
no `gamma()` shorthand. The default is `ConstantDrift(GAMMA)`.
|
||
|
||
### Per-competitor drift
|
||
|
||
A `History` has one drift model, but individual competitors can scale it.
|
||
`Member::with_drift_scale(s)` multiplies the drift *variance* that competitor
|
||
accumulates, so `s` is in the same units as `gamma`: `ConstantDrift(g)` at
|
||
scale `s` behaves exactly as `ConstantDrift(g * s)` would, for that competitor
|
||
alone.
|
||
|
||
`0.0` pins a competitor still. That is what makes a **fixed reference point**
|
||
expressible in the same graph as moving competitors — a bot at a known
|
||
strength, a rating floor, a course difficulty:
|
||
|
||
```rust
|
||
use trueskill_tt::{ConstantDrift, Event, History, Member, Outcome, Team};
|
||
|
||
let mut h = History::builder().drift(ConstantDrift(0.1)).build();
|
||
|
||
h.add_events(vec![Event {
|
||
time: 0,
|
||
teams: [
|
||
Team::with_members([Member::new("player")]),
|
||
// A course does not improve. Pin it, and the round's evidence
|
||
// lands on the player instead of being split between the two.
|
||
Team::with_members([Member::new("layout_7").with_drift_scale(0.0)]),
|
||
]
|
||
.into_iter()
|
||
.collect(),
|
||
outcome: Outcome::winner(0, 2),
|
||
}])
|
||
.unwrap();
|
||
|
||
h.converge().unwrap();
|
||
```
|
||
|
||
Like `with_prior`, the scale is **competitor configuration, not a per-event
|
||
value**: it applies to the competitor for the whole history, and it applies
|
||
whenever it is supplied — including on a key the history already knows.
|
||
Configuring one late still refits the whole history rather than taking effect
|
||
only from that event onward, because `converge` refits from competitor state.
|
||
Repeating the same value is inert; supplying two *different* values for one
|
||
competitor within a single batch is `InferenceError::ConflictingCompetitorConfig`,
|
||
since events in a batch have no order. The scale must be finite and
|
||
non-negative; ingestion otherwise fails with `InferenceError::InvalidParameter`.
|
||
|
||
The fluent `EventBuilder` reaches this too: `.team([...])` is the common case
|
||
and leaves both unset, while `.members([...])` takes `Member` values directly,
|
||
so `h.event(t).members([Member::new("layout_7").with_drift_scale(0.0)])` is
|
||
equivalent to the typed shape above.
|
||
|
||
## Scored outcomes
|
||
|
||
Use `Outcome::scores([...])` when you have continuous per-team scores rather
|
||
than just ranks. Adjacent score margins flow into a `MarginFactor` that adds
|
||
soft Gaussian evidence about the latent performance diff. Configure
|
||
`HistoryBuilder::score_sigma(σ)` to control how much you trust the margins
|
||
(smaller σ = more trust).
|
||
|
||
```rust
|
||
use trueskill_tt::History;
|
||
|
||
let mut h = History::builder().score_sigma(2.0).build();
|
||
h.event(1)
|
||
.team(["alice"])
|
||
.team(["bob"])
|
||
.scores([21.0, 9.0])
|
||
.commit()
|
||
.unwrap();
|
||
h.converge().unwrap();
|
||
```
|
||
|
||
## Prediction
|
||
|
||
`predict_outcome` gives the full distribution over finishing orders. Each entry
|
||
is a rank vector in the same shape `Outcome::ranking` takes — equal ranks mean a
|
||
tie — so an outcome feeds straight back into inference.
|
||
|
||
```rust
|
||
use trueskill_tt::History;
|
||
|
||
let mut h = History::builder().p_draw(0.1).build();
|
||
h.record_winner(&"alice", &"bob", 1).unwrap();
|
||
h.converge().unwrap();
|
||
|
||
let p = h.predict_outcome(&[&[&"alice"], &[&"bob"]]).unwrap();
|
||
|
||
// Probabilities are exhaustive and disjoint, so they sum to one.
|
||
assert!((p.total() - 1.0).abs() < 1e-6);
|
||
|
||
let (best, likelihood) = p.most_likely().unwrap();
|
||
println!("most likely: {best:?} at {likelihood:.3}");
|
||
println!("draw: {:.3}", p.probability_of(&[0, 0]));
|
||
```
|
||
|
||
Supports any number of teams. Because the outcome space grows factorially, the
|
||
full distribution is capped at `MAX_PREDICTED_TEAMS`; two cheaper entry points
|
||
stay available at any size:
|
||
|
||
- `predict_win_probabilities(teams)` — `P(team i finishes strictly first)`,
|
||
quadratic in team count.
|
||
- `predict_ranking(teams, ranks)` — one specific finishing order.
|
||
|
||
Unknown keys are an error by default, not a silent omission: a team the history
|
||
has never seen cannot produce a confident-looking probability. The error names
|
||
the key, and every key must already be known — pre-filter with `lookup` or
|
||
`current_skill` if your caller cannot guarantee that.
|
||
|
||
If predicting for competitors you have never seen is the point rather than a
|
||
mistake, say so once:
|
||
|
||
```rust
|
||
use trueskill_tt::{History, UnknownKeys};
|
||
|
||
let h = History::builder().unknown_keys(UnknownKeys::Prior).build();
|
||
```
|
||
|
||
An unknown competitor is then answered from the configured prior, which is the
|
||
honest reading — you have no evidence about them — and correctly *widens* a team
|
||
that contains one. There is deliberately no "skip the member" mode: a team's
|
||
performance is the sum of its members, so dropping one would make the model more
|
||
certain because it knows less.
|
||
|
||
### Asking about one competitor
|
||
|
||
`Gaussian` answers tail questions directly, which is what a stopping rule needs:
|
||
|
||
```rust
|
||
use trueskill_tt::History;
|
||
|
||
let mut h = History::builder().build();
|
||
h.record_winner(&"alice", &"bob", 1).unwrap();
|
||
let _ = h.converge().unwrap();
|
||
|
||
let skill = h.current_skill(&"alice").unwrap();
|
||
|
||
// "How sure am I that this is below the cutoff?" — a probability, not a
|
||
// `mu + z * sigma` band whose confidence drifts as sigma changes.
|
||
let _ = skill.probability_below(20.0);
|
||
|
||
// Use this rather than `1.0 - probability_below(x)`: the complement cancels
|
||
// away every digit in the upper tail, which is where a stopping rule lives.
|
||
let _ = skill.probability_above(30.0);
|
||
```
|
||
|
||
## Which match to play next
|
||
|
||
`quality()` measures whether a matchup is *fair*. That is not the same as
|
||
whether it is *informative*, and the two only coincide for two evenly matched
|
||
competitors. When each observation costs something, ask
|
||
`expected_information_gain` instead — the outcome-weighted divergence between
|
||
what you believe now and what you would believe afterwards.
|
||
|
||
```rust
|
||
use trueskill_tt::History;
|
||
|
||
let mut h = History::builder().build();
|
||
for t in 1..=10 {
|
||
h.record_winner(&"veteran", &"regular", t).unwrap();
|
||
h.record_winner(&"regular", &"veteran", t + 100).unwrap();
|
||
}
|
||
h.record_winner(&"veteran", &"newcomer", 500).unwrap();
|
||
h.converge().unwrap();
|
||
|
||
let settled = h.expected_information_gain(&[&[&"veteran"], &[&"regular"]]).unwrap();
|
||
let unknown = h.expected_information_gain(&[&[&"veteran"], &[&"newcomer"]]).unwrap();
|
||
|
||
// Playing the newcomer teaches you more than replaying a settled rivalry.
|
||
assert!(unknown > settled);
|
||
```
|
||
|
||
The result is in nats, and is bounded by the entropy of the outcome: at most
|
||
`ln 2 ≈ 0.693` for a two-way result, `ln 3` once draws are possible, `ln k` for
|
||
`k` outcomes. A value near zero means you already know how it ends.
|
||
|
||
This costs one full inference pass **per possible outcome**, so it is far more
|
||
expensive than `quality()`. Scoring every pairing among `n` competitors is
|
||
`O(n² × outcomes)` passes — shortlist with `quality()` or
|
||
`predict_win_probabilities` first, then score only the shortlist.
|
||
|
||
## Todo
|
||
|
||
- [x] Implement approx for Gaussian
|
||
- [x] Add more tests from `TrueSkillThroughTime.jl`
|
||
- [x] Generalise a time axis — `Time` is now a trait (`Untimed`, `i64`), not an enum
|
||
- [x] Add examples (`examples/atp.rs`, `examples/scored.rs`)
|
||
- [x] Add Observer (`Observer` / `NullObserver`)
|
||
- [x] Benchmark the inference loop (`benches/batch.rs`, `benches/history_converge.rs`, `benches/ingest.rs`)
|
||
- [x] N-team `predict_outcome` with draw mass, and `expected_information_gain`
|
||
- [x] Cross-check `quality()` against [sublee/trueskill](https://github.com/sublee/trueskill/tree/master) — N identical teams follow the closed form `(1/5)^((n-1)/2)` for the conventional parameters, asserted for n = 2..10, and the n=3/n=5 values (0.200, 0.040) match the reference package
|
||
|
||
## License
|
||
|
||
Licensed under either of
|
||
|
||
- Apache License, Version 2.0 ([LICENSE-APACHE](LICENSE-APACHE) or
|
||
<http://www.apache.org/licenses/LICENSE-2.0>)
|
||
- MIT license ([LICENSE-MIT](LICENSE-MIT) or
|
||
<http://opensource.org/licenses/MIT>)
|
||
|
||
at your option.
|
||
|
||
### Contribution
|
||
|
||
Unless you explicitly state otherwise, any contribution intentionally submitted
|
||
for inclusion in the work by you, as defined in the Apache-2.0 license, shall be
|
||
dual licensed as above, without any additional terms or conditions.
|