# TrueSkill - Through Time Rust port of [TrueSkillThroughTime.py](https://github.com/glandfried/TrueSkillThroughTime.py). ## Other implementations - [ttt-scala](https://github.com/ankurdave/ttt-scala) - [ChessAnalysis #F](https://github.com/lucasmaystre/ChessAnalysis) - [TrueSkillThroughTime.jl](https://github.com/glandfried/TrueSkillThroughTime.jl) - [TrueSkillThroughTime.R](https://github.com/glandfried/TrueSkillThroughTime.R) - [TrueSkill Through Time: Revisiting the History of Chess](https://www.microsoft.com/en-us/research/wp-content/uploads/2008/01/NIPS2007_0931.pdf) - [TrueSkill Through Time. The full scientific documentation](https://glandfried.github.io/publication/landfried2021-learning/) ## Drift Skill drift models how a competitor's true skill can change between appearances. Each time they reappear after a gap, their skill uncertainty is widened by the drift model before the new evidence is incorporated. Drift is represented by the `Drift` trait (`src/drift.rs`), generic over the history's time type: ```text pub trait Drift: Copy + Debug + Send + Sync { fn variance_delta(&self, from: &T, to: &T) -> f64; fn variance_for_elapsed(&self, elapsed: i64) -> f64; } ``` Both methods return the amount to add to `σ²`, not to `σ`. `variance_delta` works from two timestamps; `variance_for_elapsed` takes an already-computed elapsed count, and is used on the paths that cache it. `Gaussian::forget` applies the result entirely in variance space — `from_mv(mu, variance() + variance_delta)` — taking no square root. That block is a quotation rather than a doctest. The custom-drift example below is compiled by CI, so it is what actually pins the signature. ### ConstantDrift The built-in `ConstantDrift` implements a linear random walk — skill uncertainty grows proportionally to time: ```text variance_delta = elapsed * γ² ``` This is the standard TrueSkill Through Time model. Pass a `ConstantDrift(gamma)` when constructing a `Rating`: ```rust use trueskill_tt::{ConstantDrift, Gaussian, Rating}; // gamma = 0.1 means skill can shift ~0.1 per time unit. let rating: Rating = Rating::new(Gaussian::from_ms(0.0, 6.0), 1.0, ConstantDrift(0.1)); assert_eq!(rating.drift().0, 0.1); ``` The type annotation is load-bearing: `ConstantDrift` implements `Drift` for every `T: Time`, so without it `T` is ambiguous. ### Custom drift Implement `Drift` to express any other model. For example, a drift that saturates after a long absence, with uncertainty growing as the square root of elapsed time instead of linearly: ```rust use trueskill_tt::{Drift, Gaussian, History, Rating, Time}; #[derive(Clone, Copy, Debug)] struct SqrtDrift { gamma: f64, } impl Drift for SqrtDrift { fn variance_delta(&self, from: &T, to: &T) -> f64 { let elapsed = from.elapsed_to(to).max(0) as f64; elapsed.sqrt() * self.gamma * self.gamma } fn variance_for_elapsed(&self, elapsed: i64) -> f64 { (elapsed.max(0) as f64).sqrt() * self.gamma * self.gamma } } // On a single Rating: let rating: Rating = Rating::new(Gaussian::from_ms(0.0, 6.0), 1.0, SqrtDrift { gamma: 0.5 }); // Or for a whole History, via the builder: let history = History::builder().drift(SqrtDrift { gamma: 0.5 }).build(); assert_eq!(rating.beta(), 1.0); assert_eq!(history.log_evidence(), 0.0); ``` `HistoryBuilder::drift` is the only way to set a history's drift model; there is no `gamma()` shorthand. The default is `ConstantDrift(GAMMA)`. ### Per-competitor drift A `History` has one drift model, but individual competitors can scale it. `Member::with_drift_scale(s)` multiplies the drift *variance* that competitor accumulates, so `s` is in the same units as `gamma`: `ConstantDrift(g)` at scale `s` behaves exactly as `ConstantDrift(g * s)` would, for that competitor alone. `0.0` pins a competitor still. That is what makes a **fixed reference point** expressible in the same graph as moving competitors — a bot at a known strength, a rating floor, a course difficulty: ```rust use trueskill_tt::{ConstantDrift, Event, History, Member, Outcome, Team}; let mut h = History::builder().drift(ConstantDrift(0.1)).build(); h.add_events(vec![Event { time: 0, teams: [ Team::with_members([Member::new("player")]), // A course does not improve. Pin it, and the round's evidence // lands on the player instead of being split between the two. Team::with_members([Member::new("layout_7").with_drift_scale(0.0)]), ] .into_iter() .collect(), outcome: Outcome::winner(0, 2), }]) .unwrap(); h.converge().unwrap(); ``` Like `with_prior`, the scale is **competitor configuration, not a per-event value**: it applies to the competitor for the whole history, and it applies whenever it is supplied — including on a key the history already knows. Configuring one late still refits the whole history rather than taking effect only from that event onward, because `converge` refits from competitor state. Repeating the same value is inert; supplying two *different* values for one competitor within a single batch is `InferenceError::ConflictingCompetitorConfig`, since events in a batch have no order. The scale must be finite and non-negative; ingestion otherwise fails with `InferenceError::InvalidParameter`. Note that the fluent `EventBuilder` (`h.event(t).team([...])`) sets weights but not `drift_scale` or `prior`; those need the typed `Event` / `Team` / `Member` shape shown above. ## Scored outcomes Use `Outcome::scores([...])` when you have continuous per-team scores rather than just ranks. Adjacent score margins flow into a `MarginFactor` that adds soft Gaussian evidence about the latent performance diff. Configure `HistoryBuilder::score_sigma(σ)` to control how much you trust the margins (smaller σ = more trust). ```rust use trueskill_tt::History; let mut h = History::builder().score_sigma(2.0).build(); h.event(1) .team(["alice"]) .team(["bob"]) .scores([21.0, 9.0]) .commit() .unwrap(); h.converge().unwrap(); ``` ## Prediction `predict_outcome` gives the full distribution over finishing orders. Each entry is a rank vector in the same shape `Outcome::ranking` takes — equal ranks mean a tie — so an outcome feeds straight back into inference. ```rust use trueskill_tt::History; let mut h = History::builder().p_draw(0.1).build(); h.record_winner(&"alice", &"bob", 1).unwrap(); h.converge().unwrap(); let p = h.predict_outcome(&[&[&"alice"], &[&"bob"]]).unwrap(); // Probabilities are exhaustive and disjoint, so they sum to one. assert!((p.total() - 1.0).abs() < 1e-6); let (best, likelihood) = p.most_likely().unwrap(); println!("most likely: {best:?} at {likelihood:.3}"); println!("draw: {:.3}", p.probability_of(&[0, 0])); ``` Supports any number of teams. Because the outcome space grows factorially, the full distribution is capped at `MAX_PREDICTED_TEAMS`; two cheaper entry points stay available at any size: - `predict_win_probabilities(teams)` — `P(team i finishes strictly first)`, quadratic in team count. - `predict_ranking(teams, ranks)` — one specific finishing order. Unknown keys are an error by default, not a silent omission: a team the history has never seen cannot produce a confident-looking probability. The error names the key, and every key must already be known — pre-filter with `lookup` or `current_skill` if your caller cannot guarantee that. If predicting for competitors you have never seen is the point rather than a mistake, say so once: ```rust use trueskill_tt::{History, UnknownKeys}; let h = History::builder().unknown_keys(UnknownKeys::Prior).build(); ``` An unknown competitor is then answered from the configured prior, which is the honest reading — you have no evidence about them — and correctly *widens* a team that contains one. There is deliberately no "skip the member" mode: a team's performance is the sum of its members, so dropping one would make the model more certain because it knows less. ### Asking about one competitor `Gaussian` answers tail questions directly, which is what a stopping rule needs: ```rust use trueskill_tt::History; let mut h = History::builder().build(); h.record_winner(&"alice", &"bob", 1).unwrap(); let _ = h.converge().unwrap(); let skill = h.current_skill(&"alice").unwrap(); // "How sure am I that this is below the cutoff?" — a probability, not a // `mu + z * sigma` band whose confidence drifts as sigma changes. let _ = skill.probability_below(20.0); // Use this rather than `1.0 - probability_below(x)`: the complement cancels // away every digit in the upper tail, which is where a stopping rule lives. let _ = skill.probability_above(30.0); ``` ## Which match to play next `quality()` measures whether a matchup is *fair*. That is not the same as whether it is *informative*, and the two only coincide for two evenly matched competitors. When each observation costs something, ask `expected_information_gain` instead — the outcome-weighted divergence between what you believe now and what you would believe afterwards. ```rust use trueskill_tt::History; let mut h = History::builder().build(); for t in 1..=10 { h.record_winner(&"veteran", &"regular", t).unwrap(); h.record_winner(&"regular", &"veteran", t + 100).unwrap(); } h.record_winner(&"veteran", &"newcomer", 500).unwrap(); h.converge().unwrap(); let settled = h.expected_information_gain(&[&[&"veteran"], &[&"regular"]]).unwrap(); let unknown = h.expected_information_gain(&[&[&"veteran"], &[&"newcomer"]]).unwrap(); // Playing the newcomer teaches you more than replaying a settled rivalry. assert!(unknown > settled); ``` The result is in nats, and is bounded by the entropy of the outcome: at most `ln 2 ≈ 0.693` for a two-way result, `ln 3` once draws are possible, `ln k` for `k` outcomes. A value near zero means you already know how it ends. This costs one full inference pass **per possible outcome**, so it is far more expensive than `quality()`. Scoring every pairing among `n` competitors is `O(n² × outcomes)` passes — shortlist with `quality()` or `predict_win_probabilities` first, then score only the shortlist. ## Todo - [x] Implement approx for Gaussian - [x] Add more tests from `TrueSkillThroughTime.jl` - [x] Generalise a time axis — `Time` is now a trait (`Untimed`, `i64`), not an enum - [x] Add examples (`examples/atp.rs`, `examples/scored.rs`) - [x] Add Observer (`Observer` / `NullObserver`) - [x] Benchmark the inference loop (`benches/batch.rs`, `benches/history_converge.rs`, `benches/ingest.rs`) - [x] N-team `predict_outcome` with draw mass, and `expected_information_gain` - [x] Cross-check `quality()` against [sublee/trueskill](https://github.com/sublee/trueskill/tree/master) — N identical teams follow the closed form `(1/5)^((n-1)/2)` for the conventional parameters, asserted for n = 2..10, and the n=3/n=5 values (0.200, 0.040) match the reference package ## License Licensed under either of - Apache License, Version 2.0 ([LICENSE-APACHE](LICENSE-APACHE) or ) - MIT license ([LICENSE-MIT](LICENSE-MIT) or ) at your option. ### Contribution Unless you explicitly state otherwise, any contribution intentionally submitted for inclusion in the work by you, as defined in the Apache-2.0 license, shall be dual licensed as above, without any additional terms or conditions.