logaritmiskandClaude Opus 5 c12bc830a5 feat!: name the unknown key, expose tail probabilities, flag short fits
Three issues from two downstream consumers, all small, all sharing a
theme: the crate had the information and would not hand it over.

#44 — `UnknownKey { team: 0, member: 0 }` did not say which key. A
consumer upgrading 0.1.2 -> 0.4.1 had every one of 5591 predictions
return this error, fell back to a neutral 0.5, and lost its entire
metadata model for a day. Nothing crashed and nothing logged; it was
found by sweeping an unrelated parameter and noticing the output did not
move. The 0.4.0 change that made unknown keys an error was right — the
error was just too anonymous to act on. It now carries the key's `Debug`
rendering, and its `Display` says what to do about it. The precondition
is documented on every prediction entry point, which the reporter said
would alone have saved the day.

#43 — `cdf` was `pub(crate)`, so a consumer asking "is this competitor
below the cutoff" approximated it with a `mu + z * sigma` band and had no
way to say what confidence any `z` bought. Adds
`Gaussian::probability_below` / `probability_above`. The second is
separate on purpose: `1 - cdf` collapses to exactly zero past ~8.3
sigma, and a stopping rule is evaluated precisely there. Both route
through the survival function added in 0.4.1, so this is visibility
rather than new numerics.

#50 — `ConvergenceReport` was not `#[must_use]`, so the one signal that
a fit stopped short was trivially discarded. It now is, and that
immediately found 78 sites doing exactly that — including this crate's
own ATP example, which was capped at 10 sweeps when the history needs
30. The example now reads the report and says so.

`ITERATIONS = 30` is documented as the floor it is, with the three
measurements to hand: 400 events over 100 competitors already stops
there at ~7e-3 against a 1e-6 tolerance, the ATP example needs 30 at a
much looser one, and a consumer's 2000-node model needs 76 to 161.

BREAKING CHANGE: `InferenceError::UnknownKey` gains a `key` field, and
the prediction methods now require `K: Debug` in order to fill it.

Closes #43, #50. Refs #44 — its third ask, an opt-in `UnknownKeys::Skip`
mode, is a live API question and deliberately not answered here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-07 23:41:03 +02:00

TrueSkill - Through Time

Rust port of TrueSkillThroughTime.py.

Other implementations

Drift

Skill drift models how a competitor's true skill can change between appearances. Each time they reappear after a gap, their skill uncertainty is widened by the drift model before the new evidence is incorporated.

Drift is represented by the Drift trait (src/drift.rs), generic over the history's time type:

pub trait Drift<T: Time>: Copy + Debug + Send + Sync {
    fn variance_delta(&self, from: &T, to: &T) -> f64;
    fn variance_for_elapsed(&self, elapsed: i64) -> f64;
}

Both methods return the amount to add to σ², not to σ. variance_delta works from two timestamps; variance_for_elapsed takes an already-computed elapsed count, and is used on the paths that cache it. Gaussian::forget applies the result entirely in variance space — from_mv(mu, variance() + variance_delta) — taking no square root.

That block is a quotation rather than a doctest. The custom-drift example below is compiled by CI, so it is what actually pins the signature.

ConstantDrift

The built-in ConstantDrift implements a linear random walk — skill uncertainty grows proportionally to time:

variance_delta = elapsed * γ²

This is the standard TrueSkill Through Time model. Pass a ConstantDrift(gamma) when constructing a Rating:

use trueskill_tt::{ConstantDrift, Gaussian, Rating};

// gamma = 0.1 means skill can shift ~0.1 per time unit.
let rating: Rating<i64, ConstantDrift> =
    Rating::new(Gaussian::from_ms(0.0, 6.0), 1.0, ConstantDrift(0.1));

assert_eq!(rating.drift().0, 0.1);

The type annotation is load-bearing: ConstantDrift implements Drift<T> for every T: Time, so without it T is ambiguous.

Custom drift

Implement Drift<T> to express any other model. For example, a drift that saturates after a long absence, with uncertainty growing as the square root of elapsed time instead of linearly:

use trueskill_tt::{Drift, Gaussian, History, Rating, Time};

#[derive(Clone, Copy, Debug)]
struct SqrtDrift {
    gamma: f64,
}

impl<T: Time> Drift<T> for SqrtDrift {
    fn variance_delta(&self, from: &T, to: &T) -> f64 {
        let elapsed = from.elapsed_to(to).max(0) as f64;
        elapsed.sqrt() * self.gamma * self.gamma
    }

    fn variance_for_elapsed(&self, elapsed: i64) -> f64 {
        (elapsed.max(0) as f64).sqrt() * self.gamma * self.gamma
    }
}

// On a single Rating:
let rating: Rating<i64, SqrtDrift> =
    Rating::new(Gaussian::from_ms(0.0, 6.0), 1.0, SqrtDrift { gamma: 0.5 });

// Or for a whole History, via the builder:
let history = History::builder().drift(SqrtDrift { gamma: 0.5 }).build();

assert_eq!(rating.beta(), 1.0);
assert_eq!(history.log_evidence(), 0.0);

HistoryBuilder::drift is the only way to set a history's drift model; there is no gamma() shorthand. The default is ConstantDrift(GAMMA).

Per-competitor drift

A History has one drift model, but individual competitors can scale it. Member::with_drift_scale(s) multiplies the drift variance that competitor accumulates, so s is in the same units as gamma: ConstantDrift(g) at scale s behaves exactly as ConstantDrift(g * s) would, for that competitor alone.

0.0 pins a competitor still. That is what makes a fixed reference point expressible in the same graph as moving competitors — a bot at a known strength, a rating floor, a course difficulty:

use trueskill_tt::{ConstantDrift, Event, History, Member, Outcome, Team};

let mut h = History::builder().drift(ConstantDrift(0.1)).build();

h.add_events(vec![Event {
    time: 0,
    teams: [
        Team::with_members([Member::new("player")]),
        // A course does not improve. Pin it, and the round's evidence
        // lands on the player instead of being split between the two.
        Team::with_members([Member::new("layout_7").with_drift_scale(0.0)]),
    ]
    .into_iter()
    .collect(),
    outcome: Outcome::winner(0, 2),
}])
.unwrap();

h.converge().unwrap();

Like with_prior, the scale is competitor configuration captured at first appearance — setting it on a key the history already knows has no effect. It must be finite and non-negative; ingestion otherwise fails with InferenceError::InvalidParameter.

Note that the fluent EventBuilder (h.event(t).team([...])) sets weights but not drift_scale or prior; those need the typed Event / Team / Member shape shown above.

Scored outcomes

Use Outcome::scores([...]) when you have continuous per-team scores rather than just ranks. Adjacent score margins flow into a MarginFactor that adds soft Gaussian evidence about the latent performance diff. Configure HistoryBuilder::score_sigma(σ) to control how much you trust the margins (smaller σ = more trust).

use trueskill_tt::History;

let mut h = History::builder().score_sigma(2.0).build();
h.event(1)
    .team(["alice"])
    .team(["bob"])
    .scores([21.0, 9.0])
    .commit()
    .unwrap();
h.converge().unwrap();

Prediction

predict_outcome gives the full distribution over finishing orders. Each entry is a rank vector in the same shape Outcome::ranking takes — equal ranks mean a tie — so an outcome feeds straight back into inference.

use trueskill_tt::History;

let mut h = History::builder().p_draw(0.1).build();
h.record_winner(&"alice", &"bob", 1).unwrap();
h.converge().unwrap();

let p = h.predict_outcome(&[&[&"alice"], &[&"bob"]]).unwrap();

// Probabilities are exhaustive and disjoint, so they sum to one.
assert!((p.total() - 1.0).abs() < 1e-6);

let (best, likelihood) = p.most_likely().unwrap();
println!("most likely: {best:?} at {likelihood:.3}");
println!("draw:        {:.3}", p.probability_of(&[0, 0]));

Supports any number of teams. Because the outcome space grows factorially, the full distribution is capped at MAX_PREDICTED_TEAMS; two cheaper entry points stay available at any size:

  • predict_win_probabilities(teams)P(team i finishes strictly first), quadratic in team count.
  • predict_ranking(teams, ranks) — one specific finishing order.

Unknown keys are an error, not a silent omission: a team the history has never seen cannot produce a confident-looking probability. The error names the key, and every key must already be known — pre-filter with lookup or current_skill if your caller cannot guarantee that.

Asking about one competitor

Gaussian answers tail questions directly, which is what a stopping rule needs:

use trueskill_tt::History;

let mut h = History::builder().build();
h.record_winner(&"alice", &"bob", 1).unwrap();
let _ = h.converge().unwrap();

let skill = h.current_skill(&"alice").unwrap();

// "How sure am I that this is below the cutoff?" — a probability, not a
// `mu + z * sigma` band whose confidence drifts as sigma changes.
let _ = skill.probability_below(20.0);

// Use this rather than `1.0 - probability_below(x)`: the complement cancels
// away every digit in the upper tail, which is where a stopping rule lives.
let _ = skill.probability_above(30.0);

Which match to play next

quality() measures whether a matchup is fair. That is not the same as whether it is informative, and the two only coincide for two evenly matched competitors. When each observation costs something, ask expected_information_gain instead — the outcome-weighted divergence between what you believe now and what you would believe afterwards.

use trueskill_tt::History;

let mut h = History::builder().build();
for t in 1..=10 {
    h.record_winner(&"veteran", &"regular", t).unwrap();
    h.record_winner(&"regular", &"veteran", t + 100).unwrap();
}
h.record_winner(&"veteran", &"newcomer", 500).unwrap();
h.converge().unwrap();

let settled = h.expected_information_gain(&[&[&"veteran"], &[&"regular"]]).unwrap();
let unknown = h.expected_information_gain(&[&[&"veteran"], &[&"newcomer"]]).unwrap();

// Playing the newcomer teaches you more than replaying a settled rivalry.
assert!(unknown > settled);

The result is in nats, and is bounded by the entropy of the outcome: at most ln 2 ≈ 0.693 for a two-way result, ln 3 once draws are possible, ln k for k outcomes. A value near zero means you already know how it ends.

This costs one full inference pass per possible outcome, so it is far more expensive than quality(). Scoring every pairing among n competitors is O(n² × outcomes) passes — shortlist with quality() or predict_win_probabilities first, then score only the shortlist.

Todo

  • Implement approx for Gaussian
  • Add more tests from TrueSkillThroughTime.jl
  • Generalise a time axis — Time is now a trait (Untimed, i64), not an enum
  • Add examples (examples/atp.rs, examples/scored.rs)
  • Add Observer (Observer / NullObserver)
  • Benchmark the inference loop (benches/batch.rs, benches/history_converge.rs, benches/ingest.rs)
  • N-team predict_outcome with draw mass, and expected_information_gain
  • Cross-check quality() against sublee/trueskill — N identical teams follow the closed form (1/5)^((n-1)/2) for the conventional parameters, asserted for n = 2..10, and the n=3/n=5 values (0.200, 0.040) match the reference package

License

Licensed under either of

at your option.

Contribution

Unless you explicitly state otherwise, any contribution intentionally submitted for inclusion in the work by you, as defined in the Apache-2.0 license, shall be dual licensed as above, without any additional terms or conditions.

S
Description
No description provided
Readme
12 MiB
Languages
Rust 99.6%
Just 0.4%