logaritmiskandClaude Opus 5 8c087ad015 fix!: apply competitor configuration whenever it is supplied
`Member::with_prior` and `with_drift_scale` were consumed only on the
branch that *creates* a competitor — `priors.remove` sat inside
`if !self.agents.contains(..)`. Supplying either for a key the history
already knew did nothing at all: no error, no warning, and output
computed from the default prior. A prior applied on a competitor's very
first event and was silently discarded ever after.

Configuration now applies whenever supplied. Two details this forced:

Configuration is tracked per *field* rather than as a merged `Rating`.
A member setting only `drift_scale` must not also assert the default
prior, or it would silently undo a prior seeded on an earlier event.

Slice state has to be refreshed. `drift_scale` is re-derived on every
forward pass, but a prior is written into the competitor's earliest
slice once, at ingestion, and `iteration` refreshes only slices after
the first. Without the refresh a late prior would reach the drift terms
and nothing else — a subtler version of the drop being fixed. This was
caught by a test, not by reading the code.

Conflicting values for one competitor within a single batch are now
`ConflictingCompetitorConfig` rather than resolved by iteration order.
Events in a batch are unordered, so "last one wins" would make the
result depend on traversal — and `tests/ingestion_equivalence.rs` exists
to rule exactly that out. Repeating the same value stays inert, which is
the shape callers get when configuration is a property of the domain.

That invariant turned out to be tested only for *unconfigured*
competitors: every helper in that file built members with `Member::new`.
Extended to cover configured ones, including a check that configuration
changes the fit at all, so the order tests cannot pass vacuously.

`with_prior` had no coverage under `tests/` whatsoever, which is how
this survived. Adds `tests/competitor_config.rs`.

`drift_scale_is_ignored_after_first_appearance` asserted the old
behaviour and now asserts the new one. It was written as a deliberate
change-detector — "moving the capture would be a visible break, not a
silent one" — so it inverted rather than being deleted.

Also removes `InferenceError::ConvergenceFailed` and `NegativePrecision`,
which no code path ever constructed: public variants advertising failure
modes no caller could observe. Partial #20 — its other items were
already resolved, except `Outcome::winner` still panicking.

BREAKING CHANGE: `prior` and `drift_scale` now take effect for
competitors the history already knows, where they were previously
ignored; a batch supplying conflicting values for one competitor is now
an error. `InferenceError::ConvergenceFailed` and
`InferenceError::NegativePrecision` are removed.

Closes #10. Refs #20.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-07 15:42:31 +02:00

TrueSkill - Through Time

Rust port of TrueSkillThroughTime.py.

Other implementations

Drift

Skill drift models how a competitor's true skill can change between appearances. Each time they reappear after a gap, their skill uncertainty is widened by the drift model before the new evidence is incorporated.

Drift is represented by the Drift trait (src/drift.rs), generic over the history's time type:

pub trait Drift<T: Time>: Copy + Debug + Send + Sync {
    fn variance_delta(&self, from: &T, to: &T) -> f64;
    fn variance_for_elapsed(&self, elapsed: i64) -> f64;
}

Both methods return the amount to add to σ², not to σ. variance_delta works from two timestamps; variance_for_elapsed takes an already-computed elapsed count, and is used on the paths that cache it. Gaussian::forget applies the result entirely in variance space — from_mv(mu, variance() + variance_delta) — taking no square root.

That block is a quotation rather than a doctest. The custom-drift example below is compiled by CI, so it is what actually pins the signature.

ConstantDrift

The built-in ConstantDrift implements a linear random walk — skill uncertainty grows proportionally to time:

variance_delta = elapsed * γ²

This is the standard TrueSkill Through Time model. Pass a ConstantDrift(gamma) when constructing a Rating:

use trueskill_tt::{ConstantDrift, Gaussian, Rating};

// gamma = 0.1 means skill can shift ~0.1 per time unit.
let rating: Rating<i64, ConstantDrift> =
    Rating::new(Gaussian::from_ms(0.0, 6.0), 1.0, ConstantDrift(0.1));

assert_eq!(rating.drift().0, 0.1);

The type annotation is load-bearing: ConstantDrift implements Drift<T> for every T: Time, so without it T is ambiguous.

Custom drift

Implement Drift<T> to express any other model. For example, a drift that saturates after a long absence, with uncertainty growing as the square root of elapsed time instead of linearly:

use trueskill_tt::{Drift, Gaussian, History, Rating, Time};

#[derive(Clone, Copy, Debug)]
struct SqrtDrift {
    gamma: f64,
}

impl<T: Time> Drift<T> for SqrtDrift {
    fn variance_delta(&self, from: &T, to: &T) -> f64 {
        let elapsed = from.elapsed_to(to).max(0) as f64;
        elapsed.sqrt() * self.gamma * self.gamma
    }

    fn variance_for_elapsed(&self, elapsed: i64) -> f64 {
        (elapsed.max(0) as f64).sqrt() * self.gamma * self.gamma
    }
}

// On a single Rating:
let rating: Rating<i64, SqrtDrift> =
    Rating::new(Gaussian::from_ms(0.0, 6.0), 1.0, SqrtDrift { gamma: 0.5 });

// Or for a whole History, via the builder:
let history = History::builder().drift(SqrtDrift { gamma: 0.5 }).build();

assert_eq!(rating.beta(), 1.0);
assert_eq!(history.log_evidence(), 0.0);

HistoryBuilder::drift is the only way to set a history's drift model; there is no gamma() shorthand. The default is ConstantDrift(GAMMA).

Per-competitor drift

A History has one drift model, but individual competitors can scale it. Member::with_drift_scale(s) multiplies the drift variance that competitor accumulates, so s is in the same units as gamma: ConstantDrift(g) at scale s behaves exactly as ConstantDrift(g * s) would, for that competitor alone.

0.0 pins a competitor still. That is what makes a fixed reference point expressible in the same graph as moving competitors — a bot at a known strength, a rating floor, a course difficulty:

use trueskill_tt::{ConstantDrift, Event, History, Member, Outcome, Team};

let mut h = History::builder().drift(ConstantDrift(0.1)).build();

h.add_events(vec![Event {
    time: 0,
    teams: [
        Team::with_members([Member::new("player")]),
        // A course does not improve. Pin it, and the round's evidence
        // lands on the player instead of being split between the two.
        Team::with_members([Member::new("layout_7").with_drift_scale(0.0)]),
    ]
    .into_iter()
    .collect(),
    outcome: Outcome::winner(0, 2),
}])
.unwrap();

h.converge().unwrap();

Like with_prior, the scale is competitor configuration captured at first appearance — setting it on a key the history already knows has no effect. It must be finite and non-negative; ingestion otherwise fails with InferenceError::InvalidParameter.

Note that the fluent EventBuilder (h.event(t).team([...])) sets weights but not drift_scale or prior; those need the typed Event / Team / Member shape shown above.

Scored outcomes

Use Outcome::scores([...]) when you have continuous per-team scores rather than just ranks. Adjacent score margins flow into a MarginFactor that adds soft Gaussian evidence about the latent performance diff. Configure HistoryBuilder::score_sigma(σ) to control how much you trust the margins (smaller σ = more trust).

use trueskill_tt::History;

let mut h = History::builder().score_sigma(2.0).build();
h.event(1)
    .team(["alice"])
    .team(["bob"])
    .scores([21.0, 9.0])
    .commit()
    .unwrap();
h.converge().unwrap();

Prediction

predict_outcome gives the full distribution over finishing orders. Each entry is a rank vector in the same shape Outcome::ranking takes — equal ranks mean a tie — so an outcome feeds straight back into inference.

use trueskill_tt::History;

let mut h = History::builder().p_draw(0.1).build();
h.record_winner(&"alice", &"bob", 1).unwrap();
h.converge().unwrap();

let p = h.predict_outcome(&[&[&"alice"], &[&"bob"]]).unwrap();

// Probabilities are exhaustive and disjoint, so they sum to one.
assert!((p.total() - 1.0).abs() < 1e-6);

let (best, likelihood) = p.most_likely().unwrap();
println!("most likely: {best:?} at {likelihood:.3}");
println!("draw:        {:.3}", p.probability_of(&[0, 0]));

Supports any number of teams. Because the outcome space grows factorially, the full distribution is capped at MAX_PREDICTED_TEAMS; two cheaper entry points stay available at any size:

  • predict_win_probabilities(teams)P(team i finishes strictly first), quadratic in team count.
  • predict_ranking(teams, ranks) — one specific finishing order.

Unknown keys are an error, not a silent omission: a team the history has never seen cannot produce a confident-looking probability.

Which match to play next

quality() measures whether a matchup is fair. That is not the same as whether it is informative, and the two only coincide for two evenly matched competitors. When each observation costs something, ask expected_information_gain instead — the outcome-weighted divergence between what you believe now and what you would believe afterwards.

use trueskill_tt::History;

let mut h = History::builder().build();
for t in 1..=10 {
    h.record_winner(&"veteran", &"regular", t).unwrap();
    h.record_winner(&"regular", &"veteran", t + 100).unwrap();
}
h.record_winner(&"veteran", &"newcomer", 500).unwrap();
h.converge().unwrap();

let settled = h.expected_information_gain(&[&[&"veteran"], &[&"regular"]]).unwrap();
let unknown = h.expected_information_gain(&[&[&"veteran"], &[&"newcomer"]]).unwrap();

// Playing the newcomer teaches you more than replaying a settled rivalry.
assert!(unknown > settled);

The result is in nats, and is bounded by the entropy of the outcome: at most ln 2 ≈ 0.693 for a two-way result, ln 3 once draws are possible, ln k for k outcomes. A value near zero means you already know how it ends.

This costs one full inference pass per possible outcome, so it is far more expensive than quality(). Scoring every pairing among n competitors is O(n² × outcomes) passes — shortlist with quality() or predict_win_probabilities first, then score only the shortlist.

Todo

  • Implement approx for Gaussian
  • Add more tests from TrueSkillThroughTime.jl
  • Generalise a time axis — Time is now a trait (Untimed, i64), not an enum
  • Add examples (examples/atp.rs, examples/scored.rs)
  • Add Observer (Observer / NullObserver)
  • Benchmark the inference loop (benches/batch.rs, benches/history_converge.rs, benches/ingest.rs)
  • N-team predict_outcome with draw mass, and expected_information_gain
  • Cross-check quality() against sublee/trueskill — N-group support works and is covered by invariants, but no reference values are asserted

License

Licensed under either of

at your option.

Contribution

Unless you explicitly state otherwise, any contribution intentionally submitted for inclusion in the work by you, as defined in the Apache-2.0 license, shall be dual licensed as above, without any additional terms or conditions.

S
Description
No description provided
Readme
12 MiB
Languages
Rust 99.6%
Just 0.4%