Files
trueskill-tt/CLAUDE.md
T
logaritmiskandClaude Opus 5 633a503900 refactor!: remove the factor-graph surface nothing used, add try_winner
Two decisions taken before cutting 0.5.0.

#42 — the `Schedule` trait was public API the engine never called. Its
only call site was `Game::custom`, itself `#[doc(hidden)]`, and
`EpsilonOrMax` was never constructed anywhere. Removing that surface
showed the problem was larger than the issue described: with `custom`
gone, the compiler found `Factor`, `BuiltinFactor`, `RankDiffFactor` and
`TeamSumFactor` all dead too.

`Game::run_chain` drives a local `DiffFactor` enum and bypasses the
whole T1 abstraction — it has done since it was written. So this is not
just an unused extension point but the machinery it was built on, and
`CLAUDE.md` was documenting it as live architecture.

Removed: `graph` module, `Schedule`, `EpsilonOrMax`, `ScheduleReport`,
`Game::custom`, `Factor`, `BuiltinFactor`, `RankDiffFactor`,
`TeamSumFactor`. `TruncFactor`, `MarginFactor`, `VarStore` and `VarId`
stay — inference uses those. The measurement behind choosing removal
over wiring is in #42: the within-game loop converges in 1 to 8
iterations against a cap of 30, so a `Residual` schedule has no headroom
to reclaim, and `Damped` already shipped as `ConvergenceOptions::alpha`.

#20 — `Outcome::winner` panicking on an out-of-range index. Kept, and
the reasoning is now on the method. It is the only constructor here that
validates, which looks inconsistent until you try deferring like its
siblings: `winner(5, 2)` produces ranks `[1, 1]`, an all-tied draw that
ingestion accepts without complaint when `p_draw > 0`. Asking "team 5
won" and silently getting "everyone drew" is the exact failure this
crate keeps removing, so the check belongs where the mistake is.

Adds `Outcome::try_winner` for indices that are computed or parsed
rather than written literally, following the `new`/`try_new` convention.
That is additive; the panicking form stays because every call site in
this repo, its tests and its README passes literals, where a `?` would
be noise.

BREAKING CHANGE: the `graph` module and everything it exported are
removed, as are `Game::custom` and `ScheduleReport`.

Closes #42. Closes #20.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 02:14:43 +02:00

6.9 KiB

CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

Commands

just test                            # Full suite across every feature combination CI checks
just check                           # Fast inner loop: cargo test --features approx
just lint                            # clippy, warnings denied
just fmt                             # ALWAYS nightly — rustfmt.toml uses nightly-only options
just determinism                     # Bit-identical posteriors at RAYON_NUM_THREADS 1/2/4/8
just ci                              # Everything CI runs
cargo test --lib <test_name>         # A single test by name
cargo bench                          # Criterion benchmarks

Run tests in release too. debug_assert! is compiled out there, and that is where several defects have hidden — a debug-only run is not evidence. just test includes a release job.

Feature flags

  • approxapprox::AbsDiffEq etc. for Gaussian. Most numerical goldens need it.
  • rayon — opt-in parallel within-slice sweep and per-slice query passes.

Working rules

  • Investigate before implementing. Measure the actual behaviour first — against an analytic reference where one exists. Several "obvious" fixes in this repo turned out to be wrong in sign or unnecessary, and the measurement is what caught them.
  • Fix the root issue, not the symptom. A clamp that hides an underflow, or a tolerance loosened to make a test pass, is a defect deferred.
  • Scout crates.io before hand-rolling numerics. Check accuracy against an independent reference rather than trusting downloads: puruspe has 1.4M downloads and is 346 ULP off in the tail, where libm is 1. Fewer dependencies is preferable, not mandatory — take the dependency when it is measurably better.

Architecture

A Rust port of TrueSkillThroughTime.py: Bayesian skill rating that infers skill at every point in time, propagating evidence both forward and backward across a history.

Data flow

Ingestion (public types, event.rs):

Event<T, K>  →  Team<K>[]  →  Member<K>[]

History::add_events flattens that into indices; teams survive only as grouping, not as a value. Inference then runs on the internal shapes:

History  →  TimeSlice[]  →  Event[]  →  Item[]
                                ↓
                   Game (factor graph) → Schedule → BuiltinFactor[]
  • History (history.rs) — top level. Interns keys, groups events into TimeSlices by time, runs the forward/backward sweep in converge(), and answers learning_curves(), current_skill(), log_evidence(), predict_quality(), predict_outcome(). Built via HistoryBuilder.
  • TimeSlice (time_slice.rs) — all events at one time. Owns a SkillStore and a ScratchArena; iteration() sweeps its events, using ColorGroups to partition independent ones.
  • Event — two distinct types, do not confuse them. The public ingestion Event<T, K> is in event.rs (with Team/Member); the internal pub(crate) Event in time_slice.rs is one match during inference, where compute() runs inference reading skills immutably and apply() folds the result back. That split is what lets a color group run in parallel with no unsafe.
  • Game (game.rs) — a single match's factor graph. run_chain builds the diff chain between rank-adjacent teams and drives it to convergence.
  • Gaussian (gaussian.rs) — natural parameters (pi = 1/sigma², tau = mu/sigma²). Mul/Div are the EP product/cavity: pure adds and subtracts. Variance-space ops (Add, Sub, exclude, forget) go through from_mv/variance() and take no square root.
  • factor/TruncFactor (ranked) and MarginFactor (scored) over a flat VarStore. Game::run_chain drives them directly through a local DiffFactor enum; there is no Schedule indirection and no generic Factor trait. Both were removed once measurement showed nothing had ever used them — see #42.
  • Competitor (competitor.rs) — per-history temporal state (message, last_time). Rating (rating.rs) — static config (prior, beta, drift).
  • storage/SkillStore (per slice, pub(crate)) and CompetitorStore (per history, public), both indexed by Index. The module is pub, but only CompetitorStore is reachable from outside the crate.
  • KeyTable (key_table.rs) — user key ↔ Index, both directions O(1).
  • Drift (drift.rs) / Time (time.rs) — traits. Time is a trait (i64, Untimed), not an enum.
  • lib.rs — public exports, global defaults (MU, SIGMA, BETA, GAMMA, P_DRAW, EPSILON, ITERATIONS), and the standalone quality(). The cdf() / erfc() helpers live here too but are pub(crate) and private respectively — not public API.

Invariants worth knowing

  • A tie needs p_draw > 0. With p_draw == 0.0 the truncation margin is zero and the two-sided tie update evaluates 0/0. Ingestion rejects such events with InferenceError::TieWithoutDrawProbability. This includes Outcome::winner(w, n) for n >= 3, which ties every loser.
  • NaN is never convergence. Comparisons against NaN are all false, so tuple_gt reads NaN as "below epsilon". Use step_converged / step_is_finite, never !tuple_gt(..) alone.
  • Evidence accumulates in log space. A linear product over a long diff chain underflows to zero, and ln(0) is -inf.
  • Colors are contiguous. recompute_color_groups reorders events so each color occupies one range; ColorGroups::groups_are_contiguous asserts it.
  • Transcendentals go through libm, not std. IEEE 754 pins the basic operations and sqrt but says nothing about exp/log/erf, and std delegates to the system math library — measured, f64::exp and libm::exp disagree on 9.7% of inputs by one ULP. Since inference is an iterative fixed point, one ULP can change an iteration count. Use libm::exp / libm::log in inference code; f64::sqrt is fine (IEEE specifies it). Tests may use either.
  • The crate is #![forbid(unsafe_code)]. Keep it that way.
  • Ingestion order must not change the answer. Events added one at a time must converge to the same fixed point as the same events batched — see tests/ingestion_equivalence.rs.

Testing notes

  • Numerical goldens are cross-validated against the Python/Julia reference. Some are convergence residuals, not exact values; treat a small movement as suspicious but check whether the new value is closer to the analytic truth (symmetric fixtures converge to their prior mean exactly) before assuming a regression.
  • tests/degenerate_inputs.rs covers empty/boundary/error paths, tests/ingestion_equivalence.rs covers batching order, tests/quality.rs covers N-group quality, tests/determinism.rs covers thread counts.