Follow-on from #41, which added `libm` for `erfc`. Surveying what else the dependency offers: its unique surface over `std` is `erf`/`erfc`, `lgamma`/`tgamma` and Bessel functions, and only the first was ever needed. But the survey found something better than another special function. `std`'s `exp` and `ln` delegate to the *system* math library. IEEE 754 specifies the basic operations and `sqrt` exactly and says nothing about transcendentals, so those differ per platform. Measured here over 200k inputs: exp: 19425/200000 differ from libm (worst 1 ulp) log: 9932/200000 differ Inference is an iterative fixed point, so a one-ULP difference can change an iteration count and move the answer by more than one ULP. Routing every transcendental through `libm` makes a fit reproducible across platforms — a stronger guarantee than `tests/determinism.rs`, which only covers thread counts. It costs nothing. `Batch::iteration` measured -2.7% [-5.7%, -0.3%] with the whole set swapped, and not one golden moved. Also switches the two places that combined sigmas as `sqrt(a^2 + b^2)` to `hypot`. Squaring overflows to infinity above ~1.3e154 and flushes to zero below ~1.5e-154 — measured, the naive form returns `inf` where `hypot` returns 1.41e160 — and `Gaussian`'s constructors are public, so a caller can reach both ends. Deliberately not done: rewriting the KL divergence's `ln` of a ratio via `ln_1p`. The cancellation is real as the ratio approaches one, but measured absolute error is at most ~1e-11 in a quantity of order 0.4 nats, so it changes nothing. The invariant is recorded in `CLAUDE.md` and on `erfc`'s own docs, since nothing enforces it mechanically. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
121 lines
6.2 KiB
Markdown
121 lines
6.2 KiB
Markdown
# CLAUDE.md
|
|
|
|
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
|
|
|
|
## Commands
|
|
|
|
```bash
|
|
just test # Full suite across every feature combination CI checks
|
|
just check # Fast inner loop: cargo test --features approx
|
|
just lint # clippy, warnings denied
|
|
just fmt # ALWAYS nightly — rustfmt.toml uses nightly-only options
|
|
just determinism # Bit-identical posteriors at RAYON_NUM_THREADS 1/2/4/8
|
|
just ci # Everything CI runs
|
|
cargo test --lib <test_name> # A single test by name
|
|
cargo bench # Criterion benchmarks
|
|
```
|
|
|
|
**Run tests in release too.** `debug_assert!` is compiled out there, and that
|
|
is where several defects have hidden — a debug-only run is not evidence.
|
|
`just test` includes a release job.
|
|
|
|
### Feature flags
|
|
|
|
- `approx` — `approx::AbsDiffEq` etc. for `Gaussian`. Most numerical goldens need it.
|
|
- `rayon` — opt-in parallel within-slice sweep and per-slice query passes.
|
|
|
|
## Architecture
|
|
|
|
A Rust port of [TrueSkillThroughTime.py](https://github.com/glandfried/TrueSkillThroughTime.py):
|
|
Bayesian skill rating that infers skill at every point in time, propagating
|
|
evidence both forward and backward across a history.
|
|
|
|
### Data flow
|
|
|
|
Ingestion (public types, `event.rs`):
|
|
|
|
```
|
|
Event<T, K> → Team<K>[] → Member<K>[]
|
|
```
|
|
|
|
`History::add_events` flattens that into indices; teams survive only as
|
|
grouping, not as a value. Inference then runs on the internal shapes:
|
|
|
|
```
|
|
History → TimeSlice[] → Event[] → Item[]
|
|
↓
|
|
Game (factor graph) → Schedule → BuiltinFactor[]
|
|
```
|
|
|
|
- **`History`** (`history.rs`) — top level. Interns keys, groups events into
|
|
`TimeSlice`s by time, runs the forward/backward sweep in `converge()`, and
|
|
answers `learning_curves()`, `current_skill()`, `log_evidence()`,
|
|
`predict_quality()`, `predict_outcome()`. Built via `HistoryBuilder`.
|
|
- **`TimeSlice`** (`time_slice.rs`) — all events at one time. Owns a
|
|
`SkillStore` and a `ScratchArena`; `iteration()` sweeps its events, using
|
|
`ColorGroups` to partition independent ones.
|
|
- **`Event`** — two distinct types, do not confuse them. The *public* ingestion
|
|
`Event<T, K>` is in `event.rs` (with `Team`/`Member`); the *internal*
|
|
`pub(crate) Event` in `time_slice.rs` is one match during inference, where
|
|
`compute()` runs inference reading skills immutably and `apply()` folds the
|
|
result back. That split is what lets a color group run in parallel with no
|
|
`unsafe`.
|
|
- **`Game`** (`game.rs`) — a single match's factor graph. `run_chain` builds the
|
|
diff chain between rank-adjacent teams and drives it to convergence.
|
|
- **`Gaussian`** (`gaussian.rs`) — natural parameters (`pi = 1/sigma²`,
|
|
`tau = mu/sigma²`). `Mul`/`Div` are the EP product/cavity: pure adds and
|
|
subtracts. Variance-space ops (`Add`, `Sub`, `exclude`, `forget`) go through
|
|
`from_mv`/`variance()` and take no square root.
|
|
- **`factor/`** — `TeamSumFactor`, `RankDiffFactor`, `TruncFactor` (ranked),
|
|
`MarginFactor` (scored), over a flat `VarStore`. `BuiltinFactor` dispatches
|
|
by enum rather than `dyn`.
|
|
- **`Schedule`** (`schedule.rs`) — drives factor propagation. `EpsilonOrMax` is
|
|
the only implementation.
|
|
- **`Competitor`** (`competitor.rs`) — per-history temporal state (`message`,
|
|
`last_time`). **`Rating`** (`rating.rs`) — static config (prior, `beta`, drift).
|
|
- **`storage/`** — `SkillStore` (per slice, `pub(crate)`) and `CompetitorStore`
|
|
(per history, public), both indexed by `Index`. The module is `pub`, but only
|
|
`CompetitorStore` is reachable from outside the crate.
|
|
- **`KeyTable`** (`key_table.rs`) — user key ↔ `Index`, both directions O(1).
|
|
- **`Drift`** (`drift.rs`) / **`Time`** (`time.rs`) — traits. `Time` is a *trait*
|
|
(`i64`, `Untimed`), not an enum.
|
|
- **`lib.rs`** — public exports, global defaults (`MU`, `SIGMA`, `BETA`,
|
|
`GAMMA`, `P_DRAW`, `EPSILON`, `ITERATIONS`), and the standalone `quality()`.
|
|
The `cdf()` / `erfc()` helpers live here too but are `pub(crate)` and private
|
|
respectively — not public API.
|
|
|
|
### Invariants worth knowing
|
|
|
|
- **A tie needs `p_draw > 0`.** With `p_draw == 0.0` the truncation margin is
|
|
zero and the two-sided tie update evaluates `0/0`. Ingestion rejects such
|
|
events with `InferenceError::TieWithoutDrawProbability`. This includes
|
|
`Outcome::winner(w, n)` for `n >= 3`, which ties every loser.
|
|
- **NaN is never convergence.** Comparisons against NaN are all false, so
|
|
`tuple_gt` reads NaN as "below epsilon". Use `step_converged` /
|
|
`step_is_finite`, never `!tuple_gt(..)` alone.
|
|
- **Evidence accumulates in log space.** A linear product over a long diff
|
|
chain underflows to zero, and `ln(0)` is `-inf`.
|
|
- **Colors are contiguous.** `recompute_color_groups` reorders events so each
|
|
color occupies one range; `ColorGroups::groups_are_contiguous` asserts it.
|
|
- **Transcendentals go through `libm`, not `std`.** IEEE 754 pins the basic
|
|
operations and `sqrt` but says nothing about `exp`/`log`/`erf`, and `std`
|
|
delegates to the *system* math library — measured, `f64::exp` and `libm::exp`
|
|
disagree on 9.7% of inputs by one ULP. Since inference is an iterative fixed
|
|
point, one ULP can change an iteration count. Use `libm::exp` / `libm::log` in
|
|
inference code; `f64::sqrt` is fine (IEEE specifies it). Tests may use either.
|
|
- **The crate is `#![forbid(unsafe_code)]`.** Keep it that way.
|
|
- **Ingestion order must not change the answer.** Events added one at a time
|
|
must converge to the same fixed point as the same events batched — see
|
|
`tests/ingestion_equivalence.rs`.
|
|
|
|
### Testing notes
|
|
|
|
- Numerical goldens are cross-validated against the Python/Julia reference.
|
|
Some are *convergence residuals*, not exact values; treat a small movement
|
|
as suspicious but check whether the new value is closer to the analytic
|
|
truth (symmetric fixtures converge to their prior mean exactly) before
|
|
assuming a regression.
|
|
- `tests/degenerate_inputs.rs` covers empty/boundary/error paths,
|
|
`tests/ingestion_equivalence.rs` covers batching order, `tests/quality.rs`
|
|
covers N-group quality, `tests/determinism.rs` covers thread counts.
|