trueskill-tt

Author	SHA1	Message	Date
logaritmisk	0705986929	feat(game): plumb ConvergenceOptions through to run_chain Game and OwnedGame gain a convergence: ConvergenceOptions field set at construction. Game::{ranked,scored} forward options.convergence into OwnedGame::{new,new_scored} (previously dropped on the floor). {ranked,scored}_with_arena take it as a parameter. run_chain reads self.convergence.{epsilon, max_iter, alpha} instead of hardcoded 1e-6 / 10 / undamped. DiffFactor::propagate gains an alpha parameter and dispatches into Trunc/MarginFactor::propagate_with_alpha. In-tree callsites in src/time_slice.rs and src/history.rs pass ConvergenceOptions::default(). Pre-existing T2 fallout in tests, benches, and the atp example (struct literals missing the new alpha field) is fixed by adding alpha: 1.0 so the workspace builds clean. Default alpha is 1.0, so all 96 lib + 27 integration test goldens remain bit-equal.	2026-05-08 15:10:35 +02:00
logaritmisk	aacaa60baa	feat(factor): add MarginFactor::propagate_with_alpha for EP damping Mirrors TruncFactor: inherent damped-propagate method, trait impl delegates with α=1.0. Existing goldens unchanged because cavity*new_msg equals the previous marginal write when α=1.0.	2026-05-08 15:03:45 +02:00
logaritmisk	fcfe0ffe37	feat(factor): add TruncFactor::propagate_with_alpha for EP damping Inherent method that applies α-damping to the outgoing message via Gaussian::damp_natural. The Factor trait impl delegates with α=1.0, preserving today's behavior bit-equal. Variable write switched from `trunc` to `cavity * damped` — algebraically identical when α=1.0 (cavity * new_msg = trunc by construction); reflects partial-update math when α<1.0.	2026-05-08 15:02:09 +02:00
logaritmisk	0fa4e7d277	feat(convergence): add ConvergenceOptions::alpha damping field Adds an EP damping coefficient defaulting to 1.0 (undamped). Will be read by run_chain in a follow-up commit. By itself this commit changes no behavior — existing constructors using ..Default::default() pick up the new field automatically.	2026-05-08 15:00:34 +02:00
logaritmisk	0dd7dab266	feat(gaussian): add damp_natural helper for EP damping Computes α·new + (1−α)·self in natural-parameter space. Will be used by TruncFactor and MarginFactor to support opt-in EP damping via ConvergenceOptions::alpha.	2026-05-08 14:59:18 +02:00
logaritmisk	43cc6d82f9	docs: implementation plan for game-local Damped EP Six tasks: Gaussian::damp_natural helper, ConvergenceOptions::alpha field, TruncFactor and MarginFactor propagate_with_alpha pair, DiffFactor + Game integration (the big task — must land atomically), and end-to-end tests for max_iter and alpha behavior.	2026-05-08 14:57:41 +02:00
logaritmisk	48a6049dc6	docs: spec for game-local Damped EP Smallest-scope realisation of spec §"Built-in schedules" Damped: a ConvergenceOptions::alpha field plumbed through run_chain to a new Gaussian::damp_natural helper applied inside TruncFactor and MarginFactor's propagate. alpha=1.0 default keeps every existing golden bit-equal; alpha<1.0 stabilises oscillating fixed-point loops on hard graphs. Defers Schedule trait integration, nat-param convergence switch, oscillation auto-detect, Residual/OneShot, and Synergy/ScoreFactor — each gets its own future plan.	2026-05-08 14:52:36 +02:00
logaritmisk	1445c08896	docs: fix stale numerics in t4-margin-factor plan The plan's prose quoted Z_cav ≈ 0.046827 and log_evidence ≈ -3.0613, which diverged from the values asserted by the shipped test in src/factor/mod.rs (-3.062235327364623). Update prose and the matching code comment to 0.04678 / -3.0622.	2026-05-08 14:37:58 +02:00
logaritmisk	f6a83e4dc6	refactor: make BuiltinFactor::log_evidence match exhaustive Replace the `_ => 0.0` wildcard with explicit `Self::TeamSum(_) \| Self::RankDiff(_) => 0.0`. No behavioral change; future variants now produce a compile error instead of being silently absorbed by the wildcard.	2026-05-08 14:37:13 +02:00
logaritmisk	68b589b965	refactor: dedupe Game::likelihoods and likelihoods_scored via run_chain Both methods were 95-line near-duplicates differing only in the closure that builds the per-diff DiffFactor. Extract the shared body as a private run_chain<F>(&self, arena, make_link) helper that returns (evidence, likelihoods); the two callers shrink to ~10 lines each. Pure code-shape change: posteriors and evidence remain bit-equal; all existing tests (lib + integration) pass unchanged.	2026-05-08 14:36:35 +02:00
logaritmisk	7481c31ad8	docs: implementation plan for post-T4-MarginFactor tech debt cleanup Three-task plan covering the run_chain dedup, exhaustive BuiltinFactor log_evidence match, and stale-numerics fix in the T4 plan doc.	2026-05-08 14:28:10 +02:00
logaritmisk	a69a3004b2	docs: spec for post-T4-MarginFactor tech debt cleanup Three independent cleanups: dedupe Game::likelihoods and likelihoods_scored via a run_chain helper taking a make_link closure, make BuiltinFactor's log_evidence match exhaustive, and fix stale numerics in the T4 plan doc.	2026-05-08 14:24:48 +02:00
logaritmisk	dbaad0e7d2	fix: release generated CHANGELOG at the wrong location	2026-04-27 09:02:38 +02:00
logaritmisk	8069941a81	chore: Release trueskill-tt version 0.1.1 v0.1.1	2026-04-27 09:01:46 +02:00
logaritmiskandClaude Opus 4.7	8b53cacd64	T4 (MarginFactor): scored outcomes via Gaussian-margin EP evidence Adds soft Gaussian-observation evidence on the per-pair diff variable, enabling continuous score margins as a richer alternative to ranks. Public API: - `Outcome::Scored([scores])` (non-breaking enum extension under `#[non_exhaustive]`). - `Game::scored(teams, outcome, options)` constructor parallel to `Game::ranked`. - `EventBuilder::scores([...])` fluent helper. - `HistoryBuilder::score_sigma(σ)` knob (default 1.0, validated > 0). - `GameOptions::score_sigma`. - `EventKind` re-exported from `lib.rs` (annotated `#[non_exhaustive]`). - New `InferenceError::InvalidParameter { name, value }` variant. Internals: - `MarginFactor` (`factor/margin.rs`): Gaussian observation factor that closes in one EP step; cavity-cached log-evidence mirrors `TruncFactor`. - `BuiltinFactor::Margin` dispatch arm. - `DiffFactor` enum in `game.rs` lets `Game::likelihoods` and the new `likelihoods_scored` share the per-pair link abstraction. - Per-event `EventKind { Ranked, Scored { score_sigma } }` routed through `TimeSlice::add_events`, `iteration_direct`, and `log_evidence`. Tests: 88 lib + 27 integration (4 new in `tests/scored.rs`); existing goldens byte-identical. Bench: `benches/scored.rs` baseline ~960µs for 60 events × 20-player pool with default convergence. Plan: docs/superpowers/plans/2026-04-27-t4-margin-factor.md Spec item marked Done. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>	2026-04-27 08:47:36 +02:00
logaritmisk	6bf3e7e294	T3: rayon-backed concurrency (opt-in) (#2 ) Implements T3 of `docs/superpowers/specs/2026-04-23-trueskill-engine-redesign-design.md` Section 6. Plan: `docs/superpowers/plans/2026-04-24-t3-concurrency.md` (11 tasks). ## Summary ### Breaking - `Send + Sync` bounds added to public traits: `Time`, `Drift<T>`, `Observer<T>`, `Factor`, `Schedule`. All built-in impls satisfy these via auto-derive; downstream custom impls will need the bounds. ### New - Opt-in `rayon` cargo feature. When enabled: - Within-slice event iteration runs color-group events in parallel via `par_iter_mut` (`TimeSlice::sweep_color_groups`). - `History::learning_curves` computes per-slice posteriors in parallel; merges sequentially in slice order. - `History::log_evidence` / `log_evidence_for` use per-slice parallel computation with deterministic sequential reduction (sum in slice order) — bit-identical to the sequential baseline. - `ColorGroups` infrastructure (`src/color_group.rs`) with greedy graph coloring. Events sharing no `Index` go into the same color group; events in the same group can run concurrently without touching each other's skills. - `tests/determinism.rs` asserts bit-identical posteriors across `RAYON_NUM_THREADS={1, 2, 4, 8}`. - `benches/history_converge.rs` measures end-to-end convergence on three workload shapes. ## Performance ### Sequential (no rayon, default build) \| Metric \| Before T3 \| After T3 \| Delta \| \|---\|---\|---\|---\| \| `Batch::iteration` \| 22.88 µs \| 23.23 µs \| +1.5% (noise) \| \| `Gaussian::` \| ≈218–264 ps \| ≈236 ps \| within noise \| No sequential regression.* Default build is as fast as T2. ### Parallel (`--features rayon`, Apple M5 Pro, auto thread count) \| Workload \| Sequential \| Parallel \| Speedup \| \|---\|---:\|---:\|---:\| \| 500 events / 100 competitors / 10 per slice \| 4.03 ms \| 4.24 ms \| 1.0× \| \| 2000 events / 200 competitors / 20 per slice \| 20.18 ms \| 19.82 ms \| 1.0× \| \| 5000 events / 50000 competitors / 1 slice \| 11.88 ms \| 9.10 ms \| 1.3× \| ### ⚠️ The spec's >=2× target was not met on realistic workloads. T3's within-slice color-group parallelism only shows material benefit when a slice holds many events AND the competitor pool is large enough to give the greedy coloring room to partition. Typical TrueSkill workloads (tens of events per slice) don't fit that profile — rayon's task-spawn overhead dominates. Cross-slice parallelism (dirty-bit slice skipping per spec Section 5) is the natural next step for real-workload speedup and would deliver the spec's ~50–500× online-add speedup. Deferred to a future tier. ## Determinism `tests/determinism.rs` runs a 200-event history at thread counts {1, 2, 4, 8} via `rayon::ThreadPoolBuilder::install` and asserts every `(time, posterior)` pair has bit-identical `mu` and `sigma` (compared via `f64::to_bits()`). Passes. ## Internals - Parallel path uses an `unsafe` block to concurrently write to `SkillStore` from color-group-disjoint events. Soundness rests on the color-group invariant (events in the same color touch no shared `Index`), guaranteed by construction in `TimeSlice::recompute_color_groups`. Sequential path unchanged from T2. - `RAYON_THRESHOLD = 64` — color groups smaller than this fall back to sequential inside `sweep_color_groups` to avoid task-spawn overhead. - Thread-local `ScratchArena` per rayon worker thread. ## Test plan - [x] `cargo test --features approx` — 96 tests pass (74 lib + 22 integration) - [x] `cargo test --features approx,rayon` — 97 tests pass (+1 determinism) - [x] `cargo clippy --all-targets --features approx -- -D warnings` — clean - [x] `cargo clippy --all-targets --features approx,rayon -- -D warnings` — clean - [x] `cargo +nightly fmt --check` — clean - [x] `cargo bench --bench batch --features approx` — 23.23 µs (no regression vs T2) - [x] `cargo bench --bench history_converge --features approx,rayon` — runs on all three workloads - [x] Bit-identical posteriors across `RAYON_NUM_THREADS={1, 2, 4, 8}` — verified ## Commit history 13 commits on `t3-concurrency`. Each task is self-contained and bisectable. See `git log main..t3-concurrency` for the full list. ## Deferred - Cross-slice parallelism (dirty-bit slice skipping) — the path that would actually speed up typical TrueSkill workloads. - Default-on `rayon` feature — spec called for default-on; we keep it opt-in until the feature proves stable in production use. - Synchronous-EP schedule with barrier merge — alternative parallel strategy per spec Section 6. - `MarginFactor` / `Outcome::Scored` — T4. - `Damped` / `Residual` schedules — T4. - N-team `predict_outcome` — T4. - `Game::custom` full ergonomics — T4. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Reviewed-on: #2 Co-authored-by: Anders Olsson <anders.e.olsson@gmail.com> Co-committed-by: Anders Olsson <anders.e.olsson@gmail.com>	2026-04-24 13:01:01 +00:00
logaritmisk	d2aab82c1e	T0 + T1 + T2: engine redesign through new API surface (#1 ) Implements tiers T0, T1, T2 of `docs/superpowers/specs/2026-04-23-trueskill-engine-redesign-design.md`. All three tiers have landed together on this branch because they build on one another; this PR rolls them up for a single review pass. Per-tier plans: - T0: `docs/superpowers/plans/2026-04-23-t0-numerical-parity.md` - T1: `docs/superpowers/plans/2026-04-24-t1-factor-graph.md` - T2: `docs/superpowers/plans/2026-04-24-t2-new-api-surface.md` ## Summary ### T0 — Numerical parity (internal) - `Gaussian` switched to natural-parameter storage `(pi, tau)`; mul/div now ~7× faster (218 ps vs 1.57 ns). - `HashMap<Index, _>` → dense `Vec<_>` keyed by `Index.0` (via `AgentStore<D>`, `SkillStore`). - `ScratchArena` eliminates per-event allocations in `Game::likelihoods`. - `InferenceError` seed type added (1 variant). - 38 → 53 tests passing through T1. - Benchmark: `Batch::iteration` 29.84 → 21.25 µs. ### T1 — Factor graph machinery (internal) - `Factor` trait + `BuiltinFactor` enum (TeamSum / RankDiff / Trunc) driving within-game inference. - `VarStore` flat storage for variable marginals. - `Schedule` trait + `EpsilonOrMax` impl replacing the hand-rolled EP loop. - `Game::likelihoods` rebuilt on the factor-graph machinery; iteration counts and goldens preserved to within 1e-6. - 53 tests passing. - Benchmark: `Batch::iteration` 23.01 µs (slight regression absorbed in T2). ### T2 — New API surface (breaking) Renames: - `IndexMap → KeyTable`, `Player → Rating`, `Agent → Competitor`, `Batch → TimeSlice` New types: - `Time` trait with `Untimed` ZST and `i64` impls; `Drift<T>`, `Rating<T, D>`, `Competitor<T, D>`, `TimeSlice<T>`, `History<T, D, O, K>` all generic. - `Event<T, K>`, `Team<K>`, `Member<K>`, `Outcome` (`Ranked` variant; `#[non_exhaustive]`). - `Observer<T>` trait + `NullObserver`. - `ConvergenceOptions`, `ConvergenceReport`. - `GameOptions`, `OwnedGame<T, D>`. Three-tier ingestion: - `history.record_winner(&K, &K, T)` / `record_draw(&K, &K, T)` — 1v1 convenience. - `history.add_events(iter)` — typed bulk. - `history.event(T).team([...]).weights([...]).ranking([...]).commit()` — fluent. Query API: `current_skill`, `learning_curve`, `learning_curves` (keyed on `K`), `log_evidence`, `log_evidence_for`, `predict_quality`, `predict_outcome`. Game constructors: `ranked`, `one_v_one`, `free_for_all`, `custom` — all returning `Result<_, InferenceError>`. `factors` module: `Factor`, `Schedule`, `VarStore`, `VarId`, `BuiltinFactor`, `EpsilonOrMax`, `ScheduleReport`, `TeamSumFactor`, `RankDiffFactor`, `TruncFactor` now public. Errors: `InferenceError` gains `MismatchedShape`, `InvalidProbability`, `ConvergenceFailed`; boundary panics converted to `Result`. Removed (breaking): `History::convergence(iters, eps, verbose)`, `HistoryBuilder::gamma(f64)`, `HistoryBuilder::time(bool)`, `History.time: bool`, `learning_curves_by_index`, nested-Vec public `add_events`. ## Behavior change (documented in CHANGELOG) `Time = Untimed` has `elapsed_to → 0`, so no drift accumulates between slices. The old `time=false` mode implicitly forced `elapsed=1` on reappearance via an `i64::MAX` sentinel — that quirk is not reproducible under a typed time axis. Tests that depended on it now use `History::<i64, _>` with explicit `1..=n` timestamps. One test (`test_env_ttt`) had 3 Gaussian goldens updated to reflect the corrected semantics; documented in commit `33a7d90`. ## Final numbers \| Metric \| Before T0 \| After T2 \| Delta \| \|---\|---\|---\|---\| \| `Batch::iteration` \| 29.84 µs \| 21.36 µs \| -28% \| \| `Gaussian::mul` \| 1.57 ns \| 219 ps \| -86% \| \| `Gaussian::div` \| 1.57 ns \| 219 ps \| -86% \| \| Tests passing \| 38 \| 90 \| +52 \| All other Gaussian ops unchanged (~219 ps add/sub, ~264 ps pi/tau reads). ## Test plan - [x] `cargo test --features approx` — 90/90 pass (68 lib + 10 api_shape + 6 game + 4 record_winner + 2 equivalence) - [x] `cargo clippy --all-targets --features approx -- -D warnings` — clean - [x] `cargo +nightly fmt --check` — clean - [x] `cargo bench --bench batch` — 21.36 µs - [x] `cargo bench --bench gaussian` — unchanged from T1 - [x] `cargo run --example atp --features approx` — rewritten in new API, runs clean - [x] Historical Game-level goldens preserved in `tests/equivalence.rs` - [x] Public API matches spec Section 4 (verified by integration tests in `tests/api_shape.rs`) ## Commit history ~45 commits total across T0 + T1 + T2. Each task is self-contained and individually tested; the branch is bisectable. See `git log main..t2-new-api-surface` for the full list. ## Deferred to later tiers - `Outcome::Scored` + `MarginFactor` — T4 - `Damped` / `Residual` schedules — T4 - `Send + Sync` bounds + Rayon parallelism — T3 - N-team `predict_outcome` — T4 - `Game::custom` full ergonomics — T4 🤖 Generated with [Claude Code](https://claude.com/claude-code) Reviewed-on: #1 Co-authored-by: Anders Olsson <anders.e.olsson@gmail.com> Co-committed-by: Anders Olsson <anders.e.olsson@gmail.com>	2026-04-24 11:20:04 +00:00
logaritmisk	a14df02089	chore: do not publish v0.1.0	2026-04-23 20:26:52 +02:00
logaritmisk	0d266b4428	chore: make cargo release add CHANGELOG.md before commit	2026-04-23 20:26:16 +02:00
logaritmisk	a4b4e5e8fa	chore: clean up	2026-04-23 20:24:10 +02:00
logaritmisk	04d5478ee4	style: cargo fmt	2026-04-23 20:23:13 +02:00
logaritmisk	480467ac32	chore: added cliff.toml, release.toml and rustfmt.toml	2026-04-23 20:22:27 +02:00
logaritmisk	dc47964310	added benchmark	2026-03-23 14:55:18 +01:00
logaritmisk	61a5507f5c	remove notepad	2026-03-23 14:21:23 +01:00
logaritmisk	a1f282a1c8	feat: added a Drift trait and a "default" ConstantDrift implementation	2026-03-16 12:06:04 +01:00
logaritmisk	853f177fa8	Small changes for new 2024 edition	2025-02-21 14:09:58 +01:00
logaritmisk	fc0efcdc52	Update edition	2025-02-21 14:06:28 +01:00
logaritmisk	3bbddb168f	Ignore temp folder	2024-04-03 14:43:54 +02:00
logaritmisk	2366c45f6a	Basic test for quality	2024-04-03 10:25:10 +02:00
logaritmisk	3a22b20a17	Added todo to readme, and documentation for quality function	2024-04-03 09:53:07 +02:00
logaritmisk	02ae2f0977	Change assert to debug_assert	2024-04-03 09:44:41 +02:00
Anders Olsson	db743bc417	Improve performance	2023-10-31 10:02:07 +01:00
Anders Olsson	7e2576085f	Make quality a free standing function instead	2023-10-26 11:11:54 +02:00
Anders Olsson	062c9d3765	Added quality function	2023-10-26 11:09:30 +02:00
Anders Olsson	755a5ea668	Move stuff around	2023-10-26 11:01:14 +02:00
Anders Olsson	72e06eb536	Rename variables	2023-10-26 08:26:28 +02:00
Anders Olsson	e3eebb507c	Refactor history	2023-10-26 08:18:15 +02:00
Anders Olsson	d8dfbba251	Fix clippy warning	2023-10-25 08:16:45 +02:00
Anders Olsson	d152e356f1	Remove unnecessary allocations	2023-10-24 16:10:40 +02:00
Anders Olsson	59c256edad	Dry my eyes	2023-10-24 09:50:16 +02:00
Anders Olsson	efa235be59	Clean up	2023-10-24 09:44:42 +02:00
Anders Olsson	aea3df285a	Update crates	2023-06-28 09:42:12 +02:00
logaritmisk	a491f8de8d	Fix broken link in README	2023-01-09 14:37:07 +01:00
logaritmisk	0b4b07d60e	Added more links to readme	2023-01-09 14:35:34 +01:00
logaritmisk	05f178641c	Rename d to diff, and t to team	2022-12-29 20:38:36 +01:00
logaritmisk	e3906aebaa	Small refactor	2022-12-27 22:37:12 +01:00
logaritmisk	8e25826f91	More rustifying	2022-12-27 22:30:20 +01:00
logaritmisk	9b6cb9e7eb	Make it more rusty	2022-12-27 22:24:01 +01:00
logaritmisk	fdddf56156	Remove unused mut reference	2022-12-27 22:11:04 +01:00
logaritmisk	b93194f762	Added default implementation for TeamMessage	2022-12-20 11:08:27 +01:00

1 2