13 Commits
Author SHA1 Message Date
logaritmisk 07285283b6 chore: Release trueskill-tt version 0.2.0 2026-08-27 17:37:31 +02:00
logaritmiskandClaude Opus 5 b73cf0145a chore: dual-license MIT OR Apache-2.0
The Rust ecosystem convention, and what kickscore, xy and saphyr already use —
kickscore being the closest sibling to this crate.

LICENSE-APACHE is the canonical 201-line Apache-2.0 text with the appendix
left as the unfilled template, which is the form Rust crates ship. LICENSE-MIT
carries the copyright line. Both are picked up by cargo automatically and
appear in the packaged crate.

This also unblocks the release workflow. `cargo publish` does not check the
license field when the target is an alternative registry, but cargo-release
does, and refuses outright:

    error: trueskill-tt is missing the following fields:
             license || license-file

--no-verify does not bypass it. So `just release` was inert without this,
whatever cargo publish alone would have accepted.

README gains the conventional dual-license section and the contribution note
that dedicates inbound contributions under the same terms.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T5SYDExxL4vZgvunrcNSMc
2026-08-27 17:35:42 +02:00
logaritmiskandClaude Opus 5 56ff01074f docs(cargo): correct the licence note — kellnr does not require one
The note claimed `cargo publish` rejects a crate without `license`. That is
true only for crates.io; publishing to an alternative registry does not check
it, verified by a dry run against kellnr that packages and verifies cleanly.

Staying unlicensed is a deliberate choice, so the note now says that and states
the actual consequence — all-rights-reserved by default — rather than a
mechanical blocker that does not exist.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T5SYDExxL4vZgvunrcNSMc
2026-08-27 17:30:59 +02:00
logaritmiskandClaude Opus 5 8d47e54a8a chore: keep the 48 MB ATP dataset out of the published crate
`cargo publish --dry-run` packaged 48.1 MiB (8.6 MiB compressed) for a library
whose source is 312 KB. All of it was examples/atp.csv, a tennis dataset the
atp example reads.

examples/atp.rs opens it by relative path at runtime rather than include_str!,
so excluding the data still compiles and `cargo package --verify` still builds
every target — the example just needs the file fetched from the repo to run.

Packaged size is now 360.8 KiB / 83.1 KiB compressed, a 133x reduction. Every
consumer would otherwise have paid that download.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T5SYDExxL4vZgvunrcNSMc
2026-08-27 17:29:44 +02:00
logaritmiskandClaude Opus 5 7de092ba12 chore: target releases at the private kellnr registry
Mirrors the textus setup, adapted for a single crate rather than a workspace.

Cargo.toml gains publish = ["kellnr"], which does double duty: it points
cargo-release at the private registry and makes an accidental `cargo publish`
to crates.io a hard error rather than an irreversible mistake.

.cargo/config.toml is committed rather than left to a per-user
~/.cargo/config.toml. Without it a fresh clone, a new machine, or CI fails
with "registry index was not found in any configuration: kellnr" before
compiling anything. The index URL is not a secret; the token stays in
~/.cargo/credentials.toml, or CARGO_REGISTRIES_KELLNR_TOKEN in CI.

release.toml flips publish from false to true and pins push = false, so the
Justfile recipe pushes last — after tags and publish have both succeeded.
The git-cliff pre-release hook is unchanged.

cliff.toml gained a Breaking Changes group. Its commit_parsers matched on type
alone with conventional_commits = false, so `refactor!: remove the inert online
flag` rendered as an ordinary Refactor bullet and the break was invisible in
the generated changelog. The new parsers match a `!` subject and a
BREAKING CHANGE body, and must precede the type parsers because the first match
wins. The unreleased section now opens with the break, which matters because
the next release is the one that removes HistoryBuilder::online.

The release recipe runs `just ci` before cutting: cargo-release only
verify-compiles the packaged crate and publishing cannot be undone, and the
release profile is where this crate's defects have historically hidden.

Still unpublishable: Cargo.toml has no `license`. That is a deliberate TODO,
not an oversight.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T5SYDExxL4vZgvunrcNSMc
2026-08-27 17:28:51 +02:00
logaritmiskandClaude Opus 5 eeb43e3be1 fix: close out four small issues and pin #27's repro
#29 — log_evidence and log_evidence_for took &mut self while mutating
nothing. Loosening them to &self is not source-breaking for ordinary callers
(a &mut reborrows as & transparently) and brings them in line with the
filtered_* accessors added last week.

Not the mechanical change it looked like: under the rayon feature the closure
in log_evidence_internal captured all of &self rather than just the competitor
store, which drags KeyTable<K> in and demands K: Sync from every caller. That
compiled while the method took &mut self and stopped compiling the moment it
did not. Binding `let agents = &self.agents;` before the closure narrows the
capture; the comment there says why, because the next person to inline it will
reintroduce the bound.

#31 — TimeSlice::add_events constructed Skill with ..Default::default() while
filtered_step spells every field out. The design relies on a new Skill field
being a compile error at construction sites rather than a silent default, and
that tripwire only fired at one of the two. Now both.

#28 — log_evidence_internal's `forward` flag is a genuine forward-only
quantity only on a history that has never been converged, because iteration
alternates sweeps and the likelihood feeding the forward message absorbs
backward information from the second iteration onward. Documented, with a
pointer to filtered_log_evidence for the quantity that survives convergence.
That trap is one function away from the one #19 was about.

#23 — color_greedy carried #[allow(dead_code)] despite being called by
recompute_color_groups: a mute button on a live function, which is the
specific complaint in that issue.

#27 was already fixed — the guard landed in f4e2922 and the issue was filed
against 7742b2b, which merge-base confirms predates it — but nothing pinned
it. Added the issue's own reproduction, which matters because the two profiles
fail differently and a debug-only test would miss the release path. Removing
both guards reproduces the issue verbatim: "attempt to subtract with overflow"
in debug, "index out of bounds: the len is 0 but the index is
18446744073709551615" in release.

Also amended the filtered-estimates spec (#30): the tolerance-not-bit-identity
caveat is conservative. Forcing the scratch onto the sequential sweep instead
of the grouped one — a far larger perturbation than a permuted event order —
still agrees within 1e-8 under tight convergence.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T5SYDExxL4vZgvunrcNSMc
2026-08-27 17:21:05 +02:00
logaritmiskandClaude Opus 5 69ddebe21d docs: state filtered accessor cost and evidence semantics precisely
filtered_learning_curve's signature mirrors learning_curve, which is cheap
per key — so the mirroring trained callers to assume this one is too. It is
a full forward pass per call, making the natural loop over competitors
O(competitors * events). The doc now says so in complexity terms and points
multi-key callers at the plural form.

filtered_log_evidence claimed each event is scored "using only what was
known before it". That is exact for a slice holding one event, but events
sharing a timestamp inform each other through the within-slice sweep, so
the honest claim is "before that time". The behaviour is deliberate and
matches log_evidence's own convention; only the promise was too strong.

This branch exists because a feature's documentation was quietly false.
Shipping it with two more overstated doc comments would be a poor joke.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T5SYDExxL4vZgvunrcNSMc
2026-08-27 16:59:42 +02:00
logaritmisk 9c39d1e681 test: pin the invariants that make filtered estimates trustworthy
The bracket test proves the feature works on one fixture. These pin the bug
class:

- Invariance to converge(). This is the one that matters. Reading
  skill.forward instead of the carried message makes it fail immediately,
  because converge() alternates sweeps and contaminates skill.forward with
  backward information from the second iteration onward. That is the
  property a stored field cannot have, and the reason issue #19's proposed
  fix would not have worked.
- Invariance to ingestion order, the crate's standing invariant.
- One slice has no future to propagate back, so filtered equals smoothed.
- Empty history yields zero and empty maps.

Agreement is to 1e-8 under tight convergence rather than bit-identity:
iteration recomputes the colour partition only when from == 0, so an
incrementally built slice keeps insertion order until the first converge()
reorders it, and the scratch clone inherits whichever order it finds. Same
fixed point, different path to it.
2026-08-27 16:37:21 +02:00
logaritmisk 50e11cfbfa feat: add filtered learning curves
learning_curve returns post-convergence posteriors, so every point is
smoothed: the estimate at a given date incorporates rounds played years
later. On ustat's data that starts six players' curves already spread apart
at sigma 0.9-1.6 against a prior of 6.0, barely moving thereafter.

filtered_learning_curve plots the same competitor on forward-only
information, so everyone starts at the prior and fans out. It could not be
reconstructed from the public API before: a caller could only refit over
events[0..k] for every k, which is O(n^2) fits for something one forward
pass already computes.
2026-08-27 16:30:19 +02:00
logaritmisk d4af048914 feat: add filtered_log_evidence
Scores every event on what was known before it, rather than on priors that
carry information from events which had not happened yet. This is the
quantity HistoryBuilder::online promised and never delivered.

The pass walks slices in time order carrying its own forward messages, and
per slice runs the unmodified production sweep on a scratch copy whose
backward message is left improper. Reusing iterate_to_convergence rather
than reimplementing inference means a competitor playing twice at one time
is handled by the same within-slice EP that converge() uses, instead of
being approximated the way the old evidence paths approximated it.

Nothing is stored on Skill and nothing on self is mutated, so the result is
independent of whether converge() has run — the property a stored field
cannot have.
2026-08-27 16:21:01 +02:00
logaritmisk bf9d964cae refactor!: remove the inert online flag
Skill.online was initialised to N_INF and assigned nowhere, so
HistoryBuilder::online(true) made every rating improper and log_evidence()
reported n * ln(0.5) — every game scored as a coin flip. The value is finite
and plausible, which is why it went unnoticed.

The default was false, so no existing result changes. A working replacement
lands next; a stored field cannot hold the quantity, because converge()
alternates sweeps and contaminates skill.forward with backward information
from the second iteration onward.

Also renames a test binding from ..._online to ..._forward: it passes the
forward flag, and the two senses being conflated is how this survived.
2026-08-27 16:11:37 +02:00
logaritmiskandClaude Opus 5 187aede924 docs: implementation plan for filtered estimates
Five tasks: delete the inert online machinery, add filtered_log_evidence,
add the two learning-curve methods, pin the invariants, record the API break.

Two spec corrections fell out of writing it. The spec claimed filtered results
would be bit-identical before and after converge(); they cannot be. iteration
recomputes the colour partition only when from == 0, so a slice built by
repeated appends keeps insertion order until the first converge() reorders it,
and the scratch clone inherits whichever order it finds — same fixed point,
different path. Corrected to agreement within 1e-8 under tight convergence,
matching the house pattern in tests/ingestion_equivalence.rs. The spec also
declared filtered_pass as Vec<(T, Vec<(Index, Gaussian)>)>, which cannot carry
the evidence its own step 3 harvests; it returns Vec<(T, FilteredStep)>.

CHANGELOG.md is generated by git-cliff, so the spec's "CHANGELOG records the
API break" cannot be satisfied by editing the file — it regenerates. Task 5
records the break through the commit subject and verifies the generated output
instead. cliff.toml has no breaking-change parser at all, which the task is
told to report rather than work around.

An adversarial reviewer checked the plan against the source before this commit
and found four real defects, all in plan text, none in the design:

- Two prescribed mutations provably could not fail their named tests. The
  learning-curve mutation altered only what filtered_pass writes after a slice,
  while the test inspected filtered[0], which is computed from an empty message
  map. Fixed by asserting monotonic mu across the whole curve.
- The ingestion-order fixture used four distinct timestamps, giving one event
  per slice — the exact degenerate shape ingestion_equivalence.rs documents as
  the weak case, making the assertion true by construction. Fixed to several
  events per timestamp with shared competitors.
- filtered_learning_curves was never asserted for content, only for emptiness
  on an empty history.
- A doc comment restated learning_curves' claim that key(idx) is O(n) and the
  method O(n^2). KeyTable::key is self.reverse.get(idx.0) — O(1) — and the
  type's own doc says so. The claim predates reverse becoming a Vec. The plan
  now corrects the original at history.rs:323 rather than copying it.

The reviewer confirmed the central claim by tracing the call graph: N_INF is
{pi: 0, tau: 0} and Mul is a natural-parameter add, so it is an exact
multiplicative identity, and the only write to skill.backward in the crate is
in new_backward_info, reachable only from History::iteration and never from
iterate_to_convergence under either rayon cfg.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T5SYDExxL4vZgvunrcNSMc
2026-08-27 16:04:16 +02:00
logaritmiskandClaude Opus 5 4fde482e48 docs: spec for filtered (forward-only) estimates
`HistoryBuilder::online(true)` is inert: it flips a flag that reaches
`Item::within_prior`, which reads `Skill.online` — a field initialised to
`N_INF` and assigned nowhere. So `log_evidence()` under that setting reports
`n * ln(0.5)`, every game scored as a coin flip. The number is finite and
plausible, which is why nothing caught it.

Issue #19 proposed populating the field during the forward pass. That does
not work, and the reason shapes the whole design. `new_forward_info` sets
`skill.forward` from the previous slice's `forward_prior_out`, which is
`skill.forward * skill.likelihood`; `History::iteration` alternates backward
and forward sweeps, so from the second iteration onward that likelihood has
already absorbed backward information. After `converge()`, `skill.forward`
is a smoothed quantity — and so is anything written from it.

The same reasoning condemns the neighbouring `forward: bool` flag, which is
a filtering quantity only on a history that was never converged. That is why
the test at history.rs:1183 can assert the two evidences are equal. Left
alone here; recorded as a follow-up.

The design is a read-only forward-only pass instead: walk slices in time
order carrying their own forward messages, and per slice build a scratch
clone whose `backward` is `N_INF`, then run the unmodified production sweep
on it. Reusing `iterate_to_convergence` rather than reimplementing inference
means a competitor playing twice at one time is handled by the same
within-slice EP that `converge()` uses, instead of being approximated the
way today's evidence paths approximate it. Nothing is stored on `Skill`,
which drops 16 bytes and helps #17 regardless.

Three methods ship — `filtered_log_evidence`, `filtered_learning_curves`,
`filtered_learning_curve` — all taking `&self`. The second consumer is
ustat, whose learning curves start already collapsed to sigma 0.9-1.6
against a prior of 6.0 because every point is smoothed; the filtered view
cannot be reconstructed from the public API today except by O(n^2) refits.

The red test brackets the issue's own fixture strictly between 5*ln(0.5) and
the batch evidence, so neither "still inert" nor "accidentally smoothed"
passes. The invariant that would have caught this bug class is that filtered
results are identical before and after `converge()` — exactly what a stored
field cannot give.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T5SYDExxL4vZgvunrcNSMc
2026-08-27 15:42:54 +02:00
16 changed files with 2494 additions and 59 deletions
+15
View File
@@ -0,0 +1,15 @@
# `Cargo.toml` sets `publish = ["kellnr"]`, so `cargo publish` targets the
# private registry and refuses crates.io. Cargo needs that registry's index
# declared to resolve the name.
#
# Committed rather than left to a per-user `~/.cargo/config.toml` so the repo
# is self-contained: a fresh clone, a new machine, or CI would otherwise fail
# with
#
# error: registry index was not found in any configuration: `kellnr`
#
# Index URL only — it is not a secret. Publish tokens live in
# `~/.cargo/credentials.toml` (per-user, never committed) or, in CI, in
# `CARGO_REGISTRIES_KELLNR_TOKEN`.
[registries.kellnr]
index = "sparse+https://crates.aceofba.se/api/v1/crates/"
+55
View File
@@ -2,6 +2,57 @@
All notable changes to this project will be documented in this file.
## 0.2.0 - 2026-08-27
### Breaking Changes
- refactor!: remove the inert online flag
### Bug Fixes
- fix: reject ties without draw probability; never report NaN as converged
- fix(quality): support any number of rating groups
- fix(evidence): accumulate in log space and floor the per-link value
- fix(history): stop reprocessing the slice that was just appended to
- fix(rayon): remove the aliasing unsafe from the parallel sweep
- fix: close out four small issues and pin #27's repro
### Documentation
- docs: refresh README and CLAUDE.md; add ingest benchmark
- docs: spec for filtered (forward-only) estimates
- docs: implementation plan for filtered estimates
- docs: state filtered accessor cost and evidence semantics precisely
- docs(cargo): correct the licence note — kellnr does not require one
### Features
- feat: add filtered_log_evidence
- feat: add filtered learning curves
### Miscellaneous Tasks
- chore: add CI, crate metadata, and crate-level documentation
- chore: target releases at the private kellnr registry
- chore: keep the 48 MB ATP dataset out of the published crate
- chore: dual-license MIT OR Apache-2.0
### Performance
- perf(gaussian): drop the sqrt round-trip from variance-space operations
### Refactor
- refactor: unify convergence defaults, validate builders, clear dead code
### Styling
- style: make NaN rejection explicit in score_sigma validation
### Testing
- test: pin the invariants that make filtered estimates trustworthy
## 0.1.2 - 2026-06-12
### Bug Fixes
@@ -32,6 +83,10 @@ All notable changes to this project will be documented in this file.
- feat(outcome): per-event score_sigma override on Outcome::Scored
- feat(event_builder): expose scores_with_sigma fluent method
### Miscellaneous Tasks
- chore: Release trueskill-tt version 0.1.2
### Refactor
- refactor: dedupe Game::likelihoods and likelihoods_scored via run_chain
+18 -6
View File
@@ -1,18 +1,30 @@
[package]
name = "trueskill-tt"
version = "0.1.2"
version = "0.2.0"
edition = "2024"
rust-version = "1.85"
description = "TrueSkill Through Time: Bayesian skill rating that tracks how skill evolves over time, via Gaussian message passing"
repository = "https://git.aceofba.se/logaritmisk/trueskill-tt"
authors = ["Anders Olsson"]
# Publishing is restricted to the private kellnr registry; this also makes
# an accidental `cargo publish` to crates.io a hard error rather than a
# irreversible mistake. Index is declared in `.cargo/config.toml`.
publish = ["kellnr"]
readme = "README.md"
keywords = ["trueskill", "rating", "bayesian", "elo", "skill"]
categories = ["algorithms", "science", "game-development"]
# TODO: pick a licence before publishing. `cargo publish` rejects a crate
# without `license` (or `license-file`), and without one the source carries no
# stated terms. The Rust convention is `license = "MIT OR Apache-2.0"` plus the
# matching LICENSE-MIT / LICENSE-APACHE files.
exclude = ["/docs", "/benches/*.txt", "/temp", "/.gitea"]
license = "MIT OR Apache-2.0"
# `examples/atp.csv` is a 48 MB tennis dataset — 99% of the packaged crate,
# for a library whose source is 312 KB. `examples/atp.rs` opens it by
# relative path at runtime, so excluding the data still compiles; the
# example just needs the file fetched from the repo to run.
exclude = [
"/docs",
"/benches/*.txt",
"/temp",
"/.gitea",
"/examples/atp.csv",
]
[lib]
bench = false
+46
View File
@@ -43,3 +43,49 @@ bench:
flame:
cargo flamegraph --root --example atp
# ---------------------------------------------------------------------------
# Release workflow
#
# Publishing goes to the private kellnr registry only: `Cargo.toml` sets
# `publish = ["kellnr"]`, so an accidental `cargo publish` to crates.io is a
# hard error rather than an irreversible mistake. The index is declared in the
# committed `.cargo/config.toml`; the token is per-user and lives in
# `~/.cargo/credentials.toml` (`cargo login --registry kellnr`).
#
# Step 1: just release-plan [level] — dry run, no writes
# Step 2: just release [level] — bump, changelog, tag, publish, push
#
# LEVEL is the cargo-release bump level (default `minor`). On 0.x:
# minor -> breaking bump (0.1.2 -> 0.2.0) <- any public-API change
# patch -> additive only (0.1.2 -> 0.1.3)
# major -> reserved for the 1.0.0 jump
#
# `release.toml` regenerates CHANGELOG.md with git-cliff in a pre-release hook
# and keeps push = false; this recipe pushes last, after publish has succeeded.
# ---------------------------------------------------------------------------
# Dry-run preview of the next release. Inspect the version bump and the
# "Publishing ..." line before running `just release`.
release-plan level="minor":
cargo release {{level}}
# Cut a release from a clean main: gate -> bump -> tag -> publish -> push.
release level="minor":
#!/usr/bin/env bash
set -euo pipefail
if [[ "$(git branch --show-current)" != "main" ]]; then
echo "error: run 'just release' from the 'main' branch" >&2; exit 1
fi
if [[ -n "$(git status --porcelain)" ]]; then
echo "error: working tree is dirty — commit or stash first" >&2; exit 1
fi
# cargo-release only verify-compiles the packaged crate; it does not run the
# suite, and publishing is irreversible. Run the same gate CI does, which
# includes the release profile where debug_assert! is compiled out.
just ci
cargo release {{level}} --execute --no-confirm
git push --follow-tags
+201
View File
@@ -0,0 +1,201 @@
Apache License
Version 2.0, January 2004
http://www.apache.org/licenses/
TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
1. Definitions.
"License" shall mean the terms and conditions for use, reproduction,
and distribution as defined by Sections 1 through 9 of this document.
"Licensor" shall mean the copyright owner or entity authorized by
the copyright owner that is granting the License.
"Legal Entity" shall mean the union of the acting entity and all
other entities that control, are controlled by, or are under common
control with that entity. For the purposes of this definition,
"control" means (i) the power, direct or indirect, to cause the
direction or management of such entity, whether by contract or
otherwise, or (ii) ownership of fifty percent (50%) or more of the
outstanding shares, or (iii) beneficial ownership of such entity.
"You" (or "Your") shall mean an individual or Legal Entity
exercising permissions granted by this License.
"Source" form shall mean the preferred form for making modifications,
including but not limited to software source code, documentation
source, and configuration files.
"Object" form shall mean any form resulting from mechanical
transformation or translation of a Source form, including but
not limited to compiled object code, generated documentation,
and conversions to other media types.
"Work" shall mean the work of authorship, whether in Source or
Object form, made available under the License, as indicated by a
copyright notice that is included in or attached to the work
(an example is provided in the Appendix below).
"Derivative Works" shall mean any work, whether in Source or Object
form, that is based on (or derived from) the Work and for which the
editorial revisions, annotations, elaborations, or other modifications
represent, as a whole, an original work of authorship. For the purposes
of this License, Derivative Works shall not include works that remain
separable from, or merely link (or bind by name) to the interfaces of,
the Work and Derivative Works thereof.
"Contribution" shall mean any work of authorship, including
the original version of the Work and any modifications or additions
to that Work or Derivative Works thereof, that is intentionally
submitted to Licensor for inclusion in the Work by the copyright owner
or by an individual or Legal Entity authorized to submit on behalf of
the copyright owner. For the purposes of this definition, "submitted"
means any form of electronic, verbal, or written communication sent
to the Licensor or its representatives, including but not limited to
communication on electronic mailing lists, source code control systems,
and issue tracking systems that are managed by, or on behalf of, the
Licensor for the purpose of discussing and improving the Work, but
excluding communication that is conspicuously marked or otherwise
designated in writing by the copyright owner as "Not a Contribution."
"Contributor" shall mean Licensor and any individual or Legal Entity
on behalf of whom a Contribution has been received by Licensor and
subsequently incorporated within the Work.
2. Grant of Copyright License. Subject to the terms and conditions of
this License, each Contributor hereby grants to You a perpetual,
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
copyright license to reproduce, prepare Derivative Works of,
publicly display, publicly perform, sublicense, and distribute the
Work and such Derivative Works in Source or Object form.
3. Grant of Patent License. Subject to the terms and conditions of
this License, each Contributor hereby grants to You a perpetual,
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
(except as stated in this section) patent license to make, have made,
use, offer to sell, sell, import, and otherwise transfer the Work,
where such license applies only to those patent claims licensable
by such Contributor that are necessarily infringed by their
Contribution(s) alone or by combination of their Contribution(s)
with the Work to which such Contribution(s) was submitted. If You
institute patent litigation against any entity (including a
cross-claim or counterclaim in a lawsuit) alleging that the Work
or a Contribution incorporated within the Work constitutes direct
or contributory patent infringement, then any patent licenses
granted to You under this License for that Work shall terminate
as of the date such litigation is filed.
4. Redistribution. You may reproduce and distribute copies of the
Work or Derivative Works thereof in any medium, with or without
modifications, and in Source or Object form, provided that You
meet the following conditions:
(a) You must give any other recipients of the Work or
Derivative Works a copy of this License; and
(b) You must cause any modified files to carry prominent notices
stating that You changed the files; and
(c) You must retain, in the Source form of any Derivative Works
that You distribute, all copyright, patent, trademark, and
attribution notices from the Source form of the Work,
excluding those notices that do not pertain to any part of
the Derivative Works; and
(d) If the Work includes a "NOTICE" text file as part of its
distribution, then any Derivative Works that You distribute must
include a readable copy of the attribution notices contained
within such NOTICE file, excluding those notices that do not
pertain to any part of the Derivative Works, in at least one
of the following places: within a NOTICE text file distributed
as part of the Derivative Works; within the Source form or
documentation, if provided along with the Derivative Works; or,
within a display generated by the Derivative Works, if and
wherever such third-party notices normally appear. The contents
of the NOTICE file are for informational purposes only and
do not modify the License. You may add Your own attribution
notices within Derivative Works that You distribute, alongside
or as an addendum to the NOTICE text from the Work, provided
that such additional attribution notices cannot be construed
as modifying the License.
You may add Your own copyright statement to Your modifications and
may provide additional or different license terms and conditions
for use, reproduction, or distribution of Your modifications, or
for any such Derivative Works as a whole, provided Your use,
reproduction, and distribution of the Work otherwise complies with
the conditions stated in this License.
5. Submission of Contributions. Unless You explicitly state otherwise,
any Contribution intentionally submitted for inclusion in the Work
by You to the Licensor shall be under the terms and conditions of
this License, without any additional terms or conditions.
Notwithstanding the above, nothing herein shall supersede or modify
the terms of any separate license agreement you may have executed
with Licensor regarding such Contributions.
6. Trademarks. This License does not grant permission to use the trade
names, trademarks, service marks, or product names of the Licensor,
except as required for reasonable and customary use in describing the
origin of the Work and reproducing the content of the NOTICE file.
7. Disclaimer of Warranty. Unless required by applicable law or
agreed to in writing, Licensor provides the Work (and each
Contributor provides its Contributions) on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
implied, including, without limitation, any warranties or conditions
of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
PARTICULAR PURPOSE. You are solely responsible for determining the
appropriateness of using or redistributing the Work and assume any
risks associated with Your exercise of permissions under this License.
8. Limitation of Liability. In no event and under no legal theory,
whether in tort (including negligence), contract, or otherwise,
unless required by applicable law (such as deliberate and grossly
negligent acts) or agreed to in writing, shall any Contributor be
liable to You for damages, including any direct, indirect, special,
incidental, or consequential damages of any character arising as a
result of this License or out of the use or inability to use the
Work (including but not limited to damages for loss of goodwill,
work stoppage, computer failure or malfunction, or any and all
other commercial damages or losses), even if such Contributor
has been advised of the possibility of such damages.
9. Accepting Warranty or Additional Liability. While redistributing
the Work or Derivative Works thereof, You may choose to offer,
and charge a fee for, acceptance of support, warranty, indemnity,
or other liability obligations and/or rights consistent with this
License. However, in accepting such obligations, You may act only
on Your own behalf and on Your sole responsibility, not on behalf
of any other Contributor, and only if You agree to indemnify,
defend, and hold each Contributor harmless for any liability
incurred by, or claims asserted against, such Contributor by reason
of your accepting any such warranty or additional liability.
END OF TERMS AND CONDITIONS
APPENDIX: How to apply the Apache License to your work.
To apply the Apache License to your work, attach the following
boilerplate notice, with the fields enclosed by brackets "[]"
replaced with your own identifying information. (Don't include
the brackets!) The text should be enclosed in the appropriate
comment syntax for the file format. We also recommend that a
file or class name and description of purpose be included on the
same "printed page" as the copyright notice for easier
identification within third-party archives.
Copyright [yyyy] [name of copyright owner]
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.
+19
View File
@@ -0,0 +1,19 @@
Copyright (c) 2026 Anders Olsson
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in
all copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN
THE SOFTWARE.
+17
View File
@@ -101,3 +101,20 @@ h.converge().unwrap();
- [x] Add Observer (`Observer` / `NullObserver`)
- [x] Benchmark the inference loop (`benches/batch.rs`, `benches/history_converge.rs`, `benches/ingest.rs`)
- [ ] Cross-check `quality()` against [sublee/trueskill](https://github.com/sublee/trueskill/tree/master) — N-group support works and is covered by invariants, but no reference values are asserted
## License
Licensed under either of
- Apache License, Version 2.0 ([LICENSE-APACHE](LICENSE-APACHE) or
<http://www.apache.org/licenses/LICENSE-2.0>)
- MIT license ([LICENSE-MIT](LICENSE-MIT) or
<http://opensource.org/licenses/MIT>)
at your option.
### Contribution
Unless you explicitly state otherwise, any contribution intentionally submitted
for inclusion in the work by you, as defined in the Apache-2.0 license, shall be
dual licensed as above, without any additional terms or conditions.
+5
View File
@@ -44,6 +44,11 @@ split_commits = false
# Assigns commits to groups.
# Optionally sets the commit's scope and can decide to exclude commits from further processing.
commit_parsers = [
# Must precede the type parsers below: a `feat!`/`fix!`/`refactor!` subject
# matches those too, and the first match wins. Without this a breaking
# change renders as an ordinary line of its own type.
{ message = "^[a-z]+(\\(.+\\))?!:", group = "Breaking Changes" },
{ body = "BREAKING CHANGE", group = "Breaking Changes" },
{ message = "^feat", group = "Features" },
{ message = "^fix", group = "Bug Fixes" },
{ message = "^doc", group = "Documentation" },
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,342 @@
# Filtered (Forward-Only) Estimates
Closes [#19](https://git.aceofba.se/logaritmisk/trueskill-tt/issues/19).
## Summary
`HistoryBuilder::online(true)` is inert. It flips a flag that reaches
`Item::within_prior` (`src/time_slice.rs:70-71`), which reads
`Skill.online` (`src/time_slice.rs:25`) — a field initialised to `N_INF`
(`src/time_slice.rs:41`) and never assigned anywhere. The online path
therefore builds every rating from the improper Gaussian, and
`log_evidence()` silently reports `n × ln(0.5)`: every game scored as a
coin flip, finite and plausible-looking.
This spec replaces the field and the flag with a **read-only forward-only
pass** over the converged history, exposed as three new public methods.
The pass reuses the production within-slice sweep verbatim rather than
reimplementing inference, and stores nothing on `Skill`.
## Background
### Why a stored field cannot hold this quantity
The issue proposes populating `skill.online` during the forward pass,
alongside `new_forward_info` (`src/time_slice.rs:576`). That would not
work, and understanding why determines the whole design.
`new_forward_info` sets `skill.forward` from
`agents[a].receive_for_elapsed(...)`, whose `message` was written by the
previous slice's `forward_prior_out` (`src/time_slice.rs:549`):
```rust
skill.forward * skill.likelihood
```
`History::iteration` (`src/history.rs:255`) alternates a backward sweep
over slices and a forward sweep. From the second iteration onward, the
`skill.likelihood` feeding that message has already absorbed backward
information from the preceding backward sweep. So after `converge()`,
**`skill.forward` is a smoothed quantity, not a filtering one** — and any
field written from it inherits the same contamination on every sweep
after the first.
### The neighbouring trap
The same reasoning applies to the existing `forward: bool` parameter on
`log_evidence_internal` (`src/history.rs:395`). It is a genuine filtering
quantity only on a history that has never been converged. That is why the
test at `src/history.rs:1183` can assert
```rust
assert_ulps_eq!(trueskill_log_evidence, trueskill_log_evidence_online, epsilon = 1e-6);
```
— the fixture is never converged, so the forward message still equals the
cavity prior. (Note also that the local binding is named `..._online`
while the flag it passes is `forward`; the two senses were already
muddled.)
Fixing `forward: bool` is **out of scope** here; see *Out-of-scope
follow-ups*.
### Why this is worth implementing rather than deleting
The forward-only estimate has a second consumer beyond prequential model
comparison. `learning_curve()` returns post-convergence posteriors, so
every point is smoothed — the estimate at a given date incorporates
rounds played years later. On [ustat](https://git.aceofba.se/logaritmisk/ustat)'s
real data (prior μ=0, σ=6) that produces curves which start already
spread apart and barely move:
```
player first point final point
Eskil mu +3.72 sigma 1.17 mu +4.61 sigma 1.21
Anders Olsson mu +1.61 sigma 0.90 mu +1.16 sigma 0.82
LUDVIGSSON mu -2.09 sigma 1.08 mu -2.61 sigma 1.13
Anners mu -2.85 sigma 1.27 mu -2.86 sigma 1.26
```
σ at the *first* plotted point is 0.901.60 against a prior of 6.00. A
caller cannot reconstruct the filtered view from the public API today
except by refitting over `events[0..k]` for every k — O(n²) fits for
something one forward pass already computes.
## Scope
### What ships
1. A read-only forward-only pass on `History`, walking slices in time
order and carrying its own forward messages.
2. Three public methods: `filtered_log_evidence`,
`filtered_learning_curves`, `filtered_learning_curve`.
3. Removal of `Skill.online`, `History.online`, `HistoryBuilder.online`,
`HistoryBuilder::online()`, and the `online: bool` parameter threaded
through `Item::within_prior`, `Event::within_priors`, and
`TimeSlice::log_evidence`.
4. `#[derive(Clone)]` on `Event`, `Team`, `Item`; `iterate_to_convergence`
loses its `#[cfg(test)]` gate.
5. A CHANGELOG entry recording the API break.
### What does not ship
- No change to `log_evidence()`, `log_evidence_for()`, `learning_curve()`,
`learning_curves()`, or `current_skill()`. Their values are unchanged
by this work.
- No fix to the `forward: bool` flag described above.
- No caching of pass results. Each call runs a full pass; the doc
comments say so.
- No `rayon` parallelism across slices — the pass is sequentially
dependent by construction.
- No prior-predictive accessor. The pass computes the pre-event forward
message internally, but only the filtered posterior is exposed until a
second caller needs otherwise.
## Design
### Naming
`filtered_*`, not `online_*`. "Filtered" is the standard term for the
forward-only estimate, and the crate already uses "online" for a second,
unrelated thing — incremental ingestion, which `benches/baseline.txt:128`
calls the "online-add" path. Two senses of one word in one crate is how
the present bug reads as plausible.
### The pass
```rust
pub(crate) struct FilteredStep {
log_evidence: f64,
posteriors: Vec<(Index, Gaussian)>,
}
fn filtered_pass(&self) -> Vec<(T, FilteredStep)>
```
`posteriors` doubles as the outgoing forward message: the scratch sweep never
writes `backward`, so it stays `N_INF`, and `Skill::posterior()` and
`forward_prior_out` are then the same product.
Walk `self.time_slices` in order, carrying
`messages: HashMap<Index, Gaussian>` — the forward message out of each
competitor's most recent appearance. For each slice:
1. **Build a scratch clone.** Same `time`, `p_draw`, `convergence`, and
cloned `events` with every `item.likelihood` reset to `N_INF`. Fresh
`SkillStore` in which, for each agent present in the real slice:
```rust
forward = match messages.get(&agent) {
Some(msg) => msg.forget(rating.drift.variance_for_elapsed(skill.elapsed)),
None => rating.prior,
}
backward = N_INF
likelihood = N_INF
elapsed = skill.elapsed // copied from the real slice
```
This mirrors `Competitor::receive_for_elapsed` (`src/competitor.rs:39`)
exactly, including its `message != N_INF` fallback to the prior.
`skill.elapsed` is reused rather than recomputed: it is maintained by
`add_events_with_prior` across out-of-order ingestion, and production
convergence already trusts it.
2. **Run the real sweep.** `scratch.iterate_to_convergence(agents)`
(`src/time_slice.rs:516`), unmodified. Fidelity comes from reusing the
production path rather than a parallel reimplementation — in
particular, a competitor appearing in two events at the same time is
handled by the same within-slice EP that `converge()` uses, not
approximated the way the current `online`/`forward` evidence paths are
(they run each event independently and sum).
3. **Harvest.** With `backward == N_INF` acting as the multiplicative
identity, `Skill::posterior()` is exactly forward × likelihood — the
filtered posterior. Slice evidence is
`scratch.events.iter().map(|e| e.log_evidence).sum()`; `apply`
(`src/time_slice.rs:162`) writes that field on every event during the
sweep.
4. **Carry forward.** `messages.insert(a, scratch.forward_prior_out(&a))`
for each agent in the slice.
Steps 14 are the forward half of `History::iteration`
(`src/history.rs:283-297`) with the backward half never run. The pass
touches no field of `self`.
### Public API
```rust
impl<T: Time, D: Drift<T>, O: Observer<T>, K: Eq + Hash + Clone> History<T, D, O, K> {
pub fn filtered_log_evidence(&self) -> f64;
pub fn filtered_learning_curves(&self) -> HashMap<K, Vec<(T, Gaussian)>>;
pub fn filtered_learning_curve<Q>(&self, key: &Q) -> Vec<(T, Gaussian)>
where
K: Borrow<Q>,
Q: Hash + Eq + ?Sized;
}
```
All take `&self` — the pass mutates nothing. Shapes deliberately mirror
`learning_curve` / `learning_curves` (`src/history.rs:325`, `:381`) so a
caller can plot smoothed and filtered curves on one chart with the same
handling code.
`filtered_learning_curve` runs the same full pass as the plural form and
collects one key; the cost is identical, only the collection differs.
Callers wanting several keys should use the plural form. Documented on
both methods.
Because the pass carries its own messages and re-runs inference, its
results **do not depend on whether `converge()` has been called**. That
is the property a stored field cannot have, and it is asserted as a test.
### Removal inventory
| Location | Change |
|---|---|
| `src/time_slice.rs:25` | delete `pub(crate) online: Gaussian` |
| `src/time_slice.rs:41` | delete `online: N_INF` from `Default` |
| `src/time_slice.rs:62,70-73` | drop `online` param and its branch from `Item::within_prior` |
| `src/time_slice.rs:110,120` | drop `online` param from `Event::within_priors` |
| `src/time_slice.rs:585,597,626,634` | drop `online` param from `TimeSlice::log_evidence`; `online \|\| forward` becomes `forward` |
| `src/history.rs:32,63,138,158,174,199,226` | delete the two `online` field declarations (`:32`, `:199`) and the five struct-literal copies |
| `src/history.rs:90-93` | delete `HistoryBuilder::online()` |
| `src/history.rs:402,410` | drop the `self.online` argument |
| `src/history.rs:1183-1189` | the `..._online` assertion becomes a `forward`-flag assertion; rename the binding to match what it tests |
`Skill` loses 16 bytes, which is a small independent win for #17.
## Testing strategy
Every new test is mutation-proved before it counts: break the production
line it names, watch it fail for the *right* assertion, restore. A test
never observed failing is not evidence.
### The red test
On the issue's own fixture — five 1v1 games, same winner each time —
`filtered_log_evidence()` must land strictly between the two known
endpoints:
```
5 × ln(0.5) = -3.4657... (today's inert value)
< filtered
< -0.4012... (batch / smoothed evidence)
```
Two-sided, so neither "still inert" nor "accidentally smoothed" can pass.
The lower bound is right for a real reason: game one genuinely *is* a
coin flip under filtering, games two through five are not.
### Invariants
1. **Invariant to `converge()`** — `filtered_log_evidence()` and
`filtered_learning_curves()` agree before and after `converge()`. This
is exactly what `skill.forward` fails, and what makes a stored field
the wrong mechanism.
Agreement is to tolerance, not bit-identity, and the reason is worth
recording. `iteration` calls `recompute_color_groups`
(`src/time_slice.rs:369`) only when `from == 0`, so a slice built by
repeated appends keeps insertion order until the first `converge()`
reorders it. The scratch clone inherits whichever order it finds, and
greedy coloring over a permuted input can group differently, giving a
different within-slice sweep order — same EP fixed point, different
path to it. Follow the house pattern in
`tests/ingestion_equivalence.rs`: converge tightly (`max_iter: 2_000`,
`epsilon: 1e-12`) and compare within `1e-8`.
2. **Invariant to ingestion order** — events added one at a time produce
the same filtered results as the same events batched. Extends the
existing invariant in `tests/ingestion_equivalence.rs`.
3. **Single-slice exactness** — for a history with one time slice there
is no future to propagate back, so filtered results equal smoothed
results exactly.
4. **Uncertainty ordering** — for a competitor with many later games, σ
at the first filtered point is greater than σ at the first smoothed
point, and less than the prior σ. This is the ustat complaint restated
as an assertion.
5. **Degenerate inputs** — empty history yields `0.0` and empty maps;
unknown key yields an empty curve. Added to
`tests/degenerate_inputs.rs`.
### Regression net
The existing suite must be unchanged by the removals: `log_evidence()`,
`log_evidence_for()`, and every numerical golden keep their current
values, since the default `online` was already `false` and the flag was
inert.
## Verification gates
- `just test` — full matrix, including the release job. `debug_assert!`
is compiled out in release, and that is where defects in this crate
have hidden before.
- `just lint` — clippy, warnings denied.
- `just fmt` — nightly.
- `just determinism` — the new pass must not perturb bit-identical
posteriors across `RAYON_NUM_THREADS` 1/2/4/8.
- `#![forbid(unsafe_code)]` stays.
## Risks
- **Clone cost.** One slice's events are cloned per slice visited. At
ustat scale this is negligible, but the pass is O(events) allocation on
top of O(events) inference. Accepted: fidelity to the production sweep
is worth more than avoiding the clone, and no caller is on a hot path.
- **`iterate_to_convergence` leaving test-only status.** Its doc comment
claims "only used by tests"; that comment must be updated, or it
becomes the next piece of load-bearing prose that is quietly false.
- **Event order is inherited, not normalised.** The scratch clone takes
the real slice's current event order, which differs pre- and
post-`converge()` for incrementally-ingested slices (see *Invariants*).
Results agree to within convergence tolerance rather than exactly.
Normalising the order in the scratch builder would buy bit-identity at
the cost of diverging from what the real sweep does; not worth it.
**Measured after implementation, this risk is smaller than stated.**
Flipping the scratch's `color_groups_dirty` from `true` to `false`
switches it between the grouped sweep (`sweep_color_groups`) and the
sequential fallback across its entire convergence loop — a far larger
perturbation than a permuted event order — and the ingestion-order
invariance test stays green at `1e-8` under `max_iter: 2_000`,
`epsilon: 1e-12`. EP reaches the same fixed point regardless of sweep
order once driven far enough. The tolerance caveat is correct but
conservative. Note the flag itself is load-bearing: with it `false` the
scratch would take the sequential path always, diverging from the
production sweep it exists to mirror.
- **Divergence risk.** If `TimeSlice`'s sweep gains state that the
scratch construction does not initialise, the pass silently reads a
default. The scratch builder must construct `Skill` field-by-field
rather than via `..Default::default()`, so adding a field to `Skill`
is a compile error here rather than a silent wrong answer.
## Out-of-scope follow-ups
File as separate issues:
1. **`forward: bool` is only a filtering quantity pre-convergence**
(`src/history.rs:395`). Either document the constraint or fold the
flag into the new pass and delete it.
2. **`log_evidence` takes `&mut self`** (`src/history.rs:416`) but
mutates nothing. The new `filtered_*` methods take `&self`; the
asymmetry is worth removing.
+5 -1
View File
@@ -1,2 +1,6 @@
publish = false
# Publish to the registry named in Cargo.toml's `publish` list (kellnr).
publish = true
# Hold off pushing until tags and publish have both succeeded; `just release`
# pushes last.
push = false
pre-release-hook = ["sh", "-c", "git cliff -o CHANGELOG.md --tag {{version}} && git add CHANGELOG.md"]
-1
View File
@@ -103,7 +103,6 @@ impl ColorGroups {
/// `Index` values that event touches. The returned `ColorGroups` has one
/// inner `Vec<usize>` per color, containing event indices in the order
/// they were assigned.
#[allow(dead_code)]
pub(crate) fn color_greedy<I, F>(n_events: usize, index_set: F) -> ColorGroups
where
F: Fn(usize) -> I,
+116 -29
View File
@@ -13,7 +13,7 @@ use crate::{
sort_time,
storage::CompetitorStore,
time::Time,
time_slice::{self, EventKind, TimeSlice},
time_slice::{self, EventKind, FilteredStep, TimeSlice},
tuple_gt, tuple_max,
};
@@ -29,7 +29,6 @@ pub struct HistoryBuilder<
beta: f64,
drift: D,
p_draw: f64,
online: bool,
score_sigma: f64,
convergence: ConvergenceOptions,
observer: O,
@@ -60,7 +59,6 @@ impl<T: Time, D: Drift<T>, O: Observer<T>, K: Eq + Hash + Clone> HistoryBuilder<
sigma: self.sigma,
beta: self.beta,
p_draw: self.p_draw,
online: self.online,
score_sigma: self.score_sigma,
convergence: self.convergence,
observer: self.observer,
@@ -87,11 +85,6 @@ impl<T: Time, D: Drift<T>, O: Observer<T>, K: Eq + Hash + Clone> HistoryBuilder<
self
}
pub fn online(mut self, online: bool) -> Self {
self.online = online;
self
}
/// Default observation noise for scored outcomes.
///
/// # Panics
@@ -135,7 +128,6 @@ impl<T: Time, D: Drift<T>, O: Observer<T>, K: Eq + Hash + Clone> HistoryBuilder<
beta: self.beta,
drift: self.drift,
p_draw: self.p_draw,
online: self.online,
score_sigma: self.score_sigma,
convergence: self.convergence,
observer,
@@ -155,7 +147,6 @@ impl<T: Time, D: Drift<T>, O: Observer<T>, K: Eq + Hash + Clone> HistoryBuilder<
beta: self.beta,
drift: self.drift,
p_draw: self.p_draw,
online: self.online,
score_sigma: self.score_sigma,
convergence: self.convergence,
observer: self.observer,
@@ -171,7 +162,6 @@ impl Default for HistoryBuilder<i64, ConstantDrift, NullObserver, &'static str>
beta: BETA,
drift: ConstantDrift(GAMMA),
p_draw: P_DRAW,
online: false,
score_sigma: 1.0,
convergence: ConvergenceOptions::default(),
observer: NullObserver,
@@ -196,7 +186,6 @@ pub struct History<
beta: f64,
drift: D,
p_draw: f64,
online: bool,
score_sigma: f64,
convergence: ConvergenceOptions,
observer: O,
@@ -223,7 +212,6 @@ impl<K: Eq + Hash + Clone> History<i64, ConstantDrift, NullObserver, K> {
beta: BETA,
drift: ConstantDrift(GAMMA),
p_draw: P_DRAW,
online: false,
score_sigma: 1.0,
convergence: ConvergenceOptions::default(),
observer: NullObserver,
@@ -319,9 +307,6 @@ impl<T: Time, D: Drift<T>, O: Observer<T>, K: Eq + Hash + Clone> History<T, D, O
}
/// Learning curves for all competitors, keyed by their user-facing key.
///
/// Note: `key(idx)` is O(n) per lookup; this method is therefore O(n²)
/// in the number of competitors. Acceptable for T2; T3 may optimize.
pub fn learning_curves(&self) -> HashMap<K, Vec<(T, Gaussian)>> {
#[cfg(feature = "rayon")]
{
@@ -392,14 +377,79 @@ impl<T: Time, D: Drift<T>, O: Observer<T>, K: Eq + Hash + Clone> History<T, D, O
.collect()
}
pub(crate) fn log_evidence_internal(&mut self, forward: bool, targets: &[Index]) -> f64 {
/// Filtered learning curves for all competitors, keyed by user-facing key.
///
/// Each point is the posterior using only events up to and including that
/// time — "what we knew then". Contrast `learning_curves`, whose points
/// are smoothed and so incorporate rounds played later.
///
/// Runs a full forward pass per call and caches nothing. This is the
/// entry point for multi-key work — see `filtered_learning_curve` for
/// why calling that once per key is far more expensive.
pub fn filtered_learning_curves(&self) -> HashMap<K, Vec<(T, Gaussian)>> {
let mut data: HashMap<K, Vec<(T, Gaussian)>> = HashMap::new();
for (time, step) in self.filtered_pass() {
for (agent, posterior) in step.posteriors {
if let Some(key) = self.keys.key(agent).cloned() {
data.entry(key).or_default().push((time, posterior));
}
}
}
data
}
/// Filtered learning curve for a single key: (time, posterior) pairs in
/// time order.
///
/// Despite mirroring `learning_curve`'s signature, this is not the cheap
/// per-key lookup that method is: it runs a full forward pass, O(events),
/// discarding every posterior but the requested key's. N keys fetched
/// this way costs O(N * events); use `filtered_learning_curves` for
/// multi-key work instead — it computes the same pass once.
pub fn filtered_learning_curve<Q>(&self, key: &Q) -> Vec<(T, Gaussian)>
where
K: Borrow<Q>,
Q: Hash + Eq + ?Sized,
{
let Some(idx) = self.keys.get(key) else {
return Vec::new();
};
self.filtered_pass()
.into_iter()
.filter_map(|(time, step)| {
step.posteriors
.iter()
.find(|(agent, _)| *agent == idx)
.map(|&(_, posterior)| (time, posterior))
})
.collect()
}
/// Sum per-slice evidence.
///
/// `forward` selects `skill.forward` as each event's prior instead of the
/// cavity. That is a genuine forward-only (filtering) quantity ONLY on a
/// history that has never been converged: `iteration` alternates backward
/// and forward sweeps, so from the second iteration onward the likelihood
/// feeding the forward message has already absorbed backward information.
/// For a filtering quantity that holds after convergence, use
/// `filtered_log_evidence`.
pub(crate) fn log_evidence_internal(&self, forward: bool, targets: &[Index]) -> f64 {
// Bound before the closure so it captures the store rather than all of
// `&self`: capturing `&History` would drag `KeyTable<K>` in and demand
// `K: Sync` from every caller, which the key type need not satisfy.
let agents = &self.agents;
#[cfg(feature = "rayon")]
{
use rayon::prelude::*;
let per_slice: Vec<f64> = self
.time_slices
.par_iter()
.map(|ts| ts.log_evidence(self.online, targets, forward, &self.agents))
.map(|ts| ts.log_evidence(targets, forward, agents))
.collect();
per_slice.into_iter().sum()
}
@@ -407,19 +457,19 @@ impl<T: Time, D: Drift<T>, O: Observer<T>, K: Eq + Hash + Clone> History<T, D, O
{
self.time_slices
.iter()
.map(|ts| ts.log_evidence(self.online, targets, forward, &self.agents))
.map(|ts| ts.log_evidence(targets, forward, agents))
.sum()
}
}
/// Total log-evidence across the history.
pub fn log_evidence(&mut self) -> f64 {
pub fn log_evidence(&self) -> f64 {
self.log_evidence_internal(false, &[])
}
/// Log-evidence restricted to time slices containing at least one of the
/// given keys. Useful for leave-one-out cross-validation.
pub fn log_evidence_for<Q>(&mut self, keys: &[&Q]) -> f64
pub fn log_evidence_for<Q>(&self, keys: &[&Q]) -> f64
where
K: std::borrow::Borrow<Q>,
Q: std::hash::Hash + Eq + ?Sized,
@@ -428,6 +478,48 @@ impl<T: Time, D: Drift<T>, O: Observer<T>, K: Eq + Hash + Clone> History<T, D, O
self.log_evidence_internal(false, &targets)
}
/// Walk the slices in time order carrying forward messages only.
///
/// This is the forward half of `iteration` with the backward half never
/// run. It reads `self` and mutates nothing.
fn filtered_pass(&self) -> Vec<(T, FilteredStep)> {
let mut messages: HashMap<Index, Gaussian> = HashMap::new();
let mut pass = Vec::with_capacity(self.time_slices.len());
for slice in &self.time_slices {
let step = slice.filtered_step(&messages, &self.agents);
for &(agent, posterior) in &step.posteriors {
messages.insert(agent, posterior);
}
pass.push((slice.time, step));
}
pass
}
/// Total log-evidence under forward-only (filtering) information.
///
/// Each event is scored using only what was known before that *time*,
/// which is the right quantity for prequential scoring and model
/// comparison. Events sharing a timestamp still inform each other
/// through the within-slice sweep, so within one slice this is not a
/// guarantee that event A is scored independently of simultaneous event
/// B. Contrast `log_evidence`, whose per-event priors carry information
/// from events that had not happened yet.
///
/// Runs a full forward pass per call and caches nothing. The result does
/// not depend on whether `converge` has been called.
#[must_use]
pub fn filtered_log_evidence(&self) -> f64 {
self.filtered_pass()
.iter()
.map(|(_, step)| step.log_evidence)
.sum()
}
/// Draw-probability quality metric for the given teams (key slices).
///
/// Values range roughly [0, 1]; 1 == perfectly matched. Supports any
@@ -935,12 +1027,7 @@ mod tests {
let w = [vec![1.0], vec![1.0]];
let p = Game::ranked_with_arena(
h.time_slices[1].events[0].within_priors(
false,
false,
&h.time_slices[1].skills,
&h.agents,
),
h.time_slices[1].events[0].within_priors(false, &h.time_slices[1].skills, &h.agents),
&[0.0, 1.0],
&w,
P_DRAW,
@@ -1180,11 +1267,11 @@ mod tests {
let f = h.keys.get("f").unwrap();
let trueskill_log_evidence = h.log_evidence_internal(false, &[]);
let trueskill_log_evidence_online = h.log_evidence_internal(true, &[]);
let trueskill_log_evidence_forward = h.log_evidence_internal(true, &[]);
assert_ulps_eq!(
trueskill_log_evidence,
trueskill_log_evidence_online,
trueskill_log_evidence_forward,
epsilon = 1e-6
);
+90 -21
View File
@@ -22,7 +22,6 @@ pub(crate) struct Skill {
backward: Gaussian,
likelihood: Gaussian,
pub(crate) elapsed: i64,
pub(crate) online: Gaussian,
}
impl Skill {
@@ -38,7 +37,6 @@ impl Default for Skill {
backward: N_INF,
likelihood: N_INF,
elapsed: 0,
online: N_INF,
}
}
}
@@ -50,7 +48,7 @@ pub enum EventKind {
Scored { score_sigma: f64 },
}
#[derive(Debug)]
#[derive(Clone, Debug)]
struct Item {
agent: Index,
likelihood: Gaussian,
@@ -59,7 +57,6 @@ struct Item {
impl Item {
fn within_prior<T: Time, D: Drift<T>>(
&self,
online: bool,
forward: bool,
skills: &SkillStore,
agents: &CompetitorStore<T, D>,
@@ -67,9 +64,7 @@ impl Item {
let r = &agents[self.agent].rating;
let skill = skills.get(self.agent).unwrap();
if online {
Rating::new(skill.online, r.beta, r.drift)
} else if forward {
if forward {
Rating::new(skill.forward, r.beta, r.drift)
} else {
Rating::new(skill.posterior() / self.likelihood, r.beta, r.drift)
@@ -77,13 +72,13 @@ impl Item {
}
}
#[derive(Debug)]
#[derive(Clone, Debug)]
struct Team {
items: Vec<Item>,
output: f64,
}
#[derive(Debug)]
#[derive(Clone, Debug)]
pub(crate) struct Event {
teams: Vec<Team>,
log_evidence: f64,
@@ -107,7 +102,6 @@ impl Event {
pub(crate) fn within_priors<T: Time, D: Drift<T>>(
&self,
online: bool,
forward: bool,
skills: &SkillStore,
agents: &CompetitorStore<T, D>,
@@ -117,7 +111,7 @@ impl Event {
.map(|team| {
team.items
.iter()
.map(|item| item.within_prior(online, forward, skills, agents))
.map(|item| item.within_prior(forward, skills, agents))
.collect::<Vec<_>>()
})
.collect::<Vec<_>>()
@@ -136,7 +130,7 @@ impl Event {
convergence: crate::ConvergenceOptions,
arena: &mut ScratchArena,
) -> EventUpdate {
let teams = self.within_priors(false, false, skills, agents);
let teams = self.within_priors(false, skills, agents);
let result = self.outputs();
let g = match self.kind {
EventKind::Ranked => {
@@ -195,6 +189,17 @@ struct EventUpdate {
likelihoods: Vec<Vec<Gaussian>>,
}
/// One slice's worth of forward-only inference.
///
/// `posteriors` doubles as the outgoing forward message: the scratch sweep
/// never writes `backward`, so it stays `N_INF`, and `Skill::posterior()`
/// and `forward_prior_out` are then the same product.
#[derive(Debug)]
pub(crate) struct FilteredStep {
pub(crate) log_evidence: f64,
pub(crate) posteriors: Vec<(Index, Gaussian)>,
}
#[derive(Debug)]
pub struct TimeSlice<T: Time = i64> {
pub(crate) events: Vec<Event>,
@@ -300,8 +305,9 @@ impl<T: Time> TimeSlice<T> {
*idx,
Skill {
forward: agents[*idx].receive(&self.time),
backward: N_INF,
likelihood: N_INF,
elapsed,
..Default::default()
},
);
}
@@ -372,7 +378,7 @@ impl<T: Time> TimeSlice<T> {
if from > 0 || self.color_groups.is_empty() {
// Initial pass (add_events) or no color groups yet: simple sequential sweep.
for event in self.events.iter_mut().skip(from) {
let teams = event.within_priors(false, false, &self.skills, agents);
let teams = event.within_priors(false, &self.skills, agents);
let result = event.outputs();
let g = match event.kind {
@@ -506,13 +512,13 @@ impl<T: Time> TimeSlice<T> {
/// Iterate this slice alone until its posteriors stop moving, returning
/// the number of iterations taken.
///
/// Only used by tests: production convergence is driven across slices by
/// `History::converge`.
/// Used by `filtered_step` to drive a scratch copy of the slice, and by
/// tests. Production convergence across slices is driven by
/// `History::converge`, which calls `iteration` directly.
///
/// Honours `self.convergence`; it previously hard-coded an epsilon and a
/// 20-iteration cap that matched neither `ConvergenceOptions` nor the
/// schedule default.
#[cfg(test)]
pub(crate) fn iterate_to_convergence<D: Drift<T>>(
&mut self,
agents: &CompetitorStore<T, D>,
@@ -580,9 +586,72 @@ impl<T: Time> TimeSlice<T> {
self.iteration(0, agents);
}
/// Run this slice's events on forward (filtering) information alone.
///
/// `incoming` holds each competitor's forward message out of their
/// previous appearance; a competitor absent from it starts at their
/// configured prior. The sweep runs on a scratch copy, so the real slice
/// is untouched — which is what makes the filtered estimates independent
/// of whether `History::converge` has run.
pub(crate) fn filtered_step<D: Drift<T>>(
&self,
incoming: &HashMap<Index, Gaussian>,
agents: &CompetitorStore<T, D>,
) -> FilteredStep {
let mut scratch = TimeSlice {
events: self.events.clone(),
skills: SkillStore::new(),
time: self.time,
p_draw: self.p_draw,
convergence: self.convergence,
arena: ScratchArena::new(),
color_groups: ColorGroups::new(),
color_groups_dirty: true,
};
for event in &mut scratch.events {
for team in &mut event.teams {
for item in &mut team.items {
item.likelihood = N_INF;
}
}
event.log_evidence = 0.0;
}
for (agent, skill) in self.skills.iter() {
let rating = &agents[agent].rating;
let forward = match incoming.get(&agent) {
Some(message) => message.forget(rating.drift.variance_for_elapsed(skill.elapsed)),
None => rating.prior,
};
scratch.skills.insert(
agent,
Skill {
forward,
backward: N_INF,
likelihood: N_INF,
elapsed: skill.elapsed,
},
);
}
scratch.iterate_to_convergence(agents);
FilteredStep {
log_evidence: scratch.events.iter().map(|event| event.log_evidence).sum(),
posteriors: scratch
.skills
.iter()
.map(|(agent, skill)| (agent, skill.posterior()))
.collect(),
}
}
pub(crate) fn log_evidence<D: Drift<T>>(
&self,
online: bool,
targets: &[Index],
forward: bool,
agents: &CompetitorStore<T, D>,
@@ -594,7 +663,7 @@ impl<T: Time> TimeSlice<T> {
let mut arena = ScratchArena::new();
let run_event = |event: &Event, arena: &mut ScratchArena| -> f64 {
let teams = event.within_priors(online, forward, &self.skills, agents);
let teams = event.within_priors(forward, &self.skills, agents);
let result = event.outputs();
match event.kind {
EventKind::Ranked => {
@@ -623,7 +692,7 @@ impl<T: Time> TimeSlice<T> {
};
if targets.is_empty() {
if online || forward {
if forward {
self.events
.iter()
.map(|event| run_event(event, &mut arena))
@@ -631,7 +700,7 @@ impl<T: Time> TimeSlice<T> {
} else {
self.events.iter().map(|event| event.log_evidence).sum()
}
} else if online || forward {
} else if forward {
self.events
.iter()
.filter(|event| {
+26 -1
View File
@@ -5,7 +5,7 @@
use trueskill_tt::{
ConstantDrift, ConvergenceOptions, Game, GameOptions, Gaussian, History, InferenceError,
Outcome, Rating,
NullObserver, Outcome, Rating,
};
type R = Rating<i64, ConstantDrift>;
@@ -127,6 +127,20 @@ fn empty_history_converges_trivially() {
assert!(report.converged);
}
/// Issue #27's exact reproduction: a non-default key type reaching `converge`
/// with no events at all. The underflow it reported trapped in debug and
/// indexed out of bounds in release, so this must run in both profiles.
#[test]
fn converge_on_an_empty_history_with_owned_keys() {
let mut history: History<i64, ConstantDrift, NullObserver, String> =
History::builder_with_key().score_sigma(5.0).build();
let report = history.converge().unwrap();
assert_eq!(report.iterations, 0);
assert!(report.converged);
}
#[test]
fn empty_event_stream_then_converge() {
let mut h = History::default();
@@ -245,3 +259,14 @@ fn log_evidence_finite_for_near_certain_outcome() {
upset.log_evidence()
);
}
#[test]
fn empty_history_has_no_filtered_estimates() {
let history: History = History::builder().build();
assert_eq!(history.filtered_log_evidence(), 0.0);
assert!(history.filtered_learning_curves().is_empty());
assert!(history.filtered_learning_curve("nobody").is_empty());
}
+254
View File
@@ -0,0 +1,254 @@
//! Forward-only (filtering) estimates: what the model knew at the time,
//! as opposed to the smoothed posteriors `learning_curve` reports.
use smallvec::smallvec;
use trueskill_tt::{ConvergenceOptions, Event, History, Member, Outcome, Team};
/// `games` one-on-one matches at successive times, won by "a" every time,
/// built with the given convergence options.
fn repeated_winner_with(games: i64, convergence: ConvergenceOptions) -> History {
let mut history = History::builder().convergence(convergence).build();
for time in 1..=games {
history
.add_events([Event {
time,
teams: smallvec![
Team::with_members([Member::new("a")]),
Team::with_members([Member::new("b")]),
],
outcome: Outcome::winner(0, 2),
}])
.unwrap();
}
history
}
/// `games` one-on-one matches at successive times, won by "a" every time.
///
/// This is the fixture from issue #19, where `online(true)` reported
/// `games * ln(0.5)`.
fn repeated_winner(games: i64) -> History {
repeated_winner_with(games, ConvergenceOptions::default())
}
/// The default 30-iteration cap leaves a residual around 1e-6, which would
/// swamp these comparisons. Drive both sides well past the fixed point.
fn tight() -> ConvergenceOptions {
ConvergenceOptions {
max_iter: 2_000,
epsilon: 1e-12,
..ConvergenceOptions::default()
}
}
#[test]
fn filtered_evidence_sits_between_coin_flip_and_batch() {
let mut history = repeated_winner(5);
history.converge().unwrap();
let coin_flip = 5.0 * 0.5f64.ln();
let batch = history.log_evidence();
let filtered = history.filtered_log_evidence();
assert!(
filtered > coin_flip,
"filtered evidence {filtered} is at or below {coin_flip}, the all-coin-flip \
value the inert online flag reported; game one is a coin flip but games two \
through five are not"
);
assert!(
filtered < batch,
"filtered evidence {filtered} is not below the smoothed {batch}; filtering \
scores each game on strictly less information than smoothing does"
);
}
#[test]
fn filtered_first_point_is_less_certain_than_smoothed() {
let mut history = repeated_winner(12);
history.converge().unwrap();
let smoothed = history.learning_curve("a");
let filtered = history.filtered_learning_curve("a");
assert_eq!(
smoothed.len(),
filtered.len(),
"both curves must cover the same time points"
);
let (smoothed_time, first_smoothed) = smoothed[0];
let (filtered_time, first_filtered) = filtered[0];
assert_eq!(smoothed_time, filtered_time);
assert!(
first_filtered.sigma() > first_smoothed.sigma(),
"filtered sigma {} at the first point is not above smoothed {}; the smoother \
collapses uncertainty before the first round is drawn, which is the whole \
reason this method exists",
first_filtered.sigma(),
first_smoothed.sigma()
);
assert!(
first_filtered.sigma() < trueskill_tt::SIGMA,
"filtered sigma {} at the first point is not below the prior {}; one game was \
played, so some uncertainty must have been resolved",
first_filtered.sigma(),
trueskill_tt::SIGMA
);
for pair in filtered.windows(2) {
assert!(
pair[1].1.mu() > pair[0].1.mu(),
"filtered mu must climb at every step for a competitor who wins every \
game: t={} mu={} then t={} mu={}",
pair[0].0,
pair[0].1.mu(),
pair[1].0,
pair[1].1.mu()
);
}
}
#[test]
fn filtered_curves_plural_agrees_with_singular() {
let mut history = repeated_winner(4);
history.converge().unwrap();
let curves = history.filtered_learning_curves();
assert_eq!(
curves["b"],
history.filtered_learning_curve("b"),
"the plural form must agree with the singular for the same key"
);
}
#[test]
fn filtered_evidence_is_invariant_to_convergence() {
let mut history = repeated_winner_with(6, tight());
let before = history.filtered_log_evidence();
let report = history.converge().unwrap();
assert!(
report.converged,
"fixture must converge: {:?}",
report.final_step
);
let after = history.filtered_log_evidence();
assert!(
(before - after).abs() < 1e-8,
"filtered evidence moved across converge(): {before} -> {after}. The pass must \
carry its own forward messages; anything reading skill.forward shows exactly \
this drift, because converge() contaminates it with backward information."
);
}
#[test]
fn single_slice_filtered_matches_smoothed() {
let mut history = History::builder().convergence(tight()).build();
history
.add_events([
Event {
time: 1,
teams: smallvec![
Team::with_members([Member::new("a")]),
Team::with_members([Member::new("b")]),
],
outcome: Outcome::winner(0, 2),
},
Event {
time: 1,
teams: smallvec![
Team::with_members([Member::new("c")]),
Team::with_members([Member::new("d")]),
],
outcome: Outcome::winner(0, 2),
},
])
.unwrap();
history.converge().unwrap();
let smoothed = history.learning_curve("a");
let filtered = history.filtered_learning_curve("a");
assert_eq!(smoothed.len(), 1);
assert_eq!(filtered.len(), 1);
assert!(
(smoothed[0].1.mu() - filtered[0].1.mu()).abs() < 1e-8
&& (smoothed[0].1.sigma() - filtered[0].1.sigma()).abs() < 1e-8,
"one slice has no future to propagate back, so filtered and smoothed must \
agree: smoothed mu={} sigma={}, filtered mu={} sigma={}",
smoothed[0].1.mu(),
smoothed[0].1.sigma(),
filtered[0].1.mu(),
filtered[0].1.sigma()
);
}
#[test]
fn filtered_curves_do_not_depend_on_ingestion_order() {
let events = |time: i64, winner: &'static str, loser: &'static str| Event {
time,
teams: smallvec![
Team::with_members([Member::new(winner)]),
Team::with_members([Member::new(loser)]),
],
outcome: Outcome::winner(0, 2),
};
let all = vec![
events(1, "a", "b"),
events(1, "c", "d"),
events(1, "a", "c"),
events(1, "b", "d"),
events(2, "a", "d"),
events(2, "b", "c"),
events(2, "a", "b"),
];
let mut batched = History::builder().convergence(tight()).build();
batched.add_events(all.clone()).unwrap();
batched.converge().unwrap();
let mut incremental = History::builder().convergence(tight()).build();
for event in all {
incremental.add_events([event]).unwrap();
}
incremental.converge().unwrap();
let from_batched = batched.filtered_learning_curve("a");
let from_incremental = incremental.filtered_learning_curve("a");
assert_eq!(from_batched.len(), from_incremental.len());
for ((time_b, gaussian_b), (time_i, gaussian_i)) in
from_batched.iter().zip(from_incremental.iter())
{
assert_eq!(time_b, time_i);
assert!(
(gaussian_b.mu() - gaussian_i.mu()).abs() < 1e-8
&& (gaussian_b.sigma() - gaussian_i.sigma()).abs() < 1e-8,
"at t={time_b}: batched mu={} sigma={}, incremental mu={} sigma={}",
gaussian_b.mu(),
gaussian_b.sigma(),
gaussian_i.mu(),
gaussian_i.sigma()
);
}
}