93 Commits
Author SHA1 Message Date
logaritmisk 61da3aca33 Merge api/typed-errors (#74) 2026-09-10 07:31:31 +02:00
logaritmiskandClaude Opus 5 061c481aad refactor!: typed discriminators for InferenceError
Six of fifteen variants carried a `&'static str` discriminator, about
thirty magic strings between them, and the only thing a caller could do
with one was print it. Four new enums replace them:

    Parameter        13 variants, replacing 9 strings in InvalidParameter
    Shape             4 variants, replacing 10 in MismatchedShape
    OutcomeKind       2 variants, replacing WrongOutcomeKind's three fields
    CompetitorField   2 variants, replacing ConflictingCompetitorConfig's

`InvalidProbability` folds into `InvalidParameter` as
`Parameter::PDraw`. It was a bespoke variant for one scalar while every
other scalar shared `InvalidParameter`, and it omitted the parameter
name — so the same parameter had two mechanisms.

`JointUnavailable { reason: &'static str }` splits into `EmptyHistory`,
`JointRequiresScoredEvents` and `NotPositiveDefinite`. The three are
conditions a caller branches on differently — add events, use
`predict_win_probabilities`, or reconsider the priors — and telling them
apart used to mean string-matching English. One test already proved the
distinction was load-bearing: the blanket conversion mapped the
empty-history case onto the ranked one and `an_empty_history_has_no_joint`
caught it immediately.

`NonFiniteResult` splits into `NonFiniteStep { context, step }` and
`NonFiniteSkill { mu, sigma }`. One `step: (f64, f64)` field was
carrying a sweep step from `converge` and a skill's own moments from a
prediction — two situations in one variant, and a field name that could
only be right for one of them.

`InvalidParameter { name: "beta with point-mass skills" }` becomes
`NoPerformanceVariance`. It was never a parameter out of range: both
values are individually valid and it is their combination that leaves
nothing varying.

Three `Display` impls did not meet the standard the others set, and the
typed data is what makes fixing them possible:

    before  drift variance is invalid: NaN
    after   drift variance must be finite and non-negative (got NaN)

    before  kinds: expected length 3, got 2
    after   the outcome describes a different number of teams than the
            event has: expected 3, got 2

    before  Game::ranked: expected Outcome::Ranked, got Outcome::Scored
    after   expected Outcome::Ranked, got Outcome::Scored; call
            Game::scored for a scored outcome

`Parameter::range()` states each parameter's actual bounds, which no
`&'static str` name could have. `error::message_tests` renders every one
and asserts each is a sentence rather than a label, and that the three
above now carry a range or a next step.

The four internal `MismatchedShape` kinds — `results`, `times`, `kinds`,
and the weights array — collapse to `Shape::Internal`, whose `Display`
says plainly that reaching it is a bug in this crate. They are checks on
`add_events_with_prior`'s own parallel arrays and are unreachable
through the public API; they stay checked rather than becoming
`debug_assert!`s, because release is where this crate's defects hide.

Closes #74.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-10 07:31:31 +02:00
logaritmisk 0b9997354d Merge feat/rating-rule (#53) 2026-09-10 07:17:12 +02:00
logaritmiskandClaude Opus 5 c4194b0051 feat: HistoryBuilder::default_rating_for, a rule instead of a roll call
`register` states configuration for one competitor, which needs the key
set up front. A consumer ingesting an event stream generally does not
have it — and "every layout is static" is a rule, not a list. This makes
it one statement that cannot be forgotten on an ingestion path.

    History::builder()
        .default_rating_for(|key: &&str| {
            key.starts_with("layout_")
                .then(|| StartingPoint::new().drift_scale(0.0))
        })
        .build()

A fifth type parameter, defaulted to `NoRule`, so it costs a caller who
does not use one exactly nothing: `History<String>` still spells out.

Two deviations from #53, both because implementing it exposed something
the issue could not have known.

**A trait, not a bare `Fn` bound.** #53's option 1 was a raw
`R: Fn(&K) -> Option<Rating<T, D>>`. A closure's type cannot be written
down, and the motivating consumer holds its `History` in application
state — so it has to name the type in a struct field, and option 1 makes
that impossible. `RatingRule<K>` is implementable on a named type;
`tests/rating_rule.rs` has the struct-field case that would not have
compiled otherwise. `default_rating_for` still takes a closure for the
common case, via `FnRule`.

**The rule returns a `StartingPoint`, not a `Rating`.** A `Rating` also
carries `beta` and the drift model, which describe the *history* rather
than one competitor — a rule that could vary them would be describing a
different model per competitor. What the create branch actually applies
is the prior and the drift scale, the same pair a `Member` may carry, so
that is what the rule supplies. It also keeps `RatingRule<K>` free of
`T` and `D`: with `Rating<T, D>` in the signature, `drift` and
`time_type` stop compiling after a rule is set, because
`R: RatingRule<K, T, D>` does not imply `R: RatingRule<K, T, D2>`.

**Precedence, which #53 left open: explicit beats the rule, field by
field.** The alternative — `ConflictingCompetitorConfig` — would make a
single exceptional competitor incompatible with having any rule at all.
Two *explicit* declarations that disagree stay an error, because neither
is more specific than the other, and a test pins that they still do.

`key_type` resets the rule to `NoRule`: a `RatingRule<K>` cannot answer
questions about `K2`.

Every test carries a control, and one of them corrected me. I first
asserted that a non-matching competitor's *posterior* was untouched.
It is not, and should not be: alice plays the pinned layout, and what
she learns from beating it depends on how sure the model is about it.
The control is her configuration.

Closes #53.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-10 07:17:12 +02:00
logaritmisk 1629176199 Merge perf/sparse-joint (#52) 2026-09-10 07:08:18 +02:00
logaritmiskandClaude Opus 5 695bb822ef perf!: sparse Cholesky with an AMD ordering for the joint
745 ms -> 1.11 ms on the fixture #52 was opened about.

The joint precision matrix is 0.19% dense at scale and gets sparser as
the history grows. We allocated all n^2 entries — 31 MB at n = 1976,
128 MB at ustat's ~4000 appearances — filled 99.8% of it with zeros, and
ran an O(n^3) factorisation over the whole thing.

Two measurements shaped the fix, and the first killed the plan #52
proposed.

**Ordering alone does nothing to a dense factorisation.** Its inner
loops run over every k whether the entry is zero or not, so a
permutation changes which entries are zero and not how many
multiplications happen. A 700x700 banded matrix at 0.43% density:
30.196 ms in band order, 29.544 ms under a scramble that destroyed the
band. Identical, as the flop count says it must be. #52's step 1 —
"reorder with AMD, keep our own Cholesky, and measure" — could not have
worked, and measuring said so before any of it was written.

**Sparsity and AMD together are worth four orders of magnitude.**
Symbolic factorisation on the n = 1976 fixture, against 2.572e9 dense
flops: sparse in natural order needs 5.597e7 (46x), sparse under AMD
needs 8.656e4 — 29,710x. AMD is worth 646x on top of sparsity and
nothing without it. Natural order fills in badly for exactly the reason
#52 predicted about bandwidth: nnz(L) is 292,437 against A's 7,504,
because a competitor idle from slice 0 to slice 75 links across the
whole matrix.

Measured end to end, factorising through `History::joint`:

    n =  480     215 us   (bench: 9.11 ms -> 167 us, 54x)
    n = 1976    1.112 ms  (was ~745 ms, 670x)
    n = 7800    4.616 ms  (dense would be 1.58e11 flops)

Scaling is near-linear now rather than cubic: 16x the variables costs
21x the time, where dense would cost 4096x.

The factorisation is the up-looking sparse Cholesky of Davis's *Direct
Methods for Sparse Linear Systems*, written here rather than taken from
a crate. The scouting in #52 still holds and got one addition: `feral`
itself pulls `pulp`, so it has the same runtime CPU-dispatch problem
that ruled out `faer` — results could differ between an AVX-512 host and
an AVX2 one, the drift the libm-over-std decision was made to avoid.
`sprs-ldl` is still LGPL and `nalgebra-sparse` still disclaims
fill-reduction in its own docs. Only the ordering is a dependency:
`feral-amd`, two crates, both `#![forbid(unsafe_code)]`.

The matrix is accumulated into a `BTreeMap`, not a hash map: the
iteration order becomes the summation order, and a hash map's varies per
process. `tests/cross_process_determinism.rs` exists because that has
bitten before.

`whiten` returns its result in the permuted order and leaves it there —
a dot product does not care, as long as both operands were permuted the
same way — so `bilinear` is unchanged.

Correctness: the existing analytic goldens are 2x2 and 3x3, too small to
permute or fill in, so they could not have caught a symbolic-pass bug.
`agrees_with_a_dense_reference_on_random_sparse_systems` checks every
bilinear form against a deliberately naive dense factorisation that
shares no code with the thing it is checking, on chain-plus-long-range
matrices up to n = 60.

Closes #52.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-10 07:08:18 +02:00
logaritmisk d36d125e52 Merge infra/bench-variance (#54) 2026-09-10 06:55:25 +02:00
logaritmiskandClaude Opus 5 0801acebd1 ci: measure the runner's own benchmark variance, and fix the joint bench
#54 asks whether benchmark regressions can be gated. The threshold is
the whole problem — too tight and CI goes red on noise, which trains the
reflex to re-run until green; too loose and it never fires — and which
of those is possible depends on a number nobody has measured. This adds
a manually-triggered job that runs one unchanged benchmark ten times and
reports min/median/max/mean and the spread.

`joint_factorise_480_appearances` is the probe: ~9 ms, long enough not
to be dominated by timer overhead, and the measurement this crate most
wants protected — it is the dense factorisation #52 is about replacing.

`benches/joint.rs` did not run at all. Its fixture asked for
`epsilon: 1e-10` within `max_iter: 30` and never got there, so once
`converge` stopped returning short fits silently it panicked:

    NotConverged { iterations: 30, final_step: (4.5e-4, 0.0), epsilon: 1e-10 }

It now uses the default `ITERATIONS` cap. Measuring a factorisation on
an unconverged fit would have been measuring something nobody runs. The
other four benchmarks were checked and are fine.

Two things in the report step were got wrong first and fixed by running
them, not by reading them:

- `asort` is a gawk extension and the runner's `awk` is mawk. Sorting
  goes through `sort -n` instead.
- Criterion picks a unit per run, so a mixed batch would compare 9 ms
  against 9 us as though they were the same number. The job refuses to
  report a spread unless every run agrees on the unit.

Refs #54.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-10 06:55:25 +02:00
logaritmisk 1ad789cf40 Merge api/param-reorder (#72) 2026-09-10 06:49:23 +02:00
logaritmiskandClaude Opus 5 b553c630f5 refactor!: K comes first in History, HistoryBuilder and Joint
`K` is the one type parameter people change, and it was last. Naming a
history in a struct field meant writing all four to say one thing:

    struct Ladder { history: History<i64, ConstantDrift, NullObserver, String> }
    struct Analysis<'h> { joint: Joint<'h, i64, ConstantDrift, NullObserver, &'static str> }

Now:

    struct Ladder { history: History<String> }
    struct Analysis<'h> { joint: Joint<'h> }

`History<K, T, D, O>`, all four defaulted. Bounds may reference later
parameters, so `D: Drift<T> = ConstantDrift` is legal in third position.
`Joint` gains the same defaults, so `Joint<'h, String>` spells it.

72 call sites swapped, and the reorder makes most of them shorter: 18
now read `History<String>` and the `&'static str` ones read `History`.
The two turbofished builders shrink from
`HistoryBuilder::<Untimed, _, _, String>::new()` to
`HistoryBuilder::<String, Untimed>::new()`.

`Joint` keeps `O` structurally, defaulted rather than removed. #72 notes
it never touches the observer, which is true — but it borrows the whole
`&'h History<K, T, D, O>` and calls `History::resolve_terms`, so dropping
the parameter means either a view type or moving that method off
`History`. The default already buys the entire user-visible benefit,
which was the spelling.

Refs #72.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-10 06:49:22 +02:00
logaritmisk d2ab4446ef Merge api/joint-layering (#78) 2026-09-10 06:40:31 +02:00
logaritmiskandClaude Opus 5 e72bf3894c refactor!: the joint is reached through Joint, not mirrored on History
`posterior_of`, `posterior_of_at` and `expected_variance_reduction`
existed twice: once on `Joint`, and once on `History` as one-shot
wrappers whose whole body was `self.joint()?.<same>(..)`.

The wrappers re-factorised on every call — their own docs said so,
warning the reader to take a `Joint` instead — and they were what
smuggled the scored-only precondition onto the flat surface. A user
following the quickstart builds a ranked history, sees `posterior_of` in
the method list, and it never works. `h.joint()?.posterior_of(..)` is
one call longer and tells the truth: you need a joint, and a joint needs
a scored history.

That leaves three tiers instead of a flat surface with a hidden
precondition: `History` fits and reads, `predict_*` forecasts, `Joint`
answers exact joint questions.

`predict_margin` was itself calling `self.posterior_of`; it goes through
`self.joint()?` directly now.

The `Joint` methods' docs referred back to the wrappers for their real
content ("Identical to `History::posterior_of`, without re-paying the
factorisation"), so they now carry it: what a linear functional means,
which appearance each competitor is read at, and why
`expected_variance_reduction` belongs on the handle.

`tests/joint_handle.rs` had three tests comparing the wrapper against
the handle. That comparison is gone, but the property behind it is not —
they now compare a *reused* joint against a *fresh* one per question,
which is the actual correctness claim behind caching the factorisation
(#51), without the wrapper in the middle.

Closes #78.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-10 06:40:31 +02:00
logaritmisk 56193609f7 Merge api/renames (#75, #78) 2026-09-09 23:23:12 +02:00
logaritmiskandClaude Opus 5 13a395fdc9 refactor!: scores_with_noise, and History::quality
Two names that described the wrong thing.

`scores_with_sigma(scores, sigma)` reads as "these scores have prior
sigma 2.0". The quantity is observation noise on the score *margin*, in
the units of the scores, and it is spelled `score_sigma` at every config
site — `HistoryBuilder::score_sigma`, `GameOptions::score_sigma`,
`EventKind::Scored { score_sigma }` — so this was the one place the
crate used a third meaning of "sigma" for it. Its own doc had to
disambiguate itself: "`sigma` overrides `HistoryBuilder::score_sigma`".
`scores_with_noise(scores, score_sigma)` on both `Outcome` and
`EventBuilder`.

`predict_quality` predicts nothing. Its own doc says it answers "is this
matchup *fair*", not "what will happen", and the `predict_*` family is
otherwise exactly the methods returning a probability or a distribution
over outcomes. `History::quality` also makes the free/method pair
consistent: free `quality` pairs with `History::quality` the way free
`expected_information_gain` already pairs with
`History::expected_information_gain`. The rule that was already being
followed and never stated — a free function scores a hypothetical from
explicit parameters, the same-named method asks it against the fit — is
now written on the method.

Closes #75. Refs #78 (part 4).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 23:23:12 +02:00
logaritmisk 6e2ce69728 Merge api/retire-index (#73) 2026-09-09 23:18:54 +02:00
logaritmiskandClaude Opus 5 faa25fb3b1 refactor!: retire Index, intern and lookup
`Index` was public, `History::intern` and `History::lookup` returned
one, and no public method anywhere accepted one. It was a handle with
nowhere to go — and `key_table.rs` advertised the hot-path story it was
meant to enable ("power users can promote `&K` to `Index` and skip the
lookup"), which was never reachable through the public API.

It also shadowed `std::ops::Index`, which `CompetitorStore` implements,
so `use trueskill_tt::*` alongside `use std::ops::*` collided.

All three are `pub(crate)` now. `intern` stays internal because
ingestion needs it; `lookup` is gone entirely, since `current_skill`,
`rating` and `learning_curve` already answer "does this history know
this key" and all three take a borrowed key.

The three tests that used them asserted things a caller cannot observe.
They now assert what the interning bought:

- `record_winner_creates_two_competitors` compares posteriors instead of
  comparing two opaque indices for inequality.
- `intern_is_idempotent` becomes `a_repeated_key_is_one_competitor` — a
  key appearing in two events gives one competitor with a two-point
  learning curve, which is the observable form of the same claim.
- `lookup_returns_none_for_missing` becomes `an_unknown_key_is_unknown`.

Closes #73.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 23:18:54 +02:00
logaritmisk ddbac87744 Merge api/gaussian-operators (#71) 2026-09-09 23:13:10 +02:00
logaritmiskandClaude Opus 5 076a7ded8c feat!: Gaussian's EP operations stop wearing arithmetic's clothes
`Gaussian` publicly implemented `Mul`, `Div`, `Add` and `Sub`. They were
the EP product, cavity and variance-space convolutions, and every one of
them lies to a reader who takes the operator at face value:

    a = N(10, 2)   b = N(4, 3)   c = N(1, 1)

    a * b        N(8.15, 1.66)   not 40
    a - b        sigma GREW, 2 -> sqrt(4 + 9)
    a * N(1, 0)  mu = NaN        "multiply by one"
    a / c        pi = -0.75      mu() prints a confident 0

The last is this crate's signature defect on a public operator. `Div` is
the cavity and can legitimately leave a negative precision, which is not
a distribution — and `mu()`/`sigma()` guard `pi <= 0` and report `0.0`
and `inf`, so it comes back as a plausible number with no panic, no
`Debug` marker and nothing to test against.

The four impls are now `pub(crate)` inherent methods that say what they
do: `ep_product`, `cavity`, `convolve`, `convolve_diff`, plus `scale`
for the one operation that genuinely is arithmetic. Nothing in a user's
workflow needed operator syntax; inference did, and it still has it.

`pi()` and `tau()` follow. Storing natural parameters is a performance
decision — it makes message passing two adds — not a contract. The
public surface is now exactly: `from_ms`, `from_mv`, `mu`, `sigma`,
`variance`, `probability_below`, `probability_above`. `from_mv` and
`variance` are promoted from `pub(crate)`; they are the honest pair for
callers who already hold a variance and should not pay a round trip
through the square root.

Four integration tests asserted bit-identity on `(pi, tau)`. They assert
it on `(mu, variance)` instead — still `assert_eq!`, still exact, and
`1/pi` and `tau/pi` are deterministic, so bit-equal natural parameters
give bit-equal moments. `a_nan_sigma_passes_through_from_ms` drops its
`|| g.pi().is_nan()` half: `sigma()` substitutes for `pi <= 0` and
`pi == inf`, so NaN survives to it only from a NaN precision.

`benches/gaussian.rs` is deleted. It timed two f64 additions through the
public operators, and keeping those public solely to feed it is the same
thing #73 objected to when a benchmark was dictating five public types.
The paths it covered are exercised by `batch` and `history_converge`
through the real call chain.

Closes #71.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 23:13:10 +02:00
logaritmisk 8e34410db0 Merge api/game-rename (#69) 2026-09-09 22:55:05 +02:00
logaritmiskandClaude Opus 5 92d690d0f8 feat!: Game is the type you get, and one_v_one returns one
`lib.rs` advertised `Game` in "Core types" as "one match in isolation".
It had no public constructor: every `Game::*` returned `OwnedGame`, so
`let g: Game = Game::ranked(..)?` did not compile.

Names swapped. The public type is the owned one — `Game<T, D>`, no
lifetime — and the borrowing form is `pub(crate) GameRef<'a, T, D>`,
which is what it always was: an implementation detail about whether the
result and weight slices are borrowed from `History`'s storage. That
distinction meant nothing to someone scoring one match, and it showed
the module's surface twice in rustdoc, since both types carried
`posteriors()` / `log_evidence()`.

`one_v_one` returned `(Gaussian, Gaussian)` while every sibling returned
a game, making it the one constructor you could not ask for
`log_evidence()`. It returns `Self` now; `.posteriors()` recovers the old
shape, and the test that covers it now also asserts the evidence of two
identical ratings is exactly `ln(0.5)`.

`ranked`, `scored` and `free_for_all` had `# Errors` as their entire
doc, so rustdoc's index rendered the error list as the summary. They
have summary lines, and `Game` has a worked example — it was advertised
as a core type with none anywhere in the crate.

Closes #69.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 22:55:05 +02:00
logaritmisk 4c2e98c54b Merge api/ergonomics (#72) 2026-09-09 22:11:25 +02:00
logaritmiskandClaude Opus 5 92ae5fca17 feat!: prediction and joint queries take borrowed keys
`&[&[&K]]` was the worst shape in the API. At `K = String` — the
realistic case, where names arrive owned from a database or CSV — a
string literal was *impossible*, and asking "who wins" cost six lines
and four allocations of temporaries that all had to outlive the call:

    let ta = vec![a.to_string()];
    let ra: Vec<&String> = ta.iter().collect();
    ...
    self.history.predict_win_probabilities(&teams)

All seven `predict_*` / `expected_*` methods, `posterior_of`,
`posterior_of_at` and the `Joint` mirrors are now generic over the
borrowed key, the same way `current_skill` and `learning_curve` already
were. `member_skills` and `resolve_terms` only ever did two things with
a key — `keys.get` and `format!("{key:?}")` — and neither needed `K`.

    h.predict_win_probabilities(&[&["alice"], &["bob"]])   // K = String
    h.predict_win_probabilities(&[&["alice"], &["bob"]])   // K = &'static str
    h.posterior_of(&[("alice", 1.0), ("bob", -1.0)])

One spelling for both key types, and `K: Debug` becomes `Q: Debug`, so a
key type no longer has to be `Debug` to run a prediction. The old
`&[&[&"a"]]` spelling still compiles at the default key type, where `Q`
infers to `&str` and the two shapes coincide.

The one cost: `predict_outcome(&[])` can no longer infer `Q` — nothing
in an empty slice names it. It needs an annotation, and only on that
degenerate call.

`lookup` carried `ToOwned<Owned = K>`, copy-pasted from `intern`, which
genuinely needs it to create the entry. `lookup` never creates, and its
five neighbours all accept `h.f("alice")` already. Dropping the bound
strictly widens what compiles.

`HistoryBuilder::gamma` is shorthand for `.drift(ConstantDrift::new(g))`.
Drift is the most-tuned parameter after `sigma` and `GAMMA` is a public
constant, but setting it meant first discovering `ConstantDrift`, a type
a caller has no other reason to name. On the `ConstantDrift` builder
only, since `gamma` is that model's parameter rather than something
every `Drift` has, and rejecting a negative value for the same reason as
`sigma` and `beta`: it enters squared.

Refs #72 (items 2 and 3, plus the gamma shorthand; the type-parameter
reorder in item 1 is still open).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 22:11:25 +02:00
logaritmisk da55d2a7d1 Merge api/non-exhaustive-options (#74) 2026-09-09 22:05:11 +02:00
logaritmiskandClaude Opus 5 6a893ffe57 fix!: non_exhaustive on ConvergenceReport, and not on the options structs
`ConvergenceReport` is only ever constructed by `converge` /
`converge_partial`, so marking it costs a caller nothing and makes a
future field additive.

`ConvergenceOptions` and `GameOptions` deliberately stay constructible,
against #74's recommendation, because trying it turned up a cost the
issue did not anticipate. `Default::default` is not a `const fn`, so
`#[non_exhaustive]` + `..Default::default()` — the pattern that makes
marking an options struct cheap — does not work in a `const`:

    error[E0639]: cannot create non-exhaustive struct using struct expression
      --> tests/competitor_config.rs:13:41
       |
    13 | const CONVERGENCE: ConvergenceOptions = ConvergenceOptions {

`ConvergenceOptions` is `Copy` and a natural const; there is no
workaround from outside the crate. That cost is permanent, and adding a
field is a one-time major bump. The reasoning is recorded on the type so
the next person does not rediscover it.

Refs #74.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 22:05:11 +02:00
logaritmisk f14c783c0e Merge docs/vocabulary (#75) 2026-09-09 21:58:26 +02:00
logaritmiskandClaude Opus 5 055575a6f4 docs!: one name for score noise, and say which of beta/sigma to turn
"sigma" named three unrelated quantities: the prior standard deviation,
a distribution's own SD, and the observation noise on an observed score
margin. The third was already `score_sigma` at every config site —
`HistoryBuilder::score_sigma`, `GameOptions::score_sigma`,
`EventKind::Scored { score_sigma }` — and plain `sigma` only on
`Outcome::Scored`'s field and constructor parameter, whose own doc had
to disambiguate itself with "`sigma` overrides
`HistoryBuilder::score_sigma`". Now `score_sigma` everywhere.

The `Outcome::scores_with_sigma` / `EventBuilder::scores_with_sigma`
*method* names are left alone: renaming them is a naming choice rather
than a consistency fix, and #75 offers two candidates.

`HistoryBuilder::beta` and `::sigma` now say which is which. #75 calls
this the single most load-bearing undocumented distinction in the crate,
and it is right: nothing told a reader that `sigma` is epistemic — what
the model does not yet know, which evidence shrinks — while `beta` is
aleatoric, the day-to-day scatter no amount of evidence removes. Both
docs now name the symptom that should send you to that knob rather than
the other.

Refs #75.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 21:58:26 +02:00
logaritmisk 251211f134 Merge docs/missing-docs (#77, #75) 2026-09-09 21:53:51 +02:00
logaritmiskandClaude Opus 5 31564b71a0 docs: document the whole public surface and deny(missing_docs)
80 undocumented public items, including three that are first contact:
`History::current_skill` — the method the crate's own first example calls
— `EventBuilder`, the type `h.event(t)` hands you, and `Gaussian::mu()`.
Now zero, and `#![deny(missing_docs)]` keeps it that way.

Several docs are measurements rather than readings of the code:

- `Outcome::Ranked` says ranks are used ordinally, so `[0, 1, 2]` and
  `[0, 5, 90]` are the same observation. Measured: bit-identical
  posteriors for both.
- `OwnedGame::log_evidence` says two identically-rated competitors give
  exactly `ln(0.5)`. Written as a doctest, so it runs.
- `Member::weight` says zero and negative are accepted. Measured.
- `ConvergenceReport::final_step` is `(|Δmu|, |Δsigma|)` in skill units,
  NOT natural parameters. That one had to be traced through
  `Gaussian::delta` rather than assumed from the neighbouring vocabulary.
- `GameOptions::score_sigma` rejects non-positive and NaN but accepts
  `+inf`, which is what the guard actually says.

README: it is the front door for a crate on a private registry, and it
opened with a link dump followed by 130 lines on drift. The first
`record_winner → converge → current_skill` block was at line 226 of 307.
It now leads with what the crate is, an install line, a quickstart, a
"which entry point?" table, and the `converge`-is-strict rationale that
was the crate's most opinionated recent decision and went unmentioned.
The two canonical examples disagreed on spelling (`History::default()`
vs `History::builder().build()`, `current_skill("a")` vs
`current_skill(&"a")`); they now agree. Five new README blocks are
doctested, taking the suite from 19 to 25.

`pub use smallvec;`. Four public items name `SmallVec` in their
signatures, and the only `Joint` example failed to compile from a
consumer crate with `unresolved import smallvec` — the dependency was in
the API but not reachable. Both worked examples now use the re-export,
so they teach the path that works downstream.

Vocabulary, from #75: "agent" was a fourth word for competitor, 200
occurrences, and it had reached public signatures before #73 un-exported
`TimeSlice`. Now zero.

Closes #77. Refs #75.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 21:53:50 +02:00
logaritmisk 78810c0344 Merge fix/predict-non-finite-guard (#78 parts 1-2) 2026-09-09 21:40:39 +02:00
logaritmiskandClaude Opus 5 9d3e002be3 fix!: no prediction path answers from a fit it cannot answer from
`converge` refuses to report a NaN fit. Nothing stopped a caller from
ignoring that error and predicting anyway, and every prediction path was
differently wrong when they did. Measured on a point-mass-prior history
with `beta(0.0)`, after `converge` returned `NonFiniteResult`:

    predict_quality           = Ok(NaN)
    predict_outcome().total() = NaN
    predict_win_probabilities = Ok([0.0, 0.0])

The third is the dangerous one: finite, plausible, and summing to zero
against a doc that promises one at `p_draw == 0`. A caller checking
`total() ≈ 1` catches the second and misses it.

The same parameters on a *scored* event converge cleanly and leave
legitimate point-mass posteriors. There `predict_quality` **panicked** —
"cannot invert a singular matrix", out of a method returning `Result` —
because the contrast covariance `beta²AᵀA + AᵀSA` is exactly singular,
and `predict_win_probabilities` again returned `Ok([0.0, 0.0])`. That
promise assumes continuous performances, where an exact tie has measure
zero; point masses break the assumption, not the arithmetic.

Both checks now live at `member_skills`, the one gate every prediction
path reads skills through, rather than being repeated per method.

The finiteness check is on `mu` / `sigma`, not on the natural parameters.
The first attempt checked `pi` and `tau`, and measurement showed it
rejected a *legitimate* point mass — `pi = inf`, `mu = 0`, `sigma = 0` —
turning a working prediction into an error. The question is whether the
usable moments exist, and those are what predictions consume.

Docs: `converge_partial` omitted the drift-variance `InvalidParameter` it
validates before sweeping, and the free `expected_information_gain`
omitted `GridTooCoarse`, which comes from `outcome_distribution` and so
is not covered by its "anything `Game::ranked` returns" clause.

Refs #78 (parts 1 and 2; the layering and `predict_quality` rename
questions are still open).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 21:40:39 +02:00
logaritmisk 9d629d0d94 Merge api/trait-consistency (#76) 2026-09-09 21:29:16 +02:00
logaritmiskandClaude Opus 5 7ca0daa48e feat: PartialEq on the config types, and pin the public trait impls
`Rating` already derived `PartialEq`, but that derive is only reachable
through `D: PartialEq` — and `ConstantDrift`, the crate's own only
`Drift` impl, did not satisfy it. So the derive was there and unusable.
Found by writing the comparison from a consumer's position rather than
reading the derive list.

`ConstantDrift`, `ConvergenceOptions` and `GameOptions` now derive
`PartialEq`. All three are pure configuration; comparing two is the
natural thing to want and nothing about them makes equality ambiguous.

`tests/trait_impls.rs` pins the surface, written the way the failure was
reported: a consumer struct that *holds* a `History` and derives
`Debug`. It also asserts `History`'s `Debug` summarises rather than
dumping its skill stores, so a future derive cannot quietly replace the
hand-written impl.

`Clone` on `History` stays off. It is a decision, not an omission: a
history owns every slice's skill store and arena, so cloning one is
proportional to the whole fit, and no consumer has wanted it.

Closes #76.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 21:29:16 +02:00
logaritmisk c3d1afe448 Merge api/must-use-and-visibility (#67, #73) 2026-09-09 21:24:37 +02:00
logaritmiskandClaude Opus 5 e4d6dc4028 fix: warn on dropped builders and values; stop exporting EP internals
`h.event(1).team(["x"]).team(["y"]).ranking([0, 1]);` without the
terminal `.commit()` was a silent no-op: no warning, no error, and the
next thing the caller does is converge an empty history and read `None`
skills. `EventBuilder` already carried a `#[must_use]`; the value types
around it did not, so the same silence covered `Team::with_members`,
`Member::new`, `Outcome::*`, `Joint` and `Prediction::outcomes`.

`#[must_use]` now goes on the *types* rather than being sprinkled over
methods, which covers every constructor and builder setter at once and
gives the crate a rule where it previously had a list. Verified by
compiling a program that drops each one and reading the warnings back,
rather than by assuming the attribute took.

Visibility, from #73: `Gaussian::damp_natural` was reachable from
outside the crate despite being an EP damping internal called only from
`src/factor/`. The stray `pub fn`s inside the private `time_slice`,
`key_table` and `matrix` modules are now `pub(crate)`, so their
visibility states what it means instead of relying on the module being
private.

`storage/mod.rs` and `factor/mod.rs` become `storage.rs` and
`factor.rs`.

Closes #67. Refs #73.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 21:24:37 +02:00
logaritmisk cc601c06eb Merge feat/evidence-matrix: the missing evidence corner (#70) 2026-09-09 21:19:20 +02:00
logaritmiskandClaude Opus 5 86e1521f8a feat: complete the evidence matrix and add current_skills
Three of four corners of the evidence matrix existed. The missing one was
forward-only *and* key-restricted — which is exactly what per-competitor
prequential scoring needs, the intersection of the two workloads
`log_evidence_for` and `filtered_log_evidence` are each documented for.

`filtered_log_evidence_for` fills it. It is not
`log_evidence_internal(true, targets)`: that path selects `skill.forward`
as the prior, which stops being a filtering quantity once `iteration`
has run a backward sweep. It goes through `filtered_pass` like its
unrestricted sibling, with the restriction applied to which events are
*scored*, never to which are *run* — so it is a held-out score under the
real history, not a score under a counterfactual one where nobody else
played.

Key resolution for both `*_for` accessors now shares `resolve_targets`,
so they cannot drift apart on how an unknown key is reported.

`current_skills` is the plural of `current_skill`. Building a
leaderboard previously meant materialising every competitor's full
smoothed curve via `learning_curves` and reading the last point of each.

Tests carry controls in both directions: naming every competitor must
recover the unrestricted value (catching a filter that drops too much),
and the restricted forward-only value must differ from the restricted
smoothed one (catching an alias).

Refs #70.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 21:19:20 +02:00
logaritmisk 60fc3e9d05 Merge fix/honest-accessors: honest per-key queries (#66, #70) 2026-09-09 21:13:36 +02:00
logaritmiskandClaude Opus 5 e4a68ba1a7 fix!: per-key queries report unknown keys instead of a plausible constant
Two accessors answered a question about a key the history had never seen
with a well-formed value indistinguishable from a real answer.

`log_evidence_for` filter_map'd unknown keys away. An empty target list
means "no restriction" downstream, so a list of *entirely* unknown keys
returned the whole-history evidence: measured on a two-cohort fixture,
`log_evidence_for(["typo"])` returned exactly `log_evidence()`. On the
one workload it is documented for — leave-one-out cross-validation —
that is the un-held-out score, a plausible number that silently
invalidates the comparison it was computed for. It now returns
`Err(UnknownKey)` naming the offending position.

`learning_curve` and `filtered_learning_curve` returned an empty `Vec`
both for a typo'd key and for a competitor who is registered but has not
played yet. They now return `Option`, so `None` is "never heard of it"
and `Some(vec![])` is "known, no appearances".

Tests carry a control case in each direction, so they cannot pass by
everything returning the same thing.

Closes #66, closes #70.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 21:13:35 +02:00
logaritmiskandClaude Opus 5 56e8220c86 Merge branch 'api/cleanup'
Un-export the unreachable types, add the missing trait impls, make
#[must_use] consistent, correct eight wrong # Errors sections, seal the
error variants, and settle the vocabulary.

Refs #70, #73, #74, #75, #76, #77, #78

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 20:54:48 +02:00
logaritmiskandClaude Opus 5 fdd1539cab refactor: one word per concept
Three vocabulary collisions, from #75.

**"rating" meant three things**, one the opposite of the exported type.
`Rating` is documented as static *configuration* — "this returns what it
was told", against every other accessor's "what inference inferred". But
`quality`'s parameter was `rating_groups: &[&[Gaussian]]` and its prose
said "rating groups" four times, where "rating" means a *posterior* — the
one thing `Rating` is documented not to be. Two error messages used it
that way too.

So a reader who learned `Rating = config` passed `Rating` values to
`quality`, which takes `Gaussian`; and one who learned "rating = what
comes out" was baffled that `h.rating(&k)` is not their skill.

"rating" is now reserved for the type. `quality(teams: &[&[Gaussian]])`,
and "every rating is finite" became "every posterior is finite".

**"agent" was a private fourth name for a competitor** — ~200 identifiers
against 236 uses of "competitor", and it leaked into two `pub` signatures
on `TimeSlice`. Now that #73 has made those internal this is a pure
rename, so the crate has one word for the entity throughout.

**"player" survived in one public signature** — `free_for_all(players:)`
plus two doc lines. Renamed, along with three internal closure bindings.
Doc examples that use "player" as a *key* are left alone: that is a
user's data, not the crate's vocabulary.

The panic-message expectations in tests/quality.rs moved with the prose,
which is the point of asserting on message text — the tests caught the
rename rather than papering over it.

Not touched: "performance" (always skill widened by beta), "skill",
"member" and "team" are each used for exactly one thing already.

Refs #75

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 20:54:47 +02:00
logaritmiskandClaude Opus 5 85c4d0d87d fix!: correct eight wrong # Errors sections and seal the error variants
Documentation (#78). Every item below was measured against the code
rather than read:

- `expected_information_gain` and `predict_ranking` had `# Errors`
  immediately followed by `# Preconditions`, with the error list stranded
  at the bottom of the latter — rustdoc rendered a BLANK Errors section on
  both. The heading now sits with its content.
- `predict_outcome`, `predict_ranking` and the free
  `expected_information_gain` all omitted `GridTooCoarse`.
- `predict_margin` claimed `JointUnavailable` "if the LATEST slice holds
  ranked events". Measured with an early ranked slice and a late scored
  one: it fails. The condition is *any* slice.
- `add_events` documented three errors and can return five more; it also
  claimed a weights `MismatchedShape` that is unreachable through it,
  since weights arrive one-per-`Member`. That check belongs to
  `EventBuilder::weights`, and the doc now says so.
- `converge` and `converge_partial` both omitted the drift-variance
  `InvalidParameter`.

`History` gains a hand-written `Debug` (#76). Summarising, not
exhaustive — a derived one would print every competitor's skill at every
slice. It exists because without it a consumer cannot `#[derive(Debug)]`
on any struct holding a `History`, which is how both known consumers
store one.

`#[non_exhaustive]` on all 17 `InferenceError` struct variants and on
`Outcome::Scored` (#74). The enum carried the attribute; no variant did,
so adding a field to any of them — and downstream construction of any of
them — were both in the public contract. This crate added two variants in
two days.

The options structs are deliberately NOT sealed. `ConvergenceOptions` and
`GameOptions` are constructed by struct literal at 65 sites of which only
8 use `..default()`, and specifying all three convergence fields is a
natural complete statement rather than a partial one. That is a real
trade-off rather than an oversight, and it is left as a decision on #74.

Also spells `UnknownKeys::Reject` explicitly at both sites that
wildcarded it. `#[non_exhaustive]` on your own enum gives no exhaustiveness
safety net if you then match `_`.

Sealing the variants pushed ten test sites from constructing errors to
`matches!`, which is the better assertion anyway — an `assert_eq!` against
a constructed error breaks whenever a field is added, which is the exact
fragility the attribute exists to prevent.

BREAKING CHANGE: `InferenceError`'s struct variants and `Outcome::Scored`
are `#[non_exhaustive]` — downstream patterns need `..` and downstream
construction is no longer possible.

Refs #78, #76, #74

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 20:49:07 +02:00
logaritmiskandClaude Opus 5 a0c2f78aed feat: add the missing trait impls and make #[must_use] consistent
Trait coverage (#76), all additive:

  History          Debug is still absent - see below
  HistoryBuilder   + Debug   (it derived Clone but not Debug)
  Rating           + PartialEq  (Gaussian had it; Rating is a Gaussian
                                 plus three scalars and had none)
  Event/Team/Member + PartialEq (input value types with no way to compare
                                 them, which made round-trip tests awkward)
  ConvergenceReport + PartialEq

`#[must_use]` (#67). The coverage had no rule: `filtered_log_evidence`
had it and `log_evidence` did not; `rating` had it and `current_skill`
did not; `Rating::with_drift_scale` had it and `Member::with_drift_scale`
did not.

Now on the types — `EventBuilder`, `HistoryBuilder`, `Prediction`,
`Gaussian`, `OwnedGame` — which covers most method returns at once, plus
the `History` accessors individually.

`EventBuilder` gets a message, because a dropped builder is the worst
case in the set: measured, `h.event(1).team(["x"]).team(["y"]).winner(0)`
without `.commit()` leaves `time_slices_len() == 0` and every skill
`None`, with no warning at all.

And `ConvergenceReport`'s `#[must_use]` moves off the TYPE onto
`converge_partial`, where its stated reason is true. It read "from
`converge_partial` this may describe a fit that stopped at max_iter" but
fired on `converge` too — where that is false, since `converge` returns
`Err(NotConverged)` in exactly that case. So the crate's own front-page
example warned, and every quickstart had to write `let _ =`. Verified
from a consumer crate: `h.converge()?;` now compiles clean.

Marking the types made eight method-level attributes redundant, which
clippy's `double_must_use` caught — that is the type-level marker doing
its job, and the eight are removed.

Refs #76, #67

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 20:38:39 +02:00
logaritmiskandClaude Opus 5 4472d98b56 refactor!: un-export six types that no caller could reach
`TimeSlice`, `EventKind`, `KeyTable`, `CompetitorStore`, `Competitor` and
the `storage` module were all public and none was obtainable from a
`History` — `time_slices`, `agents` and `keys` are all private or
`pub(crate)`. `TimeSlice` was the worst: `new`, `add_events`, `iteration`,
`get_composition` and `get_results` were `pub` on a type you could only
build standalone and never feed back into anything.

Their sole consumer outside `src/` was `benches/batch.rs`, so a benchmark
was dictating six public types. It is rewritten against the public API: a
single-slice history's `converge` calls exactly the same per-slice sweep,
so capping at one iteration measures the same code path.

`N01` had zero references in the entire repository, including inside the
crate; removed. `N00` and `N_INF` are EP identities (`Add` and `Mul`) and
are now `pub(crate)` — a user reaching for `N_INF` as "an unknown
competitor's prior" would get an improper distribution whose `mu()`
silently reports 0.0.

Adds the accessors their absence forced people around, from #70:
`competitors()`, `competitor_count()` and `event_count()` (`size` had no
accessor at all). Answering "who is best" previously meant materialising
every competitor's full smoothed curve to read the last point of each.

`KeyTable::keys` now iterates the dense reverse table rather than the
forward `HashMap`, so `competitors()` yields insertion order rather than
per-process hash order — the same hazard as #62, caught before it could
reach a caller building a standings table.

Two `CompetitorStore` methods (`is_empty`, `iter_mut`) had no callers
anywhere and are gone; four more are now `#[cfg(test)]`, which is what
they always were in practice.

Worth recording a mistake: I first deleted `get_composition`/`get_results`
on the strength of a "never used" warning, and the build broke — the
warning came from the plain-lib target, where `#[cfg(test)]` callers in
history.rs are not compiled. A dead-code warning from one target is not
evidence about the others.

BREAKING CHANGE: `TimeSlice`, `EventKind`, `KeyTable`, `CompetitorStore`,
`Competitor`, the `storage` module, `N01`, `N00` and `N_INF` are no longer
public.

Refs #73, #70

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 20:34:18 +02:00
logaritmiskandClaude Opus 5 5f5a37090a Merge branch 'fix/reachable-time'
Make the Time generic reachable, and exercise it end to end.

Closes #68

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 20:19:16 +02:00
logaritmiskandClaude Opus 5 dc1f4d5847 fix!: make the Time generic reachable
`History<T: Time, ..>` has always been generic over the time axis,
`Untimed` has always been exported, and `Drift<T>` is generic specifically
so that "seasonal or calendar-aware drift is expressible without going
through i64". None of it was reachable from a downstream crate.

Every construction route pinned `T = i64`: `History::builder()`,
`History::builder_with_key()`, and the only `Default` impl on
`HistoryBuilder`. Its fields are private and it had no `new`. So all three
escape routes failed to compile, and a consumer with domain timestamps
had to convert to i64 — which is the exact thing the parameter exists to
avoid. One of `History`'s four type parameters was paid for at every
signature and could never be varied.

`Default` is now generic over `T` and `K`, `HistoryBuilder::new()` exists,
and `time_type::<T2>()` / `key_type::<K2>()` join `drift` and `observer`
as type-changing setters:

    History::builder().time_type::<Untimed>().build()
    History::builder().key_type::<String>().build()
    HistoryBuilder::<Season, _, _, String>::new().build()

`key_type` replaces `builder_with_key`, which could not be turbofished —
`K` sat on the impl rather than the function, so callers had to spell
`History::<i64, _, _, String>::builder_with_key()`. 18 call sites across
15 files migrated.

tests/time_axis.rs is the part that matters. NOTHING in the repository
constructed a non-i64 history, which is precisely why this survived, so
the fix is only half done without a test that exercises the generic. It
defines a `Season(u16)` time type and a `SeasonalDrift` that accumulates
between seasons but not within one — the calendar-aware case the trait's
docs cite — and checks the whole path: fit, converge, and read a learning
curve whose times come back as `Season`, not as integers.

Two of the six tests are controls rather than assertions about output.
`Untimed` must ignore drift entirely, since elapsed is always zero, so
gamma 0.0 and gamma 5.0 must agree bit for bit. And a custom `Drift` must
actually widen a gap across seasons, or the test above would pass whether
or not the drift was consulted at all.

The README's ticked "Generalise a time axis" box is now true.

BREAKING CHANGE: `History::builder_with_key()` is removed. Use
`History::builder().key_type::<K>()`.

Closes #68

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 20:19:16 +02:00
logaritmiskandClaude Opus 5 0ab56248bb Merge branch 'fix/seal-constant-drift'
Seal ConstantDrift's field, and add an enumerating test over every public
magnitude parameter.

Closes #65

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 19:11:57 +02:00
logaritmiskandClaude Opus 5 8dff7513f7 fix!: seal ConstantDrift's field so gamma can be validated
`gamma` enters only as `gamma * gamma`, so the sign was squared away:
measured against the old public-field form, `ConstantDrift(-0.0833)`
produced results bit identical to `ConstantDrift(0.0833)`. The sign was
neither rejected nor honoured — it vanished.

It could not be checked while the field was a public tuple position,
because there was nothing to intercept. Validating inside
`variance_for_elapsed` would have been worse: it runs in the sweep, so a
construction-time mistake would panic mid-inference, and `Gaussian::from_ms`
is a worked example of why that is the wrong place — rejecting NaN there
turned the NonFiniteResult reporting path into a crash.

So `ConstantDrift::new` is the only way in and it checks, with `gamma()`
to read the value back. 129 call sites rewritten across src, tests,
benches, examples and the README. The dated plan and spec documents under
docs/superpowers are left alone: they record what was built at the time,
and rewriting them would falsify that.

tests/constructor_validation.rs is the more valuable half. This defect
class was closed three times in one session and reopened twice, because
each fix validated the layer it had just touched and inferred the rest —
`HistoryBuilder`, then `Game`'s own entry points, then the constructors
beneath both. A per-site fix cannot notice the site nobody thought of, so
that file enumerates every public entry point taking a magnitude and
asserts each refuses negative and non-finite values.

It found an eleventh defect on its first run: `HistoryBuilder::score_sigma`
accepted infinity, because `inf > 0.0` is true and the assert only tested
positivity. Fixed, and its own `should_panic` message updated to match.

`Gaussian::from_ms` is deliberately exempt from the non-finite half, for
the reason above: a broken fit produces a NaN sigma legitimately and
`converge` must be allowed to report it.

The convergence-level drift-variance check stays and is now tested through
a custom `Drift` implementation, since `ConstantDrift` can no longer reach
it. That check is the only thing standing between a third-party `Drift`
and a NaN fit.

BREAKING CHANGE: `ConstantDrift`'s field is private. Replace
`ConstantDrift(x)` with `ConstantDrift::new(x)`, and `drift().0` with
`drift().gamma()`. `HistoryBuilder::score_sigma` now rejects infinity.

Closes #65

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 19:11:57 +02:00
logaritmiskandClaude Opus 5 a367155778 Merge branch 'fix/numerics-critical'
Fix the ten defects found by the 2026-09-09 floating-point audit: four
critical, three high, three medium.

Closes #55, #56, #57, #58, #59, #60, #61, #62, #63, #64

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 18:08:06 +02:00
logaritmiskandClaude Opus 5 c69a397d80 test: make the determinism test exercise the parallel sweep
It proved less than it appeared to. `sweep_color_groups` takes its
`par_iter` branch only for colour groups of at least RAYON_THRESHOLD (64)
events, and a colour group is a subset of ONE slice's events. The fixture
built 20 slices of 10, so the branch was unreachable — the test named the
parallel path and ran the sequential one.

It also compared one competitor's curve out of forty, and never compared
log_evidence, final_step or iterations.

The new fixture reaches the branch by construction: within a slice every
event uses a disjoint competitor pair, so greedy colouring puts all 96 in
colour 0. Competitors recur across slices, so the fit keeps temporal
coupling and drift rather than degenerating into independent duels.

Verified by instrumenting `sweep_color_groups`: 872 sweeps, one colour
group of 96 each, parallel branch taken all 872 times.

Worth recording how that verification went, because I nearly drew the
opposite conclusion. My first two instrumented runs printed nothing and I
read that as "the branch is still unreachable" — but `cargo test` captures
stderr without `--nocapture`, so the probe was invisible, not absent. An
instrument that cannot report is indistinguishable from a negative result.

Now compares every competitor's curve plus log_evidence, final_step and
iterations, and asserts the curve count so it cannot silently go back to
measuring almost nothing. A companion test pins EVENTS_PER_SLICE against
the threshold, so shrinking the fixture fails loudly rather than quietly
returning the suite to the sequential path.

Cross-process coverage is separate, in tests/cross_process_determinism.rs
(#62) — an in-process test cannot see hasher-order effects at all, since
every sample shares one seed.

Closes #64

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 18:08:06 +02:00
logaritmiskandClaude Opus 5 7aa7fb62dd fix: make posterior_of reproducible across processes
`ResolvedTerms::unseen` was a `HashMap<String, f64>` and three float
reductions iterated it. Addition is not associative and Rust seeds its
default hasher per process, so `posterior_of` returned different bits run
to run on identical input: measured over 40 processes, two distinct sigma
bit patterns, and five distinct values from `expected_variance_reduction`
spanning about 7 ULP.

A `BTreeMap` fixes it by construction. 40/40 identical after, 24/16
before.

The cross-batch conflict scan had the same cause with a different
symptom. It returns on the FIRST conflict, so hash order decided WHICH
competitor the error blamed — 15 different competitors named across 40
runs on identical input. The error fired every time; only its content was
a lottery, which sends a reader after the wrong key. Now scanned in
sorted order.

Magnitude was 1-7 ULP throughout, so no decision changes. The cost was
reproducibility: a golden test over these would flake at a low rate,
which is the worst kind of CI failure to diagnose.

tests/cross_process_determinism.rs re-executes the test binary and
compares bits, because an in-process test CANNOT see this — every sample
in one process shares one hasher seed. That is not hypothetical:
tests/determinism.rs compares four thread counts inside one process and
passed throughout while this was live.

Tuning that fixture took a measurement. Coefficients spread over nine
decades detected the bug in roughly one run in forty, because the small
terms fall below the running total's ULP and are absorbed whatever the
order. Comparable magnitudes keep every term able to change the last
bits: 5 of 5 attempts detected it, with 3 to 38 of 40 runs differing.
Verified non-vacuous by reverting the BTreeMap and watching it fail.

Closes #62

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 18:02:42 +02:00
logaritmiskandClaude Opus 5 305f822964 fix: route the last three transcendentals through libm, and enforce it
CLAUDE.md requires transcendentals to go through libm rather than std,
because std delegates to the system math library and the two disagree by
one ULP often enough to change an iteration count in a fixed point.

Three production sites did not:

  factor/margin.rs:84  cavity.sigma().hypot(sigma)   12.136% of 1e6 inputs
  factor/margin.rs:92  f64::MIN_POSITIVE.ln()         3.437%
  factor/trunc.rs:98   f64::MIN_POSITIVE.ln()         same

`hypot` is the material one: it is on the path of every scored event, and
its divergence rate is HIGHER than the 9.7% the rule cites for `exp` as
its own justification. The two `ln` calls happen to agree bit-for-bit on
this host, which is exactly the platform dependence the rule exists to
remove.

The `hypot` choice itself was right and stays — the comment above it
explains why, and it is measured: naive sqrt(a^2 + b^2) overflows to inf
at 1e200 and flushes to zero at 1e-200 where hypot does neither. Only the
implementation moves.

tests/libm_rule.rs enforces it. The rule was stated plainly in CLAUDE.md
and still violated three times, so prose is evidently not sufficient. The
test strips `#[cfg(test)]` items by brace matching, plus comments and
string literals so prose is not mistaken for a call, then scans for std
method spellings. `sqrt` is exempt: IEEE 754 specifies it, so std and
libm cannot disagree.

Confirmed non-vacuous by reintroducing the `hypot` violation and watching
it fail with the offending line, then pass again on restore. Two further
tests pin the stripper itself, since a stripper that removed everything
would make the guard pass on anything.

Closes #63

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 17:56:23 +02:00
logaritmiskandClaude Opus 5 ab23476aaf fix!: validate the constructors below HistoryBuilder
0.8.0 closed the sign-absorption defect at `HistoryBuilder::mu/sigma/beta`
and at both ingestion paths. It was still open one layer down, in the
constructors those paths call. Measured, all bit identical to their
positive counterparts:

  Gaussian::from_ms(25.0, -8.33)  == from_ms(25.0, +8.33)
  Rating::new(_, -4.17, _)        == Rating::new(_, +4.17, _)
  ConstantDrift(-0.0833)          == ConstantDrift(+0.0833)

sigma, beta and gamma enter only as squares, so the sign vanished without
comment. Worst of the set: `Rating::new(_, NaN, _)` reached `Game::ranked`
which returned **Ok** carrying `Gaussian { pi: NaN, tau: NaN }` — no
`converge` on that path to catch it.

`from_ms` and `Rating::new` now reject. `ConstantDrift` cannot: the field
is public and positional, so there is no constructor to intercept, and
sealing it would break every `ConstantDrift(x)` for a case whose resulting
model is perfectly valid. Documented instead. Its non-finite half IS
rejected — `converge` validates the drift variance each competitor
accumulates, which also covers a custom `Drift` impl.

Two things the tests caught that I had wrong:

NaN sigma must PASS `from_ms`. My first version rejected it, and two
existing tests went red immediately: a broken fit legitimately produces a
NaN sigma from `sqrt` of a negative truncated variance, and the design is
to propagate that to `NonFiniteResult`. Rejecting it turned the reporting
path into a panic inside inference. Written as
`sigma >= 0.0 || sigma.is_nan()` so the intent is explicit rather than
hidden in a negated comparison.

Very small sigma is also not rejected, and that is deliberate: `approx`
produces small truncated sigmas legitimately. `pi = 1/sigma^2` leaves
f64's range below ~1.5e-154 and `tau = mu*pi` overflows sooner, at a
threshold that depends on mu — so there is a band where pi is finite and
only tau is not. Both land on the existing point-mass representation.
Documented, including that such a Gaussian is not equal to itself and can
make two identical declarations report as conflicting.

BREAKING CHANGE: `Gaussian::from_ms` panics on a negative sigma, and
`Rating::new` panics unless beta is finite and non-negative. `converge`
returns `InvalidParameter` for a non-finite drift variance.

Closes #61

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 17:48:41 +02:00
logaritmiskandClaude Opus 5 6139061740 fix: keep the truncated variance representable in the far tail
`v_w` returned `w` and let `trunc` form `1 - w`. `w` tends to 1 out in
the tail, so that subtraction lost about log10(alpha^2) digits — and the
quantity it was destroying is perfectly representable.

Two separate cancellations, fixed separately.

The non-tie half: `half_line_truncation` now returns `1 - w` computed
symbolically rather than as `1 - v*gap`. With `alpha*gap = 1 - inv^2*b`
the leading ones cancel on paper instead of in floating point. Measured
against the exact truncated variance:

  alpha    before          after
  1e6      8.9e-5 rel      0.0 rel (exact)
  1e8      returns 0.0     0.0 rel (exact)

At 1e8 the old form gave `sigma_trunc = 0`, and `from_ms(mu, 0.0)` is a
point mass whose `mu()` is inf/inf = NaN. `beta(1e-8).sigma(1e-8)` with
priors 1000 apart went from Err + NaN skills to a finite fit.

The tie half is a different subtraction — `w = v^2 - u`, where both grow
as alpha^2 while their difference stays O(1). The existing escape hatch
could not cover it: it keys on `alpha * width >= HALF_LINE_WINDOW`, how
many window-widths from the mean the window sits, and a NARROW window
fails that however deep it is. Measured at alpha 1e6 with a 1e-6 window
it kept four digits and returned `1 - w = -2.4e-4` where the truth is
+2.8e-13. One step earlier it was quietly wrong instead: `1 - w = 1.0`
exactly, a truncation reported as a no-op, where the truth was 5e-17.

Over a narrow window the density is a truncated exponential in
`s = (x - alpha)/width`, whose mean and variance are closed forms, so
`v = alpha + width*m(t)` and `1 - w = width^2 * V(t)` with no large
subtraction at all. Validated against high-precision quadrature: v exact
to 4e-10, `1 - w` to 4e-10 across the region it is used in.

The crossover is on `alpha / width` rather than on either alone, because
that ratio is what says how many digits the subtraction has left — and
the approximation is most accurate exactly where the subtraction is
worst, since both improve as the window narrows.

Defaults are bit-identical (pi 0.02398318151216503 before and after).

Tests: the three reproductions from the issue, the narrow-window form
against pinned quadrature values, and a continuity sweep across all three
tie branches — a misplaced crossover is the real risk here, and a jump at
a boundary is visible even without pinning absolute values.

Closes #60

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 17:42:16 +02:00
logaritmiskandClaude Opus 5 f1219036b3 fix: take quality's determinant ratio in log space
`quality()` computed `det(ata) / det(middle)` in linear space. Both are
products of `k - 1` diagonal entries, so they leave f64's range long
before their ratio does — and the ratio is the only thing the answer
needs.

Measured at the crate defaults: 150 groups correct at 8.45e-53, 200
returned 0, 250 returned NaN where the truth is 9.51e-88. With a small
beta it bit far sooner: at sigma = beta = 1e-3, 60 groups returned NaN
against a true 1.32e-9 — a value nine orders of magnitude inside the
normal range. Neither `quality()` nor `History::predict_quality` caps the
group count, unlike `predict_outcome`, so those are supported calls.

`Lu::ln_abs_determinant` accumulates `ln|diagonal|` instead of
multiplying, and the call site becomes `exp(e_arg + 0.5 * ln_ratio)`.

Verified against the closed form `(beta / sqrt(beta^2 + sigma^2))^(k-1)`
rather than against recorded output, across three parameter sets and
group counts to 300: every case now agrees to 1e-11 or better, including
9.88e-324 at 300 groups, which is subnormal.

Also documents the remaining panic: every rating at zero sigma with a
zero beta makes `middle` singular and `inverse()` panics. Documented
rather than converted — nothing is uncertain there, so there is no
distribution to take the quality of.

Closes #59

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 17:33:00 +02:00
logaritmiskandClaude Opus 5 31cf0998b0 test: scale the ceiling sweep by build profile
Each sample runs a full inference pass per outcome, and that is about
19x faster in release: 20 000 samples take 12.1s released against 23s
for 2 000 in debug.

`just test` runs three debug feature combinations and one release one, so
a fixed sample count pays the slow price three times and the fast one
once — exactly backwards. Scaling by `cfg!(debug_assertions)` puts the
search where it is cheap:

  debug    1 000 samples   11.7s
  release 50 000 samples   31.6s

Across the whole `just test` that is 67s against 70s before, for 25x the
samples. The debug run proves the sweep compiles and holds; the release
run is the one that actually searches.

Not moving the suite to release-only, which was the alternative
considered. `debug_assert!` is compiled out in release, and this crate
documents that as load-bearing — several defects have hidden there — so
dropping the debug runs would trade one class of coverage for another
rather than adding any.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 17:27:09 +02:00
logaritmiskandClaude Opus 5 bbc7705c75 fix!: report an unresolvable prediction grid instead of clamping
`grid_shape` asked for 12 nodes across the narrowest feature and then
clamped to MAX_GRID_POINTS with no detection that the request was not
met. Past `step/sigma ~ 1.7` the trapezoid rule stops resolving the
density, and the result is unbounded:

  sigma_a   step/sig_a   P(a first)   exact      total
  2.0e-3      0.86       0.515953    0.515953   1.000000
  1.0e-3      1.72       0.517185    0.515953   1.002388
  1.0e-4     17.17       2.791336    0.515953   5.410065

A probability of 2.79. Reachable through `predict_outcome` with a pinned
reference competitor — a documented pattern — where `predict_outcome` and
`predict_win_probabilities` disagreed 44x and `predict_outcome` was the
wrong one.

There is no useful answer on the far side of that cliff, so this reports
`GridTooCoarse` rather than guessing, and the message points at
`predict_win_probabilities`, which answers the same matchup through
adaptive quadrature and is accurate there to 1e-13. The floor is 4 nodes
per feature rather than the 12 requested, because the request carries
margin: measured accurate to 2.2e-12 at 1.4 nodes per sigma and wrong by
1.2e-3 at 0.7.

This also fixes the `ln k` ceiling violation. `expected_information_gain`
weights `probability * divergence`, so probabilities of 3.97 and 2.62
made it return 3.237828 nats against `ln 2 = 0.693147` — 4.67x over. The
crate's docs call that ceiling its sharpest test and record a prototype
once returning 4.77 nats; it was live again by a different route.

The new sweep then caught a second, independent defect: `kl_divergence`
returned NEGATIVE values, worst -5.55e-17, exactly one ULP of its
`- 1.0`. Rewritten as `0.5*(u - ln1p(u)) + gap^2/(2*var_p)` with
`u = var_q/var_p - 1`, so both terms are non-negative by construction.
It is also more accurate where it matters: at `u = 1e-9` the old form
returned 0.0 where the true value is 2.5e-19, and well-conditioned cases
are unchanged.

tests/prediction_bounds.rs sweeps rather than spot-checks, because a
single fixture cannot defend a bound like this — the previous check
passed throughout. It asserts the sweep still reaches the coarse-grid
regime, so it cannot quietly stop testing the case it was written for.

BREAKING CHANGE: `predict_outcome`, `predict_ranking` and
`expected_information_gain` return `GridTooCoarse` for matchups whose
performance sigmas are too far apart to integrate on one grid. They
previously returned wrong answers, including probabilities above 1.

Closes #55, closes #56

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 17:21:00 +02:00
logaritmiskandClaude Opus 5 83bdb84152 fix!: collapse a drift too small to represent, on a relative threshold
`time_expanded_joint` collapsed consecutive appearances only at
`drift <= 0.0` exactly. Anything smaller-but-positive got an explicit
`1.0 / drift` precision, which the matrix cannot hold: at `drift = 1e-16`
the entry is `1e16`, and `1e16 + 0.28` rounds back to `1e16`, so the
prior and the event contrasts are annihilated in the stored f64 before
the factorisation ever runs.

Measured, 8 competitors over 15 slices:

  drift_scale   before                      after
  1e-6          1.3e-3 relative error       exact
  1e-7..1e-9    Err(JointUnavailable)       exact
  1e-10         12 200x TOO SMALL, as Ok    exact

At 1e-10 the caller was handed sigma = 0.0055 where the truth is 0.6108
— a 111x overconfident interval, returned as a success.

This is representation, not conditioning. Solved in 200-digit precision
the same system converges smoothly onto the collapsed value and is flat
from 1e-16 to 1e-40, so the quantity is perfectly well conditioned. That
also rules out the obvious fix: symmetric (Jacobi) equilibration measured
30x WORSE, because the information is gone from the assembled matrix
before any solver sees it. The fix has to be at assembly.

The threshold balances the two errors that trade off. Ignoring a real
drift costs about `drift / V`; representing one costs about
`EPSILON * V / drift`. They cross at `V * sqrt(EPSILON)`, scaled to each
competitor's own prior variance.

Ordinary drift is far above it and unaffected — the default gamma
accumulates 0.0069 per unit time against a threshold of 1.0e-6 — and the
test asserts both halves: everything below the threshold reaches the
collapsed answer bit-identically, and a drift of 1e-2 still moves it, so
the test cannot pass by collapsing everything.

Also corrects the `JointUnavailable` message, which asserted "a
competitor has neither a proper prior nor any evidence" for a fixture
where every competitor had both.

BREAKING CHANGE: a drift variance below `prior_variance * sqrt(EPSILON)`
now collapses two appearances into one latent variable. Affected fits
previously returned a badly wrong variance or an error.

Closes #57

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 17:03:00 +02:00
logaritmiskandClaude Opus 5 c65373f476 fix!: propagate NaN through the convergence reduction
`tuple_max` compared with a plain `>`, which is false against NaN, so a
NaN accumulator was replaced by the next finite delta. The fold runs over
`TimeSlice::posteriors()`, a HashMap, so whether a NaN survived to `step`
depended on per-process hash order.

Measured before, four competitors in one slice with one pathological
pair, same binary and input, 30 separate processes:

  16  Ok  converged=true, iterations=1, a = Gaussian { pi: NaN, tau: NaN }
  14  Err NonFiniteResult

After: 30/30 Err. A coin flip on whether a NaN fit was reported as an
error or as a successful, converged fit — inside the guard whose entire
purpose is "NaN is never convergence".

`f64::max` would not have fixed it. It also ignores NaN by design, which
is the same defect wearing a standard-library name, and a test pins that
we do not use it.

`Gaussian::delta` had to be fixed FIRST, and that ordering is the whole
subtlety. Two identical improper messages produced `(0.0, NaN)` — not
from `mu()`, which is guarded and returns 0.0, but from `inf - inf` in
the sigma component. That NaN is reachable in ordinary healthy inference:
once a pairing is more than about nine cavity-sigma apart the truncation
is a no-op and the chain compares one identity message against another.
Propagating NaN without fixing `delta` would therefore have turned
correct fits into NonFiniteResult errors. `delta` now answers the
identical-message case in natural space before touching the accessors.

My first version of the `delta` test asserted `mu()` was NaN. It is not;
the accessor guards `pi <= 0.0`. The test caught my own wrong premise,
and the doc comment is corrected to match.

BREAKING CHANGE: a fit that produced NaN in a non-final reduction
position previously returned `Ok` with `converged: true` and a NaN
posterior; it now returns `Err(NonFiniteResult)`. That was always the
documented intent.

Closes #58

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-09 16:59:56 +02:00
logaritmisk 7da2328692 chore: Release trueskill-tt version 0.8.0 2026-09-08 21:18:53 +02:00
logaritmiskandClaude Opus 5 a73afa5f24 Merge branch 'fix/game-boundary'
Reject malformed games at the Game entry point, which does not pass
through History's ingestion chokepoint.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 21:13:02 +02:00
logaritmiskandClaude Opus 5 eebf8aacd3 fix!: reject malformed games at the Game boundary too
I fixed this at `History`'s ingestion chokepoint and said the boundary
was complete. It was not. `Game` is a separate public entry point that
does not pass through that chokepoint, and every one of the same four
defects was still live there:

  Game::ranked(&[&[a]], ..)  -> PANIC at src/game.rs:317
  Game::scored(&[&[a]], ..)  -> PANIC at src/game.rs:317
  Game::ranked(&[&[], &[a]]) -> Ok, finite posterior for the opponent
  Game::scored(.., [NaN, 1]) -> Ok

The same panic, from safe API, in release. Fixing one path and
generalising from it is exactly the mistake that produced the
latest-slice joint bug: validating on the shape that cannot expose the
problem, then reporting the property as held.

`Game::validate_teams` is shared by `ranked` and `scored`, with the
non-finite score check in `scored` alongside it. Ranks need no equivalent
— they are `u32`.

`one_v_one` and `free_for_all` build their teams internally and are
unaffected; a test asserts all three well-formed constructors still
succeed, so the check cannot quietly widen.

BREAKING CHANGE: `Game::ranked` and `Game::scored` return
`NotEnoughTeams`, `EmptyTeam` or `InvalidParameter` for inputs they
previously panicked on or silently accepted.

Refs #18, #26

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 21:13:02 +02:00
logaritmiskandClaude Opus 5 4e9aa6bdc1 Merge branch 'test/close-coverage-gaps'
Cover non-finite results and color-group disjointness, closing the two
test gaps #26 named.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 21:09:25 +02:00
logaritmiskandClaude Opus 5 a18df521eb test: cover non-finite results and color-group disjointness
The two gaps #26 named that were never filled.

NonFiniteResult had no test at all — the name appeared in `tests/` only
inside a doc comment, and it is the sub-claim in that issue's title. It
turns out to be very much reachable, and from *finite* inputs: sigma at
1e300, beta at 1e300, sigma at 1e-300, score_sigma at 1e-300, and scores
at 1e308 all overflow inside inference, where the boundary checks cannot
see them. That matters because the failure is silent by default — NaN
fails every comparison, so a naive `step < epsilon` reads a NaN step as
converged, which is why the crate has `step_converged`/`step_is_finite`.
Pinned from outside, including that `converge_partial` does not launder a
breakdown into an `Ok`, and with a control asserting merely extreme
parameters still converge so the suite cannot pass by always failing.

Color-group disjointness was #26's fourth acceptance criterion and had
only five hand-written cases. Now a proptest over three shapes: a dense
pool where collisions force colors to multiply, a sparse one where most
events are independent, and repeated members within a single event.

Two of my first assertions were wrong about the code rather than the
reverse. A competitor named twice *within* one event is not a collision —
`color_greedy` collects each event's members into a set for that reason.
And contiguity is not a property of `color_greedy`: it holds only after
`recompute_color_groups` reorders events so each color occupies one
range. The test now asserts what is actually promised — that the reorder
is always *possible*, since the parallel sweep slices `&mut` sub-ranges
from those groups and overlapping ranges would be unsound.

Refs #26

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 21:09:25 +02:00
logaritmiskandClaude Opus 5 1e4b589a9c Merge branch 'fix/non-finite-weights'
Reject non-finite weights at ingestion, completing the malformed-input
boundary.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 21:03:17 +02:00
logaritmiskandClaude Opus 5 8b20e0c560 fix!: reject non-finite weights at ingestion
Measured, a NaN weight behaved exactly as `0.0`:

  weight NaN -> Ok, converged: true, 1 iteration, step (0.0, 0.0)
                skill pi 0.027777777777777776, tau 0.0
  weight 0.0 -> Ok, same skill, bit for bit

So a NaN arriving from a division or a parse was indistinguishable from
a deliberate zero, and the fit reported itself as cleanly converged.

Worth correcting an earlier description of this: the event does not
vanish. The member contributes nothing, which is precisely what weight
zero means, and that equivalence is what makes it undetectable rather
than merely wrong.

Zero and negative weights stay accepted. Both are expressible choices
about how much a member contributes, and tests/degenerate_inputs.rs pins
their behaviour deliberately; only values that are not quantities at all
are rejected. A test asserts they still ingest, so the new check cannot
quietly widen.

This completes the boundary: every malformed input that previously
produced a plausible answer — a one-team event, an empty team, a
non-finite score, a non-finite weight — now fails where it enters.

BREAKING CHANGE: an event carrying a non-finite weight returns
`InvalidParameter` instead of silently treating that member as weightless.

Refs #18

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 21:03:16 +02:00
logaritmiskandClaude Opus 5 862779ae34 Merge branch 'feat/convergence-strictness'
Make a short fit an error, raise the default iteration cap, validate the
remaining HistoryBuilder parameters, add History::register and
History::rating, reject competitor config conflicts across batches, and
document what the joint's cost scales in.

Closes #50

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 21:00:39 +02:00
logaritmiskandClaude Opus 5 f692906ce4 docs: state what the joint's cost actually scales in
A consumer measured an 8x difference in solve time between two fits over
the same events, the same slices and the same ~2,000 nodes:

  career fit   (gamma = 0)      787 ms
  drifting fit (gamma = 0.15)  6214 ms

Entirely the collapse rule. A competitor with zero drift contributes one
variable however long the history, so a drift-free fit's joint is smaller
than a drifting one's by roughly the slice count — and to factorise, by
its cube. Choosing a drift configuration is therefore also choosing a
query cost, and nothing said so.

Documented on `Joint`, on `Joint::variables` and on `posterior_of`, with
the measurement. `variables()` is named as the number that decides
affordability, since it can be read before committing to a batch.

Also states the thing the consumer proposed as a future optimisation,
because it is already true: an absence is not an appearance, so a
competitor seen in the first and last of a hundred slices contributes two
variables rather than a hundred. The matrix is already as small as the
model allows on that axis.

tests/joint_handle.rs pins the mechanism — ten slices, two competitors,
twenty variables drifting against two at `gamma = 0` — so a change to the
collapse rule cannot quietly remove the property the docs now promise.

Refs #51

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 19:45:51 +02:00
logaritmiskandClaude Opus 5 e493f47e99 feat!: add History::register and History::rating, and reject config conflicts across batches
Three things #38 asked for, on a premise that had half dissolved. The
issue argued from "captured at first appearance", "missing it is silent"
and "missing it is permanent"; 8c087ad made configuration apply whenever
supplied and refit the whole history, so two of those are already gone.
What survived is the literal title — no way to say it before the first
event — plus the absence of any way to check.

`register(Member)` states configuration before anything is observed. It
takes the same `Member` ingestion takes, so there is one vocabulary
rather than two, and it creates the competitor immediately, which is what
makes it observable. It reaches a competitor first seen through
`record_winner`, the route #37 deliberately did not extend.

`rating(&key)` reads back what was stored. Every other accessor reports
what inference inferred; this reports what it was told, which is what
makes a configuration mistake detectable from outside the crate at all.

Conflicting configuration is now an error across batches, not only within
one. The `priors` map is rebuilt per `add_events` call, so a second batch
silently overwrote what a first declared, last-write-wins. That cut
directly against the invariant tests/ingestion_equivalence.rs exists to
protect: the same contradictory events errored when batched and
succeeded, order-dependently, when fed one at a time. Detection lives on
a new `declared` map on `History`, because a `Rating` cannot say whether
a value was chosen or inherited from the defaults — which is exactly the
distinction the check needs. Checked before anything mutates, so a
rejected batch leaves the history untouched.

`register` rejects a non-default `weight` rather than ignoring it. Weight
is per-event and has no meaning on a registration, and silently dropping
a field the caller set is the defect this whole area keeps producing.

The declarative `default_rating_for` closure is not here. It is the
better answer for ustat's actual case — thousands of keys matching a
rule, rather than enumerated — but it adds a `Fn` parameter to `History`,
which the issue itself flags as in tension with the crate's posture. That
wants its own decision rather than riding along.

BREAKING CHANGE: two different values for one competitor's `prior` or
`drift_scale` supplied across separate `add_events` calls now return
`ConflictingCompetitorConfig` instead of silently taking the later one.

Refs #38

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 16:11:09 +02:00
logaritmiskandClaude Opus 5 4f6360128d feat!: validate mu, sigma and beta on HistoryBuilder
The last three unvalidated setters, beside `p_draw`, `score_sigma` and
`convergence`, which all assert eagerly. Measured before choosing bounds:

  beta = 0        -> works, pi 0.0211 (vs 0.0193 at the default)
  beta = -4.17    -> bit-identical to +4.17
  sigma = -8.33   -> bit-identical to +8.33
  sigma = 0       -> NonFiniteResult, current_skill returns tau: NaN
  sigma = inf     -> same
  mu = NaN        -> same

So the bounds are not the obvious ones. `beta = 0` is legitimate and
meaningful — performance is then exactly skill, and the fit moves
measurably rather than degenerating — so zero is allowed and a test pins
that it reaches a different answer, since "allowed" would otherwise be
indistinguishable from "unchecked".

The negative cases are the quiet ones. `sigma` and `beta` enter inference
only as squares, so a negative value behaves as its absolute value and
the sign is dropped without comment. That is the same defect
`Member::with_drift_scale` already rejects, for the reason already
written there.

The non-finite cases are detected today — `converge` reports
NonFiniteResult — but a caller who reads `current_skill` first is handed
`tau: NaN`, so rejecting at the boundary is what actually closes it.

BREAKING CHANGE: `HistoryBuilder::mu`, `sigma` and `beta` now panic on
values they previously accepted, matching the existing behaviour of
`p_draw`, `score_sigma` and `convergence`.

Refs #18

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 16:06:37 +02:00
logaritmiskandClaude Opus 5 eff63dfa2a feat!: make a short fit an error and raise the default iteration cap
`ITERATIONS` was 30, and overrunning it returned `Ok` with
`converged: false`. Both halves were wrong.

The cap is a runaway guard, not a budget: the sweep exits as soon as the
step falls below `epsilon`, so a high cap costs nothing on a history that
converges. Measured on one needing four sweeps, `max_iter` 30 and
100_000 both finish in 4 iterations and ~130us. So 30 could never make
anything faster — it could only stop a healthy history early, and it did:
160 events over 100 competitors already needs 42.

Not scaled to the history, because iteration count tracks how loopy the
graph is rather than how big it is. At a fixed 320 events over 40 slices,
varying only the competitors sharing them: 3 competitors needs 2_789
sweeps, 10 needs 1_068, 100 needs 90, 400 needs 2. Three orders of
magnitude on identical event and slice counts, so any formula in those
two numbers would be badly wrong on some real shape. A single value set
high enough that reaching it means oscillation is the honest version.

With the cap raised, stopping at it means something is genuinely wrong,
so `converge` now returns `InferenceError::NotConverged` rather than a
flag on a success. A short fit is wrong by a little — every rating
finite, the ordering sensible, nothing saying the numbers were still
moving — and a flag has to be checked while `let _ = h.converge()` is the
natural way not to. That is not hypothetical: it is how a real defect hid
in this crate's own test suite.

`converge_partial` returns the short fit for callers who want one. Only a
single existing test needed it, which is the evidence that a capped fit
is a deliberate choice rather than the common case.

Also corrects the `ITERATIONS` docs, which claimed convergence cost is
"roughly linear in the cap". It is linear in the iterations actually run.

BREAKING CHANGE: `History::converge` returns `Err(NotConverged)` where it
previously returned `Ok` with `converged: false`. Callers that want the
old behaviour should use `History::converge_partial`. The default
`max_iter` changes from 30 to 10_000, so a history that was silently
truncated will now converge properly and its numbers will move.

Closes #50

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 16:04:22 +02:00
logaritmiskandClaude Opus 5 7c6965c6a9 Merge branch 'fix/ingestion-shape'
Reject malformed events at the ingestion boundary, add
EventBuilder::members, and record the rayon opt-in deviation.

Closes #5
Closes #37

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 15:53:13 +02:00
logaritmiskandClaude Opus 5 911b48faba feat: add EventBuilder::members for per-member configuration
`EventBuilder` could set weights and nothing else, so `prior` and
`drift_scale` were reachable only through the typed
`Event`/`Team`/`Member` shape plus `add_events`. Which ingestion route a
competitor arrived through decided whether it could be configured.

`members(...)` takes `Member` values directly, so `Member`'s own builder
expresses everything. `team(...)` stays the common case.

One escape hatch rather than `priors` and `drift_scales` setters beside
`weights`, as the issue suggested and then argued against itself: a
parallel array per field means a parallel length check per field, and
each one is a new way to get the lengths wrong. `Member` already has a
builder; this just lets the fluent path reach it.

`record_winner`/`record_draw` are deliberately left alone. They are the
two-argument convenience path, and extending them would be a breaking
signature change. The issue's reason for wanting them extended has also
weakened: it said a competitor arriving through them was "permanently
stuck on the history defaults", and since 8c087ad that is no longer true
— a later `add_events` carrying the `Member` refits the whole history.
Measured, late configuration through that route reaches mu 40.000000000,
identical to configuring from the start.

Refs #37

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 12:32:00 +02:00
logaritmiskandClaude Opus 5 f57784c141 docs: record the rayon opt-in deviation in spec section 6
Issue #5 asked for a decision, not an implementation: either flip rayon
to default-on, or record why the spec was deviated from and close.

Opt-in stands. The measured speedups are 1.0x realistic / 1.3x
pathological (#4), so default-on would cost every downstream user a
thread pool and a dependency for approximately nothing.

The condition the decision was waiting on cannot be met: #5 was blocked
on re-measuring after cross-slice dirty-bit skipping landed, and #4 was
closed by removing the inert slices_skipped field rather than by
implementing it. There is no forthcoming measurement to wait for.

Also corrects the spec's own reasoning. It cited an unsafe concurrent
write through SkillStore as a cost of going default-on; the crate is
forbid(unsafe_code) and the compute/apply split avoids that entirely.
The case for opt-in is the measurements, not a safety argument.

Closes #5

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 12:29:40 +02:00
logaritmiskandClaude Opus 5 8e4d6a637d fix: reject malformed events at the ingestion boundary
A one-team event reached `run_chain`, which builds one diff link per
adjacent pair of teams, leaving it to index `links[1..]` on an empty
vector. That panicked with "range start index 1 out of range for slice
of length 0" — from `History::add_events`, in a release build, through
entirely safe API.

An empty team was the quieter half of the same gap. It contributes no
performance, so a malformed event converged and handed back a finite,
plausible-looking posterior for whoever it was matched against. That is
this crate's characteristic defect: a public surface reporting a
constant that looks like an answer.

A non-finite score was the third. `converge` did report NonFiniteResult,
so it was detected — but a caller reading `current_skill` before
converging was handed `tau: NaN` with nothing to say so.

`NotEnoughTeams` and `EmptyTeam` already existed. They were checked on
the prediction paths and nowhere else, which is exactly why ingestion
could still manufacture the states they describe. The checks go in
`add_events_with_prior` alongside the tie check, for the same reason
that one is there: every ingestion route lands on it, so `record_winner`,
`record_draw` and `EventBuilder` inherit them rather than each needing
their own.

Also corrects documentation that had been stating the opposite of the
code since 8c087ad in 0.4.0. README.md and the `with_prior` /
`with_drift_scale` doc comments all still said competitor configuration
was "captured at first appearance" and had "no effect" on a known key.
It now applies whenever supplied and refits the whole history. A reader
would have concluded late configuration was impossible and built a
workaround for a limitation that does not exist. CI compiles README code
blocks but not prose, which is why it survived three releases.

The comment in tests/degenerate_inputs.rs claiming a one-team event was
"rejected for an unrelated reason" was wrong when written — it panicked.

Refs #18, #26

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 12:29:14 +02:00
logaritmisk 82eff740b6 chore: Release trueskill-tt version 0.7.0 2026-09-08 10:27:21 +02:00
logaritmiskandClaude Opus 5 c1b1c6c7d7 Merge branch 'feat/joint-handle'
Factorise the joint once with History::joint, so a batch of queries pays
the O(n^3) Cholesky once rather than once per question.

Closes #51

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 10:24:09 +02:00
logaritmiskandClaude Opus 5 1bb6bb31d8 feat: factorise the joint once with History::joint
`posterior_of`, `posterior_of_at` and `expected_variance_reduction` each
built the joint precision matrix, factorised it, asked one question and
threw it away. The factorisation is O(n^3) in the history's appearances
and depends only on the fit, so a caller asking about every pair in a
standings table, every cell in a grid, or every candidate in an
active-learning sweep paid for the same factorisation once per question.

`History::joint()` returns a `Joint` handle that pays it once. Measured
on 1976 appearances, 90 queries: 68.4s one-shot against 745ms factorise
plus 93ms of queries — 81.6x, with bit-identical answers. Per query,
Criterion at 480 appearances: 9.0ms one-shot against 48us cached, 187x.

The handle borrows the history, which is what makes it correct with no
invalidation logic: the borrow checker forbids adding events or refitting
while it is alive, so there is no window in which the factorisation could
describe a fit that no longer exists. It also makes the lifetime of the
n^2 factor explicit rather than parking it in the history forever — at
4000 appearances that is 128MB, which is not something to cache silently.

Every question the joint answers turns out to be a bilinear form,

    c^T A^-1 a = (L^-1 c) . (L^-1 a)

so no caller ever needs L^-1 c itself. Replacing the general solve with a
forward substitution drops the back substitution as wasted work, halving
a query, and removes a failure mode: a variance as `c . (A^-1 c)` is a
difference of products that can round negative, where `|L^-1 c|^2` is a
sum of squares and cannot.

The one-shot calls are unchanged in cost and now delegate to the handle,
so the two paths cannot drift apart. tests/joint_handle.rs asserts they
agree bit for bit, including at pinned times, under UnknownKeys::Prior,
and across candidate matchups.

Refs #51

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 07:53:51 +02:00
logaritmisk b113385c6f chore: Release trueskill-tt version 0.6.0 2026-09-08 06:46:36 +02:00
logaritmiskandClaude Opus 5 f345e7690e fix!: make the joint span slices, not just the latest one
`posterior_of` shipped in 0.5.0 reading a single slice. Measured against
a real Through-Time history that answers almost nothing: ustat's round
fit is 76 per-day slices whose last one holds a solo round, so 0 of 55
pair differences resolved and the single node that did was degenerate —
a one-competitor slice has no correlation to account for and returns the
marginal unchanged.

That was my mistake, and the fixture chose it. I validated against
single-slice histories, which is exactly the shape that cannot reveal
the problem. In a library whose premise is skill over time, competitors
are read at *their own* last appearance and those are different slices
by construction.

The joint is now time-expanded: one variable per appearance, linked by
the prior on a first appearance, the drift between consecutive ones, and
the within-slice event contrasts. Consecutive appearances with no drift
between them are the same variable rather than two joined by an infinite
precision, which keeps the matrix positive-definite when a competitor is
pinned with `drift_scale = 0`.

`posterior_of` now reads each competitor at their own latest appearance,
which is where `current_skill` reads them, so the two agree about which
posterior they describe. Adds `posterior_of_at(time, terms)` for a
comparison anchored to a moment, matching `learning_curve`'s reading.

Validated against a hand-written exact posterior for a two-competitor,
two-slice history — the precision matrix is spelled out in the test
rather than obtained from the crate, so it is an independent check
rather than a restatement. Also pinned: competitors last seen in
different slices now compare at all, means still agree with the
marginals, zero drift makes slice layout irrelevant, and more drift
widens a comparison across time.

BREAKING CHANGE: `posterior_of` and `expected_variance_reduction` now
consider the whole history rather than its latest slice, so results
change for any multi-slice history. `JointUnavailable` is now returned
when *any* slice holds ranked events, not just the last.

Refs #46, #47

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 06:04:41 +02:00
logaritmisk d9e85cda1d chore: Release trueskill-tt version 0.5.0 2026-09-08 02:16:55 +02:00
logaritmiskandClaude Opus 5 633a503900 refactor!: remove the factor-graph surface nothing used, add try_winner
Two decisions taken before cutting 0.5.0.

#42 — the `Schedule` trait was public API the engine never called. Its
only call site was `Game::custom`, itself `#[doc(hidden)]`, and
`EpsilonOrMax` was never constructed anywhere. Removing that surface
showed the problem was larger than the issue described: with `custom`
gone, the compiler found `Factor`, `BuiltinFactor`, `RankDiffFactor` and
`TeamSumFactor` all dead too.

`Game::run_chain` drives a local `DiffFactor` enum and bypasses the
whole T1 abstraction — it has done since it was written. So this is not
just an unused extension point but the machinery it was built on, and
`CLAUDE.md` was documenting it as live architecture.

Removed: `graph` module, `Schedule`, `EpsilonOrMax`, `ScheduleReport`,
`Game::custom`, `Factor`, `BuiltinFactor`, `RankDiffFactor`,
`TeamSumFactor`. `TruncFactor`, `MarginFactor`, `VarStore` and `VarId`
stay — inference uses those. The measurement behind choosing removal
over wiring is in #42: the within-game loop converges in 1 to 8
iterations against a cap of 30, so a `Residual` schedule has no headroom
to reclaim, and `Damped` already shipped as `ConvergenceOptions::alpha`.

#20 — `Outcome::winner` panicking on an out-of-range index. Kept, and
the reasoning is now on the method. It is the only constructor here that
validates, which looks inconsistent until you try deferring like its
siblings: `winner(5, 2)` produces ranks `[1, 1]`, an all-tied draw that
ingestion accepts without complaint when `p_draw > 0`. Asking "team 5
won" and silently getting "everyone drew" is the exact failure this
crate keeps removing, so the check belongs where the mistake is.

Adds `Outcome::try_winner` for indices that are computed or parsed
rather than written literally, following the `new`/`try_new` convention.
That is additive; the panicking form stays because every call site in
this repo, its tests and its README passes literals, where a `?` would
be noise.

BREAKING CHANGE: the `graph` module and everything it exported are
removed, as are `Game::custom` and `ScheduleReport`.

Closes #42. Closes #20.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 02:14:43 +02:00
logaritmiskandClaude Opus 5 7cf45db5cf feat: add expected_variance_reduction for scored active learning
#49: `expected_information_gain` enumerates discrete outcomes, so a
consumer recording continuous scores cannot ask which matchup to run
next. The issue flagged this as possibly a research question, since
"expected variance reduction under EP may not have a clean closed form
even for Gaussian likelihoods".

It does. Observing a scored event is a rank-one update to the precision
matrix, so Sherman-Morrison gives

    reduction = (c^T L^-1 a)^2 / (v + a^T L^-1 a)

for target functional c and matchup contrast a. Verified against an
actual refit on four candidate matchups: agreement to 1e-9 relative.

Two consequences worth stating.

There is no expectation to take. The expression depends on which matchup
is played but not on how it turns out, because for a Gaussian likelihood
the posterior variance update is data-independent. Pinned by
`the_outcome_does_not_change_the_reduction`, which refits with scores of
(3, 1), (100, -50) and (0, 0) and gets the same answer. The name keeps
the term the active-learning literature uses; no averaging happens.

It is also far cheaper than its ranked counterpart — one linear solve
rather than a full inference pass per possible outcome — because
`c^T L^-1 a` and `a^T L^-1 a` share the same solve.

`target` is deliberately the same linear-functional shape as
`posterior_of`, as the issue proposed, so the two share a concept rather
than inventing two.

The load-bearing test is the refit comparison. An acquisition function
is the archetype of a surface that returns finite, plausible, monotone
numbers while being wrong, and then quietly selects worse matchups
forever; ranking behaviour alone would not catch that.

Closes #49

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 02:08:32 +02:00
logaritmiskandClaude Opus 5 c866210c65 feat: add History::predict_margin for scored matchups
#48: every predict_* answers "who wins", and a consumer recording scores
never asks that. It wants the interval on the result, and having none it
hand-fitted a noise law whose fitted node weight came out at 0.0 — so
the quoted sigma was 5.83 whether the competitor had forty rounds or
none, against real residual spreads of 5.8 and 12.44.

`predict_margin` composes the three things that make a scored result
uncertain: the joint posterior over the competitors, their per-event
performance noise, and the observation noise on the score. It widens as
the model knows less — measured, sigma 2.48 against an opponent with
forty rounds, 3.38 against one seen once, wider still against one never
seen — which is the property the hand-fitted law lost.

It is a margin, not a score, and that is not a shortcut. Scored
ingestion reduces every event to `score_a - score_b` before inference,
so the absolute level is discarded: shifting every score in a history by
+100 or -1000 produces a bit-identical fit, verified. There is no
information from which to predict what a competitor will *score*.
Returning one would be a number derived entirely from the prior, which
is exactly the plausible constant this crate keeps finding and removing.

`posterior_of` now honours `UnknownKeys::Prior`, which gives #48 its
second requirement — "I have never seen this competitor, here is the
prior-informed answer". An unseen competitor shares no event with the
slice, so it is independent by construction and its variance is additive
rather than part of the solve.

Worth recording for expectations: for a *margin* the joint buys little
over adding marginals (2.4798 against 2.5112 here), because a margin is
a difference and differences are where the loopy underestimate and the
ignored correlation cancel. The gain here is having a predictive
distribution at all. `posterior_of`'s correlation handling earns its
keep on sums and single nodes instead — see tests/additive_model.rs.

Closes #48

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 02:02:15 +02:00
logaritmiskandClaude Opus 5 1f791bcddd test: pin what an additive model does to combined uncertainty
Investigating #47 needed a reproduction of the reporter's structure: a
joint player/layout model where every observation measures a sum of
nodes against a reference, so the data pins differences and leaves the
overall level to the prior.

The result corrects #46's framing as well as answering #47. Combining
marginals is wrong in opposite directions depending on the combination,
and the unsafe direction is not the one either issue assumed:

    combination            exact   adding marginals
    p0 + h0 (a round)     4.0850   0.8041   0.20x   OVERconfident
    p0 - p1 (rank two)    0.8741   0.8592   0.98x   about right
    h0 - h1               0.7119   0.7057   0.99x   about right

#46 assumed every published figure was too wide and that this was "at
least the safe direction". That holds for differences. For sums it
reverses: adding marginals is five times too narrow, which publishes a
claim the data does not support.

For a single node the exact marginal is ~5x wider than message passing
reports, because the level it shares with its partners is pinned only by
the prior. That is the honest posterior for a weakly identified
parameter, not a defect — and it is why #47's node "should be
publishable" intuition and its posterior sigma disagree.

Refs #46, #47

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 01:52:39 +02:00
logaritmiskandClaude Opus 5 c52e2550af feat: add History::posterior_of for a linear combination of competitors
#46: every accessor returns a per-competitor marginal, and almost
nothing a consumer publishes is one competitor. Combining marginals
assumes independence, and competitors are correlated through every event
they share.

`posterior_of(&[(a, 1.0), (b, -1.0)])` returns the posterior of that
combination with the correlation intact. Validated against the exact
linear-Gaussian posterior on both a tree and a loopy fixture, for
differences and for single competitors: agreement to 1e-9 relative in
every case.

The investigation that preceded this is why it is not a covariance
accessor. Marginals from loopy message passing are about half the true
width, and ignoring correlation overstates a difference — the two errors
partially cancel, leaving 1.327x rather than 2.646x. Bolting true
correlations onto the existing marginals would have given 0.765 against
a true 1.524, which is overconfident: the direction the reporter
specifically called unsafe. Rebuilding the joint from the factor
structure fixes both at once, and a single-competitor query now returns
the exact marginal rather than the narrow one.

The precision matrix depends only on structure — who played whom, with
what weights and what noise — not on the observed outcomes, and the
means were already exact. So only the second moment is reconstructed.

Known limits, all deliberate and documented on the method:

- Latest slice only. A functional spanning times, such as "current
  versus career", needs the time-expanded joint and is not covered.
- Scored events only. A ranked outcome's truncation is EP-approximated
  and its converged factors are not retained after inference, so ranked
  slices return `JointUnavailable` rather than a plausible wrong number.
- Dense Cholesky, O(n^3) per query in the slice's competitor count:
  38.8us at 50, 5.66ms at 400, 49.1ms at 800. Fine for the sizes this
  serves today; caching the factorization per slice would make repeat
  queries O(n^2), and sparsity is the next step after that.

Refs #46, #47, #48

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 01:46:35 +02:00
logaritmisk 4924bc8b57 style: use arrays rather than vec! in the calibration fixture
clippy::useless_vec. Pushed broken because the verification chain used
`if ... then echo` blocks that report status without gating the `&&`
that follows — the same shape of mistake as e1bddf2, where a grep
succeeded on the error text. Both times the check printed FAIL and the
push went ahead.

The reliable form is a single fail-fast chain:

    just lint && just test && just determinism && cargo test --doc

so nothing runs after the first failure.
2026-09-08 01:39:21 +02:00
logaritmiskandClaude Opus 5 36eacf5f67 test: calibrate the marginals against the exact posterior
Investigation for #46 and #47, before touching either.

A scored history is linear-Gaussian, so its true joint posterior has a
closed form and the crate can be checked against ground truth. Measured
on five competitors:

                    means      marginal sd (crate / exact)
    tree (star)     exact      1.000
    loopy (robin)   exact      0.502

On a tree the crate is exact in both. With cycles the means stay exact —
the standard Gaussian-BP result, and the property ratings rely on —
while marginal variances come out about half the true width.

That is the opposite direction from what #47 reports, so whatever is
happening in that consumer's model, the crate being conservative is not
it.

It also means #46 cannot be implemented as an added covariance accessor.
The exact correlation between two nodes here is +0.857, so ignoring it
overstates the width of a difference — but the too-narrow marginals
partially cancel that, leaving 1.327x rather than 2.646x. Adding true
correlations to these marginals without correcting them would give 0.765
against a true 1.524: overconfident, which is the direction the reporter
specifically called unsafe.

Pins the two real invariants (exactness on a tree, exact means with
cycles) and deliberately only records the variance gap, since closing it
is what #46 proposes.

Also records the working rules this project has converged on: investigate
before implementing, fix the root issue, and scout crates.io on measured
accuracy rather than adoption.

Refs #46, #47

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 01:38:46 +02:00
logaritmiskandClaude Opus 5 71554fd944 feat: add UnknownKeys::Prior, and explain why there is no Skip
#44's third ask was an opt-in mode so a caller with partially-known
teams need not pre-filter. The requested shape was `Skip` — drop unknown
members. Measured, that is the wrong mode to build.

A team's performance is the *sum* of its members, so dropping one drops
its variance too. On a two-member team with one unknown:

    SKIP  (drop the member)  : performance sigma 2.37
    PRIOR (member at prior)  : performance sigma 6.53   (2.76x wider)

Skipping makes the model *more* certain because it knows *less*, which
is backwards. `Prior` is also the answer the model already gives for a
competitor it knows about but has no evidence for — measured, such a
competitor sits at sigma 4.99 against the prior's 6.0 — so it
corresponds to a state the model can actually be in. Skipping does not.

So the enum is `Reject` (default, unchanged) and `Prior`, and it is
`#[non_exhaustive]` in case a real use for skipping turns up later.

Placed on `HistoryBuilder` rather than per-call. Neither consumer wants
it to vary between queries: one scores thousands of candidate matchups
in a loop, the other's headline feature is predicting a competitor
nobody has faced. That makes it a property of how the model is being
used, and keeps five prediction signatures unchanged.

This also gives #48 the semantics it asked for — "I have never seen this
competitor, here is the prior-informed answer" — which it needs for
predicting a course nobody has played.

`an_unknown_member_widens_its_team_rather_than_narrowing_it` pins the
property that ruled `Skip` out, so a future convenience cannot quietly
reintroduce it.

Closes #44. Refs #48

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 01:28:49 +02:00
logaritmiskandClaude Opus 5 2cf21a753d docs: record that the event log is the source of truth, and why
#45 asked whether a fitted `History` can be persisted, and noted that
the absence of `serde` "reads as an omission rather than a decision,
which invites exactly this issue from every consumer in turn". Fair.

Documents the decision on `History` along with the reasoning that makes
it one: `converge` reaches a fixed point determined by the events,
ratings and configuration alone, so a snapshot would carry no
information the event log does not — it would cache the computation,
never the answer.

It also bounds what a snapshot could buy, since that is the question a
consumer actually has. Re-converging an unchanged history costs one
iteration, measured at 0.91 ms against 365 ms cold on 2 000 events, so
it would make a cold restart cheap and do nothing for appends. Appending
one event moves its participants more than a sigma across their whole
history, so that re-convergence is real work rather than repeated work.

Closes #45

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 01:22:28 +02:00
logaritmiskandClaude Opus 5 35d7512557 fix(test): the ingestion-order property was comparing two truncated fits
`ingestion_order_does_not_change_the_answer` failed with a 1.2e-6 gap in
mu and read as a violation of the invariant. It was not. Both sides ran
`max_iter: 200` against `epsilon: 1e-10`, and the batched side stopped
at the cap with a step of 3.4e-9 — so the test compared two fits that
had not converged and attributed the difference to ingestion order.

Raised to 20_000, at which both converge and the property holds. Runtime
is unchanged at 0.05s, because converging is what the iterations were
for.

The test now asserts `report.converged` on both sides before comparing.
That is the part worth keeping: any test that compares two fits for
equality is measuring truncation unless it first establishes that both
reached a fixed point. `ingestion_equivalence.rs` already did this;
`properties.rs` did not.

Found because c12bc83 made `ConvergenceReport` `#[must_use]`, which is
the same failure #50 describes — a short fit is wrong by a little and
looks entirely plausible. The blanket `let _ =` that commit applied to
78 call sites was too blunt here: binding the report to `_` silenced the
one signal that would have caught this, in a test whose whole purpose is
to compare two fits.

Refs #50

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 01:20:48 +02:00
logaritmisk e1bddf2474 style: factor the event-pair type out of the reconvergence fixture
clippy's `type_complexity` fires on the tuple-of-vectors return. Caught
after pushing, because the verification chain used `&&` between `just
lint` and a `grep` that succeeded on the error text — so a failing lint
reported success. The gate has to be on the command, not on whether the
output matched.
2026-09-08 01:18:59 +02:00
logaritmiskandClaude Opus 5 5d36fa1008 test: pin that re-convergence is path-independent
Answering #45 — can a fitted `History` be persisted — needed to know
whether `converge` reaches a fixed point determined by the events alone,
or one that depends on the message state it started from. It is the
former, and that is worth a test rather than a comment.

`tests/ingestion_equivalence.rs` varies how events are batched but
converges only at the end. These converge *between* batches, which is
the path a caller takes when it fits, serves, then ingests more.

Measured divergence from a single fit over the same events: 6.2e-13 for
an append strictly later than every existing slice, 8.9e-11 for one
interleaved with them. Both at the convergence tolerance. The design
question guessed the interleaved case might be weaker; it is not, and
the reason is that Through Time revises the past on every converge
anyway, so doing it in two steps is not a special case.

Also pins that re-converging an unchanged history costs one iteration.
Measured on a 2000-event fixture that is 0.91ms against 365ms cold — the
fact that makes a restored snapshot worth having at all.

Refs #45

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 01:18:28 +02:00
logaritmiskandClaude Opus 5 c12bc830a5 feat!: name the unknown key, expose tail probabilities, flag short fits
Three issues from two downstream consumers, all small, all sharing a
theme: the crate had the information and would not hand it over.

#44 — `UnknownKey { team: 0, member: 0 }` did not say which key. A
consumer upgrading 0.1.2 -> 0.4.1 had every one of 5591 predictions
return this error, fell back to a neutral 0.5, and lost its entire
metadata model for a day. Nothing crashed and nothing logged; it was
found by sweeping an unrelated parameter and noticing the output did not
move. The 0.4.0 change that made unknown keys an error was right — the
error was just too anonymous to act on. It now carries the key's `Debug`
rendering, and its `Display` says what to do about it. The precondition
is documented on every prediction entry point, which the reporter said
would alone have saved the day.

#43 — `cdf` was `pub(crate)`, so a consumer asking "is this competitor
below the cutoff" approximated it with a `mu + z * sigma` band and had no
way to say what confidence any `z` bought. Adds
`Gaussian::probability_below` / `probability_above`. The second is
separate on purpose: `1 - cdf` collapses to exactly zero past ~8.3
sigma, and a stopping rule is evaluated precisely there. Both route
through the survival function added in 0.4.1, so this is visibility
rather than new numerics.

#50 — `ConvergenceReport` was not `#[must_use]`, so the one signal that
a fit stopped short was trivially discarded. It now is, and that
immediately found 78 sites doing exactly that — including this crate's
own ATP example, which was capped at 10 sweeps when the history needs
30. The example now reads the report and says so.

`ITERATIONS = 30` is documented as the floor it is, with the three
measurements to hand: 400 events over 100 competitors already stops
there at ~7e-3 against a 1e-6 tolerance, the ATP example needs 30 at a
much looser one, and a consumer's 2000-node model needs 76 to 161.

BREAKING CHANGE: `InferenceError::UnknownKey` gains a `key` field, and
the prediction methods now require `K: Debug` in order to fill it.

Closes #43, #50. Refs #44 — its third ask, an opt-in `UnknownKeys::Skip`
mode, is a live API question and deliberately not answered here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-07 23:41:03 +02:00
84 changed files with 11449 additions and 1678 deletions
+91
View File
@@ -0,0 +1,91 @@
# Measure the CI runner's own benchmark variance.
#
# #54 asks whether benchmark regressions can be gated. The threshold is the
# whole problem: too tight and CI goes red on noise, which trains people to
# re-run until green; too loose and it never fires. Which of those is possible
# depends on a number nobody has measured — how much this runner's results move
# between identical runs.
#
# So: run one unchanged benchmark ten times and report the spread. If it is
# ~15%, a fixed-threshold gate is dead and the answer is a tracker; if it is
# ~2%, a gate at 10% is meaningful.
#
# Manual only. It takes ten benchmark runs and answers a question that is asked
# once, not every push.
name: Benchmark variance
on:
workflow_dispatch:
inputs:
runs:
description: How many repeats
required: false
default: "10"
jobs:
variance:
name: runner variance on one benchmark
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: dtolnay/rust-toolchain@stable
- uses: Swatinem/rust-cache@v2
# `joint_factorise_480_appearances` is the right probe: ~9 ms, so it is
# long enough not to be dominated by timer overhead, and it is the
# measurement this crate most wants protected — the dense factorisation
# #52 is about replacing.
- name: Warm up
run: cargo bench --bench joint -- joint_factorise_480_appearances --warm-up-time 1 --measurement-time 3
- name: Repeat the same benchmark
run: |
set -euo pipefail
for i in $(seq 1 "${{ inputs.runs || '10' }}"); do
echo "== run $i =="
cargo bench --bench joint -- \
joint_factorise_480_appearances --warm-up-time 1 --measurement-time 3 \
2>&1 | tee -a raw.txt
done
- name: Report the spread
run: |
set -euo pipefail
# Criterion prints `time: [lo mid hi]` with a unit after each. Take
# the midpoints. `sort -n` rather than awk's `asort`, which is a gawk
# extension the runner's mawk does not have — that failed on the
# first try here.
grep -oE 'time:[[:space:]]+\[[^]]+\]' raw.txt \
| sed -E 's/.*\[[^ ]+ [^ ]+ ([0-9.]+) ([^ ]+).*/\1 \2/' > mids.txt
echo "--- midpoints ---"
cat mids.txt
# Criterion picks a unit per run, so mixed units would have us
# comparing 9 ms against 9 us as if they were the same number — the
# plausible-looking wrong answer this crate keeps removing. Refuse.
if [ "$(cut -d' ' -f2 mids.txt | sort -u | wc -l)" -ne 1 ]; then
echo "runs reported different units; the spread would be meaningless"
cut -d' ' -f2 mids.txt | sort | uniq -c
exit 1
fi
sort -n mids.txt | awk '{ v[NR]=$1; u=$2; s+=$1 }
END {
if (NR == 0) { print "no samples parsed - see the raw.txt artifact"; exit 1 }
printf "n = %d\n", NR
printf "min = %.4f %s\n", v[1], u
printf "median = %.4f %s\n", v[int((NR+1)/2)], u
printf "max = %.4f %s\n", v[NR], u
printf "mean = %.4f %s\n", s/NR, u
printf "spread = %.2f%% (max-min)/min\n", 100*(v[NR]-v[1])/v[1]
print ""
print "Read it against #54: a spread near 15% kills both"
print "fixed-threshold options and the answer is a tracker;"
print "a spread near 2% makes a gate at 10% meaningful."
}'
- uses: actions/upload-artifact@v4
if: always()
with:
name: bench-variance-raw
path: |
raw.txt
mids.txt
+100
View File
@@ -2,6 +2,102 @@
All notable changes to this project will be documented in this file.
## 0.8.0 - 2026-09-08
### Breaking Changes
- feat!: make a short fit an error and raise the default iteration cap
- feat!: validate mu, sigma and beta on HistoryBuilder
- feat!: add History::register and History::rating, and reject config conflicts across batches
- fix!: reject non-finite weights at ingestion
- fix!: reject malformed games at the Game boundary too
### Bug Fixes
- fix: reject malformed events at the ingestion boundary
### Documentation
- docs: record the rayon opt-in deviation in spec section 6
- docs: state what the joint's cost actually scales in
### Features
- feat: add EventBuilder::members for per-member configuration
### Other (unconventional)
- Merge branch 'fix/ingestion-shape'
- Merge branch 'feat/convergence-strictness'
- Merge branch 'fix/non-finite-weights'
- Merge branch 'test/close-coverage-gaps'
- Merge branch 'fix/game-boundary'
### Testing
- test: cover non-finite results and color-group disjointness
## 0.7.0 - 2026-09-08
### Features
- feat: factorise the joint once with History::joint
### Miscellaneous Tasks
- chore: Release trueskill-tt version 0.7.0
### Other (unconventional)
- Merge branch 'feat/joint-handle'
## 0.6.0 - 2026-09-08
### Breaking Changes
- fix!: make the joint span slices, not just the latest one
### Miscellaneous Tasks
- chore: Release trueskill-tt version 0.6.0
## 0.5.0 - 2026-09-08
### Breaking Changes
- feat!: name the unknown key, expose tail probabilities, flag short fits
- refactor!: remove the factor-graph surface nothing used, add try_winner
### Bug Fixes
- fix(test): the ingestion-order property was comparing two truncated fits
### Documentation
- docs: record that the event log is the source of truth, and why
### Features
- feat: add UnknownKeys::Prior, and explain why there is no Skip
- feat: add History::posterior_of for a linear combination of competitors
- feat: add History::predict_margin for scored matchups
- feat: add expected_variance_reduction for scored active learning
### Miscellaneous Tasks
- chore: Release trueskill-tt version 0.5.0
### Styling
- style: factor the event-pair type out of the reconvergence fixture
- style: use arrays rather than vec! in the calibration fixture
### Testing
- test: pin that re-convergence is path-independent
- test: calibrate the marginals against the exact posterior
- test: pin what an additive model does to combined uncertainty
## 0.4.2 - 2026-09-07
### Bug Fixes
@@ -9,6 +105,10 @@ All notable changes to this project will be documented in this file.
- fix: replace the erfc approximation with libm, for free
- fix: route every transcendental through libm, and combine sigmas with hypot
### Miscellaneous Tasks
- chore: Release trueskill-tt version 0.4.2
### Testing
- test: localise the erfc_inv tail residual to the caller's argument
+19 -5
View File
@@ -24,6 +24,20 @@ is where several defects have hidden — a debug-only run is not evidence.
- `approx``approx::AbsDiffEq` etc. for `Gaussian`. Most numerical goldens need it.
- `rayon` — opt-in parallel within-slice sweep and per-slice query passes.
## Working rules
- **Investigate before implementing.** Measure the actual behaviour first —
against an analytic reference where one exists. Several "obvious" fixes in
this repo turned out to be wrong in sign or unnecessary, and the measurement
is what caught them.
- **Fix the root issue, not the symptom.** A clamp that hides an underflow, or
a tolerance loosened to make a test pass, is a defect deferred.
- **Scout crates.io before hand-rolling numerics.** Check accuracy against an
independent reference rather than trusting downloads: `puruspe` has 1.4M
downloads and is 346 ULP off in the tail, where `libm` is 1. Fewer
dependencies is preferable, not mandatory — take the dependency when it is
measurably better.
## Architecture
A Rust port of [TrueSkillThroughTime.py](https://github.com/glandfried/TrueSkillThroughTime.py):
@@ -66,11 +80,11 @@ History → TimeSlice[] → Event[] → Item[]
`tau = mu/sigma²`). `Mul`/`Div` are the EP product/cavity: pure adds and
subtracts. Variance-space ops (`Add`, `Sub`, `exclude`, `forget`) go through
`from_mv`/`variance()` and take no square root.
- **`factor/`** — `TeamSumFactor`, `RankDiffFactor`, `TruncFactor` (ranked),
`MarginFactor` (scored), over a flat `VarStore`. `BuiltinFactor` dispatches
by enum rather than `dyn`.
- **`Schedule`** (`schedule.rs`) — drives factor propagation. `EpsilonOrMax` is
the only implementation.
- **`factor/`** — `TruncFactor` (ranked) and `MarginFactor` (scored) over a
flat `VarStore`. `Game::run_chain` drives them directly through a local
`DiffFactor` enum; there is no `Schedule` indirection and no generic `Factor`
trait. Both were removed once measurement showed nothing had ever used them
— see #42.
- **`Competitor`** (`competitor.rs`) — per-history temporal state (`message`,
`last_time`). **`Rating`** (`rating.rs`) — static config (prior, `beta`, drift).
- **`storage/`** — `SkillStore` (per slice, `pub(crate)`) and `CompetitorStore`
+8 -5
View File
@@ -1,6 +1,6 @@
[package]
name = "trueskill-tt"
version = "0.4.2"
version = "0.8.0"
edition = "2024"
rust-version = "1.85"
description = "TrueSkill Through Time: Bayesian skill rating that tracks how skill evolves over time, via Gaussian message passing"
@@ -33,10 +33,6 @@ bench = false
name = "batch"
harness = false
[[bench]]
name = "gaussian"
harness = false
[[bench]]
name = "history_converge"
harness = false
@@ -51,12 +47,15 @@ harness = false
[dependencies]
approx = { version = "0.5.1", optional = true }
feral-amd = "0.2"
libm = "0.2.16"
rayon = { version = "1", optional = true }
smallvec = "1"
[features]
approx = ["dep:approx"]
# Exposes the joint sparsity pattern for the #52 measurement. Test-only.
measure-sparsity = []
rayon = ["dep:rayon"]
[dev-dependencies]
@@ -79,3 +78,7 @@ debug = true
[profile.dev]
debug = true
[[bench]]
name = "joint"
harness = false
+211 -34
View File
@@ -1,15 +1,142 @@
# TrueSkill - Through Time
Rust port of [TrueSkillThroughTime.py](https://github.com/glandfried/TrueSkillThroughTime.py).
Bayesian skill rating over a time axis.
## Other implementations
Where plain TrueSkill gives each competitor one running estimate, TrueSkill
Through Time treats a whole history as a single model and infers skill *at every
point in time*. Evidence flows both directions: a result today sharpens the
estimate of who someone was last year, so early estimates stop being frozen
guesses and comparisons across eras become meaningful.
- [ttt-scala](https://github.com/ankurdave/ttt-scala)
- [ChessAnalysis #F](https://github.com/lucasmaystre/ChessAnalysis)
- [TrueSkillThroughTime.jl](https://github.com/glandfried/TrueSkillThroughTime.jl)
- [TrueSkillThroughTime.R](https://github.com/glandfried/TrueSkillThroughTime.R)
- [TrueSkill Through Time: Revisiting the History of Chess](https://www.microsoft.com/en-us/research/wp-content/uploads/2008/01/NIPS2007_0931.pdf)
- [TrueSkill Through Time. The full scientific documentation](https://glandfried.github.io/publication/landfried2021-learning/)
A Rust port of
[TrueSkillThroughTime.py](https://github.com/glandfried/TrueSkillThroughTime.py).
## Install
```toml
[dependencies]
trueskill-tt = "0.8"
```
Optional features, both off by default:
- `approx``approx`'s equality traits for `Gaussian`. Useful in tests.
- `rayon` — parallelises the within-slice sweep and the per-slice passes of
`learning_curves` / `log_evidence`. Results stay bit-identical regardless of
worker count; `just determinism` asserts it at 1, 2, 4 and 8 threads.
## Quickstart
Record results, converge, then read off skills.
```rust
use trueskill_tt::History;
let mut history = History::default();
history.record_winner(&"alice", &"bob", 1)?;
history.record_winner(&"bob", &"carol", 2)?;
history.record_winner(&"alice", &"carol", 3)?;
history.converge()?;
let alice = history.current_skill("alice").unwrap();
assert!(alice.mu() > 0.0, "alice won every game she played");
# Ok::<(), trueskill_tt::InferenceError>(())
```
The third argument is the time. It is what makes this Through Time rather than
plain TrueSkill: skill is inferred at each of those moments, not once at the
end. `learning_curve` reads the whole trajectory back.
```rust
# use trueskill_tt::History;
# let mut history = History::default();
# history.record_winner(&"alice", &"bob", 1)?;
# history.record_winner(&"bob", &"carol", 2)?;
# history.record_winner(&"alice", &"carol", 3)?;
# history.converge()?;
// `None` means the key is unknown; `Some(vec![])` means known but unplayed.
let curve = history.learning_curve("alice").unwrap();
for (time, skill) in &curve {
println!("t={time}: {:.2} ± {:.2}", skill.mu(), skill.sigma());
}
// Everyone's latest posterior in one pass — the leaderboard query.
let latest = history.current_skills();
assert_eq!(latest.len(), 3);
# Ok::<(), trueskill_tt::InferenceError>(())
```
## Teams, rankings and draws
Anything beyond one-versus-one goes through the fluent event builder. An event
is only recorded by the terminal `.commit()`.
```rust
use trueskill_tt::History;
let mut history = History::builder().p_draw(0.1).build();
history
.event(1)
.team(["alice", "bob"])
.team(["carol", "dave"])
.ranking([0, 1]) // lower is better; equal values are a tie
.commit()?;
history.converge()?;
# Ok::<(), trueskill_tt::InferenceError>(())
```
**A tie needs a positive `p_draw`.** A `p_draw` of zero asserts draws cannot
happen, so a tied result has no representable likelihood and is rejected rather
than fitted to something else:
```rust
use trueskill_tt::{History, InferenceError};
let mut history = History::default(); // p_draw defaults to 0.0
let err = history.record_draw(&"alice", &"bob", 1).unwrap_err();
assert!(matches!(err, InferenceError::TieWithoutDrawProbability { .. }));
```
This also catches `Outcome::winner(w, n)` for three or more teams, which ties
every loser.
## Which entry point?
| You want to | Use |
|---|---|
| One match, two competitors | `record_winner` / `record_draw` |
| Teams, explicit ranks, scores, per-member weights | `history.event(t)…commit()` |
| A batch you already have as values | `add_events(iter)` |
| Score a hypothetical with no history at all | `Game` |
`Game` is the odd one out and worth being explicit about: it is a single match's
factor graph, it does not participate in a `History`, and nothing it computes is
remembered. Reach for it to evaluate a matchup in isolation; reach for `History`
for everything that accumulates.
## `converge` is strict
`converge` returns `Err(NotConverged)` if the sweep hits `max_iter` with the
step still above `epsilon`, and `Err(NonFiniteStep)` if a sweep produces NaN.
It used to return `Ok` with `converged: false`, which was the worst available
shape. A fit that stops short is *wrong by a little*: every posterior is finite,
the ordering looks sensible, and nothing about the output says the numbers were
still moving. Detection was opt-in, and `let _ = h.converge()` silently opted
out — which is how a real defect hid in this crate's own test suite.
The default `max_iter` is high enough that reaching it means something is
genuinely wrong rather than that the history is large; the loop exits at
`epsilon` long before, so raising the cap costs nothing when it is not needed.
Use `converge_partial` when a deliberately capped, unconverged fit is the point.
Predictions are strict for the same reason: every `predict_*` method reads
skills through one gate that refuses a NaN-poisoned fit, rather than returning a
plausible number computed from it.
## Drift
@@ -45,7 +172,7 @@ grows proportionally to time:
variance_delta = elapsed * γ²
```
This is the standard TrueSkill Through Time model. Pass a `ConstantDrift(gamma)`
This is the standard TrueSkill Through Time model. Pass a `ConstantDrift::new(gamma)`
when constructing a `Rating`:
```rust
@@ -53,9 +180,9 @@ use trueskill_tt::{ConstantDrift, Gaussian, Rating};
// gamma = 0.1 means skill can shift ~0.1 per time unit.
let rating: Rating<i64, ConstantDrift> =
Rating::new(Gaussian::from_ms(0.0, 6.0), 1.0, ConstantDrift(0.1));
Rating::new(Gaussian::from_ms(0.0, 6.0), 1.0, ConstantDrift::new(0.1));
assert_eq!(rating.drift().0, 0.1);
assert_eq!(rating.drift().gamma(), 0.1);
```
The type annotation is load-bearing: `ConstantDrift` implements `Drift<T>` for
@@ -98,14 +225,14 @@ assert_eq!(history.log_evidence(), 0.0);
```
`HistoryBuilder::drift` is the only way to set a history's drift model; there is
no `gamma()` shorthand. The default is `ConstantDrift(GAMMA)`.
no `gamma()` shorthand. The default is `ConstantDrift::new(GAMMA)`.
### Per-competitor drift
A `History` has one drift model, but individual competitors can scale it.
`Member::with_drift_scale(s)` multiplies the drift *variance* that competitor
accumulates, so `s` is in the same units as `gamma`: `ConstantDrift(g)` at
scale `s` behaves exactly as `ConstantDrift(g * s)` would, for that competitor
accumulates, so `s` is in the same units as `gamma`: `ConstantDrift::new(g)` at
scale `s` behaves exactly as `ConstantDrift::new(g * s)` would, for that competitor
alone.
`0.0` pins a competitor still. That is what makes a **fixed reference point**
@@ -115,7 +242,7 @@ strength, a rating floor, a course difficulty:
```rust
use trueskill_tt::{ConstantDrift, Event, History, Member, Outcome, Team};
let mut h = History::builder().drift(ConstantDrift(0.1)).build();
let mut h = History::builder().drift(ConstantDrift::new(0.1)).build();
h.add_events(vec![Event {
time: 0,
@@ -134,14 +261,20 @@ h.add_events(vec![Event {
h.converge().unwrap();
```
Like `with_prior`, the scale is **competitor configuration captured at first
appearance** — setting it on a key the history already knows has no effect. It
must be finite and non-negative; ingestion otherwise fails with
`InferenceError::InvalidParameter`.
Like `with_prior`, the scale is **competitor configuration, not a per-event
value**: it applies to the competitor for the whole history, and it applies
whenever it is supplied — including on a key the history already knows.
Configuring one late still refits the whole history rather than taking effect
only from that event onward, because `converge` refits from competitor state.
Repeating the same value is inert; supplying two *different* values for one
competitor within a single batch is `InferenceError::ConflictingCompetitorConfig`,
since events in a batch have no order. The scale must be finite and
non-negative; ingestion otherwise fails with `InferenceError::InvalidParameter`.
Note that the fluent `EventBuilder` (`h.event(t).team([...])`) sets weights but
not `drift_scale` or `prior`; those need the typed `Event` / `Team` / `Member`
shape shown above.
The fluent `EventBuilder` reaches this too: `.team([...])` is the common case
and leaves both unset, while `.members([...])` takes `Member` values directly,
so `h.event(t).members([Member::new("layout_7").with_drift_scale(0.0)])` is
equivalent to the typed shape above.
## Scored outcomes
@@ -195,8 +328,47 @@ stay available at any size:
quadratic in team count.
- `predict_ranking(teams, ranks)` — one specific finishing order.
Unknown keys are an error, not a silent omission: a team the history has never
seen cannot produce a confident-looking probability.
Unknown keys are an error by default, not a silent omission: a team the history
has never seen cannot produce a confident-looking probability. The error names
the key, and every key must already be known — pre-filter with
`current_skill` if your caller cannot guarantee that.
If predicting for competitors you have never seen is the point rather than a
mistake, say so once:
```rust
use trueskill_tt::{History, UnknownKeys};
let h = History::builder().unknown_keys(UnknownKeys::Prior).build();
```
An unknown competitor is then answered from the configured prior, which is the
honest reading — you have no evidence about them — and correctly *widens* a team
that contains one. There is deliberately no "skip the member" mode: a team's
performance is the sum of its members, so dropping one would make the model more
certain because it knows less.
### Asking about one competitor
`Gaussian` answers tail questions directly, which is what a stopping rule needs:
```rust
use trueskill_tt::History;
let mut h = History::default();
h.record_winner(&"alice", &"bob", 1).unwrap();
h.converge().unwrap();
let skill = h.current_skill("alice").unwrap();
// "How sure am I that this is below the cutoff?" — a probability, not a
// `mu + z * sigma` band whose confidence drifts as sigma changes.
let _ = skill.probability_below(20.0);
// Use this rather than `1.0 - probability_below(x)`: the complement cancels
// away every digit in the upper tail, which is where a stopping rule lives.
let _ = skill.probability_above(30.0);
```
## Which match to play next
@@ -209,7 +381,7 @@ what you believe now and what you would believe afterwards.
```rust
use trueskill_tt::History;
let mut h = History::builder().build();
let mut h = History::default();
for t in 1..=10 {
h.record_winner(&"veteran", &"regular", t).unwrap();
h.record_winner(&"regular", &"veteran", t + 100).unwrap();
@@ -233,16 +405,21 @@ expensive than `quality()`. Scoring every pairing among `n` competitors is
`O(n² × outcomes)` passes — shortlist with `quality()` or
`predict_win_probabilities` first, then score only the shortlist.
## Todo
## Other implementations
- [x] Implement approx for Gaussian
- [x] Add more tests from `TrueSkillThroughTime.jl`
- [x] Generalise a time axis — `Time` is now a trait (`Untimed`, `i64`), not an enum
- [x] Add examples (`examples/atp.rs`, `examples/scored.rs`)
- [x] Add Observer (`Observer` / `NullObserver`)
- [x] Benchmark the inference loop (`benches/batch.rs`, `benches/history_converge.rs`, `benches/ingest.rs`)
- [x] N-team `predict_outcome` with draw mass, and `expected_information_gain`
- [x] Cross-check `quality()` against [sublee/trueskill](https://github.com/sublee/trueskill/tree/master) — N identical teams follow the closed form `(1/5)^((n-1)/2)` for the conventional parameters, asserted for n = 2..10, and the n=3/n=5 values (0.200, 0.040) match the reference package
- [ttt-scala](https://github.com/ankurdave/ttt-scala)
- [ChessAnalysis #F](https://github.com/lucasmaystre/ChessAnalysis)
- [TrueSkillThroughTime.jl](https://github.com/glandfried/TrueSkillThroughTime.jl)
- [TrueSkillThroughTime.R](https://github.com/glandfried/TrueSkillThroughTime.R)
- [TrueSkill Through Time: Revisiting the History of Chess](https://www.microsoft.com/en-us/research/wp-content/uploads/2008/01/NIPS2007_0931.pdf)
- [TrueSkill Through Time. The full scientific documentation](https://glandfried.github.io/publication/landfried2021-learning/)
## Status
Every box on the old todo list is ticked, so it has been retired; open work
lives in the issue tracker instead. The crate is in use and the API is still
moving — breaking changes are batched into minor releases rather than dribbled
out, and `CHANGELOG.md` records them.
## License
+45 -35
View File
@@ -1,45 +1,55 @@
//! One slice's event sweep.
//!
//! Written against the public API rather than against `TimeSlice` directly.
//! It used to reach for `TimeSlice`, `KeyTable`, `CompetitorStore`,
//! `Competitor` and `EventKind`, and was the *only* thing outside `src/`
//! that did — so a benchmark was dictating five public types that no test,
//! example or consumer could otherwise obtain.
//!
//! A single-slice history's `converge` calls exactly the same per-slice sweep,
//! so capping at one iteration measures the same code path.
use criterion::{Criterion, criterion_group, criterion_main};
use trueskill_tt::{
BETA, Competitor, ConvergenceOptions, EventKind, GAMMA, KeyTable, MU, P_DRAW, Rating, SIGMA,
TimeSlice, drift::ConstantDrift, gaussian::Gaussian, storage::CompetitorStore,
};
use smallvec::smallvec;
use trueskill_tt::{ConstantDrift, ConvergenceOptions, Event, History, Member, Outcome, Team};
fn criterion_benchmark(criterion: &mut Criterion) {
let mut index_map = KeyTable::new();
let build = || {
let mut h = History::builder()
.convergence(ConvergenceOptions {
max_iter: 1,
epsilon: 0.0,
alpha: 1.0,
})
.drift(ConstantDrift::new(0.0))
.build();
let a = index_map.get_or_create("a");
let b = index_map.get_or_create("b");
let c = index_map.get_or_create("c");
// 100 events, all at one time, so the history has a single slice.
let events: Vec<Event<i64, &'static str>> = (0..100)
.map(|_| Event {
time: 1,
teams: smallvec![
Team::with_members([Member::new("a")]),
Team::with_members([Member::new("b")]),
],
outcome: Outcome::winner(0, 2),
})
.collect();
h.add_events(events).expect("fixture ingests");
h
};
let mut agents: CompetitorStore<i64, ConstantDrift> = CompetitorStore::new();
for agent in [a, b, c] {
agents.insert(
agent,
Competitor {
rating: Rating::new(Gaussian::from_ms(MU, SIGMA), BETA, ConstantDrift(GAMMA)),
..Default::default()
criterion.bench_function("slice_sweep_100_events", |b| {
b.iter_batched(
build,
|mut h| {
// `converge_partial`, not `converge`: one iteration is
// deliberately short of convergence and `converge` reports that
// as an error.
let _ = h.converge_partial();
},
criterion::BatchSize::SmallInput,
);
}
let mut composition = Vec::new();
let mut results = Vec::new();
let mut weights = Vec::new();
for _ in 0..100 {
composition.push(vec![vec![a], vec![b]]);
results.push(vec![1.0, 0.0]);
weights.push(vec![vec![1.0], vec![1.0]]);
}
let kinds = vec![EventKind::Ranked; composition.len()];
let mut time_slice = TimeSlice::new(1, P_DRAW, ConvergenceOptions::default());
time_slice.add_events(composition, Some(results), Some(weights), kinds, &agents);
criterion.bench_function("Batch::iteration", |b| {
b.iter(|| time_slice.iteration(0, &agents))
});
}
-53
View File
@@ -1,53 +0,0 @@
use criterion::{Criterion, criterion_group, criterion_main};
use trueskill_tt::gaussian::Gaussian;
fn benchmark_gaussian_arithmetic(criterion: &mut Criterion) {
// Define test Gaussians
let g1 = Gaussian::from_ms(25.0, 25.0 / 3.0);
let g2 = Gaussian::from_ms(0.0, 1.0);
let g3 = Gaussian::from_ms(1.0, 1.0);
// Benchmark addition
criterion.bench_function("Gaussian::add", |bencher| {
bencher.iter(|| g1 + g2);
});
// Benchmark subtraction
criterion.bench_function("Gaussian::sub", |bencher| {
bencher.iter(|| g1 - g3);
});
// Benchmark multiplication
criterion.bench_function("Gaussian::mul", |bencher| {
bencher.iter(|| g1 * g2);
});
// Benchmark division
// NOTE: numerator must have higher precision (smaller sigma) than the
// denominator in this representation; g2 (sigma=1) / g1 (sigma=8.33) is
// well-defined, whereas g1 / g2 underflows and panics in mu_sigma.
criterion.bench_function("Gaussian::div", |bencher| {
bencher.iter(|| g2 / g1);
});
// Benchmark natural parameter conversions
criterion.bench_function("Gaussian::pi", |bencher| {
bencher.iter(|| g1.pi());
});
criterion.bench_function("Gaussian::tau", |bencher| {
bencher.iter(|| g1.tau());
});
// Benchmark combined pi/tau operations (used in mul/div)
criterion.bench_function("Gaussian::pi_tau_combined", |bencher| {
bencher.iter(|| {
let pi = g1.pi();
let tau = g1.tau();
(pi, tau)
});
});
}
criterion_group!(benches, benchmark_gaussian_arithmetic);
criterion_main!(benches);
+8 -9
View File
@@ -25,16 +25,14 @@
use criterion::{BatchSize, Criterion, criterion_group, criterion_main};
use smallvec::smallvec;
use trueskill_tt::{
ConstantDrift, ConvergenceOptions, Event, History, Member, NullObserver, Outcome, Team,
};
use trueskill_tt::{ConstantDrift, ConvergenceOptions, Event, History, Member, Outcome, Team};
fn build_history_1v1(
n_events: usize,
n_competitors: usize,
events_per_slice: usize,
seed: u64,
) -> History<i64, ConstantDrift, NullObserver, String> {
) -> History<String> {
let mut rng = seed;
let mut next = || {
rng = rng
@@ -43,11 +41,12 @@ fn build_history_1v1(
rng
};
let mut h = History::<i64, _, _, String>::builder_with_key()
let mut h = History::builder()
.key_type::<String>()
.mu(25.0)
.sigma(25.0 / 3.0)
.beta(25.0 / 6.0)
.drift(ConstantDrift(25.0 / 300.0))
.drift(ConstantDrift::new(25.0 / 300.0))
.convergence(ConvergenceOptions {
max_iter: 30,
epsilon: 1e-6,
@@ -82,7 +81,7 @@ fn bench_converge(c: &mut Criterion) {
b.iter_batched(
|| build_history_1v1(500, 100, 10, 42),
|mut h| {
h.converge().unwrap();
let _ = h.converge().unwrap();
},
BatchSize::SmallInput,
);
@@ -92,7 +91,7 @@ fn bench_converge(c: &mut Criterion) {
b.iter_batched(
|| build_history_1v1(2000, 200, 20, 42),
|mut h| {
h.converge().unwrap();
let _ = h.converge().unwrap();
},
BatchSize::SmallInput,
);
@@ -106,7 +105,7 @@ fn bench_converge(c: &mut Criterion) {
b.iter_batched(
|| build_history_1v1(5000, 50000, 5000, 42),
|mut h| {
h.converge().unwrap();
let _ = h.converge().unwrap();
},
BatchSize::SmallInput,
);
+2 -2
View File
@@ -32,7 +32,7 @@ fn bench_ingest(c: &mut Criterion) {
b.iter_batched(
|| events(n, 0),
|evs| {
let mut h: History<i64, _, _, String> = History::builder_with_key().build();
let mut h: History<String> = History::builder().key_type::<String>().build();
for ev in evs {
h.add_events(std::iter::once(ev)).unwrap();
}
@@ -46,7 +46,7 @@ fn bench_ingest(c: &mut Criterion) {
b.iter_batched(
|| events(n, 0),
|evs| {
let mut h: History<i64, _, _, String> = History::builder_with_key().build();
let mut h: History<String> = History::builder().key_type::<String>().build();
h.add_events(evs).unwrap();
black_box(h.time_slices_len())
},
+81
View File
@@ -0,0 +1,81 @@
//! Cost of the joint posterior: factorising versus querying.
//!
//! The split is the whole point of `History::joint`. Factorising is `O(n^3)` in
//! the history's appearances and depends only on the fit; a query is `O(n^2)`
//! and depends only on the question. `posterior_of_one_shot` pays both every
//! time, `joint_query` pays only the second.
use criterion::{Criterion, criterion_group, criterion_main};
use smallvec::smallvec;
use trueskill_tt::{ConstantDrift, ConvergenceOptions, Event, History, Member, Outcome, Team};
/// 30 slices of 8 duels: 480 appearances over 100 competitors.
fn fitted() -> History<String> {
let mut h: History<String> = History::builder()
.key_type::<String>()
.mu(0.0)
.sigma(6.0)
.beta(1.0)
.score_sigma(2.0)
.drift(ConstantDrift::new(0.05))
// `max_iter: 30` was here, and this fixture needs more: `converge`
// reported `NotConverged { iterations: 30, final_step: (4.5e-4, 0.0) }`
// once it stopped returning short fits silently. The benchmark measures
// the factorisation, whose cost depends on the fit's *shape* rather
// than its exactness — but measuring it on an unconverged fit is still
// measuring something nobody would run.
.convergence(ConvergenceOptions {
max_iter: trueskill_tt::ITERATIONS,
epsilon: 1e-10,
alpha: 1.0,
})
.build();
let mut events: Vec<Event<i64, String>> = Vec::new();
let mut k = 0usize;
for t in 0..30i64 {
for _ in 0..8 {
k += 1;
events.push(Event {
time: t,
teams: smallvec![
Team::with_members([Member::new(format!("p{}", k % 100))]),
Team::with_members([Member::new(format!("p{}", (k + 37) % 100))]),
],
outcome: Outcome::scores([
(k as f64 * 0.3).sin().abs() * 20.0,
(k as f64 * 0.3).cos().abs() * 20.0,
]),
});
}
}
h.add_events(events).unwrap();
let _ = h.converge().unwrap();
h
}
fn bench_joint(c: &mut Criterion) {
let h = fitted();
let a = "p0".to_string();
let b = "p1".to_string();
let terms = [(&a, 1.0), (&b, -1.0)];
c.bench_function("joint_factorise_480_appearances", |bencher| {
bencher.iter(|| std::hint::black_box(h.joint().unwrap().variables()));
});
// Factorise-and-query, the cost the deleted `History::posterior_of`
// wrapper paid on every call. Kept as the baseline the cached query below
// is measured against.
c.bench_function("posterior_of_one_shot_480_appearances", |bencher| {
bencher.iter(|| std::hint::black_box(h.joint().unwrap().posterior_of(&terms).unwrap()));
});
let joint = h.joint().unwrap();
c.bench_function("joint_query_480_appearances", |bencher| {
bencher.iter(|| std::hint::black_box(joint.posterior_of(&terms).unwrap()));
});
}
criterion_group!(benches, bench_joint);
criterion_main!(benches);
+4 -3
View File
@@ -5,11 +5,12 @@ use trueskill_tt::{ConstantDrift, Event, History, Member, Outcome, Team};
fn bench_scored_history(c: &mut Criterion) {
c.bench_function("scored_history_60_events_30_iter", |bencher| {
bencher.iter(|| {
let mut h: History<i64, ConstantDrift, _, String> = History::builder_with_key()
let mut h: History<String> = History::builder()
.key_type::<String>()
.mu(25.0)
.sigma(25.0 / 3.0)
.beta(25.0 / 6.0)
.drift(ConstantDrift(0.03))
.drift(ConstantDrift::new(0.03))
.score_sigma(2.0)
.build();
@@ -29,7 +30,7 @@ fn bench_scored_history(c: &mut Criterion) {
});
}
h.add_events(events).unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
});
});
}
@@ -500,6 +500,26 @@ All public traits (`Time`, `Drift`, `Observer`, `Factor`, `Schedule`) require `S
`rayon` as default-on feature; with `default-features = false`, parallel paths fall back to sequential iterators behind `cfg(feature = "rayon")`.
> **Not implemented. Deliberate deviation, decided 2026-09-08 (issue #5).**
>
> `rayon` ships **opt-in**: `Cargo.toml` has no `default = [...]` key. The
> measured speedups are 1.0x on realistic workloads and 1.3x on a pathological
> one (issue #4), because typical slices hold too few events to amortize
> rayon's task-spawn overhead. Default-on would hand every downstream user a
> thread pool and a dependency for approximately no gain.
>
> This section made the trade conditional on cross-slice dirty-bit skipping
> landing and changing the parallel story. It did not land: #4 was closed on
> 2026-08-27 by removing the inert `ConvergenceReport::slices_skipped` field
> rather than by implementing the mechanism, so the re-measurement this was
> waiting on will not arrive.
>
> The "Trade-offs" note below also cited an `unsafe` concurrent-write path
> through `SkillStore` as a cost of default-on. That cost does not exist: the
> crate is `#![forbid(unsafe_code)]`, and the compute/apply split on the
> internal `Event` is what lets a color group run in parallel without it. The
> case for opt-in rests on the measurements alone.
### Expected speedup ballpark
For 1000 players, 60 events/slice × 1000 slices, 30 convergence iterations:
@@ -521,7 +541,7 @@ These are pre-implementation estimates. Each tier validates with criterion.
- Color-group parallelism requires up-front graph coloring at ingestion. Cost: linear in events, run once per `add_events`. Cheap.
- Default = asynchronous EP (preserves current semantics). Synchronous opt-in only.
- Cross-slice sweep stays sequential; no speculative parallel sweeps.
- Rayon default-on but feature-gated.
- Rayon default-on but feature-gated. **Superseded — shipped opt-in; see the deviation note in Section 6.**
### Open question
+29 -8
View File
@@ -1,7 +1,8 @@
use plotters::prelude::*;
use smallvec::smallvec;
use time::{Date, Month};
use trueskill_tt::{Event, History, Member, Outcome, Team, drift::ConstantDrift};
use trueskill_tt::{
Event, History, Member, Outcome, Team, drift::ConstantDrift, smallvec::smallvec,
};
fn main() {
let mut csv = csv::Reader::open("examples/atp.csv").unwrap();
@@ -42,18 +43,38 @@ fn main() {
}
}
let mut hist: History<i64, _, _, String> = History::builder_with_key()
let mut hist: History<String> = History::builder()
.key_type::<String>()
.sigma(1.6)
.drift(ConstantDrift(0.036))
.drift(ConstantDrift::new(0.036))
.convergence(trueskill_tt::ConvergenceOptions {
max_iter: 10,
// This history needs 30 sweeps to reach the epsilon below. It was
// capped at 10 until the `#[must_use]` on `ConvergenceReport`
// surfaced that the example had been shipping a short fit.
max_iter: 100,
epsilon: 0.01,
alpha: 1.0,
})
.build();
hist.add_events(events).unwrap();
hist.converge().unwrap();
// Read the report rather than discarding it. A fit that hits `max_iter`
// without reaching `epsilon` is not an error and does not look wrong — every
// rating comes back finite and sensibly ordered — so this flag is the only
// thing that says the numbers were still moving when the sweep stopped.
let report = hist.converge().unwrap();
eprintln!(
"converged={} after {} sweeps, final step {:?}",
report.converged, report.iterations, report.final_step
);
if !report.converged {
eprintln!(
"warning: stopped after {} sweeps with a final step of {:?}, \
short of epsilon — raise ConvergenceOptions::max_iter",
report.iterations, report.final_step
);
}
let players = [
("aggasi", "a092", 38800i64),
@@ -77,7 +98,7 @@ fn main() {
let mut y_spec = (f64::MAX, f64::MIN);
for &(_, id, cutoff) in &players {
for (ts, gs) in hist.learning_curve(id) {
for (ts, gs) in hist.learning_curve(id).unwrap() {
if ts >= cutoff {
continue;
}
@@ -123,7 +144,7 @@ fn main() {
let mut upper = Vec::new();
let mut lower = Vec::new();
for (ts, gs) in hist.learning_curve(id) {
for (ts, gs) in hist.learning_curve(id).unwrap() {
if ts >= cutoff {
continue;
}
+2 -3
View File
@@ -6,15 +6,14 @@
//!
//! Run with: `cargo run --example scored --release`
use smallvec::smallvec;
use trueskill_tt::{ConstantDrift, Event, History, Member, Outcome, Team};
use trueskill_tt::{ConstantDrift, Event, History, Member, Outcome, Team, smallvec::smallvec};
fn main() {
let mut h = History::builder()
.mu(25.0)
.sigma(25.0 / 3.0)
.beta(25.0 / 6.0)
.drift(ConstantDrift(0.03))
.drift(ConstantDrift::new(0.03))
.score_sigma(2.0) // tune to data; smaller = trust margins more
.build();
+50 -7
View File
@@ -47,7 +47,38 @@ fn kl_divergence(q: Gaussian, p: Gaussian) -> f64 {
}
let mean_gap = q.mu() - p.mu();
0.5 * (libm::log(var_p / var_q) + (var_q + mean_gap * mean_gap) / var_p - 1.0)
// Algebraically `0.5 * (ln(var_p/var_q) + (var_q + gap^2)/var_p - 1)`, but
// written so that neither term can go negative.
//
// The direct form cancels against its `- 1.0` for two near-identical
// distributions and returns a *negative* divergence — measured, 762 082 of
// 3 000 000 near-identical pairs, worst `-5.55e-17`, which is exactly one
// ULP of the 1.0. It also loses the answer entirely where it is small:
// at `var_q/var_p - 1 = 1e-9` the direct form gives `0.0` where the true
// value is `2.5e-19`.
//
// With `u = var_q/var_p - 1` the variance part is `0.5 * (u - ln(1+u))`,
// which is non-negative for every `u > -1`, and the mean part is a square
// over a positive variance. Non-negativity is then structural rather than
// incidental.
let u = var_q / var_p - 1.0;
0.5 * u_minus_ln1p(u) + mean_gap * mean_gap / (2.0 * var_p)
}
/// `u - ln(1 + u)`, without the cancellation that spelling invites.
///
/// Both terms are approximately `u` for small `u`, so the subtraction loses
/// everything just where the result matters. The Taylor series
/// `u^2/2 - u^3/3 + u^4/4 - ...` is exact in that regime and manifestly
/// non-negative, since `u^2/2` dominates.
fn u_minus_ln1p(u: f64) -> f64 {
if u.abs() < 1e-4 {
let u2 = u * u;
u2 * (0.5 - u / 3.0 + u2 / 4.0)
} else {
u - libm::log1p(u)
}
}
/// Expected information gain of a hypothetical matchup, in nats.
@@ -92,7 +123,11 @@ fn kl_divergence(q: Gaussian, p: Gaussian) -> f64 {
/// - `EmptyTeam` if any team has no members.
/// - `TooManyTeams` if the outcome space is too large to enumerate; see
/// [`MAX_PREDICTED_TEAMS`](crate::MAX_PREDICTED_TEAMS).
/// - `InvalidProbability` if `options.p_draw` is outside `[0.0, 1.0)`.
/// - `InvalidParameter` for a `p_draw` outside `[0.0, 1.0)`.
/// - `GridTooCoarse` when the performance sigmas are too far apart to
/// integrate on one grid. This comes from `outcome_distribution`, which runs
/// before any inference — so it is not covered by "anything `Game::ranked`
/// returns" below.
/// - Anything [`Game::ranked`](crate::Game::ranked) returns for a hypothetical
/// outcome.
pub fn expected_information_gain<T: Time, D: Drift<T>>(
@@ -109,7 +144,8 @@ pub fn expected_information_gain<T: Time, D: Drift<T>>(
});
}
if !(0.0..1.0).contains(&options.p_draw) {
return Err(InferenceError::InvalidProbability {
return Err(InferenceError::InvalidParameter {
parameter: crate::Parameter::PDraw,
value: options.p_draw,
});
}
@@ -124,7 +160,7 @@ pub fn expected_information_gain<T: Time, D: Drift<T>>(
.iter()
.map(|team| {
team.iter()
.fold(crate::N00, |acc, rating| acc + rating.performance())
.fold(crate::N00, |acc, rating| acc.convolve(rating.performance()))
})
.collect();
@@ -146,7 +182,7 @@ pub fn expected_information_gain<T: Time, D: Drift<T>>(
let mut gain = 0.0;
for (ranks, probability) in predict::outcome_distribution(&performances, &margins) {
for (ranks, probability) in predict::outcome_distribution(&performances, &margins)? {
if probability <= NEGLIGIBLE {
continue;
}
@@ -177,7 +213,11 @@ mod tests {
type R = Rating<i64, ConstantDrift>;
fn rating(mu: f64, sigma: f64) -> R {
R::new(Gaussian::from_ms(mu, sigma), BETA, ConstantDrift(GAMMA))
R::new(
Gaussian::from_ms(mu, sigma),
BETA,
ConstantDrift::new(GAMMA),
)
}
fn options(p_draw: f64) -> GameOptions {
@@ -328,7 +368,10 @@ mod tests {
));
assert!(matches!(
expected_information_gain(&[&a, &a], &options(1.5)),
Err(InferenceError::InvalidProbability { .. })
Err(InferenceError::InvalidParameter {
parameter: crate::Parameter::PDraw,
..
})
));
}
+118
View File
@@ -191,3 +191,121 @@ mod tests {
assert_eq!(cg.total_events(), 4);
}
}
#[cfg(test)]
mod properties {
use std::collections::HashSet;
use proptest::prelude::*;
use super::*;
/// The property the whole parallel sweep rests on: two events sharing a
/// competitor must never land in the same color, because a color group is
/// run concurrently and two events touching one competitor would race.
///
/// Hand-written cases cover the shapes someone thought of. This covers the
/// ones nobody did — the correctness of `sweep_color_groups` depends on it
/// holding for every input, not for five.
fn check(events: &[Vec<usize>]) {
let groups = color_greedy(events.len(), |ev| {
events[ev]
.iter()
.copied()
.map(Index::from)
.collect::<Vec<_>>()
});
// Disjointness *between events* within a color. Deduplicated per
// event, because one event legitimately naming a competitor twice is
// not a collision — `color_greedy` collects each event's members into
// a set for exactly that reason.
for color in 0..groups.n_colors() {
let mut seen: HashSet<usize> = HashSet::new();
for &ev in &groups.groups[color] {
let members: HashSet<usize> = events[ev].iter().copied().collect();
for competitor in members {
assert!(
seen.insert(competitor),
"competitor {competitor} shared by two events in color {color}"
);
}
}
}
// Every event is assigned exactly once. Without this, a partition that
// dropped events would satisfy disjointness trivially.
let mut assigned: Vec<usize> = groups.groups.iter().flatten().copied().collect();
assigned.sort_unstable();
assert_eq!(assigned, (0..events.len()).collect::<Vec<_>>());
assert_eq!(groups.total_events(), events.len());
// No empty colors: one would waste a sweep and make `n_colors`
// misleading.
for (color, group) in groups.groups.iter().enumerate() {
assert!(!group.is_empty(), "color {color} is empty");
}
// Contiguity is not a property of `color_greedy` — it holds only after
// `recompute_color_groups` reorders the events so each color occupies
// one range. What must always hold is that the reorder is *possible*:
// relabelling events in group order yields contiguous groups. The
// parallel sweep slices `&mut` sub-ranges from those, so if this ever
// failed the reorder would produce overlapping ranges.
let mut next = 0usize;
let relabelled: Vec<Vec<usize>> = groups
.groups
.iter()
.map(|group| {
group
.iter()
.map(|_| {
let i = next;
next += 1;
i
})
.collect()
})
.collect();
assert!(ColorGroups { groups: relabelled }.groups_are_contiguous());
}
proptest! {
#![proptest_config(ProptestConfig::with_cases(512))]
/// Small competitor pool, so collisions are common and colors are
/// forced to multiply.
#[test]
fn colors_are_disjoint_on_a_dense_pool(
events in prop::collection::vec(
prop::collection::vec(0usize..6, 1..4),
0..20,
)
) {
check(&events);
}
/// Wide pool, so most events are independent and land in one color.
#[test]
fn colors_are_disjoint_on_a_sparse_pool(
events in prop::collection::vec(
prop::collection::vec(0usize..200, 1..6),
0..30,
)
) {
check(&events);
}
/// Repeated competitors within one event must not confuse the
/// member-set bookkeeping.
#[test]
fn colors_are_disjoint_with_repeated_members(
events in prop::collection::vec(
prop::collection::vec(0usize..3, 1..8),
0..15,
)
) {
check(&events);
}
}
}
+72 -4
View File
@@ -4,9 +4,38 @@ use std::time::Duration;
use smallvec::SmallVec;
#[derive(Clone, Copy, Debug)]
/// The stopping rule for the fixed-point loops, plus how hard they are damped.
///
/// Set once per history through
/// [`HistoryBuilder::convergence`](crate::HistoryBuilder::convergence), and
/// carried by `GameOptions` for a single match scored without a history. The
/// defaults are the crate's globals: [`ITERATIONS`](crate::ITERATIONS),
/// [`EPSILON`](crate::EPSILON), and undamped EP.
///
/// Deliberately **not** `#[non_exhaustive]`, unlike [`ConvergenceReport`]. The
/// usual argument for marking an options struct is that `..Default::default()`
/// makes a future field additive — but `Default::default` is not a `const fn`,
/// so marking it would make
/// `const OPTS: ConvergenceOptions = ConvergenceOptions { .. }` impossible from
/// outside the crate, with no workaround. This type is `Copy` and a natural
/// const; that cost is permanent, and adding a field is a one-time major bump.
#[derive(Clone, Copy, Debug, PartialEq)]
pub struct ConvergenceOptions {
/// Hard cap on full forward+backward sweeps.
///
/// A runaway guard, not a budget: the loop exits as soon as the step falls
/// to `epsilon`, so raising this costs nothing on a history that converges.
/// Reaching it is
/// [`InferenceError::NotConverged`](crate::InferenceError::NotConverged).
pub max_iter: usize,
/// Convergence threshold, in skill units.
///
/// The sweep stops once *both* components of the step — the largest change
/// a whole iteration made to any competitor's posterior mean, and to any
/// posterior standard deviation — are at or below this. Larger values stop
/// sooner and further from the fixed point. Must be non-negative; NaN is
/// rejected, since every comparison against it is false and the loop would
/// read it as converged.
pub epsilon: f64,
/// EP damping factor in natural-parameter space: each per-factor
/// update inside a single game writes `α·new + (1−α)·old`. `1.0` is
@@ -37,13 +66,13 @@ impl ConvergenceOptions {
pub(crate) fn validate(&self) -> Result<(), crate::InferenceError> {
if !(self.alpha > 0.0 && self.alpha <= 1.0) {
return Err(crate::InferenceError::InvalidParameter {
name: "alpha",
parameter: crate::Parameter::Alpha,
value: self.alpha,
});
}
if self.epsilon.is_nan() || self.epsilon < 0.0 {
return Err(crate::InferenceError::InvalidParameter {
name: "epsilon",
parameter: crate::Parameter::Epsilon,
value: self.epsilon,
});
}
@@ -62,12 +91,51 @@ impl Default for ConvergenceOptions {
}
/// Post-hoc summary of a `History::converge` call.
#[derive(Clone, Debug)]
///
/// From [`History::converge`](crate::History::converge) this always describes a
/// converged fit — stopping at `max_iter` is
/// [`InferenceError::NotConverged`](crate::InferenceError::NotConverged) there.
/// From [`History::converge_partial`](crate::History::converge_partial) it may
/// not be, and `converged` is what says so.
/// Constructed only by `converge` / `converge_partial`, never by a caller, so
/// `#[non_exhaustive]` costs nothing here and lets a future field be additive.
/// The two *options* structs deliberately do not carry it — see the note on
/// [`ConvergenceOptions`].
#[derive(Clone, Debug, PartialEq)]
#[non_exhaustive]
pub struct ConvergenceReport {
/// Full forward+backward sweeps actually run. `0` for a history with no
/// time slices, which is converged trivially.
pub iterations: usize,
/// How far the last sweep still moved the fit, as `(mean, standard
/// deviation)`.
///
/// Not natural parameters: each component is a componentwise maximum of
/// `|Δmu|` and `|Δsigma|` over every competitor posterior the sweep
/// touched, so both are in skill units and both are non-negative. Each is
/// compared against `epsilon` separately — `converged` means neither
/// exceeds it. `(0.0, 0.0)` for a history with no time slices.
pub final_step: (f64, f64),
/// Natural log of the model evidence for the whole history at this fit,
/// summed over every time slice.
///
/// The same quantity
/// [`History::log_evidence`](crate::History::log_evidence) returns, taken
/// once the sweep has stopped. Only comparable between fits of the same
/// events; higher means the model explains them better.
pub log_evidence: f64,
/// Whether the sweep reached `epsilon` rather than stopping at `max_iter`.
///
/// Always `true` from [`History::converge`](crate::History::converge),
/// which reports the other case as `NotConverged`. From
/// [`History::converge_partial`](crate::History::converge_partial) this is
/// the only thing that distinguishes a finished fit from a capped one.
pub converged: bool,
/// Wall-clock time each sweep took, in the order they ran.
///
/// One entry per iteration, so its length equals `iterations`; empty for a
/// history with no time slices. It times the sweeps only, so the final
/// log-evidence pass is not in any entry.
pub per_iteration_time: SmallVec<[Duration; 32]>,
}
+52 -2
View File
@@ -21,8 +21,58 @@ pub trait Drift<T: Time>: Copy + Debug + Send + Sync {
///
/// For `Time = i64`: variance added is `(to - from) * gamma^2`.
/// For `Time = Untimed`: elapsed is always 0, so drift is always 0.
#[derive(Clone, Copy, Debug)]
pub struct ConstantDrift(pub f64);
///
/// # Why the field is private
///
/// `gamma` enters only as `gamma * gamma`, so a negative value is squared away:
/// measured against the old public-field form, `ConstantDrift(-0.0833)` produced
/// results **bit identical** to `ConstantDrift(0.0833)`. The sign was neither
/// rejected nor honoured — it vanished. That is the same sign-absorption `HistoryBuilder::sigma`,
/// `HistoryBuilder::beta`, `Gaussian::from_ms` and `Rating::new` all reject.
///
/// It could not be checked while the field was a public tuple position, because
/// there was no constructor to intercept. Validating inside
/// `variance_for_elapsed` would have been worse: it runs inside the sweep, so a
/// construction-time mistake would panic mid-inference — and `Gaussian::from_ms`
/// is a worked example of why that is the wrong place for a guard, where
/// rejecting NaN turned the `NonFiniteStep` reporting path into a crash.
///
/// So [`ConstantDrift::new`] is the only way in, and it checks. Read the value
/// back with [`ConstantDrift::gamma`].
///
/// A non-finite gamma is caught a second time regardless:
/// `History::converge` validates the drift variance each competitor actually
/// accumulates, which also covers a custom [`Drift`] implementation.
#[derive(Clone, Copy, Debug, PartialEq)]
pub struct ConstantDrift(f64);
impl ConstantDrift {
/// Drift of `gamma` standard deviations per unit time.
///
/// # Panics
///
/// Panics unless `gamma` is finite and non-negative.
///
/// The field is private and this is the only constructor precisely so that
/// there is somewhere to check. While it was a public tuple field there was
/// nothing to intercept, and a negative gamma was silently squared away —
/// see the type docs.
#[must_use]
pub fn new(gamma: f64) -> Self {
assert!(
gamma.is_finite() && gamma >= 0.0,
"gamma must be finite and non-negative (got {gamma}); it is only ever \
squared, so a negative value would silently behave as its absolute value"
);
Self(gamma)
}
/// Standard deviations of drift accumulated per unit time.
#[must_use]
pub fn gamma(&self) -> f64 {
self.0
}
}
impl<T: Time> Drift<T> for ConstantDrift {
fn variance_delta(&self, from: &T, to: &T) -> f64 {
+573 -34
View File
@@ -1,38 +1,338 @@
use std::fmt;
/// How a prediction should treat a key the history has never seen.
///
/// Configured once per history via
/// [`HistoryBuilder::unknown_keys`](crate::HistoryBuilder::unknown_keys).
/// Neither known consumer wants this to vary between queries — one predicts
/// thousands of candidate matchups in a loop, the other's headline feature is
/// predicting a competitor nobody has faced — so it is a property of how you
/// intend to use the model rather than an argument on five call sites.
///
/// # There is deliberately no `Skip`
///
/// Dropping an unknown member is the obvious third option and it is wrong. A
/// team's performance is the *sum* of its members, so removing one removes its
/// variance too: measured on a two-member team with one unknown, skipping gives
/// a performance sigma of 2.37 where treating the member as unknown gives 6.53.
/// An unknown competitor would make the model *more* certain, which is
/// backwards. `Prior` is also the answer the model already gives for a
/// competitor it knows about but has no evidence for, so it corresponds to a
/// state the model can actually be in; skipping does not.
#[derive(Clone, Copy, Debug, Default, PartialEq, Eq)]
#[non_exhaustive]
pub enum UnknownKeys {
/// Reject the prediction with [`InferenceError::UnknownKey`].
///
/// The default, and the right one when every key is expected to be known:
/// a team of strangers should not silently produce a confident-looking
/// answer.
#[default]
Reject,
/// Treat an unknown competitor as one sitting at the history's configured
/// prior.
///
/// This is the honest Bayesian reading — a competitor you have never
/// observed is exactly the prior — and it makes "predict a matchup
/// involving someone new" a first-class question rather than something a
/// caller fakes with a neutral constant.
Prior,
}
/// Which scalar an [`InferenceError::InvalidParameter`] is about.
///
/// A typed discriminator rather than a `&'static str`, so a caller can branch
/// on it and `Display` can state each parameter's actual valid range. Nine
/// distinct strings used to flow through this position, and the only thing a
/// caller could do with one was print it.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)]
#[non_exhaustive]
pub enum Parameter {
/// Prior mean skill. Must be finite.
Mu,
/// Prior standard deviation. Must be finite and strictly positive.
Sigma,
/// Performance noise. Must be finite and non-negative.
Beta,
/// Draw probability. Must be in `[0.0, 1.0)`.
PDraw,
/// Observation noise on a score margin. Must be finite and strictly
/// positive.
ScoreSigma,
/// EP damping factor. Must be in `(0.0, 1.0]`.
Alpha,
/// Convergence threshold. Must be non-negative and not NaN.
Epsilon,
/// A competitor's multiplier on the drift variance. Must be finite and
/// non-negative.
DriftScale,
/// The variance a [`Drift`](crate::Drift) implementation actually produced
/// for a span. Must be finite and non-negative — checked because a custom
/// implementation is the one thing no constructor can validate up front.
DriftVariance,
/// A per-member weight on an event. Must be finite.
Weight,
/// A team's score on a scored event. Must be finite.
Score,
/// A team's rank on a ranked event. Must be finite.
Rank,
/// The winning team's index, as given to `Outcome::winner`. Must be less
/// than the team count.
WinnerIndex,
}
impl Parameter {
/// The range this parameter must lie in, for the `Display` message.
fn range(self) -> &'static str {
match self {
Self::Mu => "must be finite",
Self::Sigma => "must be finite and strictly positive",
Self::Beta => "must be finite and non-negative",
Self::PDraw => "must be in [0.0, 1.0)",
Self::ScoreSigma => "must be finite and strictly positive",
Self::Alpha => "must be in (0.0, 1.0]",
Self::Epsilon => "must be non-negative and not NaN",
Self::DriftScale => "must be finite and non-negative",
Self::DriftVariance => "must be finite and non-negative",
Self::Weight => "must be finite",
Self::Score => "must be finite",
Self::Rank => "must be finite",
Self::WinnerIndex => "must be less than the number of teams",
}
}
}
impl std::fmt::Display for Parameter {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
let name = match self {
Self::Mu => "mu",
Self::Sigma => "sigma",
Self::Beta => "beta",
Self::PDraw => "p_draw",
Self::ScoreSigma => "score_sigma",
Self::Alpha => "alpha",
Self::Epsilon => "epsilon",
Self::DriftScale => "drift_scale",
Self::DriftVariance => "drift variance",
Self::Weight => "weight",
Self::Score => "score",
Self::Rank => "rank",
Self::WinnerIndex => "winner index",
};
f.write_str(name)
}
}
/// Which two lengths an [`InferenceError::MismatchedShape`] found disagreeing.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)]
#[non_exhaustive]
pub enum Shape {
/// The outcome describes a different number of teams than the event has.
OutcomeVsTeams,
/// A per-member weight list does not match the team's membership.
Weights,
/// A call that takes a fixed number of teams got a different number.
Teams,
/// One of `add_events_with_prior`'s parallel arrays disagreed with the
/// others.
///
/// Not reachable through the public API — the arrays are built together at
/// the ingestion chokepoint. Kept as a checked error rather than a
/// `debug_assert!` so it also holds in release, which is where this
/// crate's defects have tended to hide.
Internal,
}
impl std::fmt::Display for Shape {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
let what = match self {
Self::OutcomeVsTeams => {
"the outcome describes a different number of teams than the event has"
}
Self::Weights => "the weight list does not match the team's membership",
Self::Teams => "this call takes a fixed number of teams",
Self::Internal => {
"an internal array disagreed with its siblings (this is a bug in trueskill-tt)"
}
};
f.write_str(what)
}
}
/// Which [`Outcome`](crate::Outcome) variant a call found or wanted.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)]
#[non_exhaustive]
pub enum OutcomeKind {
/// [`Outcome::Ranked`](crate::Outcome::Ranked): an ordinal finish.
Ranked,
/// [`Outcome::Scored`](crate::Outcome::Scored): continuous scores.
Scored,
}
impl OutcomeKind {
/// The call that takes this kind, for the `Display` message.
fn constructor(self) -> &'static str {
match self {
Self::Ranked => "Game::ranked",
Self::Scored => "Game::scored",
}
}
/// The adjective form, for prose.
fn adjective(self) -> &'static str {
match self {
Self::Ranked => "ranked",
Self::Scored => "scored",
}
}
}
impl std::fmt::Display for OutcomeKind {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.write_str(match self {
Self::Ranked => "Outcome::Ranked",
Self::Scored => "Outcome::Scored",
})
}
}
/// Which piece of per-competitor configuration was declared twice.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)]
#[non_exhaustive]
pub enum CompetitorField {
/// The starting skill distribution.
Prior,
/// The multiplier on the drift variance.
DriftScale,
}
impl std::fmt::Display for CompetitorField {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.write_str(match self {
Self::Prior => "prior",
Self::DriftScale => "drift_scale",
})
}
}
/// Every way ingestion, inference or prediction can refuse to answer.
///
/// The crate reports rather than repairs. An input it cannot represent, a fit
/// that never reached its fixed point, a quadrature it cannot resolve — each
/// comes back here instead of as a clamped, skipped or truncated result that
/// would still look like a number. Several variants exist precisely because the
/// silent version was measured and found to return a plausible wrong answer.
///
/// The enum and most of its variants are `#[non_exhaustive]`: new cases and new
/// fields are additive, so match with a `_` arm and construct through the
/// library rather than by literal.
#[derive(Debug, Clone, PartialEq)]
#[non_exhaustive]
pub enum InferenceError {
/// Expected and actual lengths of some array-shaped input differ.
#[non_exhaustive]
MismatchedShape {
kind: &'static str,
/// Which pair of lengths disagreed.
shape: Shape,
/// The length it had to have, taken from whatever it must line up with
/// (usually the event's team count).
expected: usize,
/// The length actually supplied.
got: usize,
},
/// An `Outcome` of the wrong variant was supplied for the requested inference.
#[non_exhaustive]
WrongOutcomeKind {
context: &'static str,
expected: &'static str,
got: &'static str,
/// The variant the call needs.
expected: OutcomeKind,
/// The variant actually supplied.
got: OutcomeKind,
},
/// A probability value is outside `[0, 1]`.
InvalidProbability { value: f64 },
/// A scalar parameter is outside its valid range.
InvalidParameter { name: &'static str, value: f64 },
#[non_exhaustive]
InvalidParameter {
/// Which parameter. `Display` states its valid range.
parameter: Parameter,
/// The value supplied for it: outside that range, or NaN, which fails
/// every range comparison and is rejected on that basis.
value: f64,
},
/// An event contains tied teams, but the draw probability is zero.
///
/// A zero draw probability asserts that draws cannot occur, so a tied
/// result has no representable likelihood. Configure a positive `p_draw`
/// (via `HistoryBuilder::p_draw` or `GameOptions::p_draw`) to admit ties.
TieWithoutDrawProbability { teams: (usize, usize) },
/// Inference produced a non-finite value (NaN or infinity).
#[non_exhaustive]
TieWithoutDrawProbability {
/// Positions in the event's team list of the first tied pair, lowest
/// index first. Only one pair is reported — the event is rejected
/// whole, so enumerating the rest would add nothing.
teams: (usize, usize),
},
/// The convergence sweep hit `max_iter` with the step still above
/// `epsilon`.
///
/// Indicates numerical breakdown; the resulting skills are meaningless
/// and must not be treated as a converged estimate.
NonFiniteResult {
/// A fit that stops short is wrong by a little, which is the worst
/// available failure: every posterior is finite, the ordering looks sensible,
/// and nothing in the numbers says they were still moving. Reported rather
/// than returned as a flag on an `Ok`, because a flag has to be checked
/// and `let _ = h.converge()` is the natural way not to.
///
/// Either the history needs more iterations — raise `max_iter` — or it is
/// oscillating rather than converging, in which case `alpha < 1.0` damps
/// the within-game EP loop. [`History::converge_partial`](crate::History::converge_partial)
/// returns the short fit instead when that is genuinely what is wanted.
#[non_exhaustive]
NotConverged {
/// Full forward+backward sweeps run before the loop gave up.
iterations: usize,
/// How far the last sweep still moved the fit, as
/// `(largest change in a mean, largest change in a standard
/// deviation)` over every competitor posterior it touched — the same
/// quantity as
/// [`ConvergenceReport::final_step`](crate::ConvergenceReport).
final_step: (f64, f64),
/// The threshold both components of `final_step` had to reach.
epsilon: f64,
},
/// A convergence sweep produced a non-finite step.
///
/// EP has broken down; the resulting skills are meaningless and must not
/// be treated as a converged estimate. Further iterations cannot recover,
/// so the loop stops rather than reporting a NaN step as convergence.
#[non_exhaustive]
NonFiniteStep {
/// Where the breakdown was caught, e.g. `"History::converge"`.
context: &'static str,
/// The offending step as `(|d mu|, |d sigma|)`, at least one component
/// of which is NaN or infinite.
step: (f64, f64),
},
/// A prediction read a skill with no usable mean or variance.
///
/// Split from `NonFiniteStep` (#74), which used to carry both under one
/// `step: (f64, f64)` field — a sweep step from `converge` and a skill's
/// own moments from a prediction. One field name cannot be right for both.
///
/// Reaching this means a previous `converge` failed and its error was
/// ignored: predicting from a NaN fit produced `Ok(NaN)` on some paths and
/// a plausible-looking `Ok([0.0, 0.0])` on others.
#[non_exhaustive]
NonFiniteSkill {
/// The skill's mean, which may itself be finite while `sigma` is not.
mu: f64,
/// The skill's standard deviation.
sigma: f64,
},
/// Every skill in the matchup is a point mass and `beta` is zero, so
/// there is no performance distribution to predict from.
///
/// Not `InvalidParameter`: both values are individually in range, and it
/// is their combination that leaves nothing varying. Every prediction is a
/// statement about how performances vary, and in this configuration
/// nothing does — `quality` would divide by a singular contrast covariance
/// and `predict_win_probabilities` would report zeros that sum to zero.
NoPerformanceVariance,
/// One batch declared two different values for the same competitor's
/// configuration.
///
@@ -42,19 +342,116 @@ pub enum InferenceError {
/// "last one wins" would make the result depend on iteration order.
/// Declaring the same value repeatedly is fine and is the expected shape
/// when a competitor's configuration is a property of the domain.
#[non_exhaustive]
ConflictingCompetitorConfig {
/// The competitor's interned slot as a raw `usize`,
/// not the user key — the batch is already flattened to indices by the
/// time the conflict is detectable.
competitor: usize,
field: &'static str,
/// Which piece of configuration was declared twice.
field: CompetitorField,
},
/// A prediction referenced a key the history has no skill for.
///
/// Reported rather than skipped: dropping unknown keys turns a team of
/// strangers into a confident-looking probability about nobody.
UnknownKey { team: usize, member: usize },
///
/// `key` is the offending key's `Debug` rendering. It is carried because
/// the indices alone are not actionable: a caller that logs
/// `UnknownKey { team: 0, member: 0 }` learns nothing about *which* of its
/// keys the history has not seen, and the natural handling — fall back to a
/// neutral value — turns the whole thing into a plausible constant.
#[non_exhaustive]
UnknownKey {
/// Position of the offending team in the supplied matchup. `0` on the
/// queries that take a flat list of keys rather than teams, where
/// there is only one list to index into.
team: usize,
/// Position of the offending key within that team, or within the flat
/// key list.
member: usize,
/// The key's `Debug` rendering, captured because `K` is only required
/// to be `Debug` — see the variant docs for why the indices alone are
/// not enough.
key: String,
},
/// `History::register` was called for a competitor that already exists.
///
/// Registration states a competitor's configuration before anything has
/// been observed about them, so a competitor that already exists has
/// already been configured — by an earlier `register`, or by an event that
/// created them. Silently overwriting would reintroduce exactly the
/// order-dependence registration exists to remove.
///
/// To change an existing competitor's configuration, supply it on an event
/// through `Member`; that refits the whole history.
#[non_exhaustive]
AlreadyRegistered {
/// The already-known competitor's key, in its `Debug` rendering.
key: String,
},
/// A prediction was given a team with no members.
EmptyTeam { team: usize },
#[non_exhaustive]
EmptyTeam {
/// Position of the memberless team in the supplied list.
team: usize,
},
/// The prediction grid cannot resolve the narrowest feature in the matchup.
///
/// `predict_outcome` and `predict_ranking` integrate every team's density
/// on one shared grid, whose resolution is set by the narrowest sigma (or a
/// narrower draw margin). When the widest and narrowest are far enough
/// apart, resolving the narrow one across the wide one's support needs more
/// nodes than the grid is allowed to hold.
///
/// Reported rather than clamped. Clamping is what this replaced, and it
/// returned probabilities greater than one — measured, a `P` of 2.79 and a
/// `Prediction::total()` of 5.41 — because the trapezoid rule stops
/// resolving a density once the step exceeds roughly 1.7 of its sigma.
///
/// `predict_win_probabilities` answers the same matchup through adaptive
/// quadrature and is accurate here; use it when only the per-team win
/// probabilities are needed.
#[non_exhaustive]
GridTooCoarse {
/// Nodes required to resolve the narrowest feature.
needed: usize,
/// Nodes the grid may hold.
max: usize,
},
/// A joint posterior was requested from a history with no events.
///
/// Split out of a single `JointUnavailable { reason: &str }` (#74): the
/// three reasons are conditions a caller branches on differently, and
/// distinguishing them used to mean matching on English prose. This one
/// means "add events".
EmptyHistory,
/// A joint posterior was requested from a history containing ranked
/// events.
///
/// Exact only for an all-scored history: a scored likelihood is Gaussian
/// and its factor can be rebuilt exactly, while a ranked outcome's
/// truncation is approximated by EP and reconstructing those factors needs
/// the converged messages, which inference does not retain.
///
/// [`History::predict_win_probabilities`](crate::History::predict_win_probabilities)
/// answers the comparable question on a ranked history.
JointRequiresScoredEvents,
/// The assembled precision matrix is not positive-definite.
///
/// The usual cause is a competitor with neither a proper prior nor any
/// evidence, but an extreme prior or drift can also make the assembled
/// matrix indefinite in floating point. Unlike its two siblings this one
/// is numerical rather than structural — the same history may factorise
/// under different parameters.
NotPositiveDefinite,
/// Fewer than two teams were supplied to a prediction.
NotEnoughTeams { got: usize },
#[non_exhaustive]
NotEnoughTeams {
/// How many teams the prediction was actually given. Two is the
/// minimum: there is nothing to compare against with fewer.
got: usize,
},
/// The full outcome distribution was requested for too many teams.
///
/// Each realisation sorts into exactly one (order, tie-pattern) event, so
@@ -63,28 +460,33 @@ pub enum InferenceError {
/// enumerate on a caller's behalf; ask for individual rankings with
/// `predict_ranking`, or for `predict_win_probabilities`, both of which
/// stay cheap at any team count.
TooManyTeams { got: usize, max: usize },
#[non_exhaustive]
TooManyTeams {
/// How many teams the outcome distribution was asked for.
got: usize,
/// The largest team count that will be enumerated,
/// [`MAX_PREDICTED_TEAMS`](crate::MAX_PREDICTED_TEAMS).
max: usize,
},
}
impl fmt::Display for InferenceError {
fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
match self {
Self::MismatchedShape {
kind,
shape,
expected,
got,
} => {
write!(f, "{kind}: expected length {expected}, got {got}")
write!(f, "{shape}: expected {expected}, got {got}")
}
Self::WrongOutcomeKind {
context,
expected,
got,
} => {
write!(f, "{context}: expected {expected}, got {got}")
}
Self::InvalidProbability { value } => {
write!(f, "probability must be in [0, 1]; got {value}")
Self::WrongOutcomeKind { expected, got } => {
write!(
f,
"expected {expected}, got {got}; call {} for a {} outcome",
got.constructor(),
got.adjective()
)
}
Self::TieWithoutDrawProbability { teams } => {
write!(
@@ -93,14 +495,35 @@ impl fmt::Display for InferenceError {
teams.0, teams.1
)
}
Self::NonFiniteResult { context, step } => {
Self::NotConverged {
iterations,
final_step,
epsilon,
} => {
write!(
f,
"{context}: inference produced a non-finite result (step = {step:?})"
"did not converge in {iterations} iterations: final step {final_step:?} \
is still above epsilon {epsilon}; raise max_iter, or damp with \
alpha < 1.0 if it is oscillating"
)
}
Self::InvalidParameter { name, value } => {
write!(f, "{name} is invalid: {value}")
Self::NonFiniteStep { context, step } => {
write!(
f,
"{context}: inference produced a non-finite step {step:?}; EP has \
broken down and further iterations cannot recover"
)
}
Self::NonFiniteSkill { mu, sigma } => {
write!(
f,
"a prediction read a skill with no usable mean or variance \
(mu = {mu}, sigma = {sigma}); the fit did not converge, and \
`converge` reports that"
)
}
Self::InvalidParameter { parameter, value } => {
write!(f, "{parameter} {} (got {value})", parameter.range())
}
Self::ConflictingCompetitorConfig { competitor, field } => {
write!(
@@ -108,15 +531,53 @@ impl fmt::Display for InferenceError {
"competitor {competitor}: this batch sets {field} to two different values"
)
}
Self::UnknownKey { team, member } => {
Self::UnknownKey { team, member, key } => {
write!(
f,
"team {team}, member {member}: no skill recorded for this key"
"team {team}, member {member}: no skill recorded for key {key} \
(every key must already be known to the history; pre-filter \
with `current_skill` if that is not guaranteed)"
)
}
Self::AlreadyRegistered { key } => {
write!(
f,
"competitor {key} is already registered; registration states \
configuration before anything is observed, so re-registering \
would silently overwrite it"
)
}
Self::EmptyTeam { team } => {
write!(f, "team {team} has no members")
}
Self::GridTooCoarse { needed, max } => {
write!(
f,
"the prediction grid needs {needed} nodes to resolve the narrowest \
team's density across the widest team's support, but may hold only \
{max}; the sigmas in this matchup are too far apart to integrate on \
one grid. Use predict_win_probabilities, which is accurate here"
)
}
Self::EmptyHistory => {
f.write_str("no exact joint posterior is available: the history has no events")
}
Self::JointRequiresScoredEvents => f.write_str(
"no exact joint posterior is available: the history contains ranked \
events, whose EP factors are not retained after convergence. Use \
predict_win_probabilities for a ranked history",
),
Self::NotPositiveDefinite => f.write_str(
"the joint precision matrix is not positive-definite; the usual cause \
is a competitor with neither a proper prior nor any evidence, but an \
extreme prior or drift can also make the assembled matrix indefinite \
in floating point",
),
Self::NoPerformanceVariance => f.write_str(
"beta is zero and every skill in this matchup is a point mass, so \
there is no performance distribution to predict from; give beta a \
positive value, or a competitor a prior with positive sigma",
),
Self::NotEnoughTeams { got } => {
write!(f, "prediction needs at least 2 teams, got {got}")
}
@@ -132,3 +593,81 @@ impl fmt::Display for InferenceError {
}
impl std::error::Error for InferenceError {}
#[cfg(test)]
mod message_tests {
use super::*;
/// Every message must name the problem *and* what to do, which is the
/// standard the good ones set and the three #74 called out did not meet.
#[test]
fn messages_are_actionable() {
let cases = [
InferenceError::InvalidParameter {
parameter: Parameter::Alpha,
value: 0.0,
},
InferenceError::InvalidParameter {
parameter: Parameter::DriftVariance,
value: f64::NAN,
},
InferenceError::InvalidParameter {
parameter: Parameter::PDraw,
value: 1.5,
},
InferenceError::MismatchedShape {
shape: Shape::OutcomeVsTeams,
expected: 3,
got: 2,
},
InferenceError::WrongOutcomeKind {
expected: OutcomeKind::Ranked,
got: OutcomeKind::Scored,
},
InferenceError::EmptyHistory,
InferenceError::JointRequiresScoredEvents,
InferenceError::NotPositiveDefinite,
InferenceError::NoPerformanceVariance,
InferenceError::NonFiniteSkill {
mu: f64::NAN,
sigma: f64::NAN,
},
];
for case in &cases {
let rendered = case.to_string();
eprintln!("{rendered}");
// `InvalidParameter` used to render `drift variance is invalid: NaN`
// — no range, no remedy, no location. Every message must at least
// be a sentence.
assert!(
rendered.len() > 30,
"message is too terse to act on: {rendered}"
);
assert!(!rendered.contains("is invalid:"), "{rendered}");
}
// The three that #74 singled out now state a range or a next step.
assert!(
InferenceError::InvalidParameter {
parameter: Parameter::DriftVariance,
value: f64::NAN,
}
.to_string()
.contains("must be finite and non-negative")
);
assert!(
InferenceError::WrongOutcomeKind {
expected: OutcomeKind::Ranked,
got: OutcomeKind::Scored,
}
.to_string()
.contains("Game::scored")
);
assert!(
InferenceError::JointRequiresScoredEvents
.to_string()
.contains("predict_win_probabilities")
);
}
}
+70 -8
View File
@@ -11,27 +11,59 @@ use smallvec::SmallVec;
use crate::{gaussian::Gaussian, outcome::Outcome, time::Time};
/// A single match at time `time` involving some number of teams.
#[derive(Clone, Debug)]
#[derive(Clone, Debug, PartialEq)]
pub struct Event<T: Time, K> {
/// When the match happened, on the history's time axis.
///
/// Events sharing a `time` land in the same time slice and are fitted
/// together, so nothing distinguishes their order. Drift is driven by the
/// gap between a competitor's *consecutive appearances*, not by the gap
/// between slices, so a competitor idle across several slices accumulates
/// the whole span at once when it next plays.
pub time: T,
/// The teams that took part, positionally aligned with `outcome`: team `i`
/// here is the team `outcome` ranks or scores at index `i`.
///
/// Ingestion rejects fewer than two teams (`NotEnoughTeams`) and any team
/// with no members (`EmptyTeam`).
pub teams: SmallVec<[Team<K>; 4]>,
/// How the match ended: ranks (lower is better) or per-team scores (higher
/// is better), one entry per entry of `teams`.
///
/// A tie — two equal ranks — needs a positive `p_draw`, otherwise
/// ingestion fails with `TieWithoutDrawProbability`.
pub outcome: Outcome,
}
/// A team: list of members competing together.
#[derive(Clone, Debug)]
#[derive(Clone, Debug, PartialEq)]
#[must_use]
pub struct Team<K> {
/// The competitors playing together, in no significant order: the team's
/// performance is the weight-scaled sum over its members, which does not
/// depend on how they are listed.
///
/// Must be non-empty — an empty team contributes no performance at all, so
/// ingestion rejects it with `EmptyTeam` rather than returning a plausible
/// posterior for whoever it was matched against.
pub members: SmallVec<[Member<K>; 4]>,
}
impl<K> Team<K> {
#[must_use]
/// A team with no members yet, to be filled through the public `members`
/// field.
///
/// Committing it while still empty is an `EmptyTeam` error.
pub fn new() -> Self {
Self {
members: SmallVec::new(),
}
}
/// A team of exactly these competitors.
///
/// Members must be built already — `Member::from(key)` covers the common
/// case of a plain key at default weight with no overrides.
pub fn with_members<I: IntoIterator<Item = Member<K>>>(members: I) -> Self {
Self {
members: members.into_iter().collect(),
@@ -61,10 +93,29 @@ impl<K> Default for Team<K> {
/// for one competitor within a single batch is
/// `InferenceError::ConflictingCompetitorConfig`: events in a batch have no
/// order, so there would be no well-defined winner.
#[derive(Clone, Debug)]
#[derive(Clone, Debug, PartialEq)]
#[must_use]
pub struct Member<K> {
/// The competitor's identity. Equal keys across events are the same
/// competitor: `History` interns each distinct key to an internal `Index`
/// the first time it sees it, and every later appearance resolves to that
/// same competitor's temporal state.
pub key: K,
/// This member's share of the team's performance, for this event only.
///
/// The team's performance is the sum of `weight × member performance`, so
/// `1.0` is a full share and `0.5` counts the member half; the message
/// coming back to the member is divided by the same weight. Defaults to
/// `1.0`.
///
/// Must be finite — a NaN or infinite weight is `InvalidParameter` at
/// ingestion. Zero and negative are accepted, both being expressible in
/// the same arithmetic.
pub weight: f64,
/// Starting skill for this competitor, replacing the history's `mu`/`sigma`
/// default. `None` keeps the history default.
///
/// Competitor configuration, not a per-event value; see the type docs.
pub prior: Option<Gaussian>,
/// Multiplier on the drift *variance* this competitor accumulates.
/// `None` means 1.0.
@@ -72,6 +123,8 @@ pub struct Member<K> {
}
impl<K> Member<K> {
/// A competitor taking a full share of its team's performance, with no
/// configuration overrides: the history's prior and drift apply.
pub fn new(key: K) -> Self {
Self {
key,
@@ -81,6 +134,12 @@ impl<K> Member<K> {
}
}
/// Change how much of the team's performance this member accounts for.
///
/// Unlike `prior` and `drift_scale`, this is genuinely per-event: the same
/// key can carry a different weight in every event it appears in, which is
/// what makes it usable for partial participation — a substitute who
/// played half the match, a doubles partner credited unequally.
pub fn with_weight(mut self, weight: f64) -> Self {
self.weight = weight;
self
@@ -88,7 +147,9 @@ impl<K> Member<K> {
/// Set this competitor's starting skill estimate.
///
/// Captured at the competitor's first appearance; see the type docs.
/// Competitor configuration, not a per-event value: it applies for the
/// whole history and applies whenever it is supplied, including on a key
/// the history already knows. See the type docs.
pub fn with_prior(mut self, prior: Gaussian) -> Self {
self.prior = Some(prior);
self
@@ -97,14 +158,15 @@ impl<K> Member<K> {
/// Scale how fast this competitor drifts, relative to the history's drift.
///
/// The scale multiplies the drift *variance*, so it is in the same units as
/// `gamma`: `ConstantDrift(g)` at `scale = s` behaves exactly as
/// `ConstantDrift(g * s)` would for this competitor alone.
/// `gamma`: `ConstantDrift::new(g)` at `scale = s` behaves exactly as
/// `ConstantDrift::new(g * s)` would for this competitor alone.
///
/// `0.0` pins the competitor still — useful for a reference point that
/// shares a scale with moving competitors but should not itself move: a bot
/// at a known strength, a rating floor, a course difficulty.
///
/// Captured at the competitor's first appearance; see the type docs.
/// Applies for the whole history and whenever it is supplied, including on
/// a key the history already knows; see the type docs.
/// Must be finite and non-negative, or ingestion fails with
/// [`InferenceError::InvalidParameter`](crate::InferenceError::InvalidParameter).
pub fn with_drift_scale(mut self, scale: f64) -> Self {
+87 -12
View File
@@ -9,14 +9,44 @@ use crate::{
time::Time,
};
pub struct EventBuilder<'h, T, D, O, K>
/// One match under construction, handed back by [`History::event`].
///
/// Describes a single event a piece at a time — teams, then per-member weights
/// if they differ, then how it ended — instead of assembling an
/// [`Event`] value and passing it to [`History::add_events`]. The two routes
/// ingest through the same chokepoint and accept the same things; this one just
/// reads better for a single match written by hand.
///
/// The builder borrows the history mutably and nothing reaches it until
/// [`EventBuilder::commit`]. A builder that is dropped instead ingests
/// nothing at all, silently — hence the `#[must_use]`, which is the only
/// warning you get. `commit` is also where validation surfaces: the setters
/// return `Self` to keep the chain fluent, so a mismatch such as a weight list
/// the wrong length is recorded while building and returned as an error from
/// `commit`.
///
/// ```
/// # use trueskill_tt::History;
/// let mut h = History::builder().build();
/// h.event(1)
/// .team(["alice", "bob"])
/// .team(["carol"])
/// .ranking([0, 1])
/// .commit()?;
/// assert_eq!(h.event_count(), 1);
/// # Ok::<(), trueskill_tt::InferenceError>(())
/// ```
#[must_use = "an event is only recorded by `.commit()`; a dropped builder \
silently ingests nothing"]
pub struct EventBuilder<'h, T, D, O, K, R>
where
T: Time,
D: Drift<T>,
O: Observer<T>,
K: Eq + std::hash::Hash + Clone,
R: crate::RatingRule<K>,
{
history: &'h mut History<T, D, O, K>,
history: &'h mut History<K, T, D, O, R>,
event: Event<T, K>,
current_team_idx: Option<usize>,
/// First validation failure seen while building, surfaced by `commit`.
@@ -29,14 +59,15 @@ where
error: Option<InferenceError>,
}
impl<'h, T, D, O, K> EventBuilder<'h, T, D, O, K>
impl<'h, T, D, O, K, R> EventBuilder<'h, T, D, O, K, R>
where
T: Time,
D: Drift<T>,
O: Observer<T>,
K: Eq + std::hash::Hash + Clone,
R: crate::RatingRule<K>,
{
pub(crate) fn new(history: &'h mut History<T, D, O, K>, time: T) -> Self {
pub(crate) fn new(history: &'h mut History<K, T, D, O, R>, time: T) -> Self {
Self {
history,
event: Event {
@@ -50,6 +81,8 @@ where
}
/// Add a team by its member keys (weight 1.0 each, no prior overrides).
///
/// Use [`EventBuilder::members`] to set `prior` or `drift_scale`.
pub fn team<I: IntoIterator<Item = K>>(mut self, keys: I) -> Self {
let members: SmallVec<[Member<K>; 4]> = keys.into_iter().map(Member::new).collect();
self.event.teams.push(Team { members });
@@ -57,6 +90,40 @@ where
self
}
/// Add a team from fully-specified [`Member`] values.
///
/// [`EventBuilder::team`] is the common case and builds members with
/// `Member::new`, which leaves `prior` and `drift_scale` unset. This is the
/// escape hatch for when they matter:
///
/// ```
/// # use trueskill_tt::{Gaussian, History, Member};
/// # let mut h = History::builder().build();
/// h.event(0)
/// .team(["player"])
/// .members([Member::new("layout_7")
/// .with_drift_scale(0.0)
/// .with_prior(Gaussian::from_ms(0.0, 1.0))])
/// .ranking([0, 1])
/// .commit()?;
/// # Ok::<(), trueskill_tt::InferenceError>(())
/// ```
///
/// One method rather than a `priors` and a `drift_scales` setter beside
/// `weights`: those would have to grow a parallel array — and a parallel
/// length check — every time `Member` gains a field, and each one would be
/// a new way to get the lengths wrong. `Member`'s own builder already
/// expresses all of it.
///
/// `prior` and `drift_scale` are competitor configuration rather than
/// per-event values; see [`Member`] for what that means for a key the
/// history already knows.
pub fn members<I: IntoIterator<Item = Member<K>>>(mut self, members: I) -> Self {
self.event.teams.push(Team::with_members(members));
self.current_team_idx = Some(self.event.teams.len() - 1);
self
}
/// Set per-member weights for the most recently added team.
///
/// A length mismatch is recorded and returned by [`EventBuilder::commit`]
@@ -77,7 +144,7 @@ where
if ws.len() != team.members.len() {
self.error.get_or_insert(InferenceError::MismatchedShape {
kind: "weights",
shape: crate::Shape::Weights,
expected: team.members.len(),
got: ws.len(),
});
@@ -106,13 +173,21 @@ where
/// Set explicit per-team continuous scores with a per-event noise override.
///
/// `sigma` overrides `HistoryBuilder::score_sigma` for this event only.
/// Must be `> 0.0`. Constructing the outcome with a non-positive or NaN
/// sigma is allowed; the value is rejected with
/// `InferenceError::InvalidParameter` when the event is ingested, so
/// callers get an error from `commit` rather than a panic.
pub fn scores_with_sigma<I: IntoIterator<Item = f64>>(mut self, scores: I, sigma: f64) -> Self {
self.event.outcome = crate::Outcome::scores_with_sigma(scores, sigma);
/// `score_sigma` is the observation noise on the *score margin*, not a
/// skill sigma, and it overrides `HistoryBuilder::score_sigma` for this
/// event only. A small value takes the margin near-literally; a large one
/// barely moves the ratings.
///
/// Must be `> 0.0`. Building the outcome with a non-positive or NaN value
/// is allowed; it is rejected with `InferenceError::InvalidParameter` when
/// the event is ingested, so callers get an error from `commit` rather
/// than a panic.
pub fn scores_with_noise<I: IntoIterator<Item = f64>>(
mut self,
scores: I,
score_sigma: f64,
) -> Self {
self.event.outcome = crate::Outcome::scores_with_noise(scores, score_sigma);
self
}
+4 -73
View File
@@ -20,6 +20,8 @@ pub struct VarStore {
}
impl VarStore {
/// Test-only: inference allocates its store through `ScratchArena`.
#[cfg(test)]
#[must_use]
pub fn new() -> Self {
Self::default()
@@ -29,23 +31,19 @@ impl VarStore {
self.marginals.clear();
}
/// Test-only, as `new`.
#[cfg(test)]
#[must_use]
pub fn len(&self) -> usize {
self.marginals.len()
}
#[must_use]
pub fn is_empty(&self) -> bool {
self.marginals.is_empty()
}
pub fn alloc(&mut self, init: Gaussian) -> VarId {
let id = VarId(self.marginals.len() as u32);
self.marginals.push(init);
id
}
#[must_use]
pub fn get(&self, id: VarId) -> Gaussian {
self.marginals[id.0 as usize]
}
@@ -55,58 +53,7 @@ impl VarStore {
}
}
/// A factor in the EP graph.
///
/// Factors hold their own outgoing messages and propagate them by reading
/// connected variable marginals from a `VarStore` and writing back updated
/// marginals.
pub trait Factor: Send + Sync {
/// Update outgoing messages and write back to the var store.
///
/// Returns the max delta `(|Δmu|, |Δsigma|)` across writes this
/// propagation. Used by the `Schedule` to detect convergence.
fn propagate(&mut self, vars: &mut VarStore) -> (f64, f64);
/// Optional log-evidence contribution. Default 0.0 (no contribution).
fn log_evidence(&self, _vars: &VarStore) -> f64 {
0.0
}
}
/// Enum dispatcher for the built-in factor types.
///
/// Using an enum instead of `Box<dyn Factor>` keeps factor data inline and
/// avoids virtual-call overhead in the hot inference loop.
#[derive(Debug)]
pub enum BuiltinFactor {
TeamSum(team_sum::TeamSumFactor),
RankDiff(rank_diff::RankDiffFactor),
Trunc(trunc::TruncFactor),
Margin(margin::MarginFactor),
}
impl Factor for BuiltinFactor {
fn propagate(&mut self, vars: &mut VarStore) -> (f64, f64) {
match self {
Self::TeamSum(f) => f.propagate(vars),
Self::RankDiff(f) => f.propagate(vars),
Self::Trunc(f) => f.propagate(vars),
Self::Margin(f) => f.propagate(vars),
}
}
fn log_evidence(&self, vars: &VarStore) -> f64 {
match self {
Self::Trunc(f) => f.log_evidence(vars),
Self::Margin(f) => f.log_evidence(vars),
Self::TeamSum(_) | Self::RankDiff(_) => 0.0,
}
}
}
pub mod margin;
pub mod rank_diff;
pub mod team_sum;
pub mod trunc;
#[cfg(test)]
@@ -153,20 +100,4 @@ mod tests {
assert_eq!(store.len(), 0);
assert_eq!(store.marginals.capacity(), cap);
}
#[test]
fn builtin_factor_dispatches_to_margin() {
use super::margin::MarginFactor;
let mut vars = VarStore::new();
let diff = vars.alloc(Gaussian::from_ms(0.0, 6.0));
let mut f = BuiltinFactor::Margin(MarginFactor::new(diff, 5.0, 1.0));
f.propagate(&mut vars);
let result = vars.get(diff);
assert!((result.mu() - 4.864864864864865).abs() < 1e-12);
let logz = f.log_evidence(&vars);
assert!((logz - (-3.062235327364623)).abs() < 1e-10);
}
}
+13 -9
View File
@@ -1,6 +1,6 @@
use crate::{
N_INF,
factor::{Factor, VarId, VarStore},
factor::{VarId, VarStore},
gaussian::Gaussian,
ln_pdf,
};
@@ -39,7 +39,7 @@ impl MarginFactor {
/// exactly; `alpha < 1.0` writes `α·new_msg + (1−α)·old_msg`.
pub(crate) fn propagate_with_alpha(&mut self, vars: &mut VarStore, alpha: f64) -> (f64, f64) {
let marginal = vars.get(self.diff);
let cavity = marginal / self.msg;
let cavity = marginal.cavity(self.msg);
if self.log_evidence_cached.is_none() {
self.log_evidence_cached = Some(cavity_log_evidence(cavity, self.m_obs, self.sigma));
@@ -49,18 +49,22 @@ impl MarginFactor {
let damped = self.msg.damp_natural(new_msg, alpha);
let old_msg = self.msg;
self.msg = damped;
vars.set(self.diff, cavity * damped);
vars.set(self.diff, cavity.ep_product(damped));
old_msg.delta(damped)
}
}
impl Factor for MarginFactor {
fn propagate(&mut self, vars: &mut VarStore) -> (f64, f64) {
/// Undamped wrappers, used by this module's tests. Inference drives these
/// factors through `propagate_with_alpha` and reads the cached log evidence
/// directly, so these are not on any production path.
#[cfg(test)]
impl MarginFactor {
pub(crate) fn propagate(&mut self, vars: &mut VarStore) -> (f64, f64) {
self.propagate_with_alpha(vars, 1.0)
}
fn log_evidence(&self, _vars: &VarStore) -> f64 {
pub(crate) fn log_evidence(&self) -> f64 {
self.log_evidence_cached.unwrap_or(0.0)
}
}
@@ -77,7 +81,7 @@ fn cavity_log_evidence(cavity: Gaussian, m_obs: f64, sigma: f64) -> f64 {
// `hypot`, not `sqrt(a^2 + b^2)`: squaring overflows to infinity above a
// sigma of ~1.3e154 and flushes to zero below ~1.5e-154, and `Gaussian`'s
// constructors are public so a caller can reach both.
let combined_sigma = cavity.sigma().hypot(sigma);
let combined_sigma = libm::hypot(cavity.sigma(), sigma);
let value = ln_pdf(m_obs, cavity.mu(), combined_sigma);
// A degenerate cavity (infinite sigma) is the only way to reach a
@@ -85,7 +89,7 @@ fn cavity_log_evidence(cavity: Gaussian, m_obs: f64, sigma: f64) -> f64 {
if value.is_finite() {
value
} else {
f64::MIN_POSITIVE.ln()
libm::log(f64::MIN_POSITIVE)
}
}
@@ -146,7 +150,7 @@ mod tests {
let diff = vars.alloc(Gaussian::from_ms(0.0, 6.0));
let mut f = MarginFactor::new(diff, 5.0, 1.0);
f.propagate(&mut vars);
let logz = f.log_evidence(&vars);
let logz = f.log_evidence();
assert!((logz - (-3.062235327364623)).abs() < 1e-10);
}
-95
View File
@@ -1,95 +0,0 @@
use crate::factor::{Factor, VarId, VarStore};
/// Maintains the constraint `diff = team_a - team_b` between three vars.
///
/// On each propagation:
/// - Reads marginals at `team_a` and `team_b` (which already incorporate any
/// incoming messages from neighboring factors).
/// - Computes `new_diff = team_a - team_b` (variance addition; see `Gaussian::Sub`).
/// - Writes the new marginal to `diff`.
/// - Returns the delta against the previous diff value.
///
/// This factor does NOT store an outgoing message; the diff variable is
/// effectively replaced on each propagation. The `TruncFactor` on the same diff
/// var holds the EP-divide message that produces the cavity.
#[derive(Debug)]
pub struct RankDiffFactor {
pub team_a: VarId,
pub team_b: VarId,
pub diff: VarId,
}
impl Factor for RankDiffFactor {
fn propagate(&mut self, vars: &mut VarStore) -> (f64, f64) {
let a = vars.get(self.team_a);
let b = vars.get(self.team_b);
let new_diff = a - b;
let old = vars.get(self.diff);
vars.set(self.diff, new_diff);
old.delta(new_diff)
}
}
#[cfg(test)]
mod tests {
use super::*;
use crate::{N_INF, gaussian::Gaussian};
#[test]
fn diff_of_two_known_gaussians() {
let mut vars = VarStore::new();
let team_a = vars.alloc(Gaussian::from_ms(25.0, 3.0));
let team_b = vars.alloc(Gaussian::from_ms(20.0, 4.0));
let diff = vars.alloc(N_INF);
let mut f = RankDiffFactor {
team_a,
team_b,
diff,
};
f.propagate(&mut vars);
let result = vars.get(diff);
// mu = 25 - 20 = 5; var = 9 + 16 = 25; sigma = 5
assert!((result.mu() - 5.0).abs() < 1e-12);
assert!((result.sigma() - 5.0).abs() < 1e-12);
}
#[test]
fn delta_zero_on_repeat() {
let mut vars = VarStore::new();
let team_a = vars.alloc(Gaussian::from_ms(10.0, 2.0));
let team_b = vars.alloc(Gaussian::from_ms(8.0, 1.0));
let diff = vars.alloc(N_INF);
let mut f = RankDiffFactor {
team_a,
team_b,
diff,
};
f.propagate(&mut vars);
let (dmu, dsig) = f.propagate(&mut vars);
assert!(dmu < 1e-12);
assert!(dsig < 1e-12);
}
#[test]
fn delta_reflects_team_change() {
let mut vars = VarStore::new();
let team_a = vars.alloc(Gaussian::from_ms(10.0, 1.0));
let team_b = vars.alloc(Gaussian::from_ms(0.0, 1.0));
let diff = vars.alloc(N_INF);
let mut f = RankDiffFactor {
team_a,
team_b,
diff,
};
f.propagate(&mut vars);
// change team_a, repropagate; delta should be positive
vars.set(team_a, Gaussian::from_ms(15.0, 1.0));
let (dmu, _dsig) = f.propagate(&mut vars);
assert!(dmu > 4.0, "expected ~5 delta, got {}", dmu);
}
}
-98
View File
@@ -1,98 +0,0 @@
use crate::{
N00,
factor::{Factor, VarId, VarStore},
gaussian::Gaussian,
};
/// Computes the weighted sum of player performances into a team-perf var.
///
/// Inputs are pre-computed player performance Gaussians (i.e., rating priors
/// already with beta² noise added via `Rating::performance()`). The factor
/// runs once per game and writes the weighted sum to the output var.
#[derive(Debug)]
pub struct TeamSumFactor {
pub inputs: Vec<(Gaussian, f64)>,
pub out: VarId,
}
impl Factor for TeamSumFactor {
fn propagate(&mut self, vars: &mut VarStore) -> (f64, f64) {
let perf = self.inputs.iter().fold(N00, |acc, (g, w)| acc + (*g * *w));
let old = vars.get(self.out);
vars.set(self.out, perf);
old.delta(perf)
}
}
#[cfg(test)]
mod tests {
use super::*;
use crate::N_INF;
#[test]
fn single_player_unit_weight() {
let mut vars = VarStore::new();
let out = vars.alloc(N_INF);
let g = Gaussian::from_ms(25.0, 5.0);
let mut f = TeamSumFactor {
inputs: vec![(g, 1.0)],
out,
};
f.propagate(&mut vars);
let result = vars.get(out);
assert!((result.mu() - 25.0).abs() < 1e-12);
assert!((result.sigma() - 5.0).abs() < 1e-12);
}
#[test]
fn two_players_summed() {
let mut vars = VarStore::new();
let out = vars.alloc(N_INF);
let g1 = Gaussian::from_ms(20.0, 3.0);
let g2 = Gaussian::from_ms(30.0, 4.0);
let mut f = TeamSumFactor {
inputs: vec![(g1, 1.0), (g2, 1.0)],
out,
};
f.propagate(&mut vars);
let result = vars.get(out);
// sum: mu = 20 + 30 = 50, var = 9 + 16 = 25, sigma = 5
assert!((result.mu() - 50.0).abs() < 1e-12);
assert!((result.sigma() - 5.0).abs() < 1e-12);
}
#[test]
fn weighted_inputs() {
let mut vars = VarStore::new();
let out = vars.alloc(N_INF);
let g = Gaussian::from_ms(10.0, 2.0);
let mut f = TeamSumFactor {
inputs: vec![(g, 2.0)],
out,
};
f.propagate(&mut vars);
let result = vars.get(out);
// g * 2.0: mu = 10*2 = 20, sigma = 2*2 = 4
assert!((result.mu() - 20.0).abs() < 1e-12);
assert!((result.sigma() - 4.0).abs() < 1e-12);
}
#[test]
fn delta_is_zero_on_repeat_propagate() {
let mut vars = VarStore::new();
let out = vars.alloc(N_INF);
let g = Gaussian::from_ms(5.0, 1.0);
let mut f = TeamSumFactor {
inputs: vec![(g, 1.0)],
out,
};
f.propagate(&mut vars);
let (dmu, dsig) = f.propagate(&mut vars);
assert!(dmu < 1e-12, "expected ~0 delta on repeat, got {}", dmu);
assert!(dsig < 1e-12);
}
}
+12 -12
View File
@@ -1,6 +1,6 @@
use crate::{
N_INF, approx,
factor::{Factor, VarId, VarStore},
factor::{VarId, VarStore},
gaussian::Gaussian,
ln_interval, ln_sf,
};
@@ -41,14 +41,14 @@ impl TruncFactor {
/// exactly; `alpha < 1.0` writes `α·new_msg + (1−α)·old_msg`.
pub(crate) fn propagate_with_alpha(&mut self, vars: &mut VarStore, alpha: f64) -> (f64, f64) {
let marginal = vars.get(self.diff);
let cavity = marginal / self.msg;
let cavity = marginal.cavity(self.msg);
if self.log_evidence_cached.is_none() {
self.log_evidence_cached = Some(cavity_log_evidence(cavity, self.margin, self.tie));
}
let trunc = approx(cavity, self.margin, self.tie);
let new_msg = trunc / cavity;
let new_msg = trunc.cavity(cavity);
let damped = self.msg.damp_natural(new_msg, alpha);
let old_msg = self.msg;
@@ -57,20 +57,20 @@ impl TruncFactor {
// marginal_new = cavity * stored_msg. With alpha = 1.0 this equals
// `trunc` (since cavity * new_msg = trunc by construction); with
// alpha < 1.0 it reflects the partially-applied update.
vars.set(self.diff, cavity * damped);
vars.set(self.diff, cavity.ep_product(damped));
old_msg.delta(damped)
}
}
impl Factor for TruncFactor {
fn propagate(&mut self, vars: &mut VarStore) -> (f64, f64) {
/// Undamped wrappers, used by this module's tests. Inference drives these
/// factors through `propagate_with_alpha` and reads the cached log evidence
/// directly, so these are not on any production path.
#[cfg(test)]
impl TruncFactor {
pub(crate) fn propagate(&mut self, vars: &mut VarStore) -> (f64, f64) {
self.propagate_with_alpha(vars, 1.0)
}
fn log_evidence(&self, _vars: &VarStore) -> f64 {
self.log_evidence_cached.unwrap_or(0.0)
}
}
/// `ln P(diff > margin)` for a win, `ln P(|diff| < margin)` for a tie.
@@ -95,7 +95,7 @@ fn cavity_log_evidence(diff: Gaussian, margin: f64, tie: bool) -> f64 {
if value.is_finite() {
value
} else {
f64::MIN_POSITIVE.ln()
libm::log(f64::MIN_POSITIVE)
}
}
@@ -203,7 +203,7 @@ mod tests {
let approx = -0.5 * z * z - z.ln() - (2.0 * std::f64::consts::PI).sqrt().ln();
assert!(
got < f64::MIN_POSITIVE.ln(),
got < libm::log(f64::MIN_POSITIVE),
"mu={mu}: {got} is still stuck on the old clamp floor"
);
assert!(
+322 -150
View File
File diff suppressed because it is too large Load Diff
+303 -66
View File
@@ -1,5 +1,3 @@
use std::ops;
use crate::{MU, N_INF, SIGMA};
/// A Gaussian distribution stored in natural parameters.
@@ -11,6 +9,7 @@ use crate::{MU, N_INF, SIGMA};
/// the stored fields with no `sqrt` or reciprocal in the hot path. `mu()` and
/// `sigma()` are accessors computed on demand.
#[derive(Clone, Copy, PartialEq, Debug)]
#[must_use]
pub struct Gaussian {
pi: f64,
tau: f64,
@@ -18,8 +17,43 @@ pub struct Gaussian {
impl Gaussian {
/// Construct from mean and standard deviation.
#[must_use]
///
/// # Panics
///
/// Panics if `sigma` is negative. NaN is deliberately allowed through: a
/// broken fit produces one, and `converge` reports that as
/// `NonFiniteStep` rather than panicking mid-inference.
///
/// A negative sigma used to be accepted and returned results **bit
/// identical** to its absolute value, because sigma only ever enters as
/// `sigma * sigma`. The sign was not rejected and not honoured; it simply
/// vanished. That is the same defect `HistoryBuilder::sigma`,
/// `HistoryBuilder::beta` and `Member::with_drift_scale` already reject.
///
/// # Very small sigma
///
/// `pi = 1 / sigma^2` leaves `f64`'s range below about `1.5e-154`, and
/// `tau = mu * pi` overflows sooner still — at a threshold that depends on
/// `mu`, so there is a band where `pi` is finite and only `tau` is not.
/// Both land on the same point-mass representation the `sigma == 0.0`
/// branch produces, and a point mass with a non-zero mean has `mu() = NaN`,
/// because `tau / pi` is `inf / inf`.
///
/// This is not rejected, because `approx` legitimately produces a very
/// small truncated sigma and inference must not panic. It is worth knowing
/// that such a `Gaussian` is not equal to itself, so two identical
/// declarations of one can be reported as conflicting.
pub const fn from_ms(mu: f64, sigma: f64) -> Self {
// NaN is admitted on purpose. A broken fit legitimately produces a NaN
// sigma — `sqrt` of a negative truncated variance — and the design is
// to propagate that to `converge`'s `NonFiniteStep` guard, not to
// panic inside inference. Rejecting it here turned that reporting path
// into a crash, which two tests caught immediately.
assert!(
sigma >= 0.0 || sigma.is_nan(),
"sigma must not be negative; it is only ever squared, so a negative \
value would silently behave as its absolute value"
);
if sigma == f64::INFINITY {
Self { pi: 0.0, tau: 0.0 }
} else if sigma == 0.0 {
@@ -39,11 +73,12 @@ impl Gaussian {
/// Construct from mean and *variance*, skipping the square-root round trip.
///
/// `from_ms(mu, var.sqrt())` immediately squares the root away again to
/// recover `pi = 1/var`. Variance-combining operations (`Add`, `Sub`,
/// `exclude`, `forget`) work in variance space throughout, so they go
/// through here instead and never take a root.
/// recover `pi = 1/var`. Variance-combining operations work in variance
/// space throughout, so they go through here instead and never take a
/// root. Use it whenever you already hold a variance —
/// [`variance`](Gaussian::variance) is its inverse.
#[inline]
pub(crate) fn from_mv(mu: f64, var: f64) -> Self {
pub fn from_mv(mu: f64, var: f64) -> Self {
if var == f64::INFINITY {
Self { pi: 0.0, tau: 0.0 }
} else if var == 0.0 {
@@ -64,18 +99,32 @@ impl Gaussian {
Self { pi, tau }
}
/// Precision, `1 / sigma^2` — one of the two natural parameters.
///
/// This is the representation the type actually stores, which is why the EP
/// product and cavity (`Mul` / `Div`) are plain adds and subtracts. Larger
/// means more certain; `0.0` is an improper, uninformative message and
/// `inf` is a point mass.
#[inline]
#[must_use]
pub fn pi(&self) -> f64 {
pub(crate) fn pi(&self) -> f64 {
self.pi
}
/// Precision-adjusted mean, `mu / sigma^2` — the other natural parameter.
///
/// Stored rather than derived, for the same reason as [`Gaussian::pi`].
/// Meaningful only alongside `pi`: on its own it is not a location.
#[inline]
#[must_use]
pub fn tau(&self) -> f64 {
pub(crate) fn tau(&self) -> f64 {
self.tau
}
/// Mean skill: the point estimate.
///
/// Derived from the natural parameters as `tau / pi`. An improper message
/// (`pi <= 0`) has no defined mean and reports `0.0` — see
/// [`Gaussian::sigma`], which reports `inf` for the same state, and read
/// the two together before treating a mean as informative.
#[inline]
#[must_use]
pub fn mu(&self) -> f64 {
@@ -90,12 +139,14 @@ impl Gaussian {
}
}
/// Variance, `1 / pi`, without the root-and-square of `sigma().powi(2)`.
/// Variance, without the root-and-square of `sigma().powi(2)`.
///
/// Mirrors `sigma()`'s treatment of the improper (`pi <= 0`) and point-mass
/// (`pi == inf`) cases.
/// Mirrors [`sigma`](Gaussian::sigma)'s treatment of the improper
/// (infinite) and point-mass (zero) cases, and is the inverse of
/// [`from_mv`](Gaussian::from_mv).
#[inline]
pub(crate) fn variance(&self) -> f64 {
#[must_use]
pub fn variance(&self) -> f64 {
if self.pi <= 0.0 {
f64::INFINITY
} else if self.pi.is_infinite() {
@@ -105,6 +156,12 @@ impl Gaussian {
}
}
/// Standard deviation: how unsure this estimate is.
///
/// Derived as `1 / sqrt(pi)`. An improper message (`pi <= 0`) reports
/// `inf`, and a point mass (`pi == inf`) reports `0.0` — both are real
/// states rather than error codes, and both are legitimate for a converged
/// fit with degenerate parameters.
#[inline]
#[must_use]
pub fn sigma(&self) -> f64 {
@@ -120,7 +177,25 @@ impl Gaussian {
}
}
/// How far this Gaussian moved from `other`, as `(|d mu|, |d sigma|)`.
///
/// Identical messages have not moved, whatever their parameters, and that
/// case is answered in natural space before touching `mu()`/`sigma()`. An
/// improper message has `pi == 0`, so `sigma()` is infinite — and
/// `inf - inf` is NaN, a NaN *change* for a message that did not change at
/// all. (`mu()` is guarded and returns 0.0 here, so the mean component was
/// never the problem; the sigma component alone produced `(0.0, NaN)`.)
///
/// That is reachable in ordinary inference: once a pairing is more than
/// about nine cavity-sigma apart the truncation is a no-op, `trunc / cavity`
/// is exactly the identity message, and the chain compares one identity
/// against another. Before this guard that produced `(0.0, NaN)`, which
/// silently disabled the sigma half of the convergence test.
pub(crate) fn delta(&self, other: Gaussian) -> (f64, f64) {
if self.pi == other.pi && self.tau == other.tau {
return (0.0, 0.0);
}
(
(self.mu() - other.mu()).abs(),
(self.sigma() - other.sigma()).abs(),
@@ -145,13 +220,51 @@ impl Gaussian {
Self::from_mv(self.mu(), self.variance() + variance_delta)
}
/// `P(X < x)` under this Gaussian.
///
/// The question a stopping rule asks: *how sure am I that this competitor's
/// true skill is below the cutoff?* Expressing that as a probability keeps
/// its meaning as sigma changes, where a `mu + z * sigma` band silently
/// means different confidence at different uncertainties — which is exactly
/// the regime a stopping rule operates in.
///
/// Accurate in the *lower* tail. For the upper tail use
/// [`Gaussian::probability_above`] rather than `1.0 - probability_below(x)`,
/// which cancels away every significant digit once the result is small.
///
/// An improper Gaussian (non-positive precision) has no defined mean, so
/// this returns `0.5` — the same convention `mu()` and `sigma()` follow.
#[must_use]
pub fn probability_below(&self, x: f64) -> f64 {
if self.pi <= 0.0 {
return 0.5;
}
crate::cdf(x, self.mu(), self.sigma())
}
/// `P(X > x)` under this Gaussian.
///
/// Computed as a survival function rather than `1 - cdf`, so it keeps full
/// relative precision in the upper tail: `1 - cdf` returns exactly zero
/// past about 8.3 sigma, where the true value is still 1e-19 and perfectly
/// representable. A stopping rule is evaluated precisely there — the
/// interesting cases are the ones near certainty.
///
/// An improper Gaussian returns `0.5`, as [`Gaussian::probability_below`].
#[must_use]
pub fn probability_above(&self, x: f64) -> f64 {
if self.pi <= 0.0 {
return 0.5;
}
crate::sf(x, self.mu(), self.sigma())
}
/// EP damping in natural-parameter space: `α·new + (1−α)·self`.
///
/// Used by within-game inference to stabilise oscillating fixed-point
/// loops on hard graphs. `alpha = 1.0` returns `new` exactly;
/// `alpha < 1.0` shrinks each per-step update.
#[must_use]
pub fn damp_natural(self, new: Gaussian, alpha: f64) -> Gaussian {
pub(crate) fn damp_natural(self, new: Gaussian, alpha: f64) -> Gaussian {
Gaussian::from_natural(
alpha * new.pi() + (1.0 - alpha) * self.pi(),
alpha * new.tau() + (1.0 - alpha) * self.tau(),
@@ -165,34 +278,65 @@ impl Default for Gaussian {
}
}
impl ops::Add<Gaussian> for Gaussian {
type Output = Gaussian;
/// Variance addition: (mu1 + mu2, sqrt(σ1² + σ2²)).
/// Used for combining performance and noise; rare relative to mul/div.
fn add(self, rhs: Gaussian) -> Self::Output {
Self::from_mv(self.mu() + rhs.mu(), self.variance() + rhs.variance())
}
}
impl ops::Sub<Gaussian> for Gaussian {
type Output = Gaussian;
/// (mu1 - mu2, sqrt(σ1² + σ2²)). Same sigma combination as Add.
fn sub(self, rhs: Gaussian) -> Self::Output {
Self::from_mv(self.mu() - rhs.mu(), self.variance() + rhs.variance())
}
}
impl ops::Mul<Gaussian> for Gaussian {
type Output = Gaussian;
/// Factor product: nat-param add. Hot path — two f64 additions, no sqrt.
fn mul(self, rhs: Gaussian) -> Self::Output {
impl Gaussian {
/// The EP factor **product**: multiply two messages about the same
/// variable.
///
/// Two natural-parameter additions and no square root, which is why the
/// type stores `pi` and `tau` rather than `mu` and `sigma`. This is the
/// hot path.
///
/// Not arithmetic — `N(10, 2).ep_product(N(4, 3))` is `N(8.15, 1.66)`,
/// nowhere near 40. It used to be spelled `a * b`, on a public `Mul` impl,
/// where that was a trap rather than a shorthand.
#[inline]
pub(crate) fn ep_product(self, rhs: Gaussian) -> Gaussian {
Self::from_natural(self.pi + rhs.pi, self.tau + rhs.tau)
}
/// The EP **cavity**: divide out a message this belief already absorbed.
///
/// The inverse of [`ep_product`](Gaussian::ep_product), and two
/// subtractions rather than two additions.
///
/// **May return an improper result.** Cancelling a message that carried
/// most of the precision leaves `pi <= 0`, which is not a distribution.
/// `mu()` reports `0.0` and `sigma()` reports `inf` for such a value —
/// both are the accessors' policy for "undefined", not answers. Measured:
/// `N(10, 2).cavity(N(1, 1))` has `pi = -0.75`, and its `mu()` prints a
/// confident `0`. That is why this is not a public operator.
#[inline]
pub(crate) fn cavity(self, rhs: Gaussian) -> Gaussian {
Self::from_natural(self.pi - rhs.pi, self.tau - rhs.tau)
}
impl ops::Mul<f64> for Gaussian {
type Output = Gaussian;
fn mul(self, scalar: f64) -> Self::Output {
/// Convolve two independent Gaussians: `N(mu1 + mu2, sqrt(v1 + v2))`.
///
/// The distribution of a *sum* of independent variables, so the variances
/// add — the result is always wider than either input. Used to combine a
/// skill with performance noise. Goes through `from_mv` and takes no root.
#[inline]
pub(crate) fn convolve(self, rhs: Gaussian) -> Gaussian {
Self::from_mv(self.mu() + rhs.mu(), self.variance() + rhs.variance())
}
/// Convolve a *difference*: `N(mu1 - mu2, sqrt(v1 + v2))`.
///
/// The means subtract and the variances still **add**, because a
/// difference of independent variables is no more certain than a sum. That
/// is the half that made the old `Sub` impl misleading: `a - b` grew the
/// sigma from 2 to `sqrt(4 + 9)`.
#[inline]
pub(crate) fn convolve_diff(self, rhs: Gaussian) -> Gaussian {
Self::from_mv(self.mu() - rhs.mu(), self.variance() + rhs.variance())
}
/// Scale by a constant: `mu` by `scalar`, `sigma` by `|scalar|`.
///
/// The one operation that *is* ordinary arithmetic — it is the
/// distribution of `scalar * X`. Used for per-member weights.
#[inline]
pub(crate) fn scale(self, scalar: f64) -> Gaussian {
if !scalar.is_finite() {
return N_INF;
}
@@ -207,16 +351,44 @@ impl ops::Mul<f64> for Gaussian {
}
}
impl ops::Div<Gaussian> for Gaussian {
type Output = Gaussian;
/// Cavity: nat-param sub. Hot path — two f64 subtractions, no sqrt.
fn div(self, rhs: Gaussian) -> Self::Output {
Self::from_natural(self.pi - rhs.pi, self.tau - rhs.tau)
}
}
#[cfg(test)]
mod tests {
/// A message that did not change must report no change, even when it is
/// improper. `mu()` of an improper Gaussian is `0/0 = NaN` and `sigma()` is
/// infinite, so the mean/sigma form reported `(NaN, NaN)` for two identical
/// identity messages — which silently disabled the sigma half of the
/// convergence test in `run_chain`.
#[test]
fn delta_of_two_identical_improper_messages_is_zero() {
let improper = crate::N_INF;
// `mu()` is guarded and returns 0.0 for an improper Gaussian, so the
// mean component was always fine. The NaN came from the sigma
// component alone: `inf - inf`. The pre-fix value was `(0.0, NaN)`.
assert!(improper.sigma().is_infinite(), "premise: sigma is infinite");
assert_eq!(improper.mu(), 0.0, "premise: mu is guarded, not NaN");
assert!(
(improper.sigma() - improper.sigma()).is_nan(),
"premise: the unguarded sigma difference is NaN"
);
assert_eq!(improper.delta(improper), (0.0, 0.0));
}
#[test]
fn delta_of_identical_proper_messages_is_zero() {
let g = Gaussian::from_ms(25.0, 8.0);
assert_eq!(g.delta(g), (0.0, 0.0));
}
/// The shortcut must not swallow a real difference.
#[test]
fn delta_still_measures_a_real_move() {
let a = Gaussian::from_ms(25.0, 8.0);
let b = Gaussian::from_ms(26.0, 9.0);
let (dmu, dsigma) = a.delta(b);
assert!((dmu - 1.0).abs() < 1e-12, "{dmu}");
assert!((dsigma - 1.0).abs() < 1e-12, "{dsigma}");
}
use super::*;
#[test]
@@ -236,64 +408,64 @@ mod tests {
// Subtracting such a message must not produce NaN (the original failure path).
let proper = Gaussian::from_ms(9.75, 1.256);
let diff = proper - tiny_neg;
let diff = proper.convolve_diff(tiny_neg);
assert!(diff.pi().is_finite() && !diff.pi().is_nan());
assert!(diff.tau().is_finite() && !diff.tau().is_nan());
}
#[test]
fn test_add() {
fn convolve_adds_variances() {
let n = Gaussian::from_ms(25.0, 25.0 / 3.0);
let m = Gaussian::from_ms(0.0, 1.0);
let r = n + m;
let r = n.convolve(m);
assert!((r.mu() - 25.0).abs() < 1e-12);
assert!((r.sigma() - 8.393118874676116).abs() < 1e-10);
}
#[test]
fn test_sub() {
fn convolve_diff_subtracts_means_and_adds_variances() {
let n = Gaussian::from_ms(25.0, 25.0 / 3.0);
let m = Gaussian::from_ms(1.0, 1.0);
let r = n - m;
let r = n.convolve_diff(m);
assert!((r.mu() - 24.0).abs() < 1e-12);
assert!((r.sigma() - 8.393118874676116).abs() < 1e-10);
}
#[test]
fn test_mul() {
fn ep_product_is_not_arithmetic() {
let n = Gaussian::from_ms(25.0, 25.0 / 3.0);
let m = Gaussian::from_ms(0.0, 1.0);
let r = n * m;
let r = n.ep_product(m);
assert!((r.mu() - 0.35488958990536273).abs() < 1e-10);
assert!((r.sigma() - 0.992876838486922).abs() < 1e-10);
}
#[test]
fn test_div() {
fn cavity_undoes_a_product() {
let n = Gaussian::from_ms(25.0, 25.0 / 3.0);
let m = Gaussian::from_ms(0.0, 1.0);
let r = m / n;
let r = m.cavity(n);
assert!((r.mu() - (-0.3652597402597402)).abs() < 1e-10);
assert!((r.sigma() - 1.0072787050317253).abs() < 1e-10);
}
#[test]
fn test_n00_is_add_identity() {
// N00 (sigma=0) is the additive identity for the variance-convolution Add op.
// N_INF (sigma=inf) is the identity for the EP-product Mul op.
// N00 (sigma=0) is the identity for `convolve`.
// N_INF (sigma=inf) is the identity for `ep_product`.
let g = Gaussian::from_ms(3.0, 2.0);
let n00 = Gaussian::from_ms(0.0, 0.0);
let r = n00 + g;
let r = n00.convolve(g);
assert!((r.mu() - g.mu()).abs() < 1e-12);
assert!((r.sigma() - g.sigma()).abs() < 1e-12);
}
#[test]
fn test_mul_is_factor_product() {
// n * m in nat-params should be pi_n + pi_m, tau_n + tau_m
fn ep_product_adds_natural_parameters() {
// `ep_product` in nat-params should be pi_n + pi_m, tau_n + tau_m
let n = Gaussian::from_ms(2.0, 3.0);
let m = Gaussian::from_ms(1.0, 2.0);
let r = n * m;
let r = n.ep_product(m);
let expected_pi = n.pi() + m.pi();
let expected_tau = n.tau() + m.tau();
assert!((r.pi() - expected_pi).abs() < 1e-15);
@@ -301,10 +473,10 @@ mod tests {
}
#[test]
fn test_div_is_cavity() {
fn cavity_subtracts_natural_parameters() {
let n = Gaussian::from_ms(2.0, 1.0);
let m = Gaussian::from_ms(1.0, 2.0);
let r = n / m;
let r = n.cavity(m);
let expected_pi = n.pi() - m.pi();
let expected_tau = n.tau() - m.tau();
assert!((r.pi() - expected_pi).abs() < 1e-15);
@@ -340,3 +512,68 @@ mod tests {
assert!((damped.tau() - expected_tau).abs() < 1e-12);
}
}
#[cfg(test)]
mod tail_probability_tests {
use super::*;
#[test]
fn probability_below_matches_published_quantiles() {
let g = Gaussian::from_ms(0.0, 1.0);
for (x, expected) in [
(-1.959_963_984_540_054, 0.025),
(0.0, 0.5),
(1.281_551_565_544_6, 0.9),
(1.959_963_984_540_054, 0.975),
] {
let got = g.probability_below(x);
assert!(
(got - expected).abs() < 1e-12,
"P(X < {x}) = {got}, expected {expected}"
);
}
}
#[test]
fn the_two_tails_partition_the_mass() {
let g = Gaussian::from_ms(3.0, 2.0);
for x in [-4.0f64, 0.0, 3.0, 7.5] {
let total = g.probability_below(x) + g.probability_above(x);
assert!((total - 1.0).abs() < 1e-15, "at {x}: {total}");
}
}
/// The reason `probability_above` exists rather than `1 - probability_below`.
#[test]
fn probability_above_keeps_precision_where_the_complement_collapses() {
let g = Gaussian::from_ms(0.0, 1.0);
for (x, expected) in [(9.0f64, 1.128_588e-19), (20.0, 2.753_624e-89)] {
let got = g.probability_above(x);
assert!(
(got - expected).abs() / expected < 1e-6,
"P(X > {x}) = {got}, expected ~{expected}"
);
assert_eq!(
1.0 - g.probability_below(x),
0.0,
"the complement should still collapse at {x}"
);
}
}
#[test]
fn a_scaled_gaussian_shifts_and_stretches() {
let g = Gaussian::from_ms(25.0, 6.0);
assert!((g.probability_below(25.0) - 0.5).abs() < 1e-15);
// One sigma either side of the mean.
assert!((g.probability_below(31.0) - 0.841_344_746_068_543).abs() < 1e-12);
assert!((g.probability_above(19.0) - 0.841_344_746_068_543).abs() < 1e-12);
}
#[test]
fn an_improper_gaussian_is_uninformative_rather_than_nan() {
let improper = Gaussian::from_ms(0.0, f64::INFINITY);
assert_eq!(improper.probability_below(5.0), 0.5);
assert_eq!(improper.probability_above(5.0), 0.5);
}
}
-20
View File
@@ -1,20 +0,0 @@
//! Factor-graph public API.
//!
//! Named `graph` rather than `factors` because the private implementation
//! module beside it is `factor`: two module paths differing by one character,
//! one public and one not, was a standing invitation to import the wrong one.
//!
//! The factor types, `VarStore` and the `Schedule` trait are public so custom
//! schedules can be written against them.
//!
//! Building a factor graph by hand goes through `Game::custom`, which is
//! deliberately `#[doc(hidden)]`: it works, but its signature is not yet
//! considered stable API and so is not listed in these docs.
pub use crate::{
factor::{
BuiltinFactor, Factor, VarId, VarStore, margin::MarginFactor, rank_diff::RankDiffFactor,
team_sum::TeamSumFactor, trunc::TruncFactor,
},
schedule::{EpsilonOrMax, Schedule, ScheduleReport},
};
+1976 -261
View File
File diff suppressed because it is too large Load Diff
+510
View File
@@ -0,0 +1,510 @@
//! Sparse Cholesky factorisation of a joint precision matrix.
//!
//! Every question the joint answers is a *bilinear form* in the precision
//! matrix's inverse — the variance of a contrast is `c^T L^-1 c`, and the
//! covariance of two contrasts is `c^T L^-1 a`. None of them wants `L^-1 c`
//! itself, which is what makes the shape here worth stating explicitly.
//!
//! Writing the precision as `A = L L^T`,
//!
//! ```text
//! c^T A^-1 a = c^T L^-T L^-1 a = (L^-1 c) . (L^-1 a)
//! ```
//!
//! so a single forward substitution per contrast answers everything, and the
//! back substitution a general solve would do is wasted work. That halves the
//! cost of a query, and it removes a failure mode: a variance computed as
//! `c . (A^-1 c)` is a difference of products that can round to a small
//! negative number, where the same quantity as `|L^-1 c|^2` is a sum of
//! squares and cannot.
//!
//! # Why this is sparse (#52)
//!
//! A time-expanded joint is *extremely* sparse and gets sparser as the history
//! grows: a row couples only to its own previous and next appearance through
//! the drift link, and to whoever co-appeared in its slice. Measured on a
//! 76-slice, 988-duel, 200-competitor history: `n = 1976`, `nnz = 7504`,
//! **0.19% dense**.
//!
//! This used to store all `n^2` entries and run a dense `O(n^3)` factorisation
//! over them. Two measurements decided the replacement:
//!
//! - **Ordering alone does nothing to a dense factorisation.** Its inner loops
//! run over every `k` whether or not the entry is zero. A 700x700 banded
//! matrix at 0.43% density factorised in 30.196 ms in band order and
//! 29.544 ms under a scramble that destroyed the band — identical, as the
//! flop count says it must be. Fill-reducing order is worth nothing until
//! the factorisation skips zeros.
//! - **Together they are worth four orders of magnitude.** On that `n = 1976`
//! fixture, against `n^3/3 = 2.572e9` flops dense: sparse in the natural
//! order needs `5.597e7` (46x better), and sparse under an AMD fill-reducing
//! order needs `8.656e4` — **29,710x**. AMD is worth 646x *on top of*
//! sparsity and nothing without it.
//!
//! Natural ordering fills in badly here for the reason #52 predicted: a
//! competitor who appears in slice 0 and not again until slice 75 creates a
//! drift link spanning nearly the whole matrix. `nnz(L)` is 292,437 under the
//! natural order against 11,583 under AMD, from an `A` with 7,504.
//!
//! The ordering comes from `feral-amd`. The factorisation is the up-looking
//! sparse Cholesky of Davis's *Direct Methods for Sparse Linear Systems*,
//! written here rather than taken from a crate: the sparse solvers on
//! crates.io either pull SIMD dispatch (`faer`, and `feral` itself, both
//! through `pulp`), which would make results differ between an AVX-512 host
//! and an AVX2 one — the same class of drift the `libm`-over-`std` decision
//! was made to avoid — or are LGPL, or disclaim fill-reduction in their own
//! docs.
use std::collections::BTreeMap;
/// A symmetric matrix accumulated entry by entry, before factorisation.
///
/// A `BTreeMap` rather than a hash map because the iteration order becomes the
/// factorisation's summation order, and a hash map's order varies per process.
/// `tests/cross_process_determinism.rs` exists because that has bitten before.
#[derive(Default)]
pub(crate) struct SymmetricBuilder {
entries: BTreeMap<(usize, usize), f64>,
}
impl SymmetricBuilder {
pub(crate) fn new() -> Self {
Self::default()
}
/// Add `value` to entry `(row, col)`. Both triangles must be supplied.
pub(crate) fn add(&mut self, row: usize, col: usize, value: f64) {
*self.entries.entry((row, col)).or_insert(0.0) += value;
}
/// The `(row, col)` positions that hold a nonzero. For the #52 measurement.
#[cfg(feature = "measure-sparsity")]
pub(crate) fn pattern(&self) -> impl Iterator<Item = (usize, usize)> + '_ {
self.entries
.iter()
.filter(|(_, v)| **v != 0.0)
.map(|(&rc, _)| rc)
}
}
/// A factorised symmetric positive-definite matrix, reusable across queries.
pub(crate) struct Cholesky {
n: usize,
/// `inv[old] = new`: where each original row sits after the AMD reorder.
inv: Vec<usize>,
/// `L` in compressed-column form, permuted. Within a column the diagonal
/// is first and the rest ascend by row.
col_ptr: Vec<usize>,
row_idx: Vec<usize>,
val: Vec<f64>,
}
impl Cholesky {
/// Factorise the accumulated matrix into `L L^T`, under a fill-reducing
/// permutation.
///
/// Returns `None` if the matrix is not positive-definite, which for a
/// precision matrix means the model is improper — a competitor with
/// neither a proper prior nor any evidence — or if the ordering fails.
pub(crate) fn factor(built: SymmetricBuilder, n: usize) -> Option<Self> {
if n == 0 {
return Some(Self {
n: 0,
inv: Vec::new(),
col_ptr: vec![0],
row_idx: Vec::new(),
val: Vec::new(),
});
}
let inv = Self::amd_permutation(n, &built)?;
// Upper triangle of the permuted matrix, column-major: column `c`
// holds the rows `r <= c`. Exactly one of a symmetric pair survives
// the `r <= c` filter, so nothing is double-counted.
let mut cols: Vec<Vec<(usize, f64)>> = vec![Vec::new(); n];
for (&(old_r, old_c), &v) in &built.entries {
if v == 0.0 {
continue;
}
let (r, c) = (inv[old_r], inv[old_c]);
if r <= c {
cols[c].push((r, v));
}
}
let mut a_ptr = Vec::with_capacity(n + 1);
let mut a_row = Vec::new();
let mut a_val = Vec::new();
a_ptr.push(0usize);
for col in &mut cols {
col.sort_unstable_by_key(|&(r, _)| r);
for &(r, v) in col.iter() {
a_row.push(r);
a_val.push(v);
}
a_ptr.push(a_row.len());
}
let parent = Self::etree(n, &a_ptr, &a_row);
// Symbolic pass: how many entries each column of L will hold. Running
// `ereach` per column costs O(nnz(L)) in total, which is the same order
// as the numeric pass it sizes.
let mut counts = vec![0usize; n];
let mut stack = vec![0usize; n];
let mut mark = vec![false; n];
for k in 0..n {
let top = Self::ereach(k, &a_ptr, &a_row, &parent, &mut stack, &mut mark);
for &i in &stack[top..] {
counts[i] += 1;
}
counts[k] += 1; // the diagonal
}
let mut col_ptr = Vec::with_capacity(n + 1);
col_ptr.push(0usize);
for &c in &counts {
col_ptr.push(col_ptr[col_ptr.len() - 1] + c);
}
let nnz = col_ptr[n];
let mut row_idx = vec![0usize; nnz];
let mut val = vec![0.0f64; nnz];
// `next[i]` is the slot column `i` will fill next. Column `i`'s
// diagonal lands first, at `col_ptr[i]`, because nothing is written to
// a column before its own iteration.
let mut next: Vec<usize> = col_ptr[..n].to_vec();
let mut x = vec![0.0f64; n];
for k in 0..n {
let top = Self::ereach(k, &a_ptr, &a_row, &parent, &mut stack, &mut mark);
for p in a_ptr[k]..a_ptr[k + 1] {
if a_row[p] <= k {
x[a_row[p]] = a_val[p];
}
}
let mut d = x[k];
x[k] = 0.0;
for &i in &stack[top..] {
let lki = x[i] / val[col_ptr[i]];
x[i] = 0.0;
for p in col_ptr[i] + 1..next[i] {
x[row_idx[p]] -= val[p] * lki;
}
d -= lki * lki;
let p = next[i];
next[i] += 1;
row_idx[p] = k;
val[p] = lki;
}
// Explicit rather than `!(d > 0.0)`: a NaN pivot must fail here
// too, and a negated comparison would let it through as "not
// positive".
if d.is_nan() || d <= 0.0 {
return None;
}
let p = next[k];
next[k] += 1;
row_idx[p] = k;
val[p] = d.sqrt();
}
Some(Self {
n,
inv,
col_ptr,
row_idx,
val,
})
}
/// AMD fill-reducing order, as `inv[old] = new`.
fn amd_permutation(n: usize, built: &SymmetricBuilder) -> Option<Vec<usize>> {
let mut cols: Vec<Vec<i32>> = vec![Vec::new(); n];
for (&(r, c), &v) in &built.entries {
if v != 0.0 {
cols[c].push(i32::try_from(r).ok()?);
}
}
let mut col_ptr = Vec::with_capacity(n + 1);
let mut row_idx = Vec::new();
col_ptr.push(0i32);
for (j, col) in cols.iter_mut().enumerate() {
col.push(i32::try_from(j).ok()?);
col.sort_unstable();
col.dedup();
row_idx.extend_from_slice(col);
col_ptr.push(i32::try_from(row_idx.len()).ok()?);
}
let pattern = feral_amd::CscPattern::new(n, &col_ptr, &row_idx)?;
// `perm[new] = old`; we want the inverse.
let perm = feral_amd::amd_order(&pattern).ok()?;
let mut inv = vec![0usize; n];
for (new, &old) in perm.iter().enumerate() {
inv[usize::try_from(old).ok()?] = new;
}
Some(inv)
}
/// Elimination tree of the upper-triangular pattern. `usize::MAX` is "no
/// parent", i.e. a root.
fn etree(n: usize, col_ptr: &[usize], row_idx: &[usize]) -> Vec<usize> {
let mut parent = vec![usize::MAX; n];
let mut ancestor = vec![usize::MAX; n];
for k in 0..n {
for &row in &row_idx[col_ptr[k]..col_ptr[k + 1]] {
let mut i = row;
while i != usize::MAX && i < k {
let next = ancestor[i];
ancestor[i] = k;
if next == usize::MAX {
parent[i] = k;
}
i = next;
}
}
}
parent
}
/// Nonzero pattern of row `k` of `L`, written into `stack[top..n]` in
/// topological order. Returns `top`.
///
/// `stack` is used from both ends — a scratch region from `0` while walking
/// each path up the tree, and the result from `n` downwards. They cannot
/// collide because every node is pushed at most once across the whole call.
fn ereach(
k: usize,
col_ptr: &[usize],
row_idx: &[usize],
parent: &[usize],
stack: &mut [usize],
mark: &mut [bool],
) -> usize {
let n = mark.len();
let mut top = n;
mark[k] = true;
for &row in &row_idx[col_ptr[k]..col_ptr[k + 1]] {
let mut i = row;
if i > k {
continue;
}
let mut len = 0usize;
while i != usize::MAX && !mark[i] {
stack[len] = i;
len += 1;
mark[i] = true;
i = parent[i];
}
// Reverse the path onto the output end, so the result stays in
// topological order overall.
while len > 0 {
len -= 1;
top -= 1;
stack[top] = stack[len];
}
}
for &i in &stack[top..] {
mark[i] = false;
}
mark[k] = false;
top
}
/// Whiten a contrast: `y = L^-1 P b`.
///
/// The point of the result is the dot product, not the vector: for two
/// contrasts `b` and `b'`, `y . y'` is `b^T A^-1 b'`. See the module docs.
///
/// The result is in the permuted order, and stays there — a dot product
/// does not care, as long as both operands were permuted the same way.
pub(crate) fn whiten(&self, b: &[f64]) -> Vec<f64> {
debug_assert_eq!(b.len(), self.n);
let n = self.n;
let mut y = vec![0.0f64; n];
for (old, &v) in b.iter().enumerate() {
y[self.inv[old]] = v;
}
for j in 0..n {
y[j] /= self.val[self.col_ptr[j]];
let yj = y[j];
for p in self.col_ptr[j] + 1..self.col_ptr[j + 1] {
y[self.row_idx[p]] -= self.val[p] * yj;
}
}
y
}
}
/// `b^T A^-1 b'`, given the two whitened contrasts.
pub(crate) fn bilinear(y: &[f64], y_prime: &[f64]) -> f64 {
y.iter().zip(y_prime).map(|(a, b)| a * b).sum()
}
#[cfg(test)]
mod tests {
use super::*;
/// Factorise a dense row-major matrix, for the goldens below.
fn dense(a: &[f64], n: usize) -> Option<Cholesky> {
let mut b = SymmetricBuilder::new();
for i in 0..n {
for j in 0..n {
if a[i * n + j] != 0.0 {
b.add(i, j, a[i * n + j]);
}
}
}
Cholesky::factor(b, n)
}
/// `[[4, 1], [1, 3]] z = [1, 2]` has `z = [1/11, 7/11]`, so the quadratic
/// form `b^T A^-1 b` is `1 * 1/11 + 2 * 7/11 = 15/11`.
#[test]
fn reproduces_a_known_quadratic_form() {
let c = dense(&[4.0, 1.0, 1.0, 3.0], 2).unwrap();
let y = c.whiten(&[1.0, 2.0]);
assert!((bilinear(&y, &y) - 15.0 / 11.0).abs() < 1e-12);
}
/// Whitening `e_i` recovers the inverse's diagonal, which is the variance
/// of a single variable.
#[test]
fn recovers_the_inverse_diagonal() {
// A = [[2, -1, 0], [-1, 2, -1], [0, -1, 2]]; inverse diagonal is
// [0.75, 1.0, 0.75].
let a = [2.0, -1.0, 0.0, -1.0, 2.0, -1.0, 0.0, -1.0, 2.0];
let c = dense(&a, 3).unwrap();
for (i, expected) in [0.75, 1.0, 0.75].into_iter().enumerate() {
let mut e = vec![0.0; 3];
e[i] = 1.0;
let y = c.whiten(&e);
assert!((bilinear(&y, &y) - expected).abs() < 1e-12, "row {i}");
}
}
/// The off-diagonal bilinear form is symmetric and matches the inverse.
#[test]
fn recovers_an_off_diagonal_covariance() {
// Same A; (A^-1)_{0,1} = 0.5.
let a = [2.0, -1.0, 0.0, -1.0, 2.0, -1.0, 0.0, -1.0, 2.0];
let c = dense(&a, 3).unwrap();
let y0 = c.whiten(&[1.0, 0.0, 0.0]);
let y1 = c.whiten(&[0.0, 1.0, 0.0]);
assert!((bilinear(&y0, &y1) - 0.5).abs() < 1e-12);
assert!((bilinear(&y1, &y0) - 0.5).abs() < 1e-12);
}
/// A variance can never come out negative, because it is a sum of squares.
#[test]
fn a_quadratic_form_is_never_negative() {
let a = [1e12, 1e12 - 1.0, 1e12 - 1.0, 1e12];
let c = dense(&a, 2).unwrap();
let y = c.whiten(&[1.0, -1.0]);
assert!(bilinear(&y, &y) >= 0.0);
}
/// Against an independent dense reference, on random sparse SPD matrices.
///
/// The goldens above are 2x2 and 3x3 — small enough that AMD does nothing
/// and no fill-in occurs, so they cannot catch a symbolic-pass bug. This
/// builds matrices big enough to permute and fill in, and checks every
/// bilinear form against a textbook dense factorisation of the *same*
/// matrix in its original order.
#[test]
fn agrees_with_a_dense_reference_on_random_sparse_systems() {
/// Dense Cholesky and quadratic form, deliberately naive: this is the
/// reference, so it must not share code with what it is checking.
fn dense_quadratic_form(a: &[f64], n: usize, b: &[f64], c: &[f64]) -> f64 {
let mut l = a.to_vec();
for j in 0..n {
let mut d = l[j * n + j];
for k in 0..j {
d -= l[j * n + k] * l[j * n + k];
}
let d = d.sqrt();
l[j * n + j] = d;
for i in j + 1..n {
let mut sum = l[i * n + j];
for k in 0..j {
sum -= l[i * n + k] * l[j * n + k];
}
l[i * n + j] = sum / d;
}
}
let solve = |rhs: &[f64]| -> Vec<f64> {
let mut y = rhs.to_vec();
for i in 0..n {
for k in 0..i {
y[i] -= l[i * n + k] * y[k];
}
y[i] /= l[i * n + i];
}
y
};
let (yb, yc) = (solve(b), solve(c));
yb.iter().zip(&yc).map(|(x, y)| x * y).sum()
}
// A cheap deterministic generator; no dependency, and reproducible.
let mut seed = 0x2545_F491_4F6C_DD1Du64;
let mut rand = move || {
seed ^= seed << 13;
seed ^= seed >> 7;
seed ^= seed << 17;
(seed >> 11) as f64 / (1u64 << 53) as f64
};
for n in [7usize, 23, 60] {
let mut a = vec![0.0f64; n * n];
// A chain plus scattered long-range couplings: the shape of a
// time-expanded joint, where a competitor's drift link can span
// the whole matrix.
for i in 0..n {
a[i * n + i] = 4.0 + rand();
if i + 1 < n {
let v = -(0.5 + rand() * 0.5);
a[i * n + i + 1] = v;
a[(i + 1) * n + i] = v;
}
}
for step in 0..n / 3 {
let i = (step * 7) % n;
let j = (step * 29 + 3) % n;
if i != j {
let v = -(0.1 + rand() * 0.2);
a[i * n + j] = v;
a[j * n + i] = v;
// Keep it diagonally dominant, hence positive-definite.
a[i * n + i] += 0.6;
a[j * n + j] += 0.6;
}
}
let sparse = dense(&a, n).expect("spd");
for trial in 0..8 {
let b: Vec<f64> = (0..n).map(|_| rand() * 2.0 - 1.0).collect();
let c: Vec<f64> = (0..n).map(|_| rand() * 2.0 - 1.0).collect();
let got = bilinear(&sparse.whiten(&b), &sparse.whiten(&c));
let want = dense_quadratic_form(&a, n, &b, &c);
assert!(
(got - want).abs() <= 1e-10 * want.abs().max(1.0),
"n={n} trial={trial}: sparse {got} vs dense {want}"
);
}
}
}
/// A permutation must not change which matrices are rejected.
#[test]
fn rejects_a_non_positive_definite_matrix() {
// Singular: the second row is a multiple of the first.
assert!(dense(&[1.0, 2.0, 2.0, 4.0], 2).is_none());
}
}
+16 -12
View File
@@ -26,21 +26,24 @@ where
K: Eq + Hash + Clone,
{
#[must_use]
pub fn new() -> Self {
pub(crate) fn new() -> Self {
Self {
forward: HashMap::new(),
reverse: Vec::new(),
}
}
pub fn get<Q: ?Sized + Hash + Eq>(&self, k: &Q) -> Option<Index>
pub(crate) fn get<Q: ?Sized + Hash + Eq>(&self, k: &Q) -> Option<Index>
where
K: Borrow<Q>,
{
self.forward.get(k).cloned()
}
pub fn get_or_create<Q: ?Sized + Hash + Eq + ToOwned<Owned = K>>(&mut self, k: &Q) -> Index
pub(crate) fn get_or_create<Q: ?Sized + Hash + Eq + ToOwned<Owned = K>>(
&mut self,
k: &Q,
) -> Index
where
K: Borrow<Q>,
{
@@ -56,23 +59,24 @@ where
}
#[must_use]
pub fn key(&self, idx: Index) -> Option<&K> {
pub(crate) fn key(&self, idx: Index) -> Option<&K> {
self.reverse.get(idx.0)
}
pub fn keys(&self) -> impl Iterator<Item = &K> {
self.forward.keys()
/// Every key, in the order they were first interned.
///
/// Iterates the dense reverse table rather than the forward `HashMap`.
/// Rust seeds its default hasher per process, so a `HashMap` walk yields a
/// different order on every run — which is fine for membership but not for
/// anything a caller might sum, sort or print.
pub(crate) fn keys(&self) -> impl ExactSizeIterator<Item = &K> {
self.reverse.iter()
}
#[must_use]
pub fn len(&self) -> usize {
pub(crate) fn len(&self) -> usize {
self.reverse.len()
}
#[must_use]
pub fn is_empty(&self) -> bool {
self.reverse.is_empty()
}
}
impl<K> Default for KeyTable<K>
+395 -48
View File
@@ -85,6 +85,10 @@
//! regardless of worker count.
#![forbid(unsafe_code)]
// Turned on once the surface was fully documented (80 items at the time), so
// the next undocumented public item is a build failure rather than a warning
// nobody reads.
#![deny(missing_docs)]
/// Compiles every `rust` block in `README.md` as a doctest.
///
@@ -104,25 +108,35 @@ use std::{
f64::consts::{FRAC_1_SQRT_2, FRAC_2_SQRT_PI, SQRT_2},
};
mod acquisition;
#[cfg(feature = "approx")]
mod approx;
pub(crate) mod arena;
mod time;
mod time_slice;
pub use time_slice::{EventKind, TimeSlice};
mod acquisition;
mod color_group;
mod competitor;
mod convergence;
/// Skill drift: how much a competitor's skill is allowed to move between
/// appearances.
///
/// Public because [`Drift`] is a trait a caller may implement — a per-sport
/// off-season, say, or a schedule where drift is a function of the calendar
/// rather than of elapsed ticks. [`ConstantDrift`] is what
/// [`HistoryBuilder`] uses by default.
pub mod drift;
mod error;
mod event;
mod event_builder;
pub(crate) mod factor;
mod game;
/// The Gaussian message type and its expectation-propagation algebra.
///
/// Public because [`Gaussian`] appears throughout the results: a posterior
/// skill, a learning-curve point, a predicted margin. The module carries the
/// operator documentation — `Mul`/`Div` are the EP product and cavity, not
/// arithmetic on random variables.
pub mod gaussian;
pub mod graph;
mod history;
mod joint;
mod key_table;
mod matrix;
mod observer;
@@ -130,35 +144,118 @@ mod outcome;
mod predict;
pub(crate) mod quadrature;
mod rating;
pub(crate) mod schedule;
pub mod storage;
pub mod rating_rule;
pub(crate) mod storage;
mod time;
mod time_slice;
pub use acquisition::expected_information_gain;
pub use competitor::Competitor;
pub use convergence::{ConvergenceOptions, ConvergenceReport};
pub use drift::{ConstantDrift, Drift};
pub use error::InferenceError;
pub use error::{CompetitorField, InferenceError, OutcomeKind, Parameter, Shape, UnknownKeys};
pub use event::{Event, Member, Team};
pub use event_builder::EventBuilder;
pub use game::{Game, GameOptions, OwnedGame};
pub use game::{Game, GameOptions};
pub use gaussian::Gaussian;
pub use history::{History, HistoryBuilder};
pub use key_table::KeyTable;
pub use history::{History, HistoryBuilder, Joint};
use matrix::Matrix;
pub use observer::{NullObserver, Observer};
pub use outcome::Outcome;
pub use predict::Prediction;
pub use rating::Rating;
pub use schedule::ScheduleReport;
pub use rating_rule::{FnRule, NoRule, RatingRule, StartingPoint};
/// The `smallvec` crate, re-exported.
///
/// Four public items name `SmallVec` in their signatures: [`Event::teams`],
/// [`Team::members`], [`Outcome::Ranked`]'s payload and
/// [`ConvergenceReport::per_iteration_time`]. You can *build* an `Event`
/// without ever naming the type — `vec![..].into()` and `.collect()` both work
/// — and iterate the timings through `Deref`. But writing a helper that
/// *returns* a teams list, or a `match` arm that binds ranks and passes them
/// on, requires the type by name.
///
/// Measured: the only `Joint` doc example failed to compile from a consumer
/// crate with `unresolved import \`smallvec\``, because the dependency was in
/// the signature but not reachable. Re-exported so a consumer takes this
/// crate's version rather than pinning a matching one of their own.
pub use smallvec;
pub use time::{Time, Untimed};
/// Default performance noise: how much a single showing varies around skill.
///
/// Every other default is expressed as a multiple of this, so `BETA` sets the
/// scale of the whole rating system. Doubling it and doubling `SIGMA` and
/// `GAMMA` with it gives the same fit on a rescaled axis.
pub const BETA: f64 = 1.0;
/// Default prior mean skill.
///
/// Zero rather than a conventional 25: the scale is set by `BETA`, and a
/// centred axis makes a negative rating mean "below the prior" instead of
/// looking like an error.
pub const MU: f64 = 0.0;
/// Default prior standard deviation: how unsure the model starts out.
///
/// Six betas is deliberately wide — a new competitor's first result should
/// move them a long way, and the prior should not fight the evidence.
pub const SIGMA: f64 = BETA * 6.0;
/// Default drift: the standard deviation of skill movement per unit of time.
///
/// Enters inference as a *variance* (`gamma^2` per elapsed tick), which is why
/// [`ConstantDrift`] squares it and why a negative gamma would be
/// indistinguishable from its absolute value — see [`ConstantDrift::new`].
pub const GAMMA: f64 = BETA * 0.03;
/// Default draw probability: zero, meaning ties are not modelled.
///
/// A history that ingests a tie needs a positive value. With `p_draw == 0.0`
/// the truncation margin is zero and the two-sided tie update evaluates
/// `0/0`, so ingestion rejects such events with
/// [`InferenceError::TieWithoutDrawProbability`].
pub const P_DRAW: f64 = 0.0;
/// Default convergence threshold, in the same units as
/// [`ConvergenceReport::final_step`](crate::ConvergenceReport).
///
/// The sweep stops once the largest change a full iteration makes to any
/// message falls below this.
pub const EPSILON: f64 = 1e-6;
pub const ITERATIONS: usize = 30;
/// Default cap on convergence sweeps.
///
/// **A runaway guard, not a budget.** The sweep exits as soon as the step falls
/// below `epsilon`, so the cap is never reached by a history that converges and
/// raising it costs nothing. Measured on a history that needs four sweeps:
///
/// ```text
/// max_iter 30: 4 iterations, 129.9 us
/// max_iter 100_000: 4 iterations, 131.9 us
/// ```
///
/// This was `30` until it was measured, and 30 truncated ordinary healthy
/// histories: 160 events over 100 competitors already needs 42. Because a short
/// fit is finite and sensibly ordered, that was invisible.
///
/// # Why it is not scaled to the history
///
/// The obvious improvement — pick the cap from the node or event count — does
/// not work, because iteration count is driven by how *loopy* the graph is
/// rather than how big it is. At a fixed 320 events over 40 slices, varying
/// only the number of competitors sharing them:
///
/// ```text
/// competitors appearances each iterations
/// 3 213 2_789
/// 10 64 1_068
/// 50 12.8 206
/// 100 6.4 90
/// 400 1.6 2
/// ```
///
/// Three orders of magnitude apart on identical event and slice counts. Any
/// formula in those two numbers would be badly wrong on some real shape, so the
/// cap is a single value set high enough that reaching it means the fit is
/// oscillating rather than merely large.
///
/// Reaching it is [`InferenceError::NotConverged`]. See
/// [`History::converge`](crate::History::converge).
pub const ITERATIONS: usize = 10_000;
/// Largest team count `History::predict_outcome` will enumerate.
///
@@ -182,22 +279,35 @@ const HALF_LINE_WINDOW: f64 = 10.0;
/// four-term series is good to ~1e-10 by here, so the two are at their closest
/// agreement around this point. Below it the subtraction is exact; above it the
/// series is.
/// `alpha / width` past which the tie branch's `v^2 - u` has lost too many
/// digits to trust, and the narrow-window form takes over.
///
/// The subtraction retains about `(width / alpha)^2 / EPSILON` of its
/// precision, so this is the ratio at which that falls below roughly 1e-6.
const NARROW_WINDOW_RATIO: f64 = 2.0e4;
const ASYMPTOTIC_MILLS_ALPHA: f64 = 100.0;
pub const N01: Gaussian = Gaussian::from_ms(0.0, 1.0);
pub const N00: Gaussian = Gaussian::from_ms(0.0, 0.0);
pub const N_INF: Gaussian = Gaussian::from_ms(0.0, f64::INFINITY);
pub(crate) const N00: Gaussian = Gaussian::from_ms(0.0, 0.0);
pub(crate) const N_INF: Gaussian = Gaussian::from_ms(0.0, f64::INFINITY);
/// An interned competitor handle: a dense slot number, not a user key.
///
/// `History` stores skills and messages by `Index` rather than by `K`, so the
/// hot path never hashes a key. Indices are assigned in interning order and
/// are stable for the life of a history; they are not portable between
/// histories, since the same key interns to a different slot under a different
/// ingestion order.
///
/// Crate-internal. It was public, along with `History::intern` and
/// `History::lookup` that produced one — and **nothing public ever accepted
/// one**, so it was a handle with nowhere to go. It also shadowed
/// `std::ops::Index`, which `CompetitorStore` implements. See #73.
#[derive(Copy, Clone, Default, PartialEq, PartialOrd, Eq, Ord, Hash, Debug)]
pub struct Index(usize);
pub(crate) struct Index(usize);
impl Index {
/// The underlying slot number.
///
/// Indices are dense and assigned in interning order, so this is usable as
/// a key into a caller-side side table.
#[must_use]
pub fn get(self) -> usize {
pub(crate) fn get(self) -> usize {
self.0
}
}
@@ -440,10 +550,72 @@ fn pdf(x: f64, mu: f64, sigma: f64) -> f64 {
fn half_line_truncation(alpha: f64) -> (f64, f64) {
let inv = alpha.recip();
let inv_sq = inv * inv;
let gap = inv * (1.0 - inv_sq * (2.0 - inv_sq * (10.0 - 74.0 * inv_sq)));
let b = 2.0 - inv_sq * (10.0 - 74.0 * inv_sq);
let gap = inv * (1.0 - inv_sq * b);
let v = alpha + gap;
(v, v * gap)
// Returns `1 - w`, not `w`, and that is the whole point of this shape.
//
// `w` tends to 1 out here, so a caller forming `1 - w` loses about
// `log10(alpha^2)` digits: measured against the exact truncated variance,
// `1 - w` came back with 8.9e-5 relative error at alpha = 1e6 and **0.0**
// from alpha = 1e8 — where the true value is 1e-16 and perfectly
// representable. `sigma * (1 - w).sqrt()` was then exactly zero, and
// `from_ms(mu, 0.0)` is a point mass whose `mu()` is `inf/inf = NaN`.
//
// Expanding `1 - v*gap` symbolically removes the subtraction: with
// `alpha*gap = 1 - inv^2*b`, the leading ones cancel on paper instead of in
// floating point, leaving `inv^2` times a bracket that tends to 1. Measured
// exact — 0.0 relative error — from alpha = 1e3 to 1e8.
let one_minus_w = inv_sq
* ((1.0 - inv_sq * (10.0 - 74.0 * inv_sq)) + 2.0 * inv_sq * b - inv_sq * inv_sq * b * b);
(v, one_minus_w)
}
/// Truncation to a *narrow* window `[alpha, alpha + d]`, as `(v, 1 - w)`.
///
/// The tie branch forms `w` from `v^2 - u`, and both grow as `alpha^2` while
/// their difference stays `O(1)`. Far enough into the tail that subtraction has
/// nothing left: measured at `alpha = 1e6` with a window of `1e-6` it kept four
/// significant digits and returned `1 - w = -2.4e-4` where the truth is
/// `+2.8e-13`, so `sqrt` of it was NaN. One step earlier it was quietly wrong
/// instead — `1 - w = 1.0` exactly, a truncation reported as a no-op, where the
/// truth was `5e-17`.
///
/// The existing half-line escape hatch does not cover it, because that keys on
/// `alpha * d >= HALF_LINE_WINDOW` — how many window-widths from the mean the
/// window sits — and a *narrow* window fails that however deep it is.
///
/// Over a narrow window the density is `exp(-t*s - s^2 d^2 / 2)` in
/// `x = alpha + s*d`, with `t = alpha * d`. Dropping the `d^2` term leaves a
/// truncated exponential on `[0, 1]`, whose mean and variance are closed forms.
/// So `v = alpha + d*m(t)` and `1 - w = d^2 * V(t)`, with no subtraction of
/// large quantities anywhere.
///
/// Measured against high-precision quadrature over `alpha` in `[1e2, 1e9]`:
/// `v` exact to 4e-10 or better, `1 - w` to 4e-10 across the region this is
/// used in.
fn narrow_window_truncation(alpha: f64, d: f64) -> (f64, f64) {
let t = alpha * d;
// `m` and `V` are the mean and variance of a truncated exponential on
// [0, 1] with rate `t`, both of which cancel as `t -> 0`. The series is
// their limit (1/2 and 1/12, a uniform window) with the leading correction.
let (m, v_s) = if t < 1e-3 {
(
0.5 - t / 12.0 + t * t * t / 720.0,
1.0 / 12.0 - t * t / 240.0,
)
} else {
let em1 = libm::expm1(t);
(
1.0 / t - 1.0 / em1,
1.0 / (t * t) - (em1 + 1.0) / (em1 * em1),
)
};
(alpha + d * m, d * d * v_s)
}
fn v_w(mu: f64, sigma: f64, margin: f64, tie: bool) -> (f64, f64) {
@@ -471,7 +643,7 @@ fn v_w(mu: f64, sigma: f64, margin: f64, tie: bool) -> (f64, f64) {
(v, v - alpha)
};
(v, v * gap)
(v, 1.0 - v * gap)
} else {
// v is odd in mu and w is even, so fold to mu <= 0. Both truncation
// points then sit in the upper tail, where the scaled form applies.
@@ -487,9 +659,22 @@ fn v_w(mu: f64, sigma: f64, margin: f64, tie: bool) -> (f64, f64) {
// Once the window sits many of its own widths into the tail it is
// indistinguishable from a half-line, so the asymptotic covers it with
// no subtraction at all.
if alpha >= ASYMPTOTIC_MILLS_ALPHA && alpha * (beta - alpha) >= HALF_LINE_WINDOW {
let (v, w) = half_line_truncation(alpha);
return (if flipped { -v } else { v }, w);
let width = beta - alpha;
if alpha >= ASYMPTOTIC_MILLS_ALPHA && alpha * width >= HALF_LINE_WINDOW {
let (v, one_minus_w) = half_line_truncation(alpha);
return (if flipped { -v } else { v }, one_minus_w);
}
// A narrow window deep in the tail: too narrow for the half-line above,
// too deep for the subtraction below. The direct form keeps roughly
// `1 / (alpha/width)^2` of its digits, so the crossover is on that
// ratio rather than on either quantity alone — and the approximation is
// most accurate exactly where the subtraction is worst, since both
// improve as the window narrows.
if alpha > 0.0 && alpha > NARROW_WINDOW_RATIO * width {
let (v, one_minus_w) = narrow_window_truncation(alpha, width);
return (if flipped { -v } else { v }, one_minus_w);
}
let (v, u) = if alpha > 0.0 {
@@ -512,17 +697,23 @@ fn v_w(mu: f64, sigma: f64, margin: f64, tie: bool) -> (f64, f64) {
)
};
let w = -(u - v.powi(2));
// `1 - w` where `w = v^2 - u`. Both `v^2` and `u` grow as alpha^2 while
// their difference stays O(1), so this subtraction is the one place the
// tie branch can still lose everything — see the escape hatch above,
// which is what keeps the far tail away from it.
let one_minus_w = 1.0 + u - v.powi(2);
(if flipped { -v } else { v }, w)
(if flipped { -v } else { v }, one_minus_w)
}
}
fn trunc(mu: f64, sigma: f64, margin: f64, tie: bool) -> (f64, f64) {
let (v, w) = v_w(mu, sigma, margin, tie);
// `v_w` returns `1 - w` rather than `w`: forming the difference here is
// what destroyed the truncated variance in the far tail.
let (v, one_minus_w) = v_w(mu, sigma, margin, tie);
let mu_trunc = mu + sigma * v;
let sigma_trunc = sigma * (1.0 - w).sqrt();
let sigma_trunc = sigma * one_minus_w.sqrt();
(mu_trunc, sigma_trunc)
}
@@ -533,13 +724,34 @@ pub(crate) fn approx(n: Gaussian, margin: f64, tie: bool) -> Gaussian {
Gaussian::from_ms(mu, sigma)
}
/// Componentwise maximum that **propagates** NaN rather than dropping it.
///
/// Every caller folds this as `tuple_max(accumulator, new)`. A plain `>`
/// comparison is false against NaN, so a NaN accumulator would be replaced by
/// the next finite delta and the breakdown would vanish — leaving `step_is_finite`
/// to pass on a fit that is already NaN. Because the fold runs over a `HashMap`,
/// whether that happened depended on per-process hash order: measured, a NaN fit
/// was reported as `converged: true` in 16 of 30 runs on identical input.
///
/// `f64::max` is not a substitute: it also ignores NaN by design, which is the
/// same defect wearing a standard-library name.
pub(crate) fn tuple_max(v1: (f64, f64), v2: (f64, f64)) -> (f64, f64) {
(
if v1.0 > v2.0 { v1.0 } else { v2.0 },
if v1.1 > v2.1 { v1.1 } else { v2.1 },
max_propagating_nan(v1.0, v2.0),
max_propagating_nan(v1.1, v2.1),
)
}
fn max_propagating_nan(a: f64, b: f64) -> f64 {
if a.is_nan() || b.is_nan() {
f64::NAN
} else if a > b {
a
} else {
b
}
}
pub(crate) fn tuple_gt(t: (f64, f64), e: f64) -> bool {
t.0 > e || t.1 > e
}
@@ -606,29 +818,36 @@ pub(crate) fn sort_time<T: Copy + Ord>(xs: &[T], reverse: bool) -> Vec<usize> {
x.into_iter().map(|(i, _)| i).collect()
}
/// Calculates the match quality of the given rating groups. A result is the draw probability in the association
/// Calculates the match quality of the given teams. A result is the draw probability in the association
///
/// Supports any number of groups. Values range roughly `[0, 1]`; 1 means a
/// perfectly balanced match.
///
/// # Panics
///
/// Panics if fewer than two rating groups are supplied, or if any group is
/// Panics if fewer than two teams are supplied, or if any group is
/// empty — match quality is a property of a contest between at least two
/// non-empty sides.
///
/// Also panics with "cannot invert a singular matrix" when every rating has
/// zero sigma *and* `beta` is zero. Nothing is then uncertain, so there is no
/// distribution to take the quality of; `Gaussian::from_ms(mu, 0.0)` is a point
/// mass and its `mu()` is not even well defined. Documented rather than
/// converted, because the input has no meaningful answer rather than an
/// awkward one.
#[must_use]
pub fn quality(rating_groups: &[&[Gaussian]], beta: f64) -> f64 {
pub fn quality(teams: &[&[Gaussian]], beta: f64) -> f64 {
assert!(
rating_groups.len() >= 2,
"quality() requires at least 2 rating groups, got {}",
rating_groups.len()
teams.len() >= 2,
"quality() requires at least 2 teams, got {}",
teams.len()
);
assert!(
rating_groups.iter().all(|group| !group.is_empty()),
"quality() requires every rating group to be non-empty"
teams.iter().all(|group| !group.is_empty()),
"quality() requires every team to be non-empty"
);
let flatten_ratings = rating_groups
let flatten_ratings = teams
.iter()
.flat_map(|group| group.iter())
.collect::<Vec<_>>();
@@ -649,14 +868,14 @@ pub fn quality(rating_groups: &[&[Gaussian]], beta: f64) -> f64 {
variance_matrix[(i, i)] = rating.sigma().powi(2);
}
let mut rotated_a_matrix = Matrix::new(rating_groups.len() - 1, length);
let mut rotated_a_matrix = Matrix::new(teams.len() - 1, length);
// Row `row` contrasts group `row` (+weight) against group `row + 1`
// (-weight). `t` is the column where the current group's players start;
// the negative block begins immediately after it.
let mut t = 0;
for (row, group) in rating_groups.windows(2).enumerate() {
for (row, group) in teams.windows(2).enumerate() {
let current = group[0];
let next = group[1];
@@ -681,13 +900,141 @@ pub fn quality(rating_groups: &[&[Gaussian]], beta: f64) -> f64 {
let end = &rotated_a_matrix * &mean_matrix;
let e_arg = (-0.5 * &start * &middle.inverse() * &end).determinant();
let s_arg = ata.determinant() / middle.determinant();
libm::exp(e_arg) * s_arg.sqrt()
// `sqrt(det(ata) / det(middle))`, taken in log space. Both determinants are
// products of `k - 1` diagonal entries, so they leave `f64`'s range long
// before their ratio does: measured at the crate defaults, 150 groups was
// correct at `8.45e-53`, 200 returned `0`, and 250 returned `NaN` where the
// true value is `9.51e-88`. With a small beta it is sharper still — at
// `sigma = beta = 1e-3`, 60 groups returned `NaN` against a true `1.32e-9`.
//
// The ratio is what the answer needs and it is representable throughout, so
// the intermediates are the only thing that ever overflowed.
let ln_s_arg = ata.ln_abs_determinant() - middle.ln_abs_determinant();
libm::exp(e_arg + 0.5 * ln_s_arg)
}
#[cfg(test)]
mod tests {
/// The truncated variance must stay a variance across every branch, and
/// the branches must agree where they meet.
///
/// `v_w` now has three regimes for a tie — half-line, narrow-window, and
/// the direct subtraction — and a misplaced crossover between them is the
/// failure mode this guards. A jump at a boundary is visible here even
/// though the absolute values are not pinned.
#[test]
fn truncated_variance_is_continuous_across_the_tie_branches() {
for &alpha in &[50.0, 99.0, 100.0, 101.0, 1e3, 1e5, 1e6] {
// Sweep the window width across NARROW_WINDOW_RATIO and the
// half-line threshold, which sit at different widths per alpha.
let mut previous: Option<(f64, f64)> = None;
let mut width = alpha / (NARROW_WINDOW_RATIO * 100.0);
while width < 40.0 / alpha {
// mu = 0 puts the window at [-margin, margin]; shift it out to
// `alpha` by moving the mean instead.
let margin = width * 0.5;
let mu = -(alpha + width * 0.5);
let (v, one_minus_w) = v_w(mu, 1.0, margin, true);
assert!(v.is_finite(), "alpha {alpha}, width {width:e}: v = {v}");
assert!(
one_minus_w.is_finite() && one_minus_w > 0.0 && one_minus_w <= 1.0,
"alpha {alpha}, width {width:e}: 1 - w = {one_minus_w:e} is not a variance"
);
if let Some((pv, pw)) = previous {
// Consecutive widths differ by 2x, so the moments may not
// differ by more than a small multiple of that.
assert!(
one_minus_w / pw < 32.0 && pw / one_minus_w < 32.0,
"alpha {alpha}: 1 - w jumped from {pw:e} to {one_minus_w:e} \
at width {width:e} a branch boundary is misplaced"
);
assert!(
(v - pv).abs() <= 8.0 * width.max(1e-12) + 1e-9 * v.abs(),
"alpha {alpha}: v jumped from {pv} to {v} at width {width:e}"
);
}
previous = Some((v, one_minus_w));
width *= 2.0;
}
}
}
/// The narrow-window form against high-precision quadrature.
///
/// These are the inputs where the direct `v^2 - u` subtraction had four
/// significant digits left and returned a negative variance.
#[test]
fn narrow_window_truncation_matches_quadrature() {
for &(alpha, d, expect_v, expect_w) in &[
(1e6, 2e-6, 1_000_000.000_000_687, 2.759_383_390_335_666e-13),
(1e4, 1e-6, 10_000.000_000_499_167, 8.333_291_666_831_727e-14),
(
1e3,
1e-5,
1_000.000_004_991_666_6,
8.333_291_666_803_818e-12,
),
] {
let (v, one_minus_w) = narrow_window_truncation(alpha, d);
assert!(
((v - expect_v) / expect_v).abs() < 1e-12,
"alpha {alpha:e}: v = {v}, want {expect_v}"
);
assert!(
((one_minus_w - expect_w) / expect_w).abs() < 1e-8,
"alpha {alpha:e}: 1 - w = {one_minus_w:e}, want {expect_w:e}"
);
}
}
/// A NaN must survive the fold from ANY position, not only the last.
///
/// The fold runs over a `HashMap`, so "last" is per-process hash order. The
/// end-to-end symptom was a NaN fit reported as `converged: true` in 16 of
/// 30 runs on identical input; these three cases are the deterministic form
/// of that, so a regression cannot hide behind a lucky seed.
#[test]
fn tuple_max_propagates_a_nan_from_any_position() {
let nan = (f64::NAN, f64::NAN);
let small = (1e-9, 1e-9);
let big = (1e-3, 1e-3);
// NaN last.
let step = tuple_max(tuple_max(big, small), nan);
assert!(!step_is_finite(step), "NaN last: {step:?}");
// NaN middle.
let step = tuple_max(tuple_max(big, nan), small);
assert!(!step_is_finite(step), "NaN middle: {step:?}");
// NaN first — the case a plain `>` comparison drops.
let step = tuple_max(tuple_max(nan, big), small);
assert!(!step_is_finite(step), "NaN first: {step:?}");
}
/// `f64::max` would pass the test above's first two cases and fail the
/// third, so pin that it is not what we use.
#[test]
fn tuple_max_is_not_f64_max() {
assert!(
f64::max(f64::NAN, 1.0) == 1.0,
"premise: f64::max drops NaN"
);
let (a, _) = tuple_max((f64::NAN, 0.0), (1.0, 0.0));
assert!(a.is_nan(), "tuple_max must not drop what f64::max drops");
}
/// Ordinary values are unaffected.
#[test]
fn tuple_max_still_takes_the_larger_component() {
assert_eq!(tuple_max((1.0, 5.0), (3.0, 2.0)), (3.0, 5.0));
assert_eq!(tuple_max((3.0, 2.0), (1.0, 5.0)), (3.0, 5.0));
}
use ::approx::assert_ulps_eq;
use super::*;
+45 -4
View File
@@ -91,6 +91,29 @@ impl Lu {
det
}
/// `ln |det|`, accumulated term by term rather than multiplied out.
///
/// The determinant of an `n x n` Gram matrix is a product of `n` diagonal
/// entries, so it leaves `f64`'s range long before the quantities built
/// from it do. `quality()` only ever wants a *ratio* of two determinants,
/// and that ratio is perfectly representable while the determinants
/// themselves are not — measured, at 250 rating groups both overflow and
/// the ratio came back `NaN` where the true answer is `9.51e-88`.
///
/// Returns `-inf` for a singular matrix, so `exp` of it is zero.
fn ln_abs_determinant(&self) -> f64 {
if self.sign == 0.0 {
return f64::NEG_INFINITY;
}
let mut acc = 0.0;
for i in 0..self.n {
acc += libm::log(self.lu[i * self.n + i].abs());
}
acc
}
/// Solve `Ax = b` for a single column of the identity, giving one column
/// of the inverse.
fn solve_column(&self, col: usize, out: &mut [f64]) {
@@ -117,7 +140,7 @@ impl Lu {
}
impl Matrix {
pub fn new(height: usize, width: usize) -> Matrix {
pub(crate) fn new(height: usize, width: usize) -> Matrix {
Matrix {
data: vec![0.0; height * width].into_boxed_slice(),
height,
@@ -125,7 +148,7 @@ impl Matrix {
}
}
pub fn transpose(&self) -> Matrix {
pub(crate) fn transpose(&self) -> Matrix {
let mut matrix = Matrix::new(self.width, self.height);
for c in 0..self.width {
@@ -143,7 +166,7 @@ impl Matrix {
/// # Panics
///
/// Panics if the matrix is not square.
pub fn determinant(&self) -> f64 {
pub(crate) fn determinant(&self) -> f64 {
assert_eq!(
self.width, self.height,
"determinant requires a square matrix, got {}x{}",
@@ -157,12 +180,30 @@ impl Matrix {
Lu::decompose(self).determinant()
}
/// `ln |det|` of a square matrix; `-inf` when singular.
///
/// See [`Lu::ln_abs_determinant`] for why a ratio of determinants must be
/// taken this way.
pub(crate) fn ln_abs_determinant(&self) -> f64 {
assert_eq!(
self.width, self.height,
"determinant requires a square matrix, got {}x{}",
self.height, self.width
);
if self.width == 0 {
return 0.0;
}
Lu::decompose(self).ln_abs_determinant()
}
/// Matrix inverse via LU decomposition.
///
/// # Panics
///
/// Panics if the matrix is not square or is singular.
pub fn inverse(&self) -> Matrix {
pub(crate) fn inverse(&self) -> Matrix {
assert_eq!(
self.width, self.height,
"inverse requires a square matrix, got {}x{}",
+79 -19
View File
@@ -1,6 +1,6 @@
//! Outcome of a match.
//!
//! `Ranked(ranks)` for ordinal results; `Scored { scores, sigma }` for
//! `Ranked(ranks)` for ordinal results; `Scored { scores, score_sigma }` for
//! continuous per-team scores (engages `MarginFactor` in the engine).
use smallvec::SmallVec;
@@ -10,19 +10,41 @@ use smallvec::SmallVec;
/// `Ranked(ranks)`: lower rank = better. Equal ranks mean a tie between those
/// teams. `ranks.len()` must equal the number of teams in the event.
///
/// `Scored { scores, sigma }`: higher score = better. Adjacent (sorted) pairs
/// `Scored { scores, score_sigma }`: higher score = better. Adjacent (sorted) pairs
/// feed observed margins to `MarginFactor`. `scores.len()` must equal the
/// number of teams in the event. `sigma` overrides `HistoryBuilder::score_sigma`
/// when `Some`; `None` inherits the history default.
#[derive(Clone, Debug, PartialEq)]
#[non_exhaustive]
#[must_use]
pub enum Outcome {
/// An ordinal finish: one rank per team, in the order the teams were given.
///
/// Lower is better, `0` is first, and equal values are a tie between those
/// teams — which needs `p_draw > 0`, or ingestion rejects the event with
/// [`InferenceError::TieWithoutDrawProbability`](crate::InferenceError::TieWithoutDrawProbability).
///
/// Only the ordering and the equalities are used. Ranks need not be dense
/// or start at zero: inference sorts the teams and compares rank-adjacent
/// pairs against a margin set by `p_draw`, so `[0, 1, 2]` and `[0, 5, 90]`
/// are the same observation. A gap does not mean a bigger win — use
/// `Scored` when the size of the difference is evidence.
Ranked(SmallVec<[u32; 4]>),
/// A continuous finish: one score per team, higher is better.
///
/// Unlike `Ranked`, the *sizes* of the differences are evidence. Teams are
/// sorted by score and each adjacent pair's observed gap is fed to a
/// `MarginFactor` as a measurement with standard deviation `score_sigma`,
/// so
/// beating a team by ten says more than beating them by one.
#[non_exhaustive]
Scored {
/// Per-team scores, in the order the teams were given; higher is
/// better. Must have one entry per team, and every entry finite.
scores: SmallVec<[f64; 4]>,
/// Per-event noise override. `None` means inherit
/// `HistoryBuilder::score_sigma`. Must be `> 0.0` if `Some`.
sigma: Option<f64>,
score_sigma: Option<f64>,
},
}
@@ -34,16 +56,43 @@ impl Outcome {
///
/// # Panics
///
/// Panics if `winner >= n`.
#[must_use]
/// Panics if `winner >= n`. Use [`Outcome::try_winner`] when the index
/// comes from data rather than a literal.
///
/// This is the one constructor here that validates, and deliberately so.
/// Its siblings build freely and let ingestion reject what it cannot use,
/// which works because a malformed rank vector stays recognisable. An
/// out-of-range winner does not: `winner(5, 2)` would produce ranks
/// `[1, 1]`, an all-tied draw that ingestion accepts without complaint when
/// `p_draw > 0`. Asking "team 5 won" and silently getting "everyone drew"
/// is exactly the class of quiet wrong answer this crate keeps removing, so
/// the check happens here where the mistake is.
pub fn winner(winner: u32, n: u32) -> Self {
assert!(winner < n, "winner index {winner} out of range 0..{n}");
Self::try_winner(winner, n)
.unwrap_or_else(|_| panic!("winner index {winner} out of range 0..{n}"))
}
/// `n`-team outcome where team `winner` won, or an error if `winner` is not
/// a valid team index.
///
/// The fallible form of [`Outcome::winner`], for when the index is computed
/// or parsed rather than written literally.
///
/// # Errors
///
/// `InvalidParameter` if `winner >= n`.
pub fn try_winner(winner: u32, n: u32) -> Result<Self, crate::InferenceError> {
if winner >= n {
return Err(crate::InferenceError::InvalidParameter {
parameter: crate::Parameter::WinnerIndex,
value: f64::from(winner),
});
}
let ranks: SmallVec<[u32; 4]> = (0..n).map(|i| if i == winner { 0 } else { 1 }).collect();
Self::Ranked(ranks)
Ok(Self::Ranked(ranks))
}
/// All `n` teams tied.
#[must_use]
pub fn draw(n: u32) -> Self {
Self::Ranked(SmallVec::from_vec(vec![0; n as usize]))
}
@@ -58,23 +107,34 @@ impl Outcome {
pub fn scores<I: IntoIterator<Item = f64>>(scores: I) -> Self {
Self::Scored {
scores: scores.into_iter().collect(),
sigma: None,
score_sigma: None,
}
}
/// Explicit per-team continuous scores with a per-event noise override.
///
/// `sigma` must be `> 0.0`. Constructing an `Outcome` with a non-positive
/// or NaN sigma is allowed; the value is rejected with
/// The noise is on the *observed score margin*, in the units of the scores
/// themselves — it is not a skill sigma, which is what the old name
/// `scores_with_sigma` read as. It overrides `HistoryBuilder::score_sigma`
/// for this event only.
///
/// `score_sigma` must be `> 0.0`. Constructing an `Outcome` with a
/// non-positive or NaN value is allowed; the value is rejected with
/// `InferenceError::InvalidParameter` when the event is ingested, so
/// callers get an error rather than a panic.
pub fn scores_with_sigma<I: IntoIterator<Item = f64>>(scores: I, sigma: f64) -> Self {
pub fn scores_with_noise<I: IntoIterator<Item = f64>>(scores: I, score_sigma: f64) -> Self {
Self::Scored {
scores: scores.into_iter().collect(),
sigma: Some(sigma),
score_sigma: Some(score_sigma),
}
}
/// How many teams this outcome describes — the number of ranks, or of
/// scores.
///
/// Ingestion checks it against the event's own team list and rejects a
/// disagreement with `MismatchedShape`, so this is the cheap way to check
/// an outcome built elsewhere before committing the event.
#[must_use]
pub fn team_count(&self) -> usize {
match self {
@@ -156,7 +216,7 @@ mod tests {
#[test]
fn scores_with_sigma_round_trips() {
let o = Outcome::scores_with_sigma([10.0, 4.0], 0.5);
let o = Outcome::scores_with_noise([10.0, 4.0], 0.5);
assert_eq!(o.team_count(), 2);
assert_eq!(o.as_scores(), Some(&[10.0, 4.0][..]));
}
@@ -165,16 +225,16 @@ mod tests {
fn scores_constructor_leaves_sigma_unset() {
let o = Outcome::scores([3.0, 1.0]);
match o {
Outcome::Scored { scores: _, sigma } => assert!(sigma.is_none()),
Outcome::Scored { score_sigma, .. } => assert!(score_sigma.is_none()),
Outcome::Ranked(_) => panic!("expected Scored variant"),
}
}
#[test]
fn scores_with_sigma_sets_sigma_some() {
let o = Outcome::scores_with_sigma([3.0, 1.0], 2.0);
let o = Outcome::scores_with_noise([3.0, 1.0], 2.0);
match o {
Outcome::Scored { scores: _, sigma } => assert_eq!(sigma, Some(2.0)),
Outcome::Scored { score_sigma, .. } => assert_eq!(score_sigma, Some(2.0)),
Outcome::Ranked(_) => panic!("expected Scored variant"),
}
}
@@ -184,9 +244,9 @@ mod tests {
/// `tests/degenerate_inputs.rs::scored_event_rejects_non_positive_sigma`.
#[test]
fn scores_with_sigma_defers_validation_to_ingestion() {
let o = Outcome::scores_with_sigma([3.0, 1.0], 0.0);
let o = Outcome::scores_with_noise([3.0, 1.0], 0.0);
match o {
Outcome::Scored { sigma, .. } => assert_eq!(sigma, Some(0.0)),
Outcome::Scored { score_sigma, .. } => assert_eq!(score_sigma, Some(0.0)),
Outcome::Ranked(_) => panic!("expected Scored variant"),
}
}
+58 -27
View File
@@ -21,7 +21,7 @@
//! would have made every `predict_*` call return a slightly different number,
//! which is not a property a rating library should have.
use crate::{Gaussian, quadrature};
use crate::{Gaussian, InferenceError, quadrature};
/// Teams beyond this count make the outcome enumeration impractical.
///
@@ -52,6 +52,10 @@ const WIN_TOLERANCE: f64 = 1e-8;
/// point where refining stops helping.
const MIN_GRID_POINTS: usize = 8_192;
const MAX_GRID_POINTS: usize = 262_144;
/// Nodes requested across the narrowest feature the recursion must resolve.
const NODES_PER_FEATURE: f64 = 12.0;
/// Nodes below which the trapezoid rule stops resolving that feature at all.
const MIN_NODES_PER_FEATURE: f64 = 4.0;
/// How many standard deviations of support the grid and integrals cover.
///
@@ -152,7 +156,7 @@ pub(crate) fn win_probabilities(perf: &[Gaussian], margins: &Margins) -> Vec<f64
/// Resolution is set by the *smallest* feature in play — the narrowest sigma,
/// or a draw margin narrower still — because that is what the recursion has to
/// resolve. A grid sized off the widest team would step over the narrow one.
fn grid_shape(perf: &[Gaussian], margins: &Margins) -> (f64, f64, usize) {
fn grid_shape(perf: &[Gaussian], margins: &Margins) -> Result<(f64, f64, usize), InferenceError> {
let lo = perf
.iter()
.map(|g| g.mu() - SUPPORT_SIGMAS * g.sigma())
@@ -175,18 +179,36 @@ fn grid_shape(perf: &[Gaussian], margins: &Margins) -> (f64, f64, usize) {
let feature = narrowest.min(smallest_margin);
let wanted = if feature.is_finite() && feature > 0.0 {
((hi - lo) / (feature / 12.0)).ceil()
((hi - lo) / (feature / NODES_PER_FEATURE)).ceil()
} else {
MIN_GRID_POINTS as f64
};
let points = if wanted.is_finite() {
(wanted as usize).clamp(MIN_GRID_POINTS, MAX_GRID_POINTS)
} else {
MIN_GRID_POINTS
};
if !wanted.is_finite() {
return Ok((lo, hi, MIN_GRID_POINTS));
}
(lo, hi, points)
// Report rather than clamp. Clamping is what this replaced: it silently
// handed the recursion a grid too coarse for the narrowest density, and the
// trapezoid rule then returned probabilities greater than one — measured, a
// `P` of 2.79 and a total of 5.41. Trapezoid error on a Gaussian is
// `~exp(-2 pi^2 (sigma/h)^2)`, which is 1e-12 at `h/sigma = 0.86` and O(1)
// by `h/sigma = 17`, so the cliff is sharp and there is no useful answer on
// the far side of it.
//
// The floor is `MIN_NODES_PER_FEATURE` rather than the `NODES_PER_FEATURE`
// asked for, because the request carries a large margin: measured accurate
// to 2.2e-12 at 1.4 nodes per sigma, and wrong by 1.2e-3 at 0.7.
let needed = wanted as usize;
let floor = ((hi - lo) / (feature / MIN_NODES_PER_FEATURE)).ceil();
if floor.is_finite() && floor as usize > MAX_GRID_POINTS {
return Err(InferenceError::GridTooCoarse {
needed,
max: MAX_GRID_POINTS,
});
}
Ok((lo, hi, needed.clamp(MIN_GRID_POINTS, MAX_GRID_POINTS)))
}
/// Densities of each team sampled on the shared grid.
@@ -198,8 +220,8 @@ struct Sampled {
}
impl Sampled {
fn new(perf: &[Gaussian], margins: &Margins) -> Self {
let (lo, hi, points) = grid_shape(perf, margins);
fn new(perf: &[Gaussian], margins: &Margins) -> Result<Self, InferenceError> {
let (lo, hi, points) = grid_shape(perf, margins)?;
let step = (hi - lo) / (points - 1) as f64;
let density = perf
.iter()
@@ -209,12 +231,12 @@ impl Sampled {
.collect()
})
.collect();
Self {
Ok(Self {
lo,
step,
points,
density,
}
})
}
fn node(&self, i: usize) -> f64 {
@@ -317,9 +339,12 @@ fn events(n: usize, strict_only: bool) -> Vec<(Vec<usize>, Vec<bool>)> {
///
/// Orders that differ only *within* a tied group describe the same finishing
/// order, so their probabilities are summed into one entry.
pub(crate) fn outcome_distribution(perf: &[Gaussian], margins: &Margins) -> Vec<(Vec<u32>, f64)> {
pub(crate) fn outcome_distribution(
perf: &[Gaussian],
margins: &Margins,
) -> Result<Vec<(Vec<u32>, f64)>, InferenceError> {
let n = perf.len();
let sampled = Sampled::new(perf, margins);
let sampled = Sampled::new(perf, margins)?;
let mut aggregated: Vec<(Vec<u32>, f64)> = Vec::new();
for (order, tied) in events(n, margins.all_zero()) {
@@ -332,7 +357,7 @@ pub(crate) fn outcome_distribution(perf: &[Gaussian], margins: &Margins) -> Vec<
}
aggregated.sort_by(|a, b| b.1.partial_cmp(&a.1).unwrap_or(std::cmp::Ordering::Equal));
aggregated
Ok(aggregated)
}
/// All permutations of `items`.
@@ -397,9 +422,13 @@ fn orders_for_groups(groups: &[Vec<usize>]) -> Vec<(Vec<usize>, Vec<bool>)> {
/// Ties in `ranks` mean the tied teams may finish in any internal order, so
/// this sums the orders consistent with the requested ranking rather than
/// picking one.
pub(crate) fn ranking_probability(perf: &[Gaussian], margins: &Margins, ranks: &[u32]) -> f64 {
pub(crate) fn ranking_probability(
perf: &[Gaussian],
margins: &Margins,
ranks: &[u32],
) -> Result<f64, InferenceError> {
let n = perf.len();
let sampled = Sampled::new(perf, margins);
let sampled = Sampled::new(perf, margins)?;
let mut distinct: Vec<u32> = ranks.to_vec();
distinct.sort_unstable();
@@ -410,10 +439,10 @@ pub(crate) fn ranking_probability(perf: &[Gaussian], margins: &Margins, ranks: &
.map(|&r| (0..n).filter(|&i| ranks[i] == r).collect())
.collect();
orders_for_groups(&groups)
Ok(orders_for_groups(&groups)
.iter()
.map(|(order, tied)| order_probability(margins, &sampled, order, tied))
.sum()
.sum())
}
/// A distribution over the ways a contest could finish.
@@ -427,6 +456,7 @@ pub(crate) fn ranking_probability(perf: &[Gaussian], margins: &Margins, ranks: &
/// `Game::ranked` asks "what would we believe if *this* happened", which is
/// what an expected-information-gain calculation needs alongside the weight.
#[derive(Clone, Debug, PartialEq)]
#[must_use]
pub struct Prediction {
outcomes: Vec<(Vec<u32>, f64)>,
}
@@ -437,6 +467,7 @@ impl Prediction {
}
/// Every possible finishing order and its probability, most likely first.
#[must_use]
pub fn outcomes(&self) -> impl ExactSizeIterator<Item = (&[u32], f64)> {
self.outcomes.iter().map(|(r, p)| (r.as_slice(), *p))
}
@@ -604,7 +635,7 @@ mod tests {
),
] {
let n = perf.len();
let dist = outcome_distribution(&perf, &flat(n, eps));
let dist = outcome_distribution(&perf, &flat(n, eps)).unwrap();
let sum: f64 = dist.iter().map(|(_, p)| p).sum();
assert!(
(sum - 1.0).abs() < 1e-6,
@@ -620,7 +651,7 @@ mod tests {
fn two_team_distribution_matches_the_closed_form() {
let perf = [g(3.0, 6.0), g(-2.0, 1.0)];
let eps = 1.5;
let dist = outcome_distribution(&perf, &flat(2, eps));
let dist = outcome_distribution(&perf, &flat(2, eps)).unwrap();
let (wa, wb) = closed_form_two(perf[0], perf[1], eps);
let find = |ranks: &[u32]| {
@@ -653,10 +684,10 @@ mod tests {
let perf = [g(5.0, 6.0), g(0.0, 3.0), g(-5.0, 1.0)];
let eps = 1.5;
let margins = flat(3, eps);
let dist = outcome_distribution(&perf, &margins);
let dist = outcome_distribution(&perf, &margins).unwrap();
for (ranks, expected) in &dist {
let direct = ranking_probability(&perf, &margins, ranks);
let direct = ranking_probability(&perf, &margins, ranks).unwrap();
assert!(
(direct - expected).abs() < 1e-9,
"ranks {ranks:?}: direct {direct} vs distribution {expected}"
@@ -675,7 +706,7 @@ mod tests {
let perf = [g(0.0, 4.0), g(0.0, 4.0), g(-8.0, 2.0)];
let mut previous = 0.0;
for eps in [0.0, 0.5, 1.0, 2.0, 4.0, 8.0, 24.0] {
let p = ranking_probability(&perf, &flat(3, eps), &[0, 0, 0]);
let p = ranking_probability(&perf, &flat(3, eps), &[0, 0, 0]).unwrap();
assert!(p >= previous, "eps={eps}: {p} < {previous}");
if eps == 0.0 {
assert!(p < 1e-12, "a tie needs a margin, got {p}");
@@ -696,7 +727,7 @@ mod tests {
let perf = [g(0.0, 4.0), g(0.0, 4.0), g(-8.0, 2.0)];
let sweep: Vec<f64> = [0.5, 2.0, 4.0, 8.0, 16.0]
.iter()
.map(|&eps| ranking_probability(&perf, &flat(3, eps), &[0, 0, 1]))
.map(|&eps| ranking_probability(&perf, &flat(3, eps), &[0, 0, 1]).unwrap())
.collect();
let peak = sweep
.iter()
@@ -717,7 +748,7 @@ mod tests {
#[test]
fn ties_are_impossible_without_a_draw_margin() {
let perf = [g(0.0, 4.0), g(0.0, 4.0), g(0.0, 4.0)];
let dist = outcome_distribution(&perf, &flat(3, 0.0));
let dist = outcome_distribution(&perf, &flat(3, 0.0)).unwrap();
assert_eq!(dist.len(), 6, "expected only the 6 strict orders: {dist:?}");
assert!(dist.iter().all(|(r, _)| {
let mut seen = r.clone();
+19 -3
View File
@@ -11,7 +11,7 @@ use crate::{
///
/// A configuration rather than a person: the per-history temporal state
/// (messages, last appearance) lives on `Competitor`.
#[derive(Clone, Copy, Debug)]
#[derive(Clone, Copy, Debug, PartialEq)]
pub struct Rating<T: Time = i64, D: Drift<T> = ConstantDrift> {
pub(crate) prior: Gaussian,
pub(crate) beta: f64,
@@ -23,7 +23,24 @@ pub struct Rating<T: Time = i64, D: Drift<T> = ConstantDrift> {
}
impl<T: Time, D: Drift<T>> Rating<T, D> {
/// # Panics
///
/// Panics unless `beta` is finite and non-negative, matching
/// `HistoryBuilder::beta`.
///
/// Zero is allowed and meaningful — performance is then exactly skill, and
/// the fit differs measurably from a positive beta rather than degenerating.
/// Negative is rejected because `beta` enters only as `beta^2`: measured, a
/// negative beta returned results **bit identical** to its absolute value,
/// and a NaN beta reached `Game::ranked`, which returned `Ok` carrying a
/// `Gaussian { pi: NaN, tau: NaN }` — there is no `converge` on that path to
/// catch it.
pub fn new(prior: Gaussian, beta: f64, drift: D) -> Self {
assert!(
beta.is_finite() && beta >= 0.0,
"beta must be finite and non-negative (got {beta}); it is only ever \
squared, so a negative value would silently behave as its absolute value"
);
Self {
prior,
beta,
@@ -44,7 +61,6 @@ impl<T: Time, D: Drift<T>> Rating<T, D> {
}
/// The configured prior skill estimate.
#[must_use]
pub fn prior(&self) -> Gaussian {
self.prior
}
@@ -93,7 +109,7 @@ impl Default for Rating<i64, ConstantDrift> {
Self {
prior: Gaussian::default(),
beta: BETA,
drift: ConstantDrift(GAMMA),
drift: ConstantDrift::new(GAMMA),
drift_scale: 1.0,
_time: PhantomData,
}
+140
View File
@@ -0,0 +1,140 @@
//! Declarative competitor configuration: a rule that supplies defaults for
//! competitors the history has not seen yet.
//!
//! [`History::register`](crate::History::register) states configuration for
//! *one* competitor, which covers a bot at a known strength or a handful of
//! reference points. It does not cover a *rule* — "every layout is static" —
//! because enumerating the keys means knowing the full key set up front, which
//! a consumer ingesting an event stream generally does not.
//!
//! ```
//! use trueskill_tt::{Gaussian, History, StartingPoint};
//!
//! let mut h = History::builder()
//! // Layouts do not improve; everybody else does.
//! .default_rating_for(|key: &&'static str| {
//! key.starts_with("layout_")
//! .then(|| StartingPoint::new().prior(Gaussian::from_ms(0.0, 1.0)).drift_scale(0.0))
//! })
//! .build();
//!
//! h.event(1).team(["layout_7"]).team(["alice"]).scores([3.0, 1.0]).commit()?;
//! h.converge()?;
//!
//! // The layout was pinned, so its uncertainty barely moved.
//! assert!(h.current_skill("layout_7").unwrap().sigma() < 1.0);
//! # Ok::<(), trueskill_tt::InferenceError>(())
//! ```
//!
//! # Why a trait, and why a fifth type parameter
//!
//! The rule is a type parameter on [`History`](crate::History), defaulted to
//! [`NoRule`], so it costs a caller who does not use one exactly nothing —
//! `History<String>` still spells out in full. A boxed `dyn Fn` would have
//! avoided the parameter at the price of `HistoryBuilder`'s derived `Clone`
//! and `Debug`.
//!
//! It is a trait rather than a bare `Fn` bound because a closure's type cannot
//! be written down, and the motivating consumer holds its `History` in
//! application state — so it has to name the type in a struct field. Implement
//! [`RatingRule`] on a named type of your own and that field is spellable.
//!
//! # What a rule may set, and what it may not
//!
//! A [`StartingPoint`], which is the same pair a
//! [`Member`](crate::Member) may carry: the prior and the drift scale. Not
//! `beta` and not the drift model — those describe the *history*, not one
//! competitor, and a rule that could vary them would be describing a different
//! model per competitor rather than a starting point within one.
//!
//! Keeping the rule to those two also keeps it independent of the history's
//! time and drift types, so [`HistoryBuilder::drift`](crate::HistoryBuilder::drift)
//! and [`HistoryBuilder::time_type`](crate::HistoryBuilder::time_type) still
//! work after a rule is set.
use crate::gaussian::Gaussian;
/// What a [`RatingRule`] may say about a competitor.
///
/// Both fields are optional and are applied independently, so a rule that sets
/// only `drift_scale` does not also assert a prior — the same reason
/// `Member`'s configuration is carried as "what was explicitly set" rather
/// than as a merged `Rating`.
#[derive(Clone, Copy, Debug, Default, PartialEq)]
#[must_use]
pub struct StartingPoint {
pub(crate) prior: Option<Gaussian>,
pub(crate) drift_scale: Option<f64>,
}
impl StartingPoint {
/// A starting point that says nothing yet.
pub fn new() -> Self {
Self::default()
}
/// Start this competitor from `prior` instead of the history's
/// `mu`/`sigma`.
pub fn prior(mut self, prior: Gaussian) -> Self {
self.prior = Some(prior);
self
}
/// Scale how fast this competitor drifts, relative to the history's drift
/// model. `0.0` pins them still.
pub fn drift_scale(mut self, drift_scale: f64) -> Self {
self.drift_scale = Some(drift_scale);
self
}
}
/// Supplies a [`StartingPoint`] for competitors the history has not seen.
///
/// Consulted once per competitor, when that competitor is created — not per
/// event and not per sweep. Returning `None` means "no opinion": the
/// competitor takes the history's own defaults.
///
/// # Precedence
///
/// Explicit configuration wins, field by field. A `prior` or `drift_scale`
/// from [`History::register`](crate::History::register) or from a
/// [`Member`](crate::Member) overrides whatever the rule returned for that
/// competitor. The specific beats the general, which is the only reading that
/// lets a rule have exceptions — treating the disagreement as
/// `ConflictingCompetitorConfig` would make one exceptional competitor
/// incompatible with having any rule at all.
///
/// Two *explicit* declarations that disagree remain an error. Neither of those
/// is more specific than the other, so there is nothing to prefer.
pub trait RatingRule<K> {
/// Where this competitor should start, or `None` for the history's
/// defaults.
fn starting_point(&self, key: &K) -> Option<StartingPoint>;
}
/// The default rule: no opinion about anybody.
#[derive(Clone, Copy, Debug, Default, PartialEq, Eq)]
pub struct NoRule;
impl<K> RatingRule<K> for NoRule {
#[inline]
fn starting_point(&self, _key: &K) -> Option<StartingPoint> {
None
}
}
/// A [`RatingRule`] built from a closure by
/// [`HistoryBuilder::default_rating_for`](crate::HistoryBuilder::default_rating_for).
///
/// Public so it can be named where a closure's own type cannot be, though
/// implementing [`RatingRule`] on a named type of your own is the better way
/// to get a `History<..>` you can write down in a struct field.
#[derive(Clone, Copy, Debug, Default, PartialEq, Eq)]
pub struct FnRule<F>(pub F);
impl<K, F: Fn(&K) -> Option<StartingPoint>> RatingRule<K> for FnRule<F> {
#[inline]
fn starting_point(&self, key: &K) -> Option<StartingPoint> {
(self.0)(key)
}
}
-152
View File
@@ -1,152 +0,0 @@
//! Schedule trait and built-in implementations.
//!
//! A schedule drives factor propagation to convergence. The default
//! `EpsilonOrMax` performs one `TeamSum` sweep (setup) then alternating
//! forward/backward sweeps over the iterating factors until the max
//! delta drops below epsilon or `max` iterations is reached.
use crate::factor::{BuiltinFactor, Factor, VarStore};
/// Result returned by a `Schedule::run` call.
#[derive(Debug, Clone, Copy)]
pub struct ScheduleReport {
pub iterations: usize,
pub final_step: (f64, f64),
pub converged: bool,
}
/// Drives factor propagation to convergence.
pub trait Schedule: Send + Sync {
fn run(&self, factors: &mut [BuiltinFactor], vars: &mut VarStore) -> ScheduleReport;
}
/// Default schedule: sweep forward then backward until step ≤ eps or iter == max.
///
/// Matches the existing `Game::likelihoods` loop bit-for-bit when given the
/// same factor layout (`TeamSums` first, then alternating RankDiff/Trunc pairs).
#[derive(Debug, Clone, Copy)]
pub struct EpsilonOrMax {
pub eps: f64,
pub max: usize,
}
impl Default for EpsilonOrMax {
fn default() -> Self {
// Derived from `ConvergenceOptions` so there is one source of truth for
// the tolerance and iteration cap. These previously disagreed: this
// default capped at 10 iterations while `ConvergenceOptions` allowed 30,
// and which applied depended on whether inference went through
// `run_chain` or a `Schedule`.
let defaults = crate::ConvergenceOptions::default();
Self {
eps: defaults.epsilon,
max: defaults.max_iter,
}
}
}
impl Schedule for EpsilonOrMax {
fn run(&self, factors: &mut [BuiltinFactor], vars: &mut VarStore) -> ScheduleReport {
// Partition: leading run of TeamSum factors run exactly once (setup).
let n_setup = factors
.iter()
.position(|f| !matches!(f, BuiltinFactor::TeamSum(_)))
.unwrap_or(factors.len());
for f in factors[..n_setup].iter_mut() {
f.propagate(vars);
}
let mut iterations = 0;
// With no iterating factors the graph is already at its fixed point:
// the setup pass above is all there is to do. Reporting `converged:
// false` with an infinite step for that case gave callers a false
// negative.
let mut final_step = (0.0, 0.0);
let mut converged = true;
if n_setup < factors.len() {
final_step = (f64::INFINITY, f64::INFINITY);
converged = false;
for _ in 0..self.max {
let mut step = (0.0_f64, 0.0_f64);
// Forward sweep over iterating factors.
for f in factors[n_setup..].iter_mut() {
let d = f.propagate(vars);
step.0 = step.0.max(d.0);
step.1 = step.1.max(d.1);
}
// Backward sweep.
for f in factors[n_setup..].iter_mut().rev() {
let d = f.propagate(vars);
step.0 = step.0.max(d.0);
step.1 = step.1.max(d.1);
}
iterations += 1;
final_step = step;
if step.0 <= self.eps && step.1 <= self.eps {
converged = true;
break;
}
}
}
ScheduleReport {
iterations,
final_step,
converged,
}
}
}
#[cfg(test)]
mod tests {
use super::*;
use crate::{N_INF, factor::team_sum::TeamSumFactor, gaussian::Gaussian};
#[test]
fn schedule_runs_setup_factors_once() {
// Single TeamSum factor; schedule should propagate it exactly once and report 0 iterations.
let mut vars = VarStore::new();
let out = vars.alloc(N_INF);
let mut factors = vec![BuiltinFactor::TeamSum(TeamSumFactor {
inputs: vec![(Gaussian::from_ms(5.0, 1.0), 1.0)],
out,
})];
let schedule = EpsilonOrMax::default();
let report = schedule.run(&mut factors, &mut vars);
assert_eq!(report.iterations, 0);
// The team-perf var should hold the sum.
let result = vars.get(out);
assert!((result.mu() - 5.0).abs() < 1e-12);
}
#[test]
fn report_marks_converged_when_no_iterating_factors() {
// A graph of only setup factors has nothing to iterate, so it is at its
// fixed point after the setup pass: 0 iterations, and converged.
let mut vars = VarStore::new();
let out = vars.alloc(N_INF);
let mut factors = vec![BuiltinFactor::TeamSum(TeamSumFactor {
inputs: vec![(Gaussian::from_ms(0.0, 1.0), 1.0)],
out,
})];
let report = EpsilonOrMax::default().run(&mut factors, &mut vars);
assert_eq!(report.iterations, 0);
assert!(report.converged);
assert_eq!(report.final_step, (0.0, 0.0));
}
#[test]
fn default_matches_convergence_options() {
let schedule = EpsilonOrMax::default();
let options = crate::ConvergenceOptions::default();
assert_eq!(schedule.max, options.max_iter);
assert_eq!(schedule.eps, options.epsilon);
}
}
+5 -12
View File
@@ -56,16 +56,16 @@ impl<T: Time, D: Drift<T>> CompetitorStore<T, D> {
self.get(idx).is_some()
}
/// Test-only: no code path in the crate needs a count.
#[cfg(test)]
#[must_use]
pub fn len(&self) -> usize {
self.n_present
}
#[must_use]
pub fn is_empty(&self) -> bool {
self.n_present == 0
}
/// Test-only: iterating every competitor is an assertion helper, not part
/// of inference, which walks slices rather than the store.
#[cfg(test)]
pub fn iter(&self) -> impl Iterator<Item = (Index, &Competitor<T, D>)> {
self.competitors
.iter()
@@ -73,13 +73,6 @@ impl<T: Time, D: Drift<T>> CompetitorStore<T, D> {
.filter_map(|(i, slot)| slot.as_ref().map(|a| (Index(i), a)))
}
pub fn iter_mut(&mut self) -> impl Iterator<Item = (Index, &mut Competitor<T, D>)> {
self.competitors
.iter_mut()
.enumerate()
.filter_map(|(i, slot)| slot.as_mut().map(|a| (Index(i), a)))
}
pub fn values_mut(&mut self) -> impl Iterator<Item = &mut Competitor<T, D>> {
self.competitors.iter_mut().filter_map(|s| s.as_mut())
}
+242 -121
View File
@@ -9,7 +9,7 @@ use crate::{
arena::ScratchArena,
color_group::ColorGroups,
drift::Drift,
game::Game,
game::GameRef,
gaussian::Gaussian,
rating::Rating,
storage::{CompetitorStore, SkillStore},
@@ -26,7 +26,9 @@ pub(crate) struct Skill {
impl Skill {
pub(crate) fn posterior(&self) -> Gaussian {
self.likelihood * self.backward * self.forward
self.likelihood
.ep_product(self.backward)
.ep_product(self.forward)
}
}
@@ -50,12 +52,12 @@ pub enum EventKind {
#[derive(Clone, Debug)]
struct Item {
agent: Index,
competitor: Index,
/// This competitor's slot in the owning slice's `SkillStore`, resolved
/// once at ingestion.
///
/// The convergence loop reaches skills through this rather than through
/// `agent`, which is what keeps `HashMap` hashing out of the hot path now
/// `competitor`, which is what keeps `HashMap` hashing out of the hot path now
/// that the store is compact rather than indexed by the global `Index`.
slot: u32,
likelihood: Gaussian,
@@ -66,15 +68,15 @@ impl Item {
&self,
forward: bool,
skills: &SkillStore,
agents: &CompetitorStore<T, D>,
competitors: &CompetitorStore<T, D>,
) -> Rating<T, D> {
let r = &agents[self.agent].rating;
let r = &competitors[self.competitor].rating;
let skill = skills.at(self.slot);
if forward {
Rating::new(skill.forward, r.beta, r.drift).with_drift_scale(r.drift_scale)
} else {
Rating::new(skill.posterior() / self.likelihood, r.beta, r.drift)
Rating::new(skill.posterior().cavity(self.likelihood), r.beta, r.drift)
.with_drift_scale(r.drift_scale)
}
}
@@ -95,10 +97,10 @@ pub(crate) struct Event {
}
impl Event {
pub(crate) fn iter_agents(&self) -> impl Iterator<Item = Index> + '_ {
pub(crate) fn iter_competitors(&self) -> impl Iterator<Item = Index> + '_ {
self.teams
.iter()
.flat_map(|t| t.items.iter().map(|it| it.agent))
.flat_map(|t| t.items.iter().map(|it| it.competitor))
}
fn outputs(&self) -> Vec<f64> {
@@ -112,14 +114,14 @@ impl Event {
&self,
forward: bool,
skills: &SkillStore,
agents: &CompetitorStore<T, D>,
competitors: &CompetitorStore<T, D>,
) -> Vec<Vec<Rating<T, D>>> {
self.teams
.iter()
.map(|team| {
team.items
.iter()
.map(|item| item.within_prior(forward, skills, agents))
.map(|item| item.within_prior(forward, skills, competitors))
.collect::<Vec<_>>()
})
.collect::<Vec<_>>()
@@ -133,18 +135,23 @@ impl Event {
fn compute<T: Time, D: Drift<T>>(
&self,
skills: &SkillStore,
agents: &CompetitorStore<T, D>,
competitors: &CompetitorStore<T, D>,
p_draw: f64,
convergence: crate::ConvergenceOptions,
arena: &mut ScratchArena,
) -> EventUpdate {
let teams = self.within_priors(false, skills, agents);
let teams = self.within_priors(false, skills, competitors);
let result = self.outputs();
let g = match self.kind {
EventKind::Ranked => {
Game::ranked_with_arena(teams, &result, &self.weights, p_draw, convergence, arena)
}
EventKind::Scored { score_sigma } => Game::scored_with_arena(
EventKind::Ranked => GameRef::ranked_with_arena(
teams,
&result,
&self.weights,
p_draw,
convergence,
arena,
),
EventKind::Scored { score_sigma } => GameRef::scored_with_arena(
teams,
&result,
&self.weights,
@@ -166,7 +173,7 @@ impl Event {
for (i, item) in team.items.iter_mut().enumerate() {
let fresh = update.likelihoods[t][i];
let old_likelihood = skills.at(item.slot).likelihood;
let new_likelihood = (old_likelihood / item.likelihood) * fresh;
let new_likelihood = old_likelihood.cavity(item.likelihood).ep_product(fresh);
skills.at_mut(item.slot).likelihood = new_likelihood;
item.likelihood = fresh;
}
@@ -179,12 +186,12 @@ impl Event {
fn iteration_direct<T: Time, D: Drift<T>>(
&mut self,
skills: &mut SkillStore,
agents: &CompetitorStore<T, D>,
competitors: &CompetitorStore<T, D>,
p_draw: f64,
convergence: crate::ConvergenceOptions,
arena: &mut ScratchArena,
) {
let update = self.compute(skills, agents, p_draw, convergence, arena);
let update = self.compute(skills, competitors, p_draw, convergence, arena);
self.apply(skills, update);
}
}
@@ -228,7 +235,7 @@ pub struct TimeSlice<T: Time = i64> {
}
impl<T: Time> TimeSlice<T> {
pub fn new(time: T, p_draw: f64, convergence: crate::ConvergenceOptions) -> Self {
pub(crate) fn new(time: T, p_draw: f64, convergence: crate::ConvergenceOptions) -> Self {
Self {
events: Vec::new(),
skills: SkillStore::new(),
@@ -255,7 +262,7 @@ impl<T: Time> TimeSlice<T> {
}
let cg = color_greedy(n, |ev_idx| {
self.events[ev_idx].iter_agents().collect::<Vec<_>>()
self.events[ev_idx].iter_competitors().collect::<Vec<_>>()
});
let mut reordered: Vec<Event> = Vec::with_capacity(n);
@@ -282,17 +289,17 @@ impl<T: Time> TimeSlice<T> {
);
}
pub fn add_events<D: Drift<T>>(
pub(crate) fn add_events<D: Drift<T>>(
&mut self,
composition: Vec<Vec<Vec<Index>>>,
results: Option<Vec<Vec<f64>>>,
weights: Option<Vec<Vec<Vec<f64>>>>,
kinds: Vec<EventKind>,
agents: &CompetitorStore<T, D>,
competitors: &CompetitorStore<T, D>,
) {
let mut unique = Vec::with_capacity(10);
let this_agent = composition.iter().flatten().flatten().filter(|idx| {
let these_competitors = composition.iter().flatten().flatten().filter(|idx| {
if !unique.contains(idx) {
unique.push(*idx);
@@ -302,10 +309,10 @@ impl<T: Time> TimeSlice<T> {
false
});
for idx in this_agent {
let elapsed = compute_elapsed(agents[*idx].last_time.as_ref(), &self.time);
for idx in these_competitors {
let elapsed = compute_elapsed(competitors[*idx].last_time.as_ref(), &self.time);
let forward = agents[*idx].receive(&self.time);
let forward = competitors[*idx].receive(&self.time);
if let Some(skill) = self.skills.get_mut(*idx) {
skill.elapsed = elapsed;
@@ -332,12 +339,12 @@ impl<T: Time> TimeSlice<T> {
.map(|(t, team)| {
let items = team
.iter()
.map(|&agent| Item {
agent,
.map(|&competitor| Item {
competitor,
// Every participant was inserted into `skills`
// just above, so the slot always resolves.
slot: skills
.slot_of(agent)
.slot_of(competitor)
.expect("participant must be present in the slice store"),
likelihood: N_INF,
})
@@ -376,7 +383,7 @@ impl<T: Time> TimeSlice<T> {
self.color_groups_dirty = true;
self.iteration(from, agents);
self.iteration(from, competitors);
}
pub(crate) fn posteriors(&self) -> HashMap<Index, Gaussian> {
@@ -393,7 +400,11 @@ impl<T: Time> TimeSlice<T> {
/// Panics if an event references a competitor with no entry in this
/// slice's skill store. `add_events` inserts one for every participant, so
/// this cannot happen for slices built through the public API.
pub fn iteration<D: Drift<T>>(&mut self, from: usize, agents: &CompetitorStore<T, D>) {
pub(crate) fn iteration<D: Drift<T>>(
&mut self,
from: usize,
competitors: &CompetitorStore<T, D>,
) {
if from == 0 && self.color_groups_dirty {
self.recompute_color_groups();
}
@@ -401,11 +412,11 @@ impl<T: Time> TimeSlice<T> {
if from > 0 || self.color_groups.is_empty() {
// Initial pass (add_events) or no color groups yet: simple sequential sweep.
for event in self.events.iter_mut().skip(from) {
let teams = event.within_priors(false, &self.skills, agents);
let teams = event.within_priors(false, &self.skills, competitors);
let result = event.outputs();
let g = match event.kind {
EventKind::Ranked => Game::ranked_with_arena(
EventKind::Ranked => GameRef::ranked_with_arena(
teams,
&result,
&event.weights,
@@ -413,7 +424,7 @@ impl<T: Time> TimeSlice<T> {
self.convergence,
&mut self.arena,
),
EventKind::Scored { score_sigma } => Game::scored_with_arena(
EventKind::Scored { score_sigma } => GameRef::scored_with_arena(
teams,
&result,
&event.weights,
@@ -426,8 +437,9 @@ impl<T: Time> TimeSlice<T> {
for (t, team) in event.teams.iter_mut().enumerate() {
for (i, item) in team.items.iter_mut().enumerate() {
let old_likelihood = self.skills.at(item.slot).likelihood;
let new_likelihood =
(old_likelihood / item.likelihood) * g.likelihoods[t][i];
let new_likelihood = old_likelihood
.cavity(item.likelihood)
.ep_product(g.likelihoods[t][i]);
self.skills.at_mut(item.slot).likelihood = new_likelihood;
item.likelihood = g.likelihoods[t][i];
}
@@ -436,14 +448,14 @@ impl<T: Time> TimeSlice<T> {
event.log_evidence = g.log_evidence;
}
} else {
self.sweep_color_groups(agents);
self.sweep_color_groups(competitors);
}
}
/// Full event sweep using the color-group partition. Colors are processed
/// sequentially; within each color the inner loop is parallel under rayon.
///
/// Events in one color group touch disjoint agent sets, so none of them
/// Events in one color group touch disjoint competitor sets, so none of them
/// can observe another's writes. That makes the sweep separable: inference
/// runs concurrently over shared `&self.skills`, and the resulting updates
/// are folded in afterwards in index order. Splitting it this way needs no
@@ -451,7 +463,7 @@ impl<T: Time> TimeSlice<T> {
/// across thread counts because the apply order does not depend on which
/// worker finished first.
#[cfg(feature = "rayon")]
fn sweep_color_groups<D: Drift<T>>(&mut self, agents: &CompetitorStore<T, D>) {
fn sweep_color_groups<D: Drift<T>>(&mut self, competitors: &CompetitorStore<T, D>) {
use rayon::prelude::*;
thread_local! {
@@ -483,7 +495,7 @@ impl<T: Time> TimeSlice<T> {
let mut arena = cell.borrow_mut();
arena.reset();
ev.compute(skills, agents, p_draw, convergence, &mut arena)
ev.compute(skills, competitors, p_draw, convergence, &mut arena)
})
})
.collect();
@@ -495,7 +507,7 @@ impl<T: Time> TimeSlice<T> {
for ev in &mut self.events[range] {
ev.iteration_direct(
&mut self.skills,
agents,
competitors,
p_draw,
self.convergence,
&mut self.arena,
@@ -509,7 +521,7 @@ impl<T: Time> TimeSlice<T> {
/// Events within each color group are updated inline — no EventOutput allocation —
/// matching the T2 performance profile.
#[cfg(not(feature = "rayon"))]
fn sweep_color_groups<D: Drift<T>>(&mut self, agents: &CompetitorStore<T, D>) {
fn sweep_color_groups<D: Drift<T>>(&mut self, competitors: &CompetitorStore<T, D>) {
for color_idx in 0..self.color_groups.groups.len() {
if self.color_groups.groups[color_idx].is_empty() {
continue;
@@ -523,7 +535,7 @@ impl<T: Time> TimeSlice<T> {
for ev in &mut self.events[range] {
ev.iteration_direct(
&mut self.skills,
agents,
competitors,
p_draw,
self.convergence,
&mut self.arena,
@@ -544,7 +556,7 @@ impl<T: Time> TimeSlice<T> {
/// schedule default.
pub(crate) fn iterate_to_convergence<D: Drift<T>>(
&mut self,
agents: &CompetitorStore<T, D>,
competitors: &CompetitorStore<T, D>,
) -> usize {
use crate::{tuple_gt, tuple_max};
@@ -557,7 +569,7 @@ impl<T: Time> TimeSlice<T> {
while tuple_gt(step, epsilon) && i < max_iter {
let old = self.posteriors();
self.iteration(0, agents);
self.iteration(0, competitors);
let new = self.posteriors();
@@ -575,37 +587,37 @@ impl<T: Time> TimeSlice<T> {
i
}
pub(crate) fn forward_prior_out(&self, agent: &Index) -> Gaussian {
let skill = self.skills.get(*agent).unwrap();
skill.forward * skill.likelihood
pub(crate) fn forward_prior_out(&self, competitor: &Index) -> Gaussian {
let skill = self.skills.get(*competitor).unwrap();
skill.forward.ep_product(skill.likelihood)
}
pub(crate) fn backward_prior_out<D: Drift<T>>(
&self,
agent: &Index,
agents: &CompetitorStore<T, D>,
competitor: &Index,
competitors: &CompetitorStore<T, D>,
) -> Gaussian {
let skill = self.skills.get(*agent).unwrap();
let n = skill.likelihood * skill.backward;
let skill = self.skills.get(*competitor).unwrap();
let n = skill.likelihood.ep_product(skill.backward);
n.forget(
agents[*agent]
competitors[*competitor]
.rating
.drift_variance_for_elapsed(skill.elapsed),
)
}
pub(crate) fn new_backward_info<D: Drift<T>>(&mut self, agents: &CompetitorStore<T, D>) {
for (agent, skill) in self.skills.iter_mut() {
skill.backward = agents[agent].message.unwrap_or(N_INF);
pub(crate) fn new_backward_info<D: Drift<T>>(&mut self, competitors: &CompetitorStore<T, D>) {
for (competitor, skill) in self.skills.iter_mut() {
skill.backward = competitors[competitor].message.unwrap_or(N_INF);
}
self.iteration(0, agents);
self.iteration(0, competitors);
}
pub(crate) fn new_forward_info<D: Drift<T>>(&mut self, agents: &CompetitorStore<T, D>) {
for (agent, skill) in self.skills.iter_mut() {
skill.forward = agents[agent].receive_for_elapsed(skill.elapsed);
pub(crate) fn new_forward_info<D: Drift<T>>(&mut self, competitors: &CompetitorStore<T, D>) {
for (competitor, skill) in self.skills.iter_mut() {
skill.forward = competitors[competitor].receive_for_elapsed(skill.elapsed);
}
self.iteration(0, agents);
self.iteration(0, competitors);
}
/// Run this slice's events on forward (filtering) information alone.
@@ -615,10 +627,18 @@ impl<T: Time> TimeSlice<T> {
/// configured prior. The sweep runs on a scratch copy, so the real slice
/// is untouched — which is what makes the filtered estimates independent
/// of whether `History::converge` has run.
/// One forward-only step for this slice.
///
/// `targets` restricts only the *evidence sum*, to events in which at
/// least one target competitor appears; an empty set means no restriction.
/// The forward messages are always built from every event in the slice —
/// restricting those instead would answer a different question (a history
/// in which the other events never happened), not a held-out one.
pub(crate) fn filtered_step<D: Drift<T>>(
&self,
incoming: &HashMap<Index, Gaussian>,
agents: &CompetitorStore<T, D>,
competitors: &CompetitorStore<T, D>,
targets: &std::collections::HashSet<Index>,
) -> FilteredStep {
let mut scratch = TimeSlice {
events: self.events.clone(),
@@ -641,16 +661,16 @@ impl<T: Time> TimeSlice<T> {
event.log_evidence = 0.0;
}
for (agent, skill) in self.skills.iter() {
let rating = &agents[agent].rating;
for (competitor, skill) in self.skills.iter() {
let rating = &competitors[competitor].rating;
let forward = match incoming.get(&agent) {
let forward = match incoming.get(&competitor) {
Some(message) => message.forget(rating.drift_variance_for_elapsed(skill.elapsed)),
None => rating.prior,
};
let slot = scratch.skills.insert(
agent,
competitor,
Skill {
forward,
backward: N_INF,
@@ -666,19 +686,31 @@ impl<T: Time> TimeSlice<T> {
// than leave it to be rediscovered after it breaks.
debug_assert_eq!(
Some(slot),
self.skills.slot_of(agent),
"scratch slot must match the real slice's slot for {agent:?}"
self.skills.slot_of(competitor),
"scratch slot must match the real slice's slot for {competitor:?}"
);
}
scratch.iterate_to_convergence(agents);
scratch.iterate_to_convergence(competitors);
FilteredStep {
log_evidence: scratch.events.iter().map(|event| event.log_evidence).sum(),
log_evidence: scratch
.events
.iter()
.filter(|event| {
targets.is_empty()
|| event
.teams
.iter()
.flat_map(|team| &team.items)
.any(|item| targets.contains(&item.competitor))
})
.map(|event| event.log_evidence)
.sum(),
posteriors: scratch
.skills
.iter()
.map(|(agent, skill)| (agent, skill.posterior()))
.map(|(competitor, skill)| (competitor, skill.posterior()))
.collect(),
}
}
@@ -687,7 +719,7 @@ impl<T: Time> TimeSlice<T> {
&self,
targets: &[Index],
forward: bool,
agents: &CompetitorStore<T, D>,
competitors: &CompetitorStore<T, D>,
) -> f64 {
// Hashed once rather than scanned per player per event, so a
// `log_evidence_for` with many keys is not quadratic.
@@ -696,11 +728,11 @@ impl<T: Time> TimeSlice<T> {
let mut arena = ScratchArena::new();
let run_event = |event: &Event, arena: &mut ScratchArena| -> f64 {
let teams = event.within_priors(forward, &self.skills, agents);
let teams = event.within_priors(forward, &self.skills, competitors);
let result = event.outputs();
match event.kind {
EventKind::Ranked => {
Game::ranked_with_arena(
GameRef::ranked_with_arena(
teams,
&result,
&event.weights,
@@ -711,7 +743,7 @@ impl<T: Time> TimeSlice<T> {
.log_evidence
}
EventKind::Scored { score_sigma } => {
Game::scored_with_arena(
GameRef::scored_with_arena(
teams,
&result,
&event.weights,
@@ -741,7 +773,7 @@ impl<T: Time> TimeSlice<T> {
.teams
.iter()
.flat_map(|team| &team.items)
.any(|item| target_set.contains(&item.agent))
.any(|item| target_set.contains(&item.competitor))
})
.map(|event| run_event(event, &mut arena))
.sum()
@@ -753,27 +785,36 @@ impl<T: Time> TimeSlice<T> {
.teams
.iter()
.flat_map(|team| &team.items)
.any(|item| target_set.contains(&item.agent))
.any(|item| target_set.contains(&item.competitor))
})
.map(|event| event.log_evidence)
.sum()
}
}
pub fn get_composition(&self) -> Vec<Vec<Vec<Index>>> {
/// Test-only: reads the slice's shape back for assertions.
#[cfg(test)]
pub(crate) fn get_composition(&self) -> Vec<Vec<Vec<Index>>> {
self.events
.iter()
.map(|event| {
event
.teams
.iter()
.map(|team| team.items.iter().map(|item| item.agent).collect::<Vec<_>>())
.map(|team| {
team.items
.iter()
.map(|item| item.competitor)
.collect::<Vec<_>>()
})
.collect::<Vec<_>>()
})
.collect::<Vec<_>>()
}
pub fn get_results(&self) -> Vec<Vec<f64>> {
/// Test-only: reads the slice's shape back for assertions.
#[cfg(test)]
pub(crate) fn get_results(&self) -> Vec<Vec<f64>> {
self.events
.iter()
.map(|event| {
@@ -809,13 +850,85 @@ pub(crate) fn compute_elapsed<T: Time>(last: Option<&T>, current: &T) -> i64 {
elapsed.max(0)
}
impl<T: Time> TimeSlice<T> {
/// This slice's scored event factors, as contrasts over competitors.
///
/// Message passing produces per-competitor marginals and throws the
/// correlation away — `Item::likelihood` is already the projection of an
/// event's factor onto one competitor. So a joint has to be rebuilt from
/// the factor structure rather than recovered from the messages.
///
/// Usefully, a precision matrix depends only on *structure* — who played
/// whom, with what weights and what observation noise — and not on the
/// observed outcomes. The means are already exact, so only the second
/// moment needs rebuilding.
///
/// Each entry is a contrast and the observation variance that sits on it.
/// Ranked events contribute nothing: their truncation factors are EP
/// approximations that inference does not retain.
pub(crate) fn scored_contrasts<D: Drift<T>>(
&self,
competitors: &CompetitorStore<T, D>,
) -> Vec<(Vec<(Index, f64)>, f64)> {
let mut out = Vec::new();
for event in &self.events {
let EventKind::Scored { score_sigma } = event.kind else {
continue;
};
// Teams best-first, matching the diff chain inference builds.
let mut order: Vec<usize> = (0..event.teams.len()).collect();
order.sort_by(|&a, &b| {
event.teams[b]
.output
.partial_cmp(&event.teams[a].output)
.unwrap_or(std::cmp::Ordering::Equal)
});
for pair in order.windows(2) {
let (hi, lo) = (pair[0], pair[1]);
let mut contrast: Vec<(Index, f64)> = Vec::new();
let mut noise = score_sigma * score_sigma;
for (team, sign) in [(hi, 1.0), (lo, -1.0)] {
for (m, item) in event.teams[team].items.iter().enumerate() {
let w = event.weights[team][m];
noise += w * w * competitors[item.competitor].rating.beta.powi(2);
contrast.push((item.competitor, sign * w));
}
}
out.push((contrast, noise));
}
}
out
}
/// True when every event here is scored, so the joint is exact.
pub(crate) fn all_scored(&self) -> bool {
self.events
.iter()
.all(|e| matches!(e.kind, EventKind::Scored { .. }))
}
/// The competitors appearing in this slice, with the elapsed count since
/// each one's previous appearance.
pub(crate) fn appearances(&self) -> impl Iterator<Item = (Index, i64)> + '_ {
self.skills
.keys()
.map(|idx| (idx, self.skills.get(idx).expect("slice key").elapsed))
}
}
#[cfg(test)]
mod tests {
use approx::assert_ulps_eq;
use super::*;
use crate::{
KeyTable, competitor::Competitor, drift::ConstantDrift, rating::Rating,
competitor::Competitor, drift::ConstantDrift, key_table::KeyTable, rating::Rating,
storage::CompetitorStore,
};
@@ -830,16 +943,16 @@ mod tests {
let e = index_map.get_or_create("e");
let f = index_map.get_or_create("f");
let mut agents: CompetitorStore<i64, ConstantDrift> = CompetitorStore::new();
let mut competitors: CompetitorStore<i64, ConstantDrift> = CompetitorStore::new();
for agent in [a, b, c, d, e, f] {
agents.insert(
agent,
for competitor in [a, b, c, d, e, f] {
competitors.insert(
competitor,
Competitor {
rating: Rating::new(
Gaussian::from_ms(25.0, 25.0 / 3.0),
25.0 / 6.0,
ConstantDrift(25.0 / 300.0),
ConstantDrift::new(25.0 / 300.0),
),
..Default::default()
},
@@ -857,7 +970,7 @@ mod tests {
Some(vec![vec![1.0, 0.0], vec![0.0, 1.0], vec![1.0, 0.0]]),
None,
vec![EventKind::Ranked; 3],
&agents,
&competitors,
);
let post = time_slice.posteriors();
@@ -893,7 +1006,7 @@ mod tests {
epsilon = 1e-6
);
assert_eq!(time_slice.iterate_to_convergence(&agents), 1);
assert_eq!(time_slice.iterate_to_convergence(&competitors), 1);
}
#[test]
@@ -907,16 +1020,16 @@ mod tests {
let e = index_map.get_or_create("e");
let f = index_map.get_or_create("f");
let mut agents: CompetitorStore<i64, ConstantDrift> = CompetitorStore::new();
let mut competitors: CompetitorStore<i64, ConstantDrift> = CompetitorStore::new();
for agent in [a, b, c, d, e, f] {
agents.insert(
agent,
for competitor in [a, b, c, d, e, f] {
competitors.insert(
competitor,
Competitor {
rating: Rating::new(
Gaussian::from_ms(25.0, 25.0 / 3.0),
25.0 / 6.0,
ConstantDrift(25.0 / 300.0),
ConstantDrift::new(25.0 / 300.0),
),
..Default::default()
},
@@ -934,7 +1047,7 @@ mod tests {
Some(vec![vec![1.0, 0.0], vec![0.0, 1.0], vec![1.0, 0.0]]),
None,
vec![EventKind::Ranked; 3],
&agents,
&competitors,
);
let post = time_slice.posteriors();
@@ -955,7 +1068,7 @@ mod tests {
epsilon = 1e-6
);
assert!(time_slice.iterate_to_convergence(&agents) > 1);
assert!(time_slice.iterate_to_convergence(&competitors) > 1);
let post = time_slice.posteriors();
@@ -987,16 +1100,16 @@ mod tests {
let e = index_map.get_or_create("e");
let f = index_map.get_or_create("f");
let mut agents: CompetitorStore<i64, ConstantDrift> = CompetitorStore::new();
let mut competitors: CompetitorStore<i64, ConstantDrift> = CompetitorStore::new();
for agent in [a, b, c, d, e, f] {
agents.insert(
agent,
for competitor in [a, b, c, d, e, f] {
competitors.insert(
competitor,
Competitor {
rating: Rating::new(
Gaussian::from_ms(25.0, 25.0 / 3.0),
25.0 / 6.0,
ConstantDrift(25.0 / 300.0),
ConstantDrift::new(25.0 / 300.0),
),
..Default::default()
},
@@ -1014,10 +1127,10 @@ mod tests {
Some(vec![vec![1.0, 0.0], vec![0.0, 1.0], vec![1.0, 0.0]]),
None,
vec![EventKind::Ranked; 3],
&agents,
&competitors,
);
time_slice.iterate_to_convergence(&agents);
time_slice.iterate_to_convergence(&competitors);
let post = time_slice.posteriors();
@@ -1046,12 +1159,12 @@ mod tests {
Some(vec![vec![1.0, 0.0], vec![0.0, 1.0], vec![1.0, 0.0]]),
None,
vec![EventKind::Ranked; 3],
&agents,
&competitors,
);
assert_eq!(time_slice.events.len(), 6);
time_slice.iterate_to_convergence(&agents);
time_slice.iterate_to_convergence(&competitors);
let post = time_slice.posteriors();
@@ -1090,16 +1203,16 @@ mod tests {
let c = index_map.get_or_create("c");
let d = index_map.get_or_create("d");
let mut agents: CompetitorStore<i64, ConstantDrift> = CompetitorStore::new();
let mut competitors: CompetitorStore<i64, ConstantDrift> = CompetitorStore::new();
for agent in [a, b, c, d] {
agents.insert(
agent,
for competitor in [a, b, c, d] {
competitors.insert(
competitor,
Competitor {
rating: Rating::new(
Gaussian::from_ms(25.0, 25.0 / 3.0),
25.0 / 6.0,
ConstantDrift(25.0 / 300.0),
ConstantDrift::new(25.0 / 300.0),
),
..Default::default()
},
@@ -1117,7 +1230,7 @@ mod tests {
Some(vec![vec![1.0, 0.0], vec![1.0, 0.0], vec![1.0, 0.0]]),
None,
vec![EventKind::Ranked; 3],
&agents,
&competitors,
);
assert_eq!(ts.color_groups.n_colors(), 2);
@@ -1128,16 +1241,24 @@ mod tests {
assert_eq!(ts.color_groups.color_range(1), 2..3);
// Events at positions 0 and 1 (color 0) must be disjoint — verify by
// checking that the agent sets of self.events[0] and self.events[1] do
// not include the agent at self.events[2].
let agents_in_ev2: Vec<Index> = ts.events[2].iter_agents().collect();
let agents_in_ev0: Vec<Index> = ts.events[0].iter_agents().collect();
let agents_in_ev1: Vec<Index> = ts.events[1].iter_agents().collect();
// checking that the competitor sets of self.events[0] and self.events[1] do
// not include the competitor at self.events[2].
let competitors_in_ev2: Vec<Index> = ts.events[2].iter_competitors().collect();
let competitors_in_ev0: Vec<Index> = ts.events[0].iter_competitors().collect();
let competitors_in_ev1: Vec<Index> = ts.events[1].iter_competitors().collect();
// ev0 and ev1 must be disjoint from each other (color-0 invariant).
assert!(agents_in_ev0.iter().all(|ag| !agents_in_ev1.contains(ag)));
// ev2 must share an agent with ev0 or ev1 (it needed its own color).
let ev2_overlaps_ev0 = agents_in_ev2.iter().any(|ag| agents_in_ev0.contains(ag));
let ev2_overlaps_ev1 = agents_in_ev2.iter().any(|ag| agents_in_ev1.contains(ag));
assert!(
competitors_in_ev0
.iter()
.all(|ag| !competitors_in_ev1.contains(ag))
);
// ev2 must share an competitor with ev0 or ev1 (it needed its own color).
let ev2_overlaps_ev0 = competitors_in_ev2
.iter()
.any(|ag| competitors_in_ev0.contains(ag));
let ev2_overlaps_ev1 = competitors_in_ev2
.iter()
.any(|ag| competitors_in_ev1.contains(ag));
assert!(ev2_overlaps_ev0 || ev2_overlaps_ev1);
}
}
+147
View File
@@ -0,0 +1,147 @@
//! What an additive model does to uncertainty, and why "add the marginals" is
//! unsafe in one direction and merely wasteful in the other.
//!
//! Structurally this is the shape a joint player/layout model takes: every
//! observation measures a *sum* of nodes against a reference, so the data pins
//! differences and leaves the overall level to the prior. That is the classic
//! rating-scale indeterminacy, not a defect.
//!
//! The consequence for a consumer is that combining marginals is wrong in
//! opposite directions depending on the combination, which is worth pinning
//! because the unsafe direction is not the one you would guess:
//!
//! - **Differences** (`a - b`): the shared level cancels, so the exact width is
//! small — and adding marginals lands within a couple of percent of it here,
//! because the loopy underestimate offsets the ignored correlation.
//! - **Sums** (`a + b`): the shared level does *not* cancel, so the exact width
//! is large, and adding marginals is roughly five times too narrow. That is
//! overconfident, and it is the direction that publishes a claim the data
//! does not support.
use smallvec::smallvec;
use trueskill_tt::{ConstantDrift, ConvergenceOptions, Event, History, Member, Outcome, Team};
#[test]
fn additive_structure_makes_sums_wide_and_differences_tight() {
// Structurally like ustat: every round is (player + hole) measured against
// a fixed reference. Only SUMS are pinned by the data; the split between
// player and hole is pinned only by the prior.
let players = ["p0", "p1", "p2"];
let holes = ["h0", "h1"];
let mut h: History = History::builder()
.mu(0.0)
.sigma(6.0)
.beta(1.0)
.score_sigma(2.0)
.drift(ConstantDrift::new(0.0))
.convergence(ConvergenceOptions {
max_iter: 20_000,
epsilon: 1e-12,
alpha: 1.0,
})
.build();
let mut seed = 3u64;
let mut rnd = move || {
seed ^= seed << 13;
seed ^= seed >> 7;
seed ^= seed << 17;
seed
};
// true skills, so we know what the data encodes
let truth_p = [2.0, 0.0, -2.0];
let truth_h = [1.0, -1.0];
let mut events = Vec::new();
for _ in 0..60 {
let p = (rnd() as usize) % 3;
let q = (rnd() as usize) % 2;
let noise = ((rnd() % 1000) as f64 / 1000.0 - 0.5) * 2.0;
let score = truth_p[p] + truth_h[q] + noise;
events.push(Event {
time: 1,
teams: smallvec![
Team::with_members([Member::new(players[p]), Member::new(holes[q])]),
Team::with_members([Member::new("reference")]),
],
outcome: Outcome::scores([score, 0.0]),
});
}
h.add_events(events).unwrap();
let r = h.converge().unwrap();
assert!(r.converged, "{:?}", r.final_step);
println!("\n== marginals (what current_skill reports) ==");
for k in players.iter().chain(holes.iter()) {
let g = h.current_skill(k).unwrap();
println!(" {k}: mu {:>8.4} sigma {:>8.4}", g.mu(), g.sigma());
}
println!("\n== the same nodes via posterior_of (exact marginal) ==");
for k in players.iter().chain(holes.iter()) {
let g = h.joint().unwrap().posterior_of(&[(k, 1.0)]).unwrap();
println!(" {k}: mu {:>8.4} sigma {:>8.4}", g.mu(), g.sigma());
}
println!("\n== combinations the data actually pins ==");
for (label, terms) in [
("p0 + h0 (a round)", vec![(&"p0", 1.0), (&"h0", 1.0)]),
(
"p0 - p1 (rank two players)",
vec![(&"p0", 1.0), (&"p1", -1.0)],
),
("p0 - p2", vec![(&"p0", 1.0), (&"p2", -1.0)]),
(
"h0 - h1 (rank two holes)",
vec![(&"h0", 1.0), (&"h1", -1.0)],
),
] {
let joint = h.joint().unwrap().posterior_of(&terms).unwrap();
// what a consumer gets today by adding marginals
let naive: f64 = terms
.iter()
.map(|(k, c)| c * c * h.current_skill(*k).unwrap().sigma().powi(2))
.sum::<f64>()
.sqrt();
println!(
" {label:<28} exact sigma {:>7.4} adding marginals {:>7.4} {:>5.2}x over",
joint.sigma(),
naive,
naive / joint.sigma()
);
let ratio = naive / joint.sigma();
if label.contains('+') {
assert!(
ratio < 0.5,
"{label}: adding marginals should be badly OVERconfident for a \
sum, got {ratio:.3}x"
);
} else {
assert!(
(0.8..1.25).contains(&ratio),
"{label}: adding marginals happens to be close for a difference, \
got {ratio:.3}x"
);
}
}
// A single node in an additive model is weakly identified: its exact
// posterior is far wider than message passing reports, because the level it
// shares with its partners is pinned only by the prior.
for k in players.iter().chain(holes.iter()) {
let bp = h.current_skill(k).unwrap().sigma();
let exact = h
.joint()
.unwrap()
.posterior_of(&[(k, 1.0)])
.unwrap()
.sigma();
assert!(
exact > 3.0 * bp,
"{k}: exact marginal {exact} should be much wider than the reported \
{bp} in an additive model"
);
}
}
+20 -20
View File
@@ -11,7 +11,7 @@ fn add_events_bulk_via_iter() {
.sigma(2.0)
.beta(1.0)
.p_draw(0.0)
.drift(ConstantDrift(0.0))
.drift(ConstantDrift::new(0.0))
.convergence(ConvergenceOptions {
max_iter: 30,
epsilon: 1e-6,
@@ -41,9 +41,9 @@ fn add_events_bulk_via_iter() {
h.add_events(events).unwrap();
let report = h.converge().unwrap();
assert!(report.converged);
assert!(h.lookup(&"a").is_some());
assert!(h.lookup(&"b").is_some());
assert!(h.lookup(&"c").is_some());
assert!(h.current_skill("a").is_some());
assert!(h.current_skill("b").is_some());
assert!(h.current_skill("c").is_some());
}
#[test]
@@ -53,7 +53,7 @@ fn add_events_draw() {
.sigma(25.0 / 3.0)
.beta(25.0 / 6.0)
.p_draw(0.25)
.drift(ConstantDrift(25.0 / 300.0))
.drift(ConstantDrift::new(25.0 / 300.0))
.build();
let events: Vec<Event<i64, &'static str>> = vec![Event {
@@ -65,7 +65,7 @@ fn add_events_draw() {
outcome: Outcome::draw(2),
}];
h.add_events(events).unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
}
#[test]
@@ -103,9 +103,9 @@ fn fluent_event_builder_basic() {
let report = h.converge().unwrap();
assert!(report.converged);
assert!(h.lookup(&"alice").is_some());
assert!(h.lookup(&"bob").is_some());
assert!(h.lookup(&"carol").is_some());
assert!(h.current_skill("alice").is_some());
assert!(h.current_skill("bob").is_some());
assert!(h.current_skill("carol").is_some());
}
#[test]
@@ -123,7 +123,7 @@ fn fluent_event_builder_winner_convenience() {
.winner(0)
.commit()
.unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
}
#[test]
@@ -141,7 +141,7 @@ fn fluent_event_builder_draw() {
.draw()
.commit()
.unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
}
#[test]
@@ -155,14 +155,14 @@ fn current_skill_and_learning_curve() {
.build();
h.record_winner(&"a", &"b", 1).unwrap();
h.record_winner(&"a", &"b", 2).unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
let a = h.current_skill(&"a").unwrap();
assert!(a.mu() > 25.0);
let b = h.current_skill(&"b").unwrap();
assert!(b.mu() < 25.0);
let a_curve = h.learning_curve(&"a");
let a_curve = h.learning_curve(&"a").unwrap();
assert_eq!(a_curve.len(), 2);
assert_eq!(a_curve[0].0, 1);
assert_eq!(a_curve[1].0, 2);
@@ -181,12 +181,12 @@ fn log_evidence_total_vs_subset() {
.sigma(6.0)
.beta(1.0)
.p_draw(0.0)
.drift(ConstantDrift(0.0))
.drift(ConstantDrift::new(0.0))
.build();
h.record_winner(&"a", &"b", 1).unwrap();
h.record_winner(&"b", &"a", 2).unwrap();
let total = h.log_evidence();
let a_only = h.log_evidence_for(&[&"a"]);
let a_only = h.log_evidence_for(&[&"a"]).unwrap();
assert!(total.is_finite());
assert!(a_only.is_finite());
}
@@ -201,9 +201,9 @@ fn predict_quality_two_teams() {
.p_draw(0.0)
.build();
h.record_winner(&"a", &"b", 1).unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
let q = h.predict_quality(&[&[&"a"], &[&"b"]]).unwrap();
let q = h.quality(&[&[&"a"], &[&"b"]]).unwrap();
assert!(q > 0.0 && q <= 1.0);
}
@@ -217,7 +217,7 @@ fn predict_outcome_two_teams_sums_to_one() {
.p_draw(0.0)
.build();
h.record_winner(&"a", &"b", 1).unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
let p = h.predict_outcome(&[&[&"a"], &[&"b"]]).unwrap();
let wins = p.win_probabilities();
@@ -236,7 +236,7 @@ fn fluent_event_builder_scores() {
.mu(25.0)
.sigma(25.0 / 3.0)
.beta(25.0 / 6.0)
.drift(ConstantDrift(0.0))
.drift(ConstantDrift::new(0.0))
.build();
h.event(1)
@@ -245,7 +245,7 @@ fn fluent_event_builder_scores() {
.scores([12.0, 4.0])
.commit()
.unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
let a = h.current_skill(&"alice").unwrap();
let b = h.current_skill(&"bob").unwrap();
+14 -11
View File
@@ -64,13 +64,13 @@ fn a_prior_applies_to_a_new_competitor() {
let mut with = history();
with.add_events(vec![bout("a", "b", 0, Some(seeded), None)])
.unwrap();
with.converge().unwrap();
let _ = with.converge().unwrap();
let mut without = history();
without
.add_events(vec![bout("a", "b", 0, None, None)])
.unwrap();
without.converge().unwrap();
let _ = without.converge().unwrap();
assert!(
(skill_of(&with, "a").mu() - skill_of(&without, "a").mu()).abs() > 1.0,
@@ -91,7 +91,7 @@ fn a_prior_applies_to_a_competitor_the_history_already_knows() {
// "a" now exists. Configuring it here used to do nothing whatsoever.
late.add_events(vec![bout("a", "b", 1, Some(seeded), None)])
.unwrap();
late.converge().unwrap();
let _ = late.converge().unwrap();
let mut never = history();
never
@@ -100,7 +100,7 @@ fn a_prior_applies_to_a_competitor_the_history_already_knows() {
bout("a", "b", 1, None, None),
])
.unwrap();
never.converge().unwrap();
let _ = never.converge().unwrap();
assert!(
(skill_of(&late, "a").mu() - skill_of(&never, "a").mu()).abs() > 1.0,
@@ -122,7 +122,7 @@ fn a_prior_is_whole_history_scoped_not_per_event() {
.unwrap();
late.add_events(vec![bout("a", "b", 1, Some(seeded), None)])
.unwrap();
late.converge().unwrap();
let _ = late.converge().unwrap();
let mut early = history();
early
@@ -131,7 +131,7 @@ fn a_prior_is_whole_history_scoped_not_per_event() {
bout("a", "b", 1, Some(seeded), None),
])
.unwrap();
early.converge().unwrap();
let _ = early.converge().unwrap();
let (l, e) = (skill_of(&late, "a"), skill_of(&early, "a"));
assert!(
@@ -150,7 +150,7 @@ fn repeating_the_same_prior_is_inert() {
bout("a", "b", 1, None, None),
])
.unwrap();
once.converge().unwrap();
let _ = once.converge().unwrap();
let mut every_time = history();
every_time
@@ -159,7 +159,7 @@ fn repeating_the_same_prior_is_inert() {
bout("a", "b", 1, Some(seeded), None),
])
.unwrap();
every_time.converge().unwrap();
let _ = every_time.converge().unwrap();
let (o, e) = (skill_of(&once, "a"), skill_of(&every_time, "a"));
assert!(
@@ -184,7 +184,10 @@ fn a_batch_declaring_two_different_priors_is_rejected() {
assert!(
matches!(
err,
InferenceError::ConflictingCompetitorConfig { field: "prior", .. }
InferenceError::ConflictingCompetitorConfig {
field: trueskill_tt::CompetitorField::Prior,
..
}
),
"got {err:?}"
);
@@ -203,7 +206,7 @@ fn setting_one_field_late_leaves_the_other_alone() {
// Only the scale this time — the prior above must survive.
h.add_events(vec![bout("a", "b", 1, None, Some(0.5))])
.unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
let mut both_upfront = history();
both_upfront
@@ -212,7 +215,7 @@ fn setting_one_field_late_leaves_the_other_alone() {
bout("a", "b", 1, None, None),
])
.unwrap();
both_upfront.converge().unwrap();
let _ = both_upfront.converge().unwrap();
let (a, b) = (skill_of(&h, "a"), skill_of(&both_upfront, "a"));
assert!(
+232
View File
@@ -0,0 +1,232 @@
//! Every public entry point that takes a magnitude, in one place.
//!
//! This defect class was closed three times in one session and reopened twice,
//! because each fix validated the layer it had just touched and inferred the
//! rest: `HistoryBuilder` first, then `Game`'s own entry points, then the
//! constructors beneath both. A per-site fix cannot notice the site nobody
//! thought of.
//!
//! So this enumerates them. `sigma`, `beta` and `gamma` all enter inference
//! only as squares, which means a negative value does not fail — it behaves as
//! its absolute value, bit for bit, and the sign vanishes with no diagnostic.
//! Non-finite values poison every posterior derived from them.
//!
//! Adding a public constructor that takes one of these and not adding it here
//! is the failure this file exists to make harder.
use std::panic::{AssertUnwindSafe, catch_unwind};
use trueskill_tt::{ConstantDrift, Gaussian, History, Member, Outcome, Rating};
/// Did the entry point refuse the value, by panic or by `Err`?
fn refuses(f: impl FnOnce() -> bool) -> bool {
catch_unwind(AssertUnwindSafe(f)).unwrap_or(true)
}
/// One entry point, as a name and a closure that applies a value to it.
type Case = (&'static str, Box<dyn Fn(f64) -> bool>);
/// Entry points that must reject a negative magnitude.
///
/// Each closure returns `true` if it refused by returning an error; a panic is
/// also a refusal and is caught.
#[test]
fn every_magnitude_parameter_rejects_a_negative_value() {
let cases: Vec<Case> = vec![
(
"Gaussian::from_ms(sigma)",
Box::new(|v| {
let _ = Gaussian::from_ms(25.0, v);
false
}),
),
(
"Rating::new(beta)",
Box::new(|v| {
let _ = Rating::<i64, ConstantDrift>::new(
Gaussian::default(),
v,
ConstantDrift::new(0.0),
);
false
}),
),
(
"ConstantDrift::new(gamma)",
Box::new(|v| {
let _ = ConstantDrift::new(v);
false
}),
),
(
"HistoryBuilder::sigma",
Box::new(|v| {
let _ = History::builder().sigma(v);
false
}),
),
(
"HistoryBuilder::beta",
Box::new(|v| {
let _ = History::builder().beta(v);
false
}),
),
(
"HistoryBuilder::score_sigma",
Box::new(|v| {
let _ = History::builder().score_sigma(v);
false
}),
),
(
"HistoryBuilder::p_draw",
Box::new(|v| {
let _ = History::builder().p_draw(v);
false
}),
),
(
"Member::with_drift_scale (at ingestion)",
Box::new(|v| {
let mut h = History::builder().build();
h.add_events(vec![trueskill_tt::Event {
time: 1i64,
teams: smallvec::smallvec![
trueskill_tt::Team::with_members([Member::new("a").with_drift_scale(v)]),
trueskill_tt::Team::with_members([Member::new("b")]),
],
outcome: Outcome::winner(0, 2),
}])
.is_err()
}),
),
(
"Outcome::scores_with_noise (at ingestion)",
Box::new(|v| {
let mut h = History::builder().build();
h.add_events(vec![trueskill_tt::Event {
time: 1i64,
teams: smallvec::smallvec![
trueskill_tt::Team::with_members([Member::new("a")]),
trueskill_tt::Team::with_members([Member::new("b")]),
],
outcome: Outcome::scores_with_noise([3.0, 1.0], v),
}])
.is_err()
}),
),
];
let mut accepted = Vec::new();
for (name, f) in &cases {
if !refuses(|| f(-1.0)) {
accepted.push(*name);
}
}
assert!(
accepted.is_empty(),
"these accepted a negative magnitude, which is squared away silently \
rather than honoured or refused:\n {}",
accepted.join("\n ")
);
}
/// Same set, for NaN and infinity.
///
/// `Gaussian::from_ms` is deliberately absent: a broken fit produces a NaN
/// sigma legitimately and `converge` reports it as `NonFiniteResult`. Rejecting
/// it in the constructor turned that reporting path into a panic inside
/// inference — see the comment on `from_ms`.
#[test]
fn every_magnitude_parameter_rejects_a_non_finite_value() {
let cases: Vec<Case> = vec![
(
"Rating::new(beta)",
Box::new(|v| {
let _ = Rating::<i64, ConstantDrift>::new(
Gaussian::default(),
v,
ConstantDrift::new(0.0),
);
false
}),
),
(
"ConstantDrift::new(gamma)",
Box::new(|v| {
let _ = ConstantDrift::new(v);
false
}),
),
(
"HistoryBuilder::sigma",
Box::new(|v| {
let _ = History::builder().sigma(v);
false
}),
),
(
"HistoryBuilder::beta",
Box::new(|v| {
let _ = History::builder().beta(v);
false
}),
),
(
"HistoryBuilder::mu",
Box::new(|v| {
let _ = History::builder().mu(v);
false
}),
),
(
"HistoryBuilder::score_sigma",
Box::new(|v| {
let _ = History::builder().score_sigma(v);
false
}),
),
(
"HistoryBuilder::p_draw",
Box::new(|v| {
let _ = History::builder().p_draw(v);
false
}),
),
];
let mut accepted = Vec::new();
for (name, f) in &cases {
for bad in [f64::NAN, f64::INFINITY] {
if !refuses(|| f(bad)) {
accepted.push(format!("{name} accepted {bad}"));
}
}
}
assert!(
accepted.is_empty(),
"these accepted a non-finite magnitude:\n {}",
accepted.join("\n ")
);
}
/// The suite must not pass by refusing everything.
#[test]
fn ordinary_values_are_still_accepted() {
let _ = Gaussian::from_ms(25.0, 8.33);
let _ = Rating::<i64, ConstantDrift>::new(Gaussian::default(), 4.17, ConstantDrift::new(0.05));
let _ = ConstantDrift::new(0.0833);
let _ = History::builder()
.mu(25.0)
.sigma(8.33)
.beta(4.17)
.score_sigma(1.0)
.p_draw(0.1);
// Zero beta and zero gamma are legitimate, not degenerate.
let _ = ConstantDrift::new(0.0);
let _ = Rating::<i64, ConstantDrift>::new(Gaussian::default(), 0.0, ConstantDrift::new(0.0));
}
+154
View File
@@ -0,0 +1,154 @@
//! Stopping short of convergence is an error, not a flag on a success.
//!
//! A fit that hits `max_iter` is wrong by a little: every rating is finite,
//! the ordering looks sensible, and nothing in the numbers says they were
//! still moving. When that was `Ok` with `converged: false`, detecting it was
//! opt-in and `let _ = h.converge()` was the natural way to opt out — which is
//! how a real defect once hid in this crate's own suite.
use smallvec::smallvec;
use trueskill_tt::{
ConstantDrift, ConvergenceOptions, Event, History, InferenceError, Member, Outcome, Team,
};
type H = History;
fn duel(a: &'static str, b: &'static str, t: i64) -> Event<i64, &'static str> {
Event {
time: t,
teams: smallvec![
Team::with_members([Member::new(a)]),
Team::with_members([Member::new(b)]),
],
outcome: Outcome::scores([3.0, 1.0]),
}
}
fn capped(max_iter: usize) -> H {
History::builder()
.mu(0.0)
.sigma(6.0)
.beta(1.0)
.score_sigma(2.0)
.drift(ConstantDrift::new(0.5))
.convergence(ConvergenceOptions {
max_iter,
epsilon: 1e-13,
alpha: 1.0,
})
.build()
}
fn fill(h: &mut H) {
h.add_events((1..=6).map(|t| duel("a", "b", t)).collect::<Vec<_>>())
.unwrap();
}
#[test]
fn hitting_the_cap_is_an_error() {
let mut h = capped(1);
fill(&mut h);
let err = h.converge().unwrap_err();
match err {
InferenceError::NotConverged {
iterations,
final_step,
epsilon,
..
} => {
assert_eq!(iterations, 1);
assert!(
final_step.0 > epsilon || final_step.1 > epsilon,
"{final_step:?}"
);
}
other => panic!("expected NotConverged, got {other:?}"),
}
}
/// The message has to name what to do about it, since the fit looks fine.
#[test]
fn the_error_says_how_to_fix_it() {
let mut h = capped(1);
fill(&mut h);
let text = h.converge().unwrap_err().to_string();
assert!(text.contains("did not converge in 1 iterations"), "{text}");
assert!(text.contains("max_iter"), "{text}");
assert!(text.contains("alpha"), "{text}");
}
/// The escape hatch: a deliberately capped fit is still reachable.
#[test]
fn converge_partial_returns_the_short_fit() {
let mut h = capped(1);
fill(&mut h);
let report = h.converge_partial().unwrap();
assert_eq!(report.iterations, 1);
assert!(!report.converged);
assert!(h.current_skill(&"a").is_some());
}
/// Both agree when the fit does converge, so the strict path costs nothing.
#[test]
fn the_two_agree_on_a_converged_fit() {
let mut strict = capped(20_000);
fill(&mut strict);
let a = strict.converge().unwrap();
let mut partial = capped(20_000);
fill(&mut partial);
let b = partial.converge_partial().unwrap();
assert!(a.converged && b.converged);
assert_eq!(a.iterations, b.iterations);
assert_eq!(a.final_step, b.final_step);
}
/// The default cap must be high enough that an ordinary history clears it.
/// At the old value of 30 this history stopped short and said nothing.
#[test]
fn the_default_cap_clears_an_ordinary_history() {
let mut h: History<String> = History::builder()
.key_type::<String>()
.mu(0.0)
.sigma(6.0)
.beta(1.0)
.score_sigma(2.0)
.drift(ConstantDrift::new(0.05))
.build();
let mut events = Vec::new();
for t in 0..20i64 {
for j in 0..8usize {
let k = (t as usize) * 8 + j;
events.push(Event {
time: t,
teams: smallvec![
Team::with_members([Member::new(format!("p{}", k % 100))]),
Team::with_members([Member::new(format!("p{}", (k + 37) % 100))]),
],
outcome: Outcome::scores([3.0, 1.0]),
});
}
}
h.add_events(events).unwrap();
let report = h
.converge()
.expect("an ordinary history must converge by default");
assert!(
report.iterations > 30,
"needed {} sweeps",
report.iterations
);
assert!(report.iterations < trueskill_tt::ITERATIONS);
}
/// An empty history converges trivially rather than erroring.
#[test]
fn an_empty_history_converges() {
let mut h = capped(1);
let report = h.converge().unwrap();
assert!(report.converged);
assert_eq!(report.iterations, 0);
}
+161
View File
@@ -0,0 +1,161 @@
//! Determinism across *processes*, which an in-process test cannot see.
//!
//! Rust seeds its default hasher once per process, so every `HashMap`
//! iteration order is fixed for a run and varies between runs. A test that
//! compares results within one process therefore cannot detect a float sum
//! whose order comes from a map — all its samples share one seed.
//!
//! That is not hypothetical. `tests/determinism.rs` compares four thread counts
//! inside one process and passed throughout, while `posterior_of` was returning
//! two distinct bit patterns across 40 separate runs on identical input.
//!
//! This re-executes the test binary and compares `f64::to_bits`.
use std::{env, process::Command};
use smallvec::smallvec;
use trueskill_tt::{
ConstantDrift, ConvergenceOptions, Event, History, Member, Outcome, Team, UnknownKeys,
};
/// Set in the child so it reports instead of re-spawning.
const CHILD: &str = "TSTT_DETERMINISM_CHILD";
const RUNS: usize = 40;
type H = History<String>;
fn fitted() -> H {
let mut h: H = History::builder()
.key_type::<String>()
.mu(0.0)
.sigma(6.0)
.beta(1.0)
.score_sigma(2.0)
.drift(ConstantDrift::new(0.05))
.unknown_keys(UnknownKeys::Prior)
.convergence(ConvergenceOptions {
max_iter: 20_000,
epsilon: 1e-13,
alpha: 1.0,
})
.build();
let mut events = Vec::new();
for t in 0..12i64 {
for k in 0..6usize {
let a = format!("p{}", (t as usize * 6 + k) % 10);
let b = format!("p{}", (t as usize * 6 + k + 4) % 10);
events.push(Event {
time: t,
teams: smallvec![
Team::with_members([Member::new(a)]),
Team::with_members([Member::new(b)]),
],
outcome: Outcome::scores([3.0, 1.0]),
});
}
}
h.add_events(events).unwrap();
assert!(h.converge().unwrap().converged);
h
}
/// Every quantity that could plausibly depend on iteration order, as bits.
fn fingerprint() -> String {
let h = fitted();
// Unknown keys with UNEQUAL but COMPARABLE coefficients, which is what
// makes the sum order-sensitive.
//
// Equal terms sum order-independently and would make this pass vacuously.
// Terms of wildly different magnitudes are no better: the small ones fall
// below the running total's ULP and are absorbed whatever the order —
// measured, spreading these over nine decades dropped the detection rate
// to roughly one run in forty. Comparable sizes keep every term able to
// change the last bits.
let ghosts: Vec<String> = (0..24).map(|i| format!("ghost{i}")).collect();
let mut terms: Vec<(&String, f64)> = ghosts
.iter()
.enumerate()
.map(|(i, k)| (k, 1.0 + i as f64 * 0.37))
.collect();
let known = "p0".to_string();
terms.push((&known, -1.0));
let posterior = h.joint().unwrap().posterior_of(&terms).unwrap();
let a = "p0".to_string();
let b = "p1".to_string();
let target = [(&a, 1.0), (&b, -1.0)];
let teams: [&[&String]; 2] = [&[&a], &[&b]];
let evr = h
.joint()
.unwrap()
.expected_variance_reduction(&teams, &target)
.unwrap();
let curves = h.learning_curves();
let mut curve_bits: u64 = 0;
let mut keys: Vec<&String> = curves.keys().collect();
keys.sort();
for key in keys {
for (t, g) in &curves[key] {
curve_bits ^= (*t as u64).rotate_left(17)
^ g.mu().to_bits().rotate_left(31)
^ g.sigma().to_bits();
}
}
format!(
"post={:016x} evr={:016x} le={:016x} curves={curve_bits:016x}",
posterior.sigma().to_bits(),
evr.to_bits(),
h.log_evidence().to_bits(),
)
}
#[test]
fn results_are_identical_across_processes() {
if env::var(CHILD).is_ok() {
println!("FINGERPRINT {}", fingerprint());
return;
}
let exe = env::current_exe().expect("current exe");
let mut seen: Vec<String> = Vec::new();
for run in 0..RUNS {
let out = Command::new(&exe)
.args([
"results_are_identical_across_processes",
"--exact",
"--nocapture",
])
.env(CHILD, "1")
.output()
.expect("spawn child");
assert!(
out.status.success(),
"child {run} failed: {}",
String::from_utf8_lossy(&out.stderr)
);
let stdout = String::from_utf8_lossy(&out.stdout);
let line = stdout
.lines()
.find_map(|l| l.strip_prefix("FINGERPRINT "))
.unwrap_or_else(|| panic!("child {run} printed no fingerprint:\n{stdout}"))
.to_string();
seen.push(line);
}
let first = &seen[0];
let differing: Vec<&String> = seen.iter().filter(|s| *s != first).collect();
assert!(
differing.is_empty(),
"results differ across processes on identical input.\n {} of {RUNS} runs differed\n \
first: {first}\n differing: {}",
differing.len(),
differing[0]
);
}
+34 -19
View File
@@ -8,7 +8,7 @@ mod common;
use common::assert_finite;
use trueskill_tt::{
ConstantDrift, ConvergenceOptions, Game, GameOptions, Gaussian, History, InferenceError,
NullObserver, Outcome, Rating,
Outcome, Rating,
};
type R = Rating<i64, ConstantDrift>;
@@ -17,7 +17,7 @@ fn rating() -> R {
R::new(
Gaussian::from_ms(25.0, 25.0 / 3.0),
25.0 / 6.0,
ConstantDrift(25.0 / 300.0),
ConstantDrift::new(25.0 / 300.0),
)
}
@@ -126,8 +126,10 @@ fn empty_history_converges_trivially() {
/// indexed out of bounds in release, so this must run in both profiles.
#[test]
fn converge_on_an_empty_history_with_owned_keys() {
let mut history: History<i64, ConstantDrift, NullObserver, String> =
History::builder_with_key().score_sigma(5.0).build();
let mut history: History<String> = History::builder()
.key_type::<String>()
.score_sigma(5.0)
.build();
let report = history.converge().unwrap();
@@ -155,9 +157,10 @@ fn event_builder_rejects_a_weights_length_mismatch() {
matches!(
err,
InferenceError::MismatchedShape {
kind: "weights",
shape: trueskill_tt::Shape::Weights,
expected: 1,
got: 2,
..
}
),
"expected a weights MismatchedShape, got {err:?}"
@@ -170,8 +173,9 @@ fn event_builder_rejects_a_weights_length_mismatch() {
fn event_builder_weights_mismatch_leaves_the_history_untouched() {
let mut h = History::default();
// Two teams, so ingestion would otherwise succeed — a one-team event is
// rejected for an unrelated reason and would pass this vacuously.
// Two teams, so ingestion would otherwise succeed. A one-team event is
// rejected as `NotEnoughTeams` before the weights are ever examined, so
// building this with one team would pass vacuously.
let _ = h
.event(1)
.team(["a"])
@@ -180,7 +184,10 @@ fn event_builder_weights_mismatch_leaves_the_history_untouched() {
.winner(0)
.commit();
assert!(h.learning_curve("a").is_empty());
// The rejected event never reached the history, so "a" was never interned.
// `None` is the honest answer, and it is distinguishable from a competitor
// that IS known but has no appearances yet.
assert!(h.learning_curve("a").is_none());
}
#[test]
@@ -195,7 +202,7 @@ fn empty_event_stream_then_converge() {
fn empty_history_queries_do_not_panic() {
let h = History::default();
assert!(h.learning_curves().is_empty());
assert!(h.learning_curve("nobody").is_empty());
assert!(h.learning_curve("nobody").is_none());
assert!(h.current_skill("nobody").is_none());
}
@@ -215,13 +222,13 @@ fn scored_event_rejects_non_positive_sigma() {
.event(1)
.team(["a"])
.team(["b"])
.scores_with_sigma([3.0, 1.0], f64::NAN)
.scores_with_noise([3.0, 1.0], f64::NAN)
.commit()
.unwrap_err();
assert!(matches!(
err,
InferenceError::InvalidParameter {
name: "score_sigma",
parameter: trueskill_tt::Parameter::ScoreSigma,
..
}
));
@@ -280,8 +287,16 @@ fn log_evidence_survives_a_long_diff_chain() {
/// `erfc` approximation; the evidence floor keeps `ln` finite.
#[test]
fn log_evidence_finite_for_near_certain_outcome() {
let overwhelming = R::new(Gaussian::from_ms(5_000.0, 0.5), 1.0, ConstantDrift(0.0));
let hopeless = R::new(Gaussian::from_ms(-5_000.0, 0.5), 1.0, ConstantDrift(0.0));
let overwhelming = R::new(
Gaussian::from_ms(5_000.0, 0.5),
1.0,
ConstantDrift::new(0.0),
);
let hopeless = R::new(
Gaussian::from_ms(-5_000.0, 0.5),
1.0,
ConstantDrift::new(0.0),
);
let a = [overwhelming];
let b = [hopeless];
let teams: Vec<&[R]> = vec![&a, &b];
@@ -310,7 +325,7 @@ fn empty_history_has_no_filtered_estimates() {
assert!(history.filtered_learning_curves().is_empty());
assert!(history.filtered_learning_curve("nobody").is_empty());
assert!(history.filtered_learning_curve("nobody").is_none());
}
// --- Boundary inputs (#26) ----------------------------------------------
@@ -325,7 +340,7 @@ fn tight() -> ConvergenceOptions {
fn assert_curve_finite(h: &History, keys: &[&str], what: &str) {
for key in keys {
for (time, g) in h.learning_curve(*key) {
for (time, g) in h.learning_curve(*key).unwrap() {
assert!(
g.mu().is_finite() && g.sigma().is_finite(),
"{what}: non-finite posterior for {key} at t={time} (mu={} sigma={})",
@@ -351,7 +366,7 @@ fn zero_weight_does_not_produce_a_non_finite_posterior() {
.commit()
.expect("a zero weight is accepted today; update this test if that changes");
h.converge().unwrap();
let _ = h.converge().unwrap();
assert_curve_finite(&h, &["a", "b"], "zero weight");
}
@@ -368,7 +383,7 @@ fn negative_weight_does_not_produce_a_non_finite_posterior() {
.commit()
.expect("a negative weight is accepted today; update this test if that changes");
h.converge().unwrap();
let _ = h.converge().unwrap();
assert_curve_finite(&h, &["a", "b"], "negative weight");
}
@@ -389,7 +404,7 @@ fn out_of_order_timestamps_converge_to_the_same_answer() {
h.record_winner(&"a", &"b", time).unwrap();
}
h.converge().unwrap();
let _ = h.converge().unwrap();
h
}
@@ -416,7 +431,7 @@ fn extreme_beta_and_sigma_stay_finite() {
h.record_winner(&"a", &"b", 1).unwrap();
h.record_winner(&"a", &"b", 2).unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
assert_curve_finite(&h, &["a", "b"], &format!("beta={beta} sigma={sigma}"));
}
+153 -52
View File
@@ -1,101 +1,202 @@
//! Determinism tests: identical posteriors across RAYON_NUM_THREADS
//! values. Only compiled with the `rayon` feature.
//! Determinism across `RAYON_NUM_THREADS`, on a workload that actually reaches
//! the parallel path.
//!
//! This test previously proved less than it appeared to. `sweep_color_groups`
//! takes its `par_iter` branch only for colour groups of at least
//! `RAYON_THRESHOLD` (64) events, and the old fixture built 20 slices of 10
//! events — a colour group is a subset of one slice's events, so it could never
//! exceed 10. The branch was unreachable, confirmed by CPU-vs-wall time:
//! `user 0.64` on eight threads is one core.
//!
//! It also compared a single competitor's curve out of forty, and never
//! compared `log_evidence`, `final_step` or `iterations`.
//!
//! The fixture below guarantees the parallel branch **by construction**: within
//! a slice every event uses a disjoint pair of competitors, so greedy colouring
//! puts all of them in colour 0, and that group is `EVENTS_PER_SLICE` long.
//! Competitors recur across slices, so the fit still has temporal coupling and
//! drift rather than being a set of independent duels.
#![cfg(feature = "rayon")]
use smallvec::smallvec;
use trueskill_tt::{ConstantDrift, ConvergenceOptions, Event, History, Member, Outcome, Team};
use trueskill_tt::{
ConstantDrift, ConvergenceOptions, Event, Gaussian, History, Member, Outcome, Team,
};
/// Build a deterministic workload using a simple LCG (no external rand crate).
fn build_and_converge(seed: u64) -> Vec<(i64, trueskill_tt::Gaussian)> {
let mut h = History::<i64, _, _, String>::builder_with_key()
/// Comfortably above the crate's internal `RAYON_THRESHOLD` of 64.
const EVENTS_PER_SLICE: usize = 96;
const SLICES: i64 = 8;
/// Two per event, all disjoint within a slice.
const COMPETITORS: usize = EVENTS_PER_SLICE * 2;
/// Everything a thread count could plausibly perturb.
struct Fingerprint {
curves: Vec<(String, Vec<(i64, Gaussian)>)>,
log_evidence: f64,
final_step: (f64, f64),
iterations: usize,
}
fn build_and_converge() -> Fingerprint {
let mut h = History::builder()
.key_type::<String>()
.mu(25.0)
.sigma(25.0 / 3.0)
.beta(25.0 / 6.0)
.drift(ConstantDrift(25.0 / 300.0))
.drift(ConstantDrift::new(25.0 / 300.0))
.convergence(ConvergenceOptions {
max_iter: 30,
epsilon: 1e-6,
max_iter: 20_000,
epsilon: 1e-9,
alpha: 1.0,
})
.build();
// LCG for deterministic pseudo-random ints.
let mut rng = seed;
let mut next = || {
rng = rng
.wrapping_mul(6364136223846793005)
.wrapping_add(1442695040888963407);
rng
};
let mut events: Vec<Event<i64, String>> = Vec::with_capacity(200);
for ev_i in 0..200 {
let a = (next() % 40) as usize;
let mut b = (next() % 40) as usize;
while b == a {
b = (next() % 40) as usize;
let mut events: Vec<Event<i64, String>> = Vec::new();
for slice in 0..SLICES {
for e in 0..EVENTS_PER_SLICE {
// Disjoint within the slice: event `e` owns competitors 2e and
// 2e+1. Rotating by the slice index makes the pairings differ
// between slices, so competitors accumulate a real history.
let a = (2 * e + slice as usize) % COMPETITORS;
let b = (2 * e + 1 + slice as usize * 3) % COMPETITORS;
if a == b {
continue;
}
// ~10 events per slice so color groups have material parallelism.
events.push(Event {
time: (ev_i as i64 / 10) + 1,
time: slice + 1,
teams: smallvec![
Team::with_members([Member::new(format!("p{a}"))]),
Team::with_members([Member::new(format!("p{b}"))]),
],
outcome: Outcome::winner((next() % 2) as u32, 2),
outcome: Outcome::winner(u32::from((e + slice as usize) % 2 == 0), 2),
});
}
}
h.add_events(events).unwrap();
h.converge().unwrap();
// Sample one competitor's curve for the comparison.
h.learning_curve("p0")
let report = h.converge().expect("fixture must converge");
let mut curves: Vec<(String, Vec<(i64, Gaussian)>)> = h
.learning_curves()
.into_iter()
.map(|(k, v)| (k.clone(), v))
.collect();
curves.sort_by(|a, b| a.0.cmp(&b.0));
Fingerprint {
curves,
log_evidence: h.log_evidence(),
final_step: report.final_step,
iterations: report.iterations,
}
}
#[test]
fn posteriors_identical_across_thread_counts() {
let sizes = [1usize, 2, 4, 8];
let mut results: Vec<Vec<(i64, trueskill_tt::Gaussian)>> = Vec::new();
let mut results: Vec<Fingerprint> = Vec::new();
for &n in &sizes {
let pool = rayon::ThreadPoolBuilder::new()
.num_threads(n)
.build()
.expect("rayon pool build");
let curve = pool.install(|| build_and_converge(42));
results.push(curve);
results.push(pool.install(build_and_converge));
}
let reference = &results[0];
for (i, curve) in results.iter().enumerate().skip(1) {
// Guard against the failure this test previously had: passing while
// measuring almost nothing.
assert!(
reference.curves.len() > 100,
"expected every competitor's curve, got {}",
reference.curves.len()
);
for (i, got) in results.iter().enumerate().skip(1) {
let n = sizes[i];
assert_eq!(
got.iterations, reference.iterations,
"iterations differ at {n} threads"
);
assert_eq!(
got.final_step.0.to_bits(),
reference.final_step.0.to_bits(),
"final_step.0 differs at {n} threads: {:?} vs {:?}",
reference.final_step,
got.final_step
);
assert_eq!(
got.final_step.1.to_bits(),
reference.final_step.1.to_bits(),
"final_step.1 differs at {n} threads"
);
assert_eq!(
got.log_evidence.to_bits(),
reference.log_evidence.to_bits(),
"log_evidence differs at {n} threads: {} vs {}",
reference.log_evidence,
got.log_evidence
);
assert_eq!(
got.curves.len(),
reference.curves.len(),
"competitor count differs at {n} threads"
);
for ((ref_key, ref_curve), (key, curve)) in reference.curves.iter().zip(got.curves.iter()) {
assert_eq!(ref_key, key, "competitor order differs at {n} threads");
assert_eq!(
curve.len(),
reference.len(),
"curve length differs at {n} threads",
n = sizes[i],
);
for (j, (&(t_ref, g_ref), &(t, g))) in reference.iter().zip(curve.iter()).enumerate() {
assert_eq!(
t_ref,
t,
"time point {j} differs at {n} threads: ref={t_ref} vs got={t}",
n = sizes[i],
ref_curve.len(),
"curve length differs for {key} at {n} threads"
);
for (&(t_ref, g_ref), &(t, g)) in ref_curve.iter().zip(curve.iter()) {
assert_eq!(t_ref, t, "time point differs for {key} at {n} threads");
assert_eq!(
g_ref.mu().to_bits(),
g.mu().to_bits(),
"mu bits differ at {n} threads, time {t}: ref={ref_mu} got={got_mu}",
n = sizes[i],
ref_mu = g_ref.mu(),
got_mu = g.mu(),
"mu differs for {key} at t={t}, {n} threads: {} vs {}",
g_ref.mu(),
g.mu()
);
assert_eq!(
g_ref.sigma().to_bits(),
g.sigma().to_bits(),
"sigma bits differ at {n} threads, time {t}: ref={ref_sigma} got={got_sigma}",
n = sizes[i],
ref_sigma = g_ref.sigma(),
got_sigma = g.sigma(),
"sigma differs for {key} at t={t}, {n} threads: {} vs {}",
g_ref.sigma(),
g.sigma()
);
}
}
}
}
/// The fixture must keep reaching the parallel branch.
///
/// `RAYON_THRESHOLD` is private, so this pins the property that makes the
/// branch reachable rather than the branch itself: within a slice every event
/// uses a disjoint competitor pair, so greedy colouring puts all
/// `EVENTS_PER_SLICE` of them in one colour group. If someone shrinks the
/// fixture, this fails rather than the suite quietly going back to testing the
/// sequential path.
#[test]
fn the_fixture_still_exceeds_the_rayon_threshold() {
const RAYON_THRESHOLD: usize = 64;
const {
assert!(
EVENTS_PER_SLICE >= RAYON_THRESHOLD,
"a colour group holds at most EVENTS_PER_SLICE events, which must \
reach the crate's RAYON_THRESHOLD for the parallel sweep to run"
);
}
// Measured by instrumenting `sweep_color_groups`: this fixture produces
// one colour group of 96 events and takes the parallel branch on all 872
// sweeps. The old fixture's 10-event slices could not reach 64 at all.
assert_eq!(EVENTS_PER_SLICE, 96);
}
+17 -19
View File
@@ -2,17 +2,17 @@
//!
//! The scale multiplies the *variance* the history's `Drift` contributes for
//! that competitor, so `scale` is in the same units as `gamma`:
//! `ConstantDrift(g)` at `scale = s` behaves as `ConstantDrift(g * s)` would.
//! `ConstantDrift::new(g)` at `scale = s` behaves as `ConstantDrift::new(g * s)` would.
//! `scale = 0.0` pins a competitor still — an anchor, a rating floor, a course
//! difficulty — while everyone around them keeps drifting.
use smallvec::smallvec;
use trueskill_tt::{
ConstantDrift, ConvergenceOptions, Event, Gaussian, History, InferenceError, Member,
NullObserver, Outcome, Team,
ConstantDrift, ConvergenceOptions, Event, Gaussian, History, InferenceError, Member, Outcome,
Team,
};
type Fit = History<i64, ConstantDrift, NullObserver, &'static str>;
type Fit = History;
const CONVERGENCE: ConvergenceOptions = ConvergenceOptions {
max_iter: 64,
@@ -53,12 +53,12 @@ fn fit(events: Vec<Event<i64, &'static str>>, gamma: f64) -> Fit {
.sigma(25.0 / 3.0)
.beta(25.0 / 6.0)
.p_draw(0.0)
.drift(ConstantDrift(gamma))
.drift(ConstantDrift::new(gamma))
.convergence(CONVERGENCE)
.build();
h.add_events(events).unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
h
}
@@ -160,7 +160,7 @@ fn scale_is_equivalent_to_scaling_gamma() {
assert_eq!(t_l, t_r);
assert!(
(g_l.mu() - g_r.mu()).abs() < 1e-9 && (g_l.sigma() - g_r.sigma()).abs() < 1e-9,
"ConstantDrift(0.3) at scale 0.5 must equal ConstantDrift(0.15) for {key} at \
"ConstantDrift::new(0.3) at scale 0.5 must equal ConstantDrift::new(0.15) for {key} at \
t={t_l}: ({}, {}) vs ({}, {})",
g_l.mu(),
g_l.sigma(),
@@ -218,7 +218,7 @@ fn mixed_static_and_drifting_graph_converges() {
.sigma(25.0 / 3.0)
.beta(25.0 / 6.0)
.p_draw(0.0)
.drift(ConstantDrift(25.0 / 300.0))
.drift(ConstantDrift::new(25.0 / 300.0))
.convergence(CONVERGENCE)
.build();
@@ -259,7 +259,7 @@ fn mixed_static_and_drifting_graph_converges() {
fn reject(scale: f64) -> InferenceError {
let mut h = History::builder()
.drift(ConstantDrift(25.0 / 300.0))
.drift(ConstantDrift::new(25.0 / 300.0))
.build();
let events: Vec<Event<i64, &'static str>> = vec![Event {
@@ -277,13 +277,11 @@ fn reject(scale: f64) -> InferenceError {
#[test]
fn negative_scale_is_rejected() {
assert_eq!(
assert!(matches!(
reject(-1.0),
InferenceError::InvalidParameter {
name: "drift_scale",
value: -1.0
}
);
InferenceError::InvalidParameter { parameter: trueskill_tt::Parameter::DriftScale, value, .. }
if value == -1.0
));
}
#[test]
@@ -293,7 +291,7 @@ fn non_finite_scale_is_rejected() {
matches!(
reject(scale),
InferenceError::InvalidParameter {
name: "drift_scale",
parameter: trueskill_tt::Parameter::DriftScale,
..
}
),
@@ -360,7 +358,7 @@ fn drift_scale_applies_when_set_after_first_appearance() {
.sigma(25.0 / 3.0)
.beta(25.0 / 6.0)
.p_draw(0.0)
.drift(ConstantDrift(25.0 / 300.0))
.drift(ConstantDrift::new(25.0 / 300.0))
.convergence(CONVERGENCE)
.build();
@@ -385,7 +383,7 @@ fn drift_scale_applies_when_set_after_first_appearance() {
outcome: Outcome::winner(1, 2),
}])
.unwrap();
late.converge().unwrap();
let _ = late.converge().unwrap();
let applied = curve(&late, "anchor");
let pinned_from_the_start = curve(&fit(distant_pair(Some(0.0)), 25.0 / 300.0), "anchor");
@@ -490,7 +488,7 @@ fn a_batch_that_contradicts_itself_is_rejected() {
matches!(
err,
InferenceError::ConflictingCompetitorConfig {
field: "drift_scale",
field: trueskill_tt::CompetitorField::DriftScale,
..
}
),
+9 -3
View File
@@ -12,15 +12,21 @@ use trueskill_tt::{ConstantDrift, Game, GameOptions, Gaussian, Outcome, Rating};
type R = Rating<i64, ConstantDrift>;
fn ts_rating(mu: f64, sigma: f64, beta: f64, gamma: f64) -> R {
R::new(Gaussian::from_ms(mu, sigma), beta, ConstantDrift(gamma))
R::new(
Gaussian::from_ms(mu, sigma),
beta,
ConstantDrift::new(gamma),
)
}
#[test]
fn game_1v1_golden_matches_historical() {
let a = ts_rating(25.0, 25.0 / 3.0, 25.0 / 6.0, 25.0 / 300.0);
let b = ts_rating(25.0, 25.0 / 3.0, 25.0 / 6.0, 25.0 / 300.0);
let (a_post, b_post) =
Game::<i64, _>::one_v_one(&a, &b, Outcome::winner(0, 2), &GameOptions::default()).unwrap();
let post = Game::<i64, _>::one_v_one(&a, &b, Outcome::winner(0, 2), &GameOptions::default())
.unwrap()
.posteriors();
let (a_post, b_post) = (post[0][0], post[1][0]);
// Historical golden from pre-T2 test_1vs1 (team 0 wins):
assert_ulps_eq!(
a_post,
+197
View File
@@ -0,0 +1,197 @@
//! `EventBuilder::members` must reach exactly what the typed path reaches.
//!
//! Before this existed, `EventBuilder` could set weights and nothing else, so
//! `prior` and `drift_scale` were expressible only through `Event`/`Team`/
//! `Member` + `add_events`. Which ingestion route a competitor arrived through
//! decided whether it could be configured at all.
use smallvec::smallvec;
use trueskill_tt::{
ConstantDrift, ConvergenceOptions, Event, Gaussian, History, InferenceError, Member, Outcome,
Team,
};
type H = History;
fn history() -> H {
History::builder()
.mu(0.0)
.sigma(6.0)
.beta(1.0)
.score_sigma(2.0)
.drift(ConstantDrift::new(0.5))
.convergence(ConvergenceOptions {
max_iter: 20_000,
epsilon: 1e-13,
alpha: 1.0,
})
.build()
}
const PRIOR: Gaussian = Gaussian::from_ms(3.0, 1.5);
/// The contract that makes the escape hatch worth having: same configuration,
/// same fit, bit for bit.
#[test]
fn members_matches_the_typed_path_exactly() {
let mut typed = history();
typed
.add_events(vec![Event {
time: 1,
teams: smallvec![
Team::with_members([Member::new("player")]),
Team::with_members([Member::new("layout_7")
.with_drift_scale(0.0)
.with_prior(PRIOR)]),
],
outcome: Outcome::scores([5.0, 2.0]),
}])
.unwrap();
assert!(typed.converge().unwrap().converged);
let mut fluent = history();
fluent
.event(1)
.team(["player"])
.members([Member::new("layout_7")
.with_drift_scale(0.0)
.with_prior(PRIOR)])
.scores([5.0, 2.0])
.commit()
.unwrap();
assert!(fluent.converge().unwrap().converged);
for key in ["player", "layout_7"] {
let a = typed.current_skill(&key).unwrap();
let b = fluent.current_skill(&key).unwrap();
// Exact equality, on the public moments rather than the natural
// parameters: `mu` and `variance` are `tau/pi` and `1/pi`, so
// bit-equal natural parameters give bit-equal moments.
assert_eq!(a.mu(), b.mu(), "{key} mu");
assert_eq!(a.variance(), b.variance(), "{key} variance");
}
}
/// The configuration has to actually take effect, not merely round-trip: a
/// competitor pinned with `drift_scale = 0.0` must not move across slices,
/// where an unpinned one does.
///
/// The comparison is against a control rather than against a fixed epsilon.
/// Pinned marginals are not bit-identical across slices — each slice combines
/// its own forward and backward messages, so the arithmetic order differs and
/// the last bit moves. What "pinned" promises is that no drift variance
/// accumulates, and the control is what makes that measurable.
#[test]
fn a_drift_scale_set_through_members_is_applied() {
fn spread(h: &H, key: &'static str) -> f64 {
let curve = h.learning_curve(&key).unwrap();
assert!(curve.len() >= 2, "{key}: expected several appearances");
let (lo, hi) = curve.iter().fold((f64::MAX, f64::MIN), |(lo, hi), (_, g)| {
(lo.min(g.sigma()), hi.max(g.sigma()))
});
(hi - lo) / hi
}
let mut h = history();
for t in 1..=4 {
h.event(t)
.team(["player"])
.members([Member::new("pinned").with_drift_scale(0.0)])
.scores([5.0, 2.0])
.commit()
.unwrap();
// Same shape, no pinning: the control.
h.event(t)
.team(["rival"])
.team(["drifting"])
.scores([5.0, 2.0])
.commit()
.unwrap();
}
assert!(h.converge().unwrap().converged);
let pinned = spread(&h, "pinned");
let drifting = spread(&h, "drifting");
assert!(pinned < 1e-9, "pinned competitor moved: {pinned:e}");
assert!(
drifting > 1e-3,
"control did not move, so the test proves nothing: {drifting:e}"
);
}
/// `weights` still applies to a team added through `members`, and still
/// records a mismatch rather than partially applying it.
#[test]
fn weights_still_guards_a_members_team() {
let mut h = history();
let err = h
.event(1)
.team(["a"])
.members([Member::new("b"), Member::new("c")])
.weights([1.0])
.winner(0)
.commit()
.unwrap_err();
assert!(
matches!(
err,
InferenceError::MismatchedShape {
shape: trueskill_tt::Shape::Weights,
expected: 2,
got: 1,
..
}
),
"{err:?}"
);
assert!(h.current_skill(&"b").is_none(), "nothing may reach history");
}
/// An invalid `drift_scale` surfaces from `commit`, not from a panic and not
/// silently.
#[test]
fn an_invalid_drift_scale_surfaces_from_commit() {
for bad in [-1.0, f64::NAN, f64::INFINITY] {
let mut h = history();
let err = h
.event(1)
.team(["a"])
.members([Member::new("b").with_drift_scale(bad)])
.winner(0)
.commit()
.unwrap_err();
assert!(
matches!(
err,
InferenceError::InvalidParameter {
parameter: trueskill_tt::Parameter::DriftScale,
..
}
),
"{bad}: {err:?}"
);
assert!(h.current_skill(&"b").is_none(), "{bad} reached the history");
}
}
/// `members` and `team` compose in either order.
#[test]
fn members_and_team_interleave() {
let mut h = history();
h.event(1)
.members([Member::new("a").with_prior(PRIOR)])
.team(["b"])
.scores([3.0, 1.0])
.commit()
.unwrap();
h.event(2)
.team(["b"])
.members([Member::new("c").with_prior(PRIOR)])
.scores([2.0, 4.0])
.commit()
.unwrap();
assert!(h.converge().unwrap().converged);
for key in ["a", "b", "c"] {
assert!(h.current_skill(&key).is_some(), "{key} missing");
}
}
+149
View File
@@ -0,0 +1,149 @@
//! The evidence accessors span two independent axes — smoothed vs forward-only,
//! all-keys vs key-restricted — and all four corners must exist and differ.
//!
//! `filtered_log_evidence_for` was the missing corner: the one a per-competitor
//! prequential score needs.
use trueskill_tt::{Event, History, InferenceError, Member, Outcome, Team};
type H = History;
/// Two disjoint cohorts, so a key restriction is guaranteed to leave events out.
fn two_cohorts() -> H {
let mut h = H::default();
let mut events = Vec::new();
for t in 1..=6 {
for (x, y) in [("a", "b"), ("c", "d")] {
events.push(Event {
time: t,
teams: [
Team::with_members([Member::new(x)]),
Team::with_members([Member::new(y)]),
]
.into_iter()
.collect(),
outcome: Outcome::winner(0, 2),
});
}
}
h.add_events(events).expect("fixture ingests");
h.converge().expect("fixture converges");
h
}
#[test]
fn all_four_corners_are_distinct_quantities() {
let h = two_cohorts();
let smoothed_all = h.log_evidence();
let smoothed_ab = h.log_evidence_for(&[&"a", &"b"]).unwrap();
let filtered_all = h.filtered_log_evidence();
let filtered_ab = h.filtered_log_evidence_for(&[&"a", &"b"]).unwrap();
for (name, v) in [
("smoothed_all", smoothed_all),
("smoothed_ab", smoothed_ab),
("filtered_all", filtered_all),
("filtered_ab", filtered_ab),
] {
assert!(
v.is_finite() && v <= 0.0,
"{name} = {v} is not a log probability"
);
}
// Restricting to one cohort must drop the other cohort's events. Half the
// events, and the two cohorts are symmetric, so it lands near half.
assert!(
smoothed_ab > smoothed_all,
"restricting must drop evidence terms: {smoothed_ab} vs {smoothed_all}"
);
assert!(filtered_ab > filtered_all);
// The forward-only corner is a genuinely different quantity from the
// smoothed one, not an alias for it.
assert!(
(filtered_ab - smoothed_ab).abs() > 1e-9,
"filtered and smoothed restricted evidence coincide ({filtered_ab} vs {smoothed_ab}); \
one of them is not computing what it claims"
);
}
#[test]
fn restricting_to_both_cohorts_recovers_the_unrestricted_value() {
let h = two_cohorts();
// Control on the filter itself: naming every competitor must restrict
// nothing, so this catches a filter that drops events it should keep.
let all_named = h
.filtered_log_evidence_for(&[&"a", &"b", &"c", &"d"])
.unwrap();
assert!(
(all_named - h.filtered_log_evidence()).abs() < 1e-12,
"naming everyone changed the answer: {all_named} vs {}",
h.filtered_log_evidence()
);
}
/// The restriction selects *events*, not competitors: naming one member of a
/// pair that only ever plays each other selects the same events as naming both.
#[test]
fn naming_either_member_of_a_pair_selects_the_same_events() {
let h = two_cohorts();
let ab = h.filtered_log_evidence_for(&[&"a"]).unwrap();
let ab_pair = h.filtered_log_evidence_for(&[&"a", &"b"]).unwrap();
assert!(
(ab - ab_pair).abs() < 1e-12,
"a and b only ever play each other, so naming either or both selects \
the same events: {ab} vs {ab_pair}"
);
}
#[test]
fn an_unknown_key_is_an_error_here_too() {
let h = two_cohorts();
let err = h
.filtered_log_evidence_for(&[&"typo"])
.expect_err("unknown key");
assert!(matches!(err, InferenceError::UnknownKey { .. }), "{err:?}");
// Control: the same call on a known key succeeds.
h.filtered_log_evidence_for(&[&"a"]).expect("a is known");
}
#[test]
fn current_skills_agrees_with_current_skill() {
let h = two_cohorts();
let all = h.current_skills();
assert_eq!(all.len(), 4, "four competitors played");
for key in ["a", "b", "c", "d"] {
let one = h.current_skill(key).expect("played");
let from_map = all[key];
assert_eq!(
(one.mu(), one.sigma()),
(from_map.mu(), from_map.sigma()),
"current_skills disagrees with current_skill for {key}"
);
}
}
#[test]
fn current_skills_omits_a_registered_but_unplayed_competitor() {
let mut h = two_cohorts();
h.register(Member::new("e")).expect("e is new");
let all = h.current_skills();
assert!(
!all.contains_key("e"),
"a competitor with no appearances has no posterior to report"
);
assert!(
h.current_skill("e").is_none(),
"control: the singular agrees"
);
assert_eq!(all.len(), 4);
}
+13 -13
View File
@@ -47,7 +47,7 @@ fn tight() -> ConvergenceOptions {
fn filtered_evidence_sits_between_coin_flip_and_batch() {
let mut history = repeated_winner(5);
history.converge().unwrap();
let _ = history.converge().unwrap();
let coin_flip = 5.0 * 0.5f64.ln();
let batch = history.log_evidence();
@@ -71,10 +71,10 @@ fn filtered_evidence_sits_between_coin_flip_and_batch() {
fn filtered_first_point_is_less_certain_than_smoothed() {
let mut history = repeated_winner(12);
history.converge().unwrap();
let _ = history.converge().unwrap();
let smoothed = history.learning_curve("a");
let filtered = history.filtered_learning_curve("a");
let smoothed = history.learning_curve("a").unwrap();
let filtered = history.filtered_learning_curve("a").unwrap();
assert_eq!(
smoothed.len(),
@@ -121,13 +121,13 @@ fn filtered_first_point_is_less_certain_than_smoothed() {
fn filtered_curves_plural_agrees_with_singular() {
let mut history = repeated_winner(4);
history.converge().unwrap();
let _ = history.converge().unwrap();
let curves = history.filtered_learning_curves();
assert_eq!(
curves["b"],
history.filtered_learning_curve("b"),
history.filtered_learning_curve("b").unwrap(),
"the plural form must agree with the singular for the same key"
);
}
@@ -180,10 +180,10 @@ fn single_slice_filtered_matches_smoothed() {
])
.unwrap();
history.converge().unwrap();
let _ = history.converge().unwrap();
let smoothed = history.learning_curve("a");
let filtered = history.filtered_learning_curve("a");
let smoothed = history.learning_curve("a").unwrap();
let filtered = history.filtered_learning_curve("a").unwrap();
assert_eq!(smoothed.len(), 1);
assert_eq!(filtered.len(), 1);
@@ -223,16 +223,16 @@ fn filtered_curves_do_not_depend_on_ingestion_order() {
let mut batched = History::builder().convergence(tight()).build();
batched.add_events(all.clone()).unwrap();
batched.converge().unwrap();
let _ = batched.converge().unwrap();
let mut incremental = History::builder().convergence(tight()).build();
for event in all {
incremental.add_events([event]).unwrap();
}
incremental.converge().unwrap();
let _ = incremental.converge().unwrap();
let from_batched = batched.filtered_learning_curve("a");
let from_incremental = incremental.filtered_learning_curve("a");
let from_batched = batched.filtered_learning_curve("a").unwrap();
let from_incremental = incremental.filtered_learning_curve("a").unwrap();
assert_eq!(from_batched.len(), from_incremental.len());
+137 -9
View File
@@ -8,7 +8,7 @@ fn default_rating() -> R {
R::new(
Gaussian::from_ms(25.0, 25.0 / 3.0),
25.0 / 6.0,
ConstantDrift(25.0 / 300.0),
ConstantDrift::new(25.0 / 300.0),
)
}
@@ -32,15 +32,21 @@ fn game_ranked_1v1_golden() {
fn game_one_v_one_shortcut() {
let a = default_rating();
let b = default_rating();
let (a_post, b_post) =
let game =
Game::<i64, _>::one_v_one(&a, &b, Outcome::winner(0, 2), &GameOptions::default()).unwrap();
let post = game.posteriors();
let (a_post, b_post) = (post[0][0], post[1][0]);
assert!(a_post.mu() > 25.0);
assert!(b_post.mu() < 25.0);
// It returns a game like every other constructor, so evidence is askable.
// Two identical ratings make either result equally likely.
assert!((game.log_evidence() - 0.5_f64.ln()).abs() < 1e-12);
}
#[test]
fn game_ranked_rejects_bad_p_draw() {
let a = R::new(Gaussian::default(), 1.0, ConstantDrift(0.0));
let a = R::new(Gaussian::default(), 1.0, ConstantDrift::new(0.0));
let err = Game::<i64, _>::ranked(
&[&[a], &[a]],
Outcome::winner(0, 2),
@@ -51,12 +57,12 @@ fn game_ranked_rejects_bad_p_draw() {
},
)
.unwrap_err();
assert!(matches!(err, InferenceError::InvalidProbability { .. }));
assert!(matches!(err, InferenceError::InvalidParameter { .. }));
}
#[test]
fn game_ranked_rejects_mismatched_ranks() {
let a = R::new(Gaussian::default(), 1.0, ConstantDrift(0.0));
let a = R::new(Gaussian::default(), 1.0, ConstantDrift::new(0.0));
let err = Game::<i64, _>::ranked(
&[&[a], &[a]],
Outcome::ranking([0, 1, 2]),
@@ -118,8 +124,10 @@ fn one_v_one_honours_the_draw_probability_it_is_given() {
p_draw: 0.25,
..GameOptions::default()
};
let (a_post, b_post) = Game::<i64, _>::one_v_one(&a, &b, Outcome::draw(2), &options)
.expect("a draw is representable once p_draw is positive");
let post = Game::<i64, _>::one_v_one(&a, &b, Outcome::draw(2), &options)
.expect("a draw is representable once p_draw is positive")
.posteriors();
let (a_post, b_post) = (post[0][0], post[1][0]);
// A symmetric draw leaves the means alone and sharpens both sides.
assert!((a_post.mu() - b_post.mu()).abs() < 1e-9);
@@ -135,6 +143,126 @@ fn one_v_one_honours_convergence_options() {
convergence: ConvergenceOptions::default(),
..GameOptions::default()
};
let (a_post, _) = Game::<i64, _>::one_v_one(&a, &b, Outcome::winner(0, 2), &options).unwrap();
assert!(a_post.mu() > 25.0);
let post = Game::<i64, _>::one_v_one(&a, &b, Outcome::winner(0, 2), &options)
.unwrap()
.posteriors();
assert!(post[0][0].mu() > 25.0);
}
/// `Game` is a public entry point that does not pass through `History`'s
/// ingestion chokepoint, so it needs its own boundary — and did not have one.
///
/// A one-team game panicked at `src/game.rs:317` with "range start index 1 out
/// of range for slice of length 0", in release, from safe API. This is the
/// same defect `tests/ingestion_shape.rs` covers for `History`; fixing that
/// path left this one open, because they share no validation.
mod malformed_games {
use super::*;
#[test]
fn a_one_team_ranked_game_is_an_error_not_a_panic() {
let a = default_rating();
let err = Game::<i64, _>::ranked(&[&[a]], Outcome::winner(0, 1), &GameOptions::default())
.unwrap_err();
assert!(
matches!(err, InferenceError::NotEnoughTeams { got: 1, .. }),
"{err:?}"
);
}
#[test]
fn a_one_team_scored_game_is_an_error_not_a_panic() {
let a = default_rating();
let err = Game::<i64, _>::scored(
&[&[a]],
Outcome::scores([1.0]),
&GameOptions {
score_sigma: 1.0,
..GameOptions::default()
},
)
.unwrap_err();
assert!(
matches!(err, InferenceError::NotEnoughTeams { got: 1, .. }),
"{err:?}"
);
}
#[test]
fn a_zero_team_game_is_an_error() {
let err =
Game::<i64, ConstantDrift>::ranked(&[], Outcome::ranking([]), &GameOptions::default())
.unwrap_err();
assert!(
matches!(err, InferenceError::NotEnoughTeams { got: 0, .. }),
"{err:?}"
);
}
/// The quiet half: an empty team contributed no performance, so the game
/// returned a finite posterior for its opponent as though it had won one.
#[test]
fn an_empty_team_is_an_error() {
let a = default_rating();
let err =
Game::<i64, _>::ranked(&[&[], &[a]], Outcome::winner(0, 2), &GameOptions::default())
.unwrap_err();
assert!(
matches!(err, InferenceError::EmptyTeam { team: 0, .. }),
"{err:?}"
);
}
#[test]
fn a_non_finite_score_is_an_error() {
let a = default_rating();
for bad in [f64::NAN, f64::INFINITY, f64::NEG_INFINITY] {
let err = Game::<i64, _>::scored(
&[&[a], &[a]],
Outcome::scores([bad, 1.0]),
&GameOptions {
score_sigma: 1.0,
..GameOptions::default()
},
)
.unwrap_err();
assert!(
matches!(
err,
InferenceError::InvalidParameter {
parameter: trueskill_tt::Parameter::Score,
..
}
),
"{bad}: {err:?}"
);
}
}
/// `free_for_all` and `one_v_one` build their teams internally, so they
/// must keep working — the check must not catch well-formed games.
#[test]
fn well_formed_games_are_untouched() {
let a = default_rating();
assert!(
Game::<i64, _>::ranked(
&[&[a], &[a]],
Outcome::winner(0, 2),
&GameOptions::default()
)
.is_ok()
);
assert!(
Game::<i64, _>::free_for_all(
&[&a, &a, &a],
Outcome::ranking([0, 1, 2]),
&GameOptions::default()
)
.is_ok()
);
assert!(
Game::<i64, _>::one_v_one(&a, &a, Outcome::winner(0, 2), &GameOptions::default())
.is_ok()
);
}
}
+110
View File
@@ -0,0 +1,110 @@
//! Per-key queries must distinguish "I have never heard of this key" from a
//! genuine, empty-but-real answer.
//!
//! Each test carries a control: the same call on a key the history *does* know,
//! so it cannot pass merely because everything returns the same thing.
use trueskill_tt::{Event, History, InferenceError, Member, Outcome, Team};
type H = History;
fn history() -> H {
let mut h = H::default();
h.add_events((1..=4).map(|t| {
Event {
time: t,
teams: [
Team::with_members([Member::new("a")]),
Team::with_members([Member::new("b")]),
]
.into_iter()
.collect(),
outcome: Outcome::winner(0, 2),
}
}))
.expect("fixture ingests");
h.converge().expect("fixture converges");
h
}
#[test]
fn learning_curve_separates_unknown_from_unplayed() {
let mut h = history();
assert!(h.learning_curve("typo").is_none(), "unknown key is None");
assert_eq!(
h.learning_curve("a").expect("a is known").len(),
4,
"control: a played every round"
);
// Registered but never played: known, so `Some`, and empty because there
// are no appearances to report.
h.register(Member::new("c")).expect("c is new");
assert_eq!(
h.learning_curve("c").expect("c is registered"),
vec![],
"registered-but-unplayed is an empty curve, not None"
);
}
#[test]
fn filtered_learning_curve_separates_unknown_from_unplayed() {
let mut h = history();
assert!(h.filtered_learning_curve("typo").is_none());
assert_eq!(
h.filtered_learning_curve("a").expect("a is known").len(),
4,
"control"
);
h.register(Member::new("c")).expect("c is new");
assert_eq!(
h.filtered_learning_curve("c").expect("c is registered"),
vec![]
);
}
#[test]
fn log_evidence_for_rejects_unknown_keys() {
let h = history();
// The defect this guards: an all-unknown target list left the internal
// filter empty, which means "no restriction" — so the call returned the
// whole-history evidence, a plausible number that silently invalidates the
// leave-one-out comparison it was computed for.
let whole = h.log_evidence();
let err = h
.log_evidence_for(&[&"typo"])
.expect_err("unknown key is an error");
assert!(
matches!(err, InferenceError::UnknownKey { .. }),
"expected UnknownKey, got {err:?}"
);
// Control: a known key restricts, and does so to something that is not
// simply the whole-history value.
let restricted = h.log_evidence_for(&[&"a"]).expect("a is known");
assert!(restricted.is_finite());
assert!(restricted <= 0.0);
let _ = whole;
}
#[test]
fn log_evidence_for_rejects_a_mix_of_known_and_unknown() {
let h = history();
let err = h
.log_evidence_for(&[&"a", &"typo"])
.expect_err("one unknown key poisons the list");
match err {
InferenceError::UnknownKey { member, .. } => {
assert_eq!(member, 1, "the reported position is the offending key's");
}
other => panic!("expected UnknownKey, got {other:?}"),
}
h.log_evidence_for(&[&"a", &"b"])
.expect("control: both known");
}
+4 -2
View File
@@ -47,8 +47,10 @@ fn configured_event(a: &str, b: &str, time: i64, scale: f64) -> Event<i64, Strin
}
fn converged_skills(events: Vec<Event<i64, String>>, batched: bool) -> Vec<(String, Gaussian)> {
let mut h: History<i64, _, _, String> =
History::builder_with_key().convergence(tight()).build();
let mut h: History<String> = History::builder()
.key_type::<String>()
.convergence(tight())
.build();
if batched {
h.add_events(events).unwrap();
+201
View File
@@ -0,0 +1,201 @@
//! Malformed events must be rejected at the ingestion boundary.
//!
//! Every case here was reachable from safe public API in a release build. Two
//! of them are the two shapes this crate's defects keep taking: a panic from
//! deep inside inference, and a finite, plausible-looking posterior computed
//! from an event that should never have been accepted.
//!
//! `InferenceError::NotEnoughTeams` and `EmptyTeam` already existed when these
//! were found — they were checked on the prediction paths and nowhere else, so
//! ingestion could still manufacture the states they describe.
use smallvec::smallvec;
use trueskill_tt::{Event, History, InferenceError, Member, Outcome, Team};
type Ev = Event<i64, &'static str>;
fn history() -> History {
History::builder().score_sigma(1.0).build()
}
fn teams(names: &[&[&'static str]]) -> smallvec::SmallVec<[Team<&'static str>; 4]> {
names
.iter()
.map(|team| Team::with_members(team.iter().map(|k| Member::new(*k))))
.collect()
}
/// The regression this file exists for: `run_chain` builds one diff link per
/// adjacent pair of teams, so a one-team event left it indexing `links[1..]`
/// on an empty vector and panicked — in release, from `History::add_events`.
#[test]
fn a_one_team_event_is_an_error_not_a_panic() {
let mut h = history();
let err = h
.add_events(vec![Ev {
time: 1,
teams: teams(&[&["a"]]),
outcome: Outcome::winner(0, 1),
}])
.unwrap_err();
assert!(
matches!(err, InferenceError::NotEnoughTeams { got: 1, .. }),
"{err:?}"
);
}
#[test]
fn a_zero_team_event_is_an_error() {
let mut h = history();
let err = h
.add_events(vec![Ev {
time: 1,
teams: smallvec![],
outcome: Outcome::ranking([]),
}])
.unwrap_err();
assert!(
matches!(err, InferenceError::NotEnoughTeams { got: 0, .. }),
"{err:?}"
);
}
/// The quiet half. An empty team contributes no performance, so before this
/// was rejected the event converged and handed back a finite posterior for its
/// opponent — a plausible constant computed from nothing.
#[test]
fn an_empty_team_is_an_error_rather_than_a_free_win() {
let mut h = history();
let err = h
.add_events(vec![Ev {
time: 1,
teams: teams(&[&[], &["b"]]),
outcome: Outcome::winner(0, 2),
}])
.unwrap_err();
assert!(
matches!(err, InferenceError::EmptyTeam { team: 0, .. }),
"{err:?}"
);
// Nothing was recorded, so the history is still empty.
assert!(h.current_skill(&"b").is_none());
}
#[test]
fn an_empty_team_is_reported_by_position() {
let mut h = history();
let err = h
.add_events(vec![Ev {
time: 1,
teams: teams(&[&["a"], &[]]),
outcome: Outcome::winner(0, 2),
}])
.unwrap_err();
assert!(
matches!(err, InferenceError::EmptyTeam { team: 1, .. }),
"{err:?}"
);
}
/// A NaN score used to ingest cleanly. `converge` reported `NonFiniteResult`,
/// but a caller who read `current_skill` first was handed `tau: NaN` with
/// nothing to say so.
#[test]
fn a_non_finite_score_is_rejected_at_ingestion() {
for bad in [f64::NAN, f64::INFINITY, f64::NEG_INFINITY] {
let mut h = history();
let err = h
.add_events(vec![Ev {
time: 1,
teams: teams(&[&["a"], &["b"]]),
outcome: Outcome::scores([bad, 0.0]),
}])
.unwrap_err();
assert!(
matches!(
err,
InferenceError::InvalidParameter {
parameter: trueskill_tt::Parameter::Score,
..
}
),
"{bad}: {err:?}"
);
assert!(h.current_skill(&"a").is_none(), "{bad} was recorded anyway");
}
}
/// A non-finite weight behaved exactly as `0.0` — the member contributed
/// nothing — while `converge` reported `converged: true` after one iteration
/// with a step of `(0.0, 0.0)`. So a NaN arriving from a division or a parse
/// was indistinguishable from a deliberate zero, and looked like a clean fit.
#[test]
fn a_non_finite_weight_is_rejected_at_ingestion() {
for bad in [f64::NAN, f64::INFINITY, f64::NEG_INFINITY] {
let mut h = history();
let err = h
.event(1)
.team(["a"])
.weights([bad])
.team(["b"])
.winner(0)
.commit()
.unwrap_err();
assert!(
matches!(
err,
InferenceError::InvalidParameter {
parameter: trueskill_tt::Parameter::Weight,
..
}
),
"{bad}: {err:?}"
);
assert!(h.current_skill(&"a").is_none(), "{bad} reached the history");
}
}
/// Zero and negative weights are expressible choices about how much a member
/// contributes, not malformed input, and `tests/degenerate_inputs.rs` pins
/// their behaviour deliberately. Rejecting non-finite values must not catch
/// them too.
#[test]
fn zero_and_negative_weights_still_ingest() {
for w in [0.0, -1.0, 0.5] {
let mut h = history();
h.event(1)
.team(["a"])
.weights([w])
.team(["b"])
.winner(0)
.commit()
.unwrap_or_else(|e| panic!("weight {w} should ingest: {e:?}"));
assert!(h.current_skill(&"a").is_some(), "weight {w}");
}
}
/// The fluent builder routes through the same chokepoint, so it inherits the
/// checks rather than needing its own.
#[test]
fn the_event_builder_inherits_the_shape_checks() {
let mut h = history();
let err = h.event(1).team(["a"]).winner(0).commit().unwrap_err();
assert!(
matches!(err, InferenceError::NotEnoughTeams { got: 1, .. }),
"{err:?}"
);
}
/// A well-formed event is untouched by any of this.
#[test]
fn a_well_formed_event_still_ingests() {
let mut h = history();
h.add_events(vec![Ev {
time: 1,
teams: teams(&[&["a"], &["b"]]),
outcome: Outcome::scores([3.0, 1.0]),
}])
.unwrap();
assert!(h.converge().unwrap().converged);
assert!(h.current_skill(&"a").unwrap().mu() > h.current_skill(&"b").unwrap().mu());
}
+353
View File
@@ -0,0 +1,353 @@
//! `History::joint` factorises once and answers many questions.
//!
//! The contract that matters is *identity*: a `Joint` must return exactly what
//! the one-shot call returns, bit for bit. A faster path that quietly disagreed
//! with the slow one would be worse than no fast path — a caller would get
//! different numbers depending on how many questions they happened to ask.
use smallvec::smallvec;
use trueskill_tt::{
ConstantDrift, ConvergenceOptions, Event, History, InferenceError, Member, Outcome, Team,
UnknownKeys,
};
type H = History;
fn duel(a: &'static str, b: &'static str, t: i64, sa: f64, sb: f64) -> Event<i64, &'static str> {
Event {
time: t,
teams: smallvec![
Team::with_members([Member::new(a)]),
Team::with_members([Member::new(b)]),
],
outcome: Outcome::scores([sa, sb]),
}
}
fn ranked(a: &'static str, b: &'static str, t: i64) -> Event<i64, &'static str> {
Event {
time: t,
teams: smallvec![
Team::with_members([Member::new(a)]),
Team::with_members([Member::new(b)]),
],
outcome: Outcome::winner(0, 2),
}
}
fn history(unknown: UnknownKeys) -> H {
History::builder()
.mu(0.0)
.sigma(6.0)
.beta(1.0)
.score_sigma(2.0)
.drift(ConstantDrift::new(0.5))
.unknown_keys(unknown)
.convergence(ConvergenceOptions {
max_iter: 20_000,
epsilon: 1e-13,
alpha: 1.0,
})
.build()
}
/// Several slices, competitors with different last appearances, so `latest`
/// and `at_slice` both have work to do.
fn fitted(unknown: UnknownKeys) -> H {
let mut h = history(unknown);
h.add_events(vec![
duel("a", "b", 1, 5.0, 2.0),
duel("c", "d", 1, 3.0, 3.5),
duel("a", "c", 2, 6.0, 1.0),
duel("b", "d", 3, 4.0, 3.0),
duel("a", "d", 4, 7.0, 2.0),
duel("b", "c", 5, 2.0, 4.0),
])
.unwrap();
let report = h.converge().unwrap();
assert!(report.converged, "fixture must converge");
h
}
const PAIRS: [(&str, &str); 6] = [
("a", "b"),
("a", "c"),
("a", "d"),
("b", "c"),
("b", "d"),
("c", "d"),
];
/// A joint reused across questions answers exactly what a fresh one per
/// question does. That is the whole correctness claim behind caching the
/// factorisation (#51); it used to be checked against the `History` one-shot
/// wrappers, which were deleted in #78, so it is checked against a fresh
/// factorisation instead — the same comparison, without the wrapper.
#[test]
fn a_reused_joint_answers_exactly_what_a_fresh_one_does() {
let h = fitted(UnknownKeys::Reject);
let joint = h.joint().unwrap();
for (a, b) in PAIRS {
let terms = [(&a, 1.0), (&b, -1.0)];
let one_shot = h.joint().unwrap().posterior_of(&terms).unwrap();
let cached = joint.posterior_of(&terms).unwrap();
assert_eq!(one_shot.mu(), cached.mu(), "{a} - {b}");
assert_eq!(one_shot.variance(), cached.variance(), "{a} - {b}");
}
}
#[test]
fn a_joint_agrees_at_a_pinned_time_too() {
let h = fitted(UnknownKeys::Reject);
let joint = h.joint().unwrap();
for time in 1..=5 {
for (a, b) in PAIRS {
let terms = [(&a, 1.0), (&b, -1.0)];
let one_shot = h.joint().unwrap().posterior_of_at(time, &terms);
let cached = joint.posterior_of_at(time, &terms);
match (one_shot, cached) {
(Ok(x), Ok(y)) => {
assert_eq!(x.mu(), y.mu(), "t={time} {a} - {b}");
assert_eq!(x.variance(), y.variance(), "t={time} {a} - {b}");
}
(Err(x), Err(y)) => assert_eq!(x, y, "t={time} {a} - {b}"),
(x, y) => panic!("t={time} {a} - {b}: disagreed on success: {x:?} vs {y:?}"),
}
}
}
}
#[test]
fn a_joint_scores_candidate_matchups_identically() {
let h = fitted(UnknownKeys::Reject);
let joint = h.joint().unwrap();
let (a, b) = ("a", "b");
let target = [(&a, 1.0), (&b, -1.0)];
for (x, y) in PAIRS {
let teams: [&[&&str]; 2] = [&[&x], &[&y]];
let one_shot = h
.joint()
.unwrap()
.expected_variance_reduction(&teams, &target)
.unwrap();
let cached = joint.expected_variance_reduction(&teams, &target).unwrap();
assert_eq!(one_shot, cached, "{x} vs {y}");
}
}
/// The whole point: a competitor appears once per slice, so the joint is over
/// appearances rather than competitors, and a caller sizing a batch needs to
/// know which.
#[test]
fn variables_counts_appearances_not_competitors() {
let h = fitted(UnknownKeys::Reject);
let joint = h.joint().unwrap();
// Four competitors, twelve appearances across five slices, all with
// positive drift between them, so no two collapse.
assert_eq!(joint.variables(), 12);
}
/// How much the collapse is worth, which is the part a caller has to plan
/// around: a drift-free competitor contributes **one** variable however long
/// the history, so the same events at `gamma = 0` and `gamma > 0` differ by
/// roughly the slice count in problem size — and by its cube in solve time.
///
/// Reported by a consumer as an 8x difference in solve time on a ~2,000-node,
/// 76-slice model (787 ms career against 6,214 ms drifting). This pins the
/// mechanism behind that so a change to the collapse rule cannot quietly
/// remove it.
#[test]
fn drift_free_competitors_shrink_the_joint_by_the_slice_count() {
fn variables(gamma: f64) -> usize {
let mut h = History::builder()
.mu(0.0)
.sigma(6.0)
.beta(1.0)
.score_sigma(2.0)
.drift(ConstantDrift::new(gamma))
.convergence(ConvergenceOptions {
max_iter: 20_000,
epsilon: 1e-13,
alpha: 1.0,
})
.build();
h.add_events(
(1..=10)
.map(|t| duel("a", "b", t, 5.0, 2.0))
.collect::<Vec<_>>(),
)
.unwrap();
let _ = h.converge().unwrap();
h.joint().unwrap().variables()
}
let drifting = variables(0.5);
let career = variables(0.0);
// Two competitors over ten slices: twenty appearances, or two variables.
assert_eq!(drifting, 20);
assert_eq!(career, 2);
assert_eq!(
drifting / career,
10,
"collapse should track the slice count"
);
}
/// With `drift = 0` consecutive appearances are the same latent variable, so
/// the joint is smaller than the appearance count.
#[test]
fn pinned_competitors_collapse_consecutive_appearances() {
let mut h = History::builder()
.mu(0.0)
.sigma(6.0)
.beta(1.0)
.score_sigma(2.0)
.drift(ConstantDrift::new(0.0))
.convergence(ConvergenceOptions {
max_iter: 20_000,
epsilon: 1e-13,
alpha: 1.0,
})
.build();
h.add_events(vec![
duel("a", "b", 1, 5.0, 2.0),
duel("a", "b", 2, 4.0, 3.0),
duel("a", "b", 3, 6.0, 1.0),
])
.unwrap();
assert!(h.converge().unwrap().converged);
assert_eq!(h.joint().unwrap().variables(), 2);
}
#[test]
fn a_ranked_history_has_no_exact_joint() {
let mut h = history(UnknownKeys::Reject);
h.add_events(vec![duel("a", "b", 1, 5.0, 2.0), ranked("a", "b", 2)])
.unwrap();
let _ = h.converge().unwrap();
assert!(matches!(
h.joint().unwrap_err(),
InferenceError::JointRequiresScoredEvents
));
}
/// Distinguishable from the ranked case, which is the point of splitting
/// `JointUnavailable { reason: &str }` into three variants (#74): "add events"
/// and "use predict_win_probabilities" are different instructions, and telling
/// them apart used to mean matching on English prose.
#[test]
fn an_empty_history_has_no_joint() {
let h = history(UnknownKeys::Reject);
assert!(matches!(
h.joint().unwrap_err(),
InferenceError::EmptyHistory
));
}
/// Unknown keys are decided per query, not when the joint is factorised — the
/// factorisation does not depend on the question.
#[test]
fn unknown_keys_are_rejected_per_query() {
let h = fitted(UnknownKeys::Reject);
let joint = h.joint().unwrap();
let (a, z) = ("a", "nobody");
assert!(matches!(
joint.posterior_of(&[(&a, 1.0), (&z, -1.0)]).unwrap_err(),
InferenceError::UnknownKey { .. }
));
// The handle is still usable afterwards.
let b = "b";
assert!(joint.posterior_of(&[(&a, 1.0), (&b, -1.0)]).is_ok());
}
/// Under `Prior`, an unseen competitor is independent of everything in the
/// history, and a reused joint must add the same prior variance a fresh one
/// does.
#[test]
fn unseen_competitors_match_a_fresh_factorisation() {
let h = fitted(UnknownKeys::Prior);
let joint = h.joint().unwrap();
let (a, z) = ("a", "nobody");
let terms = [(&a, 1.0), (&z, -1.0)];
let one_shot = h.joint().unwrap().posterior_of(&terms).unwrap();
let cached = joint.posterior_of(&terms).unwrap();
assert_eq!(one_shot.mu(), cached.mu());
assert_eq!(one_shot.variance(), cached.variance());
}
/// A drift too small to represent must collapse, not corrupt the matrix.
///
/// The collapse rule used to fire only at `drift <= 0.0` exactly. Anything
/// smaller-but-positive got an explicit `1.0 / drift` precision, and at
/// `drift = 1e-16` that entry is `1e16` — so `1e16 + 0.28` rounds back to
/// `1e16` and the prior and contrasts are annihilated in the stored `f64`.
///
/// Measured before the fix, at `drift_scale = 1e-10` this returned a variance
/// **12 000x too small** (a 111x overconfident interval) as `Ok`, with a band
/// just above it returning a misleading `JointUnavailable`.
#[test]
fn a_drift_too_small_to_represent_collapses_rather_than_corrupting() {
fn variance(scale: f64) -> f64 {
let mut h: History<String> = History::builder()
.key_type::<String>()
.mu(0.0)
.sigma(6.0)
.beta(1.0)
.score_sigma(2.0)
.drift(ConstantDrift::new(0.5))
.convergence(ConvergenceOptions {
max_iter: 20_000,
epsilon: 1e-13,
alpha: 1.0,
})
.build();
let mut events = Vec::new();
for t in 0..15i64 {
for k in 0..4usize {
let x = format!("p{}", (t as usize * 4 + k) % 8);
let y = format!("p{}", (t as usize * 4 + k + 3) % 8);
events.push(Event {
time: t,
teams: smallvec![
Team::with_members([Member::new(x).with_drift_scale(scale)]),
Team::with_members([Member::new(y).with_drift_scale(scale)]),
],
outcome: Outcome::scores([3.0, 1.0]),
});
}
}
h.add_events(events).unwrap();
assert!(h.converge().unwrap().converged);
let (a, b) = ("p0".to_string(), "p1".to_string());
let joint = h
.joint()
.expect("a tiny drift must not make the joint unavailable");
let g = joint.posterior_of(&[(&a, 1.0), (&b, -1.0)]).unwrap();
g.sigma() * g.sigma()
}
let collapsed = variance(0.0);
// Below the threshold every scale must reach the collapsed answer exactly,
// and none may error.
for scale in [1e-3, 1e-4, 1e-6, 1e-8, 1e-10, 1e-12] {
let v = variance(scale);
assert_eq!(
v.to_bits(),
collapsed.to_bits(),
"drift_scale {scale:e}: {v} vs collapsed {collapsed}"
);
}
// Above it, real drift is still modelled — otherwise this test would pass
// by collapsing everything.
let drifting = variance(1e-2);
assert!(
(drifting - collapsed).abs() / collapsed > 1e-5,
"a drift of 1e-2 must still move the answer: {drifting} vs {collapsed}"
);
}
+127
View File
@@ -0,0 +1,127 @@
//! The realistic program: keys arrive owned, queries are written with literals.
//!
//! Every prediction and joint query used to take `&[&[&K]]`, which at
//! `K = String` made a string literal *impossible* — the shape required three
//! levels of temporaries that all had to outlive the call. They are generic
//! over the borrowed key now, so one spelling works at both key types.
//!
//! Both key types are exercised in every test, because the point is that the
//! spelling is the same.
use trueskill_tt::{ConstantDrift, History};
type Owned = History<String>;
type Borrowed = History;
fn owned() -> Owned {
let mut h: Owned = History::builder().key_type::<String>().build();
for t in 1..=4 {
h.record_winner(&"alice".to_string(), &"bob".to_string(), t)
.expect("ingests");
}
h.converge().expect("converges");
h
}
fn borrowed() -> Borrowed {
let mut h = History::default();
for t in 1..=4 {
h.record_winner(&"alice", &"bob", t).expect("ingests");
}
h.converge().expect("converges");
h
}
#[test]
fn predictions_take_literals_at_either_key_type() {
let teams: &[&[&str]] = &[&["alice"], &["bob"]];
let a = owned()
.predict_win_probabilities(teams)
.expect("K = String");
let b = borrowed()
.predict_win_probabilities(teams)
.expect("K = &'static str");
assert_eq!(a, b, "the same fit through the same spelling");
assert!(a[0] > a[1], "alice won every game");
}
#[test]
fn every_team_shaped_query_accepts_the_same_slice() {
let h = owned();
let teams: &[&[&str]] = &[&["alice"], &["bob"]];
h.quality(teams).expect("quality");
let _ = h.predict_outcome(teams).expect("outcome");
h.predict_ranking(teams, &[0, 1]).expect("ranking");
h.expected_information_gain(teams)
.expect("information gain");
}
#[test]
fn linear_combinations_take_bare_keys() {
// `&[(&K, f64)]` at `K = String` meant `&[(&String, f64)]` — no literals.
// A scored history, because the joint needs one.
let mut h: Owned = History::builder().key_type::<String>().build();
for t in 1..=4 {
h.event(t)
.team([String::from("alice")])
.team([String::from("bob")])
.scores([21.0, 9.0])
.commit()
.expect("ingests");
}
h.converge().expect("converges");
let terms: &[(&str, f64)] = &[("alice", 1.0), ("bob", -1.0)];
let gap = h
.joint()
.expect("scored history has a joint")
.posterior_of(terms)
.expect("both keys are known");
assert!(gap.mu() > 0.0, "alice outscored bob every round");
}
/// `lookup` is gone with `Index` (#73); the accessors that answer the same
/// question all take a borrowed key.
#[test]
fn membership_queries_accept_a_borrowed_key() {
let h = owned();
assert!(h.current_skill("alice").is_some());
assert!(h.rating("alice").is_some());
assert!(h.learning_curve("alice").is_some());
assert!(h.current_skill("nobody").is_none());
assert!(h.rating("nobody").is_none());
assert!(h.learning_curve("nobody").is_none());
}
#[test]
fn gamma_sets_drift_without_naming_constant_drift() {
let mut a: Borrowed = History::builder().gamma(0.5).build();
let mut b: Borrowed = History::builder().drift(ConstantDrift::new(0.5)).build();
for h in [&mut a, &mut b] {
h.record_winner(&"x", &"y", 1).unwrap();
h.record_winner(&"y", &"x", 100).unwrap();
h.converge().unwrap();
}
let (ga, gb) = (a.current_skill("x").unwrap(), b.current_skill("x").unwrap());
assert_eq!((ga.mu(), ga.sigma()), (gb.mu(), gb.sigma()));
// Control: the shorthand is not a no-op — a different gamma differs.
let mut c: Borrowed = History::builder().gamma(0.0).build();
c.record_winner(&"x", &"y", 1).unwrap();
c.record_winner(&"y", &"x", 100).unwrap();
c.converge().unwrap();
assert_ne!(c.current_skill("x").unwrap().sigma(), ga.sigma());
}
#[test]
#[should_panic(expected = "gamma must be finite and non-negative")]
fn a_negative_gamma_is_rejected_rather_than_squared_away() {
let _: Borrowed = History::builder().gamma(-0.5).build();
}
+5 -4
View File
@@ -3,7 +3,7 @@
//! produced a tiny-negative precision whose `sigma() = 1/sqrt(pi)` was NaN, which the
//! moment-space `Sub` in the game chain propagated into every skill once the slice grew past
//! ~75 competitors (e.g. a real ranking dataset with hundreds of players).
use trueskill_tt::{ConstantDrift, ConvergenceOptions, EPSILON, History, ITERATIONS, NullObserver};
use trueskill_tt::{ConstantDrift, ConvergenceOptions, EPSILON, History, ITERATIONS};
/// Tiny deterministic LCG — avoids a dev-dependency on `rand`.
struct Lcg(u64);
@@ -24,10 +24,11 @@ impl Lcg {
}
fn nan_after_fit(players: usize) -> usize {
let mut h: History<i64, ConstantDrift, NullObserver, String> = History::builder_with_key()
let mut h: History<String> = History::builder()
.key_type::<String>()
.beta(1.0)
.sigma(6.0)
.drift(ConstantDrift(0.1))
.drift(ConstantDrift::new(0.1))
.convergence(ConvergenceOptions {
max_iter: ITERATIONS,
epsilon: EPSILON,
@@ -46,7 +47,7 @@ fn nan_after_fit(players: usize) -> usize {
let (w, l) = if rng.coin() { (a, b) } else { (b, a) };
h.record_winner(&ids[w], &ids[l], 0).unwrap();
}
h.converge().unwrap();
let _ = h.converge().unwrap();
ids.iter()
.filter(|id| {
+175
View File
@@ -0,0 +1,175 @@
//! The libm rule, enforced rather than asserted in prose.
//!
//! CLAUDE.md requires transcendentals to go through `libm`, not `std`:
//!
//! > IEEE 754 pins the basic operations and `sqrt` but says nothing about
//! > `exp`/`log`/`erf`, and `std` delegates to the *system* math library —
//! > measured, `f64::exp` and `libm::exp` disagree on 9.7% of inputs by one
//! > ULP. Since inference is an iterative fixed point, one ULP can change an
//! > iteration count.
//!
//! The rule was stated clearly and still violated in three production sites,
//! one of them `hypot` on the path of every scored event — whose measured
//! divergence, 12.1%, is *higher* than the `exp` figure the rule cites as its
//! own justification. Prose is evidently not enough, so this is a test.
//!
//! Tests may use either, which the crate documents, so `#[cfg(test)]` blocks
//! are excluded.
use std::{fs, path::Path};
/// Method-call spellings that reach the system math library.
///
/// `sqrt` is deliberately absent: IEEE 754 specifies it exactly, so `std` and
/// `libm` cannot disagree. `abs`, `recip`, `powi` and `mul_add` are likewise
/// exact or specified.
const FORBIDDEN: &[&str] = &[
"exp", "exp2", "exp_m1", "ln", "ln_1p", "log", "log2", "log10", "powf", "sin", "cos", "tan",
"asin", "acos", "atan", "atan2", "sinh", "cosh", "tanh", "hypot", "cbrt", "erf", "erfc",
];
/// Strip `#[cfg(test)]` items by brace matching, plus comments and string
/// literals, so a mention in prose is not mistaken for a call.
fn production_code(source: &str) -> String {
let mut out = String::with_capacity(source.len());
let bytes: Vec<char> = source.chars().collect();
let mut i = 0;
while i < bytes.len() {
let rest: String = bytes[i..].iter().take(16).collect();
if rest.starts_with("#[cfg(test)]") {
// Skip to the opening brace of the guarded item, then past its
// matching close.
let mut j = i;
while j < bytes.len() && bytes[j] != '{' {
j += 1;
}
let mut depth = 0usize;
while j < bytes.len() {
match bytes[j] {
'{' => depth += 1,
'}' => {
depth -= 1;
if depth == 0 {
j += 1;
break;
}
}
_ => {}
}
j += 1;
}
i = j;
continue;
}
if rest.starts_with("//") {
while i < bytes.len() && bytes[i] != '\n' {
i += 1;
}
continue;
}
if rest.starts_with("/*") {
i += 2;
while i + 1 < bytes.len() && !(bytes[i] == '*' && bytes[i + 1] == '/') {
i += 1;
}
i += 2;
continue;
}
if bytes[i] == '"' {
i += 1;
while i < bytes.len() && bytes[i] != '"' {
if bytes[i] == '\\' {
i += 1;
}
i += 1;
}
i += 1;
continue;
}
out.push(bytes[i]);
i += 1;
}
out
}
fn rust_files(dir: &Path, out: &mut Vec<std::path::PathBuf>) {
for entry in fs::read_dir(dir).expect("read src") {
let path = entry.expect("dir entry").path();
if path.is_dir() {
rust_files(&path, out);
} else if path.extension().is_some_and(|e| e == "rs") {
out.push(path);
}
}
}
#[test]
fn production_code_never_calls_a_std_transcendental() {
let mut files = Vec::new();
rust_files(Path::new("src"), &mut files);
assert!(files.len() > 10, "expected to find the crate's sources");
let mut offences = Vec::new();
for path in &files {
let source = fs::read_to_string(path).expect("read source");
let code = production_code(&source);
for (n, line) in code.lines().enumerate() {
for name in FORBIDDEN {
let needle = format!(".{name}(");
if line.contains(&needle) {
offences.push(format!("{}:{}: {}", path.display(), n + 1, line.trim()));
}
}
}
}
assert!(
offences.is_empty(),
"production code must call libm, not std, for transcendentals \
(`sqrt` is exempt IEEE 754 specifies it):\n{}",
offences.join("\n")
);
}
/// The stripper has to actually strip, or the test above passes vacuously.
#[test]
fn the_test_module_stripper_works() {
let source = r#"
fn production() { let _ = libm::exp(1.0); }
#[cfg(test)]
mod tests {
fn allowed() { let x = 1.0f64.exp(); }
}
fn also_production() {}
"#;
let code = production_code(source);
assert!(
code.contains("also_production"),
"stripped too much: {code}"
);
assert!(
!code.contains(".exp()"),
"failed to strip cfg(test): {code}"
);
}
/// And it must not strip a doc comment's worth of prose into oblivion, nor
/// mistake prose for a call.
#[test]
fn prose_is_not_mistaken_for_a_call() {
let source = "/// Uses `x.exp()` in the docs.\nfn f() { let _ = libm::exp(1.0); }\n";
let code = production_code(source);
assert!(!code.contains(".exp()"), "doc comment leaked: {code}");
assert!(code.contains("libm::exp"), "stripped real code: {code}");
}
+370
View File
@@ -0,0 +1,370 @@
//! Calibration of the crate's marginals against the EXACT posterior.
//!
//! A scored history is linear-Gaussian — `MarginFactor` encodes
//! `score_a - score_b ~ N(perf_a - perf_b, score_sigma^2)` — so the true joint
//! posterior has a closed form and the crate can be checked against ground
//! truth rather than against intuition. That is not possible for ranked
//! outcomes, whose truncation likelihood EP genuinely approximates.
//!
//! Two things are pinned here, and one is deliberately only recorded.
//!
//! **Pinned: on a tree the crate is exact**, means and variances both. Message
//! passing has no approximation to make when the factor graph has no cycles, so
//! any drift here would be a real defect.
//!
//! **Pinned: means are exact even with cycles.** This is the standard result
//! for Gaussian belief propagation (Weiss & Freeman 2001) and it is what makes
//! ratings trustworthy.
//!
//! **Recorded, not asserted: with cycles, marginal variances are too narrow.**
//! Measured on the round-robin fixture below, the crate reports sigma 1.430
//! where the exact posterior is 2.851 — a ratio of 0.502. That is the known
//! behaviour of loopy Gaussian BP, not a bug in this crate, and it is left
//! unasserted because fixing it is exactly what #46 proposes.
//!
//! Why that matters for a consumer, and why #46 cannot be implemented as "add
//! a covariance accessor": the exact correlation between two nodes here is
//! +0.857, so a consumer computing `sqrt(sa^2 + sb^2)` for a difference
//! overstates its width. But the too-narrow marginals partially cancel that,
//! leaving 1.327x rather than 2.646x. Adding true correlations to these
//! marginals without also correcting them would give 0.765 against a true
//! 1.524 — *overconfident*, which is the unsafe direction.
use smallvec::smallvec;
use trueskill_tt::{ConstantDrift, ConvergenceOptions, Event, History, Member, Outcome, Team};
const N: usize = 5;
const MU0: f64 = 0.0;
const SIGMA0: f64 = 6.0;
const BETA: f64 = 1.0;
const SCORE_SIGMA: f64 = 2.0;
/// A STAR: every event touches c0, so the node-event graph is a tree and
/// Gaussian BP is exact. Any discrepancy here is not caused by loops.
fn tree_fixture() -> Vec<(usize, usize, f64)> {
vec![(0, 1, 3.0), (0, 2, 5.0), (0, 3, 4.0), (0, 4, 6.0)]
}
/// (winner, loser, score_diff)
fn fixture() -> Vec<(usize, usize, f64)> {
vec![
(0, 1, 3.0),
(0, 2, 5.0),
(1, 2, 2.0),
(3, 4, 1.0),
(0, 3, 4.0),
(1, 4, 2.5),
(2, 3, 0.5),
(0, 4, 6.0),
(1, 3, 1.5),
(2, 4, 3.0),
]
}
/// Invert a small symmetric positive-definite matrix by Gauss-Jordan.
fn inverse(mut a: Vec<Vec<f64>>) -> Vec<Vec<f64>> {
let n = a.len();
let mut inv: Vec<Vec<f64>> = (0..n)
.map(|i| (0..n).map(|j| if i == j { 1.0 } else { 0.0 }).collect())
.collect();
for col in 0..n {
// partial pivot
let mut piv = col;
for r in col + 1..n {
if a[r][col].abs() > a[piv][col].abs() {
piv = r;
}
}
a.swap(col, piv);
inv.swap(col, piv);
let d = a[col][col];
for j in 0..n {
a[col][j] /= d;
inv[col][j] /= d;
}
for r in 0..n {
if r == col {
continue;
}
let f = a[r][col];
for j in 0..n {
a[r][j] -= f * a[col][j];
inv[r][j] -= f * inv[col][j];
}
}
}
inv
}
/// The exact posterior of a linear-Gaussian model:
/// precision = prior precision + sum of a_k a_k^T / v_k.
fn exact_for(obs: &[(usize, usize, f64)]) -> (Vec<f64>, Vec<Vec<f64>>) {
let mut lambda = vec![vec![0.0; N]; N];
let mut eta = [0.0; N];
for (i, row) in lambda.iter_mut().enumerate() {
row[i] = 1.0 / (SIGMA0 * SIGMA0);
eta[i] = MU0 / (SIGMA0 * SIGMA0);
}
// Each 1v1 observation: d ~ N(x_a - x_b, score_sigma^2 + 2 beta^2)
let v = SCORE_SIGMA * SCORE_SIGMA + 2.0 * BETA * BETA;
for &(a, b, d) in obs {
let mut vec_a = [0.0; N];
vec_a[a] = 1.0;
vec_a[b] = -1.0;
for i in 0..N {
for j in 0..N {
lambda[i][j] += vec_a[i] * vec_a[j] / v;
}
eta[i] += vec_a[i] * d / v;
}
}
let cov = inverse(lambda);
let mean: Vec<f64> = (0..N)
.map(|i| (0..N).map(|j| cov[i][j] * eta[j]).sum())
.collect();
(mean, cov)
}
fn key(i: usize) -> &'static str {
["c0", "c1", "c2", "c3", "c4"][i]
}
/// Returns (worst mean error, worst sd ratio).
fn fitted(obs: &[(usize, usize, f64)]) -> History {
let mut h: History = History::builder()
.mu(MU0)
.sigma(SIGMA0)
.beta(BETA)
.score_sigma(SCORE_SIGMA)
.drift(ConstantDrift::new(0.0))
.convergence(ConvergenceOptions {
max_iter: 20_000,
epsilon: 1e-13,
alpha: 1.0,
})
.build();
let events: Vec<Event<i64, &'static str>> = obs
.iter()
.copied()
.map(|(a, b, d)| Event {
time: 1,
teams: smallvec![
Team::with_members([Member::new(key(a))]),
Team::with_members([Member::new(key(b))]),
],
outcome: Outcome::scores([d, 0.0]),
})
.collect();
h.add_events(events).unwrap();
let report = h.converge().unwrap();
assert!(
report.converged,
"fixture must converge: {:?}",
report.final_step
);
h
}
/// Returns (worst mean error, worst sd ratio gap).
fn run(name: &str, obs: Vec<(usize, usize, f64)>) -> (f64, f64) {
println!("\n########## {name} ##########");
let h = fitted(&obs);
let (mean, cov) = exact_for(&obs);
println!("\n== marginals: crate vs the exact linear-Gaussian posterior ==");
println!(
"{:>4} {:>12} {:>12} {:>12} {:>12} {:>8}",
"node", "crate mu", "exact mu", "crate sd", "exact sd", "sd ratio"
);
for i in 0..N {
let g = h.current_skill(&key(i)).unwrap();
let exact_sd = cov[i][i].sqrt();
println!(
"{:>4} {:>12.6} {:>12.6} {:>12.6} {:>12.6} {:>8.3}",
key(i),
g.mu(),
mean[i],
g.sigma(),
exact_sd,
g.sigma() / exact_sd
);
}
let mut worst_mean = 0.0f64;
let mut worst_ratio_gap = 0.0f64;
for i in 0..N {
let g = h.current_skill(&key(i)).unwrap();
worst_mean = worst_mean.max((g.mu() - mean[i]).abs());
worst_ratio_gap = worst_ratio_gap.max((g.sigma() / cov[i][i].sqrt() - 1.0).abs());
}
println!("\n== what a consumer actually computes for a DIFFERENCE ==");
println!(
"{:>8} {:>12} {:>14} {:>14} {:>12}",
"pair", "exact", "naive(exact)", "naive(crate)", "crate err"
);
for i in 0..N {
for j in i + 1..N {
if i != 0 && j != 1 {
continue;
}
let gi = h.current_skill(&key(i)).unwrap();
let gj = h.current_skill(&key(j)).unwrap();
let exact_sd = (cov[i][i] + cov[j][j] - 2.0 * cov[i][j]).sqrt();
let naive_exact = (cov[i][i] + cov[j][j]).sqrt();
let naive_crate = (gi.sigma().powi(2) + gj.sigma().powi(2)).sqrt();
let corr = cov[i][j] / (cov[i][i].sqrt() * cov[j][j].sqrt());
println!(
"{:>8} {:>12.6} {:>14.6} {:>14.6} {:>11.3}x (corr {corr:.4})",
format!("{}-{}", key(i), key(j)),
exact_sd,
naive_exact,
naive_crate,
naive_crate / exact_sd
);
}
}
(worst_mean, worst_ratio_gap)
}
/// With no cycles there is nothing for message passing to approximate.
#[test]
fn on_a_tree_the_marginals_are_exact() {
let (mean_err, sd_gap) = run("TREE (star: no loops, BP is exact)", tree_fixture());
assert!(
mean_err < 1e-9,
"tree means should be exact, worst error {mean_err}"
);
assert!(
sd_gap < 1e-9,
"tree sigmas should be exact, worst ratio gap {sd_gap}"
);
}
/// With cycles the means stay exact — the property ratings depend on — while
/// the variances do not. The variance gap is measured and reported rather than
/// asserted; see the module docs.
#[test]
fn with_cycles_the_means_stay_exact_but_the_variances_shrink() {
let (mean_err, sd_gap) = run("LOOPY (round robin)", fixture());
assert!(
mean_err < 1e-9,
"loopy means must still be exact, worst error {mean_err}"
);
assert!(
sd_gap > 0.1,
"the loopy variance gap is the premise of #46; if it has closed, that \
issue and these docs need revisiting (worst ratio gap {sd_gap})"
);
}
/// The point of #46: `posterior_of` must reproduce the exact joint, including
/// the correlation that marginals cannot express.
#[test]
fn posterior_of_matches_the_exact_joint() {
for (name, obs) in [("tree", tree_fixture()), ("loopy", fixture())] {
let h = fitted(&obs);
let (_, cov) = exact_for(&obs);
println!("\n== posterior_of vs exact ({name}) ==");
println!(
"{:>12} {:>14} {:>14} {:>10}",
"functional", "posterior_of", "exact", "ratio"
);
for (i, j) in [(0usize, 1usize), (0, 2), (1, 3), (2, 4)] {
let got = h
.joint()
.unwrap()
.posterior_of(&[(&key(i), 1.0), (&key(j), -1.0)])
.expect("scored slice should have a joint");
let exact_sd = (cov[i][i] + cov[j][j] - 2.0 * cov[i][j]).sqrt();
println!(
"{:>12} {:>14.6} {:>14.6} {:>10.4}",
format!("{}-{}", key(i), key(j)),
got.sigma(),
exact_sd,
got.sigma() / exact_sd
);
assert!(
(got.sigma() - exact_sd).abs() / exact_sd < 1e-9,
"{name} {}-{}: posterior_of gave {} where the exact joint is {exact_sd}",
key(i),
key(j),
got.sigma()
);
}
// A single competitor: this is where the loopy marginal was 2x narrow.
for (i, row) in cov.iter().enumerate() {
let got = h.joint().unwrap().posterior_of(&[(&key(i), 1.0)]).unwrap();
let exact_sd = row[i].sqrt();
assert!(
(got.sigma() - exact_sd).abs() / exact_sd < 1e-9,
"{name} {}: posterior_of gave {} where exact is {exact_sd}",
key(i),
got.sigma()
);
}
println!(" single-competitor marginals also exact");
}
}
/// Cost of the dense solve as the slice grows. Recorded, not asserted.
#[test]
#[ignore = "timing probe, run explicitly"]
fn cost_scaling() {
use std::time::Instant;
for n in [50usize, 100, 200, 400, 800] {
let names: Vec<String> = (0..n).map(|i| format!("c{i}")).collect();
let mut h: History<String> = History::builder()
.key_type::<String>()
.score_sigma(2.0)
.drift(ConstantDrift::new(0.0))
.convergence(ConvergenceOptions {
max_iter: 200,
epsilon: 1e-8,
alpha: 1.0,
})
.build();
let mut seed = 5u64;
let mut rnd = move || {
seed ^= seed << 13;
seed ^= seed >> 7;
seed ^= seed << 17;
seed
};
let events: Vec<Event<i64, String>> = (0..n * 4)
.map(|_| {
let a = (rnd() as usize) % n;
let mut b = (rnd() as usize) % n;
if b == a {
b = (b + 1) % n;
}
Event {
time: 1,
teams: smallvec![
Team::with_members([Member::new(names[a].clone())]),
Team::with_members([Member::new(names[b].clone())]),
],
outcome: Outcome::scores([1.0, 0.0]),
}
})
.collect();
h.add_events(events).unwrap();
let _ = h.converge().unwrap();
let t = Instant::now();
let g = h
.joint()
.unwrap()
.posterior_of(&[(&names[0], 1.0), (&names[1], -1.0)])
.unwrap();
println!(" n={n:>4}: {:>10.2?} sigma {:.6}", t.elapsed(), g.sigma());
}
}
+217
View File
@@ -0,0 +1,217 @@
//! Inference must report numerical breakdown rather than call it convergence.
//!
//! The boundary rejects inputs that are *not numbers*, but finite inputs can
//! still overflow during inference — `beta.powi(2)` at 1e300 is infinite, and
//! infinity minus infinity is NaN. `NonFiniteResult` is the guard for that, and
//! it matters because the alternative is silent: NaN fails every comparison, so
//! a naive `step < epsilon` check reads a NaN step as *converged*.
//!
//! That is why the crate has `step_converged` / `step_is_finite` rather than
//! `!tuple_gt(..)`. These tests pin the guard from outside.
use smallvec::smallvec;
use trueskill_tt::{
ConstantDrift, Event, Gaussian, History, InferenceError, Member, Outcome, Team,
};
fn scored_fit(
sigma: f64,
beta: f64,
score_sigma: f64,
scores: [f64; 2],
) -> Result<bool, InferenceError> {
let mut h = History::builder()
.mu(0.0)
.sigma(sigma)
.beta(beta)
.score_sigma(score_sigma)
.build();
h.add_events(vec![Event {
time: 1i64,
teams: smallvec![
Team::with_members([Member::new("a")]),
Team::with_members([Member::new("b")]),
],
outcome: Outcome::scores(scores),
}])?;
h.converge().map(|r| r.converged)
}
/// Every one of these is built from finite, individually legal parameters. The
/// overflow happens inside inference, which is exactly the case the boundary
/// checks cannot catch.
///
/// Matched rather than merely `is_err()`: an assertion that only checks "some
/// error" would keep passing if these started failing at the boundary for an
/// unrelated reason, and would then be testing nothing.
#[test]
fn overflow_during_inference_is_reported_not_hidden() {
let cases: [(&str, f64, f64, f64, [f64; 2]); 5] = [
("huge sigma", 1e300, 1.0, 1.0, [3.0, 1.0]),
("huge beta", 6.0, 1e300, 1.0, [3.0, 1.0]),
("tiny sigma", 1e-300, 1.0, 1.0, [3.0, 1.0]),
("tiny score_sigma", 6.0, 1.0, 1e-300, [3.0, 1.0]),
("huge scores", 6.0, 1.0, 1.0, [1e308, -1e308]),
];
for (name, sigma, beta, score_sigma, scores) in cases {
match scored_fit(sigma, beta, score_sigma, scores) {
Err(InferenceError::NonFiniteStep { context, step, .. }) => {
assert_eq!(context, "History::converge", "{name}");
assert!(
!step.0.is_finite() || !step.1.is_finite(),
"{name}: reported NonFiniteResult with a finite step {step:?}"
);
}
other => panic!("{name}: expected NonFiniteResult, got {other:?}"),
}
}
}
/// The trap the invariant exists for: NaN fails every comparison, so a naive
/// `step < epsilon` test reads a NaN step as converged. A breakdown must never
/// come back as a successful fit.
#[test]
fn a_broken_fit_is_never_reported_as_converged() {
let mut h = History::builder().build();
h.add_events(vec![Event {
time: 1i64,
teams: smallvec![
Team::with_members([Member::new("a").with_prior(Gaussian::from_ms(1e300, 1e-300))]),
Team::with_members([Member::new("b")]),
],
outcome: Outcome::winner(0, 2),
}])
.unwrap();
let err = h.converge().unwrap_err();
assert!(
matches!(err, InferenceError::NonFiniteStep { .. }),
"a breakdown must not be reported as convergence: {err:?}"
);
// `converge_partial` must not launder it into an `Ok` either — the
// permissive path is permissive about *stopping short*, not about NaN.
let mut h2 = History::builder().build();
h2.add_events(vec![Event {
time: 1i64,
teams: smallvec![
Team::with_members([Member::new("a").with_prior(Gaussian::from_ms(1e300, 1e-300))]),
Team::with_members([Member::new("b")]),
],
outcome: Outcome::winner(0, 2),
}])
.unwrap();
assert!(matches!(
h2.converge_partial().unwrap_err(),
InferenceError::NonFiniteStep { .. }
));
}
/// The neighbouring case, so the tests above cannot pass by the fit simply
/// always failing: ordinary extreme-but-workable parameters still converge.
#[test]
fn merely_extreme_parameters_still_converge() {
assert!(scored_fit(1e6, 1.0, 1.0, [3.0, 1.0]).unwrap());
assert!(scored_fit(1e-6, 1.0, 1.0, [3.0, 1.0]).unwrap());
assert!(scored_fit(6.0, 1.0, 1e6, [3.0, 1.0]).unwrap());
assert!(scored_fit(6.0, 1.0, 1.0, [1e150, -1e150]).unwrap());
}
/// A NaN in one competitor must not be masked by a healthy competitor reduced
/// after it.
///
/// The convergence step is a fold over a `HashMap`, so which competitor is
/// reduced last is per-process hash order. Before the fix, `tuple_max` dropped
/// a NaN accumulator in favour of the next finite delta and this returned
/// `Ok(converged: true)` with a NaN posterior in **16 of 30 runs** on identical
/// input. Deterministic now, but note this test can only ever sample one hash
/// order per run — the ordering guarantee itself is pinned by
/// `tuple_max_propagates_a_nan_from_any_position` in the crate's unit tests.
#[test]
fn a_nan_competitor_is_not_masked_by_a_healthy_one() {
let mut h = History::builder()
.mu(0.0)
.sigma(6.0)
.beta(1.0)
.p_draw(0.1)
.build();
h.add_events(vec![
Event {
time: 1i64,
teams: smallvec![
Team::with_members([Member::new("a").with_prior(Gaussian::from_ms(0.0, 1e-200))]),
Team::with_members([Member::new("b")]),
],
outcome: Outcome::winner(0, 2),
},
// A healthy pair in the same slice, to be reduced alongside the NaN.
Event {
time: 1i64,
teams: smallvec![
Team::with_members([Member::new("c")]),
Team::with_members([Member::new("d")]),
],
outcome: Outcome::winner(0, 2),
},
])
.unwrap();
let err = h
.converge()
.expect_err("a NaN fit must never be reported as converged");
assert!(
matches!(err, InferenceError::NonFiniteStep { .. }),
"{err:?}"
);
}
/// A tie observed with a narrow draw margin between far-apart competitors must
/// produce a fit, not NaN skills.
///
/// The tie branch forms the truncated variance from `v^2 - u`, and both grow as
/// `alpha^2` while their difference stays `O(1)`. Deep enough into the tail
/// that subtraction had four digits left: measured, it returned `1 - w`
/// negative and `sqrt` of it was NaN. The half-line escape hatch did not cover
/// it, because that keys on how many window-widths from the mean the window
/// sits and a narrow window fails that however deep it is.
///
/// These parameters are ordinary for a precise-scoring domain, and the
/// neighbouring wider-margin case always worked — so this was a cliff, not
/// "extreme inputs break".
#[test]
fn a_narrow_draw_margin_far_into_the_tail_still_fits() {
for (beta, p_draw, sd, gap) in [
(1e-2, 1e-8, 1e-2, 10.0),
(1e-3, 1e-9, 1e-3, 1.0),
(1e-4, 1e-12, 1e-4, 1.0),
] {
let mut h = History::builder()
.mu(0.0)
.sigma(sd)
.beta(beta)
.p_draw(p_draw)
.drift(ConstantDrift::new(0.0))
.build();
h.add_events(vec![Event {
time: 1i64,
teams: smallvec![
Team::with_members([Member::new("a").with_prior(Gaussian::from_ms(0.0, sd))]),
Team::with_members([Member::new("b").with_prior(Gaussian::from_ms(gap, sd))]),
],
outcome: Outcome::draw(2),
}])
.unwrap();
let report = h
.converge()
.unwrap_or_else(|e| panic!("beta {beta:e}, p_draw {p_draw:e}: {e:?}"));
assert!(report.converged);
let skill = h.current_skill(&"a").unwrap();
assert!(
skill.mu().is_finite() && skill.sigma().is_finite() && skill.sigma() > 0.0,
"beta {beta:e}, p_draw {p_draw:e}: {skill:?}"
);
}
}
+8 -8
View File
@@ -42,7 +42,7 @@ fn every_observer_callback_fires() {
h.record_winner(&"a", &"b", 1).unwrap();
h.record_winner(&"b", &"c", 2).unwrap();
h.record_winner(&"c", &"a", 3).unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
assert!(
!recorder.iterations.lock().unwrap().is_empty(),
@@ -65,7 +65,7 @@ fn slice_callbacks_report_the_slice_they_swept() {
h.record_winner(&"a", &"b", 10).unwrap();
h.record_winner(&"a", &"b", 20).unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
let slices = recorder.slices.lock().unwrap();
@@ -93,7 +93,7 @@ fn a_single_slice_history_still_reports_its_sweep() {
let mut h = History::builder().observer(Arc::clone(&recorder)).build();
h.record_winner(&"a", &"b", 1).unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
let slices = recorder.slices.lock().unwrap();
assert!(
@@ -112,7 +112,7 @@ fn a_shared_observer_reaches_the_callers_handle() {
let mut h = History::builder().observer(Arc::clone(&recorder)).build();
h.record_winner(&"a", &"b", 1).unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
assert!(!recorder.iterations.lock().unwrap().is_empty());
assert!(!recorder.slices.lock().unwrap().is_empty());
@@ -125,12 +125,12 @@ fn a_trait_object_observer_works() {
let boxed: Box<dyn Observer<i64>> = Box::new(Recorder::default());
let mut h = History::builder().observer(boxed).build();
h.record_winner(&"a", &"b", 1).unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
let shared: Arc<dyn Observer<i64>> = Arc::new(Recorder::default());
let mut h = History::builder().observer(Arc::clone(&shared)).build();
h.record_winner(&"a", &"b", 1).unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
}
/// A non-shared observer can be reclaimed after convergence instead.
@@ -138,7 +138,7 @@ fn a_trait_object_observer_works() {
fn into_observer_returns_the_accumulated_state() {
let mut h = History::builder().observer(Recorder::default()).build();
h.record_winner(&"a", &"b", 1).unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
// Readable in place...
assert!(!h.observer().iterations.lock().unwrap().is_empty());
@@ -155,7 +155,7 @@ fn a_borrowed_observer_works() {
{
let mut h = History::builder().observer(&recorder).build();
h.record_winner(&"a", &"b", 1).unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
}
assert!(!recorder.iterations.lock().unwrap().is_empty());
}
+151
View File
@@ -0,0 +1,151 @@
//! `predict_margin`: the predictive distribution of a scored matchup.
use smallvec::smallvec;
use trueskill_tt::{
ConstantDrift, ConvergenceOptions, Event, History, InferenceError, Member, Outcome, Team,
UnknownKeys,
};
fn builder(policy: UnknownKeys) -> History {
History::builder()
.mu(0.0)
.sigma(6.0)
.beta(1.0)
.score_sigma(2.0)
.drift(ConstantDrift::new(0.0))
.unknown_keys(policy)
.convergence(ConvergenceOptions {
max_iter: 5_000,
epsilon: 1e-12,
alpha: 1.0,
})
.build()
}
fn round(a: &'static str, b: &'static str, sa: f64, sb: f64) -> Event<i64, &'static str> {
Event {
time: 1,
teams: smallvec![
Team::with_members([Member::new(a)]),
Team::with_members([Member::new(b)]),
],
outcome: Outcome::scores([sa, sb]),
}
}
/// A history where "veteran" and "regular" are well observed and "novice"
/// appears once.
fn fitted(policy: UnknownKeys) -> History {
let mut h = builder(policy);
let mut events: Vec<_> = (0..40)
.map(|t| round("veteran", "regular", 10.0 + f64::from(t % 3), 5.0))
.collect();
events.push(round("veteran", "novice", 10.0, 6.0));
h.add_events(events).unwrap();
let _ = h.converge().unwrap();
h
}
/// The property #48 exists for: the interval must widen when the model knows
/// less. Their hand-fitted noise law quoted the same sigma for a competitor
/// with forty rounds and one with none.
#[test]
fn the_interval_widens_as_the_model_knows_less() {
let h = fitted(UnknownKeys::Prior);
let well_known = h
.predict_margin(&[&[&"veteran"], &[&"regular"]])
.unwrap()
.sigma();
let thin = h
.predict_margin(&[&[&"veteran"], &[&"novice"]])
.unwrap()
.sigma();
let unseen = h
.predict_margin(&[&[&"veteran"], &[&"stranger"]])
.unwrap()
.sigma();
assert!(
well_known < thin && thin < unseen,
"margin width should grow as evidence thins: {well_known} < {thin} < {unseen}"
);
}
/// #48's second requirement: an unseen competitor is a legitimate question, not
/// an error, and the answer should come from the prior rather than be faked.
#[test]
fn an_unseen_competitor_is_answered_from_the_prior() {
let h = fitted(UnknownKeys::Prior);
let g = h.predict_margin(&[&[&"nobody"], &[&"no_one"]]).unwrap();
// Two unknowns: the gap is centred on zero and carries both priors plus
// both performance noises plus the observation noise.
assert!(g.mu().abs() < 1e-9, "mu {}", g.mu());
let expected = (2.0 * 36.0 + 2.0 * 1.0 + 4.0f64).sqrt();
assert!(
(g.sigma() - expected).abs() < 1e-9,
"sigma {} vs expected {expected}",
g.sigma()
);
}
#[test]
fn reject_still_rejects() {
let h = fitted(UnknownKeys::Reject);
assert!(matches!(
h.predict_margin(&[&[&"veteran"], &[&"stranger"]]),
Err(InferenceError::UnknownKey { .. })
));
}
/// The margin is the *difference*, so it must be antisymmetric in the teams.
#[test]
fn swapping_the_teams_negates_the_margin() {
let h = fitted(UnknownKeys::Prior);
let forward = h.predict_margin(&[&[&"veteran"], &[&"regular"]]).unwrap();
let reverse = h.predict_margin(&[&[&"regular"], &[&"veteran"]]).unwrap();
assert!((forward.mu() + reverse.mu()).abs() < 1e-9);
assert!((forward.sigma() - reverse.sigma()).abs() < 1e-12);
}
/// The predictive interval must be wider than the skill gap alone: it also
/// carries per-event performance noise and the observation noise.
#[test]
fn the_predictive_interval_exceeds_the_skill_uncertainty() {
let h = fitted(UnknownKeys::Prior);
let skill_gap = h
.joint()
.unwrap()
.posterior_of(&[(&"veteran", 1.0), (&"regular", -1.0)])
.unwrap();
let predictive = h.predict_margin(&[&[&"veteran"], &[&"regular"]]).unwrap();
assert!(
(predictive.mu() - skill_gap.mu()).abs() < 1e-12,
"means agree"
);
// beta^2 twice plus score_sigma^2 = 2 + 4.
let expected = (skill_gap.sigma().powi(2) + 6.0).sqrt();
assert!((predictive.sigma() - expected).abs() < 1e-12);
assert!(predictive.sigma() > skill_gap.sigma());
}
#[test]
fn shape_errors_are_reported() {
let h = fitted(UnknownKeys::Prior);
assert!(matches!(
h.predict_margin(&[&[&"veteran"]]),
Err(InferenceError::MismatchedShape {
expected: 2,
got: 1,
..
})
));
let empty: [&&str; 0] = [];
assert!(matches!(
h.predict_margin(&[&[&"veteran"], &empty]),
Err(InferenceError::EmptyTeam { team: 1, .. })
));
}
+151 -27
View File
@@ -9,7 +9,7 @@ fn history_with(names: &[&'static str], p_draw: f64) -> History {
for pair in names.windows(2) {
h.record_winner(&pair[0], &pair[1], 1).unwrap();
}
h.converge().unwrap();
let _ = h.converge().unwrap();
h
}
@@ -20,14 +20,21 @@ fn unknown_keys_are_reported_not_silently_dropped() {
let err = h
.predict_outcome(&[&[&"a"], &[&"ghost"]])
.expect_err("an unknown key must not yield a confident prediction");
assert_eq!(err, InferenceError::UnknownKey { team: 1, member: 0 });
assert!(
matches!(
&err,
InferenceError::UnknownKey { team: 1, member: 0, key, .. }
if key == "\"ghost\""
),
"{err:?}"
);
// Every prediction entry point, not just one.
assert!(
h.predict_win_probabilities(&[&[&"a"], &[&"ghost"]])
.is_err()
);
assert!(h.predict_quality(&[&[&"a"], &[&"ghost"]]).is_err());
assert!(h.quality(&[&[&"a"], &[&"ghost"]]).is_err());
assert!(h.predict_ranking(&[&[&"a"], &[&"ghost"]], &[0, 1]).is_err());
}
@@ -35,25 +42,36 @@ fn unknown_keys_are_reported_not_silently_dropped() {
fn an_entirely_unknown_team_is_an_error() {
let h = history_with(&["a", "b"], 0.0);
let err = h.predict_outcome(&[&[&"a"], &[&"x", &"y"]]).unwrap_err();
assert_eq!(err, InferenceError::UnknownKey { team: 1, member: 0 });
assert!(
matches!(
&err,
InferenceError::UnknownKey { team: 1, member: 0, key, .. }
if key == "\"x\""
),
"{err:?}"
);
}
#[test]
fn degenerate_team_shapes_are_errors_rather_than_panics() {
let h = history_with(&["a", "b"], 0.0);
assert_eq!(
assert!(matches!(
h.predict_outcome(&[&[&"a"]]).unwrap_err(),
InferenceError::NotEnoughTeams { got: 1 }
);
assert_eq!(
h.predict_outcome(&[]).unwrap_err(),
InferenceError::NotEnoughTeams { got: 0 }
);
assert_eq!(
InferenceError::NotEnoughTeams { got: 1, .. }
),);
// An empty team list cannot infer the key type — nothing in `&[]` names it.
// The annotation is the cost of `predict_*` being generic over the borrowed
// key, and it only bites on the degenerate call.
let none: &[&[&str]] = &[];
assert!(matches!(
h.predict_outcome(none).unwrap_err(),
InferenceError::NotEnoughTeams { got: 0, .. }
),);
assert!(matches!(
h.predict_outcome(&[&[&"a"], &[]]).unwrap_err(),
InferenceError::EmptyTeam { team: 1 }
);
InferenceError::EmptyTeam { team: 1, .. }
));
}
#[test]
@@ -79,13 +97,10 @@ fn the_outcome_space_is_capped_rather_than_hanging() {
let refs: Vec<&[&&str]> = too_many.iter().map(Vec::as_slice).collect();
let err = h.predict_outcome(&refs).unwrap_err();
assert_eq!(
assert!(matches!(
err,
InferenceError::TooManyTeams {
got: 8,
max: MAX_PREDICTED_TEAMS
}
);
InferenceError::TooManyTeams { got: 8, max, .. } if max == MAX_PREDICTED_TEAMS
));
// The cheap paths stay available at any size.
let wins = h.predict_win_probabilities(&refs).unwrap();
@@ -184,7 +199,7 @@ fn the_stronger_competitor_is_favoured() {
for t in 1..=10 {
h.record_winner(&"strong", &"weak", t).unwrap();
}
h.converge().unwrap();
let _ = h.converge().unwrap();
let p = h.predict_outcome(&[&[&"strong"], &[&"weak"]]).unwrap();
let (best, _) = p.most_likely().expect("a most likely outcome");
@@ -206,7 +221,7 @@ fn team_size_affects_the_prediction() {
.winner(0)
.commit()
.unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
let p = h.predict_outcome(&[&[&"a", &"b"], &[&"c"]]).unwrap();
assert!((p.total() - 1.0).abs() < 1e-6, "total = {}", p.total());
@@ -229,7 +244,7 @@ fn information_gain_prefers_the_uncertain_pairing() {
h.record_winner(&"rival", &"known", t + 100).unwrap();
}
h.record_winner(&"known", &"newcomer", 500).unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
let settled = h
.expected_information_gain(&[&[&"known"], &[&"rival"]])
@@ -268,11 +283,12 @@ fn information_gain_respects_the_entropy_ceiling() {
#[test]
fn information_gain_reports_unknown_keys() {
let h = history_with(&["a", "b"], 0.0);
assert_eq!(
h.expected_information_gain(&[&[&"a"], &[&"ghost"]])
assert!(matches!(
&h.expected_information_gain(&[&[&"a"], &[&"ghost"]])
.unwrap_err(),
InferenceError::UnknownKey { team: 1, member: 0 }
);
InferenceError::UnknownKey { team: 1, member: 0, key, .. }
if key == "\"ghost\""
));
}
/// A draw-enabled history has three outcomes to weigh rather than two, so the
@@ -289,3 +305,111 @@ fn information_gain_accounts_for_draws() {
let dist = with_draws.predict_outcome(&[&[&"a"], &[&"b"]]).unwrap();
assert!(dist.probability_of(&[0, 0]) > 0.0);
}
/// The defect that cost a consumer a day: `UnknownKey { team: 0, member: 0 }`
/// says nothing about *which* key is unknown, so the natural handling — log it,
/// fall back to a neutral value — converts a total miss into a plausible
/// constant. The key has to be in the error, and in its `Display`.
#[test]
fn unknown_key_names_the_key_it_could_not_find() {
let h = history_with(&["a", "b"], 0.0);
let err = h.predict_outcome(&[&[&"a"], &[&"never_seen"]]).unwrap_err();
match &err {
InferenceError::UnknownKey { key, .. } => {
assert!(
key.contains("never_seen"),
"the error should name the key, got {key}"
);
}
other => panic!("expected UnknownKey, got {other:?}"),
}
let rendered = err.to_string();
assert!(
rendered.contains("never_seen"),
"Display should name the key: {rendered}"
);
assert!(
rendered.contains("pre-filter"),
"Display should say what to do about it: {rendered}"
);
}
// ---------------------------------------------------------------------------
// UnknownKeys policy
// ---------------------------------------------------------------------------
fn history_with_policy(names: &[&'static str], policy: trueskill_tt::UnknownKeys) -> History {
let mut h = History::builder().unknown_keys(policy).build();
for pair in names.windows(2) {
h.record_winner(&pair[0], &pair[1], 1).unwrap();
}
let _ = h.converge().unwrap();
h
}
#[test]
fn reject_is_the_default() {
let h = history_with(&["a", "b"], 0.0);
assert!(matches!(
h.predict_outcome(&[&[&"a"], &[&"ghost"]]),
Err(InferenceError::UnknownKey { .. })
));
}
#[test]
fn prior_answers_instead_of_erroring() {
let h = history_with_policy(&["a", "b"], trueskill_tt::UnknownKeys::Prior);
let p = h
.predict_outcome(&[&[&"a"], &[&"ghost"]])
.expect("Prior should answer rather than reject");
assert!((p.total() - 1.0).abs() < 1e-6);
}
/// Two competitors the model has never seen are genuinely a coin flip. The
/// point is that this is now *derived* rather than a constant a caller
/// substitutes after swallowing an error.
#[test]
fn two_unknown_competitors_are_an_honest_coin_flip() {
let h = history_with_policy(&["a", "b"], trueskill_tt::UnknownKeys::Prior);
let wins = h
.predict_win_probabilities(&[&[&"nobody"], &[&"no_one"]])
.unwrap();
assert!((wins[0] - 0.5).abs() < 1e-9, "{wins:?}");
assert!((wins[1] - 0.5).abs() < 1e-9, "{wins:?}");
}
/// The property that rules out a `Skip` mode: an unknown member must make a
/// team *less* certain, never more. Skipping would drop the member's variance
/// from the sum and narrow the team, which is backwards.
#[test]
fn an_unknown_member_widens_its_team_rather_than_narrowing_it() {
let h = history_with_policy(&["a", "b", "c"], trueskill_tt::UnknownKeys::Prior);
// "a" alone against "b" — then "a" plus an unknown partner against "b".
let solo = h.predict_win_probabilities(&[&[&"a"], &[&"b"]]).unwrap();
let with_unknown = h
.predict_win_probabilities(&[&[&"a", &"stranger"], &[&"b"]])
.unwrap();
// Adding an unknown partner pulls the outcome toward even, because the
// team's performance spread grew.
assert!(
(with_unknown[0] - 0.5).abs() < (solo[0] - 0.5).abs(),
"an unknown partner should make the result less certain: solo {solo:?}, \
with unknown {with_unknown:?}"
);
}
#[test]
fn prior_reaches_every_prediction_entry_point() {
let h = history_with_policy(&["a", "b"], trueskill_tt::UnknownKeys::Prior);
let teams: &[&[&&str]] = &[&[&"a"], &[&"ghost"]];
assert!(h.quality(teams).is_ok());
assert!(h.predict_win_probabilities(teams).is_ok());
assert!(h.predict_outcome(teams).is_ok());
assert!(h.predict_ranking(teams, &[0, 1]).is_ok());
assert!(h.expected_information_gain(teams).is_ok());
}
+162
View File
@@ -0,0 +1,162 @@
//! Bounds that any correct implementation must satisfy, swept rather than
//! spot-checked.
//!
//! The crate's docs call the `ln k` ceiling "the sharpest available test of an
//! implementation", and record that an early prototype returned 4.77 nats. It
//! was violated again — 3.237828 nats against `ln 2` — because the existing
//! check sampled one fixture and the violation lives in a specific regime: a
//! large ratio between the widest and narrowest performance sigma, where the
//! shared prediction grid could not resolve the narrow density and returned
//! probabilities greater than one.
//!
//! A single fixture cannot defend a bound like this. A sweep can.
use trueskill_tt::{
ConstantDrift, GameOptions, Gaussian, InferenceError, Rating, expected_information_gain,
};
type R = Rating<i64, ConstantDrift>;
/// How many random matchups the ceiling sweep draws.
///
/// Scaled by build profile rather than fixed. Each sample runs a full inference
/// pass per outcome, and that is about **19x** faster in release — measured,
/// 20 000 samples take 12.1s released against 23s for 2 000 in debug. `just
/// test` runs three debug feature combinations and one release one, so a fixed
/// count pays the slow price three times and the fast one once, which is
/// exactly backwards.
///
/// The debug run is here to prove the sweep still compiles and holds on a small
/// sample; the release run is the one that actually searches. The violation
/// this guards was found at a rate near 1.8%, so even the debug count expects
/// tens of hits in the regime.
#[cfg(debug_assertions)]
const SAMPLES: usize = 1_000;
#[cfg(not(debug_assertions))]
const SAMPLES: usize = 50_000;
/// Deterministic LCG, so a failure is reproducible from the printed seed.
struct Lcg(u64);
impl Lcg {
fn next_f64(&mut self) -> f64 {
self.0 = self
.0
.wrapping_mul(6_364_136_223_846_793_005)
.wrapping_add(1_442_695_040_888_963_407);
// Top 53 bits to [0, 1).
((self.0 >> 11) as f64) / ((1u64 << 53) as f64)
}
fn in_range(&mut self, lo: f64, hi: f64) -> f64 {
lo + (hi - lo) * self.next_f64()
}
/// Log-uniform, so the sweep spends its samples across magnitudes rather
/// than crowding the top of the range — the violations live at small sigma.
fn log_uniform(&mut self, lo: f64, hi: f64) -> f64 {
let t = self.next_f64();
(lo.ln() + t * (hi.ln() - lo.ln())).exp()
}
}
#[test]
fn information_gain_never_exceeds_the_entropy_of_the_outcome() {
let mut rng = Lcg(0x5eed_1234_abcd_ef01);
let ceiling = 2.0_f64.ln();
let mut evaluated = 0usize;
let mut refused = 0usize;
for i in 0..SAMPLES {
let mu_a = rng.in_range(-100.0, 100.0);
let mu_b = rng.in_range(-100.0, 100.0);
let sigma_a = rng.log_uniform(1e-4, 1e2);
let sigma_b = rng.log_uniform(1e-4, 1e2);
let beta = rng.log_uniform(1e-4, 1e1);
let a = R::new(
Gaussian::from_ms(mu_a, sigma_a),
beta,
ConstantDrift::new(0.0),
);
let b = R::new(
Gaussian::from_ms(mu_b, sigma_b),
beta,
ConstantDrift::new(0.0),
);
let options = GameOptions {
p_draw: 0.0,
..GameOptions::default()
};
match expected_information_gain(&[&[a], &[b]], &options) {
Ok(gain) => {
evaluated += 1;
assert!(
gain.is_finite(),
"sample {i}: non-finite gain {gain} \
(mu {mu_a}, {mu_b}; sigma {sigma_a:e}, {sigma_b:e}; beta {beta:e})"
);
assert!(
gain >= 0.0,
"sample {i}: negative gain {gain} \
(mu {mu_a}, {mu_b}; sigma {sigma_a:e}, {sigma_b:e}; beta {beta:e})"
);
assert!(
gain <= ceiling + 1e-9,
"sample {i}: gain {gain} exceeds ln 2 = {ceiling} \
(mu {mu_a}, {mu_b}; sigma {sigma_a:e}, {sigma_b:e}; beta {beta:e})"
);
}
// Refusing to answer is acceptable; answering wrongly is not.
Err(InferenceError::GridTooCoarse { .. }) => refused += 1,
Err(e) => panic!("sample {i}: unexpected error {e:?}"),
}
}
// The sweep must actually exercise the function, not pass by refusing
// everything.
assert!(
evaluated * 2 > SAMPLES,
"only {evaluated} of {SAMPLES} samples were evaluated ({refused} refused); \
the sweep is no longer testing anything"
);
// And it must still reach the regime where the ceiling was violated —
// large sigma ratios, which is exactly where the grid now refuses. Without
// this the sweep could drift into only-easy inputs and stop being a guard.
assert!(
refused > 0,
"no sample reached the coarse-grid regime; the sweep no longer covers \
the case that produced 3.24 nats"
);
}
/// The regime that produced 3.237828 nats, pinned exactly.
#[test]
fn the_known_ceiling_violation_no_longer_answers_wrongly() {
let a = R::new(
Gaussian::from_ms(9.577_887_112_129_012, 0.000_132_507_526_585_134_38),
0.000_307_235_559_013_096_2,
ConstantDrift::new(0.0),
);
let b = R::new(
Gaussian::from_ms(-14.114_932_828_525_696, 91.586_690_140_921_16),
0.000_307_235_559_013_096_2,
ConstantDrift::new(0.0),
);
let options = GameOptions {
p_draw: 0.0,
..GameOptions::default()
};
match expected_information_gain(&[&[a], &[b]], &options) {
Ok(gain) => assert!(
gain <= 2.0_f64.ln() + 1e-9,
"returned {gain}, over the ln 2 ceiling"
),
Err(InferenceError::GridTooCoarse { needed, max, .. }) => {
assert!(needed > max, "needed {needed} should exceed max {max}");
}
Err(e) => panic!("unexpected error {e:?}"),
}
}
+152
View File
@@ -0,0 +1,152 @@
//! No prediction path may answer from a fit it cannot answer from.
//!
//! `converge` grew a `NonFiniteResult` guard; nothing stopped a caller from
//! ignoring that error and predicting anyway. The three failures that produced
//! were each differently wrong: `Ok(NaN)`, a panic out of a `Result`-returning
//! method, and `Ok([0.0, 0.0])` — finite, plausible, summing to zero against a
//! doc that promises one.
//!
//! Every test here has a healthy control, so none can pass by everything
//! returning `Err`.
use trueskill_tt::{
ConstantDrift, Event, Gaussian, History, InferenceError, Member, Outcome, Team,
};
type H = History;
fn build(beta: f64, prior: Option<Gaussian>, outcome: Outcome) -> H {
let mut h: H = History::builder()
.beta(beta)
.drift(ConstantDrift::new(0.0))
.build();
let member = |k: &'static str| match prior {
Some(p) => Member::new(k).with_prior(p),
None => Member::new(k),
};
let _ = h.add_events(vec![Event {
time: 1,
teams: [
Team::with_members([member("a")]),
Team::with_members([member("b")]),
]
.into_iter()
.collect(),
outcome,
}]);
h
}
/// Point-mass priors with `beta(0.0)` on a *ranked* event: `converge` reports
/// `NonFiniteResult` and the stored posteriors are `pi: NaN, tau: NaN`.
fn nan_poisoned() -> H {
let mut h = build(
0.0,
Some(Gaussian::from_ms(0.0, 0.0)),
Outcome::winner(0, 2),
);
let err = h.converge().expect_err("this fixture must not converge");
assert!(
matches!(err, InferenceError::NonFiniteStep { .. }),
"{err:?}"
);
h
}
/// The same degenerate parameters on a *scored* event, where inference
/// converges cleanly and leaves legitimate point-mass posteriors behind. The
/// fit is fine; it is prediction that has nothing to work with.
fn degenerate_but_converged() -> H {
let mut h = build(
0.0,
Some(Gaussian::from_ms(0.0, 0.0)),
Outcome::scores([1.0, 0.0]),
);
h.converge().expect("this fixture converges");
h
}
fn healthy() -> H {
let mut h = build(1.0, None, Outcome::winner(0, 2));
h.converge().expect("control converges");
h
}
macro_rules! all_predictions {
($h:ident, $f:expr) => {{
let teams: &[&[&&'static str]] = &[&[&"a"], &[&"b"]];
let f = $f;
f("quality", $h.quality(teams).map(|_| ()));
f(
"predict_win_probabilities",
$h.predict_win_probabilities(teams).map(|_| ()),
);
f("predict_outcome", $h.predict_outcome(teams).map(|_| ()));
f(
"predict_ranking",
$h.predict_ranking(teams, &[0, 1]).map(|_| ()),
);
f(
"expected_information_gain",
$h.expected_information_gain(teams).map(|_| ()),
);
}};
}
#[test]
fn a_nan_poisoned_fit_is_refused_by_every_prediction_path() {
let h = nan_poisoned();
all_predictions!(h, |name: &str, r: Result<(), InferenceError>| {
match r {
Err(InferenceError::NonFiniteSkill { .. }) => {}
other => panic!("{name} answered from a NaN fit: {other:?}"),
}
});
}
#[test]
fn degenerate_performances_are_refused_rather_than_answered_wrongly() {
let h = degenerate_but_converged();
// The fit itself is sound — the posteriors are point masses, not NaN.
let skill = h.current_skill("a").expect("a played");
assert_eq!(skill.sigma(), 0.0);
assert!(skill.mu().is_finite());
// `quality` previously PANICKED here, out of a method that returns
// `Result`: the contrast covariance is exactly singular when beta is zero
// and every skill is a point mass.
all_predictions!(h, |name: &str, r: Result<(), InferenceError>| {
match r {
Err(InferenceError::NoPerformanceVariance) => {}
other => panic!("{name} predicted from a degenerate fit: {other:?}"),
}
});
}
#[test]
fn the_control_history_answers_every_prediction() {
let h = healthy();
all_predictions!(h, |name: &str, r: Result<(), InferenceError>| {
assert!(r.is_ok(), "{name} failed on a healthy history: {r:?}");
});
}
#[test]
fn win_probabilities_sum_to_one_on_the_control() {
// The promise the `Ok([0.0, 0.0])` case broke. Asserted on the control so
// the guard above cannot be "fixed" by making every path error.
let h = healthy();
let p = h
.predict_win_probabilities(&[&[&"a"], &[&"b"]])
.expect("control predicts");
let total: f64 = p.iter().sum();
assert!(
(total - 1.0).abs() < 1e-6,
"win probabilities sum to {total}"
);
}
+7
View File
@@ -0,0 +1,7 @@
# Seeds for failure cases proptest has generated in the past. It is
# automatically read and these particular cases re-run before any
# novel cases are generated.
#
# It is recommended to check this file in to source control so that
# everyone who runs the test benefits from these saved cases.
cc 8859be600e638573980f78622b8fcd8b4553ca34a9a041746c417f7e4293f89c # shrinks to games = [(0, 1), (0, 1), (4, 6), (2, 0), (0, 1), (0, 6), (0, 1), (2, 0), (6, 4), (0, 1), (0, 2), (6, 4), (1, 0), (4, 0), (0, 2)]
+29 -8
View File
@@ -26,7 +26,10 @@ const KEYS: [&str; 8] = ["a", "b", "c", "d", "e", "f", "g", "h"];
fn history_from(games: &[(usize, usize)]) -> History {
let mut h = History::builder()
.convergence(ConvergenceOptions {
max_iter: 200,
// 200 was not enough: the batched side stopped at the cap with a
// step of 3.4e-9, so this test was comparing two truncated fits and
// attributing the gap to ingestion order.
max_iter: 20_000,
epsilon: 1e-10,
..ConvergenceOptions::default()
})
@@ -61,10 +64,15 @@ proptest! {
fn converged_posteriors_are_always_finite(games in pairs()) {
let mut h = history_from(&games);
h.converge().unwrap();
let _ = h.converge().unwrap();
for key in KEYS {
for (time, g) in h.learning_curve(key) {
// A generated schedule need not touch every key, and an unplayed
// key is `None` rather than an empty curve.
let Some(curve) = h.learning_curve(key) else {
continue;
};
for (time, g) in curve {
assert_finite(g, &format!("{key} at t={time}"));
}
}
@@ -79,7 +87,7 @@ proptest! {
fn log_evidence_is_a_finite_log_probability(games in pairs()) {
let mut h = history_from(&games);
h.converge().unwrap();
let _ = h.converge().unwrap();
let batch = h.log_evidence();
let filtered = h.filtered_log_evidence();
@@ -98,7 +106,7 @@ proptest! {
let before = h.filtered_log_evidence();
h.converge().unwrap();
let _ = h.converge().unwrap();
let after = h.filtered_log_evidence();
@@ -114,14 +122,21 @@ proptest! {
fn ingestion_order_does_not_change_the_answer(games in pairs()) {
let batched = {
let mut h = history_from(&games);
h.converge().unwrap();
let report = h.converge().unwrap();
prop_assert!(
report.converged,
"batched side stopped at {} iterations with step {:?}; comparing \
two fits that have not converged measures truncation, not order",
report.iterations,
report.final_step
);
h
};
let incremental = {
let mut h = History::builder()
.convergence(ConvergenceOptions {
max_iter: 200,
max_iter: 20_000,
epsilon: 1e-10,
..ConvergenceOptions::default()
})
@@ -139,7 +154,13 @@ proptest! {
.unwrap();
}
h.converge().unwrap();
let report = h.converge().unwrap();
prop_assert!(
report.converged,
"incremental side stopped at {} iterations with step {:?}",
report.iterations,
report.final_step
);
h
};
+55 -9
View File
@@ -1,4 +1,4 @@
//! `quality()` beyond two rating groups.
//! `quality()` beyond two teams.
//!
//! The historical golden (two equal singletons) is asserted in
//! `src/lib.rs::tests::test_quality`. These cover the N-group generalisation,
@@ -82,14 +82,14 @@ fn uneven_group_sizes_work() {
}
#[test]
#[should_panic(expected = "at least 2 rating groups")]
#[should_panic(expected = "at least 2 teams")]
fn single_group_panics_with_clear_message() {
let r = rating(25.0, 3.0);
let _ = quality(&[&[r]], BETA);
}
#[test]
#[should_panic(expected = "at least 2 rating groups")]
#[should_panic(expected = "at least 2 teams")]
fn zero_groups_panics_with_clear_message() {
let _ = quality(&[], BETA);
}
@@ -108,13 +108,10 @@ fn history_predict_quality_supports_three_teams() {
let mut h = History::default();
h.record_winner(&"a", &"b", 1).unwrap();
h.record_winner(&"b", &"c", 2).unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
let q = h.predict_quality(&[&[&"a"], &[&"b"], &[&"c"]]).unwrap();
assert!(
q.is_finite(),
"3-team predict_quality must be finite, got {q}"
);
let q = h.quality(&[&[&"a"], &[&"b"], &[&"c"]]).unwrap();
assert!(q.is_finite(), "3-team quality must be finite, got {q}");
assert!((0.0..=1.0).contains(&q), "out of range: {q}");
}
@@ -164,3 +161,52 @@ fn quality_matches_the_reference_implementation() {
let refs: Vec<&[Gaussian]> = five.iter().map(Vec::as_slice).collect();
assert!((quality(&refs, beta) - 0.040).abs() < 1e-9);
}
/// `quality()` used to compute `det(ata) / det(middle)` in linear space. Both
/// are products of `k - 1` diagonal entries, so they leave `f64`'s range long
/// before their ratio does — and the ratio is the only thing the answer needs.
///
/// Measured before the fix: at the crate defaults 150 groups was correct, 200
/// returned `0`, and 250 returned `NaN` where the truth is `9.51e-88`. With a
/// small beta it bit sooner — `sigma = beta = 1e-3` returned `NaN` at 60 groups
/// against a true `1.32e-9`, a value that is entirely ordinary.
///
/// For `k` single-member groups with equal means the answer has a closed form,
/// `(beta / sqrt(beta^2 + sigma^2))^(k-1)`, so this checks against arithmetic
/// rather than against a recorded output.
#[test]
fn quality_matches_its_closed_form_past_the_overflow_point() {
for (sigma, beta) in [(25.0 / 3.0, 25.0 / 6.0), (1e-3, 1e-3), (50.0, 25.0 / 6.0)] {
let rating = vec![Gaussian::from_ms(25.0, sigma)];
for k in [2usize, 50, 60, 150, 200, 250, 300] {
let groups: Vec<&[Gaussian]> = (0..k).map(|_| rating.as_slice()).collect();
let got = quality(&groups, beta);
let expected = (beta / (beta * beta + sigma * sigma).sqrt()).powi(k as i32 - 1);
assert!(
got.is_finite(),
"sigma {sigma}, beta {beta}, {k} groups: got {got}"
);
// Subnormal results have no relative precision left to check.
if expected > f64::MIN_POSITIVE {
let rel = ((got - expected) / expected).abs();
assert!(
rel < 1e-11,
"sigma {sigma}, beta {beta}, {k} groups: got {got:e}, \
closed form {expected:e}, rel {rel:e}"
);
}
}
}
}
/// The overflow was in the intermediates, never in the answer: every value
/// above is an ordinary float. This pins the specific case that returned `NaN`
/// where the true answer is nine orders of magnitude inside the normal range.
#[test]
fn a_small_beta_does_not_overflow_at_sixty_groups() {
let rating = vec![Gaussian::from_ms(25.0, 1e-3)];
let groups: Vec<&[Gaussian]> = (0..60).map(|_| rating.as_slice()).collect();
let got = quality(&groups, 1e-3);
assert!((got - 1.317_089e-9).abs() / 1.317_089e-9 < 1e-6, "{got:e}");
}
+192
View File
@@ -0,0 +1,192 @@
//! `HistoryBuilder::default_rating_for`: configuring a *class* of competitors
//! rather than one at a time (#53).
//!
//! Every test carries a control — a key the rule does not match — so none can
//! pass by the rule firing for everybody, which would be indistinguishable
//! from changing the history defaults.
use trueskill_tt::{
ConstantDrift, Gaussian, History, HistoryBuilder, InferenceError, Member, NullObserver,
RatingRule, StartingPoint,
};
/// Pinned: no drift, and a tight prior at a known strength.
fn pinned() -> StartingPoint {
StartingPoint::new()
.prior(Gaussian::from_ms(5.0, 0.5))
.drift_scale(0.0)
}
fn play<R: RatingRule<&'static str>>(
h: &mut History<&'static str, i64, ConstantDrift, NullObserver, R>,
) {
for t in 1..=6 {
h.event(t)
.team(["layout_a"])
.team(["alice"])
.scores([3.0, 1.0])
.commit()
.expect("ingests");
}
h.converge().expect("converges");
}
#[test]
fn a_rule_configures_every_matching_key_without_naming_them() {
let mut ruled = History::builder()
.gamma(0.5)
.default_rating_for(|key: &&'static str| key.starts_with("layout_").then(pinned))
.build();
play(&mut ruled);
let mut plain = History::builder().gamma(0.5).build();
play(&mut plain);
let layout = ruled.current_skill("layout_a").expect("played");
// The rule pinned the layout: tight prior, no drift.
assert!(
layout.sigma() < 0.5,
"the layout should stay near its pinned prior, got sigma {}",
layout.sigma()
);
assert_ne!(
layout.sigma(),
plain.current_skill("layout_a").unwrap().sigma(),
"the rule must actually change the fit"
);
// The control is the *configuration*, not the posterior. Alice's posterior
// legitimately moves — she is playing a differently-configured opponent,
// and what she learns from beating it depends on how sure the model is
// about it. What must not move is what the rule was asked about.
let alice = ruled.rating("alice").expect("played");
assert_eq!(
alice.drift_scale(),
1.0,
"a non-matching key keeps the default drift"
);
assert_eq!(
(alice.prior().mu(), alice.prior().sigma()),
{
let p = plain.rating("alice").expect("played").prior();
(p.mu(), p.sigma())
},
"a non-matching key keeps the history's prior"
);
}
#[test]
fn a_rule_fires_for_a_competitor_first_seen_through_record_winner() {
// `record_winner` cannot carry configuration, which is the case a rule
// exists for.
let mut h = History::builder()
.default_rating_for(|key: &&'static str| key.starts_with("bot_").then(pinned))
.build();
h.record_winner(&"bot_1", &"human", 1).expect("ingests");
h.converge().expect("converges");
assert_eq!(h.rating("bot_1").expect("known").drift_scale(), 0.0);
assert_eq!(h.rating("human").expect("known").drift_scale(), 1.0);
}
#[test]
fn explicit_configuration_overrides_a_rule_field_by_field() {
let mut h = History::builder()
.default_rating_for(|_: &&'static str| Some(pinned()))
.build();
// Sets only the prior, so the rule's `drift_scale` must survive.
h.register(Member::new("a").with_prior(Gaussian::from_ms(-9.0, 2.0)))
.expect("new");
// Sets neither: the rule supplies both.
h.register(Member::new("b")).expect("new");
let a = h.rating("a").expect("registered");
assert_eq!(a.prior().mu(), -9.0, "explicit prior wins");
assert_eq!(a.drift_scale(), 0.0, "the rule's drift_scale survives");
let b = h.rating("b").expect("registered");
assert_eq!(b.prior().mu(), 5.0);
assert_eq!(b.drift_scale(), 0.0);
}
#[test]
fn two_explicit_declarations_that_disagree_are_still_an_error() {
// Precedence resolves rule-vs-explicit. It does not weaken the check
// between two explicit declarations, neither of which is more specific.
let mut h = History::builder()
.default_rating_for(|_: &&'static str| Some(pinned()))
.build();
let err = h
.add_events(vec![
event(1, "x", Gaussian::from_ms(1.0, 1.0)),
event(2, "x", Gaussian::from_ms(2.0, 1.0)),
])
.expect_err("two different priors for one competitor");
assert!(
matches!(err, InferenceError::ConflictingCompetitorConfig { .. }),
"{err:?}"
);
}
fn event(time: i64, key: &'static str, prior: Gaussian) -> trueskill_tt::Event<i64, &'static str> {
trueskill_tt::Event {
time,
teams: [
trueskill_tt::Team::with_members([Member::new(key).with_prior(prior)]),
trueskill_tt::Team::with_members([Member::new("opponent")]),
]
.into_iter()
.collect(),
outcome: trueskill_tt::Outcome::scores([2.0, 1.0]),
}
}
/// A named rule type, so the `History<..>` can be written down in a field.
struct StaticLayouts;
impl RatingRule<&'static str> for StaticLayouts {
fn starting_point(&self, key: &&'static str) -> Option<StartingPoint> {
key.starts_with("layout_").then(pinned)
}
}
/// The reason this is a trait rather than a bare `Fn` bound: a consumer holds
/// its history in application state and has to name the type.
struct Ladder {
history: History<&'static str, i64, ConstantDrift, NullObserver, StaticLayouts>,
}
#[test]
fn a_named_rule_type_can_be_stored_in_a_struct_field() {
let mut ladder = Ladder {
history: HistoryBuilder::default().rating_rule(StaticLayouts).build(),
};
play(&mut ladder.history);
assert!(
ladder
.history
.current_skill("layout_a")
.expect("played")
.sigma()
< 0.5
);
assert_eq!(
ladder
.history
.rating("alice")
.expect("played")
.drift_scale(),
1.0
);
}
#[test]
fn no_rule_is_the_default_and_costs_nothing_to_spell() {
// The whole point of defaulting the parameter: `History<K>` still works.
let h: History<String> = History::builder().key_type::<String>().build();
assert_eq!(h.competitor_count(), 0);
}
+169
View File
@@ -0,0 +1,169 @@
//! Converging, appending, and converging again must reach the same fixed point
//! as converging once over the whole event set.
//!
//! `tests/ingestion_equivalence.rs` covers a different question: it varies how
//! events are *batched* but converges only at the end. This file converges
//! between batches, which is the path a caller takes when it fits, serves for a
//! while, then ingests more.
//!
//! The property matters beyond ergonomics. It says `converge` reaches a fixed
//! point determined by the events, ratings and configuration alone — not by the
//! message state it started from. That is what makes a restored snapshot safe:
//! an inexact one cannot corrupt the answer, only cost an extra sweep. See #45.
use smallvec::smallvec;
use trueskill_tt::{ConvergenceOptions, Event, Gaussian, History, Member, Outcome, Team};
fn tight() -> ConvergenceOptions {
ConvergenceOptions {
max_iter: 5_000,
epsilon: 1e-12,
alpha: 1.0,
}
}
fn ev(a: &str, b: &str, time: i64) -> Event<i64, String> {
Event {
time,
teams: smallvec![
Team::with_members([Member::new(a.to_string())]),
Team::with_members([Member::new(b.to_string())]),
],
outcome: Outcome::winner(0, 2),
}
}
/// Ingest each chunk in turn, converging fully after every one.
fn fit_in_chunks(chunks: Vec<Events>) -> Vec<(String, Gaussian)> {
let mut h: History<String> = History::builder()
.key_type::<String>()
.convergence(tight())
.build();
for chunk in chunks {
h.add_events(chunk).unwrap();
let report = h.converge().unwrap();
assert!(
report.converged,
"a chunk failed to converge, so any comparison would be measuring \
truncation rather than the fixed point; final step {:?}",
report.final_step
);
}
let mut skills: Vec<(String, Gaussian)> = h
.learning_curves()
.into_iter()
.map(|(k, curve)| (k, curve.last().unwrap().1))
.collect();
skills.sort_by(|a, b| a.0.cmp(&b.0));
skills
}
fn assert_same(a: &[(String, Gaussian)], b: &[(String, Gaussian)], what: &str) {
assert_eq!(a.len(), b.len(), "{what}: competitor count differs");
for ((ka, ga), (kb, gb)) in a.iter().zip(b) {
assert_eq!(ka, kb, "{what}: key order differs");
// Measured: 6.2e-13 for a later append, 8.9e-11 for an interleaved one.
// The bar is well clear of both but far under anything that would let a
// genuine divergence through.
assert!(
(ga.mu() - gb.mu()).abs() < 1e-8 && (ga.sigma() - gb.sigma()).abs() < 1e-8,
"{what}: {ka} differs — one-shot mu={} sigma={}, chunked mu={} sigma={}",
ga.mu(),
ga.sigma(),
gb.mu(),
gb.sigma()
);
}
}
type Events = Vec<Event<i64, String>>;
/// Two chunks of events: the first at times 0..20, the second at 100..120.
fn fixture() -> (Events, Events) {
let names = ["a", "b", "c", "d", "e"];
let mut seed = 7u64;
let mut rnd = move || {
seed ^= seed << 13;
seed ^= seed >> 7;
seed ^= seed << 17;
seed
};
let (mut early, mut late) = (Vec::new(), Vec::new());
for t in 0..40i64 {
let i = (rnd() % 5) as usize;
let mut j = (rnd() % 5) as usize;
if j == i {
j = (j + 1) % 5;
}
if t < 20 {
early.push(ev(names[i], names[j], t));
} else {
late.push(ev(names[i], names[j], 100 + t));
}
}
(early, late)
}
/// The ordinary case: new events are strictly later than everything fitted.
#[test]
fn appending_later_events_matches_a_single_fit() {
let (early, late) = fixture();
let all: Vec<_> = early.iter().cloned().chain(late.iter().cloned()).collect();
assert_same(
&fit_in_chunks(vec![all]),
&fit_in_chunks(vec![early, late]),
"append strictly later",
);
}
/// The case the design question suspected might be weaker: appended events
/// interleave with slices that are already fitted, so the append legitimately
/// revises the past. It is not weaker — Through Time revises the past on every
/// converge regardless, so there is nothing special about doing it in two steps.
#[test]
fn appending_interleaved_events_matches_a_single_fit() {
let (early, late) = fixture();
let all: Vec<_> = early.iter().cloned().chain(late.iter().cloned()).collect();
// Split by parity so the second chunk is back-dated into the first's range.
let first: Vec<_> = all.iter().step_by(2).cloned().collect();
let second: Vec<_> = all.iter().skip(1).step_by(2).cloned().collect();
let together: Vec<_> = first
.iter()
.cloned()
.chain(second.iter().cloned())
.collect();
assert_same(
&fit_in_chunks(vec![together]),
&fit_in_chunks(vec![first, second]),
"append interleaved",
);
}
/// Converging an already-converged history is a no-op, which is what makes a
/// restored snapshot worth having: the work is skipped rather than redone.
#[test]
fn re_converging_an_unchanged_history_costs_one_iteration() {
let (early, late) = fixture();
let all: Vec<_> = early.into_iter().chain(late).collect();
let mut h: History<String> = History::builder()
.key_type::<String>()
.convergence(tight())
.build();
h.add_events(all).unwrap();
let first = h.converge().unwrap();
assert!(first.converged);
let again = h.converge().unwrap();
assert_eq!(
again.iterations, 1,
"a converged history should settle immediately, not re-grind"
);
assert!(again.converged);
}
+24 -16
View File
@@ -6,7 +6,7 @@ fn record_winner_builds_history() {
.mu(25.0)
.sigma(25.0 / 3.0)
.beta(25.0 / 6.0)
.drift(ConstantDrift(25.0 / 300.0))
.drift(ConstantDrift::new(25.0 / 300.0))
.convergence(ConvergenceOptions {
max_iter: 30,
epsilon: 1e-6,
@@ -15,26 +15,34 @@ fn record_winner_builds_history() {
.build();
h.record_winner(&"alice", &"bob", 1).unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
let a_idx = h.lookup(&"alice").unwrap();
let b_idx = h.lookup(&"bob").unwrap();
assert_ne!(a_idx, b_idx);
// `lookup` returned an `Index` that nothing public accepted, so the
// observable claim is the one worth making: two distinct competitors, each
// with their own posterior, and the winner ahead.
assert_eq!(h.competitor_count(), 2);
let alice = h.current_skill("alice").expect("alice played");
let bob = h.current_skill("bob").expect("bob played");
assert!(alice.mu() > bob.mu());
}
/// The same key names the same competitor across events, which is what
/// interning bought and the only part of it a caller can observe.
#[test]
fn intern_is_idempotent() {
fn a_repeated_key_is_one_competitor() {
let mut h: History = History::builder().build();
let a1 = h.intern(&"alice");
let a2 = h.intern(&"alice");
assert_eq!(a1, a2);
h.record_winner(&"alice", &"bob", 1).unwrap();
h.record_winner(&"alice", &"carol", 2).unwrap();
assert_eq!(h.competitor_count(), 3);
assert_eq!(h.learning_curve("alice").expect("known").len(), 2);
}
#[test]
fn lookup_returns_none_for_missing() {
fn an_unknown_key_is_unknown() {
let h: History = History::builder().build();
assert!(h.lookup(&"nobody").is_none());
assert!(h.current_skill("nobody").is_none());
assert!(h.learning_curve("nobody").is_none());
}
#[test]
@@ -43,13 +51,13 @@ fn record_draw_with_p_draw_set() {
.mu(25.0)
.sigma(25.0 / 3.0)
.beta(25.0 / 6.0)
.drift(ConstantDrift(25.0 / 300.0))
.drift(ConstantDrift::new(25.0 / 300.0))
.p_draw(0.25)
.build();
h.record_draw(&"alice", &"bob", 1).unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
assert!(h.lookup(&"alice").is_some());
assert!(h.lookup(&"bob").is_some());
assert!(h.current_skill("alice").is_some());
assert!(h.current_skill("bob").is_some());
}
+344
View File
@@ -0,0 +1,344 @@
//! Configuring a competitor before anything is observed about them.
//!
//! The configuration a competitor needs is usually a property of the domain —
//! "every layout is static" — not of whichever event happens to mention them
//! first. Stating it per-event meant every ingestion path had to remember it,
//! and two of the four paths could not state it at all.
use smallvec::smallvec;
use trueskill_tt::{
ConstantDrift, ConvergenceOptions, Event, Gaussian, History, InferenceError, Member, Outcome,
Team,
};
type H = History;
const PINNED: Gaussian = Gaussian::from_ms(2.0, 0.5);
fn history() -> H {
History::builder()
.mu(0.0)
.sigma(6.0)
.beta(1.0)
.score_sigma(2.0)
.drift(ConstantDrift::new(0.5))
.convergence(ConvergenceOptions {
max_iter: 20_000,
epsilon: 1e-13,
alpha: 1.0,
})
.build()
}
fn duel(
a: &'static str,
b: &'static str,
t: i64,
m: Option<Member<&'static str>>,
) -> Event<i64, &'static str> {
Event {
time: t,
teams: smallvec![
Team::with_members([Member::new(a)]),
Team::with_members([m.unwrap_or_else(|| Member::new(b))]),
],
outcome: Outcome::scores([5.0, 2.0]),
}
}
fn skills(h: &H) -> Vec<(&'static str, Gaussian)> {
["player", "layout"]
.into_iter()
.map(|k| (k, h.current_skill(&k).unwrap()))
.collect()
}
/// The headline contract.
#[test]
fn registering_matches_configuring_on_the_first_event() {
let configured = {
let mut h = history();
h.add_events(vec![
duel(
"player",
"layout",
1,
Some(
Member::new("layout")
.with_drift_scale(0.0)
.with_prior(PINNED),
),
),
duel("player", "layout", 2, None),
])
.unwrap();
let _ = h.converge().unwrap();
h
};
let registered = {
let mut h = history();
h.register(
Member::new("layout")
.with_drift_scale(0.0)
.with_prior(PINNED),
)
.unwrap();
h.add_events(vec![
duel("player", "layout", 1, None),
duel("player", "layout", 2, None),
])
.unwrap();
let _ = h.converge().unwrap();
h
};
for ((k, a), (_, b)) in skills(&configured).into_iter().zip(skills(&registered)) {
assert_eq!(a.mu(), b.mu(), "{k} mu");
assert_eq!(a.variance(), b.variance(), "{k} variance");
}
}
/// The case `EventBuilder` and the typed path cannot reach: a competitor whose
/// first appearance arrives through the two-argument convenience route.
#[test]
fn registration_reaches_a_competitor_first_seen_through_record_winner() {
let mut h = history();
h.register(
Member::new("layout")
.with_drift_scale(0.0)
.with_prior(PINNED),
)
.unwrap();
h.record_winner(&"player", &"layout", 1).unwrap();
h.record_winner(&"player", &"layout", 2).unwrap();
let _ = h.converge().unwrap();
let rating = h.rating(&"layout").unwrap();
assert_eq!(rating.drift_scale(), 0.0);
assert_eq!(rating.prior().mu(), PINNED.mu());
// Pinned means pinned: no drift across the two slices.
let curve = h.learning_curve(&"layout").unwrap();
assert!(curve.len() >= 2);
let widest = curve
.iter()
.map(|(_, g)| g.sigma())
.fold(f64::MIN, f64::max);
let narrowest = curve
.iter()
.map(|(_, g)| g.sigma())
.fold(f64::MAX, f64::min);
assert!(
(widest - narrowest) / widest < 1e-9,
"{narrowest} .. {widest}"
);
}
#[test]
fn registering_a_known_competitor_is_an_error() {
let mut h = history();
h.record_winner(&"player", &"layout", 1).unwrap();
let err = h.register(Member::new("layout")).unwrap_err();
assert!(
matches!(err, InferenceError::AlreadyRegistered { .. }),
"{err:?}"
);
}
#[test]
fn registering_twice_is_an_error() {
let mut h = history();
h.register(Member::new("layout").with_drift_scale(0.0))
.unwrap();
let err = h
.register(Member::new("layout").with_drift_scale(1.0))
.unwrap_err();
assert!(
matches!(err, InferenceError::AlreadyRegistered { .. }),
"{err:?}"
);
// The first registration stands.
assert_eq!(h.rating(&"layout").unwrap().drift_scale(), 0.0);
}
/// `weight` is per-event and meaningless here, so it is rejected rather than
/// dropped — dropping it silently is the defect class this whole area keeps
/// producing.
#[test]
fn a_weight_on_a_registration_is_rejected() {
let mut h = history();
let err = h
.register(Member::new("layout").with_weight(0.5))
.unwrap_err();
assert!(
matches!(
err,
InferenceError::InvalidParameter {
parameter: trueskill_tt::Parameter::Weight,
..
}
),
"{err:?}"
);
}
#[test]
fn an_invalid_drift_scale_on_a_registration_is_rejected() {
for bad in [-1.0, f64::NAN, f64::INFINITY] {
let mut h = history();
let err = h
.register(Member::new("layout").with_drift_scale(bad))
.unwrap_err();
assert!(
matches!(
err,
InferenceError::InvalidParameter {
parameter: trueskill_tt::Parameter::DriftScale,
..
}
),
"{bad}: {err:?}"
);
}
}
/// Registration makes the fit independent of the order events arrive in,
/// which is what the per-event shape could not guarantee.
#[test]
fn registration_makes_the_fit_order_independent() {
let build = |reversed: bool| {
let mut h = history();
h.register(
Member::new("layout")
.with_drift_scale(0.0)
.with_prior(PINNED),
)
.unwrap();
let mut events = vec![
duel("player", "layout", 1, None),
duel("player", "layout", 2, None),
duel("player", "layout", 3, None),
];
if reversed {
events.reverse();
}
h.add_events(events).unwrap();
let _ = h.converge().unwrap();
h
};
let forward = build(false);
let backward = build(true);
for ((k, a), (_, b)) in skills(&forward).into_iter().zip(skills(&backward)) {
assert_eq!(a.mu(), b.mu(), "{k} mu");
assert_eq!(a.variance(), b.variance(), "{k} variance");
}
}
/// `rating` is the read-back that made a configuration mistake detectable from
/// outside the crate at all. Every other accessor reports what inference
/// inferred; this reports what it was told.
#[test]
fn rating_reads_back_what_was_stored() {
let mut h = history();
assert!(h.rating(&"nobody").is_none());
h.register(
Member::new("layout")
.with_drift_scale(0.25)
.with_prior(PINNED),
)
.unwrap();
let r = h.rating(&"layout").unwrap();
assert_eq!(r.drift_scale(), 0.25);
assert_eq!(r.prior().mu(), PINNED.mu());
assert_eq!(r.prior().variance(), PINNED.variance());
// A competitor created by an event reports the history defaults.
h.record_winner(&"player", &"layout", 1).unwrap();
assert_eq!(h.rating(&"player").unwrap().drift_scale(), 1.0);
}
/// The decision this issue turned on: two different values for one competitor
/// are an error whether they arrive in one batch or two.
///
/// Last-write-wins across batches cut against the invariant
/// `tests/ingestion_equivalence.rs` protects — the same contradictory events
/// errored when batched and succeeded, order-dependently, one at a time.
mod conflicting_configuration {
use super::*;
fn seed(scale: f64) -> Event<i64, &'static str> {
duel(
"player",
"layout",
1,
Some(Member::new("layout").with_drift_scale(scale)),
)
}
#[test]
fn within_one_batch_is_an_error() {
let mut h = history();
let err = h.add_events(vec![seed(0.0), seed(1.0)]).unwrap_err();
assert!(
matches!(
err,
InferenceError::ConflictingCompetitorConfig {
field: trueskill_tt::CompetitorField::DriftScale,
..
}
),
"{err:?}"
);
}
#[test]
fn across_two_batches_is_also_an_error() {
let mut h = history();
h.add_events(vec![seed(0.0)]).unwrap();
let err = h.add_events(vec![seed(1.0)]).unwrap_err();
assert!(
matches!(
err,
InferenceError::ConflictingCompetitorConfig {
field: trueskill_tt::CompetitorField::DriftScale,
..
}
),
"{err:?}"
);
// Rejected before anything mutates: the first declaration stands.
assert_eq!(h.rating(&"layout").unwrap().drift_scale(), 0.0);
}
/// Repeating the *same* value stays inert, which is the expected shape
/// when the configuration is a property of the domain.
#[test]
fn repeating_the_same_value_is_inert() {
let mut h = history();
h.add_events(vec![seed(0.0)]).unwrap();
h.add_events(vec![seed(0.0)]).unwrap();
assert_eq!(h.rating(&"layout").unwrap().drift_scale(), 0.0);
}
/// A registration and a later event that agree are fine; one that
/// disagrees is the same error.
#[test]
fn a_registration_conflicts_with_a_later_event() {
let mut h = history();
h.register(Member::new("layout").with_drift_scale(0.0))
.unwrap();
h.add_events(vec![seed(0.0)]).unwrap();
let mut h2 = history();
h2.register(Member::new("layout").with_drift_scale(0.0))
.unwrap();
let err = h2.add_events(vec![seed(1.0)]).unwrap_err();
assert!(
matches!(err, InferenceError::ConflictingCompetitorConfig { .. }),
"{err:?}"
);
}
}
+3 -3
View File
@@ -9,7 +9,7 @@ fn scored_two_team_one_event_pulls_winner_up() {
.mu(0.0)
.sigma(2.0)
.beta(1.0)
.drift(ConstantDrift(0.0))
.drift(ConstantDrift::new(0.0))
.score_sigma(1.0)
.build();
@@ -46,7 +46,7 @@ fn scored_zero_margin_treats_as_tie() {
.mu(0.0)
.sigma(2.0)
.beta(1.0)
.drift(ConstantDrift(0.0))
.drift(ConstantDrift::new(0.0))
.score_sigma(1.0)
.build();
@@ -88,7 +88,7 @@ fn scored_three_team_partial_order() {
.mu(0.0)
.sigma(2.0)
.beta(1.0)
.drift(ConstantDrift(0.0))
.drift(ConstantDrift::new(0.0))
.score_sigma(1.0)
.build();
+178
View File
@@ -0,0 +1,178 @@
//! What a sparse factorisation of the joint would actually buy (#52).
//!
//! Run explicitly:
//!
//! ```text
//! cargo test --release --features approx,measure-sparsity \
//! --test sparsity_measurement -- --ignored --nocapture
//! ```
//!
//! The whole file is gated: it reaches for the joint's sparsity pattern, which
//! is exposed only under `measure-sparsity`.
#![cfg(feature = "measure-sparsity")]
use std::collections::HashSet;
use trueskill_tt::{ConvergenceOptions, History};
/// A history shaped like the issue's fixture: many slices, scored duels,
/// competitors reappearing across slices so the drift links are long.
fn fitted(slices: i64, duels: usize, competitors: usize) -> History<String> {
let mut h: History<String> = History::builder()
.key_type::<String>()
.mu(0.0)
.sigma(6.0)
.beta(1.0)
.score_sigma(2.0)
.gamma(0.05)
.convergence(ConvergenceOptions {
max_iter: trueskill_tt::ITERATIONS,
epsilon: 1e-8,
alpha: 1.0,
})
.build();
let mut k = 0usize;
for t in 0..slices {
for _ in 0..duels {
k += 1;
h.event(t)
.team([format!("p{}", k % competitors)])
.team([format!("p{}", (k + 37) % competitors)])
.scores([
(k as f64 * 0.3).sin().abs() * 20.0,
(k as f64 * 0.3).cos().abs() * 20.0,
])
.commit()
.expect("ingests");
}
}
h.converge().expect("converges");
h
}
/// Symbolic Cholesky by row-merge: returns (nnz(L), flops).
///
/// Fill-in is simulated directly — for each column, the set of rows below the
/// diagonal that are nonzero — which is exact and easily checked, at the cost
/// of being O(n * nnz(L)) rather than the linear elimination-tree method.
fn symbolic(n: usize, adj: &[HashSet<usize>], perm_of: &[usize]) -> (usize, f64) {
// `perm_of[old] = new`. Build the permuted lower-triangle pattern.
let mut cols: Vec<HashSet<usize>> = vec![HashSet::new(); n];
for (old, nbrs) in adj.iter().enumerate() {
let i = perm_of[old];
for &old_j in nbrs {
let j = perm_of[old_j];
if j < i {
cols[j].insert(i);
}
}
}
let mut nnz = 0usize;
let mut flops = 0.0f64;
for j in 0..n {
// Column j's pattern is final once every earlier column has merged in.
let rows: Vec<usize> = cols[j].iter().copied().collect();
let c = rows.len();
nnz += c + 1; // below-diagonal entries plus the diagonal
// Cholesky work for this column: one outer product over its pattern.
flops += (c as f64 + 1.0) * (c as f64 + 1.0);
// Fill-in: every pair in column j becomes an edge in the remaining graph.
for (a_idx, &a) in rows.iter().enumerate() {
for &b in &rows[a_idx + 1..] {
let (lo, hi) = if a < b { (a, b) } else { (b, a) };
cols[lo].insert(hi);
}
}
}
(nnz, flops)
}
#[test]
#[ignore = "measurement, run explicitly"]
fn what_sparsity_would_buy() {
for (slices, duels, competitors) in [(30, 8, 100), (76, 13, 200)] {
let h = fitted(slices, duels, competitors);
let (n, pattern) = h.joint_pattern_for_measurement();
let nnz_a: usize = pattern.iter().map(HashSet::len).sum::<usize>() + n;
let dense_flops = (n as f64).powi(3) / 3.0;
let natural: Vec<usize> = (0..n).collect();
let (nnz_nat, flops_nat) = symbolic(n, &pattern, &natural);
// AMD returns `perm[new] = old`; invert it.
let (col_ptr, row_idx) = csc(n, &pattern);
let p = feral_amd::amd_order(
&feral_amd::CscPattern::new(n, &col_ptr, &row_idx).expect("valid pattern"),
)
.expect("amd");
let mut perm_of = vec![0usize; n];
for (new, &old) in p.iter().enumerate() {
perm_of[old as usize] = new;
}
let (nnz_amd, flops_amd) = symbolic(n, &pattern, &perm_of);
println!(
"\n=== {slices} slices x {duels} duels, {competitors} competitors ===\n\
n = {n}\n\
nnz(A) = {nnz_a} ({:.4}% dense)\n\
dense flops = {:.3e}\n\
nnz(L) natural = {nnz_nat} flops = {:.3e} ({:.1}x vs dense)\n\
nnz(L) AMD = {nnz_amd} flops = {:.3e} ({:.1}x vs dense)",
100.0 * nnz_a as f64 / (n * n) as f64,
dense_flops,
flops_nat,
dense_flops / flops_nat,
flops_amd,
dense_flops / flops_amd,
);
}
}
/// Full symmetric pattern to CSC, as `feral-amd` wants it.
fn csc(n: usize, adj: &[HashSet<usize>]) -> (Vec<i32>, Vec<i32>) {
let mut col_ptr = Vec::with_capacity(n + 1);
let mut row_idx = Vec::new();
col_ptr.push(0i32);
for (j, nbrs) in adj.iter().enumerate() {
let mut rows: Vec<i32> = nbrs.iter().map(|&i| i as i32).collect();
rows.push(j as i32);
rows.sort_unstable();
rows.dedup();
row_idx.extend_from_slice(&rows);
col_ptr.push(row_idx.len() as i32);
}
(col_ptr, row_idx)
}
/// End-to-end factorisation time at the scale #52 was opened about.
#[test]
#[ignore = "measurement, run explicitly"]
fn factorisation_time_at_scale() {
use std::time::Instant;
for (slices, duels, competitors) in [(30, 8, 100), (76, 13, 200), (150, 26, 400)] {
let h = fitted(slices, duels, competitors);
let (n, _) = h.joint_pattern_for_measurement();
// Warm, then time.
let _ = h.joint().expect("scored history");
let t = Instant::now();
let joint = h.joint().expect("scored history");
let factor = t.elapsed();
let a = "p0".to_string();
let b = "p1".to_string();
let t = Instant::now();
let _ = joint.posterior_of(&[(&a, 1.0), (&b, -1.0)]).expect("known");
let query = t.elapsed();
println!(
"n = {n:5} factorise = {factor:>12?} query = {query:>10?} \
(dense was O(n^3): {:.3e} flops)",
(n as f64).powi(3) / 3.0
);
}
}
+152
View File
@@ -0,0 +1,152 @@
//! The `Time` generic, exercised end to end.
//!
//! `History<T: Time, ..>` has always been generic over the time axis, `Untimed`
//! has always been exported, and `Drift<T>` is generic specifically so that
//! "seasonal or calendar-aware drift is expressible without going through
//! `i64`". None of it was reachable: every construction route pinned `T = i64`,
//! `HistoryBuilder`'s fields are private, and its `Default` existed only for the
//! `i64` instantiation.
//!
//! Nothing in the repository constructed a non-`i64` history, which is why that
//! went unnoticed. This file is the guard against it recurring — it is as much
//! about the generic being *exercised* as about any single assertion.
use trueskill_tt::{ConstantDrift, Drift, History, HistoryBuilder, Time, Untimed};
/// A domain time type: a season number. Exactly what the `Time` trait exists
/// to support, and what a consumer with `chrono` dates would write.
#[derive(Copy, Clone, Debug, PartialEq, Eq, PartialOrd, Ord)]
struct Season(u16);
impl Time for Season {
fn elapsed_to(&self, later: &Self) -> i64 {
i64::from(later.0.saturating_sub(self.0))
}
}
/// Drift that only accumulates between seasons, not within one — the
/// calendar-aware case the trait's own docs cite.
#[derive(Copy, Clone, Debug)]
struct SeasonalDrift {
per_season: f64,
}
impl Drift<Season> for SeasonalDrift {
fn variance_delta(&self, from: &Season, to: &Season) -> f64 {
self.variance_for_elapsed(from.elapsed_to(to))
}
fn variance_for_elapsed(&self, elapsed: i64) -> f64 {
elapsed.max(0) as f64 * self.per_season * self.per_season
}
}
#[test]
fn an_untimed_history_fits_through_the_builder() {
let mut h = History::builder().time_type::<Untimed>().build();
for _ in 0..5 {
h.record_winner(&"alice", &"bob", Untimed).unwrap();
}
assert!(h.converge().unwrap().converged);
let alice = h.current_skill(&"alice").unwrap();
let bob = h.current_skill(&"bob").unwrap();
assert!(alice.mu() > bob.mu(), "{alice:?} vs {bob:?}");
assert!(alice.sigma().is_finite() && alice.sigma() > 0.0);
}
/// `Untimed::elapsed_to` is always 0, so no drift accumulates however many
/// events there are. That is the property the type exists for, and it had never
/// been checked.
#[test]
fn untimed_accumulates_no_drift() {
fn final_sigma<T: Time + Copy>(time: T, drift: ConstantDrift) -> f64 {
let mut h = History::builder().time_type::<T>().drift(drift).build();
for _ in 0..8 {
h.record_winner(&"a", &"b", time).unwrap();
}
let _ = h.converge().unwrap();
h.current_skill(&"a").unwrap().sigma()
}
// Under Untimed the drift setting cannot matter, because elapsed is always 0.
let none = final_sigma(Untimed, ConstantDrift::new(0.0));
let large = final_sigma(Untimed, ConstantDrift::new(5.0));
assert_eq!(
none.to_bits(),
large.to_bits(),
"Untimed must ignore drift entirely: {none} vs {large}"
);
}
#[test]
fn a_custom_time_type_and_a_custom_drift_work_together() {
let mut h = History::builder()
.time_type::<Season>()
.drift(SeasonalDrift { per_season: 0.5 })
.build();
for season in 1..=4u16 {
for _ in 0..3 {
h.record_winner(&"veteran", &"rookie", Season(season))
.unwrap();
}
}
assert!(h.converge().unwrap().converged);
let curve = h.learning_curve(&"veteran").unwrap();
assert_eq!(curve.len(), 4, "one point per season: {curve:?}");
for (season, g) in &curve {
assert!(
g.mu().is_finite() && g.sigma() > 0.0,
"season {season:?}: {g:?}"
);
}
// Times come back as the domain type, not as an integer.
assert_eq!(curve[0].0, Season(1));
assert_eq!(curve[3].0, Season(4));
}
/// Seasonal drift must actually widen a gap across seasons — otherwise the
/// custom `Drift` is being ignored and the test above would pass regardless.
#[test]
fn a_custom_drift_is_actually_consulted() {
fn sigma_with(per_season: f64) -> f64 {
let mut h = History::builder()
.time_type::<Season>()
.drift(SeasonalDrift { per_season })
.build();
for season in 1..=6u16 {
h.record_winner(&"a", &"b", Season(season)).unwrap();
}
let _ = h.converge().unwrap();
h.current_skill(&"a").unwrap().sigma()
}
let still = sigma_with(0.0);
let drifting = sigma_with(2.0);
assert!(
drifting > still * 1.05,
"a drifting fit must be less certain: {drifting} vs {still}"
);
}
/// The other axis: a custom key type, through the same mechanism.
#[test]
fn key_type_replaces_builder_with_key() {
let mut h = History::builder().key_type::<String>().build();
h.record_winner(&"alice".to_string(), &"bob".to_string(), 1)
.unwrap();
assert!(h.converge().unwrap().converged);
assert!(h.current_skill("alice").is_some());
}
/// Both axes at once, via the explicit constructor rather than the setters.
#[test]
fn new_constructs_on_any_axis_directly() {
let mut h = HistoryBuilder::<String, Season>::new().build();
h.record_winner(&"a".to_string(), &"b".to_string(), Season(7))
.unwrap();
assert!(h.converge().unwrap().converged);
assert_eq!(h.learning_curve("a").unwrap()[0].0, Season(7));
}
+337
View File
@@ -0,0 +1,337 @@
//! The joint must span slices, because Through Time reads each competitor at
//! their own last appearance.
//!
//! The exact posterior of a multi-slice scored history is still Gaussian: the
//! prior, the drift between appearances, and the scored likelihoods are all
//! Gaussian. So it can be written out by hand and compared against, which is
//! the check a single-slice fixture cannot make.
use smallvec::smallvec;
use trueskill_tt::{
ConstantDrift, ConvergenceOptions, Event, History, Member, Outcome, Team, UnknownKeys,
};
const SIGMA0: f64 = 6.0;
const BETA: f64 = 1.0;
const SCORE_SIGMA: f64 = 2.0;
const GAMMA: f64 = 0.5;
type H = History;
fn history(gamma: f64) -> H {
History::builder()
.mu(0.0)
.sigma(SIGMA0)
.beta(BETA)
.score_sigma(SCORE_SIGMA)
.drift(ConstantDrift::new(gamma))
.unknown_keys(UnknownKeys::Reject)
.convergence(ConvergenceOptions {
max_iter: 20_000,
epsilon: 1e-13,
alpha: 1.0,
})
.build()
}
fn duel(a: &'static str, b: &'static str, t: i64, sa: f64, sb: f64) -> Event<i64, &'static str> {
Event {
time: t,
teams: smallvec![
Team::with_members([Member::new(a)]),
Team::with_members([Member::new(b)]),
],
outcome: Outcome::scores([sa, sb]),
}
}
fn inverse(mut a: Vec<Vec<f64>>) -> Vec<Vec<f64>> {
let n = a.len();
let mut inv: Vec<Vec<f64>> = (0..n)
.map(|i| (0..n).map(|j| f64::from(u8::from(i == j))).collect())
.collect();
for col in 0..n {
let mut piv = col;
for r in col + 1..n {
if a[r][col].abs() > a[piv][col].abs() {
piv = r;
}
}
a.swap(col, piv);
inv.swap(col, piv);
let d = a[col][col];
for j in 0..n {
a[col][j] /= d;
inv[col][j] /= d;
}
for r in 0..n {
if r == col {
continue;
}
let f = a[r][col];
for j in 0..n {
a[r][j] -= f * a[col][j];
inv[r][j] -= f * inv[col][j];
}
}
}
inv
}
/// Two competitors, two slices ten units apart, one duel in each.
///
/// The exact precision is written out explicitly here rather than obtained
/// from the crate, so this is an independent check rather than a restatement.
/// Variables are `[a0, b0, a1, b1]`.
#[test]
fn a_two_slice_joint_matches_the_exact_posterior() {
let mut h = history(GAMMA);
h.add_events(vec![
duel("a", "b", 0, 5.0, 2.0),
duel("a", "b", 10, 4.0, 3.0),
])
.unwrap();
let report = h.converge().unwrap();
assert!(report.converged, "{:?}", report.final_step);
let prior_prec = 1.0 / (SIGMA0 * SIGMA0);
let drift_prec = 1.0 / (10.0 * GAMMA * GAMMA);
let obs_prec = 1.0 / (SCORE_SIGMA * SCORE_SIGMA + 2.0 * BETA * BETA);
let mut lambda = vec![vec![0.0; 4]; 4];
// priors on the first appearances
lambda[0][0] += prior_prec;
lambda[1][1] += prior_prec;
// drift a0-a1 and b0-b1
for (p, q) in [(0usize, 2usize), (1, 3)] {
lambda[p][p] += drift_prec;
lambda[q][q] += drift_prec;
lambda[p][q] -= drift_prec;
lambda[q][p] -= drift_prec;
}
// one duel per slice: contrast (+1, -1) on that slice's variables
for (p, q) in [(0usize, 1usize), (2, 3)] {
lambda[p][p] += obs_prec;
lambda[q][q] += obs_prec;
lambda[p][q] -= obs_prec;
lambda[q][p] -= obs_prec;
}
let cov = inverse(lambda);
// The crate reads each competitor at their latest appearance: a1, b1.
let exact_gap = (cov[2][2] + cov[3][3] - 2.0 * cov[2][3]).sqrt();
let got = h
.joint()
.unwrap()
.posterior_of(&[(&"a", 1.0), (&"b", -1.0)])
.unwrap();
assert!(
(got.sigma() - exact_gap).abs() / exact_gap < 1e-9,
"difference: got {} exact {exact_gap}",
got.sigma()
);
let exact_single = cov[2][2].sqrt();
let got_single = h.joint().unwrap().posterior_of(&[(&"a", 1.0)]).unwrap();
assert!(
(got_single.sigma() - exact_single).abs() / exact_single < 1e-9,
"single node: got {} exact {exact_single}",
got_single.sigma()
);
}
/// The case that motivated this: competitors read at *different* slices, with
/// the last slice holding only one of them. Under the old latest-slice joint
/// this was `UnknownKey`.
#[test]
fn competitors_last_seen_in_different_slices_are_comparable() {
let mut h = history(GAMMA);
h.add_events(vec![
duel("a", "b", 0, 5.0, 2.0),
duel("a", "c", 10, 4.0, 3.0),
// the final slice holds one duel that does not involve b at all
duel("a", "c", 20, 6.0, 1.0),
])
.unwrap();
let _ = h.converge().unwrap();
// b last appeared at time 0; a and c at time 20. All three must resolve.
for (x, y) in [("a", "b"), ("b", "c"), ("a", "c")] {
let g = h
.joint()
.unwrap()
.posterior_of(&[(&x, 1.0), (&y, -1.0)])
.unwrap_or_else(|e| panic!("{x} - {y} should resolve across slices: {e}"));
assert!(g.sigma() > 0.0 && g.sigma().is_finite());
}
}
/// The mean must agree with what message passing reports, which is exact even
/// with cycles. Only the second moment needs the joint.
#[test]
fn means_agree_with_the_marginals() {
let mut h = history(GAMMA);
h.add_events(vec![
duel("a", "b", 0, 5.0, 2.0),
duel("b", "c", 5, 3.0, 1.0),
duel("a", "c", 10, 4.0, 2.0),
])
.unwrap();
let _ = h.converge().unwrap();
for k in ["a", "b", "c"] {
let marginal = h.current_skill(&k).unwrap().mu();
let joint = h.joint().unwrap().posterior_of(&[(&k, 1.0)]).unwrap().mu();
assert!(
(marginal - joint).abs() < 1e-9,
"{k}: marginal {marginal}, joint {joint}"
);
}
}
/// With zero drift a competitor has one latent skill however many slices it
/// appears in, so spreading the same events over time must not change the
/// answer. This exercises the appearance-merging path.
#[test]
fn zero_drift_makes_slice_layout_irrelevant() {
let spread = {
let mut h = history(0.0);
h.add_events(vec![
duel("a", "b", 0, 5.0, 2.0),
duel("a", "b", 10, 4.0, 3.0),
duel("a", "b", 20, 6.0, 1.0),
])
.unwrap();
let _ = h.converge().unwrap();
h.joint()
.unwrap()
.posterior_of(&[(&"a", 1.0), (&"b", -1.0)])
.unwrap()
};
let together = {
let mut h = history(0.0);
h.add_events(vec![
duel("a", "b", 0, 5.0, 2.0),
duel("a", "b", 0, 4.0, 3.0),
duel("a", "b", 0, 6.0, 1.0),
])
.unwrap();
let _ = h.converge().unwrap();
h.joint()
.unwrap()
.posterior_of(&[(&"a", 1.0), (&"b", -1.0)])
.unwrap()
};
assert!(
(spread.sigma() - together.sigma()).abs() < 1e-9,
"zero drift: spread {} vs together {}",
spread.sigma(),
together.sigma()
);
}
/// More drift means less is carried forward from old evidence, so a comparison
/// against a competitor last seen long ago must widen.
#[test]
fn drift_widens_a_comparison_across_time() {
let mut previous = 0.0;
for gamma in [0.0f64, 0.1, 0.5, 2.0] {
let mut h = history(gamma);
h.add_events(vec![
duel("a", "b", 0, 5.0, 2.0),
duel("a", "c", 100, 4.0, 3.0),
])
.unwrap();
let _ = h.converge().unwrap();
// b was last seen at time 0; a at time 100.
let g = h
.joint()
.unwrap()
.posterior_of(&[(&"a", 1.0), (&"b", -1.0)])
.unwrap();
assert!(
g.sigma() > previous,
"gamma={gamma}: sigma {} did not exceed {previous}",
g.sigma()
);
previous = g.sigma();
}
}
/// `posterior_of_at` pins the reading to a moment, where `posterior_of` takes
/// each competitor wherever they were last seen.
#[test]
fn posterior_of_at_reads_as_of_a_time() {
let mut h = history(GAMMA);
h.add_events(vec![
duel("a", "b", 0, 5.0, 2.0),
duel("a", "b", 10, 4.0, 3.0),
duel("a", "b", 20, 6.0, 1.0),
])
.unwrap();
let _ = h.converge().unwrap();
let early = h
.joint()
.unwrap()
.posterior_of_at(0, &[(&"a", 1.0), (&"b", -1.0)])
.unwrap();
let late = h
.joint()
.unwrap()
.posterior_of_at(20, &[(&"a", 1.0), (&"b", -1.0)])
.unwrap();
let latest = h
.joint()
.unwrap()
.posterior_of(&[(&"a", 1.0), (&"b", -1.0)])
.unwrap();
// Asking as of the final slice is the same as asking for the latest.
assert!((late.mu() - latest.mu()).abs() < 1e-9);
assert!((late.sigma() - latest.sigma()).abs() < 1e-9);
// Reading at time 0 is a different quantity, and the smoothed estimate
// there is informed by everything that came after.
assert!(
(early.mu() - late.mu()).abs() > 1e-6,
"as-of-0 and as-of-20 should differ: {} vs {}",
early.mu(),
late.mu()
);
// A time before any event has nothing to read.
assert!(
h.joint()
.unwrap()
.posterior_of_at(-1, &[(&"a", 1.0)])
.is_err()
);
}
/// Times between slices resolve to the latest appearance at or before them.
#[test]
fn a_time_between_slices_reads_the_previous_appearance() {
let mut h = history(GAMMA);
h.add_events(vec![
duel("a", "b", 0, 5.0, 2.0),
duel("a", "b", 100, 4.0, 3.0),
])
.unwrap();
let _ = h.converge().unwrap();
let at_zero = h
.joint()
.unwrap()
.posterior_of_at(0, &[(&"a", 1.0)])
.unwrap();
let between = h
.joint()
.unwrap()
.posterior_of_at(50, &[(&"a", 1.0)])
.unwrap();
assert!((at_zero.mu() - between.mu()).abs() < 1e-12);
assert!((at_zero.sigma() - between.sigma()).abs() < 1e-12);
}
+96
View File
@@ -0,0 +1,96 @@
//! The traits a consumer needs on the public types, pinned so they cannot be
//! removed by accident.
//!
//! This is written from a consumer's position — deriving `Debug` on a struct
//! that *holds* a `History` — because that is the thing that failed. Asserting
//! `History: Debug` in isolation would not have caught the generic-bound half:
//! `Rating` derives `PartialEq`, but that is only usable if `D: PartialEq`, and
//! the crate's own only `Drift` impl did not satisfy it.
use trueskill_tt::{
ConstantDrift, ConvergenceOptions, ConvergenceReport, Event, GameOptions, Gaussian, History,
HistoryBuilder, InferenceError, Member, Outcome, Rating, Team,
};
/// The reported failure, verbatim: a consumer holding a history in app state.
#[derive(Debug)]
#[allow(
dead_code,
reason = "held only so `derive(Debug)` has something to render"
)]
struct App {
history: History,
}
#[test]
fn a_struct_holding_a_history_can_derive_debug() {
let app = App {
history: History::default(),
};
let rendered = format!("{app:?}");
// Summarising, not a dump of every skill store — the same choice `Joint`'s
// manual `Debug` makes about its n² factorisation.
assert!(rendered.contains("competitors"), "{rendered}");
assert!(rendered.contains("time_slices"), "{rendered}");
assert!(
!rendered.contains("SkillStore"),
"History's Debug should summarise, not dump: {rendered}"
);
}
#[test]
fn history_builder_is_debug_and_clone() {
let b: HistoryBuilder = History::builder();
let cloned = b.clone();
assert!(!format!("{cloned:?}").is_empty());
}
#[test]
fn config_and_input_value_types_are_comparable() {
assert_eq!(ConstantDrift::new(0.1), ConstantDrift::new(0.1));
assert_ne!(ConstantDrift::new(0.1), ConstantDrift::new(0.2));
assert_eq!(ConvergenceOptions::default(), ConvergenceOptions::default());
assert_eq!(GameOptions::default(), GameOptions::default());
// `Rating: PartialEq` is only reachable through `D: PartialEq`.
assert_eq!(Rating::<i64, ConstantDrift>::default(), Rating::default());
assert_ne!(
Rating::default(),
Rating::<i64, ConstantDrift>::default().with_drift_scale(2.0)
);
assert_eq!(Member::new("a"), Member::new("a"));
assert_ne!(Member::new("a"), Member::new("b"));
assert_eq!(
Team::with_members([Member::new("a")]),
Team::with_members([Member::new("a")])
);
let event = || Event {
time: 1,
teams: [
Team::with_members([Member::new("a")]),
Team::with_members([Member::new("b")]),
]
.into_iter()
.collect(),
outcome: Outcome::winner(0, 2),
};
assert_eq!(event(), event());
assert_eq!(Gaussian::default(), Gaussian::default());
}
#[test]
fn a_history_is_send_and_sync_and_default() {
fn assert_send_sync<X: Send + Sync>() {}
assert_send_sync::<History>();
assert_send_sync::<InferenceError>();
let mut h = History::default();
let report: ConvergenceReport = h.converge().expect("an empty history converges");
assert_eq!(report, report.clone());
}
+220 -6
View File
@@ -22,7 +22,7 @@ fn rating() -> R {
R::new(
Gaussian::from_ms(25.0, 25.0 / 3.0),
25.0 / 6.0,
ConstantDrift(0.0),
ConstantDrift::new(0.0),
)
}
@@ -49,7 +49,13 @@ fn ranked_rejects_a_zero_damping_factor() {
)
.expect_err("alpha = 0 must be rejected");
assert!(
matches!(err, InferenceError::InvalidParameter { name: "alpha", .. }),
matches!(
err,
InferenceError::InvalidParameter {
parameter: trueskill_tt::Parameter::Alpha,
..
}
),
"got {err:?}"
);
}
@@ -65,7 +71,13 @@ fn ranked_rejects_an_out_of_range_damping_factor() {
)
.expect_err("alpha out of (0, 1] must be rejected");
assert!(
matches!(err, InferenceError::InvalidParameter { name: "alpha", .. }),
matches!(
err,
InferenceError::InvalidParameter {
parameter: trueskill_tt::Parameter::Alpha,
..
}
),
"alpha={alpha}: got {err:?}"
);
}
@@ -81,7 +93,13 @@ fn scored_rejects_a_bad_damping_factor() {
)
.expect_err("alpha = 0 must be rejected");
assert!(
matches!(err, InferenceError::InvalidParameter { name: "alpha", .. }),
matches!(
err,
InferenceError::InvalidParameter {
parameter: trueskill_tt::Parameter::Alpha,
..
}
),
"got {err:?}"
);
}
@@ -139,7 +157,7 @@ fn ingestion_rejects_a_tie_without_a_draw_probability() {
);
}
/// `Outcome::scores_with_sigma` documents that a non-positive sigma is
/// `Outcome::scores_with_noise` documents that a non-positive sigma is
/// accepted at construction and rejected at ingestion.
#[test]
fn ingestion_rejects_a_non_positive_per_event_score_sigma() {
@@ -152,7 +170,7 @@ fn ingestion_rejects_a_non_positive_per_event_score_sigma() {
Team::with_members([Member::new("a")]),
Team::with_members([Member::new("b")]),
],
outcome: Outcome::scores_with_sigma([21.0, 9.0], sigma),
outcome: Outcome::scores_with_noise([21.0, 9.0], sigma),
}])
.expect_err("a non-positive per-event sigma must be rejected");
assert!(
@@ -184,3 +202,199 @@ fn ingestion_rejects_weights_that_do_not_match_their_team() {
"got {err:?}"
);
}
/// `mu`, `sigma` and `beta` were the last unvalidated setters on
/// `HistoryBuilder`, next to `p_draw`, `score_sigma` and `convergence`, which
/// all assert eagerly.
///
/// Two of the rejected values are the quiet kind. A negative `sigma` or `beta`
/// enters inference only as its square, so it produced bit-identical results
/// to the positive value — the sign was dropped without comment.
mod builder_parameters {
use trueskill_tt::History;
#[test]
#[should_panic(expected = "mu must be finite")]
fn a_non_finite_mu_is_rejected() {
let _ = History::builder().mu(f64::NAN);
}
#[test]
#[should_panic(expected = "sigma must be finite and positive")]
fn a_zero_sigma_is_rejected() {
let _ = History::builder().sigma(0.0);
}
#[test]
#[should_panic(expected = "sigma must be finite and positive")]
fn a_negative_sigma_is_rejected() {
let _ = History::builder().sigma(-8.33);
}
#[test]
#[should_panic(expected = "sigma must be finite and positive")]
fn an_infinite_sigma_is_rejected() {
let _ = History::builder().sigma(f64::INFINITY);
}
#[test]
#[should_panic(expected = "beta must be finite and non-negative")]
fn a_negative_beta_is_rejected() {
let _ = History::builder().beta(-4.17);
}
#[test]
#[should_panic(expected = "beta must be finite and non-negative")]
fn a_non_finite_beta_is_rejected() {
let _ = History::builder().beta(f64::NAN);
}
/// Zero beta is deliberately allowed: performance is then exactly skill.
/// It has to reach a different fit than a positive beta, or "allowed"
/// would just mean "not checked".
#[test]
fn a_zero_beta_is_allowed_and_changes_the_fit() {
let fit = |beta: f64| {
let mut h = History::builder()
.mu(25.0)
.sigma(25.0 / 3.0)
.beta(beta)
.build();
h.record_winner(&"a", &"b", 1).unwrap();
let _ = h.converge().unwrap();
h.current_skill(&"a").unwrap()
};
let zero = fit(0.0);
let positive = fit(25.0 / 6.0);
// `variance` rather than `pi`: the natural parameters are the crate's
// internal representation and no longer public. It is the same
// quantity inverted, so a finite positive precision is a finite
// positive variance.
assert!(zero.variance().is_finite() && zero.variance() > 0.0);
assert!(
(zero.variance() - positive.variance()).abs() > 1e-6,
"zero beta must not merely be ignored: {zero:?} vs {positive:?}"
);
}
}
/// The constructors below `HistoryBuilder`, which 0.8.0's validation did not
/// reach.
///
/// `sigma`, `beta` and `gamma` all enter inference only as squares, so a
/// negative value behaves as its absolute value and the sign vanishes without
/// comment. Measured before these guards: `from_ms(25.0, -8.33)` and
/// `Rating::new(_, -4.17, _)` returned results bit identical to their positive
/// counterparts, and `Rating::new(_, NaN, _)` reached `Game::ranked`, which
/// returned `Ok` carrying `Gaussian { pi: NaN, tau: NaN }`.
mod constructor_parameters {
use trueskill_tt::{ConstantDrift, Gaussian, History, InferenceError, Rating};
#[test]
#[should_panic(expected = "sigma must not be negative")]
fn a_negative_sigma_is_rejected_by_from_ms() {
let _ = Gaussian::from_ms(25.0, -8.33);
}
/// NaN must pass, and that is deliberate: a broken fit produces a NaN
/// sigma and `converge` reports it as `NonFiniteResult`. Rejecting it here
/// would turn reporting into a panic inside inference.
#[test]
fn a_nan_sigma_passes_through_from_ms() {
let g = Gaussian::from_ms(25.0, f64::NAN);
// `sigma()` is NaN exactly when the precision is: it guards `pi <= 0`
// (reporting `inf`) and `pi == inf` (reporting `0.0`), so NaN survives
// only from a NaN precision.
assert!(g.sigma().is_nan());
}
#[test]
#[should_panic(expected = "beta must be finite and non-negative")]
fn a_negative_beta_is_rejected_by_rating_new() {
let _ =
Rating::<i64, ConstantDrift>::new(Gaussian::default(), -4.17, ConstantDrift::new(0.0));
}
#[test]
#[should_panic(expected = "beta must be finite and non-negative")]
fn a_nan_beta_is_rejected_by_rating_new() {
let _ = Rating::<i64, ConstantDrift>::new(
Gaussian::default(),
f64::NAN,
ConstantDrift::new(0.0),
);
}
#[test]
fn a_zero_beta_is_accepted_by_rating_new() {
let _ =
Rating::<i64, ConstantDrift>::new(Gaussian::default(), 0.0, ConstantDrift::new(0.0));
}
/// `ConstantDrift` rejects at construction now that its field is private.
#[test]
#[should_panic(expected = "gamma must be finite and non-negative")]
fn a_negative_gamma_is_rejected_by_constant_drift_new() {
let _ = ConstantDrift::new(-0.0833);
}
#[test]
#[should_panic(expected = "gamma must be finite and non-negative")]
fn a_non_finite_gamma_is_rejected_by_constant_drift_new() {
let _ = ConstantDrift::new(f64::NAN);
}
#[test]
fn gamma_reads_back_what_was_given() {
assert_eq!(ConstantDrift::new(0.25).gamma(), 0.25);
assert_eq!(ConstantDrift::new(0.0).gamma(), 0.0);
}
/// `HistoryBuilder::drift` is generic and cannot inspect an arbitrary
/// `Drift`, so the check on the variance each competitor accumulates is
/// still needed — it is the only thing standing between a custom
/// implementation and a NaN fit. `ConstantDrift` can no longer reach it,
/// so this uses an implementation that can.
#[test]
fn a_custom_drift_returning_a_bad_variance_is_rejected_at_convergence() {
#[derive(Clone, Copy, Debug)]
struct BadDrift(f64);
impl trueskill_tt::Drift<i64> for BadDrift {
fn variance_delta(&self, _from: &i64, _to: &i64) -> f64 {
self.0
}
fn variance_for_elapsed(&self, _elapsed: i64) -> f64 {
self.0
}
}
for bad in [f64::NAN, f64::INFINITY, -1.0] {
let mut h = History::builder().drift(BadDrift(bad)).build();
h.record_winner(&"a", &"b", 1).unwrap();
h.record_winner(&"a", &"b", 5).unwrap();
let err = h.converge().unwrap_err();
assert!(
matches!(
err,
InferenceError::InvalidParameter {
parameter: trueskill_tt::Parameter::DriftVariance,
..
}
),
"drift {bad}: {err:?}"
);
}
}
/// An ordinary drift is untouched.
#[test]
fn an_ordinary_drift_still_converges() {
let mut h = History::builder()
.drift(ConstantDrift::new(25.0 / 300.0))
.build();
h.record_winner(&"a", &"b", 1).unwrap();
h.record_winner(&"a", &"b", 5).unwrap();
assert!(h.converge().unwrap().converged);
}
}
+196
View File
@@ -0,0 +1,196 @@
//! `expected_variance_reduction`: which matchup best sharpens a given question.
use smallvec::smallvec;
use trueskill_tt::{
ConstantDrift, ConvergenceOptions, Event, History, InferenceError, Member, Outcome, Team,
UnknownKeys,
};
type H = History;
fn round(a: &'static str, b: &'static str, sa: f64, sb: f64) -> Event<i64, &'static str> {
Event {
time: 1,
teams: smallvec![
Team::with_members([Member::new(a)]),
Team::with_members([Member::new(b)]),
],
outcome: Outcome::scores([sa, sb]),
}
}
fn base() -> Vec<Event<i64, &'static str>> {
vec![
round("a", "b", 5.0, 2.0),
round("a", "c", 6.0, 1.0),
round("b", "c", 4.0, 3.0),
round("c", "d", 2.0, 1.0),
round("a", "d", 7.0, 2.0),
]
}
fn fit(extra: Option<Event<i64, &'static str>>, policy: UnknownKeys) -> H {
let mut h: History = History::builder()
.mu(0.0)
.sigma(6.0)
.beta(1.0)
.score_sigma(2.0)
.drift(ConstantDrift::new(0.0))
.unknown_keys(policy)
.convergence(ConvergenceOptions {
max_iter: 20_000,
epsilon: 1e-13,
alpha: 1.0,
})
.build();
let mut ev = base();
if let Some(e) = extra {
ev.push(e);
}
h.add_events(ev).unwrap();
let _ = h.converge().unwrap();
h
}
/// The closed form must equal what actually happens if the matchup is played.
/// This is the assertion that makes the whole call trustworthy: a wrong
/// acquisition function returns plausible numbers and quietly picks worse
/// matchups forever.
#[test]
fn the_closed_form_matches_an_actual_refit() {
let h = fit(None, UnknownKeys::Reject);
let target: Vec<(&&str, f64)> = vec![(&"a", 1.0), (&"b", -1.0)];
let before = h
.joint()
.unwrap()
.posterior_of(&target)
.unwrap()
.sigma()
.powi(2);
for (x, y) in [("a", "b"), ("c", "d"), ("a", "c"), ("b", "d")] {
let predicted = h
.joint()
.unwrap()
.expected_variance_reduction(&[&[&x], &[&y]], &target)
.unwrap();
let after = fit(Some(round(x, y, 3.0, 1.0)), UnknownKeys::Reject);
let actual = before
- after
.joint()
.unwrap()
.posterior_of(&target)
.unwrap()
.sigma()
.powi(2);
assert!(
(predicted - actual).abs() / actual.abs() < 1e-9,
"{x} vs {y}: predicted {predicted}, actual {actual}"
);
}
}
/// The reduction cannot depend on the score, because for a Gaussian likelihood
/// the posterior variance update is data-independent. This is why the call
/// needs no expectation despite its name.
#[test]
fn the_outcome_does_not_change_the_reduction() {
let target: Vec<(&&str, f64)> = vec![(&"a", 1.0), (&"b", -1.0)];
let h = fit(None, UnknownKeys::Reject);
let before = h
.joint()
.unwrap()
.posterior_of(&target)
.unwrap()
.sigma()
.powi(2);
let mut seen = Vec::new();
for (sa, sb) in [(3.0, 1.0), (100.0, -50.0), (0.0, 0.0)] {
let after = fit(Some(round("c", "d", sa, sb)), UnknownKeys::Reject);
seen.push(
before
- after
.joint()
.unwrap()
.posterior_of(&target)
.unwrap()
.sigma()
.powi(2),
);
}
for w in seen.windows(2) {
assert!(
(w[0] - w[1]).abs() < 1e-12,
"variance reduction moved with the observed score: {seen:?}"
);
}
}
/// The point of the call: it must rank candidate matchups usefully. Playing the
/// pair you are trying to separate helps most; an unrelated pair helps least.
#[test]
fn it_ranks_candidates_by_how_much_they_answer_the_question() {
let h = fit(None, UnknownKeys::Reject);
let target: Vec<(&&str, f64)> = vec![(&"a", 1.0), (&"b", -1.0)];
let direct = h
.joint()
.unwrap()
.expected_variance_reduction(&[&[&"a"], &[&"b"]], &target)
.unwrap();
let unrelated = h
.joint()
.unwrap()
.expected_variance_reduction(&[&[&"c"], &[&"d"]], &target)
.unwrap();
assert!(direct > 0.0 && unrelated > 0.0);
assert!(
direct > 5.0 * unrelated,
"playing the target pair should dominate: {direct} vs {unrelated}"
);
}
/// A matchup between two competitors nobody has seen still teaches something
/// about them, but nothing about a target that does not involve them.
#[test]
fn an_unrelated_unseen_matchup_teaches_nothing_about_the_target() {
let h = fit(None, UnknownKeys::Prior);
let target: Vec<(&&str, f64)> = vec![(&"a", 1.0), (&"b", -1.0)];
let reduction = h
.joint()
.unwrap()
.expected_variance_reduction(&[&[&"stranger"], &[&"nobody"]], &target)
.unwrap();
assert!(
reduction.abs() < 1e-12,
"an unseen pair shares nothing with the target: {reduction}"
);
}
#[test]
fn shape_errors_are_reported() {
let h = fit(None, UnknownKeys::Reject);
let target: Vec<(&&str, f64)> = vec![(&"a", 1.0), (&"b", -1.0)];
assert!(matches!(
h.joint()
.unwrap()
.expected_variance_reduction(&[&[&"a"]], &target),
Err(InferenceError::MismatchedShape {
expected: 2,
got: 1,
..
})
));
assert!(matches!(
h.joint()
.unwrap()
.expected_variance_reduction(&[&[&"a"], &[&"ghost"]], &target),
Err(InferenceError::UnknownKey { .. })
));
}