65 Commits
Author SHA1 Message Date
logaritmisk d9e85cda1d chore: Release trueskill-tt version 0.5.0 2026-09-08 02:16:55 +02:00
logaritmiskandClaude Opus 5 633a503900 refactor!: remove the factor-graph surface nothing used, add try_winner
Two decisions taken before cutting 0.5.0.

#42 — the `Schedule` trait was public API the engine never called. Its
only call site was `Game::custom`, itself `#[doc(hidden)]`, and
`EpsilonOrMax` was never constructed anywhere. Removing that surface
showed the problem was larger than the issue described: with `custom`
gone, the compiler found `Factor`, `BuiltinFactor`, `RankDiffFactor` and
`TeamSumFactor` all dead too.

`Game::run_chain` drives a local `DiffFactor` enum and bypasses the
whole T1 abstraction — it has done since it was written. So this is not
just an unused extension point but the machinery it was built on, and
`CLAUDE.md` was documenting it as live architecture.

Removed: `graph` module, `Schedule`, `EpsilonOrMax`, `ScheduleReport`,
`Game::custom`, `Factor`, `BuiltinFactor`, `RankDiffFactor`,
`TeamSumFactor`. `TruncFactor`, `MarginFactor`, `VarStore` and `VarId`
stay — inference uses those. The measurement behind choosing removal
over wiring is in #42: the within-game loop converges in 1 to 8
iterations against a cap of 30, so a `Residual` schedule has no headroom
to reclaim, and `Damped` already shipped as `ConvergenceOptions::alpha`.

#20 — `Outcome::winner` panicking on an out-of-range index. Kept, and
the reasoning is now on the method. It is the only constructor here that
validates, which looks inconsistent until you try deferring like its
siblings: `winner(5, 2)` produces ranks `[1, 1]`, an all-tied draw that
ingestion accepts without complaint when `p_draw > 0`. Asking "team 5
won" and silently getting "everyone drew" is the exact failure this
crate keeps removing, so the check belongs where the mistake is.

Adds `Outcome::try_winner` for indices that are computed or parsed
rather than written literally, following the `new`/`try_new` convention.
That is additive; the panicking form stays because every call site in
this repo, its tests and its README passes literals, where a `?` would
be noise.

BREAKING CHANGE: the `graph` module and everything it exported are
removed, as are `Game::custom` and `ScheduleReport`.

Closes #42. Closes #20.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 02:14:43 +02:00
logaritmiskandClaude Opus 5 7cf45db5cf feat: add expected_variance_reduction for scored active learning
#49: `expected_information_gain` enumerates discrete outcomes, so a
consumer recording continuous scores cannot ask which matchup to run
next. The issue flagged this as possibly a research question, since
"expected variance reduction under EP may not have a clean closed form
even for Gaussian likelihoods".

It does. Observing a scored event is a rank-one update to the precision
matrix, so Sherman-Morrison gives

    reduction = (c^T L^-1 a)^2 / (v + a^T L^-1 a)

for target functional c and matchup contrast a. Verified against an
actual refit on four candidate matchups: agreement to 1e-9 relative.

Two consequences worth stating.

There is no expectation to take. The expression depends on which matchup
is played but not on how it turns out, because for a Gaussian likelihood
the posterior variance update is data-independent. Pinned by
`the_outcome_does_not_change_the_reduction`, which refits with scores of
(3, 1), (100, -50) and (0, 0) and gets the same answer. The name keeps
the term the active-learning literature uses; no averaging happens.

It is also far cheaper than its ranked counterpart — one linear solve
rather than a full inference pass per possible outcome — because
`c^T L^-1 a` and `a^T L^-1 a` share the same solve.

`target` is deliberately the same linear-functional shape as
`posterior_of`, as the issue proposed, so the two share a concept rather
than inventing two.

The load-bearing test is the refit comparison. An acquisition function
is the archetype of a surface that returns finite, plausible, monotone
numbers while being wrong, and then quietly selects worse matchups
forever; ranking behaviour alone would not catch that.

Closes #49

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 02:08:32 +02:00
logaritmiskandClaude Opus 5 c866210c65 feat: add History::predict_margin for scored matchups
#48: every predict_* answers "who wins", and a consumer recording scores
never asks that. It wants the interval on the result, and having none it
hand-fitted a noise law whose fitted node weight came out at 0.0 — so
the quoted sigma was 5.83 whether the competitor had forty rounds or
none, against real residual spreads of 5.8 and 12.44.

`predict_margin` composes the three things that make a scored result
uncertain: the joint posterior over the competitors, their per-event
performance noise, and the observation noise on the score. It widens as
the model knows less — measured, sigma 2.48 against an opponent with
forty rounds, 3.38 against one seen once, wider still against one never
seen — which is the property the hand-fitted law lost.

It is a margin, not a score, and that is not a shortcut. Scored
ingestion reduces every event to `score_a - score_b` before inference,
so the absolute level is discarded: shifting every score in a history by
+100 or -1000 produces a bit-identical fit, verified. There is no
information from which to predict what a competitor will *score*.
Returning one would be a number derived entirely from the prior, which
is exactly the plausible constant this crate keeps finding and removing.

`posterior_of` now honours `UnknownKeys::Prior`, which gives #48 its
second requirement — "I have never seen this competitor, here is the
prior-informed answer". An unseen competitor shares no event with the
slice, so it is independent by construction and its variance is additive
rather than part of the solve.

Worth recording for expectations: for a *margin* the joint buys little
over adding marginals (2.4798 against 2.5112 here), because a margin is
a difference and differences are where the loopy underestimate and the
ignored correlation cancel. The gain here is having a predictive
distribution at all. `posterior_of`'s correlation handling earns its
keep on sums and single nodes instead — see tests/additive_model.rs.

Closes #48

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 02:02:15 +02:00
logaritmiskandClaude Opus 5 1f791bcddd test: pin what an additive model does to combined uncertainty
Investigating #47 needed a reproduction of the reporter's structure: a
joint player/layout model where every observation measures a sum of
nodes against a reference, so the data pins differences and leaves the
overall level to the prior.

The result corrects #46's framing as well as answering #47. Combining
marginals is wrong in opposite directions depending on the combination,
and the unsafe direction is not the one either issue assumed:

    combination            exact   adding marginals
    p0 + h0 (a round)     4.0850   0.8041   0.20x   OVERconfident
    p0 - p1 (rank two)    0.8741   0.8592   0.98x   about right
    h0 - h1               0.7119   0.7057   0.99x   about right

#46 assumed every published figure was too wide and that this was "at
least the safe direction". That holds for differences. For sums it
reverses: adding marginals is five times too narrow, which publishes a
claim the data does not support.

For a single node the exact marginal is ~5x wider than message passing
reports, because the level it shares with its partners is pinned only by
the prior. That is the honest posterior for a weakly identified
parameter, not a defect — and it is why #47's node "should be
publishable" intuition and its posterior sigma disagree.

Refs #46, #47

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 01:52:39 +02:00
logaritmiskandClaude Opus 5 c52e2550af feat: add History::posterior_of for a linear combination of competitors
#46: every accessor returns a per-competitor marginal, and almost
nothing a consumer publishes is one competitor. Combining marginals
assumes independence, and competitors are correlated through every event
they share.

`posterior_of(&[(a, 1.0), (b, -1.0)])` returns the posterior of that
combination with the correlation intact. Validated against the exact
linear-Gaussian posterior on both a tree and a loopy fixture, for
differences and for single competitors: agreement to 1e-9 relative in
every case.

The investigation that preceded this is why it is not a covariance
accessor. Marginals from loopy message passing are about half the true
width, and ignoring correlation overstates a difference — the two errors
partially cancel, leaving 1.327x rather than 2.646x. Bolting true
correlations onto the existing marginals would have given 0.765 against
a true 1.524, which is overconfident: the direction the reporter
specifically called unsafe. Rebuilding the joint from the factor
structure fixes both at once, and a single-competitor query now returns
the exact marginal rather than the narrow one.

The precision matrix depends only on structure — who played whom, with
what weights and what noise — not on the observed outcomes, and the
means were already exact. So only the second moment is reconstructed.

Known limits, all deliberate and documented on the method:

- Latest slice only. A functional spanning times, such as "current
  versus career", needs the time-expanded joint and is not covered.
- Scored events only. A ranked outcome's truncation is EP-approximated
  and its converged factors are not retained after inference, so ranked
  slices return `JointUnavailable` rather than a plausible wrong number.
- Dense Cholesky, O(n^3) per query in the slice's competitor count:
  38.8us at 50, 5.66ms at 400, 49.1ms at 800. Fine for the sizes this
  serves today; caching the factorization per slice would make repeat
  queries O(n^2), and sparsity is the next step after that.

Refs #46, #47, #48

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 01:46:35 +02:00
logaritmisk 4924bc8b57 style: use arrays rather than vec! in the calibration fixture
clippy::useless_vec. Pushed broken because the verification chain used
`if ... then echo` blocks that report status without gating the `&&`
that follows — the same shape of mistake as e1bddf2, where a grep
succeeded on the error text. Both times the check printed FAIL and the
push went ahead.

The reliable form is a single fail-fast chain:

    just lint && just test && just determinism && cargo test --doc

so nothing runs after the first failure.
2026-09-08 01:39:21 +02:00
logaritmiskandClaude Opus 5 36eacf5f67 test: calibrate the marginals against the exact posterior
Investigation for #46 and #47, before touching either.

A scored history is linear-Gaussian, so its true joint posterior has a
closed form and the crate can be checked against ground truth. Measured
on five competitors:

                    means      marginal sd (crate / exact)
    tree (star)     exact      1.000
    loopy (robin)   exact      0.502

On a tree the crate is exact in both. With cycles the means stay exact —
the standard Gaussian-BP result, and the property ratings rely on —
while marginal variances come out about half the true width.

That is the opposite direction from what #47 reports, so whatever is
happening in that consumer's model, the crate being conservative is not
it.

It also means #46 cannot be implemented as an added covariance accessor.
The exact correlation between two nodes here is +0.857, so ignoring it
overstates the width of a difference — but the too-narrow marginals
partially cancel that, leaving 1.327x rather than 2.646x. Adding true
correlations to these marginals without correcting them would give 0.765
against a true 1.524: overconfident, which is the direction the reporter
specifically called unsafe.

Pins the two real invariants (exactness on a tree, exact means with
cycles) and deliberately only records the variance gap, since closing it
is what #46 proposes.

Also records the working rules this project has converged on: investigate
before implementing, fix the root issue, and scout crates.io on measured
accuracy rather than adoption.

Refs #46, #47

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 01:38:46 +02:00
logaritmiskandClaude Opus 5 71554fd944 feat: add UnknownKeys::Prior, and explain why there is no Skip
#44's third ask was an opt-in mode so a caller with partially-known
teams need not pre-filter. The requested shape was `Skip` — drop unknown
members. Measured, that is the wrong mode to build.

A team's performance is the *sum* of its members, so dropping one drops
its variance too. On a two-member team with one unknown:

    SKIP  (drop the member)  : performance sigma 2.37
    PRIOR (member at prior)  : performance sigma 6.53   (2.76x wider)

Skipping makes the model *more* certain because it knows *less*, which
is backwards. `Prior` is also the answer the model already gives for a
competitor it knows about but has no evidence for — measured, such a
competitor sits at sigma 4.99 against the prior's 6.0 — so it
corresponds to a state the model can actually be in. Skipping does not.

So the enum is `Reject` (default, unchanged) and `Prior`, and it is
`#[non_exhaustive]` in case a real use for skipping turns up later.

Placed on `HistoryBuilder` rather than per-call. Neither consumer wants
it to vary between queries: one scores thousands of candidate matchups
in a loop, the other's headline feature is predicting a competitor
nobody has faced. That makes it a property of how the model is being
used, and keeps five prediction signatures unchanged.

This also gives #48 the semantics it asked for — "I have never seen this
competitor, here is the prior-informed answer" — which it needs for
predicting a course nobody has played.

`an_unknown_member_widens_its_team_rather_than_narrowing_it` pins the
property that ruled `Skip` out, so a future convenience cannot quietly
reintroduce it.

Closes #44. Refs #48

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 01:28:49 +02:00
logaritmiskandClaude Opus 5 2cf21a753d docs: record that the event log is the source of truth, and why
#45 asked whether a fitted `History` can be persisted, and noted that
the absence of `serde` "reads as an omission rather than a decision,
which invites exactly this issue from every consumer in turn". Fair.

Documents the decision on `History` along with the reasoning that makes
it one: `converge` reaches a fixed point determined by the events,
ratings and configuration alone, so a snapshot would carry no
information the event log does not — it would cache the computation,
never the answer.

It also bounds what a snapshot could buy, since that is the question a
consumer actually has. Re-converging an unchanged history costs one
iteration, measured at 0.91 ms against 365 ms cold on 2 000 events, so
it would make a cold restart cheap and do nothing for appends. Appending
one event moves its participants more than a sigma across their whole
history, so that re-convergence is real work rather than repeated work.

Closes #45

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 01:22:28 +02:00
logaritmiskandClaude Opus 5 35d7512557 fix(test): the ingestion-order property was comparing two truncated fits
`ingestion_order_does_not_change_the_answer` failed with a 1.2e-6 gap in
mu and read as a violation of the invariant. It was not. Both sides ran
`max_iter: 200` against `epsilon: 1e-10`, and the batched side stopped
at the cap with a step of 3.4e-9 — so the test compared two fits that
had not converged and attributed the difference to ingestion order.

Raised to 20_000, at which both converge and the property holds. Runtime
is unchanged at 0.05s, because converging is what the iterations were
for.

The test now asserts `report.converged` on both sides before comparing.
That is the part worth keeping: any test that compares two fits for
equality is measuring truncation unless it first establishes that both
reached a fixed point. `ingestion_equivalence.rs` already did this;
`properties.rs` did not.

Found because c12bc83 made `ConvergenceReport` `#[must_use]`, which is
the same failure #50 describes — a short fit is wrong by a little and
looks entirely plausible. The blanket `let _ =` that commit applied to
78 call sites was too blunt here: binding the report to `_` silenced the
one signal that would have caught this, in a test whose whole purpose is
to compare two fits.

Refs #50

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 01:20:48 +02:00
logaritmisk e1bddf2474 style: factor the event-pair type out of the reconvergence fixture
clippy's `type_complexity` fires on the tuple-of-vectors return. Caught
after pushing, because the verification chain used `&&` between `just
lint` and a `grep` that succeeded on the error text — so a failing lint
reported success. The gate has to be on the command, not on whether the
output matched.
2026-09-08 01:18:59 +02:00
logaritmiskandClaude Opus 5 5d36fa1008 test: pin that re-convergence is path-independent
Answering #45 — can a fitted `History` be persisted — needed to know
whether `converge` reaches a fixed point determined by the events alone,
or one that depends on the message state it started from. It is the
former, and that is worth a test rather than a comment.

`tests/ingestion_equivalence.rs` varies how events are batched but
converges only at the end. These converge *between* batches, which is
the path a caller takes when it fits, serves, then ingests more.

Measured divergence from a single fit over the same events: 6.2e-13 for
an append strictly later than every existing slice, 8.9e-11 for one
interleaved with them. Both at the convergence tolerance. The design
question guessed the interleaved case might be weaker; it is not, and
the reason is that Through Time revises the past on every converge
anyway, so doing it in two steps is not a special case.

Also pins that re-converging an unchanged history costs one iteration.
Measured on a 2000-event fixture that is 0.91ms against 365ms cold — the
fact that makes a restored snapshot worth having at all.

Refs #45

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-08 01:18:28 +02:00
logaritmiskandClaude Opus 5 c12bc830a5 feat!: name the unknown key, expose tail probabilities, flag short fits
Three issues from two downstream consumers, all small, all sharing a
theme: the crate had the information and would not hand it over.

#44 — `UnknownKey { team: 0, member: 0 }` did not say which key. A
consumer upgrading 0.1.2 -> 0.4.1 had every one of 5591 predictions
return this error, fell back to a neutral 0.5, and lost its entire
metadata model for a day. Nothing crashed and nothing logged; it was
found by sweeping an unrelated parameter and noticing the output did not
move. The 0.4.0 change that made unknown keys an error was right — the
error was just too anonymous to act on. It now carries the key's `Debug`
rendering, and its `Display` says what to do about it. The precondition
is documented on every prediction entry point, which the reporter said
would alone have saved the day.

#43 — `cdf` was `pub(crate)`, so a consumer asking "is this competitor
below the cutoff" approximated it with a `mu + z * sigma` band and had no
way to say what confidence any `z` bought. Adds
`Gaussian::probability_below` / `probability_above`. The second is
separate on purpose: `1 - cdf` collapses to exactly zero past ~8.3
sigma, and a stopping rule is evaluated precisely there. Both route
through the survival function added in 0.4.1, so this is visibility
rather than new numerics.

#50 — `ConvergenceReport` was not `#[must_use]`, so the one signal that
a fit stopped short was trivially discarded. It now is, and that
immediately found 78 sites doing exactly that — including this crate's
own ATP example, which was capped at 10 sweeps when the history needs
30. The example now reads the report and says so.

`ITERATIONS = 30` is documented as the floor it is, with the three
measurements to hand: 400 events over 100 competitors already stops
there at ~7e-3 against a 1e-6 tolerance, the ATP example needs 30 at a
much looser one, and a consumer's 2000-node model needs 76 to 161.

BREAKING CHANGE: `InferenceError::UnknownKey` gains a `key` field, and
the prediction methods now require `K: Debug` in order to fill it.

Closes #43, #50. Refs #44 — its third ask, an opt-in `UnknownKeys::Skip`
mode, is a live API question and deliberately not answered here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-07 23:41:03 +02:00
logaritmisk 901f60972e chore: Release trueskill-tt version 0.4.2 2026-09-07 22:46:54 +02:00
logaritmiskandClaude Opus 5 8116fd081f test: localise the erfc_inv tail residual to the caller's argument
Scouting crates.io for a more accurate `erfc` than `libm` turned up a
1.16e-12 relative error in `erfc_inv` at `p_draw = 0.999999`, measured
against a 70-digit `decimal` reference. It looked like too few Newton
steps. It is not: adding a fourth changed nothing.

The error is in forming the argument. `1.0 - 0.999999` is
`1.0000000000287557e-06` — 0.999999 is not representable, and
subtracting from one cancels, leaving 2.9e-11 of relative error before
`erfc_inv` is entered. Given an exactly-representable argument it
returns 1.8e-16. So the routine was never the problem, and the extra
iteration has been reverted rather than shipped as a fix for a defect
that was not there.

`puruspe::inverfc` returns the identical wrong value for the identical
reason, which is what makes the shared upstream cause obvious.

Adds a test that separates the two, and corrects a quantile constant in
`erfc_inv_matches_known_quantiles` that was recalled rather than
computed: `Phi^-1(0.9999995)` is 4.89163847569859, not
4.891638475699099. The others were checked against the same reference
and were right.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-07 22:43:50 +02:00
logaritmiskandClaude Opus 5 17d072b2ae fix: route every transcendental through libm, and combine sigmas with hypot
Follow-on from #41, which added `libm` for `erfc`. Surveying what else
the dependency offers: its unique surface over `std` is `erf`/`erfc`,
`lgamma`/`tgamma` and Bessel functions, and only the first was ever
needed. But the survey found something better than another special
function.

`std`'s `exp` and `ln` delegate to the *system* math library. IEEE 754
specifies the basic operations and `sqrt` exactly and says nothing about
transcendentals, so those differ per platform. Measured here over 200k
inputs:

    exp: 19425/200000 differ from libm (worst 1 ulp)
    log:  9932/200000 differ

Inference is an iterative fixed point, so a one-ULP difference can change
an iteration count and move the answer by more than one ULP. Routing
every transcendental through `libm` makes a fit reproducible across
platforms — a stronger guarantee than `tests/determinism.rs`, which only
covers thread counts.

It costs nothing. `Batch::iteration` measured -2.7% [-5.7%, -0.3%] with
the whole set swapped, and not one golden moved.

Also switches the two places that combined sigmas as
`sqrt(a^2 + b^2)` to `hypot`. Squaring overflows to infinity above
~1.3e154 and flushes to zero below ~1.5e-154 — measured, the naive form
returns `inf` where `hypot` returns 1.41e160 — and `Gaussian`'s
constructors are public, so a caller can reach both ends.

Deliberately not done: rewriting the KL divergence's `ln` of a ratio via
`ln_1p`. The cancellation is real as the ratio approaches one, but
measured absolute error is at most ~1e-11 in a quantity of order 0.4
nats, so it changes nothing.

The invariant is recorded in `CLAUDE.md` and on `erfc`'s own docs, since
nothing enforces it mechanically.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-07 22:33:30 +02:00
logaritmiskandClaude Opus 5 3dd659307a fix: replace the erfc approximation with libm, for free
#41 asked whether the Numerical Recipes `erfcc` approximation — 1.2e-7
relative, and the binding accuracy constraint on the whole crate — was
worth replacing, given it sits in the inference hot loop. It is, and it
costs nothing.

Measured, against an independent incomplete-gamma reference:

    range        previous (NR)        libm
    [-3, 0]            7.95e-8     2.15e-14
    [0, 0.5]           8.69e-8     4.70e-14
    [0.5, 2]           9.38e-8     1.24e-12
    [2, 6]             1.04e-7     5.20e-14
    [6, 26]            1.07e-7     1.75e-13

    erfc(0)          1.00000003          1.0  (exactly)
    |erfc(z)+erfc(-z)-2|  6.00e-8     2.22e-16

Performance, on `benches/batch.rs`: change [-3.34% +2.26%], p = 0.89 —
no change detected.

That result is counterintuitive, because libm's erfc is 1.65x slower
when swept uniformly over [-2.5, 2.5]. The sweep was the wrong input
distribution. Capturing the arguments inference actually passes:

    |x|<0.5    96.16%
    0.5-0.84    2.05%
    0.84-1.25   1.24%
    1.25-2      0.54%
    2-6         0.00%

98% fall below 0.84375, which is exactly where FDLIBM skips the
exponential entirely — while the NR form always pays for one. On the
real trace libm is the faster of the two (3.23 vs 3.78 ns/call).

An ad-hoc `Instant` harness reported a 16% end-to-end speedup; that was
an artifact of its own setup allocating and leaking per run, and
criterion's verdict of "no change" is the one to believe.

What it bought:

- `compute_margin` against exact quantiles: 8.4e-8 -> 1.7e-16.
- `cdf(mu, mu, sigma)` is now exactly 0.5; it was 1.5e-8 out.
- `sf + cdf` sums to one within a ULP, from 3e-8.
- `erfcx`'s two branches now agree to round-off across the crossover
  rather than to 1e-7, so the log-space evidence path and the linear one
  are consistent.
- Ten test tolerances tightened from 1e-6 to 1e-13..1e-15, and the
  prediction floor is now the integrator's rather than `cdf`'s.

Five goldens moved, by 2.4e-9 to 6e-7 — the magnitude of the removed
error, and `test_env_ttt`'s mu still rounds to the same six decimals.
Re-recorded with more digits so future drift stays visible. Verified as
movement toward truth per the goldens policy: every value now derives
from a primitive checked against an independent reference and satisfying
the exact identities, which the previous one did not.

Adds `libm` — zero transitive dependencies, rust-lang maintained.

Closes #41

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-07 22:25:00 +02:00
logaritmisk 7e289ee834 chore: Release trueskill-tt version 0.4.1 2026-09-07 21:49:05 +02:00
logaritmiskandClaude Opus 5 564969ee5d test: pin quality()'s N-group closed form, closing the README cross-check
The scan for precision defects found none in `quality()` — but it did
find that N identical teams have an exact closed form, which is a much
stronger regression net than the single two-team golden that was there.

For two identical single-player teams quality is
`sqrt(2b^2 / (2b^2 + s1^2 + s2^2))`. With the conventional parameters
that ratio is exactly 1/5, and the N-group generalisation is
`(1/5)^((n-1)/2)` — one factor per adjacent pair. Measured across
n = 2..10 the implementation matches to 1e-9, so the determinant path
that #9 rebuilt is correct over the whole range, not just at n = 2.

The n=3 and n=5 values (0.200 and 0.040) are also what the `trueskill`
Python package produces for the same configuration, which is the
cross-implementation check the README Todo has been asking for since
the redesign. Asserted separately as literals so a change to the
closed-form reasoning cannot silently carry them along.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-07 21:46:46 +02:00
logaritmiskandClaude Opus 5 683813ec10 fix: correct erfc_inv's sign error and keep evidence in log space
A systematic scan for precision defects, following the tail-precision
work in 7341669. Three findings; the first is a correctness bug in a
released version.

1. `erfc_inv`'s initial guess had the wrong sign. Numerical Recipes'
   `inverfc` uses -0.70711 as the leading coefficient; this used
   +FRAC_1_SQRT_2. Since `rational - t` is negative, that put Newton on
   the mirror image of the root, and three fixed iterations could not
   cross back. Measured against exact standard-normal quantiles:

       p_draw   old rel err   new rel err
       0.50        1.46e-7       8.40e-8
       0.90        1.02e-1       8.63e-9
       0.95        3.06e-1       1.91e-8
       0.99        8.05e-1       5.89e-9

   `compute_margin` inherited it, so the draw margin was wrong for any
   `p_draw` above about 0.6 and *non-monotone* above 0.9 — it ran
   0.674, 1.476, 0.503, 0.982 as p_draw went 0.5, 0.9, 0.99, 0.999. A
   history configured for a 0.99 draw rate was being fitted at 0.385.
   Note it was slightly wrong everywhere, not only in the tail.

2. `MarginFactor` computed a density and clamped it. `pdf` underflows
   past ~38 sigma, so `ln` of the clamped zero reported -708 nats
   however far out the score actually was: 4292 nats adrift at 100
   sigma, and unbounded beyond. This is the same defect as the one
   fixed in `TruncFactor`, one file over, on the scored-outcome path.

3. `TruncFactor` still bottomed out past ~38 sigma even after 7341669
   removed the cancellation, because the linear probability itself
   underflows there.

2 and 3 are fixed the same way: factors cache a *log* evidence, built
from new `ln_pdf`, `ln_sf` and `ln_interval` helpers that factor the
shared exponential out analytically via the `erfcx` added earlier.
Nothing underflows, at any separation.

One golden moved. `test_1vs1vs1` runs at `p_draw = 0.5`, so it goes
through `compute_margin`; its 1e-6-place values shifted. Verified as
movement *toward* analytic truth by comparing both the old and new
inverse against exact quantiles, per the goldens policy in CLAUDE.md —
not re-baselined on faith.

Two test tolerances are asserted at 1e-6 rather than tighter because
above x = 2 `erfcx` uses a continued fraction accurate to ~1e-15 while
`erfc` carries ~1e-7, so the log path is the more accurate of the two
and they part company at `erfc`'s error. That floor is tracked in #41.

Refs #41

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-07 21:45:03 +02:00
logaritmisk 2a48d10aa9 chore: Release trueskill-tt version 0.4.0 2026-09-07 15:49:12 +02:00
logaritmiskandClaude Opus 5 d4f91fd221 fix: reject convergence options that silently disable inference
`Game::ranked` and `Game::scored` validated `p_draw` and `score_sigma`
but never `convergence`. `ConvergenceOptions` has public fields and
`GameOptions` carries one, so a caller could hand the engine a set that
`HistoryBuilder`'s eager asserts never saw. Past that, the only guard
was a `debug_assert!`, which is gone in the profile users ship.

An `alpha` of zero is the bad case, and it fails silently rather than
loudly. Measured in release before the fix:

    likelihoods: [[Gaussian { pi: 0.0, tau: 0.0 }],
                  [Gaussian { pi: 0.0, tau: 0.0 }]]

Every EP update unapplied, every likelihood uninformative, inference
returning the priors it was given — and an `OwnedGame` that looks
entirely ordinary to the caller. `HistoryBuilder::convergence` already
documents exactly this hazard; the `Game` constructors just did not
share the check.

Adds `ConvergenceOptions::validate`, called by both constructors.
Rejects `alpha` outside `(0.0, 1.0]` and negative `epsilon`; NaN fails
both comparisons and is rejected too.

`tests/validation.rs` states the release-mode guarantee for the whole
public surface, not just this hole, and CI already runs the suite in
release. Probing the other conditions #18 lists found five of eight
already enforced — ties without a draw probability, per-event score
sigma, weight/team dimensions, draw-probability range, score-sigma
range — so this closes the remaining gap rather than the whole issue.
The engine keeps its `debug_assert!`s as invariant documentation.

Refs #18

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-07 15:48:31 +02:00
logaritmiskandClaude Opus 5 8c087ad015 fix!: apply competitor configuration whenever it is supplied
`Member::with_prior` and `with_drift_scale` were consumed only on the
branch that *creates* a competitor — `priors.remove` sat inside
`if !self.agents.contains(..)`. Supplying either for a key the history
already knew did nothing at all: no error, no warning, and output
computed from the default prior. A prior applied on a competitor's very
first event and was silently discarded ever after.

Configuration now applies whenever supplied. Two details this forced:

Configuration is tracked per *field* rather than as a merged `Rating`.
A member setting only `drift_scale` must not also assert the default
prior, or it would silently undo a prior seeded on an earlier event.

Slice state has to be refreshed. `drift_scale` is re-derived on every
forward pass, but a prior is written into the competitor's earliest
slice once, at ingestion, and `iteration` refreshes only slices after
the first. Without the refresh a late prior would reach the drift terms
and nothing else — a subtler version of the drop being fixed. This was
caught by a test, not by reading the code.

Conflicting values for one competitor within a single batch are now
`ConflictingCompetitorConfig` rather than resolved by iteration order.
Events in a batch are unordered, so "last one wins" would make the
result depend on traversal — and `tests/ingestion_equivalence.rs` exists
to rule exactly that out. Repeating the same value stays inert, which is
the shape callers get when configuration is a property of the domain.

That invariant turned out to be tested only for *unconfigured*
competitors: every helper in that file built members with `Member::new`.
Extended to cover configured ones, including a check that configuration
changes the fit at all, so the order tests cannot pass vacuously.

`with_prior` had no coverage under `tests/` whatsoever, which is how
this survived. Adds `tests/competitor_config.rs`.

`drift_scale_is_ignored_after_first_appearance` asserted the old
behaviour and now asserts the new one. It was written as a deliberate
change-detector — "moving the capture would be a visible break, not a
silent one" — so it inverted rather than being deleted.

Also removes `InferenceError::ConvergenceFailed` and `NegativePrecision`,
which no code path ever constructed: public variants advertising failure
modes no caller could observe. Partial #20 — its other items were
already resolved, except `Outcome::winner` still panicking.

BREAKING CHANGE: `prior` and `drift_scale` now take effect for
competitors the history already knows, where they were previously
ignored; a batch supplying conflicting values for one competitor is now
an error. `InferenceError::ConvergenceFailed` and
`InferenceError::NegativePrecision` are removed.

Closes #10. Refs #20.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-07 15:42:31 +02:00
logaritmiskandClaude Opus 5 7341669d1a fix: stop destroying tail precision in evidence and truncation
`erfc` is sound — it holds ~1e-7 *relative* accuracy down to 1e-296 with
no tail degradation. Three expressions built on it threw that away by
subtracting quantities that both approach the same value.

1. `cavity_evidence` computed `1.0 - cdf(margin, ..)`, which is
   algebraically `sf(margin, ..)` and numerically a catastrophe: 7%
   error by eight sigma, and exactly zero past ~8.3, where the true
   probability is 1e-19 and perfectly representable. Clamped, that
   reached `log_evidence` as ln(f64::MIN_POSITIVE) = -708 whatever the
   truth was — off by 665 nats at nine sigma.

   `1 - cdf` is smallest precisely when the result contradicts the
   prior, so the model-comparison number was worst for upsets: the
   observation it exists to notice. Adds `sf`, the survival function,
   computed without the subtraction. The tie branch picks whichever tail
   keeps both of its terms small, for the same reason.

2. `v_w` computed the inverse Mills ratio as `pdf(-a) / cdf(-a)`. Both
   underflow together past about 39 sigma, giving `0 / 0` and putting
   NaN straight into the posterior. Adds `erfcx`, so the shared
   `exp(-alpha^2 / 2)` cancels analytically instead of being evaluated
   twice and divided.

3. With that fixed, `w = v * (v - alpha)` became the next casualty: `v`
   tends to `alpha`, so the gap lost every digit and drove `w` above 1,
   making `sqrt(1 - w)` NaN at alpha = 1e6. The gap now comes from its
   asymptotic series, which forms no difference at all. The tie branch
   had the same defect one expression over — `v * v - u` with both terms
   at 1e18 returned w = -128 — and a far-tail window is
   indistinguishable from a half-line, so it shares the asymptotic.

No public signature changes, and no existing golden moved: every one of
these only alters regions the old code got wrong. The two identity tests
are asserted at 1e-6 rather than tighter because `erfc` is not exactly
antisymmetric — `erfc(z) + erfc(-z)` differs from 2 by ~3e-8, and
`erfc(0)` returns 1.00000003. That floor is tracked separately.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-07 15:28:35 +02:00
logaritmiskandClaude Opus 5 2fff745c3b feat: let observers be shared, boxed, or borrowed
`History` takes its observer by value and never hands it back, so a
caller who wanted to read what an observer recorded had no way to keep a
handle to it. The natural spelling did not compile:

    let recorder = Arc::new(Recorder::default());
    History::builder().observer(Arc::clone(&recorder))
    // error[E0277]: `Arc<Recorder>: Observer<i64>` is not satisfied

The workaround was for every observer to wrap each of its own fields in
an `Arc` and derive `Clone` — one allocation and one lock per field, a
pattern each implementor had to rediscover, and nothing documenting it.

Adds blanket `Observer` impls for `Arc<O>`, `Box<O>` and `&O`. All are
`?Sized`, so `Arc<dyn Observer<T>>` and `Box<dyn Observer<T>>` work too
and an observer can be chosen at runtime. Also adds
`History::observer()` and `into_observer()`, so a non-shared observer's
state can be inspected in place or reclaimed after `converge` without
needing interior mutability at all.

`tests/observer.rs` is simplified to the shared spelling, so the
recommended pattern is the one demonstrated rather than the workaround.

Closes #40

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-07 15:13:28 +02:00
logaritmiskandClaude Opus 5 3c2f9ac64c feat: add expected information gain for active matchup selection
`quality()` answers "is this matchup fair". Callers picking which
comparison to run next need "is this matchup informative", and the two
coincide only for two evenly matched competitors. Without a principled
alternative, downstream code was reaching for hand-rolled heuristics
like `quality * sigma_a^2 * sigma_b^2`, which double-counts uncertainty:
the two factors are not independent.

Adds `expected_information_gain`, the outcome-weighted divergence
between current beliefs and the beliefs each result would produce:

    EIG = SUM P(outcome) * KL(posterior_after(outcome) || prior)

Available standalone over `Rating`s, and as
`History::expected_information_gain` using current skills and the
history's own beta, drift and p_draw — so the outcomes it weighs are the
ones that would actually be fitted.

This is the mutual information between the outcome and the skills, which
gives an analytic ceiling: gain cannot exceed the entropy of the thing
being observed, so at most `ln k` nats for k outcomes. That bound is the
sharpest test available, because an acquisition function is unusually
exposed to returning finite, plausible, monotone numbers while being
wrong — it would simply select slightly worse matchups forever. A
prototype of this returned 4.77 nats from a sign error while passing
every monotonicity check; `never_exceeds_the_entropy_of_the_outcome`
catches that class unconditionally.

Measured against the ceiling the values are meaningful rather than
vacuous: 0.382 nats for an even matchup between diffuse priors against
an 0.693 ceiling, falling to 0.013 for a lopsided one and 0.000 for a
hopeless one.

`disagrees_with_the_quality_times_variance_heuristic` pins down that
this is not a monotone transform of the heuristic it replaces — the two
rank a lopsided matchup and a confident even one in opposite orders — so
a later "simplification" cannot quietly revert to it.

Cost is one inference pass per possible outcome, documented on the
public API alongside the shortlist-then-score pattern, so callers do not
discover it in production.

Also folds the duplicated key-gathering in `predict_quality` and
`performances` into one validated `member_skills`.

Refs #39

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-07 15:08:12 +02:00
logaritmiskandClaude Opus 5 507894dae7 refactor!: close the remaining API gaps from #21
Three unrelated small defects, all requiring signature changes:

- `Game::one_v_one` hardcoded `GameOptions::default()`, so a 1v1 could
  never set `p_draw` or convergence options — and a drawn 1v1 was
  therefore unreachable through it, since the default `p_draw` is zero.
  It now takes `&GameOptions` like every other constructor.

- `Observer::on_batch_processed` was declared on the trait and never
  called from anywhere: implementors wired up a callback that could not
  fire. It is now called after each slice sweep, and renamed
  `on_slice_processed` to match the vocabulary the codebase adopted in
  T2 — the unit of work is a `TimeSlice`, not a batch. A slice is swept
  once travelling backward and once forward, so a multi-slice history
  fires it twice per slice per iteration; the doc comment says so.

- `pub mod factors` sat beside `pub(crate) mod factor`, two module paths
  differing by one character with only one of them importable. The
  public facade is now `graph`.

Tests cover each as a behaviour rather than a compile check: a drawn 1v1
succeeds only when p_draw is supplied, and the observer tests fail if
any callback stops firing.

BREAKING CHANGE: `Game::one_v_one` takes a fourth `&GameOptions`
argument; `Observer::on_batch_processed` is renamed
`on_slice_processed`; the `factors` module is renamed `graph`.

Closes #21

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-07 14:57:39 +02:00
logaritmiskandClaude Opus 5 bb2a845882 feat!: N-team outcome prediction with draw mass, replacing the 2-team panic
`predict_outcome` asserted `teams.len() == 2` and returned `[p, 1 - p]`,
allocating no probability to a draw even with `p_draw > 0`. For a
draw-enabled model the numbers were simply wrong, at any team count.

It now returns `Result<Prediction, InferenceError>` and supports N teams.

Two algorithms, both deterministic:

- Who finishes first. Performances are independent Gaussians, so this
  separates into a one-dimensional integral per team rather than a
  multivariate orthant probability. Adaptive Gauss-Kronrod evaluates it
  to ~1e-15, matching the exact two-team closed form.
- A specific finishing order. The factor graph only constrains
  rank-adjacent teams, so a full order is a chain of local constraints,
  not a general orthant integral. That chain collapses into a sequential
  recursion over cumulative integrals: O(teams * grid) per order.

Fixed-node Gauss-Hermite is the obvious tool for the first and is a trap:
when a rival's sigma is small the CDF product becomes a step narrower
than the node spacing, and the nodes step over it. Measured 4.4e-4 off
the closed form on a mildly skewed matchup and 1.7e-2 on a small-sigma
one, while still returning something that looks like a probability.
Adaptive refinement is what makes that case safe, and
`win_probabilities_survive_a_rival_with_a_tiny_sigma` pins it down.

The acceptance test is an identity rather than a golden: the outcome
space is exhaustive and disjoint, so the probabilities sum to one. Any
drift is integration error and nothing else. Gauss-Hermite failed it at
4.4e-4; this holds to ~1e-9.

Also from #21: unknown keys are now reported rather than dropped, so a
team of strangers can no longer produce a confident-looking prediction.
`predict_quality` returns `Result` for the same reason.

BREAKING CHANGE: `predict_outcome` returns `Result<Prediction, _>`
instead of `Vec<f64>`; `predict_quality` returns `Result<f64, _>`.

Refs #21, #39

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
2026-09-07 14:55:11 +02:00
logaritmiskandClaude Opus 5 87fca8dcca docs: correct drifted documentation and compile the README in CI
Four README code blocks no longer compiled: `Player` was renamed `Rating`
in T2, the `Drift` trait gained a `T: Time` parameter and a second method,
and two blocks were missing imports outright. The `Rating` example needed
more than a rename — with the binding unused, `T` is ambiguous because
`ConstantDrift` implements `Drift<T>` for every `T`, so it now carries an
explicit annotation.

Nothing compiled those blocks. `src/lib.rs` gains a `cfg(doctest)` struct
carrying `#[doc = include_str!("../README.md")]`, which turns every `rust`
block into a doctest without displacing the curated crate docs as the
front page. Verified it bites: reintroducing `Player` fails the build with
E0432 rather than shipping. Illustrative blocks are fenced `text` — note
that a bare fence defaults to `rust` under rustdoc, which is how the
`variance_delta = elapsed * γ²` formula became a compile error.

Prose fixes: README claimed `Gaussian::forget` takes a square root (it
works in variance space) and pointed at a `.gamma()` builder method that
does not exist. CLAUDE.md's data-flow diagram spliced the public ingestion
shape into the internal one — `Team` is not in that chain — listed
`cdf()`/`erfc()` as public when they are `pub(crate)` and private, and
called `SkillStore` public when only `CompetitorStore` escapes the crate.

Rustdoc fixes: `EventBuilder::scores_with_sigma` claimed a debug-assert
that `Outcome::scores_with_sigma` never had and whose own docs contradict;
rejection happens at ingestion as `InvalidParameter`. `event.rs` described
`add_events_with_prior` as replaced when it is still the ingestion
chokepoint. `factors.rs` advertised `Game::custom` without noting it is
`#[doc(hidden)]`. Internal T2/T4 milestone labels are dropped from public
items; the ones in the private `time_slice` module are left alone.

Closes #35

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014b6wy2q8rnFK8U8GPJVQNU
2026-09-01 19:30:32 +02:00
logaritmiskandClaude Opus 5 ef62b57a08 fix(release): skip the changelog hook during a dry run
`just release-plan` is documented as a preview that writes nothing, but
cargo-release runs pre-release hooks during a dry run too. The hook wrote
CHANGELOG.md and `git add`ed it, so the clean-tree check in `just release`
then refused to run — the repo's own two-step release workflow could not
be followed as written.

Guard the hook on DRY_RUN, which cargo-release 1.1.5 exports to the hook
environment (verified by dumping `env` from a throwaway hook; it also sets
CRATE_NAME, PREV_VERSION and NEW_VERSION). The clean-tree check itself is
left alone: it is load-bearing, because publishing is irreversible.

Closes #36

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014b6wy2q8rnFK8U8GPJVQNU
2026-09-01 19:26:22 +02:00
logaritmisk b2a7ade10c chore: Release trueskill-tt version 0.3.0 2026-09-01 06:34:31 +02:00
logaritmiskandClaude Opus 5 617bc07f6f feat: allow drift to vary per competitor via Member::with_drift_scale
Drift was a property of the History, so every competitor drifted at the
same rate and a fixed reference point could not share a graph with moving
competitors. A bot at a known strength, a rating floor, a course
difficulty — all of them drifted along with the players.

Member::with_drift_scale(s) multiplies the drift *variance* a competitor
accumulates, so s is in the same units as gamma: ConstantDrift(g) at
scale s behaves exactly as ConstantDrift(g * s) would for that competitor.
A scalar rather than a per-competitor Drift keeps History's single D type
parameter untouched and stays Copy. 0.0 pins a competitor still.

The scale lives on Rating, beside the drift it scales, and is applied
only through Rating::drift_variance_delta / drift_variance_for_elapsed.
Making those the sole entry points means a caller cannot reach the raw
drift and silently skip a competitor's scale — the filtered pass was
exactly that bug during development, caught because its test was written
before the wiring.

Like with_prior, the scale is competitor configuration captured at first
appearance rather than a per-event override; a competitor that is static
is static, and a scale that changed between events would make the skill
trajectory hard to interpret. Member's docs claimed prior was a per-event
override, which the code has never done — corrected here.

A negative scale is rejected rather than squared into its absolute value,
and a non-finite one rejected outright, both as InvalidParameter.

None means 1.0, so no existing call site changes and no existing fit
moves. Adding a public field to Member does break struct-literal
construction downstream.

Closes #34

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014b6wy2q8rnFK8U8GPJVQNU
2026-09-01 06:29:02 +02:00
logaritmiskandClaude Opus 5 1a88678384 refactor!: remove ConvergenceReport::slices_skipped
Closes #33. The field was public, hardcoded to 0 at both construction sites,
and had no route to ever being non-zero.

It was added in T3 as the reporting surface for dirty-bit slice skipping. That
feature is #4, closed as unworkable — the ceiling measured at ~6% against a
projected 5-50x, on top of three independent soundness blockers. #32, which
reattributed the cost to ingestion, is closed too: the re-convergence is
necessary work rather than waste, because appending one event genuinely moves
the involved competitors ~1.2 sigma across their whole history. Nothing left
would ever populate it.

This is the same defect class as #19, where this arc started: a public surface
that looks implemented, reports a plausible value, and is inert. A caller
reading `slices_skipped: 0` reasonably concludes "no slices were skipped this
run", not "this feature does not exist".

Removed rather than documented as reserved. Its only value was as a hook for a
plan that no longer exists, and keeping it preserves the shape of that plan.
Breaking, but ConvergenceReport is returned rather than constructed by callers,
so the only breakage is code reading a constant zero.

Also added a test asserting every remaining field carries real information —
iterations non-zero, final_step finite, log_evidence a finite negative log
probability, and per_iteration_time holding one duration per iteration.
Mutation-proved: pinning per_iteration_time to an empty SmallVec fails it. The
next always-constant member now has to survive an assertion rather than just a
reviewer's attention, which is the actual lesson of #19 and #33.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T5SYDExxL4vZgvunrcNSMc
2026-08-28 08:07:41 +02:00
logaritmisk 5f46296671 chore: ignore proptest regression seed files
proptest writes tests/*.proptest-regressions when a property fails, seeding a
replay of that exact case. Useful locally; noise in the repo when the failure
came from a deliberate mutation rather than a real defect.
2026-08-27 18:06:26 +02:00
logaritmiskandClaude Opus 5 2745fbb622 test: add property-based tests, a shared finiteness helper, and boundary inputs
Most of what remained on #26.

**Property tests (`tests/properties.rs`, proptest as a dev-dependency).** Four
invariants over generated 1v1 schedules rather than hand-written fixtures,
which is where this crate's shipped defects actually hid — a linear evidence
product that underflowed only past ~1000 teams, and a batching path no golden
exercised because every golden ingests in one call:

- converged posteriors are always finite with positive sigma
- log-evidence, batch and filtered, is finite and never above zero
- filtered evidence is invariant to whether `converge` has run
- one-at-a-time ingestion reaches the same fixed point as batched

The invariance property was mutation-proved: making `filtered_step` read
`skill.forward` instead of the carried message fails it with
`-1.1038430064192069 -> -1.1135747072822761`.

**Shared finiteness helper (`tests/common/mod.rs`).** `assert_finite` was local
to `degenerate_inputs.rs`. It now also rejects a non-positive sigma, which the
old version let through — `Gaussian::sigma` reports a non-positive precision as
improper rather than trapping, so a collapsed posterior would have passed a
finite-only check.

**Boundary inputs.** Zero and negative weights, out-of-order timestamps, and
extreme beta/sigma combinations. Worth recording that zero weight reaches
`(m - performance.exclude(..)) * (1.0 / w)` — a division by zero — and the
posterior comes out finite anyway; the test pins that rather than asserting
what ought to happen. The weight tests `expect()` the commit rather than
returning early on error, because an early return would have made them vacuous
the moment validation changed. I checked that specifically by turning the
return into a failure and confirming it did not fire.

Not done, and left on #26: benchmark regression gating. Nothing fails on a
regression today; making it fail needs a threshold chosen against how noisy the
shared runner is, which is a policy call rather than a mechanical one.

60 test binaries, up from 56. MSRV 1.85 verified with proptest in the graph.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T5SYDExxL4vZgvunrcNSMc
2026-08-27 18:06:26 +02:00
logaritmiskandClaude Opus 5 6d2573b92e perf: make the per-slice SkillStore compact instead of dense
Closes #17. Each `TimeSlice` owned a `Vec<Skill>` indexed by the GLOBAL
`Index.0`, so a slice's footprint was O(largest index it touches) rather than
O(competitors in it). Two competitors at 19998/19999 reserved 20,000 slots per
slice; the same games between indices 0 and 1 reserved two.

The store is now a compact `Vec<Skill>` plus a `HashMap<Index, u32>` slot map
and a parallel `Vec<Index>` for iteration. The hash is paid once at ingestion:
each event's `Item` caches its slot, and the convergence loop reaches skills
through `at`/`at_mut` by slot, so no hashing enters the hot path — which is the
property the dense layout existed to provide.

Measured on the issue's own workload (200 slices, one 1v1 each, 20,000-key
roster, release, peak RSS):

    indices 0 / 1          52 MB  ->  5.55 MB
    indices 19998 / 19999  309 MB ->  8.39 MB

The 257 MB gap is now 2.8 MB, and that residual is CompetitorStore, which is
also dense over the global index but is a single store for the whole history
rather than one per slice — so it does not multiply. Left alone deliberately.

Benchmarks, against the pre-change code:

    Batch::iteration        +2.4%   (regressed)
    history_converge x3     -18.8%, -21.6%, -21.7%  (improved)

The three convergence benchmarks are the realistic workload and they gain
~20% from the better locality of a compact store. The micro-benchmark loses
2.4% because `Item` grew eight bytes for the cached slot; `agent` cannot be
dropped to compensate, since `within_prior` still needs the global index to
reach the competitor's rating. I judged 2.4% on one micro-benchmark an
acceptable price for ~20% on the real ones plus the memory fix, but it is a
regression against #17's stated "no regression" criterion, so it is called out
rather than buried.

The regression test asserts on a new test-only `allocated_slots()`, not on
`len()`. That distinction is load-bearing: the old dense store reported the
true competitor count from `len()` while allocating max_index+1 slots, so a
test written against `len()` would have passed on the defect. Mutation-proved
by re-adding the dense padding, which fails it.

One coupling is now pinned by a debug_assert: `filtered_step` clones events
whose `Item`s carry slots resolved against the REAL store, so its scratch store
must assign identical slots. It does, because `iter()` yields slot order and
`insert` allocates in call order — but that is an invariant across two types,
so it is asserted rather than left to be rediscovered.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T5SYDExxL4vZgvunrcNSMc
2026-08-27 18:01:29 +02:00
logaritmiskandClaude Opus 5 1ac3b21db5 fix: enforce EventBuilder weight/team length in release
Part of #18. `EventBuilder::weights` guarded the length match with a
`debug_assert!`, so release builds accepted a mismatch, silently dropped the
weights, and ingested the event anyway. That is the exact shape #18 is about:
validation that exists only where it is least needed.

The setters return `Self` to keep the chain fluent, so they cannot return a
`Result`. The builder now records the first failure and `commit` returns it as
`MismatchedShape`. The weights are not applied on mismatch either, so a
partially-weighted team cannot reach the history by another route.

Two tests in tests/degenerate_inputs.rs, whose CI job runs in release — which
is the only place the old behaviour differed.

The second test needed strengthening before it was worth anything. As first
written it committed a ONE-team event, which ingestion rejects for an unrelated
reason, so it passed under a mutation that disabled the whole check. It now
uses two teams, so ingestion would otherwise succeed and the assertion is
actually load-bearing. Both tests were then mutation-proved together: disabling
the error path in `commit` fails both in release.

#18 stays open. The remaining debug_asserts live in `ranked_with_arena` and
`scored_with_arena`, and promoting those means threading `Result` up through
`Event::compute`, `TimeSlice::iteration`, `log_evidence` and `filtered_step` —
which lands on the public API as `log_evidence() -> Result<f64>` and
`filtered_learning_curve() -> Result<...>`. That is a trade-off about what the
query API should look like, not a mechanical change, so it is not mine to
decide.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T5SYDExxL4vZgvunrcNSMc
2026-08-27 17:53:37 +02:00
logaritmiskandClaude Opus 5 4e043364fd perf: stop cloning inference inputs in OwnedGame and ingestion
The last two items on #23.

`OwnedGame::new` and `new_scored` cloned the whole team structure to hand one
copy to `Game` and keep another. But `Game` takes the teams by value and is
dropped at the end of the constructor, so the vec can simply be taken back out
of it — the clone existed only because nobody looked at the lifetime.

`add_events_with_prior` deep-cloned each event's composition, results and
weights when chunking events into per-timestamp groups. Nothing reads those
three after the chunking loop (the agent-collection pass and the tie pre-check
both run before it), so the elements are now moved out with `mem::take`.

That soundness argument rests entirely on `o` being a permutation: visiting an
index twice would take an already-emptied vec and silently produce an event
with no teams rather than failing. Since that would be invisible, there is now
a debug_assert checking the permutation property directly, next to the comment
explaining why the code depends on it.

Verified on 1.85.0 as well as the local toolchain — an MSRV break in this
change would otherwise only surface in CI.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T5SYDExxL4vZgvunrcNSMc
2026-08-27 17:50:39 +02:00
logaritmiskandClaude Opus 5 aff3fb948d refactor!: replace emptiness-as-sentinel with Option for results and weights
Third of the five remaining items on #23.

`add_events_with_prior` and `TimeSlice::add_events` took `Vec<Vec<f64>>` and
`Vec<Vec<Vec<f64>>>` where an empty vec meant "not supplied" — so an empty
outer vec and a genuinely empty event list were the same value, and every
reader had to know which. Both are now `Option`, and the "not supplied"
branches read as `None` arms rather than `is_empty()` checks.

Two things fell out of the change that a sentinel would have hidden:

The tie pre-check iterated `results` directly. Under `Option` it needs
`.iter().flatten()`, which makes explicit that a `None` results list has no
ties to reject — previously an empty vec silently skipped the same loop and
looked identical to "checked, found nothing".

`MismatchedShape.got` could no longer be `results.len()`, because at the point
of the error there may be no vec to take a length from. It is now computed as
`map_or(0, Vec::len)` before the error is built.

MSRV note: my first draft used let-chains for the two validations. Those need
Rust 1.88 and this crate pins 1.85 — it compiled locally on 1.98 and would
have failed only in the MSRV CI job. Rewritten with `is_some_and`, and
verified by installing 1.85.0 and building against it rather than by assuming
the removal was complete.

Breaking: `TimeSlice::add_events` is public.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T5SYDExxL4vZgvunrcNSMc
2026-08-27 17:48:50 +02:00
logaritmiskandClaude Opus 5 06ed24b240 refactor!: make Competitor::message an Option, and compute_elapsed loud
Two of the five remaining items on #23.

`Competitor.message` was a `Gaussian` using the improper `N_INF` as an "unset"
sentinel, so `message != N_INF` meant "has a message" and every reader had to
know that convention. It is now `Option<Gaussian>`, which makes "no message
yet" and "a legitimately improper message" distinguishable at the type level
instead of by float comparison.

Worth noting what the change surfaced: switching the type turned every read
site into a compile error, and there were eight — two in the convergence sweep,
five in ingestion, one in new_backward_info. The last is the interesting one:
`skill.backward = agents[agent].message` needed `unwrap_or(N_INF)` rather than
an unwrap, because an absent message genuinely does mean the improper identity
there. A sentinel-based refactor would have had to find that by reading.

This is a breaking change: `message` is a public field. It rides the next
minor bump.

`compute_elapsed` clamped a negative elapsed to zero silently. Negative elapsed
means slices are being visited out of time order, which would otherwise make
drift *reduce* uncertainty. Release still clamps, so a bad timestamp degrades
to "no drift" rather than corrupting a posterior, but debug now trips — getting
there is a slice-ordering bug, not something callers can cause with ordinary
data.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T5SYDExxL4vZgvunrcNSMc
2026-08-27 17:46:12 +02:00
logaritmiskandClaude Opus 5 9b2c2b38c8 docs: complete the public API documentation contract
Closes the last open item in #25. `cargo clippy -W missing_errors_doc
-W missing_panics_doc -W must_use_candidate -W doc_markdown` went from 56
warnings to zero.

The 13 hand-written sections name the actual variants each function returns
rather than gesturing at "an error". Establishing that meant reading the error
paths — `Game::ranked` alone returns four distinct variants, and `record_draw`
can hit TieWithoutDrawProbability where `record_winner` provably cannot, since
a two-team decisive outcome has nothing to tie. Documenting those as
interchangeable would have been worse than leaving them undocumented, because
a reader would trust it.

Two existing doc comments already described panics in prose but not under a
`# Panics` heading, so neither rustdoc nor clippy surfaced them:
`Outcome::winner` and `EventBuilder::weights`. Both now carry the heading, and
`Outcome::winner` gained the note that it ties every loser, so `n >= 3` needs a
positive p_draw — the crate's easiest error to hit by accident.

The 43 mechanical fixes (31 `#[must_use]` on pure accessors, 11 missing
backticks) were applied with `cargo clippy --fix`. `#[must_use]` on Gaussian's
arithmetic and on `posteriors()` matters: discarding those results is always a
bug, and until now nothing said so.

Also documented why `[profile.release] debug = true` exists — cargo-flamegraph
needs the symbols, and library profile settings are ignored downstream, so it
reads as an oversight without the note.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T5SYDExxL4vZgvunrcNSMc
2026-08-27 17:43:31 +02:00
logaritmisk 07285283b6 chore: Release trueskill-tt version 0.2.0 2026-08-27 17:37:31 +02:00
logaritmiskandClaude Opus 5 b73cf0145a chore: dual-license MIT OR Apache-2.0
The Rust ecosystem convention, and what kickscore, xy and saphyr already use —
kickscore being the closest sibling to this crate.

LICENSE-APACHE is the canonical 201-line Apache-2.0 text with the appendix
left as the unfilled template, which is the form Rust crates ship. LICENSE-MIT
carries the copyright line. Both are picked up by cargo automatically and
appear in the packaged crate.

This also unblocks the release workflow. `cargo publish` does not check the
license field when the target is an alternative registry, but cargo-release
does, and refuses outright:

    error: trueskill-tt is missing the following fields:
             license || license-file

--no-verify does not bypass it. So `just release` was inert without this,
whatever cargo publish alone would have accepted.

README gains the conventional dual-license section and the contribution note
that dedicates inbound contributions under the same terms.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T5SYDExxL4vZgvunrcNSMc
2026-08-27 17:35:42 +02:00
logaritmiskandClaude Opus 5 56ff01074f docs(cargo): correct the licence note — kellnr does not require one
The note claimed `cargo publish` rejects a crate without `license`. That is
true only for crates.io; publishing to an alternative registry does not check
it, verified by a dry run against kellnr that packages and verifies cleanly.

Staying unlicensed is a deliberate choice, so the note now says that and states
the actual consequence — all-rights-reserved by default — rather than a
mechanical blocker that does not exist.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T5SYDExxL4vZgvunrcNSMc
2026-08-27 17:30:59 +02:00
logaritmiskandClaude Opus 5 8d47e54a8a chore: keep the 48 MB ATP dataset out of the published crate
`cargo publish --dry-run` packaged 48.1 MiB (8.6 MiB compressed) for a library
whose source is 312 KB. All of it was examples/atp.csv, a tennis dataset the
atp example reads.

examples/atp.rs opens it by relative path at runtime rather than include_str!,
so excluding the data still compiles and `cargo package --verify` still builds
every target — the example just needs the file fetched from the repo to run.

Packaged size is now 360.8 KiB / 83.1 KiB compressed, a 133x reduction. Every
consumer would otherwise have paid that download.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T5SYDExxL4vZgvunrcNSMc
2026-08-27 17:29:44 +02:00
logaritmiskandClaude Opus 5 7de092ba12 chore: target releases at the private kellnr registry
Mirrors the textus setup, adapted for a single crate rather than a workspace.

Cargo.toml gains publish = ["kellnr"], which does double duty: it points
cargo-release at the private registry and makes an accidental `cargo publish`
to crates.io a hard error rather than an irreversible mistake.

.cargo/config.toml is committed rather than left to a per-user
~/.cargo/config.toml. Without it a fresh clone, a new machine, or CI fails
with "registry index was not found in any configuration: kellnr" before
compiling anything. The index URL is not a secret; the token stays in
~/.cargo/credentials.toml, or CARGO_REGISTRIES_KELLNR_TOKEN in CI.

release.toml flips publish from false to true and pins push = false, so the
Justfile recipe pushes last — after tags and publish have both succeeded.
The git-cliff pre-release hook is unchanged.

cliff.toml gained a Breaking Changes group. Its commit_parsers matched on type
alone with conventional_commits = false, so `refactor!: remove the inert online
flag` rendered as an ordinary Refactor bullet and the break was invisible in
the generated changelog. The new parsers match a `!` subject and a
BREAKING CHANGE body, and must precede the type parsers because the first match
wins. The unreleased section now opens with the break, which matters because
the next release is the one that removes HistoryBuilder::online.

The release recipe runs `just ci` before cutting: cargo-release only
verify-compiles the packaged crate and publishing cannot be undone, and the
release profile is where this crate's defects have historically hidden.

Still unpublishable: Cargo.toml has no `license`. That is a deliberate TODO,
not an oversight.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T5SYDExxL4vZgvunrcNSMc
2026-08-27 17:28:51 +02:00
logaritmiskandClaude Opus 5 eeb43e3be1 fix: close out four small issues and pin #27's repro
#29 — log_evidence and log_evidence_for took &mut self while mutating
nothing. Loosening them to &self is not source-breaking for ordinary callers
(a &mut reborrows as & transparently) and brings them in line with the
filtered_* accessors added last week.

Not the mechanical change it looked like: under the rayon feature the closure
in log_evidence_internal captured all of &self rather than just the competitor
store, which drags KeyTable<K> in and demands K: Sync from every caller. That
compiled while the method took &mut self and stopped compiling the moment it
did not. Binding `let agents = &self.agents;` before the closure narrows the
capture; the comment there says why, because the next person to inline it will
reintroduce the bound.

#31 — TimeSlice::add_events constructed Skill with ..Default::default() while
filtered_step spells every field out. The design relies on a new Skill field
being a compile error at construction sites rather than a silent default, and
that tripwire only fired at one of the two. Now both.

#28 — log_evidence_internal's `forward` flag is a genuine forward-only
quantity only on a history that has never been converged, because iteration
alternates sweeps and the likelihood feeding the forward message absorbs
backward information from the second iteration onward. Documented, with a
pointer to filtered_log_evidence for the quantity that survives convergence.
That trap is one function away from the one #19 was about.

#23 — color_greedy carried #[allow(dead_code)] despite being called by
recompute_color_groups: a mute button on a live function, which is the
specific complaint in that issue.

#27 was already fixed — the guard landed in f4e2922 and the issue was filed
against 7742b2b, which merge-base confirms predates it — but nothing pinned
it. Added the issue's own reproduction, which matters because the two profiles
fail differently and a debug-only test would miss the release path. Removing
both guards reproduces the issue verbatim: "attempt to subtract with overflow"
in debug, "index out of bounds: the len is 0 but the index is
18446744073709551615" in release.

Also amended the filtered-estimates spec (#30): the tolerance-not-bit-identity
caveat is conservative. Forcing the scratch onto the sequential sweep instead
of the grouped one — a far larger perturbation than a permuted event order —
still agrees within 1e-8 under tight convergence.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T5SYDExxL4vZgvunrcNSMc
2026-08-27 17:21:05 +02:00
logaritmiskandClaude Opus 5 69ddebe21d docs: state filtered accessor cost and evidence semantics precisely
filtered_learning_curve's signature mirrors learning_curve, which is cheap
per key — so the mirroring trained callers to assume this one is too. It is
a full forward pass per call, making the natural loop over competitors
O(competitors * events). The doc now says so in complexity terms and points
multi-key callers at the plural form.

filtered_log_evidence claimed each event is scored "using only what was
known before it". That is exact for a slice holding one event, but events
sharing a timestamp inform each other through the within-slice sweep, so
the honest claim is "before that time". The behaviour is deliberate and
matches log_evidence's own convention; only the promise was too strong.

This branch exists because a feature's documentation was quietly false.
Shipping it with two more overstated doc comments would be a poor joke.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T5SYDExxL4vZgvunrcNSMc
2026-08-27 16:59:42 +02:00
logaritmisk 9c39d1e681 test: pin the invariants that make filtered estimates trustworthy
The bracket test proves the feature works on one fixture. These pin the bug
class:

- Invariance to converge(). This is the one that matters. Reading
  skill.forward instead of the carried message makes it fail immediately,
  because converge() alternates sweeps and contaminates skill.forward with
  backward information from the second iteration onward. That is the
  property a stored field cannot have, and the reason issue #19's proposed
  fix would not have worked.
- Invariance to ingestion order, the crate's standing invariant.
- One slice has no future to propagate back, so filtered equals smoothed.
- Empty history yields zero and empty maps.

Agreement is to 1e-8 under tight convergence rather than bit-identity:
iteration recomputes the colour partition only when from == 0, so an
incrementally built slice keeps insertion order until the first converge()
reorders it, and the scratch clone inherits whichever order it finds. Same
fixed point, different path to it.
2026-08-27 16:37:21 +02:00
logaritmisk 50e11cfbfa feat: add filtered learning curves
learning_curve returns post-convergence posteriors, so every point is
smoothed: the estimate at a given date incorporates rounds played years
later. On ustat's data that starts six players' curves already spread apart
at sigma 0.9-1.6 against a prior of 6.0, barely moving thereafter.

filtered_learning_curve plots the same competitor on forward-only
information, so everyone starts at the prior and fans out. It could not be
reconstructed from the public API before: a caller could only refit over
events[0..k] for every k, which is O(n^2) fits for something one forward
pass already computes.
2026-08-27 16:30:19 +02:00
logaritmisk d4af048914 feat: add filtered_log_evidence
Scores every event on what was known before it, rather than on priors that
carry information from events which had not happened yet. This is the
quantity HistoryBuilder::online promised and never delivered.

The pass walks slices in time order carrying its own forward messages, and
per slice runs the unmodified production sweep on a scratch copy whose
backward message is left improper. Reusing iterate_to_convergence rather
than reimplementing inference means a competitor playing twice at one time
is handled by the same within-slice EP that converge() uses, instead of
being approximated the way the old evidence paths approximated it.

Nothing is stored on Skill and nothing on self is mutated, so the result is
independent of whether converge() has run — the property a stored field
cannot have.
2026-08-27 16:21:01 +02:00
logaritmisk bf9d964cae refactor!: remove the inert online flag
Skill.online was initialised to N_INF and assigned nowhere, so
HistoryBuilder::online(true) made every rating improper and log_evidence()
reported n * ln(0.5) — every game scored as a coin flip. The value is finite
and plausible, which is why it went unnoticed.

The default was false, so no existing result changes. A working replacement
lands next; a stored field cannot hold the quantity, because converge()
alternates sweeps and contaminates skill.forward with backward information
from the second iteration onward.

Also renames a test binding from ..._online to ..._forward: it passes the
forward flag, and the two senses being conflated is how this survived.
2026-08-27 16:11:37 +02:00
logaritmiskandClaude Opus 5 187aede924 docs: implementation plan for filtered estimates
Five tasks: delete the inert online machinery, add filtered_log_evidence,
add the two learning-curve methods, pin the invariants, record the API break.

Two spec corrections fell out of writing it. The spec claimed filtered results
would be bit-identical before and after converge(); they cannot be. iteration
recomputes the colour partition only when from == 0, so a slice built by
repeated appends keeps insertion order until the first converge() reorders it,
and the scratch clone inherits whichever order it finds — same fixed point,
different path. Corrected to agreement within 1e-8 under tight convergence,
matching the house pattern in tests/ingestion_equivalence.rs. The spec also
declared filtered_pass as Vec<(T, Vec<(Index, Gaussian)>)>, which cannot carry
the evidence its own step 3 harvests; it returns Vec<(T, FilteredStep)>.

CHANGELOG.md is generated by git-cliff, so the spec's "CHANGELOG records the
API break" cannot be satisfied by editing the file — it regenerates. Task 5
records the break through the commit subject and verifies the generated output
instead. cliff.toml has no breaking-change parser at all, which the task is
told to report rather than work around.

An adversarial reviewer checked the plan against the source before this commit
and found four real defects, all in plan text, none in the design:

- Two prescribed mutations provably could not fail their named tests. The
  learning-curve mutation altered only what filtered_pass writes after a slice,
  while the test inspected filtered[0], which is computed from an empty message
  map. Fixed by asserting monotonic mu across the whole curve.
- The ingestion-order fixture used four distinct timestamps, giving one event
  per slice — the exact degenerate shape ingestion_equivalence.rs documents as
  the weak case, making the assertion true by construction. Fixed to several
  events per timestamp with shared competitors.
- filtered_learning_curves was never asserted for content, only for emptiness
  on an empty history.
- A doc comment restated learning_curves' claim that key(idx) is O(n) and the
  method O(n^2). KeyTable::key is self.reverse.get(idx.0) — O(1) — and the
  type's own doc says so. The claim predates reverse becoming a Vec. The plan
  now corrects the original at history.rs:323 rather than copying it.

The reviewer confirmed the central claim by tracing the call graph: N_INF is
{pi: 0, tau: 0} and Mul is a natural-parameter add, so it is an exact
multiplicative identity, and the only write to skill.backward in the crate is
in new_backward_info, reachable only from History::iteration and never from
iterate_to_convergence under either rayon cfg.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T5SYDExxL4vZgvunrcNSMc
2026-08-27 16:04:16 +02:00
logaritmiskandClaude Opus 5 4fde482e48 docs: spec for filtered (forward-only) estimates
`HistoryBuilder::online(true)` is inert: it flips a flag that reaches
`Item::within_prior`, which reads `Skill.online` — a field initialised to
`N_INF` and assigned nowhere. So `log_evidence()` under that setting reports
`n * ln(0.5)`, every game scored as a coin flip. The number is finite and
plausible, which is why nothing caught it.

Issue #19 proposed populating the field during the forward pass. That does
not work, and the reason shapes the whole design. `new_forward_info` sets
`skill.forward` from the previous slice's `forward_prior_out`, which is
`skill.forward * skill.likelihood`; `History::iteration` alternates backward
and forward sweeps, so from the second iteration onward that likelihood has
already absorbed backward information. After `converge()`, `skill.forward`
is a smoothed quantity — and so is anything written from it.

The same reasoning condemns the neighbouring `forward: bool` flag, which is
a filtering quantity only on a history that was never converged. That is why
the test at history.rs:1183 can assert the two evidences are equal. Left
alone here; recorded as a follow-up.

The design is a read-only forward-only pass instead: walk slices in time
order carrying their own forward messages, and per slice build a scratch
clone whose `backward` is `N_INF`, then run the unmodified production sweep
on it. Reusing `iterate_to_convergence` rather than reimplementing inference
means a competitor playing twice at one time is handled by the same
within-slice EP that `converge()` uses, instead of being approximated the
way today's evidence paths approximate it. Nothing is stored on `Skill`,
which drops 16 bytes and helps #17 regardless.

Three methods ship — `filtered_log_evidence`, `filtered_learning_curves`,
`filtered_learning_curve` — all taking `&self`. The second consumer is
ustat, whose learning curves start already collapsed to sigma 0.9-1.6
against a prior of 6.0 because every point is smoothed; the filtered view
cannot be reconstructed from the public API today except by O(n^2) refits.

The red test brackets the issue's own fixture strictly between 5*ln(0.5) and
the batch evidence, so neither "still inert" nor "accidentally smoothed"
passes. The invariant that would have caught this bug class is that filtered
results are identical before and after `converge()` — exactly what a stored
field cannot give.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T5SYDExxL4vZgvunrcNSMc
2026-08-27 15:42:54 +02:00
logaritmiskandClaude Opus 5 9e8515b7cd docs: refresh README and CLAUDE.md; add ingest benchmark
The CLAUDE.md architecture section still described the pre-redesign engine:
its data flow named `Batch`, `Agent`, `Player` and `message.rs`, none of
which have existed since T2, and the public API it listed did not match
`lib.rs`. It is the first thing a fresh session reads, so it was actively
misleading. Rewritten against the current module layout, with the invariants
that are easy to violate — ties needing a positive `p_draw`, NaN never being
convergence, log-space evidence, color contiguity, `forbid(unsafe_code)`,
and ingestion-order equivalence — written down.

The README Todo list had five entries that were already done, including
"Time needs to be an enum": `Time` has been a trait since T2, and the
`batch::compute_elapsed()` it pointed at no longer exists. The genuinely
open item — cross-checking `quality()` against sublee/trueskill — stays.

`benches/ingest.rs` measures one-event-per-call against a single batched
call. The rest of the suite only measured batched construction, which is why
the quadratic fixed earlier on this branch went unnoticed for so long.

`TimeSlice::log_evidence` also hashes its target set once instead of
scanning the slice per player per event, so `log_evidence_for` with many
keys is no longer quadratic.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DnsaJg74eNSva3PJjK2eej
2026-08-04 22:05:49 +02:00
logaritmiskandClaude Opus 5 9506fed4b3 chore: add CI, crate metadata, and crate-level documentation
There was no CI of any kind — no workflows directory at all — despite a full
release pipeline (release.toml, cliff.toml, a maintained CHANGELOG, three
tagged releases). The workflow covers the feature combinations that actually
have distinct behaviour, including a release-profile job: `debug_assert!` is
compiled out there, which is exactly where the validation this branch added
has to hold, and a debug-only suite would never have seen it. Determinism is
checked at RAYON_NUM_THREADS of 1, 2, 4 and 8.

The Justfile gains test/lint/fmt/determinism recipes so the same checks run
locally with one command, and `just ci` runs the lot.

`Cargo.toml` had only name, version and edition, so `cargo publish` would
have been rejected. Added description, repository, readme, keywords,
categories, exclude, and `rust-version = "1.85"` — the edition-2024 floor,
now verified by a CI job. Two let-chains introduced earlier on this branch
would have pushed that to 1.88; they are rewritten to keep the floor where
it was.

`src/lib.rs` had no `//!` header at all, so the docs.rs landing page would
have been a bare symbol list — conspicuous given every other module has one.
It now explains what Through Time does differently, and carries three
runnable examples (which `cargo test --doc` checks, where previously there
was nothing to check), including the draw/p_draw interaction that is the
easiest way to get an error out of this crate.

`cargo publish --dry-run` now packages and verifies cleanly. The only
remaining blocker is `license`, which is yours to choose — noted as a TODO
in the manifest rather than picked unilaterally.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DnsaJg74eNSva3PJjK2eej
2026-08-04 22:03:39 +02:00
logaritmiskandClaude Opus 5 6030dc78de refactor: unify convergence defaults, validate builders, clear dead code
Convergence configuration had two disagreeing sources of truth and one
misleading report:

- `EpsilonOrMax::default()` capped at 10 iterations while
  `ConvergenceOptions::default()` allowed 30, and which applied depended on
  whether inference went through `run_chain` or a `Schedule`. The schedule
  default now derives from `ConvergenceOptions`.
- A graph with no iterating factors reported `converged: false` with an
  infinite step, despite being at its fixed point after the setup pass. It
  now reports converged with a zero step.
- `TimeSlice::iterate_to_convergence` hard-coded an epsilon and a
  20-iteration cap matching neither. It reads `self.convergence` and is
  scoped to `#[cfg(test)]`, which is all it was ever used by.

`HistoryBuilder::p_draw` and `::convergence` now validate their arguments
like `score_sigma` already did, instead of accepting a negative `p_draw` or
an `alpha` of zero — the latter leaves every EP update unapplied, so
inference silently returns the priors.

Removing the `#[allow(dead_code)]` masks let the compiler report what they
were hiding: four `OwnedGame` fields that were stored and never read, two
`ColorGroups` helpers and three `SkillStore` helpers used only by tests, and
`iterate_to_convergence` above. Test-only items are now `#[cfg(test)]` and
the unread fields are gone.

Also exported `HistoryBuilder`, which was public but unreachable — callers
could chain `History::builder()` but could not name the type — and added
`Rating::{prior, beta, drift}` and `Index::get`, so handles the API hands
out can be read back.

Two goldens moved, both convergence residuals rather than exact values:
`iterate_to_convergence` now runs to 30 iterations instead of 20, landing
nearer the symmetric truth of 25.0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DnsaJg74eNSva3PJjK2eej
2026-08-04 22:01:49 +02:00
logaritmiskandClaude Opus 5 355cdb7e05 perf(gaussian): drop the sqrt round-trip from variance-space operations
`Add`, `Sub`, `exclude` and `forget` combined variances by way of standard
deviations: `sigma()` takes a square root, `.powi(2)` squares it away,
`var.sqrt()` takes another, and `from_ms` squares that one back. Three roots
to compute a value that is `1/pi` all along.

They now go through `variance()` and a new `from_mv(mu, var)`, which skip
both conversions. `Sub` is the hot one — `RankDiffFactor::propagate` is
`a - b`, run for every adjacent team pair on every forward and backward
sweep of every EP iteration.

`run_chain` also stopped recomputing each team's weighted performance in the
likelihood loop; the fold is already in `arena.team_prior`, indexed by the
sorted position the loop has in hand. Each `performance()` is itself a
`forget`, so the duplicate cost scaled with players per team.

Measured on this machine, before and after, same fixtures:

    Batch::iteration          23.57us -> 19.31us   (-18%)
    scored_history_60_events   1.071ms -> 983us    (-8%)

The `Gaussian::add`/`sub` microbenchmarks cannot resolve the change: they
sit at ~234ps against a ~218ps floor that `mul`/`div` also hit, so the
harness overhead dominates a single operation.

One golden moved. Two identical competitors drawing must land on their
shared prior mean exactly, by symmetry; the root-free path now returns
25.0 where the reference transcription recorded 24.999999 — that value
rounded to six decimals. Asserting a six-decimal transcription at
epsilon 1e-6 left no headroom, so the expectation is now the exact value.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DnsaJg74eNSva3PJjK2eej
2026-08-04 21:58:30 +02:00
logaritmiskandClaude Opus 5 06b6a68499 fix(rayon): remove the aliasing unsafe from the parallel sweep
The parallel color-group sweep passed a `*mut SkillStore` through a `usize`
and cast it back inside the rayon closure, so every worker materialised its
own `&mut SkillStore` to the same store. Two live `&mut` to one object is an
aliasing violation whatever the workers subsequently touch — `&mut` carries
`noalias` down to LLVM — and laundering the pointer through `usize` also
discarded provenance. The existing SAFETY comment argued element
disjointness, which is true and is why nothing miscompiled in practice, but
it is not the property the aliasing rules ask about.

Events in a color group touch disjoint agents, so none can observe another's
writes. That makes the sweep separable rather than merely safe-in-practice:
`Event::compute` runs inference over shared `&self.skills` with no mutation,
and `Event::apply` folds the results in afterwards in index order. No
`unsafe`, no aliasing argument, and the apply order does not depend on which
worker finished first, so results stay bit-identical across thread counts.

The crate now contains no `unsafe` at all, locked in with
`#![forbid(unsafe_code)]`.

Splitting compute from apply also removes the duplicated sweep body: the
`from > 0` branch of `TimeSlice::iteration` was a verbatim copy of
`iteration_direct`, and both now share one implementation.

Cost, measured on the three `history_converge` workloads (sequential vs
parallel, this machine):

    500x100@10perslice     4.02ms -> 4.21ms
    2000x200@20perslice   19.70ms -> 19.76ms
    1v1-5000x50000        11.75ms -> 10.46ms

The deferred apply gives back part of the parallel win on the only workload
where rayon ever helped (1.12x here, against the 1.3x T3 reported), and the
sequential path is unchanged. Trading a fraction of a 1.3x speedup on one
pathological shape for the removal of undefined behaviour is the right side
of that bargain.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DnsaJg74eNSva3PJjK2eej
2026-08-04 21:54:34 +02:00
logaritmiskandClaude Opus 5 c088214fed fix(history): stop reprocessing the slice that was just appended to
`add_events_with_prior` advanced `k` past the slice it had written when it
created a new one, but not when it appended to an existing one. The trailing
forward-refresh loop therefore started *on* the slice just modified and ran
`new_forward_info` over it again.

That is not merely redundant work. The loop immediately above it sets each
agent's message to `forward * likelihood` for that slice, and
`new_forward_info` then assigns `skill.forward = message.forget(drift)` —
folding the slice's own likelihood back into its own forward prior. The
skills it produced depended on how events had been batched.

Ingesting one event at a time now converges to the same fixed point as
ingesting the same events in a single call, which it previously did not:
for five events sharing a timestamp, competitor `a` converged to
mu=7.44 sigma=3.90 batched versus mu=7.99 sigma=3.10 incrementally. Both
runs had converged; the gap was not a convergence residual.

The numerical goldens never caught this because they all ingest in one call
with a distinct timestamp per event, so the append-to-existing-slice branch
is never taken. `tests/ingestion_equivalence.rs` covers it directly, and
asserts convergence before comparing so that a residual cannot be mistaken
for agreement.

Removing the redundant re-inference also removes the dominant cost of
incremental ingestion, which was quadratic in the number of events already
in the slice:

    events   before     after    speedup
       500   45.8ms     1.1ms       42x
      1000  179.5ms     2.8ms       64x
      2000  721.8ms     9.9ms       73x
      4000    2.9s     35.4ms       82x

Ingesting one at a time is now 1.8x a single batched call, down from 148x.

Two supporting changes are included:

- Color groups are rebuilt lazily rather than on every append. Nothing
  reads the partition between an append and the next full sweep, so the
  per-append rebuild was pure waste.
- `ColorGroups::groups_are_contiguous` is asserted after each rebuild and
  in `color_range`. The parallel sweep derives one `&mut` sub-slice per
  color from those ranges and relies on them being disjoint; that invariant
  was established by construction but never checked.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DnsaJg74eNSva3PJjK2eej
2026-08-04 21:51:02 +02:00
logaritmiskandClaude Opus 5 0f1a1b8911 fix(evidence): accumulate in log space and floor the per-link value
Per-link evidence was multiplied in linear space and logged only at the
end. Each link contributes a probability in (0, 1], so the product over an
n-team game decays geometrically: around a thousand links it flushes to
exactly 0.0 and `ln(0.0)` is `-inf`, which then propagates through the sum
in `History::log_evidence_internal` and takes the whole history with it.
`Game::free_for_all` builds one team per player, so this is reachable at
the competitor counts the T3 benchmarks target.

`Game`, `OwnedGame`, and `time_slice::Event` now carry `log_evidence`
directly, summed over links rather than multiplied then logged.

The cached per-link evidence is also floored at `f64::MIN_POSITIVE`. It
could legitimately reach zero or go negative: `1.0 - cdf(..)` rounds to
zero for a near-certain outcome, and the `erfc` approximation carries
~1e-7 error so `cdf` can exceed 1.0 and make the difference negative —
`ln` of which is NaN.

Existing log-evidence goldens are unchanged, confirming the accumulation
is numerically equivalent in the range where the old form worked.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DnsaJg74eNSva3PJjK2eej
2026-08-04 21:43:57 +02:00
logaritmiskandClaude Opus 5 0d32690fcc fix(quality): support any number of rating groups
`quality()` was two-group-only in three separate ways, and
`History::predict_quality` inherited all of them.

- The contrast-matrix column counter tracked two positions with two
  variables that only agree on the first row, so three or more groups wrote
  past the end of the row and panicked with an out-of-bounds index. The
  negative block always begins immediately after the positive one, so the
  second counter is unnecessary.
- `Matrix::inverse` was implemented only for the 1x1 case and otherwise
  `panic!("eh, okey")`. It now uses LU decomposition with partial pivoting,
  which also replaces the recursive cofactor `determinant` — that was O(n!)
  and allocated a `Vec` per minor, so a 10-team match needed 362,880 terms.
- Degenerate inputs (zero groups, one group, empty groups) underflowed or
  produced NaN. They now assert with a message naming the requirement.

`Matrix` also gains dimension checks on multiply/add and bounds checks on
indexing, and loses the now-unused `adjugate`/`minor` cofactor path.

The two-group golden is unchanged. N-group behaviour is covered by
invariants — permutation invariance, and quality falling as a skill gap
widens — since no reference values were available to compare against; the
sublee/trueskill cross-check remains open.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DnsaJg74eNSva3PJjK2eej
2026-08-04 21:42:04 +02:00
logaritmiskandClaude Opus 5 6b8bd786d7 style: make NaN rejection explicit in score_sigma validation
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DnsaJg74eNSva3PJjK2eej
2026-08-04 21:38:48 +02:00
logaritmiskandClaude Opus 5 f4e2922d59 fix: reject ties without draw probability; never report NaN as converged
A tie with `p_draw == 0.0` produced NaN posteriors in release builds and
`converge()` reported `converged: true`, because every comparison against
NaN is false and `tuple_gt` therefore read NaN as "below epsilon".

Two independent defects, fixed together:

- Ingestion now rejects tied outcomes when the draw probability is zero,
  promoting the existing `debug_assert!` in `Game::ranked_with_arena` to a
  real `InferenceError::TieWithoutDrawProbability`. Validation sits in
  `add_events_with_prior`, the chokepoint every route reaches — including
  `record_draw`, which bypasses `Outcome` entirely.
- `converge()` treats a non-finite step as failure and returns
  `InferenceError::NonFiniteResult` rather than claiming convergence.

Also in this change:

- `History::converge()` on an empty history returned a `usize` underflow
  panic from `0..len()-1`; it now short-circuits to a zero-iteration report.
- `Outcome::scores_with_sigma` no longer panics on a non-positive sigma;
  the value is validated at ingestion so callers get an error instead.
- `InferenceError` gains `WrongOutcomeKind`, replacing the misuse of
  `MismatchedShape` for variant mismatches (which rendered as the nonsense
  "expected length 0, got 0"), and is now `#[non_exhaustive]`.

Note `Outcome::winner(w, n)` for n >= 3 ties every loser, so those events
now require a positive `p_draw`. They previously returned NaN.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DnsaJg74eNSva3PJjK2eej
2026-08-04 21:38:29 +02:00
71 changed files with 11937 additions and 1194 deletions
+15
View File
@@ -0,0 +1,15 @@
# `Cargo.toml` sets `publish = ["kellnr"]`, so `cargo publish` targets the
# private registry and refuses crates.io. Cargo needs that registry's index
# declared to resolve the name.
#
# Committed rather than left to a per-user `~/.cargo/config.toml` so the repo
# is self-contained: a fresh clone, a new machine, or CI would otherwise fail
# with
#
# error: registry index was not found in any configuration: `kellnr`
#
# Index URL only — it is not a secret. Publish tokens live in
# `~/.cargo/credentials.toml` (per-user, never committed) or, in CI, in
# `CARGO_REGISTRIES_KELLNR_TOKEN`.
[registries.kellnr]
index = "sparse+https://crates.aceofba.se/api/v1/crates/"
+87
View File
@@ -0,0 +1,87 @@
name: CI
on:
push:
branches: [main]
pull_request:
env:
CARGO_TERM_COLOR: always
RUSTFLAGS: -D warnings
jobs:
test:
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
include:
# The build most consumers get.
- name: default
features: ""
profile: ""
# Most numerical goldens need `approx` for assert_ulps_eq.
- name: approx
features: "--features approx"
profile: ""
# The parallel path, including tests/determinism.rs.
- name: rayon
features: "--features approx,rayon"
profile: ""
# Critical: debug_assert! is compiled out here, which is where the
# tie/p_draw and score_sigma validation actually has to hold.
- name: release
features: "--features approx"
profile: "--release"
name: test (${{ matrix.name }})
steps:
- uses: actions/checkout@v4
- uses: dtolnay/rust-toolchain@stable
- uses: Swatinem/rust-cache@v2
- run: cargo test ${{ matrix.profile }} ${{ matrix.features }}
- run: cargo test ${{ matrix.profile }} ${{ matrix.features }} --doc
determinism:
name: determinism across thread counts
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: dtolnay/rust-toolchain@stable
- uses: Swatinem/rust-cache@v2
# Posteriors must be bit-identical regardless of how many rayon workers
# run the color-group sweep.
- run: |
for threads in 1 2 4 8; do
echo "== RAYON_NUM_THREADS=$threads =="
RAYON_NUM_THREADS=$threads cargo test --release \
--features approx,rayon --test determinism
done
lint:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: dtolnay/rust-toolchain@stable
with:
components: clippy
- uses: Swatinem/rust-cache@v2
- run: cargo clippy --all-targets --all-features -- -D warnings
format:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
# rustfmt.toml uses nightly-only options (imports_granularity).
- uses: dtolnay/rust-toolchain@nightly
with:
components: rustfmt
- run: cargo +nightly fmt --check
msrv:
name: minimum supported Rust version
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: dtolnay/rust-toolchain@1.85.0
- uses: Swatinem/rust-cache@v2
- run: cargo check --all-targets --features approx,rayon
+1
View File
@@ -7,3 +7,4 @@
NOTEPAD.md
/.claude
proptest-regressions/
+179
View File
@@ -2,6 +2,181 @@
All notable changes to this project will be documented in this file.
## 0.5.0 - 2026-09-08
### Breaking Changes
- feat!: name the unknown key, expose tail probabilities, flag short fits
- refactor!: remove the factor-graph surface nothing used, add try_winner
### Bug Fixes
- fix(test): the ingestion-order property was comparing two truncated fits
### Documentation
- docs: record that the event log is the source of truth, and why
### Features
- feat: add UnknownKeys::Prior, and explain why there is no Skip
- feat: add History::posterior_of for a linear combination of competitors
- feat: add History::predict_margin for scored matchups
- feat: add expected_variance_reduction for scored active learning
### Styling
- style: factor the event-pair type out of the reconvergence fixture
- style: use arrays rather than vec! in the calibration fixture
### Testing
- test: pin that re-convergence is path-independent
- test: calibrate the marginals against the exact posterior
- test: pin what an additive model does to combined uncertainty
## 0.4.2 - 2026-09-07
### Bug Fixes
- fix: replace the erfc approximation with libm, for free
- fix: route every transcendental through libm, and combine sigmas with hypot
### Miscellaneous Tasks
- chore: Release trueskill-tt version 0.4.2
### Testing
- test: localise the erfc_inv tail residual to the caller's argument
## 0.4.1 - 2026-09-07
### Bug Fixes
- fix: correct erfc_inv's sign error and keep evidence in log space
### Miscellaneous Tasks
- chore: Release trueskill-tt version 0.4.1
### Testing
- test: pin quality()'s N-group closed form, closing the README cross-check
## 0.4.0 - 2026-09-07
### Breaking Changes
- feat!: N-team outcome prediction with draw mass, replacing the 2-team panic
- refactor!: close the remaining API gaps from #21
- fix!: apply competitor configuration whenever it is supplied
### Bug Fixes
- fix(release): skip the changelog hook during a dry run
- fix: stop destroying tail precision in evidence and truncation
- fix: reject convergence options that silently disable inference
### Documentation
- docs: correct drifted documentation and compile the README in CI
### Features
- feat: add expected information gain for active matchup selection
- feat: let observers be shared, boxed, or borrowed
### Miscellaneous Tasks
- chore: Release trueskill-tt version 0.4.0
## 0.3.0 - 2026-09-01
### Breaking Changes
- refactor!: make Competitor::message an Option, and compute_elapsed loud
- refactor!: replace emptiness-as-sentinel with Option for results and weights
- refactor!: remove ConvergenceReport::slices_skipped
### Bug Fixes
- fix: enforce EventBuilder weight/team length in release
### Documentation
- docs: complete the public API documentation contract
### Features
- feat: allow drift to vary per competitor via Member::with_drift_scale
### Miscellaneous Tasks
- chore: ignore proptest regression seed files
- chore: Release trueskill-tt version 0.3.0
### Performance
- perf: stop cloning inference inputs in OwnedGame and ingestion
- perf: make the per-slice SkillStore compact instead of dense
### Testing
- test: add property-based tests, a shared finiteness helper, and boundary inputs
## 0.2.0 - 2026-08-27
### Breaking Changes
- refactor!: remove the inert online flag
### Bug Fixes
- fix: reject ties without draw probability; never report NaN as converged
- fix(quality): support any number of rating groups
- fix(evidence): accumulate in log space and floor the per-link value
- fix(history): stop reprocessing the slice that was just appended to
- fix(rayon): remove the aliasing unsafe from the parallel sweep
- fix: close out four small issues and pin #27's repro
### Documentation
- docs: refresh README and CLAUDE.md; add ingest benchmark
- docs: spec for filtered (forward-only) estimates
- docs: implementation plan for filtered estimates
- docs: state filtered accessor cost and evidence semantics precisely
- docs(cargo): correct the licence note — kellnr does not require one
### Features
- feat: add filtered_log_evidence
- feat: add filtered learning curves
### Miscellaneous Tasks
- chore: add CI, crate metadata, and crate-level documentation
- chore: target releases at the private kellnr registry
- chore: keep the 48 MB ATP dataset out of the published crate
- chore: dual-license MIT OR Apache-2.0
- chore: Release trueskill-tt version 0.2.0
### Performance
- perf(gaussian): drop the sqrt round-trip from variance-space operations
### Refactor
- refactor: unify convergence defaults, validate builders, clear dead code
### Styling
- style: make NaN rejection explicit in score_sigma validation
### Testing
- test: pin the invariants that make filtered estimates trustworthy
## 0.1.2 - 2026-06-12
### Bug Fixes
@@ -32,6 +207,10 @@ All notable changes to this project will be documented in this file.
- feat(outcome): per-event score_sigma override on Outcome::Scored
- feat(event_builder): expose scores_with_sigma fluent method
### Miscellaneous Tasks
- chore: Release trueskill-tt version 0.1.2
### Refactor
- refactor: dedupe Game::likelihoods and likelihoods_scored via run_chain
+114 -26
View File
@@ -5,42 +5,130 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co
## Commands
```bash
cargo build # Build the library
cargo test --lib # Run all library tests
cargo test --lib <test_name> # Run a single test by name
cargo test --lib -- --nocapture # Run tests with stdout output
cargo clippy # Lint
cargo bench # Run benchmarks (criterion)
just test # Full suite across every feature combination CI checks
just check # Fast inner loop: cargo test --features approx
just lint # clippy, warnings denied
just fmt # ALWAYS nightly — rustfmt.toml uses nightly-only options
just determinism # Bit-identical posteriors at RAYON_NUM_THREADS 1/2/4/8
just ci # Everything CI runs
cargo test --lib <test_name> # A single test by name
cargo bench # Criterion benchmarks
```
The `approx` feature enables `approx::AbsDiffEq` for `Gaussian`:
```bash
cargo test --features approx
```
**Run tests in release too.** `debug_assert!` is compiled out there, and that
is where several defects have hidden — a debug-only run is not evidence.
`just test` includes a release job.
### Feature flags
- `approx``approx::AbsDiffEq` etc. for `Gaussian`. Most numerical goldens need it.
- `rayon` — opt-in parallel within-slice sweep and per-slice query passes.
## Working rules
- **Investigate before implementing.** Measure the actual behaviour first —
against an analytic reference where one exists. Several "obvious" fixes in
this repo turned out to be wrong in sign or unnecessary, and the measurement
is what caught them.
- **Fix the root issue, not the symptom.** A clamp that hides an underflow, or
a tolerance loosened to make a test pass, is a defect deferred.
- **Scout crates.io before hand-rolling numerics.** Check accuracy against an
independent reference rather than trusting downloads: `puruspe` has 1.4M
downloads and is 346 ULP off in the tail, where `libm` is 1. Fewer
dependencies is preferable, not mandatory — take the dependency when it is
measurably better.
## Architecture
This is a Rust port of [TrueSkillThroughTime.py](https://github.com/glandfried/TrueSkillThroughTime.py) — a Bayesian skill rating system that tracks skill evolution over time using Gaussian message passing.
A Rust port of [TrueSkillThroughTime.py](https://github.com/glandfried/TrueSkillThroughTime.py):
Bayesian skill rating that infers skill at every point in time, propagating
evidence both forward and backward across a history.
### Data flow
Ingestion (public types, `event.rs`):
```
History → Batch[] → Game[] → teams/players
Event<T, K> → Team<K>[] → Member<K>[]
```
- **`History`** (`history.rs`) — top-level container. Organizes games by time into `Batch`es, runs forward/backward message passing across batches, and exposes `learning_curves()` and `log_evidence()`.
- **`Batch`** (`batch.rs`) — all games at a single time step. Runs `iteration()` to update skill estimates via `Game::posteriors()`, collecting `Skill` distributions per player.
- **`Game`** (`game.rs`) — a single match. Given teams (slices of `Gaussian`), computes posterior skill distributions using Gaussian factor graphs and `message.rs` helpers.
- **`Agent`** (`agent.rs`) — wraps a `Player` with temporal state (`last_time`, `message`). `receive()` applies time-decay (`gamma`) when the player reappears after a gap.
- **`Player`** (`player.rs`) — static configuration: prior `Gaussian`, `beta` (performance noise), `gamma` (skill drift per time unit).
- **`Gaussian`** (`gaussian.rs`) — core probability type. Stored as natural parameters (`pi = 1/sigma²`, `tau = mu/sigma²`). Arithmetic ops implement message multiplication/division in the factor graph.
- **`message.rs`** — `TeamMessage` and `DiffMessage`: intermediate factor graph messages used inside `Game`.
- **`MarginFactor`** (`factor/margin.rs`) — Gaussian observation factor on a diff variable; engaged by `Outcome::Scored`.
- **`lib.rs`** — exports the public API (`Game`, `Gaussian`, `History`, `Player`) and standalone functions (`quality()`, `pdf()`, `cdf()`, `erfc()`). Also defines global defaults: `MU=0.0`, `SIGMA=6.0`, `BETA=1.0`, `GAMMA=0.03`, `P_DRAW=0.0`, `EPSILON=1e-6`, `ITERATIONS=30`.
`History::add_events` flattens that into indices; teams survive only as
grouping, not as a value. Inference then runs on the internal shapes:
### Key design points
```
History → TimeSlice[] → Event[] → Item[]
Game (factor graph) → Schedule → BuiltinFactor[]
```
- `History` uses `IndexMap<K>` (defined in `lib.rs`) to map arbitrary player keys to `Agent` state.
- Convergence is measured by the maximum `delta()` across all skill distributions; iteration stops when below `EPSILON` or after `ITERATIONS` rounds.
- The `approx` feature gates `AbsDiffEq` on `Gaussian` for use in tests — the feature is optional and only needed for approximate equality assertions.
- `time` in `History`/`Batch` is currently an `f64`; the README notes it needs to become an enum to support richer temporal states.
- **`History`** (`history.rs`) top level. Interns keys, groups events into
`TimeSlice`s by time, runs the forward/backward sweep in `converge()`, and
answers `learning_curves()`, `current_skill()`, `log_evidence()`,
`predict_quality()`, `predict_outcome()`. Built via `HistoryBuilder`.
- **`TimeSlice`** (`time_slice.rs`) — all events at one time. Owns a
`SkillStore` and a `ScratchArena`; `iteration()` sweeps its events, using
`ColorGroups` to partition independent ones.
- **`Event`** — two distinct types, do not confuse them. The *public* ingestion
`Event<T, K>` is in `event.rs` (with `Team`/`Member`); the *internal*
`pub(crate) Event` in `time_slice.rs` is one match during inference, where
`compute()` runs inference reading skills immutably and `apply()` folds the
result back. That split is what lets a color group run in parallel with no
`unsafe`.
- **`Game`** (`game.rs`) — a single match's factor graph. `run_chain` builds the
diff chain between rank-adjacent teams and drives it to convergence.
- **`Gaussian`** (`gaussian.rs`) — natural parameters (`pi = 1/sigma²`,
`tau = mu/sigma²`). `Mul`/`Div` are the EP product/cavity: pure adds and
subtracts. Variance-space ops (`Add`, `Sub`, `exclude`, `forget`) go through
`from_mv`/`variance()` and take no square root.
- **`factor/`** — `TruncFactor` (ranked) and `MarginFactor` (scored) over a
flat `VarStore`. `Game::run_chain` drives them directly through a local
`DiffFactor` enum; there is no `Schedule` indirection and no generic `Factor`
trait. Both were removed once measurement showed nothing had ever used them
— see #42.
- **`Competitor`** (`competitor.rs`) — per-history temporal state (`message`,
`last_time`). **`Rating`** (`rating.rs`) — static config (prior, `beta`, drift).
- **`storage/`** — `SkillStore` (per slice, `pub(crate)`) and `CompetitorStore`
(per history, public), both indexed by `Index`. The module is `pub`, but only
`CompetitorStore` is reachable from outside the crate.
- **`KeyTable`** (`key_table.rs`) — user key ↔ `Index`, both directions O(1).
- **`Drift`** (`drift.rs`) / **`Time`** (`time.rs`) — traits. `Time` is a *trait*
(`i64`, `Untimed`), not an enum.
- **`lib.rs`** — public exports, global defaults (`MU`, `SIGMA`, `BETA`,
`GAMMA`, `P_DRAW`, `EPSILON`, `ITERATIONS`), and the standalone `quality()`.
The `cdf()` / `erfc()` helpers live here too but are `pub(crate)` and private
respectively — not public API.
### Invariants worth knowing
- **A tie needs `p_draw > 0`.** With `p_draw == 0.0` the truncation margin is
zero and the two-sided tie update evaluates `0/0`. Ingestion rejects such
events with `InferenceError::TieWithoutDrawProbability`. This includes
`Outcome::winner(w, n)` for `n >= 3`, which ties every loser.
- **NaN is never convergence.** Comparisons against NaN are all false, so
`tuple_gt` reads NaN as "below epsilon". Use `step_converged` /
`step_is_finite`, never `!tuple_gt(..)` alone.
- **Evidence accumulates in log space.** A linear product over a long diff
chain underflows to zero, and `ln(0)` is `-inf`.
- **Colors are contiguous.** `recompute_color_groups` reorders events so each
color occupies one range; `ColorGroups::groups_are_contiguous` asserts it.
- **Transcendentals go through `libm`, not `std`.** IEEE 754 pins the basic
operations and `sqrt` but says nothing about `exp`/`log`/`erf`, and `std`
delegates to the *system* math library — measured, `f64::exp` and `libm::exp`
disagree on 9.7% of inputs by one ULP. Since inference is an iterative fixed
point, one ULP can change an iteration count. Use `libm::exp` / `libm::log` in
inference code; `f64::sqrt` is fine (IEEE specifies it). Tests may use either.
- **The crate is `#![forbid(unsafe_code)]`.** Keep it that way.
- **Ingestion order must not change the answer.** Events added one at a time
must converge to the same fixed point as the same events batched — see
`tests/ingestion_equivalence.rs`.
### Testing notes
- Numerical goldens are cross-validated against the Python/Julia reference.
Some are *convergence residuals*, not exact values; treat a small movement
as suspicious but check whether the new value is closer to the analytic
truth (symmetric fixtures converge to their prior mean exactly) before
assuming a regression.
- `tests/degenerate_inputs.rs` covers empty/boundary/error paths,
`tests/ingestion_equivalence.rs` covers batching order, `tests/quality.rs`
covers N-group quality, `tests/determinism.rs` covers thread counts.
+34 -1
View File
@@ -1,7 +1,30 @@
[package]
name = "trueskill-tt"
version = "0.1.2"
version = "0.5.0"
edition = "2024"
rust-version = "1.85"
description = "TrueSkill Through Time: Bayesian skill rating that tracks how skill evolves over time, via Gaussian message passing"
repository = "https://git.aceofba.se/logaritmisk/trueskill-tt"
authors = ["Anders Olsson"]
# Publishing is restricted to the private kellnr registry; this also makes
# an accidental `cargo publish` to crates.io a hard error rather than a
# irreversible mistake. Index is declared in `.cargo/config.toml`.
publish = ["kellnr"]
readme = "README.md"
keywords = ["trueskill", "rating", "bayesian", "elo", "skill"]
categories = ["algorithms", "science", "game-development"]
license = "MIT OR Apache-2.0"
# `examples/atp.csv` is a 48 MB tennis dataset — 99% of the packaged crate,
# for a library whose source is 312 KB. `examples/atp.rs` opens it by
# relative path at runtime, so excluding the data still compiles; the
# example just needs the file fetched from the repo to run.
exclude = [
"/docs",
"/benches/*.txt",
"/temp",
"/.gitea",
"/examples/atp.csv",
]
[lib]
bench = false
@@ -22,8 +45,13 @@ harness = false
name = "scored"
harness = false
[[bench]]
name = "ingest"
harness = false
[dependencies]
approx = { version = "0.5.1", optional = true }
libm = "0.2.16"
rayon = { version = "1", optional = true }
smallvec = "1"
@@ -35,9 +63,14 @@ rayon = ["dep:rayon"]
criterion = "0.5"
plotters = { version = "0.3", default-features = false, features = ["svg_backend", "all_elements", "all_series"] }
plotters-backend = "0.3"
proptest = "1.11.0"
time = { version = "0.3", features = ["parsing"] }
trueskill-tt = { path = ".", features = ["approx"] }
# Debug symbols in release are for `just flame` (cargo-flamegraph), which needs
# them to symbolicate. Profile settings in a library are ignored by downstream
# consumers, so these only affect local builds — this is deliberate, not an
# oversight.
[profile.release]
debug = true
+81
View File
@@ -1,4 +1,39 @@
alias b := bench
alias t := test
# Run the full test suite across the feature combinations CI checks.
test:
cargo test
cargo test --features approx
cargo test --features approx,rayon
cargo test --release --features approx
# Fast inner-loop tests.
check:
cargo test --features approx
# Posteriors must be bit-identical across rayon worker counts.
determinism:
#!/usr/bin/env bash
set -euo pipefail
for threads in 1 2 4 8; do
echo "== RAYON_NUM_THREADS=$threads =="
RAYON_NUM_THREADS=$threads cargo test --release \
--features approx,rayon --test determinism
done
lint:
cargo clippy --all-targets --all-features -- -D warnings
# Always nightly: rustfmt.toml uses nightly-only options.
fmt:
cargo +nightly fmt
fmt-check:
cargo +nightly fmt --check
# Everything CI runs.
ci: fmt-check lint test determinism
store:
cargo bench -- --save-baseline base
@@ -8,3 +43,49 @@ bench:
flame:
cargo flamegraph --root --example atp
# ---------------------------------------------------------------------------
# Release workflow
#
# Publishing goes to the private kellnr registry only: `Cargo.toml` sets
# `publish = ["kellnr"]`, so an accidental `cargo publish` to crates.io is a
# hard error rather than an irreversible mistake. The index is declared in the
# committed `.cargo/config.toml`; the token is per-user and lives in
# `~/.cargo/credentials.toml` (`cargo login --registry kellnr`).
#
# Step 1: just release-plan [level] — dry run, no writes
# Step 2: just release [level] — bump, changelog, tag, publish, push
#
# LEVEL is the cargo-release bump level (default `minor`). On 0.x:
# minor -> breaking bump (0.1.2 -> 0.2.0) <- any public-API change
# patch -> additive only (0.1.2 -> 0.1.3)
# major -> reserved for the 1.0.0 jump
#
# `release.toml` regenerates CHANGELOG.md with git-cliff in a pre-release hook
# and keeps push = false; this recipe pushes last, after publish has succeeded.
# ---------------------------------------------------------------------------
# Dry-run preview of the next release. Inspect the version bump and the
# "Publishing ..." line before running `just release`.
release-plan level="minor":
cargo release {{level}}
# Cut a release from a clean main: gate -> bump -> tag -> publish -> push.
release level="minor":
#!/usr/bin/env bash
set -euo pipefail
if [[ "$(git branch --show-current)" != "main" ]]; then
echo "error: run 'just release' from the 'main' branch" >&2; exit 1
fi
if [[ -n "$(git status --porcelain)" ]]; then
echo "error: working tree is dirty — commit or stash first" >&2; exit 1
fi
# cargo-release only verify-compiles the packaged crate; it does not run the
# suite, and publishing is irreversible. Run the same gate CI does, which
# includes the release profile where debug_assert! is compiled out.
just ci
cargo release {{level}} --execute --no-confirm
git push --follow-tags
+201
View File
@@ -0,0 +1,201 @@
Apache License
Version 2.0, January 2004
http://www.apache.org/licenses/
TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
1. Definitions.
"License" shall mean the terms and conditions for use, reproduction,
and distribution as defined by Sections 1 through 9 of this document.
"Licensor" shall mean the copyright owner or entity authorized by
the copyright owner that is granting the License.
"Legal Entity" shall mean the union of the acting entity and all
other entities that control, are controlled by, or are under common
control with that entity. For the purposes of this definition,
"control" means (i) the power, direct or indirect, to cause the
direction or management of such entity, whether by contract or
otherwise, or (ii) ownership of fifty percent (50%) or more of the
outstanding shares, or (iii) beneficial ownership of such entity.
"You" (or "Your") shall mean an individual or Legal Entity
exercising permissions granted by this License.
"Source" form shall mean the preferred form for making modifications,
including but not limited to software source code, documentation
source, and configuration files.
"Object" form shall mean any form resulting from mechanical
transformation or translation of a Source form, including but
not limited to compiled object code, generated documentation,
and conversions to other media types.
"Work" shall mean the work of authorship, whether in Source or
Object form, made available under the License, as indicated by a
copyright notice that is included in or attached to the work
(an example is provided in the Appendix below).
"Derivative Works" shall mean any work, whether in Source or Object
form, that is based on (or derived from) the Work and for which the
editorial revisions, annotations, elaborations, or other modifications
represent, as a whole, an original work of authorship. For the purposes
of this License, Derivative Works shall not include works that remain
separable from, or merely link (or bind by name) to the interfaces of,
the Work and Derivative Works thereof.
"Contribution" shall mean any work of authorship, including
the original version of the Work and any modifications or additions
to that Work or Derivative Works thereof, that is intentionally
submitted to Licensor for inclusion in the Work by the copyright owner
or by an individual or Legal Entity authorized to submit on behalf of
the copyright owner. For the purposes of this definition, "submitted"
means any form of electronic, verbal, or written communication sent
to the Licensor or its representatives, including but not limited to
communication on electronic mailing lists, source code control systems,
and issue tracking systems that are managed by, or on behalf of, the
Licensor for the purpose of discussing and improving the Work, but
excluding communication that is conspicuously marked or otherwise
designated in writing by the copyright owner as "Not a Contribution."
"Contributor" shall mean Licensor and any individual or Legal Entity
on behalf of whom a Contribution has been received by Licensor and
subsequently incorporated within the Work.
2. Grant of Copyright License. Subject to the terms and conditions of
this License, each Contributor hereby grants to You a perpetual,
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
copyright license to reproduce, prepare Derivative Works of,
publicly display, publicly perform, sublicense, and distribute the
Work and such Derivative Works in Source or Object form.
3. Grant of Patent License. Subject to the terms and conditions of
this License, each Contributor hereby grants to You a perpetual,
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
(except as stated in this section) patent license to make, have made,
use, offer to sell, sell, import, and otherwise transfer the Work,
where such license applies only to those patent claims licensable
by such Contributor that are necessarily infringed by their
Contribution(s) alone or by combination of their Contribution(s)
with the Work to which such Contribution(s) was submitted. If You
institute patent litigation against any entity (including a
cross-claim or counterclaim in a lawsuit) alleging that the Work
or a Contribution incorporated within the Work constitutes direct
or contributory patent infringement, then any patent licenses
granted to You under this License for that Work shall terminate
as of the date such litigation is filed.
4. Redistribution. You may reproduce and distribute copies of the
Work or Derivative Works thereof in any medium, with or without
modifications, and in Source or Object form, provided that You
meet the following conditions:
(a) You must give any other recipients of the Work or
Derivative Works a copy of this License; and
(b) You must cause any modified files to carry prominent notices
stating that You changed the files; and
(c) You must retain, in the Source form of any Derivative Works
that You distribute, all copyright, patent, trademark, and
attribution notices from the Source form of the Work,
excluding those notices that do not pertain to any part of
the Derivative Works; and
(d) If the Work includes a "NOTICE" text file as part of its
distribution, then any Derivative Works that You distribute must
include a readable copy of the attribution notices contained
within such NOTICE file, excluding those notices that do not
pertain to any part of the Derivative Works, in at least one
of the following places: within a NOTICE text file distributed
as part of the Derivative Works; within the Source form or
documentation, if provided along with the Derivative Works; or,
within a display generated by the Derivative Works, if and
wherever such third-party notices normally appear. The contents
of the NOTICE file are for informational purposes only and
do not modify the License. You may add Your own attribution
notices within Derivative Works that You distribute, alongside
or as an addendum to the NOTICE text from the Work, provided
that such additional attribution notices cannot be construed
as modifying the License.
You may add Your own copyright statement to Your modifications and
may provide additional or different license terms and conditions
for use, reproduction, or distribution of Your modifications, or
for any such Derivative Works as a whole, provided Your use,
reproduction, and distribution of the Work otherwise complies with
the conditions stated in this License.
5. Submission of Contributions. Unless You explicitly state otherwise,
any Contribution intentionally submitted for inclusion in the Work
by You to the Licensor shall be under the terms and conditions of
this License, without any additional terms or conditions.
Notwithstanding the above, nothing herein shall supersede or modify
the terms of any separate license agreement you may have executed
with Licensor regarding such Contributions.
6. Trademarks. This License does not grant permission to use the trade
names, trademarks, service marks, or product names of the Licensor,
except as required for reasonable and customary use in describing the
origin of the Work and reproducing the content of the NOTICE file.
7. Disclaimer of Warranty. Unless required by applicable law or
agreed to in writing, Licensor provides the Work (and each
Contributor provides its Contributions) on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
implied, including, without limitation, any warranties or conditions
of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
PARTICULAR PURPOSE. You are solely responsible for determining the
appropriateness of using or redistributing the Work and assume any
risks associated with Your exercise of permissions under this License.
8. Limitation of Liability. In no event and under no legal theory,
whether in tort (including negligence), contract, or otherwise,
unless required by applicable law (such as deliberate and grossly
negligent acts) or agreed to in writing, shall any Contributor be
liable to You for damages, including any direct, indirect, special,
incidental, or consequential damages of any character arising as a
result of this License or out of the use or inability to use the
Work (including but not limited to damages for loss of goodwill,
work stoppage, computer failure or malfunction, or any and all
other commercial damages or losses), even if such Contributor
has been advised of the possibility of such damages.
9. Accepting Warranty or Additional Liability. While redistributing
the Work or Derivative Works thereof, You may choose to offer,
and charge a fee for, acceptance of support, warranty, indemnity,
or other liability obligations and/or rights consistent with this
License. However, in accepting such obligations, You may act only
on Your own behalf and on Your sole responsibility, not on behalf
of any other Contributor, and only if You agree to indemnify,
defend, and hold each Contributor harmless for any liability
incurred by, or claims asserted against, such Contributor by reason
of your accepting any such warranty or additional liability.
END OF TERMS AND CONDITIONS
APPENDIX: How to apply the Apache License to your work.
To apply the Apache License to your work, attach the following
boilerplate notice, with the fields enclosed by brackets "[]"
replaced with your own identifying information. (Don't include
the brackets!) The text should be enclosed in the appropriate
comment syntax for the file format. We also recommend that a
file or class name and description of purpose be included on the
same "printed page" as the copyright notice for easier
identification within third-party archives.
Copyright [yyyy] [name of copyright owner]
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.
+19
View File
@@ -0,0 +1,19 @@
Copyright (c) 2026 Anders Olsson
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in
all copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN
THE SOFTWARE.
+226 -28
View File
@@ -13,64 +13,136 @@ Rust port of [TrueSkillThroughTime.py](https://github.com/glandfried/TrueSkillTh
## Drift
Skill drift models how a player's true skill can change between appearances. Each time a player reappears after a gap, their skill uncertainty is widened by the drift model before the new evidence is incorporated.
Skill drift models how a competitor's true skill can change between appearances.
Each time they reappear after a gap, their skill uncertainty is widened by the
drift model before the new evidence is incorporated.
Drift is represented by the `Drift` trait:
Drift is represented by the `Drift` trait (`src/drift.rs`), generic over the
history's time type:
```rust
pub trait Drift: Copy + Debug {
fn variance_delta(&self, elapsed: i64) -> f64;
```text
pub trait Drift<T: Time>: Copy + Debug + Send + Sync {
fn variance_delta(&self, from: &T, to: &T) -> f64;
fn variance_for_elapsed(&self, elapsed: i64) -> f64;
}
```
`variance_delta` returns the amount to add to `σ²` given the elapsed time since the player last played. Internally, `Gaussian::forget` uses this to compute the new sigma: `σ_new = sqrt(σ² + variance_delta)`.
Both methods return the amount to add to `σ²`, not to `σ`. `variance_delta`
works from two timestamps; `variance_for_elapsed` takes an already-computed
elapsed count, and is used on the paths that cache it. `Gaussian::forget`
applies the result entirely in variance space — `from_mv(mu, variance() +
variance_delta)` — taking no square root.
That block is a quotation rather than a doctest. The custom-drift example below
is compiled by CI, so it is what actually pins the signature.
### ConstantDrift
The built-in `ConstantDrift` implements a linear random walk — skill uncertainty grows proportionally to time:
The built-in `ConstantDrift` implements a linear random walk — skill uncertainty
grows proportionally to time:
```
```text
variance_delta = elapsed * γ²
```
This is the standard TrueSkill Through Time model. Use it by passing a `ConstantDrift(gamma)` when constructing a `Player`:
This is the standard TrueSkill Through Time model. Pass a `ConstantDrift(gamma)`
when constructing a `Rating`:
```rust
use trueskill_tt::{Player, Gaussian, drift::ConstantDrift};
use trueskill_tt::{ConstantDrift, Gaussian, Rating};
// gamma = 0.1 means skill can shift ~0.1 per time unit
let player = Player::new(Gaussian::from_ms(0.0, 6.0), 1.0, ConstantDrift(0.1));
// gamma = 0.1 means skill can shift ~0.1 per time unit.
let rating: Rating<i64, ConstantDrift> =
Rating::new(Gaussian::from_ms(0.0, 6.0), 1.0, ConstantDrift(0.1));
assert_eq!(rating.drift().0, 0.1);
```
The type annotation is load-bearing: `ConstantDrift` implements `Drift<T>` for
every `T: Time`, so without it `T` is ambiguous.
### Custom drift
Implement `Drift` to express any other model. For example, a drift that saturates after a long absence (uncertainty grows with the square root of elapsed time instead of linearly):
Implement `Drift<T>` to express any other model. For example, a drift that
saturates after a long absence, with uncertainty growing as the square root of
elapsed time instead of linearly:
```rust
use trueskill_tt::drift::Drift;
use trueskill_tt::{Drift, Gaussian, History, Rating, Time};
#[derive(Clone, Copy, Debug)]
struct SqrtDrift {
gamma: f64,
}
impl Drift for SqrtDrift {
fn variance_delta(&self, elapsed: i64) -> f64 {
(elapsed as f64).sqrt() * self.gamma * self.gamma
impl<T: Time> Drift<T> for SqrtDrift {
fn variance_delta(&self, from: &T, to: &T) -> f64 {
let elapsed = from.elapsed_to(to).max(0) as f64;
elapsed.sqrt() * self.gamma * self.gamma
}
fn variance_for_elapsed(&self, elapsed: i64) -> f64 {
(elapsed.max(0) as f64).sqrt() * self.gamma * self.gamma
}
}
let player = Player::new(Gaussian::from_ms(0.0, 6.0), 1.0, SqrtDrift { gamma: 0.5 });
// On a single Rating:
let rating: Rating<i64, SqrtDrift> =
Rating::new(Gaussian::from_ms(0.0, 6.0), 1.0, SqrtDrift { gamma: 0.5 });
// Or for a whole History, via the builder:
let history = History::builder().drift(SqrtDrift { gamma: 0.5 }).build();
assert_eq!(rating.beta(), 1.0);
assert_eq!(history.log_evidence(), 0.0);
```
To use a custom drift type with `History`, use the `.drift()` builder method instead of `.gamma()`:
`HistoryBuilder::drift` is the only way to set a history's drift model; there is
no `gamma()` shorthand. The default is `ConstantDrift(GAMMA)`.
### Per-competitor drift
A `History` has one drift model, but individual competitors can scale it.
`Member::with_drift_scale(s)` multiplies the drift *variance* that competitor
accumulates, so `s` is in the same units as `gamma`: `ConstantDrift(g)` at
scale `s` behaves exactly as `ConstantDrift(g * s)` would, for that competitor
alone.
`0.0` pins a competitor still. That is what makes a **fixed reference point**
expressible in the same graph as moving competitors — a bot at a known
strength, a rating floor, a course difficulty:
```rust
let h = History::builder()
.drift(SqrtDrift { gamma: 0.5 })
.build();
use trueskill_tt::{ConstantDrift, Event, History, Member, Outcome, Team};
let mut h = History::builder().drift(ConstantDrift(0.1)).build();
h.add_events(vec![Event {
time: 0,
teams: [
Team::with_members([Member::new("player")]),
// A course does not improve. Pin it, and the round's evidence
// lands on the player instead of being split between the two.
Team::with_members([Member::new("layout_7").with_drift_scale(0.0)]),
]
.into_iter()
.collect(),
outcome: Outcome::winner(0, 2),
}])
.unwrap();
h.converge().unwrap();
```
Like `with_prior`, the scale is **competitor configuration captured at first
appearance** — setting it on a key the history already knows has no effect. It
must be finite and non-negative; ingestion otherwise fails with
`InferenceError::InvalidParameter`.
Note that the fluent `EventBuilder` (`h.event(t).team([...])`) sets weights but
not `drift_scale` or `prior`; those need the typed `Event` / `Team` / `Member`
shape shown above.
## Scored outcomes
Use `Outcome::scores([...])` when you have continuous per-team scores rather
@@ -80,7 +152,7 @@ soft Gaussian evidence about the latent performance diff. Configure
(smaller σ = more trust).
```rust
use trueskill_tt::{History, Outcome};
use trueskill_tt::History;
let mut h = History::builder().score_sigma(2.0).build();
h.event(1)
@@ -92,12 +164,138 @@ h.event(1)
h.converge().unwrap();
```
## Prediction
`predict_outcome` gives the full distribution over finishing orders. Each entry
is a rank vector in the same shape `Outcome::ranking` takes — equal ranks mean a
tie — so an outcome feeds straight back into inference.
```rust
use trueskill_tt::History;
let mut h = History::builder().p_draw(0.1).build();
h.record_winner(&"alice", &"bob", 1).unwrap();
h.converge().unwrap();
let p = h.predict_outcome(&[&[&"alice"], &[&"bob"]]).unwrap();
// Probabilities are exhaustive and disjoint, so they sum to one.
assert!((p.total() - 1.0).abs() < 1e-6);
let (best, likelihood) = p.most_likely().unwrap();
println!("most likely: {best:?} at {likelihood:.3}");
println!("draw: {:.3}", p.probability_of(&[0, 0]));
```
Supports any number of teams. Because the outcome space grows factorially, the
full distribution is capped at `MAX_PREDICTED_TEAMS`; two cheaper entry points
stay available at any size:
- `predict_win_probabilities(teams)``P(team i finishes strictly first)`,
quadratic in team count.
- `predict_ranking(teams, ranks)` — one specific finishing order.
Unknown keys are an error by default, not a silent omission: a team the history
has never seen cannot produce a confident-looking probability. The error names
the key, and every key must already be known — pre-filter with `lookup` or
`current_skill` if your caller cannot guarantee that.
If predicting for competitors you have never seen is the point rather than a
mistake, say so once:
```rust
use trueskill_tt::{History, UnknownKeys};
let h = History::builder().unknown_keys(UnknownKeys::Prior).build();
```
An unknown competitor is then answered from the configured prior, which is the
honest reading — you have no evidence about them — and correctly *widens* a team
that contains one. There is deliberately no "skip the member" mode: a team's
performance is the sum of its members, so dropping one would make the model more
certain because it knows less.
### Asking about one competitor
`Gaussian` answers tail questions directly, which is what a stopping rule needs:
```rust
use trueskill_tt::History;
let mut h = History::builder().build();
h.record_winner(&"alice", &"bob", 1).unwrap();
let _ = h.converge().unwrap();
let skill = h.current_skill(&"alice").unwrap();
// "How sure am I that this is below the cutoff?" — a probability, not a
// `mu + z * sigma` band whose confidence drifts as sigma changes.
let _ = skill.probability_below(20.0);
// Use this rather than `1.0 - probability_below(x)`: the complement cancels
// away every digit in the upper tail, which is where a stopping rule lives.
let _ = skill.probability_above(30.0);
```
## Which match to play next
`quality()` measures whether a matchup is *fair*. That is not the same as
whether it is *informative*, and the two only coincide for two evenly matched
competitors. When each observation costs something, ask
`expected_information_gain` instead — the outcome-weighted divergence between
what you believe now and what you would believe afterwards.
```rust
use trueskill_tt::History;
let mut h = History::builder().build();
for t in 1..=10 {
h.record_winner(&"veteran", &"regular", t).unwrap();
h.record_winner(&"regular", &"veteran", t + 100).unwrap();
}
h.record_winner(&"veteran", &"newcomer", 500).unwrap();
h.converge().unwrap();
let settled = h.expected_information_gain(&[&[&"veteran"], &[&"regular"]]).unwrap();
let unknown = h.expected_information_gain(&[&[&"veteran"], &[&"newcomer"]]).unwrap();
// Playing the newcomer teaches you more than replaying a settled rivalry.
assert!(unknown > settled);
```
The result is in nats, and is bounded by the entropy of the outcome: at most
`ln 2 ≈ 0.693` for a two-way result, `ln 3` once draws are possible, `ln k` for
`k` outcomes. A value near zero means you already know how it ends.
This costs one full inference pass **per possible outcome**, so it is far more
expensive than `quality()`. Scoring every pairing among `n` competitors is
`O(n² × outcomes)` passes — shortlist with `quality()` or
`predict_win_probabilities` first, then score only the shortlist.
## Todo
- [x] Implement approx for Gaussian
- [x] Add more tests from `TrueSkillThroughTime.jl`
- [ ] Add tests for `quality()` (Use [sublee/trueskill](https://github.com/sublee/trueskill/tree/master) as reference)
- [ ] Benchmark Batch::iteration()
- [ ] Time needs to be an enum so we can have multiple states (see `batch::compute_elapsed()`)
- [ ] Add examples (use same TrueSkillThroughTime.(py|jl))
- [ ] Add Observer (see [argmin](https://docs.rs/argmin/latest/argmin/core/trait.Observe.html) for inspiration)
- [x] Generalise a time axis — `Time` is now a trait (`Untimed`, `i64`), not an enum
- [x] Add examples (`examples/atp.rs`, `examples/scored.rs`)
- [x] Add Observer (`Observer` / `NullObserver`)
- [x] Benchmark the inference loop (`benches/batch.rs`, `benches/history_converge.rs`, `benches/ingest.rs`)
- [x] N-team `predict_outcome` with draw mass, and `expected_information_gain`
- [x] Cross-check `quality()` against [sublee/trueskill](https://github.com/sublee/trueskill/tree/master) — N identical teams follow the closed form `(1/5)^((n-1)/2)` for the conventional parameters, asserted for n = 2..10, and the n=3/n=5 values (0.200, 0.040) match the reference package
## License
Licensed under either of
- Apache License, Version 2.0 ([LICENSE-APACHE](LICENSE-APACHE) or
<http://www.apache.org/licenses/LICENSE-2.0>)
- MIT license ([LICENSE-MIT](LICENSE-MIT) or
<http://opensource.org/licenses/MIT>)
at your option.
### Contribution
Unless you explicitly state otherwise, any contribution intentionally submitted
for inclusion in the work by you, as defined in the Apache-2.0 license, shall be
dual licensed as above, without any additional terms or conditions.
+1 -1
View File
@@ -36,7 +36,7 @@ fn criterion_benchmark(criterion: &mut Criterion) {
let kinds = vec![EventKind::Ranked; composition.len()];
let mut time_slice = TimeSlice::new(1, P_DRAW, ConvergenceOptions::default());
time_slice.add_events(composition, results, weights, kinds, &agents);
time_slice.add_events(composition, Some(results), Some(weights), kinds, &agents);
criterion.bench_function("Batch::iteration", |b| {
b.iter(|| time_slice.iteration(0, &agents))
+3 -3
View File
@@ -82,7 +82,7 @@ fn bench_converge(c: &mut Criterion) {
b.iter_batched(
|| build_history_1v1(500, 100, 10, 42),
|mut h| {
h.converge().unwrap();
let _ = h.converge().unwrap();
},
BatchSize::SmallInput,
);
@@ -92,7 +92,7 @@ fn bench_converge(c: &mut Criterion) {
b.iter_batched(
|| build_history_1v1(2000, 200, 20, 42),
|mut h| {
h.converge().unwrap();
let _ = h.converge().unwrap();
},
BatchSize::SmallInput,
);
@@ -106,7 +106,7 @@ fn bench_converge(c: &mut Criterion) {
b.iter_batched(
|| build_history_1v1(5000, 50000, 5000, 42),
|mut h| {
h.converge().unwrap();
let _ = h.converge().unwrap();
},
BatchSize::SmallInput,
);
+62
View File
@@ -0,0 +1,62 @@
//! Ingestion cost: one event per call versus one batched call.
//!
//! The rest of the suite only measures batched construction, which is why a
//! quadratic in the incremental path went unnoticed — `record_winner` and
//! `event(..).commit()` each ingest a single event, so a caller looping over a
//! match feed takes that path.
use std::hint::black_box;
use criterion::{BenchmarkId, Criterion, criterion_group, criterion_main};
use smallvec::smallvec;
use trueskill_tt::{Event, History, Member, Outcome, Team};
fn events(n: usize, time: i64) -> Vec<Event<i64, String>> {
(0..n)
.map(|i| Event {
time,
teams: smallvec![
Team::with_members([Member::new(format!("p{}", 2 * i))]),
Team::with_members([Member::new(format!("p{}", 2 * i + 1))]),
],
outcome: Outcome::winner(0, 2),
})
.collect()
}
fn bench_ingest(c: &mut Criterion) {
let mut group = c.benchmark_group("ingest");
for n in [250usize, 500, 1000] {
group.bench_with_input(BenchmarkId::new("one-at-a-time", n), &n, |b, &n| {
b.iter_batched(
|| events(n, 0),
|evs| {
let mut h: History<i64, _, _, String> = History::builder_with_key().build();
for ev in evs {
h.add_events(std::iter::once(ev)).unwrap();
}
black_box(h.time_slices_len())
},
criterion::BatchSize::SmallInput,
);
});
group.bench_with_input(BenchmarkId::new("single-batch", n), &n, |b, &n| {
b.iter_batched(
|| events(n, 0),
|evs| {
let mut h: History<i64, _, _, String> = History::builder_with_key().build();
h.add_events(evs).unwrap();
black_box(h.time_slices_len())
},
criterion::BatchSize::SmallInput,
);
});
}
group.finish();
}
criterion_group!(benches, bench_ingest);
criterion_main!(benches);
+1 -1
View File
@@ -29,7 +29,7 @@ fn bench_scored_history(c: &mut Criterion) {
});
}
h.add_events(events).unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
});
});
}
+5
View File
@@ -44,6 +44,11 @@ split_commits = false
# Assigns commits to groups.
# Optionally sets the commit's scope and can decide to exclude commits from further processing.
commit_parsers = [
# Must precede the type parsers below: a `feat!`/`fix!`/`refactor!` subject
# matches those too, and the first match wins. Without this a breaking
# change renders as an ordinary line of its own type.
{ message = "^[a-z]+(\\(.+\\))?!:", group = "Breaking Changes" },
{ body = "BREAKING CHANGE", group = "Breaking Changes" },
{ message = "^feat", group = "Features" },
{ message = "^fix", group = "Bug Fixes" },
{ message = "^doc", group = "Documentation" },
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,342 @@
# Filtered (Forward-Only) Estimates
Closes [#19](https://git.aceofba.se/logaritmisk/trueskill-tt/issues/19).
## Summary
`HistoryBuilder::online(true)` is inert. It flips a flag that reaches
`Item::within_prior` (`src/time_slice.rs:70-71`), which reads
`Skill.online` (`src/time_slice.rs:25`) — a field initialised to `N_INF`
(`src/time_slice.rs:41`) and never assigned anywhere. The online path
therefore builds every rating from the improper Gaussian, and
`log_evidence()` silently reports `n × ln(0.5)`: every game scored as a
coin flip, finite and plausible-looking.
This spec replaces the field and the flag with a **read-only forward-only
pass** over the converged history, exposed as three new public methods.
The pass reuses the production within-slice sweep verbatim rather than
reimplementing inference, and stores nothing on `Skill`.
## Background
### Why a stored field cannot hold this quantity
The issue proposes populating `skill.online` during the forward pass,
alongside `new_forward_info` (`src/time_slice.rs:576`). That would not
work, and understanding why determines the whole design.
`new_forward_info` sets `skill.forward` from
`agents[a].receive_for_elapsed(...)`, whose `message` was written by the
previous slice's `forward_prior_out` (`src/time_slice.rs:549`):
```rust
skill.forward * skill.likelihood
```
`History::iteration` (`src/history.rs:255`) alternates a backward sweep
over slices and a forward sweep. From the second iteration onward, the
`skill.likelihood` feeding that message has already absorbed backward
information from the preceding backward sweep. So after `converge()`,
**`skill.forward` is a smoothed quantity, not a filtering one** — and any
field written from it inherits the same contamination on every sweep
after the first.
### The neighbouring trap
The same reasoning applies to the existing `forward: bool` parameter on
`log_evidence_internal` (`src/history.rs:395`). It is a genuine filtering
quantity only on a history that has never been converged. That is why the
test at `src/history.rs:1183` can assert
```rust
assert_ulps_eq!(trueskill_log_evidence, trueskill_log_evidence_online, epsilon = 1e-6);
```
— the fixture is never converged, so the forward message still equals the
cavity prior. (Note also that the local binding is named `..._online`
while the flag it passes is `forward`; the two senses were already
muddled.)
Fixing `forward: bool` is **out of scope** here; see *Out-of-scope
follow-ups*.
### Why this is worth implementing rather than deleting
The forward-only estimate has a second consumer beyond prequential model
comparison. `learning_curve()` returns post-convergence posteriors, so
every point is smoothed — the estimate at a given date incorporates
rounds played years later. On [ustat](https://git.aceofba.se/logaritmisk/ustat)'s
real data (prior μ=0, σ=6) that produces curves which start already
spread apart and barely move:
```
player first point final point
Eskil mu +3.72 sigma 1.17 mu +4.61 sigma 1.21
Anders Olsson mu +1.61 sigma 0.90 mu +1.16 sigma 0.82
LUDVIGSSON mu -2.09 sigma 1.08 mu -2.61 sigma 1.13
Anners mu -2.85 sigma 1.27 mu -2.86 sigma 1.26
```
σ at the *first* plotted point is 0.901.60 against a prior of 6.00. A
caller cannot reconstruct the filtered view from the public API today
except by refitting over `events[0..k]` for every k — O(n²) fits for
something one forward pass already computes.
## Scope
### What ships
1. A read-only forward-only pass on `History`, walking slices in time
order and carrying its own forward messages.
2. Three public methods: `filtered_log_evidence`,
`filtered_learning_curves`, `filtered_learning_curve`.
3. Removal of `Skill.online`, `History.online`, `HistoryBuilder.online`,
`HistoryBuilder::online()`, and the `online: bool` parameter threaded
through `Item::within_prior`, `Event::within_priors`, and
`TimeSlice::log_evidence`.
4. `#[derive(Clone)]` on `Event`, `Team`, `Item`; `iterate_to_convergence`
loses its `#[cfg(test)]` gate.
5. A CHANGELOG entry recording the API break.
### What does not ship
- No change to `log_evidence()`, `log_evidence_for()`, `learning_curve()`,
`learning_curves()`, or `current_skill()`. Their values are unchanged
by this work.
- No fix to the `forward: bool` flag described above.
- No caching of pass results. Each call runs a full pass; the doc
comments say so.
- No `rayon` parallelism across slices — the pass is sequentially
dependent by construction.
- No prior-predictive accessor. The pass computes the pre-event forward
message internally, but only the filtered posterior is exposed until a
second caller needs otherwise.
## Design
### Naming
`filtered_*`, not `online_*`. "Filtered" is the standard term for the
forward-only estimate, and the crate already uses "online" for a second,
unrelated thing — incremental ingestion, which `benches/baseline.txt:128`
calls the "online-add" path. Two senses of one word in one crate is how
the present bug reads as plausible.
### The pass
```rust
pub(crate) struct FilteredStep {
log_evidence: f64,
posteriors: Vec<(Index, Gaussian)>,
}
fn filtered_pass(&self) -> Vec<(T, FilteredStep)>
```
`posteriors` doubles as the outgoing forward message: the scratch sweep never
writes `backward`, so it stays `N_INF`, and `Skill::posterior()` and
`forward_prior_out` are then the same product.
Walk `self.time_slices` in order, carrying
`messages: HashMap<Index, Gaussian>` — the forward message out of each
competitor's most recent appearance. For each slice:
1. **Build a scratch clone.** Same `time`, `p_draw`, `convergence`, and
cloned `events` with every `item.likelihood` reset to `N_INF`. Fresh
`SkillStore` in which, for each agent present in the real slice:
```rust
forward = match messages.get(&agent) {
Some(msg) => msg.forget(rating.drift.variance_for_elapsed(skill.elapsed)),
None => rating.prior,
}
backward = N_INF
likelihood = N_INF
elapsed = skill.elapsed // copied from the real slice
```
This mirrors `Competitor::receive_for_elapsed` (`src/competitor.rs:39`)
exactly, including its `message != N_INF` fallback to the prior.
`skill.elapsed` is reused rather than recomputed: it is maintained by
`add_events_with_prior` across out-of-order ingestion, and production
convergence already trusts it.
2. **Run the real sweep.** `scratch.iterate_to_convergence(agents)`
(`src/time_slice.rs:516`), unmodified. Fidelity comes from reusing the
production path rather than a parallel reimplementation — in
particular, a competitor appearing in two events at the same time is
handled by the same within-slice EP that `converge()` uses, not
approximated the way the current `online`/`forward` evidence paths are
(they run each event independently and sum).
3. **Harvest.** With `backward == N_INF` acting as the multiplicative
identity, `Skill::posterior()` is exactly forward × likelihood — the
filtered posterior. Slice evidence is
`scratch.events.iter().map(|e| e.log_evidence).sum()`; `apply`
(`src/time_slice.rs:162`) writes that field on every event during the
sweep.
4. **Carry forward.** `messages.insert(a, scratch.forward_prior_out(&a))`
for each agent in the slice.
Steps 14 are the forward half of `History::iteration`
(`src/history.rs:283-297`) with the backward half never run. The pass
touches no field of `self`.
### Public API
```rust
impl<T: Time, D: Drift<T>, O: Observer<T>, K: Eq + Hash + Clone> History<T, D, O, K> {
pub fn filtered_log_evidence(&self) -> f64;
pub fn filtered_learning_curves(&self) -> HashMap<K, Vec<(T, Gaussian)>>;
pub fn filtered_learning_curve<Q>(&self, key: &Q) -> Vec<(T, Gaussian)>
where
K: Borrow<Q>,
Q: Hash + Eq + ?Sized;
}
```
All take `&self` — the pass mutates nothing. Shapes deliberately mirror
`learning_curve` / `learning_curves` (`src/history.rs:325`, `:381`) so a
caller can plot smoothed and filtered curves on one chart with the same
handling code.
`filtered_learning_curve` runs the same full pass as the plural form and
collects one key; the cost is identical, only the collection differs.
Callers wanting several keys should use the plural form. Documented on
both methods.
Because the pass carries its own messages and re-runs inference, its
results **do not depend on whether `converge()` has been called**. That
is the property a stored field cannot have, and it is asserted as a test.
### Removal inventory
| Location | Change |
|---|---|
| `src/time_slice.rs:25` | delete `pub(crate) online: Gaussian` |
| `src/time_slice.rs:41` | delete `online: N_INF` from `Default` |
| `src/time_slice.rs:62,70-73` | drop `online` param and its branch from `Item::within_prior` |
| `src/time_slice.rs:110,120` | drop `online` param from `Event::within_priors` |
| `src/time_slice.rs:585,597,626,634` | drop `online` param from `TimeSlice::log_evidence`; `online \|\| forward` becomes `forward` |
| `src/history.rs:32,63,138,158,174,199,226` | delete the two `online` field declarations (`:32`, `:199`) and the five struct-literal copies |
| `src/history.rs:90-93` | delete `HistoryBuilder::online()` |
| `src/history.rs:402,410` | drop the `self.online` argument |
| `src/history.rs:1183-1189` | the `..._online` assertion becomes a `forward`-flag assertion; rename the binding to match what it tests |
`Skill` loses 16 bytes, which is a small independent win for #17.
## Testing strategy
Every new test is mutation-proved before it counts: break the production
line it names, watch it fail for the *right* assertion, restore. A test
never observed failing is not evidence.
### The red test
On the issue's own fixture — five 1v1 games, same winner each time —
`filtered_log_evidence()` must land strictly between the two known
endpoints:
```
5 × ln(0.5) = -3.4657... (today's inert value)
< filtered
< -0.4012... (batch / smoothed evidence)
```
Two-sided, so neither "still inert" nor "accidentally smoothed" can pass.
The lower bound is right for a real reason: game one genuinely *is* a
coin flip under filtering, games two through five are not.
### Invariants
1. **Invariant to `converge()`** — `filtered_log_evidence()` and
`filtered_learning_curves()` agree before and after `converge()`. This
is exactly what `skill.forward` fails, and what makes a stored field
the wrong mechanism.
Agreement is to tolerance, not bit-identity, and the reason is worth
recording. `iteration` calls `recompute_color_groups`
(`src/time_slice.rs:369`) only when `from == 0`, so a slice built by
repeated appends keeps insertion order until the first `converge()`
reorders it. The scratch clone inherits whichever order it finds, and
greedy coloring over a permuted input can group differently, giving a
different within-slice sweep order — same EP fixed point, different
path to it. Follow the house pattern in
`tests/ingestion_equivalence.rs`: converge tightly (`max_iter: 2_000`,
`epsilon: 1e-12`) and compare within `1e-8`.
2. **Invariant to ingestion order** — events added one at a time produce
the same filtered results as the same events batched. Extends the
existing invariant in `tests/ingestion_equivalence.rs`.
3. **Single-slice exactness** — for a history with one time slice there
is no future to propagate back, so filtered results equal smoothed
results exactly.
4. **Uncertainty ordering** — for a competitor with many later games, σ
at the first filtered point is greater than σ at the first smoothed
point, and less than the prior σ. This is the ustat complaint restated
as an assertion.
5. **Degenerate inputs** — empty history yields `0.0` and empty maps;
unknown key yields an empty curve. Added to
`tests/degenerate_inputs.rs`.
### Regression net
The existing suite must be unchanged by the removals: `log_evidence()`,
`log_evidence_for()`, and every numerical golden keep their current
values, since the default `online` was already `false` and the flag was
inert.
## Verification gates
- `just test` — full matrix, including the release job. `debug_assert!`
is compiled out in release, and that is where defects in this crate
have hidden before.
- `just lint` — clippy, warnings denied.
- `just fmt` — nightly.
- `just determinism` — the new pass must not perturb bit-identical
posteriors across `RAYON_NUM_THREADS` 1/2/4/8.
- `#![forbid(unsafe_code)]` stays.
## Risks
- **Clone cost.** One slice's events are cloned per slice visited. At
ustat scale this is negligible, but the pass is O(events) allocation on
top of O(events) inference. Accepted: fidelity to the production sweep
is worth more than avoiding the clone, and no caller is on a hot path.
- **`iterate_to_convergence` leaving test-only status.** Its doc comment
claims "only used by tests"; that comment must be updated, or it
becomes the next piece of load-bearing prose that is quietly false.
- **Event order is inherited, not normalised.** The scratch clone takes
the real slice's current event order, which differs pre- and
post-`converge()` for incrementally-ingested slices (see *Invariants*).
Results agree to within convergence tolerance rather than exactly.
Normalising the order in the scratch builder would buy bit-identity at
the cost of diverging from what the real sweep does; not worth it.
**Measured after implementation, this risk is smaller than stated.**
Flipping the scratch's `color_groups_dirty` from `true` to `false`
switches it between the grouped sweep (`sweep_color_groups`) and the
sequential fallback across its entire convergence loop — a far larger
perturbation than a permuted event order — and the ingestion-order
invariance test stays green at `1e-8` under `max_iter: 2_000`,
`epsilon: 1e-12`. EP reaches the same fixed point regardless of sweep
order once driven far enough. The tolerance caveat is correct but
conservative. Note the flag itself is load-bearing: with it `false` the
scratch would take the sequential path always, diverging from the
production sweep it exists to mirror.
- **Divergence risk.** If `TimeSlice`'s sweep gains state that the
scratch construction does not initialise, the pass silently reads a
default. The scratch builder must construct `Skill` field-by-field
rather than via `..Default::default()`, so adding a field to `Skill`
is a compile error here rather than a silent wrong answer.
## Out-of-scope follow-ups
File as separate issues:
1. **`forward: bool` is only a filtering quantity pre-convergence**
(`src/history.rs:395`). Either document the constraint or fold the
flag into the new pass and delete it.
2. **`log_evidence` takes `&mut self`** (`src/history.rs:416`) but
mutates nothing. The new `filtered_*` methods take `&self`; the
asymmetry is worth removing.
+21 -2
View File
@@ -46,14 +46,33 @@ fn main() {
.sigma(1.6)
.drift(ConstantDrift(0.036))
.convergence(trueskill_tt::ConvergenceOptions {
max_iter: 10,
// This history needs 30 sweeps to reach the epsilon below. It was
// capped at 10 until the `#[must_use]` on `ConvergenceReport`
// surfaced that the example had been shipping a short fit.
max_iter: 100,
epsilon: 0.01,
alpha: 1.0,
})
.build();
hist.add_events(events).unwrap();
hist.converge().unwrap();
// Read the report rather than discarding it. A fit that hits `max_iter`
// without reaching `epsilon` is not an error and does not look wrong — every
// rating comes back finite and sensibly ordered — so this flag is the only
// thing that says the numbers were still moving when the sweep stopped.
let report = hist.converge().unwrap();
eprintln!(
"converged={} after {} sweeps, final step {:?}",
report.converged, report.iterations, report.final_step
);
if !report.converged {
eprintln!(
"warning: stopped after {} sweeps with a final step of {:?}, \
short of epsilon — raise ConvergenceOptions::max_iter",
report.iterations, report.final_step
);
}
let players = [
("aggasi", "a092", 38800i64),
+14 -2
View File
@@ -1,2 +1,14 @@
publish = false
pre-release-hook = ["sh", "-c", "git cliff -o CHANGELOG.md --tag {{version}} && git add CHANGELOG.md"]
# Publish to the registry named in Cargo.toml's `publish` list (kellnr).
publish = true
# Hold off pushing until tags and publish have both succeeded; `just release`
# pushes last.
push = false
# Regenerate the changelog and stage it so it lands in the release commit.
#
# Guarded on DRY_RUN because cargo-release runs pre-release hooks during a dry
# run too (verified against cargo-release 1.1.5, which exports DRY_RUN=true,
# CRATE_NAME, PREV_VERSION and NEW_VERSION to the hook). Without the guard,
# `just release-plan` — documented as a preview that writes nothing — writes and
# `git add`s CHANGELOG.md, and the clean-tree check in `just release` then
# refuses to run. That check is load-bearing: publishing is irreversible.
pre-release-hook = ["sh", "-c", '[ "$DRY_RUN" = "true" ] || (git cliff -o CHANGELOG.md --tag {{version}} && git add CHANGELOG.md)']
+352
View File
@@ -0,0 +1,352 @@
//! Active learning: which comparison teaches you the most.
//!
//! [`quality`](crate::quality) answers "is this matchup *fair*". That is a
//! different question from "is this matchup *informative*", and the two
//! coincide only for two evenly matched competitors. When each observation
//! costs something — a human click, a scheduled fixture — the question worth
//! asking is the second one.
//!
//! The quantity here is expected information gain: the outcome-weighted
//! divergence between what you believe now and what you would believe after
//! seeing the result.
//!
//! ```text
//! EIG(matchup) = SUM P(outcome) * KL( posterior_after(outcome) || prior )
//! outcome
//! ```
//!
//! It is the mutual information between the observed outcome and the skills,
//! which is worth remembering because it pins the scale: information gain
//! cannot exceed the entropy of the thing you are about to observe. A contest
//! with `k` distinguishable outcomes can teach you at most `ln k` nats,
//! whatever the ratings. That ceiling is the sharpest available test of an
//! implementation — see [`expected_information_gain`].
use crate::{
GameOptions, Gaussian, InferenceError, Outcome, Rating, drift::Drift, predict, time::Time,
};
/// Outcomes below this probability contribute nothing measurable and are not
/// worth an inference pass.
///
/// The contribution of an outcome is `P * KL`, and `KL` is bounded in practice
/// by tens of nats, so a probability this small moves the total by less than
/// the quadrature error already present in `P` itself.
const NEGLIGIBLE: f64 = 1e-12;
/// `KL(q || p)` for two univariate Gaussians, in nats.
///
/// Both arguments are proper posteriors from inference, so the degenerate
/// cases guarded here (zero or infinite variance) indicate that inference has
/// broken down rather than anything a caller did.
fn kl_divergence(q: Gaussian, p: Gaussian) -> f64 {
let (var_q, var_p) = (q.sigma().powi(2), p.sigma().powi(2));
if !(var_q.is_finite() && var_p.is_finite()) || var_q <= 0.0 || var_p <= 0.0 {
return 0.0;
}
let mean_gap = q.mu() - p.mu();
0.5 * (libm::log(var_p / var_q) + (var_q + mean_gap * mean_gap) / var_p - 1.0)
}
/// Expected information gain of a hypothetical matchup, in nats.
///
/// Enumerates the outcomes this matchup could have, runs inference for each to
/// get the belief it would produce, and weights the resulting divergence by
/// that outcome's probability. A higher value means the result would teach you
/// more.
///
/// # Interpreting the value
///
/// Nats. The upper bound is the entropy of the outcome variable: at most
/// `ln 2 ≈ 0.693` for a two-way result, `ln 3 ≈ 1.099` once draws are
/// possible, `ln k` for `k` outcomes. A value near the ceiling means the
/// result is close to a coin flip *and* would move the posteriors a long way;
/// a value near zero means you already know what will happen, or that the
/// result would barely change your beliefs if you saw it.
///
/// This is not a monotone transform of [`quality`](crate::quality). A lopsided
/// matchup between two uncertain competitors scores well on quality-times-
/// variance heuristics and poorly here, because the near-certain outcome
/// carries almost no information.
///
/// # Cost
///
/// One full inference pass per possible outcome, so this is far more expensive
/// than `quality()` — which is one closed-form evaluation. The outcome count
/// grows quickly with team count (3 outcomes for two teams that can draw, 13
/// for three, 75 for four), and scoring every candidate pairing among `n`
/// competitors is `O(n² × outcomes)` inference passes.
///
/// For a selector over many candidates, shortlist with the cheap
/// [`quality`](crate::quality) or
/// [`predict_win_probabilities`](crate::History::predict_win_probabilities)
/// first and score only the shortlist here. The expected-variance-reduction
/// proxy sometimes suggested as a cheaper alternative is *not* cheaper: it
/// needs the same hypothetical posteriors, so it shares the dominant cost.
///
/// # Errors
///
/// - `NotEnoughTeams` if fewer than two teams are supplied.
/// - `EmptyTeam` if any team has no members.
/// - `TooManyTeams` if the outcome space is too large to enumerate; see
/// [`MAX_PREDICTED_TEAMS`](crate::MAX_PREDICTED_TEAMS).
/// - `InvalidProbability` if `options.p_draw` is outside `[0.0, 1.0)`.
/// - Anything [`Game::ranked`](crate::Game::ranked) returns for a hypothetical
/// outcome.
pub fn expected_information_gain<T: Time, D: Drift<T>>(
teams: &[&[Rating<T, D>]],
options: &GameOptions,
) -> Result<f64, InferenceError> {
if teams.len() < 2 {
return Err(InferenceError::NotEnoughTeams { got: teams.len() });
}
if teams.len() > crate::MAX_PREDICTED_TEAMS {
return Err(InferenceError::TooManyTeams {
got: teams.len(),
max: crate::MAX_PREDICTED_TEAMS,
});
}
if !(0.0..1.0).contains(&options.p_draw) {
return Err(InferenceError::InvalidProbability {
value: options.p_draw,
});
}
for (idx, team) in teams.iter().enumerate() {
if team.is_empty() {
return Err(InferenceError::EmptyTeam { team: idx });
}
}
// Prediction runs on performances: skill inflated by each member's beta.
let performances: Vec<Gaussian> = teams
.iter()
.map(|team| {
team.iter()
.fold(crate::N00, |acc, rating| acc + rating.performance())
})
.collect();
// Draw margins per pair, derived from the teams' betas exactly as
// inference derives them, so the outcomes weighted here are the outcomes
// that would actually be fitted.
let beta_sq: Vec<f64> = teams
.iter()
.map(|team| team.iter().map(|r| r.beta().powi(2)).sum())
.collect();
let p_draw = options.p_draw;
let margins = predict::Margins::new(teams.len(), |i, j| {
if p_draw == 0.0 {
0.0
} else {
crate::compute_margin(p_draw, (beta_sq[i] + beta_sq[j]).sqrt())
}
});
let mut gain = 0.0;
for (ranks, probability) in predict::outcome_distribution(&performances, &margins) {
if probability <= NEGLIGIBLE {
continue;
}
let game = crate::Game::ranked(teams, Outcome::ranking(ranks), options)?;
let posteriors = game.posteriors();
// Beliefs factorise across competitors, so the joint divergence is the
// sum of the per-competitor ones.
let divergence: f64 = teams
.iter()
.zip(&posteriors)
.flat_map(|(team, posterior)| team.iter().zip(posterior))
.map(|(rating, &after)| kl_divergence(after, rating.prior()))
.sum();
gain += probability * divergence;
}
Ok(gain)
}
#[cfg(test)]
mod tests {
use super::*;
use crate::{BETA, ConstantDrift, GAMMA};
type R = Rating<i64, ConstantDrift>;
fn rating(mu: f64, sigma: f64) -> R {
R::new(Gaussian::from_ms(mu, sigma), BETA, ConstantDrift(GAMMA))
}
fn options(p_draw: f64) -> GameOptions {
GameOptions {
p_draw,
..GameOptions::default()
}
}
fn eig(teams: &[&[R]], p_draw: f64) -> f64 {
expected_information_gain(teams, &options(p_draw)).unwrap()
}
/// The analytic ceiling. Information gain is the mutual information between
/// the outcome and the skills, so it cannot exceed the entropy of the
/// outcome variable — whatever the ratings. This is the check a subtly
/// wrong implementation fails while still returning plausible numbers: an
/// early prototype of this returned 4.77 nats from a sign error and passed
/// every monotonicity test.
#[test]
fn never_exceeds_the_entropy_of_the_outcome() {
let ceiling_two = std::f64::consts::LN_2;
for (a, b) in [
(rating(0.0, 6.0), rating(0.0, 6.0)),
(rating(0.0, 0.5), rating(0.0, 0.5)),
(rating(12.0, 6.0), rating(-12.0, 6.0)),
(rating(40.0, 1.0), rating(-40.0, 1.0)),
(rating(3.0, 6.0), rating(-2.0, 0.1)),
(rating(0.0, 25.0), rating(0.0, 25.0)),
] {
let g = eig(&[&[a], &[b]], 0.0);
assert!(
g >= 0.0 && g <= ceiling_two,
"EIG {g} outside [0, ln 2] for mu=({}, {}) sigma=({}, {})",
a.prior().mu(),
b.prior().mu(),
a.prior().sigma(),
b.prior().sigma()
);
}
}
/// With draws enabled there are three outcomes, so the ceiling rises to
/// `ln 3` — and the two-outcome bound no longer applies.
#[test]
fn the_ceiling_follows_the_outcome_count() {
let ceiling_three = 3.0f64.ln();
for sigma in [0.5, 3.0, 6.0, 25.0] {
let g = eig(&[&[rating(0.0, sigma)], &[rating(0.0, sigma)]], 0.25);
assert!(
g >= 0.0 && g <= ceiling_three,
"EIG {g} outside [0, ln 3] at sigma {sigma}"
);
}
}
/// An even matchup between uncertain competitors is the informative one.
/// A hopelessly lopsided matchup teaches you almost nothing, because you
/// already know how it ends.
#[test]
fn an_even_matchup_beats_a_lopsided_one() {
let even = eig(&[&[rating(0.0, 6.0)], &[rating(0.0, 6.0)]], 0.0);
let lopsided = eig(&[&[rating(12.0, 6.0)], &[rating(-12.0, 6.0)]], 0.0);
assert!(
even > lopsided,
"even {even} should beat lopsided {lopsided}"
);
}
/// Certainty is the thing information gain is measuring the absence of:
/// the less you know, the more there is to learn.
#[test]
fn gain_falls_as_certainty_rises() {
let mut previous = f64::INFINITY;
for sigma in [12.0, 6.0, 3.0, 1.0, 0.5, 0.1] {
let g = eig(&[&[rating(0.0, sigma)], &[rating(0.0, sigma)]], 0.0);
assert!(
g < previous,
"sigma {sigma}: {g} did not fall below {previous}"
);
previous = g;
}
assert!(previous >= 0.0);
}
/// The heuristic this replaces is `quality * sigma_a^2 * sigma_b^2`. It is
/// not a monotone transform of information gain — it ranks a lopsided
/// matchup above a confident even one, and EIG ranks them the other way.
/// Pinning the disagreement down is what stops a future "simplification"
/// from quietly reverting to the heuristic.
#[test]
fn disagrees_with_the_quality_times_variance_heuristic() {
let heuristic = |a: &R, b: &R| {
crate::quality(&[&[a.prior()], &[b.prior()]], BETA)
* a.prior().sigma().powi(2)
* b.prior().sigma().powi(2)
};
let (confident_a, confident_b) = (rating(0.0, 0.5), rating(0.0, 0.5));
let (lopsided_a, lopsided_b) = (rating(12.0, 6.0), rating(-12.0, 6.0));
assert!(
heuristic(&lopsided_a, &lopsided_b) > heuristic(&confident_a, &confident_b),
"the heuristic should prefer the lopsided matchup"
);
assert!(
eig(&[&[confident_a], &[confident_b]], 0.0) > eig(&[&[lopsided_a], &[lopsided_b]], 0.0),
"information gain should prefer the even matchup"
);
}
#[test]
fn supports_more_than_two_teams() {
let teams: Vec<Vec<R>> = vec![
vec![rating(0.0, 6.0)],
vec![rating(0.0, 6.0)],
vec![rating(0.0, 6.0)],
];
let refs: Vec<&[R]> = teams.iter().map(Vec::as_slice).collect();
let g = expected_information_gain(&refs, &options(0.0)).unwrap();
// Six distinguishable orderings with no draws.
assert!(
g > 0.0 && g <= 6.0f64.ln(),
"three-team EIG {g} out of range"
);
}
#[test]
fn multi_member_teams_are_supported() {
let a = [rating(0.0, 6.0), rating(1.0, 4.0)];
let b = [rating(0.0, 6.0)];
let g = expected_information_gain(&[&a, &b], &options(0.0)).unwrap();
assert!(g > 0.0 && g <= std::f64::consts::LN_2, "{g}");
}
#[test]
fn degenerate_shapes_are_errors() {
let a = [rating(0.0, 6.0)];
assert!(matches!(
expected_information_gain(&[&a], &options(0.0)),
Err(InferenceError::NotEnoughTeams { got: 1 })
));
let empty: [R; 0] = [];
assert!(matches!(
expected_information_gain(&[&a, &empty], &options(0.0)),
Err(InferenceError::EmptyTeam { team: 1 })
));
assert!(matches!(
expected_information_gain(&[&a, &a], &options(1.5)),
Err(InferenceError::InvalidProbability { .. })
));
}
#[test]
fn kl_divergence_is_zero_for_identical_beliefs() {
let g = Gaussian::from_ms(3.0, 2.0);
assert!(kl_divergence(g, g).abs() < 1e-15);
}
#[test]
fn kl_divergence_is_non_negative_and_grows_with_separation() {
let prior = Gaussian::from_ms(0.0, 3.0);
let mut previous = 0.0;
for mu in [0.0, 0.5, 1.0, 2.0, 4.0] {
let d = kl_divergence(Gaussian::from_ms(mu, 3.0), prior);
assert!(d >= 0.0, "negative divergence at mu {mu}: {d}");
assert!(d >= previous, "not increasing at mu {mu}");
previous = d;
}
}
}
+46 -11
View File
@@ -26,39 +26,75 @@ pub(crate) struct ColorGroups {
}
impl ColorGroups {
#[allow(dead_code)]
pub(crate) fn new() -> Self {
Self::default()
}
#[allow(dead_code)]
pub(crate) fn n_colors(&self) -> usize {
self.groups.len()
}
#[allow(dead_code)]
pub(crate) fn is_empty(&self) -> bool {
self.groups.is_empty()
}
/// Total event count across all colors.
#[allow(dead_code)]
/// Number of distinct colors in the partition. Test-only.
#[cfg(test)]
pub(crate) fn n_colors(&self) -> usize {
self.groups.len()
}
/// Total event count across all colors. Test-only.
#[cfg(test)]
pub(crate) fn total_events(&self) -> usize {
self.groups.iter().map(|g| g.len()).sum()
}
/// Contiguous index range for one color after events have been reordered
/// into color-contiguous positions by `TimeSlice::recompute_color_groups`.
#[allow(dead_code)]
pub(crate) fn color_range(&self, color_idx: usize) -> std::ops::Range<usize> {
let group = &self.groups[color_idx];
if group.is_empty() {
return 0..0;
}
let start = *group.first().unwrap();
let end = *group.last().unwrap() + 1;
debug_assert_eq!(
end - start,
group.len(),
"color {color_idx} is not contiguous; its range would overlap other colors"
);
start..end
}
/// Whether every color occupies a contiguous, ascending range of event
/// indices, and no two colors overlap.
///
/// The parallel sweep derives one `&mut` sub-slice per color from these
/// ranges and relies on them being disjoint. That disjointness is what
/// makes concurrent writes to distinct skills sound, so it is checked
/// rather than assumed.
pub(crate) fn groups_are_contiguous(&self) -> bool {
let mut expected_start = 0;
for group in &self.groups {
if group.is_empty() {
continue;
}
let ascending_run = group
.iter()
.enumerate()
.all(|(offset, &idx)| idx == group[0] + offset);
if !ascending_run || group[0] != expected_start {
return false;
}
expected_start += group.len();
}
true
}
}
/// Compute color groups greedily.
@@ -67,7 +103,6 @@ impl ColorGroups {
/// `Index` values that event touches. The returned `ColorGroups` has one
/// inner `Vec<usize>` per color, containing event indices in the order
/// they were assigned.
#[allow(dead_code)]
pub(crate) fn color_greedy<I, F>(n_events: usize, index_set: F) -> ColorGroups
where
F: Fn(usize) -> I,
+22 -16
View File
@@ -1,5 +1,4 @@
use crate::{
N_INF,
drift::{ConstantDrift, Drift},
gaussian::Gaussian,
rating::Rating,
@@ -8,12 +7,19 @@ use crate::{
/// Per-history, temporal state for someone competing.
///
/// Renamed from `Agent` in T2; the former `.player` field is now
/// `.rating` to match the `Player → Rating` rename.
/// The mutable half of a competitor: `Rating` holds their static
/// configuration, this holds what inference learns as it sweeps.
#[derive(Debug)]
pub struct Competitor<T: Time = i64, D: Drift<T> = ConstantDrift> {
pub rating: Rating<T, D>,
pub message: Gaussian,
/// The forward message carried from this competitor's last appearance, or
/// `None` before they have appeared anywhere.
///
/// Previously an improper `N_INF` served as the unset sentinel, which made
/// "no message yet" indistinguishable from "a legitimately improper
/// message" at the type level and required every reader to know the
/// convention.
pub message: Option<Gaussian>,
pub last_time: Option<T>,
}
@@ -21,14 +27,16 @@ impl<T: Time, D: Drift<T>> Competitor<T, D> {
/// Compute the message received at time `now`, with drift accumulated
/// from `self.last_time` (if any) to `now`.
pub(crate) fn receive(&self, now: &T) -> Gaussian {
if self.message != N_INF {
match self.message {
Some(message) => {
let elapsed_variance = match &self.last_time {
Some(last) => self.rating.drift.variance_delta(last, now),
Some(last) => self.rating.drift_variance_delta(last, now),
None => 0.0,
};
self.message.forget(elapsed_variance)
} else {
self.rating.prior
message.forget(elapsed_variance)
}
None => self.rating.prior,
}
}
@@ -37,11 +45,9 @@ impl<T: Time, D: Drift<T>> Competitor<T, D> {
/// Used in convergence sweeps where the elapsed was cached at slice-construction time
/// and should not be recomputed from `last_time` (which may have shifted).
pub(crate) fn receive_for_elapsed(&self, elapsed: i64) -> Gaussian {
if self.message != N_INF {
self.message
.forget(self.rating.drift.variance_for_elapsed(elapsed))
} else {
self.rating.prior
match self.message {
Some(message) => message.forget(self.rating.drift_variance_for_elapsed(elapsed)),
None => self.rating.prior,
}
}
}
@@ -50,7 +56,7 @@ impl Default for Competitor<i64, ConstantDrift> {
fn default() -> Self {
Self {
rating: Rating::default(),
message: N_INF,
message: None,
last_time: None,
}
}
@@ -63,7 +69,7 @@ where
C: Iterator<Item = &'a mut Competitor<T, D>>,
{
for c in competitors {
c.message = N_INF;
c.message = None;
if last_time {
c.last_time = None;
}
+34 -1
View File
@@ -20,6 +20,37 @@ pub struct ConvergenceOptions {
pub alpha: f64,
}
impl ConvergenceOptions {
/// Reject values that would make inference silently meaningless.
///
/// `HistoryBuilder::convergence` asserts these eagerly, but the fields are
/// public and `GameOptions` carries a `ConvergenceOptions` — so a caller
/// can hand `Game::ranked` a set the builder never saw. In release the
/// engine's `debug_assert!`s are gone, and an `alpha` of zero leaves every
/// EP update unapplied: inference returns the priors, with every likelihood
/// uninformative and nothing to indicate anything went wrong.
///
/// # Errors
///
/// `InvalidParameter` if `alpha` is outside `(0.0, 1.0]` or `epsilon` is
/// negative. NaN fails both comparisons and is rejected.
pub(crate) fn validate(&self) -> Result<(), crate::InferenceError> {
if !(self.alpha > 0.0 && self.alpha <= 1.0) {
return Err(crate::InferenceError::InvalidParameter {
name: "alpha",
value: self.alpha,
});
}
if self.epsilon.is_nan() || self.epsilon < 0.0 {
return Err(crate::InferenceError::InvalidParameter {
name: "epsilon",
value: self.epsilon,
});
}
Ok(())
}
}
impl Default for ConvergenceOptions {
fn default() -> Self {
Self {
@@ -32,13 +63,15 @@ impl Default for ConvergenceOptions {
/// Post-hoc summary of a `History::converge` call.
#[derive(Clone, Debug)]
#[must_use = "a ConvergenceReport carries `converged`, and a fit that stopped \
at `max_iter` is wrong by a little rather than loudly broken — \
check it, or bind it to `_` to say you have decided not to"]
pub struct ConvergenceReport {
pub iterations: usize,
pub final_step: (f64, f64),
pub log_evidence: f64,
pub converged: bool,
pub per_iteration_time: SmallVec<[Duration; 32]>,
pub slices_skipped: usize,
}
#[cfg(test)]
+147 -13
View File
@@ -1,6 +1,46 @@
use std::fmt;
/// How a prediction should treat a key the history has never seen.
///
/// Configured once per history via
/// [`HistoryBuilder::unknown_keys`](crate::HistoryBuilder::unknown_keys).
/// Neither known consumer wants this to vary between queries — one predicts
/// thousands of candidate matchups in a loop, the other's headline feature is
/// predicting a competitor nobody has faced — so it is a property of how you
/// intend to use the model rather than an argument on five call sites.
///
/// # There is deliberately no `Skip`
///
/// Dropping an unknown member is the obvious third option and it is wrong. A
/// team's performance is the *sum* of its members, so removing one removes its
/// variance too: measured on a two-member team with one unknown, skipping gives
/// a performance sigma of 2.37 where treating the member as unknown gives 6.53.
/// An unknown competitor would make the model *more* certain, which is
/// backwards. `Prior` is also the answer the model already gives for a
/// competitor it knows about but has no evidence for, so it corresponds to a
/// state the model can actually be in; skipping does not.
#[derive(Clone, Copy, Debug, Default, PartialEq, Eq)]
#[non_exhaustive]
pub enum UnknownKeys {
/// Reject the prediction with [`InferenceError::UnknownKey`].
///
/// The default, and the right one when every key is expected to be known:
/// a team of strangers should not silently produce a confident-looking
/// answer.
#[default]
Reject,
/// Treat an unknown competitor as one sitting at the history's configured
/// prior.
///
/// This is the honest Bayesian reading — a competitor you have never
/// observed is exactly the prior — and it makes "predict a matchup
/// involving someone new" a first-class question rather than something a
/// caller fakes with a neutral constant.
Prior,
}
#[derive(Debug, Clone, PartialEq)]
#[non_exhaustive]
pub enum InferenceError {
/// Expected and actual lengths of some array-shaped input differ.
MismatchedShape {
@@ -8,17 +48,73 @@ pub enum InferenceError {
expected: usize,
got: usize,
},
/// An `Outcome` of the wrong variant was supplied for the requested inference.
WrongOutcomeKind {
context: &'static str,
expected: &'static str,
got: &'static str,
},
/// A probability value is outside `[0, 1]`.
InvalidProbability { value: f64 },
/// A scalar parameter is outside its valid range.
InvalidParameter { name: &'static str, value: f64 },
/// Convergence exceeded `max_iter` without falling below `epsilon`.
ConvergenceFailed {
last_step: (f64, f64),
iterations: usize,
/// An event contains tied teams, but the draw probability is zero.
///
/// A zero draw probability asserts that draws cannot occur, so a tied
/// result has no representable likelihood. Configure a positive `p_draw`
/// (via `HistoryBuilder::p_draw` or `GameOptions::p_draw`) to admit ties.
TieWithoutDrawProbability { teams: (usize, usize) },
/// Inference produced a non-finite value (NaN or infinity).
///
/// Indicates numerical breakdown; the resulting skills are meaningless
/// and must not be treated as a converged estimate.
NonFiniteResult {
context: &'static str,
step: (f64, f64),
},
/// Negative precision: a Gaussian with `pi < 0` slipped into an API call.
NegativePrecision { pi: f64 },
/// One batch declared two different values for the same competitor's
/// configuration.
///
/// `prior` and `drift_scale` configure a competitor, not an event, so a
/// batch that sets one of them twice with different values has no
/// well-defined meaning: events within a batch are not ordered, so
/// "last one wins" would make the result depend on iteration order.
/// Declaring the same value repeatedly is fine and is the expected shape
/// when a competitor's configuration is a property of the domain.
ConflictingCompetitorConfig {
competitor: usize,
field: &'static str,
},
/// A prediction referenced a key the history has no skill for.
///
/// Reported rather than skipped: dropping unknown keys turns a team of
/// strangers into a confident-looking probability about nobody.
///
/// `key` is the offending key's `Debug` rendering. It is carried because
/// the indices alone are not actionable: a caller that logs
/// `UnknownKey { team: 0, member: 0 }` learns nothing about *which* of its
/// keys the history has not seen, and the natural handling — fall back to a
/// neutral value — turns the whole thing into a plausible constant.
UnknownKey {
team: usize,
member: usize,
key: String,
},
/// A prediction was given a team with no members.
EmptyTeam { team: usize },
/// A joint posterior was requested where one cannot be formed exactly.
JointUnavailable { reason: &'static str },
/// Fewer than two teams were supplied to a prediction.
NotEnoughTeams { got: usize },
/// The full outcome distribution was requested for too many teams.
///
/// Each realisation sorts into exactly one (order, tie-pattern) event, so
/// the space holds `n! * 2^(n-1)` members — 1_920 at five teams, 23_040 at
/// six, 322_560 at seven. Past `max` this stops being something to
/// enumerate on a caller's behalf; ask for individual rankings with
/// `predict_ranking`, or for `predict_win_probabilities`, both of which
/// stay cheap at any team count.
TooManyTeams { got: usize, max: usize },
}
impl fmt::Display for InferenceError {
@@ -31,23 +127,61 @@ impl fmt::Display for InferenceError {
} => {
write!(f, "{kind}: expected length {expected}, got {got}")
}
Self::WrongOutcomeKind {
context,
expected,
got,
} => {
write!(f, "{context}: expected {expected}, got {got}")
}
Self::InvalidProbability { value } => {
write!(f, "probability must be in [0, 1]; got {value}")
}
Self::TieWithoutDrawProbability { teams } => {
write!(
f,
"teams {} and {} are tied, but p_draw is 0.0; set a positive draw probability to admit ties",
teams.0, teams.1
)
}
Self::NonFiniteResult { context, step } => {
write!(
f,
"{context}: inference produced a non-finite result (step = {step:?})"
)
}
Self::InvalidParameter { name, value } => {
write!(f, "{name} is invalid: {value}")
}
Self::ConvergenceFailed {
last_step,
iterations,
} => {
Self::ConflictingCompetitorConfig { competitor, field } => {
write!(
f,
"convergence failed after {iterations} iterations; last step = {last_step:?}"
"competitor {competitor}: this batch sets {field} to two different values"
)
}
Self::NegativePrecision { pi } => {
write!(f, "precision must be non-negative; got {pi}")
Self::UnknownKey { team, member, key } => {
write!(
f,
"team {team}, member {member}: no skill recorded for key {key} \
(every key must already be known to the history; pre-filter \
with `lookup` or `current_skill` if that is not guaranteed)"
)
}
Self::EmptyTeam { team } => {
write!(f, "team {team} has no members")
}
Self::JointUnavailable { reason } => {
write!(f, "no exact joint posterior is available: {reason}")
}
Self::NotEnoughTeams { got } => {
write!(f, "prediction needs at least 2 teams, got {got}")
}
Self::TooManyTeams { got, max } => {
write!(
f,
"the outcome distribution over {got} teams is too large to enumerate (limit {max}); \
use predict_ranking or predict_win_probabilities instead"
)
}
}
}
+49 -6
View File
@@ -1,8 +1,10 @@
//! Typed event description for bulk ingestion.
//!
//! `Event<T, K>` is the new public event shape (spec Section 4). Replaces
//! the nested `Vec<Vec<Vec<Index>>>`, `Vec<Vec<f64>>`, `Vec<Vec<Vec<f64>>>`
//! that the old `add_events_with_prior` took.
//! `Event<T, K>` is the public event shape taken by `History::add_events`. It
//! is a typed front end, not a replacement: `add_events` flattens it into the
//! nested `Vec<Vec<Vec<Index>>>` / `Vec<Vec<f64>>` / `Vec<Vec<Vec<f64>>>` that
//! the internal `add_events_with_prior` chokepoint still takes, and which
//! `record_winner` and `record_draw` also route through.
use smallvec::SmallVec;
@@ -23,6 +25,7 @@ pub struct Team<K> {
}
impl<K> Team<K> {
#[must_use]
pub fn new() -> Self {
Self {
members: SmallVec::new(),
@@ -44,13 +47,28 @@ impl<K> Default for Team<K> {
/// One member of a team, identified by user key `K`.
///
/// `weight` defaults to 1.0; a per-event `prior` can override the competitor's
/// current skill estimate for this event only.
/// `weight` applies per event and defaults to 1.0.
///
/// `prior` and `drift_scale` are **competitor configuration**, not per-event
/// values. Setting either applies to the competitor for the whole history, not
/// just to this event, and applies whenever it is supplied — including on a key
/// the history already knows. Because configuration lives on the competitor and
/// `converge` refits from competitor state, configuring one late still refits
/// the whole history rather than taking effect only from that event onward.
///
/// Repeating the same value is inert, which is the expected shape when the
/// configuration is a property of the domain. Supplying two *different* values
/// for one competitor within a single batch is
/// `InferenceError::ConflictingCompetitorConfig`: events in a batch have no
/// order, so there would be no well-defined winner.
#[derive(Clone, Debug)]
pub struct Member<K> {
pub key: K,
pub weight: f64,
pub prior: Option<Gaussian>,
/// Multiplier on the drift *variance* this competitor accumulates.
/// `None` means 1.0.
pub drift_scale: Option<f64>,
}
impl<K> Member<K> {
@@ -59,6 +77,7 @@ impl<K> Member<K> {
key,
weight: 1.0,
prior: None,
drift_scale: None,
}
}
@@ -67,10 +86,31 @@ impl<K> Member<K> {
self
}
/// Set this competitor's starting skill estimate.
///
/// Captured at the competitor's first appearance; see the type docs.
pub fn with_prior(mut self, prior: Gaussian) -> Self {
self.prior = Some(prior);
self
}
/// Scale how fast this competitor drifts, relative to the history's drift.
///
/// The scale multiplies the drift *variance*, so it is in the same units as
/// `gamma`: `ConstantDrift(g)` at `scale = s` behaves exactly as
/// `ConstantDrift(g * s)` would for this competitor alone.
///
/// `0.0` pins the competitor still — useful for a reference point that
/// shares a scale with moving competitors but should not itself move: a bot
/// at a known strength, a rating floor, a course difficulty.
///
/// Captured at the competitor's first appearance; see the type docs.
/// Must be finite and non-negative, or ingestion fails with
/// [`InferenceError::InvalidParameter`](crate::InferenceError::InvalidParameter).
pub fn with_drift_scale(mut self, scale: f64) -> Self {
self.drift_scale = Some(scale);
self
}
}
/// Convenience: a member is a user key with default weight 1.0 and no prior.
@@ -91,15 +131,18 @@ mod tests {
assert_eq!(m.key, "alice");
assert_eq!(m.weight, 1.0);
assert!(m.prior.is_none());
assert!(m.drift_scale.is_none());
}
#[test]
fn member_builder_methods_chain() {
let m = Member::new("alice")
.with_weight(0.5)
.with_prior(Gaussian::from_ms(20.0, 5.0));
.with_prior(Gaussian::from_ms(20.0, 5.0))
.with_drift_scale(0.0);
assert_eq!(m.weight, 0.5);
assert!(m.prior.is_some());
assert_eq!(m.drift_scale, Some(0.0));
}
#[test]
+44 -8
View File
@@ -19,6 +19,14 @@ where
history: &'h mut History<T, D, O, K>,
event: Event<T, K>,
current_team_idx: Option<usize>,
/// First validation failure seen while building, surfaced by `commit`.
///
/// The setters return `Self` so the chain stays fluent; they cannot return
/// a `Result` without breaking that. Recording the failure and reporting it
/// at `commit` keeps the check enforced in release, where the previous
/// `debug_assert!` was compiled out and a mismatched event was ingested
/// silently.
error: Option<InferenceError>,
}
impl<'h, T, D, O, K> EventBuilder<'h, T, D, O, K>
@@ -37,6 +45,7 @@ where
outcome: Outcome::Ranked(SmallVec::new()),
},
current_team_idx: None,
error: None,
}
}
@@ -50,22 +59,36 @@ where
/// Set per-member weights for the most recently added team.
///
/// Panics in debug builds if called before `.team(...)` or if the length
/// doesn't match the team's member count.
/// A length mismatch is recorded and returned by [`EventBuilder::commit`]
/// as `InferenceError::MismatchedShape`, in both debug and release. The
/// weights are not applied in that case, so a partially-weighted team
/// cannot reach the history.
///
/// # Panics
///
/// Panics if called before any `.team(...)`.
pub fn weights<I: IntoIterator<Item = f64>>(mut self, weights: I) -> Self {
let idx = self
.current_team_idx
.expect(".weights(...) called before any .team(...)");
let ws: Vec<f64> = weights.into_iter().collect();
let team = &mut self.event.teams[idx];
debug_assert_eq!(
ws.len(),
team.members.len(),
"weights length must match team size"
);
if ws.len() != team.members.len() {
self.error.get_or_insert(InferenceError::MismatchedShape {
kind: "weights",
expected: team.members.len(),
got: ws.len(),
});
return self;
}
for (m, w) in team.members.iter_mut().zip(ws) {
m.weight = w;
}
self
}
@@ -84,7 +107,10 @@ where
/// Set explicit per-team continuous scores with a per-event noise override.
///
/// `sigma` overrides `HistoryBuilder::score_sigma` for this event only.
/// Must be `> 0.0`; debug-asserts otherwise via `Outcome::scores_with_sigma`.
/// Must be `> 0.0`. Constructing the outcome with a non-positive or NaN
/// sigma is allowed; the value is rejected with
/// `InferenceError::InvalidParameter` when the event is ingested, so
/// callers get an error from `commit` rather than a panic.
pub fn scores_with_sigma<I: IntoIterator<Item = f64>>(mut self, scores: I, sigma: f64) -> Self {
self.event.outcome = crate::Outcome::scores_with_sigma(scores, sigma);
self
@@ -103,7 +129,17 @@ where
}
/// Commit the event to the history.
///
/// # Errors
///
/// Returns the first validation failure recorded while building — see
/// [`EventBuilder::weights`] — otherwise forwards to
/// [`History::add_events`] and returns its errors.
pub fn commit(self) -> Result<(), InferenceError> {
if let Some(error) = self.error {
return Err(error);
}
self.history.add_events(std::iter::once(self.event))
}
}
+43 -19
View File
@@ -1,8 +1,8 @@
use crate::{
N_INF,
factor::{Factor, VarId, VarStore},
factor::{VarId, VarStore},
gaussian::Gaussian,
pdf,
ln_pdf,
};
/// Gaussian observation factor on a diff variable.
@@ -16,10 +16,11 @@ pub struct MarginFactor {
pub m_obs: f64,
pub sigma: f64,
pub(crate) msg: Gaussian,
pub(crate) evidence_cached: Option<f64>,
pub(crate) log_evidence_cached: Option<f64>,
}
impl MarginFactor {
#[must_use]
pub fn new(diff: VarId, m_obs: f64, sigma: f64) -> Self {
debug_assert!(sigma > 0.0, "score sigma must be positive");
Self {
@@ -27,7 +28,7 @@ impl MarginFactor {
m_obs,
sigma,
msg: N_INF,
evidence_cached: None,
log_evidence_cached: None,
}
}
}
@@ -40,8 +41,8 @@ impl MarginFactor {
let marginal = vars.get(self.diff);
let cavity = marginal / self.msg;
if self.evidence_cached.is_none() {
self.evidence_cached = Some(cavity_evidence(cavity, self.m_obs, self.sigma));
if self.log_evidence_cached.is_none() {
self.log_evidence_cached = Some(cavity_log_evidence(cavity, self.m_obs, self.sigma));
}
let new_msg = Gaussian::from_ms(self.m_obs, self.sigma);
@@ -54,19 +55,42 @@ impl MarginFactor {
}
}
impl Factor for MarginFactor {
fn propagate(&mut self, vars: &mut VarStore) -> (f64, f64) {
/// Undamped wrappers, used by this module's tests. Inference drives these
/// factors through `propagate_with_alpha` and reads the cached log evidence
/// directly, so these are not on any production path.
#[cfg(test)]
impl MarginFactor {
pub(crate) fn propagate(&mut self, vars: &mut VarStore) -> (f64, f64) {
self.propagate_with_alpha(vars, 1.0)
}
fn log_evidence(&self, _vars: &VarStore) -> f64 {
self.evidence_cached.unwrap_or(1.0).ln()
pub(crate) fn log_evidence(&self) -> f64 {
self.log_evidence_cached.unwrap_or(0.0)
}
}
fn cavity_evidence(cavity: Gaussian, m_obs: f64, sigma: f64) -> f64 {
let combined_sigma = (cavity.sigma().powi(2) + sigma.powi(2)).sqrt();
pdf(m_obs, cavity.mu(), combined_sigma)
/// `ln` of the observed margin's density under the cavity.
///
/// Computed in log space rather than as `pdf(..).ln()`. The density underflows
/// to zero past about 38 sigma of separation, and clamping that to
/// `f64::MIN_POSITIVE` reported -708 nats however far out the observation
/// actually was — 4292 nats adrift at 100 sigma, and unbounded beyond. A score
/// far from what the model expected is exactly the observation a log-evidence
/// figure exists to notice.
fn cavity_log_evidence(cavity: Gaussian, m_obs: f64, sigma: f64) -> f64 {
// `hypot`, not `sqrt(a^2 + b^2)`: squaring overflows to infinity above a
// sigma of ~1.3e154 and flushes to zero below ~1.5e-154, and `Gaussian`'s
// constructors are public so a caller can reach both.
let combined_sigma = cavity.sigma().hypot(sigma);
let value = ln_pdf(m_obs, cavity.mu(), combined_sigma);
// A degenerate cavity (infinite sigma) is the only way to reach a
// non-finite result; fall back to the old floor rather than emit -inf.
if value.is_finite() {
value
} else {
f64::MIN_POSITIVE.ln()
}
}
#[cfg(test)]
@@ -108,16 +132,16 @@ mod tests {
let mut vars = VarStore::new();
let diff = vars.alloc(Gaussian::from_ms(0.0, 6.0));
let mut f = MarginFactor::new(diff, 5.0, 1.0);
assert!(f.evidence_cached.is_none());
assert!(f.log_evidence_cached.is_none());
f.propagate(&mut vars);
let z = f.evidence_cached.unwrap();
// pdf(5, 0, sqrt(37)) 0.046783
assert!((z - 0.04678300292616668).abs() < 1e-10);
let z = f.log_evidence_cached.unwrap();
// ln pdf(5, 0, sqrt(37)) = ln(0.046783...)
assert!((z.exp() - 0.04678300292616668).abs() < 1e-10);
// Subsequent propagations don't change it.
f.propagate(&mut vars);
assert_eq!(f.evidence_cached.unwrap(), z);
assert_eq!(f.log_evidence_cached.unwrap(), z);
}
#[test]
@@ -126,7 +150,7 @@ mod tests {
let diff = vars.alloc(Gaussian::from_ms(0.0, 6.0));
let mut f = MarginFactor::new(diff, 5.0, 1.0);
f.propagate(&mut vars);
let logz = f.log_evidence(&vars);
let logz = f.log_evidence();
assert!((logz - (-3.062235327364623)).abs() < 1e-10);
}
+7 -71
View File
@@ -20,6 +20,9 @@ pub struct VarStore {
}
impl VarStore {
/// Test-only: inference allocates its store through `ScratchArena`.
#[cfg(test)]
#[must_use]
pub fn new() -> Self {
Self::default()
}
@@ -28,20 +31,20 @@ impl VarStore {
self.marginals.clear();
}
/// Test-only, as `new`.
#[cfg(test)]
#[must_use]
pub fn len(&self) -> usize {
self.marginals.len()
}
pub fn is_empty(&self) -> bool {
self.marginals.is_empty()
}
pub fn alloc(&mut self, init: Gaussian) -> VarId {
let id = VarId(self.marginals.len() as u32);
self.marginals.push(init);
id
}
#[must_use]
pub fn get(&self, id: VarId) -> Gaussian {
self.marginals[id.0 as usize]
}
@@ -51,58 +54,7 @@ impl VarStore {
}
}
/// A factor in the EP graph.
///
/// Factors hold their own outgoing messages and propagate them by reading
/// connected variable marginals from a `VarStore` and writing back updated
/// marginals.
pub trait Factor: Send + Sync {
/// Update outgoing messages and write back to the var store.
///
/// Returns the max delta `(|Δmu|, |Δsigma|)` across writes this
/// propagation. Used by the `Schedule` to detect convergence.
fn propagate(&mut self, vars: &mut VarStore) -> (f64, f64);
/// Optional log-evidence contribution. Default 0.0 (no contribution).
fn log_evidence(&self, _vars: &VarStore) -> f64 {
0.0
}
}
/// Enum dispatcher for the built-in factor types.
///
/// Using an enum instead of `Box<dyn Factor>` keeps factor data inline and
/// avoids virtual-call overhead in the hot inference loop.
#[derive(Debug)]
pub enum BuiltinFactor {
TeamSum(team_sum::TeamSumFactor),
RankDiff(rank_diff::RankDiffFactor),
Trunc(trunc::TruncFactor),
Margin(margin::MarginFactor),
}
impl Factor for BuiltinFactor {
fn propagate(&mut self, vars: &mut VarStore) -> (f64, f64) {
match self {
Self::TeamSum(f) => f.propagate(vars),
Self::RankDiff(f) => f.propagate(vars),
Self::Trunc(f) => f.propagate(vars),
Self::Margin(f) => f.propagate(vars),
}
}
fn log_evidence(&self, vars: &VarStore) -> f64 {
match self {
Self::Trunc(f) => f.log_evidence(vars),
Self::Margin(f) => f.log_evidence(vars),
Self::TeamSum(_) | Self::RankDiff(_) => 0.0,
}
}
}
pub mod margin;
pub mod rank_diff;
pub mod team_sum;
pub mod trunc;
#[cfg(test)]
@@ -149,20 +101,4 @@ mod tests {
assert_eq!(store.len(), 0);
assert_eq!(store.marginals.capacity(), cap);
}
#[test]
fn builtin_factor_dispatches_to_margin() {
use super::margin::MarginFactor;
let mut vars = VarStore::new();
let diff = vars.alloc(Gaussian::from_ms(0.0, 6.0));
let mut f = BuiltinFactor::Margin(MarginFactor::new(diff, 5.0, 1.0));
f.propagate(&mut vars);
let result = vars.get(diff);
assert!((result.mu() - 4.864864864864865).abs() < 1e-12);
let logz = f.log_evidence(&vars);
assert!((logz - (-3.062235327364623)).abs() < 1e-10);
}
}
-95
View File
@@ -1,95 +0,0 @@
use crate::factor::{Factor, VarId, VarStore};
/// Maintains the constraint `diff = team_a - team_b` between three vars.
///
/// On each propagation:
/// - Reads marginals at `team_a` and `team_b` (which already incorporate any
/// incoming messages from neighboring factors).
/// - Computes `new_diff = team_a - team_b` (variance addition; see Gaussian::Sub).
/// - Writes the new marginal to `diff`.
/// - Returns the delta against the previous diff value.
///
/// This factor does NOT store an outgoing message; the diff variable is
/// effectively replaced on each propagation. The TruncFactor on the same diff
/// var holds the EP-divide message that produces the cavity.
#[derive(Debug)]
pub struct RankDiffFactor {
pub team_a: VarId,
pub team_b: VarId,
pub diff: VarId,
}
impl Factor for RankDiffFactor {
fn propagate(&mut self, vars: &mut VarStore) -> (f64, f64) {
let a = vars.get(self.team_a);
let b = vars.get(self.team_b);
let new_diff = a - b;
let old = vars.get(self.diff);
vars.set(self.diff, new_diff);
old.delta(new_diff)
}
}
#[cfg(test)]
mod tests {
use super::*;
use crate::{N_INF, gaussian::Gaussian};
#[test]
fn diff_of_two_known_gaussians() {
let mut vars = VarStore::new();
let team_a = vars.alloc(Gaussian::from_ms(25.0, 3.0));
let team_b = vars.alloc(Gaussian::from_ms(20.0, 4.0));
let diff = vars.alloc(N_INF);
let mut f = RankDiffFactor {
team_a,
team_b,
diff,
};
f.propagate(&mut vars);
let result = vars.get(diff);
// mu = 25 - 20 = 5; var = 9 + 16 = 25; sigma = 5
assert!((result.mu() - 5.0).abs() < 1e-12);
assert!((result.sigma() - 5.0).abs() < 1e-12);
}
#[test]
fn delta_zero_on_repeat() {
let mut vars = VarStore::new();
let team_a = vars.alloc(Gaussian::from_ms(10.0, 2.0));
let team_b = vars.alloc(Gaussian::from_ms(8.0, 1.0));
let diff = vars.alloc(N_INF);
let mut f = RankDiffFactor {
team_a,
team_b,
diff,
};
f.propagate(&mut vars);
let (dmu, dsig) = f.propagate(&mut vars);
assert!(dmu < 1e-12);
assert!(dsig < 1e-12);
}
#[test]
fn delta_reflects_team_change() {
let mut vars = VarStore::new();
let team_a = vars.alloc(Gaussian::from_ms(10.0, 1.0));
let team_b = vars.alloc(Gaussian::from_ms(0.0, 1.0));
let diff = vars.alloc(N_INF);
let mut f = RankDiffFactor {
team_a,
team_b,
diff,
};
f.propagate(&mut vars);
// change team_a, repropagate; delta should be positive
vars.set(team_a, Gaussian::from_ms(15.0, 1.0));
let (dmu, _dsig) = f.propagate(&mut vars);
assert!(dmu > 4.0, "expected ~5 delta, got {}", dmu);
}
}
-98
View File
@@ -1,98 +0,0 @@
use crate::{
N00,
factor::{Factor, VarId, VarStore},
gaussian::Gaussian,
};
/// Computes the weighted sum of player performances into a team-perf var.
///
/// Inputs are pre-computed player performance Gaussians (i.e., rating priors
/// already with beta² noise added via `Rating::performance()`). The factor
/// runs once per game and writes the weighted sum to the output var.
#[derive(Debug)]
pub struct TeamSumFactor {
pub inputs: Vec<(Gaussian, f64)>,
pub out: VarId,
}
impl Factor for TeamSumFactor {
fn propagate(&mut self, vars: &mut VarStore) -> (f64, f64) {
let perf = self.inputs.iter().fold(N00, |acc, (g, w)| acc + (*g * *w));
let old = vars.get(self.out);
vars.set(self.out, perf);
old.delta(perf)
}
}
#[cfg(test)]
mod tests {
use super::*;
use crate::N_INF;
#[test]
fn single_player_unit_weight() {
let mut vars = VarStore::new();
let out = vars.alloc(N_INF);
let g = Gaussian::from_ms(25.0, 5.0);
let mut f = TeamSumFactor {
inputs: vec![(g, 1.0)],
out,
};
f.propagate(&mut vars);
let result = vars.get(out);
assert!((result.mu() - 25.0).abs() < 1e-12);
assert!((result.sigma() - 5.0).abs() < 1e-12);
}
#[test]
fn two_players_summed() {
let mut vars = VarStore::new();
let out = vars.alloc(N_INF);
let g1 = Gaussian::from_ms(20.0, 3.0);
let g2 = Gaussian::from_ms(30.0, 4.0);
let mut f = TeamSumFactor {
inputs: vec![(g1, 1.0), (g2, 1.0)],
out,
};
f.propagate(&mut vars);
let result = vars.get(out);
// sum: mu = 20 + 30 = 50, var = 9 + 16 = 25, sigma = 5
assert!((result.mu() - 50.0).abs() < 1e-12);
assert!((result.sigma() - 5.0).abs() < 1e-12);
}
#[test]
fn weighted_inputs() {
let mut vars = VarStore::new();
let out = vars.alloc(N_INF);
let g = Gaussian::from_ms(10.0, 2.0);
let mut f = TeamSumFactor {
inputs: vec![(g, 2.0)],
out,
};
f.propagate(&mut vars);
let result = vars.get(out);
// g * 2.0: mu = 10*2 = 20, sigma = 2*2 = 4
assert!((result.mu() - 20.0).abs() < 1e-12);
assert!((result.sigma() - 4.0).abs() < 1e-12);
}
#[test]
fn delta_is_zero_on_repeat_propagate() {
let mut vars = VarStore::new();
let out = vars.alloc(N_INF);
let g = Gaussian::from_ms(5.0, 1.0);
let mut f = TeamSumFactor {
inputs: vec![(g, 1.0)],
out,
};
f.propagate(&mut vars);
let (dmu, dsig) = f.propagate(&mut vars);
assert!(dmu < 1e-12, "expected ~0 delta on repeat, got {}", dmu);
assert!(dsig < 1e-12);
}
}
+115 -25
View File
@@ -1,7 +1,8 @@
use crate::{
N_INF, approx, cdf,
factor::{Factor, VarId, VarStore},
N_INF, approx,
factor::{VarId, VarStore},
gaussian::Gaussian,
ln_interval, ln_sf,
};
/// EP truncation factor on a diff variable.
@@ -15,20 +16,21 @@ pub struct TruncFactor {
pub diff: VarId,
pub margin: f64,
pub tie: bool,
/// Outgoing message to the diff variable (initial: N_INF, the EP identity).
/// Outgoing message to the diff variable (initial: `N_INF`, the EP identity).
pub(crate) msg: Gaussian,
/// Cached evidence (linear, not log) computed from the cavity on first propagation.
pub(crate) evidence_cached: Option<f64>,
pub(crate) log_evidence_cached: Option<f64>,
}
impl TruncFactor {
#[must_use]
pub fn new(diff: VarId, margin: f64, tie: bool) -> Self {
Self {
diff,
margin,
tie,
msg: N_INF,
evidence_cached: None,
log_evidence_cached: None,
}
}
}
@@ -41,8 +43,8 @@ impl TruncFactor {
let marginal = vars.get(self.diff);
let cavity = marginal / self.msg;
if self.evidence_cached.is_none() {
self.evidence_cached = Some(cavity_evidence(cavity, self.margin, self.tie));
if self.log_evidence_cached.is_none() {
self.log_evidence_cached = Some(cavity_log_evidence(cavity, self.margin, self.tie));
}
let trunc = approx(cavity, self.margin, self.tie);
@@ -61,22 +63,39 @@ impl TruncFactor {
}
}
impl Factor for TruncFactor {
fn propagate(&mut self, vars: &mut VarStore) -> (f64, f64) {
/// Undamped wrappers, used by this module's tests. Inference drives these
/// factors through `propagate_with_alpha` and reads the cached log evidence
/// directly, so these are not on any production path.
#[cfg(test)]
impl TruncFactor {
pub(crate) fn propagate(&mut self, vars: &mut VarStore) -> (f64, f64) {
self.propagate_with_alpha(vars, 1.0)
}
fn log_evidence(&self, _vars: &VarStore) -> f64 {
self.evidence_cached.unwrap_or(1.0).ln()
}
}
/// P(diff > margin) for non-tie, P(|diff| < margin) for tie.
fn cavity_evidence(diff: Gaussian, margin: f64, tie: bool) -> f64 {
if tie {
cdf(margin, diff.mu(), diff.sigma()) - cdf(-margin, diff.mu(), diff.sigma())
/// `ln P(diff > margin)` for a win, `ln P(|diff| < margin)` for a tie.
///
/// Computed in log space throughout. Two earlier shapes both lost the tail:
/// `1 - cdf(..)` cancelled away every digit of an unlikely outcome, and even
/// once that was fixed the linear probability underflows to zero past about 38
/// sigma, where clamping reported -708 nats regardless of the truth. An upset
/// is the observation a log-evidence figure exists to notice, so it has to stay
/// exact precisely where it is smallest.
fn cavity_log_evidence(diff: Gaussian, margin: f64, tie: bool) -> f64 {
let (mu, sigma) = (diff.mu(), diff.sigma());
let value = if tie {
ln_interval(-margin, margin, mu, sigma)
} else {
1.0 - cdf(margin, diff.mu(), diff.sigma())
ln_sf(margin, mu, sigma)
};
// A degenerate cavity is the only route to a non-finite result; keep the
// old floor for it rather than letting -inf poison the whole history's sum.
if value.is_finite() {
value
} else {
f64::MIN_POSITIVE.ln()
}
}
@@ -108,19 +127,90 @@ mod tests {
let diff = vars.alloc(Gaussian::from_ms(2.0, 3.0));
let mut f = TruncFactor::new(diff, 0.0, false);
assert!(f.evidence_cached.is_none());
assert!(f.log_evidence_cached.is_none());
f.propagate(&mut vars);
assert!(f.evidence_cached.is_some());
let first = f.evidence_cached.unwrap();
assert!(f.log_evidence_cached.is_some());
let first = f.log_evidence_cached.unwrap();
// Evidence should be P(diff > 0) for diff ~ N(2, 9) ≈ 0.748
assert!(first > 0.7);
assert!(first < 0.8);
assert!(first.exp() > 0.7);
assert!(first.exp() < 0.8);
// Subsequent propagations don't change it.
f.propagate(&mut vars);
assert_eq!(f.evidence_cached.unwrap(), first);
assert_eq!(f.log_evidence_cached.unwrap(), first);
}
/// The defect this guards: `1 - cdf` collapsed to zero for a surprising
/// result, the clamp turned that into `f64::MIN_POSITIVE`, and
/// `log_evidence` reported ln of *that* — about -708 whatever the truth
/// was. An upset is the observation a model-comparison score exists to
/// notice, so it was wrong exactly where it mattered.
#[test]
fn evidence_of_an_upset_is_not_flattened_to_the_clamp_floor() {
// diff ~ N(-9, 1) with margin 0: the favoured side lost by nine sigma.
let evidence = cavity_log_evidence(Gaussian::from_ms(-9.0, 1.0), 0.0, false).exp();
assert!(
evidence > f64::MIN_POSITIVE,
"evidence collapsed onto the clamp floor: {evidence}"
);
// P(X > 0) for X ~ N(-9, 1) is the standard normal tail at 9 sigma.
assert!(
(evidence - 1.128_588e-19).abs() / 1.128_588e-19 < 1e-6,
"expected ~1.13e-19, got {evidence}"
);
assert!(
(evidence.ln() + 43.628).abs() < 1e-2,
"log evidence {} should be about -43.6, not -708",
evidence.ln()
);
}
/// Evidence must stay finite and positive however extreme the mismatch,
/// since `log_evidence` sums across the whole history and one `-inf` or
/// `NaN` poisons all of it.
///
/// Finiteness alone is too weak a bar — the clamped version was finite too,
/// and wrong by hundreds of nats. `log_evidence_tracks_the_analytic_tail`
/// below is the assertion that actually holds this up.
#[test]
fn evidence_stays_positive_and_finite_at_any_separation() {
for mu in [-300.0f64, -50.0, -9.0, 0.0, 9.0, 50.0, 300.0] {
for tie in [false, true] {
let ln_e = cavity_log_evidence(Gaussian::from_ms(mu, 1.0), 1.0, tie);
assert!(
ln_e.is_finite() && ln_e <= 0.0,
"mu={mu} tie={tie}: log evidence {ln_e} is not a log-probability"
);
}
}
}
/// The clamp used to floor everything past ~38 sigma at `ln(MIN_POSITIVE)`
/// = -708, however far out the real observation was. In log space the
/// answer is a polynomial and stays exact: at 1000 sigma the truth is about
/// -500_000 nats, and -708 is not a rounding error.
#[test]
fn log_evidence_tracks_the_analytic_tail() {
for mu in [-40.0f64, -60.0, -100.0, -1000.0] {
// P(diff > 0) for diff ~ N(mu, 1), mu far below zero.
let got = cavity_log_evidence(Gaussian::from_ms(mu, 1.0), 0.0, false);
// ln Phi(mu) ~ -mu^2/2 - ln(-mu) - ln(sqrt(2 pi)) for mu << 0.
let z = -mu;
let approx = -0.5 * z * z - z.ln() - (2.0 * std::f64::consts::PI).sqrt().ln();
assert!(
got < f64::MIN_POSITIVE.ln(),
"mu={mu}: {got} is still stuck on the old clamp floor"
);
assert!(
(got - approx).abs() / approx.abs() < 1e-3,
"mu={mu}: got {got}, asymptotic expectation {approx}"
);
}
}
#[test]
@@ -132,7 +222,7 @@ mod tests {
f.propagate(&mut vars);
// For diff ~ N(0, 4), tie=true with margin=1: P(-1 < diff < 1) ≈ 0.383
let ev = f.evidence_cached.unwrap();
let ev = f.log_evidence_cached.unwrap().exp();
assert!(ev > 0.35 && ev < 0.42);
}
-13
View File
@@ -1,13 +0,0 @@
//! Factor-graph public API.
//!
//! Power users can construct custom factor graphs via `Game::custom` (T2
//! minimal; full ergonomics in T4) and drive them with custom `Schedule`
//! implementations.
pub use crate::{
factor::{
BuiltinFactor, Factor, VarId, VarStore, margin::MarginFactor, rank_diff::RankDiffFactor,
team_sum::TeamSumFactor, trunc::TruncFactor,
},
schedule::{EpsilonOrMax, Schedule, ScheduleReport},
};
+117 -78
View File
@@ -37,10 +37,17 @@ impl DiffFactor {
}
}
pub(crate) fn evidence(&self) -> f64 {
/// Log of this link's cached evidence.
///
/// Accumulating in log space keeps a long diff chain from underflowing:
/// each link contributes a probability in `(0, 1]`, so the linear product
/// over an n-team game decays geometrically and flushes to zero — and
/// `ln(0.0)` is `-inf` — well within the team counts a large free-for-all
/// reaches.
pub(crate) fn log_evidence(&self) -> f64 {
match self {
Self::Trunc(f) => f.evidence_cached.unwrap_or(1.0),
Self::Margin(f) => f.evidence_cached.unwrap_or(1.0),
Self::Trunc(f) => f.log_evidence_cached.unwrap_or(0.0),
Self::Margin(f) => f.log_evidence_cached.unwrap_or(0.0),
}
}
@@ -81,18 +88,14 @@ impl Default for GameOptions {
/// Owned variant of `Game` returned by public constructors.
///
/// Unlike `Game<'a, T, D>` (which borrows its result/weights slices from
/// History's internal state), `OwnedGame<T, D>` owns its inputs so it can
/// be returned freely from public constructors.
/// History's internal state), `OwnedGame<T, D>` owns the team ratings, so it
/// can be returned freely from public constructors. The inference inputs
/// themselves are not retained — nothing reads them back.
#[derive(Debug)]
#[allow(dead_code)]
pub struct OwnedGame<T: Time, D: Drift<T>> {
teams: Vec<Vec<Rating<T, D>>>,
result: Vec<f64>,
weights: Vec<Vec<f64>>,
p_draw: f64,
pub(crate) convergence: crate::ConvergenceOptions,
pub(crate) likelihoods: Vec<Vec<Gaussian>>,
pub(crate) evidence: f64,
pub(crate) log_evidence: f64,
}
impl<T: Time, D: Drift<T>> OwnedGame<T, D> {
@@ -104,24 +107,15 @@ impl<T: Time, D: Drift<T>> OwnedGame<T, D> {
convergence: crate::ConvergenceOptions,
) -> Self {
let mut arena = ScratchArena::new();
let g = Game::ranked_with_arena(
teams.clone(),
&result,
&weights,
p_draw,
convergence,
&mut arena,
);
let likelihoods = g.likelihoods;
let evidence = g.evidence;
// `Game` takes the teams by value and is dropped here, so take the vec
// back out of it rather than handing it a clone.
let g = Game::ranked_with_arena(teams, &result, &weights, p_draw, convergence, &mut arena);
Self {
teams,
result,
weights,
p_draw,
convergence,
likelihoods,
evidence,
teams: g.teams,
likelihoods: g.likelihoods,
log_evidence: g.log_evidence,
}
}
@@ -133,27 +127,24 @@ impl<T: Time, D: Drift<T>> OwnedGame<T, D> {
convergence: crate::ConvergenceOptions,
) -> Self {
let mut arena = ScratchArena::new();
let g = Game::scored_with_arena(
teams.clone(),
teams,
&scores,
&weights,
score_sigma,
convergence,
&mut arena,
);
let likelihoods = g.likelihoods;
let evidence = g.evidence;
Self {
teams,
result: scores,
weights,
p_draw: 0.0,
convergence,
likelihoods,
evidence,
teams: g.teams,
likelihoods: g.likelihoods,
log_evidence: g.log_evidence,
}
}
#[must_use]
pub fn posteriors(&self) -> Vec<Vec<Gaussian>> {
self.likelihoods
.iter()
@@ -162,8 +153,9 @@ impl<T: Time, D: Drift<T>> OwnedGame<T, D> {
.collect()
}
#[must_use]
pub fn log_evidence(&self) -> f64 {
self.evidence.ln()
self.log_evidence
}
}
@@ -175,7 +167,7 @@ pub struct Game<'a, T: Time = i64, D: Drift<T> = crate::drift::ConstantDrift> {
p_draw: f64,
pub(crate) convergence: crate::ConvergenceOptions,
pub(crate) likelihoods: Vec<Vec<Gaussian>>,
pub(crate) evidence: f64,
pub(crate) log_evidence: f64,
}
impl<'a, T: Time, D: Drift<T>> Game<'a, T, D> {
@@ -222,7 +214,7 @@ impl<'a, T: Time, D: Drift<T>> Game<'a, T, D> {
p_draw,
convergence,
likelihoods: Vec::new(),
evidence: 0.0,
log_evidence: 0.0,
};
this.likelihoods(arena);
@@ -261,7 +253,7 @@ impl<'a, T: Time, D: Drift<T>> Game<'a, T, D> {
p_draw: 0.0,
convergence,
likelihoods: Vec::new(),
evidence: 0.0,
log_evidence: 0.0,
};
this.likelihoods_scored(arena, score_sigma);
@@ -355,7 +347,7 @@ impl<'a, T: Time, D: Drift<T>> Game<'a, T, D> {
arena.lhood_lose[n_teams - 1] = pw_last - links[n_diffs - 1].msg();
}
let evidence: f64 = links.iter().map(|l| l.evidence()).product();
let log_evidence: f64 = links.iter().map(DiffFactor::log_evidence).sum();
// Inverse permutation: inv_buf[orig_i] = sorted_i.
arena.inv_buf.resize(n_teams, 0);
@@ -371,10 +363,9 @@ impl<'a, T: Time, D: Drift<T>> Game<'a, T, D> {
.map(|(orig_i, (players, weights))| {
let si = arena.inv_buf[orig_i];
let m = arena.lhood_win[si] * arena.lhood_lose[si];
let performance = players
.iter()
.zip(weights.iter())
.fold(N00, |p, (player, &w)| p + (player.performance() * w));
// Already folded into `team_prior` at the top of the chain,
// indexed by sorted position.
let performance = arena.team_prior[si];
players
.iter()
.zip(weights.iter())
@@ -386,11 +377,11 @@ impl<'a, T: Time, D: Drift<T>> Game<'a, T, D> {
})
.collect::<Vec<_>>();
(evidence, likelihoods)
(log_evidence, likelihoods)
}
fn likelihoods(&mut self, arena: &mut ScratchArena) {
let (evidence, likelihoods) = self.run_chain(arena, |i, sort_buf, vars| {
let (log_evidence, likelihoods) = self.run_chain(arena, |i, sort_buf, vars| {
let tie = self.result[sort_buf[i]] == self.result[sort_buf[i + 1]];
let margin = if self.p_draw == 0.0 {
0.0
@@ -405,20 +396,21 @@ impl<'a, T: Time, D: Drift<T>> Game<'a, T, D> {
let vid = vars.alloc(N_INF);
DiffFactor::Trunc(TruncFactor::new(vid, margin, tie))
});
self.evidence = evidence;
self.log_evidence = log_evidence;
self.likelihoods = likelihoods;
}
fn likelihoods_scored(&mut self, arena: &mut ScratchArena, score_sigma: f64) {
let (evidence, likelihoods) = self.run_chain(arena, |i, sort_buf, vars| {
let (log_evidence, likelihoods) = self.run_chain(arena, |i, sort_buf, vars| {
let m_obs = self.result[sort_buf[i]] - self.result[sort_buf[i + 1]];
let vid = vars.alloc(N_INF);
DiffFactor::Margin(MarginFactor::new(vid, m_obs, score_sigma))
});
self.evidence = evidence;
self.log_evidence = log_evidence;
self.likelihoods = likelihoods;
}
#[must_use]
pub fn posteriors(&self) -> Vec<Vec<Gaussian>> {
self.likelihoods
.iter()
@@ -432,17 +424,30 @@ impl<'a, T: Time, D: Drift<T>> Game<'a, T, D> {
.collect::<Vec<_>>()
}
#[must_use]
pub fn log_evidence(&self) -> f64 {
self.evidence.ln()
self.log_evidence
}
}
impl<T: Time, D: Drift<T>> Game<'_, T, D> {
/// # Errors
///
/// - `InvalidParameter` if `options.convergence` is out of range — an
/// `alpha` of zero would leave every EP update unapplied and silently
/// return the priors.
/// - `InvalidProbability` if `options.p_draw` is outside `[0.0, 1.0)`.
/// - `MismatchedShape` if the outcome's rank count differs from `teams.len()`.
/// - `WrongOutcomeKind` if `outcome` is not `Outcome::Ranked`.
/// - `TieWithoutDrawProbability` if the outcome ties two teams while
/// `p_draw` is zero: the truncation margin is then zero and the two-sided
/// tie update evaluates `0/0`.
pub fn ranked(
teams: &[&[Rating<T, D>]],
outcome: crate::Outcome,
options: &GameOptions,
) -> Result<OwnedGame<T, D>, crate::InferenceError> {
options.convergence.validate()?;
if !(0.0..1.0).contains(&options.p_draw) {
return Err(crate::InferenceError::InvalidProbability {
value: options.p_draw,
@@ -458,11 +463,22 @@ impl<T: Time, D: Drift<T>> Game<'_, T, D> {
let ranks = outcome
.as_ranks()
.ok_or(crate::InferenceError::MismatchedShape {
kind: "Game::ranked requires Outcome::Ranked",
expected: 0,
got: 0,
.ok_or(crate::InferenceError::WrongOutcomeKind {
context: "Game::ranked",
expected: "Outcome::Ranked",
got: "Outcome::Scored",
})?;
let tied = if options.p_draw == 0.0 {
crate::first_tied_pair(ranks)
} else {
None
};
if let Some(teams) = tied {
return Err(crate::InferenceError::TieWithoutDrawProbability { teams });
}
let max_rank = ranks.iter().copied().max().unwrap_or(0) as f64;
let result: Vec<f64> = ranks.iter().map(|&r| max_rank - r as f64).collect();
let teams_owned: Vec<Vec<Rating<T, D>>> = teams.iter().map(|t| t.to_vec()).collect();
@@ -477,11 +493,18 @@ impl<T: Time, D: Drift<T>> Game<'_, T, D> {
))
}
/// # Errors
///
/// - `InvalidParameter` if `options.score_sigma` is not strictly positive
/// or is NaN, or if `options.convergence` is out of range.
/// - `MismatchedShape` if the outcome's score count differs from `teams.len()`.
/// - `WrongOutcomeKind` if `outcome` is not `Outcome::Scored`.
pub fn scored(
teams: &[&[Rating<T, D>]],
outcome: crate::Outcome,
options: &GameOptions,
) -> Result<OwnedGame<T, D>, crate::InferenceError> {
options.convergence.validate()?;
if options.score_sigma <= 0.0 || options.score_sigma.is_nan() {
return Err(crate::InferenceError::InvalidParameter {
name: "score_sigma",
@@ -497,10 +520,10 @@ impl<T: Time, D: Drift<T>> Game<'_, T, D> {
}
let scores = outcome
.as_scores()
.ok_or(crate::InferenceError::MismatchedShape {
kind: "Game::scored requires Outcome::Scored",
expected: 0,
got: 0,
.ok_or(crate::InferenceError::WrongOutcomeKind {
context: "Game::scored",
expected: "Outcome::Scored",
got: "Outcome::Ranked",
})?
.to_vec();
let teams_owned: Vec<Vec<Rating<T, D>>> = teams.iter().map(|t| t.to_vec()).collect();
@@ -514,16 +537,28 @@ impl<T: Time, D: Drift<T>> Game<'_, T, D> {
))
}
/// Convenience wrapper over [`Game::ranked`] for two single-player teams.
///
/// # Errors
///
/// Delegates to [`Game::ranked`], so it returns the same errors — in
/// practice `WrongOutcomeKind` for a non-ranked outcome, or
/// `TieWithoutDrawProbability` for a draw when `options.p_draw` is zero.
pub fn one_v_one(
a: &Rating<T, D>,
b: &Rating<T, D>,
outcome: crate::Outcome,
options: &GameOptions,
) -> Result<(Gaussian, Gaussian), crate::InferenceError> {
let game = Self::ranked(&[&[*a], &[*b]], outcome, &GameOptions::default())?;
let game = Self::ranked(&[&[*a], &[*b]], outcome, options)?;
let post = game.posteriors();
Ok((post[0][0], post[1][0]))
}
/// # Errors
///
/// Wraps each player in a one-member team and delegates to
/// [`Game::ranked`], so it returns the same errors.
pub fn free_for_all(
players: &[&Rating<T, D>],
outcome: crate::Outcome,
@@ -533,15 +568,6 @@ impl<T: Time, D: Drift<T>> Game<'_, T, D> {
let team_refs: Vec<&[Rating<T, D>]> = teams.iter().map(|t| t.as_slice()).collect();
Self::ranked(&team_refs, outcome, options)
}
#[doc(hidden)]
pub fn custom<S: crate::factors::Schedule>(
factors: &mut [crate::factors::BuiltinFactor],
vars: &mut crate::factors::VarStore,
schedule: &S,
) -> crate::factors::ScheduleReport {
schedule.run(factors, vars)
}
}
#[cfg(test)]
@@ -698,9 +724,15 @@ mod tests {
let c = p[2][0];
// T1 ULP shift: mu rounds to 25.0 (was 24.999999) under natural-parameter storage.
//
// The 1e-6-place values moved when `erfc_inv`'s sign error was fixed:
// this case runs at `p_draw = 0.5`, so it goes through `compute_margin`,
// and the margin is now 8.4e-8 from the exact quantile where it was
// 1.46e-7. Verified as movement *toward* analytic truth, not a
// regression — see `erfc_inv_matches_known_quantiles`.
assert_ulps_eq!(a, Gaussian::from_ms(25.0, 6.092561), epsilon = 1e-6);
assert_ulps_eq!(b, Gaussian::from_ms(33.379314, 6.483575), epsilon = 1e-6);
assert_ulps_eq!(c, Gaussian::from_ms(16.620685, 6.483575), epsilon = 1e-6);
assert_ulps_eq!(b, Gaussian::from_ms(33.379315, 6.483576), epsilon = 1e-6);
assert_ulps_eq!(c, Gaussian::from_ms(16.620685, 6.483576), epsilon = 1e-6);
}
#[test]
@@ -730,8 +762,12 @@ mod tests {
let a = p[0][0];
let b = p[1][0];
assert_ulps_eq!(a, Gaussian::from_ms(24.999999, 6.469480), epsilon = 1e-6);
assert_ulps_eq!(b, Gaussian::from_ms(24.999999, 6.469480), epsilon = 1e-6);
// Two identical competitors drawing must land on their shared prior
// mean exactly, by symmetry. The reference transcription of 24.999999
// is that value rounded to six decimals; asserting it at epsilon 1e-6
// left no headroom. The root-free variance path now hits 25.0 exactly.
assert_ulps_eq!(a, Gaussian::from_ms(25.0, 6.469480), epsilon = 1e-6);
assert_ulps_eq!(b, Gaussian::from_ms(25.0, 6.469480), epsilon = 1e-6);
let t_a = R::new(
Gaussian::from_ms(25.0, 3.0),
@@ -1124,7 +1160,10 @@ mod tests {
&GameOptions::default(),
)
.unwrap_err();
assert!(matches!(err, crate::InferenceError::MismatchedShape { .. }));
assert!(matches!(
err,
crate::InferenceError::WrongOutcomeKind { .. }
));
}
#[test]
@@ -1206,7 +1245,7 @@ mod tests {
);
assert_ulps_eq!(
p[1][0],
Gaussian::from_ms(19.287197, 7.243465),
Gaussian::from_ms(19.287198285, 7.243465848),
epsilon = 1e-6
);
assert_ulps_eq!(
@@ -1266,7 +1305,7 @@ mod tests {
assert_ulps_eq!(
p[0][0],
Gaussian::from_ms(31.674697, 7.501180),
Gaussian::from_ms(31.674698083, 7.501180037),
epsilon = 1e-6
);
assert_ulps_eq!(
+155 -13
View File
@@ -18,6 +18,7 @@ pub struct Gaussian {
impl Gaussian {
/// Construct from mean and standard deviation.
#[must_use]
pub const fn from_ms(mu: f64, sigma: f64) -> Self {
if sigma == f64::INFINITY {
Self { pi: 0.0, tau: 0.0 }
@@ -35,6 +36,28 @@ impl Gaussian {
}
}
/// Construct from mean and *variance*, skipping the square-root round trip.
///
/// `from_ms(mu, var.sqrt())` immediately squares the root away again to
/// recover `pi = 1/var`. Variance-combining operations (`Add`, `Sub`,
/// `exclude`, `forget`) work in variance space throughout, so they go
/// through here instead and never take a root.
#[inline]
pub(crate) fn from_mv(mu: f64, var: f64) -> Self {
if var == f64::INFINITY {
Self { pi: 0.0, tau: 0.0 }
} else if var == 0.0 {
// Point mass at mu; see `from_ms` for the tau convention.
Self {
pi: f64::INFINITY,
tau: if mu == 0.0 { 0.0 } else { f64::INFINITY },
}
} else {
let pi = 1.0 / var;
Self { pi, tau: mu * pi }
}
}
/// Construct directly from natural parameters.
#[inline]
pub(crate) const fn from_natural(pi: f64, tau: f64) -> Self {
@@ -42,16 +65,19 @@ impl Gaussian {
}
#[inline]
#[must_use]
pub fn pi(&self) -> f64 {
self.pi
}
#[inline]
#[must_use]
pub fn tau(&self) -> f64 {
self.tau
}
#[inline]
#[must_use]
pub fn mu(&self) -> f64 {
// A non-positive precision is an improper (uninformative) Gaussian — its mean is
// undefined. Treat it like `pi == 0` and return 0. EP message cancellation can land
@@ -64,7 +90,23 @@ impl Gaussian {
}
}
/// Variance, `1 / pi`, without the root-and-square of `sigma().powi(2)`.
///
/// Mirrors `sigma()`'s treatment of the improper (`pi <= 0`) and point-mass
/// (`pi == inf`) cases.
#[inline]
pub(crate) fn variance(&self) -> f64 {
if self.pi <= 0.0 {
f64::INFINITY
} else if self.pi.is_infinite() {
0.0
} else {
1.0 / self.pi
}
}
#[inline]
#[must_use]
pub fn sigma(&self) -> f64 {
// A non-positive precision is improper → infinite standard deviation. Guarding
// `pi <= 0.0` (not just `== 0.0`) keeps `1.0 / pi.sqrt()` from returning NaN when EP
@@ -86,22 +128,60 @@ impl Gaussian {
}
pub(crate) fn exclude(&self, other: Gaussian) -> Self {
let var = self.sigma().powi(2) - other.sigma().powi(2);
let var = self.variance() - other.variance();
if var <= 0.0 {
// When sigma_self ≈ sigma_other (including ULP-level rounding differences
// from the pi→sigma accessor round-trip), the excluded contribution is N00.
// Computing from_ms(tiny_mu, 0.0) would give {pi:inf, tau:inf}, whose
// mu() = inf/inf = NaN. Returning N00 is correct: when both Gaussians
// carry the same variance, the residual is a point mass at 0.
return Gaussian::from_ms(0.0, 0.0);
return Gaussian::from_mv(0.0, 0.0);
}
let mu = self.mu() - other.mu();
Self::from_ms(mu, var.sqrt())
Self::from_mv(self.mu() - other.mu(), var)
}
pub(crate) fn forget(&self, variance_delta: f64) -> Self {
let var = self.sigma().powi(2) + variance_delta;
Self::from_ms(self.mu(), var.sqrt())
Self::from_mv(self.mu(), self.variance() + variance_delta)
}
/// `P(X < x)` under this Gaussian.
///
/// The question a stopping rule asks: *how sure am I that this competitor's
/// true skill is below the cutoff?* Expressing that as a probability keeps
/// its meaning as sigma changes, where a `mu + z * sigma` band silently
/// means different confidence at different uncertainties — which is exactly
/// the regime a stopping rule operates in.
///
/// Accurate in the *lower* tail. For the upper tail use
/// [`Gaussian::probability_above`] rather than `1.0 - probability_below(x)`,
/// which cancels away every significant digit once the result is small.
///
/// An improper Gaussian (non-positive precision) has no defined mean, so
/// this returns `0.5` — the same convention `mu()` and `sigma()` follow.
#[must_use]
pub fn probability_below(&self, x: f64) -> f64 {
if self.pi <= 0.0 {
return 0.5;
}
crate::cdf(x, self.mu(), self.sigma())
}
/// `P(X > x)` under this Gaussian.
///
/// Computed as a survival function rather than `1 - cdf`, so it keeps full
/// relative precision in the upper tail: `1 - cdf` returns exactly zero
/// past about 8.3 sigma, where the true value is still 1e-19 and perfectly
/// representable. A stopping rule is evaluated precisely there — the
/// interesting cases are the ones near certainty.
///
/// An improper Gaussian returns `0.5`, as [`Gaussian::probability_below`].
#[must_use]
pub fn probability_above(&self, x: f64) -> f64 {
if self.pi <= 0.0 {
return 0.5;
}
crate::sf(x, self.mu(), self.sigma())
}
/// EP damping in natural-parameter space: `α·new + (1−α)·self`.
@@ -109,6 +189,7 @@ impl Gaussian {
/// Used by within-game inference to stabilise oscillating fixed-point
/// loops on hard graphs. `alpha = 1.0` returns `new` exactly;
/// `alpha < 1.0` shrinks each per-step update.
#[must_use]
pub fn damp_natural(self, new: Gaussian, alpha: f64) -> Gaussian {
Gaussian::from_natural(
alpha * new.pi() + (1.0 - alpha) * self.pi(),
@@ -128,9 +209,7 @@ impl ops::Add<Gaussian> for Gaussian {
/// Variance addition: (mu1 + mu2, sqrt(σ1² + σ2²)).
/// Used for combining performance and noise; rare relative to mul/div.
fn add(self, rhs: Gaussian) -> Self::Output {
let mu = self.mu() + rhs.mu();
let var = self.sigma().powi(2) + rhs.sigma().powi(2);
Self::from_ms(mu, var.sqrt())
Self::from_mv(self.mu() + rhs.mu(), self.variance() + rhs.variance())
}
}
@@ -138,9 +217,7 @@ impl ops::Sub<Gaussian> for Gaussian {
type Output = Gaussian;
/// (mu1 - mu2, sqrt(σ1² + σ2²)). Same sigma combination as Add.
fn sub(self, rhs: Gaussian) -> Self::Output {
let mu = self.mu() - rhs.mu();
let var = self.sigma().powi(2) + rhs.sigma().powi(2);
Self::from_ms(mu, var.sqrt())
Self::from_mv(self.mu() - rhs.mu(), self.variance() + rhs.variance())
}
}
@@ -161,7 +238,7 @@ impl ops::Mul<f64> for Gaussian {
if scalar == 0.0 {
// Scaling by 0 collapses to a point mass at 0 (sigma' = 0, mu' = 0).
// This is N00, the additive identity, NOT N_INF.
return Gaussian::from_ms(0.0, 0.0);
return Gaussian::from_mv(0.0, 0.0);
}
// sigma' = sigma * |scalar| => pi' = pi / scalar²
// mu' = mu * scalar => tau' = tau / scalar
@@ -302,3 +379,68 @@ mod tests {
assert!((damped.tau() - expected_tau).abs() < 1e-12);
}
}
#[cfg(test)]
mod tail_probability_tests {
use super::*;
#[test]
fn probability_below_matches_published_quantiles() {
let g = Gaussian::from_ms(0.0, 1.0);
for (x, expected) in [
(-1.959_963_984_540_054, 0.025),
(0.0, 0.5),
(1.281_551_565_544_6, 0.9),
(1.959_963_984_540_054, 0.975),
] {
let got = g.probability_below(x);
assert!(
(got - expected).abs() < 1e-12,
"P(X < {x}) = {got}, expected {expected}"
);
}
}
#[test]
fn the_two_tails_partition_the_mass() {
let g = Gaussian::from_ms(3.0, 2.0);
for x in [-4.0f64, 0.0, 3.0, 7.5] {
let total = g.probability_below(x) + g.probability_above(x);
assert!((total - 1.0).abs() < 1e-15, "at {x}: {total}");
}
}
/// The reason `probability_above` exists rather than `1 - probability_below`.
#[test]
fn probability_above_keeps_precision_where_the_complement_collapses() {
let g = Gaussian::from_ms(0.0, 1.0);
for (x, expected) in [(9.0f64, 1.128_588e-19), (20.0, 2.753_624e-89)] {
let got = g.probability_above(x);
assert!(
(got - expected).abs() / expected < 1e-6,
"P(X > {x}) = {got}, expected ~{expected}"
);
assert_eq!(
1.0 - g.probability_below(x),
0.0,
"the complement should still collapse at {x}"
);
}
}
#[test]
fn a_scaled_gaussian_shifts_and_stretches() {
let g = Gaussian::from_ms(25.0, 6.0);
assert!((g.probability_below(25.0) - 0.5).abs() < 1e-15);
// One sigma either side of the mean.
assert!((g.probability_below(31.0) - 0.841_344_746_068_543).abs() < 1e-12);
assert!((g.probability_above(19.0) - 0.841_344_746_068_543).abs() < 1e-12);
}
#[test]
fn an_improper_gaussian_is_uninformative_rather_than_nan() {
let improper = Gaussian::from_ms(0.0, f64::INFINITY);
assert_eq!(improper.probability_below(5.0), 0.5);
assert_eq!(improper.probability_above(5.0), 0.5);
}
}
+1229 -138
View File
File diff suppressed because it is too large Load Diff
+99
View File
@@ -0,0 +1,99 @@
//! Posterior of a linear combination of competitors.
//!
//! Every accessor on `History` returns a per-competitor marginal, and almost
//! nothing a consumer publishes is one competitor: "can we tell these two
//! apart" is a difference, "what was this round worth" is a sum. Combining
//! marginals means assuming the competitors are independent, and they are
//! correlated through every event they share — which is the mechanism the model
//! exists to exploit.
//!
//! Measured on a five-competitor round robin, the exact correlation is +0.857,
//! so `sqrt(sa^2 + sb^2)` overstates the width of a difference by 2.6x.
/// Solve `A z = b` for a symmetric positive-definite `A`, by Cholesky.
///
/// `a` is row-major and is consumed as scratch.
///
/// Returns `None` if the matrix is not positive-definite, which for a precision
/// matrix means the model is improper — a competitor with no prior and no
/// evidence.
pub(crate) fn solve_spd(mut a: Vec<f64>, b: &[f64]) -> Option<Vec<f64>> {
let n = b.len();
debug_assert_eq!(a.len(), n * n);
// In-place Cholesky: A = L L^T, lower triangle.
for j in 0..n {
let mut d = a[j * n + j];
for k in 0..j {
d -= a[j * n + k] * a[j * n + k];
}
// Explicit rather than `!(d > 0.0)`: a NaN pivot must fail here too,
// and a negated comparison would let it through as "not positive".
if d.is_nan() || d <= 0.0 {
return None;
}
let d = d.sqrt();
a[j * n + j] = d;
for i in j + 1..n {
let mut s = a[i * n + j];
for k in 0..j {
s -= a[i * n + k] * a[j * n + k];
}
a[i * n + j] = s / d;
}
}
// Forward substitution, then back substitution.
let mut z = b.to_vec();
for i in 0..n {
let mut s = z[i];
for k in 0..i {
s -= a[i * n + k] * z[k];
}
z[i] = s / a[i * n + i];
}
for i in (0..n).rev() {
let mut s = z[i];
for k in i + 1..n {
s -= a[k * n + i] * z[k];
}
z[i] = s / a[i * n + i];
}
Some(z)
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn solves_a_known_system() {
// [[4, 1], [1, 3]] z = [1, 2] => z = [1/11, 7/11]
let a = vec![4.0, 1.0, 1.0, 3.0];
let z = solve_spd(a, &[1.0, 2.0]).unwrap();
assert!((z[0] - 1.0 / 11.0).abs() < 1e-12, "{z:?}");
assert!((z[1] - 7.0 / 11.0).abs() < 1e-12, "{z:?}");
}
#[test]
fn recovers_the_inverse_diagonal() {
// A = [[2, -1, 0], [-1, 2, -1], [0, -1, 2]]; inverse diagonal is
// [0.75, 1.0, 0.75].
let a = vec![2.0, -1.0, 0.0, -1.0, 2.0, -1.0, 0.0, -1.0, 2.0];
for (i, expected) in [0.75, 1.0, 0.75].into_iter().enumerate() {
let mut e = vec![0.0; 3];
e[i] = 1.0;
let z = solve_spd(a.clone(), &e).unwrap();
assert!((z[i] - expected).abs() < 1e-12, "row {i}: {z:?}");
}
}
#[test]
fn rejects_a_non_positive_definite_matrix() {
// Singular: the second row is a multiple of the first.
let a = vec![1.0, 2.0, 2.0, 4.0];
assert!(solve_spd(a, &[1.0, 1.0]).is_none());
}
}
+28 -15
View File
@@ -12,59 +12,72 @@ use crate::Index;
/// crate. Power users can promote `&K` to `Index` via `get_or_create` and
/// skip the lookup on subsequent hot-path calls.
#[derive(Debug)]
pub struct KeyTable<K>(HashMap<K, Index>);
pub struct KeyTable<K> {
forward: HashMap<K, Index>,
/// Reverse mapping, indexed by `Index.0`.
///
/// Indices are handed out densely and sequentially, so position *is* the
/// index and `key()` is a lookup rather than a scan over every entry.
reverse: Vec<K>,
}
impl<K> KeyTable<K>
where
K: Eq + Hash,
K: Eq + Hash + Clone,
{
#[must_use]
pub fn new() -> Self {
Self(HashMap::new())
Self {
forward: HashMap::new(),
reverse: Vec::new(),
}
}
pub fn get<Q: ?Sized + Hash + Eq>(&self, k: &Q) -> Option<Index>
where
K: Borrow<Q>,
{
self.0.get(k).cloned()
self.forward.get(k).cloned()
}
pub fn get_or_create<Q: ?Sized + Hash + Eq + ToOwned<Owned = K>>(&mut self, k: &Q) -> Index
where
K: Borrow<Q>,
{
if let Some(idx) = self.0.get(k) {
if let Some(idx) = self.forward.get(k) {
*idx
} else {
let idx = Index::from(self.0.len());
self.0.insert(k.to_owned(), idx);
let idx = Index::from(self.reverse.len());
let owned = k.to_owned();
self.reverse.push(owned.clone());
self.forward.insert(owned, idx);
idx
}
}
#[must_use]
pub fn key(&self, idx: Index) -> Option<&K> {
self.0
.iter()
.find(|&(_, value)| *value == idx)
.map(|(key, _)| key)
self.reverse.get(idx.0)
}
pub fn keys(&self) -> impl Iterator<Item = &K> {
self.0.keys()
self.forward.keys()
}
#[must_use]
pub fn len(&self) -> usize {
self.0.len()
self.reverse.len()
}
#[must_use]
pub fn is_empty(&self) -> bool {
self.0.is_empty()
self.reverse.is_empty()
}
}
impl<K> Default for KeyTable<K>
where
K: Eq + Hash,
K: Eq + Hash + Clone,
{
fn default() -> Self {
KeyTable::new()
+790 -40
View File
@@ -1,3 +1,104 @@
//! `TrueSkill` Through Time — Bayesian skill rating over a time axis.
//!
//! Where plain `TrueSkill` gives each competitor one running estimate, `TrueSkill`
//! Through Time treats a whole history as a single model and infers skill *at
//! every point in time*. Evidence flows both directions: a result today
//! sharpens the estimate of who someone was last year, so early estimates stop
//! being frozen guesses and comparisons across eras become meaningful.
//!
//! This is a Rust port of
//! [TrueSkillThroughTime.py](https://github.com/glandfried/TrueSkillThroughTime.py).
//!
//! # Getting started
//!
//! Record results, converge, then read off skills:
//!
//! ```
//! use trueskill_tt::History;
//!
//! let mut history = History::default();
//!
//! history.record_winner(&"alice", &"bob", 1)?;
//! history.record_winner(&"bob", &"carol", 2)?;
//! history.record_winner(&"alice", &"carol", 3)?;
//!
//! let report = history.converge()?;
//! assert!(report.converged);
//!
//! let alice = history.current_skill("alice").unwrap();
//! assert!(alice.mu() > 0.0, "alice won every game she played");
//! # Ok::<(), trueskill_tt::InferenceError>(())
//! ```
//!
//! Teams, weights, explicit rankings and continuous scores go through the
//! fluent event builder:
//!
//! ```
//! use trueskill_tt::History;
//!
//! let mut history = History::builder().p_draw(0.1).build();
//!
//! history
//! .event(1)
//! .team(["alice", "bob"])
//! .team(["carol", "dave"])
//! .ranking([0, 1])
//! .commit()?;
//!
//! history.converge()?;
//! # Ok::<(), trueskill_tt::InferenceError>(())
//! ```
//!
//! # Draws need a draw probability
//!
//! A `p_draw` of zero asserts that draws cannot happen, so a tied result has
//! no representable likelihood and is rejected:
//!
//! ```
//! use trueskill_tt::{History, InferenceError};
//!
//! let mut history = History::default(); // p_draw defaults to 0.0
//! let err = history.record_draw(&"alice", &"bob", 1).unwrap_err();
//! assert!(matches!(err, InferenceError::TieWithoutDrawProbability { .. }));
//! ```
//!
//! This also applies to [`Outcome::winner`] for three or more teams, which
//! ties every loser. Configure a positive `p_draw` for those.
//!
//! # Core types
//!
//! - [`History`] — the top-level container: ingests events, runs
//! forward/backward message passing, and answers queries.
//! - [`Gaussian`] — the probability type, stored in natural parameters
//! (`pi = 1/sigma²`, `tau = mu/sigma²`) so message passing is add/subtract.
//! - [`Game`] — one match in isolation, for scoring a hypothetical without a
//! history.
//! - [`Outcome`] — how a match ended: ranks, or continuous scores.
//! - [`Rating`] — a competitor's static configuration (prior, `beta`, drift).
//!
//! # Feature flags
//!
//! - `approx` — implements [`approx`](https://docs.rs/approx) equality traits
//! for [`Gaussian`]. Useful in tests.
//! - `rayon` — parallelises the within-slice sweep and the per-slice passes of
//! `learning_curves`/`log_evidence`. Opt-in; results stay bit-identical
//! regardless of worker count.
#![forbid(unsafe_code)]
/// Compiles every `rust` block in `README.md` as a doctest.
///
/// The README is not the crate's front page — the module docs above are — so it
/// is pulled in here rather than via a crate-level `#![doc = ...]`, purely so
/// its examples are type-checked. Without this nothing compiled them, and they
/// had drifted far enough that four blocks no longer built (#35). `cfg(doctest)`
/// means this type exists only while collecting doctests.
///
/// Blocks that are illustrative rather than runnable are fenced as `text`.
#[cfg(doctest)]
#[doc = include_str!("../README.md")]
pub struct ReadmeDoctests;
use std::{
cmp::Reverse,
f64::consts::{FRAC_1_SQRT_2, FRAC_2_SQRT_PI, SQRT_2},
@@ -9,6 +110,7 @@ pub(crate) mod arena;
mod time;
mod time_slice;
pub use time_slice::{EventKind, TimeSlice};
mod acquisition;
mod color_group;
mod competitor;
mod convergence;
@@ -17,33 +119,35 @@ mod error;
mod event;
mod event_builder;
pub(crate) mod factor;
pub mod factors;
mod game;
pub mod gaussian;
mod history;
mod joint;
mod key_table;
mod matrix;
mod observer;
mod outcome;
mod predict;
pub(crate) mod quadrature;
mod rating;
pub(crate) mod schedule;
pub mod storage;
pub use acquisition::expected_information_gain;
pub use competitor::Competitor;
pub use convergence::{ConvergenceOptions, ConvergenceReport};
pub use drift::{ConstantDrift, Drift};
pub use error::InferenceError;
pub use error::{InferenceError, UnknownKeys};
pub use event::{Event, Member, Team};
pub use event_builder::EventBuilder;
pub use game::{Game, GameOptions, OwnedGame};
pub use gaussian::Gaussian;
pub use history::History;
pub use history::{History, HistoryBuilder};
pub use key_table::KeyTable;
use matrix::Matrix;
pub use observer::{NullObserver, Observer};
pub use outcome::Outcome;
pub use predict::Prediction;
pub use rating::Rating;
pub use schedule::ScheduleReport;
pub use time::{Time, Untimed};
pub const BETA: f64 = 1.0;
@@ -52,9 +156,48 @@ pub const SIGMA: f64 = BETA * 6.0;
pub const GAMMA: f64 = BETA * 0.03;
pub const P_DRAW: f64 = 0.0;
pub const EPSILON: f64 = 1e-6;
/// Default cap on convergence sweeps.
///
/// **This is a floor, not a recommendation.** It is adequate for small
/// histories and is quickly outgrown: a history of 400 events over 100
/// competitors already stops here with a final step of ~7e-3 against the 1e-6
/// default tolerance — four orders of magnitude short — and a dense joint model
/// of ~2,000 nodes over ~3,300 events has been measured needing 76 to 161.
///
/// Overrunning it is not an error, and deliberately so: `converge` returns a
/// [`ConvergenceReport`] whose `converged` flag says what happened. But a fit
/// that stopped short is *wrong by a little*, which is the worst available
/// failure — every rating is finite and ordered sensibly, and nothing in the
/// numbers themselves says they were still moving. Read the report; the type is
/// `#[must_use]` for that reason.
///
/// Raise it via [`ConvergenceOptions`]. Convergence cost is roughly linear in
/// the cap, and for anything but a toy the extra sweeps are milliseconds.
pub const ITERATIONS: usize = 30;
/// Largest team count `History::predict_outcome` will enumerate.
///
/// The outcome space holds `n! * 2^(n-1)` events, so it grows factorially:
/// 1_920 at five teams, 23_040 at six, 322_560 at seven. Six is where
/// enumerating on a caller's behalf stops being reasonable.
pub const MAX_PREDICTED_TEAMS: usize = predict::MAX_TEAMS_FOR_DISTRIBUTION;
const SQRT_TAU: f64 = 2.5066282746310002;
/// `1 / sqrt(pi)`, the leading factor of the `erfcx` continued fraction.
const FRAC_1_SQRT_PI: f64 = 0.564_189_583_547_756_3;
/// `sqrt(2 / pi)`, the numerator of the inverse Mills ratio in scaled form.
const SQRT_2_OVER_PI: f64 = 0.797_884_560_802_865_4;
/// How many window widths into the tail before a tie window is treated as a
/// half-line. Beyond this the truncated mass is concentrated within `1/alpha`
/// of the near edge, so the far edge contributes nothing measurable.
const HALF_LINE_WINDOW: f64 = 10.0;
/// Where `v - alpha` switches from subtraction to its asymptotic series.
///
/// The subtraction loses roughly `eps * alpha^2` of relative precision, and the
/// four-term series is good to ~1e-10 by here, so the two are at their closest
/// agreement around this point. Below it the subtraction is exact; above it the
/// series is.
const ASYMPTOTIC_MILLS_ALPHA: f64 = 100.0;
pub const N01: Gaussian = Gaussian::from_ms(0.0, 1.0);
pub const N00: Gaussian = Gaussian::from_ms(0.0, 0.0);
@@ -63,30 +206,70 @@ pub const N_INF: Gaussian = Gaussian::from_ms(0.0, f64::INFINITY);
#[derive(Copy, Clone, Default, PartialEq, PartialOrd, Eq, Ord, Hash, Debug)]
pub struct Index(usize);
impl Index {
/// The underlying slot number.
///
/// Indices are dense and assigned in interning order, so this is usable as
/// a key into a caller-side side table.
#[must_use]
pub fn get(self) -> usize {
self.0
}
}
impl From<usize> for Index {
fn from(ix: usize) -> Self {
Self(ix)
}
}
fn erfc(x: f64) -> f64 {
let z = x.abs();
let t = 1.0 / (1.0 + z / 2.0);
let a = -0.82215223 + t * 0.17087277;
let b = 1.48851587 + t * a;
let c = -1.13520398 + t * b;
let d = 0.27886807 + t * c;
let e = -0.18628806 + t * d;
let f = 0.09678418 + t * e;
let g = 0.37409196 + t * f;
let h = 1.00002368 + t * g;
let r = t * (-z * z - 1.26551223 + t * h).exp();
if x >= 0.0 { r } else { 2.0 - r }
impl From<Index> for usize {
fn from(idx: Index) -> Self {
idx.0
}
}
/// Complementary error function.
///
/// # Why every transcendental in this crate goes through `libm`
///
/// IEEE 754 specifies the basic operations and `sqrt` exactly, but says nothing
/// about `exp`, `log` or `erf`. `std`'s versions delegate to the *system* math
/// library, so they differ between platforms: measured here, `f64::exp` and
/// `libm::exp` disagree on 9.7% of inputs and `f64::ln` / `libm::log` on 5.0%,
/// each by one ULP.
///
/// Inference is an iterative fixed point, so a one-ULP difference can change an
/// iteration count and therefore the answer by more than one ULP. Routing every
/// transcendental through `libm` makes a fit reproducible across platforms, not
/// just across thread counts as `tests/determinism.rs` already checks.
///
/// **So: use `libm::exp` / `libm::log` in inference code, never `f64::exp` /
/// `f64::ln`.** `sqrt` is exempt — IEEE specifies it exactly, so `f64::sqrt` is
/// already portable. Test code may use whichever is clearer.
///
/// It costs nothing: `Batch::iteration` measured -2.7% [-5.7%, -0.3%] with the
/// whole set swapped.
///
/// Delegates to `libm`, which is the Rust port of FDLIBM and accurate to about
/// one ULP. This replaced a Numerical Recipes `erfcc` rational approximation
/// whose documented bound was 1.2e-7 *relative* — measured at ~1e-7 across the
/// whole range, and the binding accuracy constraint on the entire crate.
///
/// The swap is free. 98% of the arguments inference passes here have
/// `|x| < 0.84375`, which is exactly where FDLIBM skips the exponential
/// entirely, so the longer polynomial costs nothing on the distribution that
/// actually occurs: `Batch::iteration` moved -1.6% [-4.7%, +0.9%], p = 0.31.
///
/// What it bought: `compute_margin` went from 8.4e-8 to 1.7e-16 against exact
/// quantiles, `cdf(mu, mu, sigma)` is now exactly 0.5, and `sf + cdf` sums to
/// one within a single ULP where it was 3e-8 out.
fn erfc(x: f64) -> f64 {
libm::erfc(x)
}
/// The previous Numerical Recipes `erfcc`, kept only so the timing test can
/// compare both in one binary. Removed once the comparison is recorded.
fn erfc_inv(mut y: f64) -> f64 {
if y >= 2.0 {
return f64::NEG_INFINITY;
@@ -102,14 +285,22 @@ fn erfc_inv(mut y: f64) -> f64 {
y = 2.0 - y;
}
let t = (-2.0 * (y / 2.0).ln()).sqrt();
let t = libm::sqrt(-2.0 * libm::log(y / 2.0));
let mut x = FRAC_1_SQRT_2 * ((2.30753 + t * 0.27061) / (1.0 + t * (0.99229 + t * 0.04481)) - t);
// The leading coefficient is NEGATIVE. `rational - t` is negative here, so
// a positive coefficient mirrors the starting point to `-x0` — the
// reflection of the root. Newton then has to cross the origin to get back,
// which a fixed iteration count does not manage: measured against the true
// value, `erfc_inv(0.1)` returned 1.044 instead of 1.16309, and the error
// grew as y shrank until `compute_margin` stopped being monotone in
// `p_draw` altogether.
let mut x =
-FRAC_1_SQRT_2 * ((2.30753 + t * 0.27061) / (1.0 + t * (0.99229 + t * 0.04481)) - t);
for _ in 0..3 {
let err = erfc(x) - y;
x += err / (FRAC_2_SQRT_PI * (-(x.powi(2))).exp() - x * err)
x += err / (FRAC_2_SQRT_PI * libm::exp(-(x * x)) - x * err)
}
if y < 1.0 { x } else { -x }
@@ -129,32 +320,216 @@ pub(crate) fn cdf(x: f64, mu: f64, sigma: f64) -> f64 {
0.5 * erfc(z)
}
/// `P(X > x)` for `X ~ N(mu, sigma^2)`.
///
/// The survival function, computed directly rather than as `1 - cdf(..)`.
///
/// The two are algebraically identical and numerically are not. `cdf` returns
/// a value approaching 1 for an upper tail, so subtracting it from 1 cancels
/// away every significant digit the tail had: measured against this function,
/// `1 - cdf` carries 7% error by four sigma past the mean and returns exactly
/// zero beyond about 8.3 sigma — where the true value is still 1e-19 and
/// perfectly representable. `erfc` holds *relative* accuracy all the way down
/// to 1e-296, so the precision is there to keep; only the subtraction threw it
/// away.
///
/// This matters most where evidence is smallest, which is exactly where an
/// upset makes it interesting: `ln` of a clamped zero is -708 regardless of
/// whether the truth was -43 or -600.
pub(crate) fn sf(x: f64, mu: f64, sigma: f64) -> f64 {
0.5 * erfc((x - mu) / (sigma * SQRT_2))
}
/// `e^(x^2) * erfc(x)`, the scaled complementary error function, for `x >= 0`.
///
/// Exists so the exponential factor common to a Gaussian density and its tail
/// integral can be cancelled *analytically* instead of being computed twice
/// and divided. Both underflow to zero past about 26 sigma, and their ratio is
/// then `0/0` — finite in the limit, `NaN` in floating point.
fn erfcx(x: f64) -> f64 {
if x < 2.0 {
// Below the crossover neither factor is extreme: erfc is O(1) and
// exp(x^2) is at most e^4, so the direct product is exact enough and
// cheaper than the continued fraction.
libm::exp(x * x) * erfc(x)
} else {
// erfcx(x) = 1/sqrt(pi) * 1/(x + (1/2)/(x + 1/(x + (3/2)/(x + ...)))),
// evaluated by backward recurrence. Converges quickly for x >= 2 and,
// unlike the product form, never touches an exponential.
let mut f = 0.0;
for n in (1..=60u32).rev() {
f = (f64::from(n) * 0.5) / (x + f);
}
FRAC_1_SQRT_PI / (x + f)
}
}
/// `ln` of the normal density at `x`.
///
/// The density itself underflows to zero past about 38 sigma, and `ln` of a
/// clamped zero is -708 whatever the truth was. The log form is a polynomial:
/// it stays exact at any separation, and the values it produces (-5001 nats at
/// 100 sigma, -500001 at 1000) are perfectly representable.
pub(crate) fn ln_pdf(x: f64, mu: f64, sigma: f64) -> f64 {
let z = (x - mu) / sigma;
-libm::log(SQRT_TAU * sigma) - 0.5 * z * z
}
/// `ln P(X > x)` for `X ~ N(mu, sigma^2)`.
///
/// In the upper tail the `exp(-z^2 / 2)` common to the tail integral is
/// factored out analytically via `erfcx`, so this never underflows — where
/// `sf(..).ln()` bottoms out at -708 once `erfc` itself reaches zero.
pub(crate) fn ln_sf(x: f64, mu: f64, sigma: f64) -> f64 {
let z = (x - mu) / sigma;
if z > 0.0 {
// ln(0.5 * erfc(z/sqrt2)) with erfc(y) = exp(-y^2) * erfcx(y).
-std::f64::consts::LN_2 - 0.5 * z * z + libm::log(erfcx(z / SQRT_2))
} else {
// The mass here is at least a half; nothing to lose.
libm::log(sf(x, mu, sigma))
}
}
/// `ln P(lo < X < hi)` for `X ~ N(mu, sigma^2)`.
///
/// When the interval sits in a tail both endpoint probabilities underflow
/// together, so their difference is taken in scaled form with the shared
/// exponential factored out. When it straddles the mean nothing is small and
/// the direct difference is exact.
pub(crate) fn ln_interval(lo: f64, hi: f64, mu: f64, sigma: f64) -> f64 {
let z_lo = (lo - mu) / sigma;
let z_hi = (hi - mu) / sigma;
if z_hi <= z_lo {
return f64::NEG_INFINITY;
}
// Fold a lower-tail interval onto the upper tail; the normal is symmetric.
let (near, far) = if z_lo >= 0.0 {
(z_lo, z_hi)
} else if z_hi <= 0.0 {
(-z_hi, -z_lo)
} else {
// Straddles the mean: the interval holds a non-negligible share of the
// mass, so neither endpoint is near enough to 1 to cancel.
return libm::log((cdf(hi, mu, sigma) - cdf(lo, mu, sigma)).max(f64::MIN_POSITIVE));
};
let (a, b) = (near / SQRT_2, far / SQRT_2);
// b > a >= 0, so this ratio of exponentials is at most 1 and cannot overflow.
let scale = libm::exp(a * a - b * b);
let bracket = erfcx(a) - scale * erfcx(b);
if bracket <= 0.0 {
return f64::NEG_INFINITY;
}
-std::f64::consts::LN_2 - a * a + libm::log(bracket)
}
fn pdf(x: f64, mu: f64, sigma: f64) -> f64 {
let normalizer = (SQRT_TAU * sigma).powi(-1);
let functional = (-((x - mu).powi(2)) / (2.0 * sigma.powi(2))).exp();
let functional = libm::exp(-((x - mu) * (x - mu)) / (2.0 * sigma * sigma));
normalizer * functional
}
/// Truncated-Gaussian correction terms `(v, w)`.
///
/// `v` shifts the mean and `w` shrinks the variance. Both are ratios whose
/// numerator and denominator underflow together in the tails, so both are
/// computed in scaled form there: the shared `exp(-alpha^2 / 2)` is cancelled
/// analytically rather than evaluated and divided out. Without that, a
/// truncation point beyond about 39 sigma produced `0 / 0` and put `NaN`
/// straight into the posterior.
/// Truncation terms for a boundary `alpha` standard deviations into the upper
/// tail, from the asymptotic expansion of the inverse Mills ratio.
///
/// `v` tends to `alpha` out here, so the gap between them cannot be obtained by
/// subtracting one from the other — the series computes the gap directly, and
/// `w = v * gap` then never forms the difference of two large near-equal
/// numbers. A far-tail *window* behaves like a half-line once it is more than a
/// few multiples of its own width from the mean, so the tie branch shares this.
fn half_line_truncation(alpha: f64) -> (f64, f64) {
let inv = alpha.recip();
let inv_sq = inv * inv;
let gap = inv * (1.0 - inv_sq * (2.0 - inv_sq * (10.0 - 74.0 * inv_sq)));
let v = alpha + gap;
(v, v * gap)
}
fn v_w(mu: f64, sigma: f64, margin: f64, tie: bool) -> (f64, f64) {
if !tie {
let alpha = (margin - mu) / sigma;
let v = pdf(-alpha, 0.0, 1.0) / cdf(-alpha, 0.0, 1.0);
let w = v * (v + (-alpha));
// v is the inverse Mills ratio, phi(alpha) / Phi(-alpha), and w needs
// the gap `v - alpha` as well as v itself. Far into the tail v tends to
// alpha, so that gap is a subtraction of two nearly equal numbers and
// loses every digit it has: at alpha = 1e6 it drove w above 1 and made
// `sqrt(1 - w)` NaN. Past the crossover the gap comes from its
// asymptotic series instead, which has no subtraction in it.
if alpha >= ASYMPTOTIC_MILLS_ALPHA {
return half_line_truncation(alpha);
}
(v, w)
let (v, gap) = if alpha > 0.0 {
// Both terms carry exp(-alpha^2 / 2); in scaled form it cancels
// and the result stays exact however far into the tail alpha sits.
let v = SQRT_2_OVER_PI / erfcx(alpha / SQRT_2);
(v, v - alpha)
} else {
// Phi(-alpha) >= 1/2 here, so the direct ratio loses nothing.
let v = pdf(-alpha, 0.0, 1.0) / cdf(-alpha, 0.0, 1.0);
(v, v - alpha)
};
(v, v * gap)
} else {
// v is odd in mu and w is even, so fold to mu <= 0. Both truncation
// points then sit in the upper tail, where the scaled form applies.
let flipped = mu > 0.0;
let mu = if flipped { -mu } else { mu };
let alpha = (-margin - mu) / sigma;
let beta = (margin - mu) / sigma;
let v = (pdf(alpha, 0.0, 1.0) - pdf(beta, 0.0, 1.0))
/ (cdf(beta, 0.0, 1.0) - cdf(alpha, 0.0, 1.0));
let u = (alpha * pdf(alpha, 0.0, 1.0) - beta * pdf(beta, 0.0, 1.0))
/ (cdf(beta, 0.0, 1.0) - cdf(alpha, 0.0, 1.0));
// `w` comes out of `v * v - u`, and both terms grow as alpha^2 while
// their difference stays O(1) — at alpha = 1e9 that subtraction had no
// digits left and returned w = -128, making `sqrt(1 - w)` nonsense.
// Once the window sits many of its own widths into the tail it is
// indistinguishable from a half-line, so the asymptotic covers it with
// no subtraction at all.
if alpha >= ASYMPTOTIC_MILLS_ALPHA && alpha * (beta - alpha) >= HALF_LINE_WINDOW {
let (v, w) = half_line_truncation(alpha);
return (if flipped { -v } else { v }, w);
}
let (v, u) = if alpha > 0.0 {
// beta > alpha > 0, so this ratio of exponentials is at most 1 and
// cannot overflow.
let scale = libm::exp(0.5 * (alpha * alpha - beta * beta));
let denominator = 0.5 * (erfcx(alpha / SQRT_2) - scale * erfcx(beta / SQRT_2));
(
(1.0 - scale) / SQRT_TAU / denominator,
(alpha - beta * scale) / SQRT_TAU / denominator,
)
} else {
// The interval straddles the mean, so nothing here is small.
let denominator = cdf(beta, 0.0, 1.0) - cdf(alpha, 0.0, 1.0);
(
(pdf(alpha, 0.0, 1.0) - pdf(beta, 0.0, 1.0)) / denominator,
(alpha * pdf(alpha, 0.0, 1.0) - beta * pdf(beta, 0.0, 1.0)) / denominator,
)
};
let w = -(u - v.powi(2));
(v, w)
(if flipped { -v } else { v }, w)
}
}
@@ -184,6 +559,56 @@ pub(crate) fn tuple_gt(t: (f64, f64), e: f64) -> bool {
t.0 > e || t.1 > e
}
/// Whether a convergence step is finite in both components.
///
/// A NaN step means EP broke down numerically. Because every comparison
/// against NaN is false, `tuple_gt` reads NaN as "below epsilon" — so
/// convergence checks must test finiteness explicitly rather than inferring
/// success from `!tuple_gt(..)`.
pub(crate) fn step_is_finite(t: (f64, f64)) -> bool {
t.0.is_finite() && t.1.is_finite()
}
/// Whether a step counts as converged: finite *and* within `epsilon`.
pub(crate) fn step_converged(t: (f64, f64), epsilon: f64) -> bool {
step_is_finite(t) && !tuple_gt(t, epsilon)
}
/// Indices of the first pair of teams sharing a rank, if any.
///
/// A tie is only representable when the draw probability is positive: with
/// `p_draw == 0.0` the truncation margin collapses to zero and the two-sided
/// tie update evaluates `0/0`. Callers use this to reject such events before
/// they reach inference.
pub(crate) fn first_tied_pair(ranks: &[u32]) -> Option<(usize, usize)> {
for (i, a) in ranks.iter().enumerate() {
for (j, b) in ranks.iter().enumerate().skip(i + 1) {
if a == b {
return Some((i, j));
}
}
}
None
}
/// As `first_tied_pair`, but over the engine's internal `f64` outputs.
///
/// Ranks reach the engine already converted to descending `f64` outputs, and
/// `Game` decides a tie by exact equality of those values — so this mirrors
/// the comparison inference itself performs.
pub(crate) fn first_tied_output(outputs: &[f64]) -> Option<(usize, usize)> {
for (i, a) in outputs.iter().enumerate() {
for (j, b) in outputs.iter().enumerate().skip(i + 1) {
if a == b {
return Some((i, j));
}
}
}
None
}
pub(crate) fn sort_time<T: Copy + Ord>(xs: &[T], reverse: bool) -> Vec<usize> {
let mut x: Vec<(usize, T)> = xs.iter().enumerate().map(|(i, &t)| (i, t)).collect();
@@ -197,7 +622,27 @@ pub(crate) fn sort_time<T: Copy + Ord>(xs: &[T], reverse: bool) -> Vec<usize> {
}
/// Calculates the match quality of the given rating groups. A result is the draw probability in the association
///
/// Supports any number of groups. Values range roughly `[0, 1]`; 1 means a
/// perfectly balanced match.
///
/// # Panics
///
/// Panics if fewer than two rating groups are supplied, or if any group is
/// empty — match quality is a property of a contest between at least two
/// non-empty sides.
#[must_use]
pub fn quality(rating_groups: &[&[Gaussian]], beta: f64) -> f64 {
assert!(
rating_groups.len() >= 2,
"quality() requires at least 2 rating groups, got {}",
rating_groups.len()
);
assert!(
rating_groups.iter().all(|group| !group.is_empty()),
"quality() requires every rating group to be non-empty"
);
let flatten_ratings = rating_groups
.iter()
.flat_map(|group| group.iter())
@@ -221,8 +666,10 @@ pub fn quality(rating_groups: &[&[Gaussian]], beta: f64) -> f64 {
let mut rotated_a_matrix = Matrix::new(rating_groups.len() - 1, length);
// Row `row` contrasts group `row` (+weight) against group `row + 1`
// (-weight). `t` is the column where the current group's players start;
// the negative block begins immediately after it.
let mut t = 0;
let mut x = 0;
for (row, group) in rating_groups.windows(2).enumerate() {
let current = group[0];
@@ -230,17 +677,13 @@ pub fn quality(rating_groups: &[&[Gaussian]], beta: f64) -> f64 {
for n in t..t + current.len() {
rotated_a_matrix[(row, n)] = flatten_weights[n];
x += 1;
}
t += current.len();
for n in x..x + next.len() {
for n in t..t + next.len() {
rotated_a_matrix[(row, n)] = -flatten_weights[n];
}
x += next.len();
}
let a_matrix = rotated_a_matrix.transpose();
@@ -255,7 +698,7 @@ pub fn quality(rating_groups: &[&[Gaussian]], beta: f64) -> f64 {
let e_arg = (-0.5 * &start * &middle.inverse() * &end).determinant();
let s_arg = ata.determinant() / middle.determinant();
e_arg.exp() * s_arg.sqrt()
libm::exp(e_arg) * s_arg.sqrt()
}
#[cfg(test)]
@@ -269,6 +712,313 @@ mod tests {
assert_eq!(sort_time(&[0i64, 1, 2, 0], true), vec![2, 1, 0, 3]);
}
/// Upper-tail values of the standard normal, from published tables. The
/// point is not the digits — these are 7-digit table values — but that a
/// number comes back at all: `1 - cdf` returned exactly zero for every one
/// of these.
#[test]
fn survival_function_survives_the_far_tail() {
for (z, expected) in [
(9.0f64, 1.128_588e-19),
(12.0, 1.776_482e-33),
(20.0, 2.753_624e-89),
(37.0, 5.725_571e-300),
] {
let got = sf(z, 0.0, 1.0);
assert!(got > 0.0, "sf({z}) collapsed to zero");
assert!(
(got - expected).abs() / expected < 1e-6, // published table values, 7 digits
"sf({z}) = {got}, expected ~{expected}"
);
assert_eq!(
1.0 - cdf(z, 0.0, 1.0),
0.0,
"the naive form should still be zero here"
);
}
}
/// Where no cancellation happens the two forms must agree exactly enough
/// that nothing else in the crate shifts.
#[test]
fn survival_function_matches_the_naive_form_where_that_form_works() {
for z in [-4.0f64, -1.0, 0.0, 0.5, 1.0, 2.0, 3.0, 4.0] {
let naive = 1.0 - cdf(z, 0.0, 1.0);
let direct = sf(z, 0.0, 1.0);
assert!(
(naive - direct).abs() < 1e-15,
"z={z}: naive {naive} vs direct {direct}"
);
}
}
#[test]
fn survival_and_cdf_partition_the_mass() {
for z in [-3.0f64, -0.5, 0.0, 1.0, 2.5] {
let total = sf(z, 1.0, 2.0) + cdf(z, 1.0, 2.0);
assert!((total - 1.0).abs() < 1e-15, "z={z}: {total}");
}
}
/// `erfcx` switches formulation at x = 2; the two sides must meet.
#[test]
fn erfcx_is_continuous_across_its_crossover() {
for x in [1.90f64, 1.99, 1.999, 2.0, 2.001, 2.01, 2.10] {
let direct = (x * x).exp() * erfc(x);
let scaled = erfcx(x);
assert!(
(direct - scaled).abs() / scaled < 1e-14,
"x={x}: direct {direct} vs erfcx {scaled}"
);
}
}
/// The whole reason `erfcx` exists: it stays finite and O(1/x) exactly
/// where `exp(x^2)` overflows and `erfc(x)` underflows.
#[test]
fn erfcx_stays_finite_where_its_factors_do_not() {
for x in [27.0f64, 50.0, 1.0e3, 1.0e8] {
let scaled = erfcx(x);
assert!(scaled.is_finite() && scaled > 0.0, "erfcx({x}) = {scaled}");
// Asymptotically erfcx(x) -> 1 / (x * sqrt(pi)).
let asymptote = 1.0 / (x * std::f64::consts::PI.sqrt());
assert!(
(scaled - asymptote).abs() / asymptote < 1e-2,
"erfcx({x}) = {scaled} strays from its asymptote {asymptote}"
);
assert!(
(x * x).exp().is_infinite(),
"x={x} should overflow the direct form"
);
}
}
/// Truncation must never produce a non-finite posterior. Before the scaled
/// formulation these returned NaN from `0 / 0` past about 39 sigma.
#[test]
fn truncation_stays_finite_arbitrarily_far_into_the_tail() {
for alpha in [0.0f64, 8.0, 38.0, 40.0, 100.0, 1.0e3, 1.0e6, 1.0e9, 1.0e15] {
for tie in [false, true] {
let (v, w) = v_w(-alpha, 1.0, if tie { 1.0 } else { 0.0 }, tie);
assert!(v.is_finite(), "alpha={alpha} tie={tie}: v = {v}");
assert!(w.is_finite(), "alpha={alpha} tie={tie}: w = {w}");
// sigma_trunc = sigma * sqrt(1 - w) must stay real.
assert!(
(0.0..=1.0).contains(&w),
"alpha={alpha} tie={tie}: w = {w} leaves sqrt(1 - w) imaginary"
);
let (mu_t, sigma_t) = trunc(-alpha, 1.0, if tie { 1.0 } else { 0.0 }, tie);
assert!(
mu_t.is_finite() && sigma_t.is_finite(),
"alpha={alpha} tie={tie}: trunc = ({mu_t}, {sigma_t})"
);
}
}
}
/// The Mills gap switches from subtraction to series at alpha = 100. Both
/// are supposed to be right there; if they disagree, the crossover is in
/// the wrong place.
#[test]
fn the_mills_gap_series_meets_the_scaled_form() {
for alpha in [50.0f64, 99.0, 100.0, 101.0, 200.0] {
let scaled = SQRT_2_OVER_PI / erfcx(alpha / SQRT_2) - alpha;
let inv = alpha.recip();
let inv_sq = inv * inv;
let series = inv * (1.0 - inv_sq * (2.0 - inv_sq * (10.0 - 74.0 * inv_sq)));
assert!(
(scaled - series).abs() / series < 1e-9,
"alpha={alpha}: scaled {scaled} vs series {series}"
);
}
}
/// Folding the tie branch to `mu <= 0` is only valid if v is odd in mu and
/// w is even. Assert the symmetry the implementation relies on.
#[test]
fn tie_truncation_is_odd_in_v_and_even_in_w() {
for mu in [0.5f64, 3.0, 20.0, 40.0, 100.0, 1.0e3] {
let (v_pos, w_pos) = v_w(mu, 1.0, 1.0, true);
let (v_neg, w_neg) = v_w(-mu, 1.0, 1.0, true);
assert!(
(v_pos + v_neg).abs() < 1e-9,
"mu={mu}: v should be odd, got {v_pos} and {v_neg}"
);
assert!(
(w_pos - w_neg).abs() < 1e-9,
"mu={mu}: w should be even, got {w_pos} and {w_neg}"
);
}
}
/// `erfc_inv`'s initial guess had the wrong sign, putting Newton on the
/// mirror image of the root. Three fixed iterations could not cross back,
/// so the error grew as the argument shrank: at `p_draw = 0.99` the margin
/// came out 0.503 where the answer is 2.576.
#[test]
fn erfc_inv_matches_known_quantiles() {
// sqrt(2) * erfc_inv(1 - p) is the standard normal quantile
// Phi^-1((1 + p) / 2).
for (p, exact) in [
(0.5f64, 0.674_489_750_196_081_7f64),
(0.9, 1.644_853_626_951_472_7),
(0.95, 1.959_963_984_540_054_2),
(0.99, 2.575_829_303_548_9),
(0.999, 3.290_526_731_491_896_4),
] {
let got = SQRT_2 * erfc_inv(1.0 - p);
assert!(
(got - exact).abs() / exact < 1e-14,
"p={p}: got {got}, exact {exact}"
);
}
}
/// The draw margin must grow with the draw probability. It did not: it ran
/// 0.674 -> 1.476 -> 0.503 -> 0.982 as `p_draw` went 0.5 -> 0.9 -> 0.99 ->
/// 0.999, which is not a rounding error but a broken function.
/// Deep in the tail the accuracy limit is the *caller's* argument, not this
/// function.
///
/// `compute_margin(0.999999, ..)` computes `1.0 - p_draw`, and 0.999999 is
/// not representable: the subtraction cancels and leaves 2.9e-11 of
/// relative error in the argument before `erfc_inv` is even entered. Given
/// an exactly-representable argument the result is good to 1.8e-16, so this
/// is inherent to taking `p_draw` near one rather than something to fix
/// here. At `p_draw = 0.999` the whole path is still accurate to 4e-16.
///
/// Worth pinning: measured against a 70-digit reference, `puruspe`'s
/// `inverfc` returns the identical wrong value for the identical reason,
/// which is what makes it clear the fault is upstream of both.
#[test]
fn erfc_inv_is_exact_given_an_exactly_representable_argument() {
// erfc(z / sqrt2) = 1e-6 exactly, so z = Phi^-1(0.9999995).
let got = SQRT_2 * erfc_inv(1e-6);
let exact = 4.891_638_475_698_59;
assert!(
(got - exact).abs() / exact < 1e-14,
"got {got}, exact {exact}"
);
}
#[test]
fn compute_margin_is_monotone_in_the_draw_probability() {
let mut previous = 0.0;
for p_draw in [
0.001f64, 0.01, 0.1, 0.25, 0.5, 0.75, 0.9, 0.99, 0.999, 0.9999,
] {
let margin = compute_margin(p_draw, 1.0);
assert!(
margin > previous,
"p_draw={p_draw}: margin {margin} did not exceed {previous}"
);
previous = margin;
}
}
/// Round-tripping the margin back through the model's own CDF must recover
/// the draw probability it was built from.
#[test]
fn compute_margin_round_trips_through_the_cdf() {
for p_draw in [0.001f64, 0.1, 0.5, 0.9, 0.99, 0.999] {
for sd in [0.5f64, 1.0, 5.892_557] {
let margin = compute_margin(p_draw, sd);
// P(|X| < margin) for X ~ N(0, sd^2).
let recovered = 1.0 - 2.0 * cdf(-margin, 0.0, sd);
assert!(
(recovered - p_draw).abs() < 1e-14,
"p_draw={p_draw} sd={sd}: recovered {recovered}"
);
}
}
}
/// `ln_pdf`, `ln_sf` and `ln_interval` exist so evidence stays exact where
/// the linear forms underflow. Past ~38 sigma the linear value is zero and
/// its log is whatever floor it was clamped to.
#[test]
fn log_space_helpers_stay_exact_where_the_linear_forms_underflow() {
for z in [40.0f64, 60.0, 100.0, 1000.0] {
assert_eq!(pdf(z, 0.0, 1.0), 0.0, "pdf should underflow at {z}");
assert_eq!(sf(z, 0.0, 1.0), 0.0, "sf should underflow at {z}");
let lp = ln_pdf(z, 0.0, 1.0);
let expected_lp = -(SQRT_TAU).ln() - 0.5 * z * z;
assert!(
(lp - expected_lp).abs() < 1e-9,
"ln_pdf({z}) = {lp}, expected {expected_lp}"
);
let ls = ln_sf(z, 0.0, 1.0);
// ln Phi(-z) ~ -z^2/2 - ln(z) - ln(sqrt(2 pi)) for large z.
let approx = -0.5 * z * z - z.ln() - SQRT_TAU.ln();
assert!(
(ls - approx).abs() / approx.abs() < 1e-3,
"ln_sf({z}) = {ls}, asymptote {approx}"
);
assert!(
ls < f64::MIN_POSITIVE.ln(),
"ln_sf({z}) still on the clamp floor"
);
}
}
/// Where nothing underflows, the log helpers must agree with the direct
/// forms exactly enough that nothing else in the crate shifts.
#[test]
fn log_space_helpers_agree_with_the_linear_forms_in_range() {
for z in [-3.0f64, -1.0, 0.0, 1.0, 2.0, 5.0, 10.0, 20.0] {
let lp = ln_pdf(z, 0.5, 2.0);
let direct_pdf = pdf(z, 0.5, 2.0);
assert!(
(lp.exp() - direct_pdf).abs() <= 1e-12 * direct_pdf,
"ln_pdf at {z}: {} vs {direct_pdf}",
lp.exp()
);
let ls = ln_sf(z, 0.5, 2.0);
let direct = sf(z, 0.5, 2.0);
assert!(
(ls.exp() - direct).abs() <= 1e-13 * direct.max(1e-300),
"ln_sf at {z}: {} vs {direct}",
ls.exp()
);
}
}
#[test]
fn ln_interval_matches_the_direct_difference_when_nothing_is_small() {
for mu in [-2.0f64, 0.0, 0.5, 2.0] {
let direct = cdf(1.0, mu, 1.0) - cdf(-1.0, mu, 1.0);
let logged = ln_interval(-1.0, 1.0, mu, 1.0).exp();
assert!(
(logged - direct).abs() <= 1e-13 * direct,
"mu={mu}: {logged} vs {direct}"
);
}
}
/// A window far out in the tail: both endpoints underflow together, so the
/// difference has to be taken in scaled form.
#[test]
fn ln_interval_survives_a_window_deep_in_the_tail() {
for mu in [-50.0f64, -100.0, -1000.0] {
let logged = ln_interval(-1.0, 1.0, mu, 1.0);
assert!(logged.is_finite(), "mu={mu}: {logged}");
assert!(
logged < f64::MIN_POSITIVE.ln(),
"mu={mu}: {logged} is stuck on the clamp floor"
);
// Dominated by the near edge: ln P ~ ln Phi(-(|mu| - 1)).
let near = ln_sf(-1.0, mu, 1.0);
assert!(
(logged - near).abs() < 5.0,
"mu={mu}: {logged} strays from the near-edge tail {near}"
);
}
}
#[test]
fn test_quality() {
let a = Gaussian::from_ms(25.0, 3.0);
+316 -119
View File
@@ -1,29 +1,13 @@
//! Minimal dense matrix used by `quality()`.
//!
//! `determinant` and `inverse` go through one LU decomposition with partial
//! pivoting — O(n³) and numerically stable. The previous implementation
//! expanded cofactors recursively (O(n!), allocating a `Vec` per minor) and
//! only implemented `inverse` for the 1×1 case, which limited `quality()` to
//! exactly two rating groups.
use std::ops;
fn det(m: &[f64], x: usize) -> f64 {
if x == 1 {
m[0]
} else if x == 2 {
m[0] * m[3] - m[1] * m[2]
} else {
let mut d = 0.0;
for n in 0..x {
let ms = m
.iter()
.enumerate()
.skip(x)
.filter(|(i, _)| (i % x) != n)
.map(|(_, v)| *v)
.collect::<Vec<_>>();
d += (-1.0f64).powi(n as i32) * m[n] * det(&ms, x - 1);
}
d
}
}
#[derive(Clone, Debug)]
pub struct Matrix {
data: Box<[f64]>,
@@ -31,6 +15,107 @@ pub struct Matrix {
width: usize,
}
/// LU decomposition with partial pivoting: `PA = LU`, stored compactly.
///
/// `lu` holds `L` below the diagonal (unit diagonal implied) and `U` on and
/// above it. `sign` is the determinant sign contributed by row swaps, or 0.0
/// when the matrix is singular.
struct Lu {
lu: Vec<f64>,
perm: Vec<usize>,
n: usize,
sign: f64,
}
impl Lu {
fn decompose(m: &Matrix) -> Self {
debug_assert_eq!(m.width, m.height, "LU requires a square matrix");
let n = m.width;
let mut lu = m.data.to_vec();
let mut perm: Vec<usize> = (0..n).collect();
let mut sign = 1.0;
for col in 0..n {
// Partial pivot: take the largest-magnitude candidate to limit
// growth of round-off in the elimination below.
let mut pivot_row = col;
let mut pivot_max = lu[col * n + col].abs();
for row in (col + 1)..n {
let candidate = lu[row * n + col].abs();
if candidate > pivot_max {
pivot_max = candidate;
pivot_row = row;
}
}
if pivot_max == 0.0 {
sign = 0.0;
continue;
}
if pivot_row != col {
for k in 0..n {
lu.swap(col * n + k, pivot_row * n + k);
}
perm.swap(col, pivot_row);
sign = -sign;
}
let pivot = lu[col * n + col];
for row in (col + 1)..n {
let factor = lu[row * n + col] / pivot;
lu[row * n + col] = factor;
for k in (col + 1)..n {
lu[row * n + k] -= factor * lu[col * n + k];
}
}
}
Self { lu, perm, n, sign }
}
fn determinant(&self) -> f64 {
if self.sign == 0.0 {
return 0.0;
}
let mut det = self.sign;
for i in 0..self.n {
det *= self.lu[i * self.n + i];
}
det
}
/// Solve `Ax = b` for a single column of the identity, giving one column
/// of the inverse.
fn solve_column(&self, col: usize, out: &mut [f64]) {
let n = self.n;
// Forward substitution through L, applying the row permutation.
for i in 0..n {
let mut sum = if self.perm[i] == col { 1.0 } else { 0.0 };
for (k, &solved) in out.iter().enumerate().take(i) {
sum -= self.lu[i * n + k] * solved;
}
out[i] = sum;
}
// Back substitution through U.
for i in (0..n).rev() {
let mut sum = out[i];
for (k, &solved) in out.iter().enumerate().skip(i + 1) {
sum -= self.lu[i * n + k] * solved;
}
out[i] = sum / self.lu[i * n + i];
}
}
}
impl Matrix {
pub fn new(height: usize, width: usize) -> Matrix {
Matrix {
@@ -52,73 +137,59 @@ impl Matrix {
matrix
}
pub fn minor(&self, row_n: usize, col_n: usize) -> Matrix {
let mut matrix = Matrix::new(self.height - 1, self.width - 1);
let mut nr = 0;
for r in 0..self.height {
if r == row_n {
continue;
}
let mut nc = 0;
for c in 0..self.width {
if c == col_n {
continue;
}
matrix[(nr, nc)] = self[(r, c)];
nc += 1;
}
nr += 1;
}
matrix
}
/// Determinant of a square matrix. The 0×0 determinant is 1 by convention
/// (the empty product).
///
/// # Panics
///
/// Panics if the matrix is not square.
pub fn determinant(&self) -> f64 {
debug_assert!(self.width == self.height);
assert_eq!(
self.width, self.height,
"determinant requires a square matrix, got {}x{}",
self.height, self.width
);
det(&self.data, self.width)
if self.width == 0 {
return 1.0;
}
pub fn adjugate(&self) -> Matrix {
debug_assert!(self.width == self.height);
let mut matrix = Matrix::new(self.height, self.width);
if matrix.height == 2 {
matrix[(0, 0)] = self[(1, 1)];
matrix[(0, 1)] = -self[(0, 1)];
matrix[(1, 0)] = -self[(1, 0)];
matrix[(1, 1)] = self[(0, 0)];
} else {
for r in 0..matrix.height {
for c in 0..matrix.width {
let sign = if (r + c) % 2 == 0 { 1.0 } else { -1.0 };
matrix[(r, c)] = self.minor(r, c).determinant() * sign;
}
}
}
matrix
Lu::decompose(self).determinant()
}
/// Matrix inverse via LU decomposition.
///
/// # Panics
///
/// Panics if the matrix is not square or is singular.
pub fn inverse(&self) -> Matrix {
let mut matrix = Matrix::new(self.width, self.height);
assert_eq!(
self.width, self.height,
"inverse requires a square matrix, got {}x{}",
self.height, self.width
);
if self.height == self.width && self.height == 1 {
matrix[(0, 0)] = 1.0 / self[(0, 0)];
} else {
panic!("eh, okey")
let n = self.width;
let mut inverse = Matrix::new(n, n);
if n == 0 {
return inverse;
}
matrix
let lu = Lu::decompose(self);
assert!(lu.sign != 0.0, "cannot invert a singular matrix");
let mut column = vec![0.0; n];
for c in 0..n {
lu.solve_column(c, &mut column);
for (r, &value) in column.iter().enumerate() {
inverse[(r, c)] = value;
}
}
inverse
}
}
@@ -126,20 +197,62 @@ impl ops::Index<(usize, usize)> for Matrix {
type Output = f64;
fn index(&self, pos: (usize, usize)) -> &Self::Output {
debug_assert!(
pos.0 < self.height && pos.1 < self.width,
"index ({}, {}) out of bounds for {}x{} matrix",
pos.0,
pos.1,
self.height,
self.width
);
&self.data[(self.width * pos.0) + pos.1]
}
}
impl ops::IndexMut<(usize, usize)> for Matrix {
fn index_mut(&mut self, pos: (usize, usize)) -> &mut Self::Output {
debug_assert!(
pos.0 < self.height && pos.1 < self.width,
"index ({}, {}) out of bounds for {}x{} matrix",
pos.0,
pos.1,
self.height,
self.width
);
&mut self.data[(self.width * pos.0) + pos.1]
}
}
impl<'a> ops::Mul<&'a Matrix> for f64 {
fn multiply(lhs: &Matrix, rhs: &Matrix) -> Matrix {
assert_eq!(
lhs.width, rhs.height,
"cannot multiply {}x{} by {}x{}",
lhs.height, lhs.width, rhs.height, rhs.width
);
let mut matrix = Matrix::new(lhs.height, rhs.width);
for r in 0..matrix.height {
for c in 0..matrix.width {
let mut value = 0.0;
for x in 0..lhs.width {
value += lhs[(r, x)] * rhs[(x, c)];
}
matrix[(r, c)] = value;
}
}
matrix
}
impl ops::Mul<&Matrix> for f64 {
type Output = Matrix;
fn mul(self, rhs: &'a Matrix) -> Matrix {
fn mul(self, rhs: &Matrix) -> Matrix {
let mut matrix = Matrix::new(rhs.height, rhs.width);
for r in 0..rhs.height {
@@ -152,54 +265,35 @@ impl<'a> ops::Mul<&'a Matrix> for f64 {
}
}
impl<'a> ops::Mul<&'a Matrix> for Matrix {
impl ops::Mul<&Matrix> for Matrix {
type Output = Matrix;
fn mul(self, rhs: &'a Matrix) -> Matrix {
let mut matrix = Matrix::new(self.height, rhs.width);
for r in 0..matrix.height {
for c in 0..matrix.width {
let mut value = 0.0;
for x in 0..self.width {
value += self[(r, x)] * rhs[(x, c)];
}
matrix[(r, c)] = value;
}
}
matrix
fn mul(self, rhs: &Matrix) -> Matrix {
multiply(&self, rhs)
}
}
impl<'a> ops::Mul<&'a Matrix> for &'a Matrix {
impl ops::Mul<&Matrix> for &Matrix {
type Output = Matrix;
fn mul(self, rhs: &'a Matrix) -> Matrix {
let mut matrix = Matrix::new(self.height, rhs.width);
for r in 0..matrix.height {
for c in 0..matrix.width {
let mut value = 0.0;
for x in 0..self.width {
value += self[(r, x)] * rhs[(x, c)];
}
matrix[(r, c)] = value;
}
}
matrix
fn mul(self, rhs: &Matrix) -> Matrix {
multiply(self, rhs)
}
}
impl<'a> ops::Add<&'a Matrix> for &'a Matrix {
impl ops::Add<&Matrix> for &Matrix {
type Output = Matrix;
fn add(self, rhs: &'a Matrix) -> Matrix {
fn add(self, rhs: &Matrix) -> Matrix {
assert!(
self.height == rhs.height && self.width == rhs.width,
"cannot add {}x{} to {}x{}",
self.height,
self.width,
rhs.height,
rhs.width
);
let mut matrix = Matrix::new(self.height, self.width);
for r in 0..matrix.height {
@@ -211,3 +305,106 @@ impl<'a> ops::Add<&'a Matrix> for &'a Matrix {
matrix
}
}
#[cfg(test)]
mod tests {
use super::*;
fn from_rows(rows: &[&[f64]]) -> Matrix {
let mut m = Matrix::new(rows.len(), rows[0].len());
for (r, row) in rows.iter().enumerate() {
for (c, &v) in row.iter().enumerate() {
m[(r, c)] = v;
}
}
m
}
#[test]
fn determinant_1x1() {
assert!((from_rows(&[&[3.0]]).determinant() - 3.0).abs() < 1e-12);
}
#[test]
fn determinant_2x2() {
let m = from_rows(&[&[1.0, 2.0], &[3.0, 4.0]]);
assert!((m.determinant() - (-2.0)).abs() < 1e-12);
}
#[test]
fn determinant_3x3() {
let m = from_rows(&[&[6.0, 1.0, 1.0], &[4.0, -2.0, 5.0], &[2.0, 8.0, 7.0]]);
assert!((m.determinant() - (-306.0)).abs() < 1e-10);
}
#[test]
fn determinant_requires_no_pivot_at_origin() {
// A zero in the top-left forces a row swap; the sign must follow.
let m = from_rows(&[&[0.0, 1.0], &[1.0, 0.0]]);
assert!((m.determinant() - (-1.0)).abs() < 1e-12);
}
#[test]
fn determinant_of_singular_is_zero() {
let m = from_rows(&[&[1.0, 2.0], &[2.0, 4.0]]);
assert!(m.determinant().abs() < 1e-12);
}
#[test]
fn inverse_1x1() {
let inv = from_rows(&[&[4.0]]).inverse();
assert!((inv[(0, 0)] - 0.25).abs() < 1e-12);
}
#[test]
fn inverse_times_original_is_identity() {
for rows in [
vec![vec![1.0, 2.0], vec![3.0, 4.0]],
vec![
vec![6.0, 1.0, 1.0],
vec![4.0, -2.0, 5.0],
vec![2.0, 8.0, 7.0],
],
vec![
vec![2.0, 0.0, 1.0, 3.0],
vec![1.0, 5.0, 2.0, 0.0],
vec![0.0, 1.0, 4.0, 1.0],
vec![3.0, 2.0, 0.0, 6.0],
],
] {
let refs: Vec<&[f64]> = rows.iter().map(|r| r.as_slice()).collect();
let m = from_rows(&refs);
let product = &m * &m.inverse();
for r in 0..product.height {
for c in 0..product.width {
let expected = if r == c { 1.0 } else { 0.0 };
assert!(
(product[(r, c)] - expected).abs() < 1e-9,
"({r},{c}) = {} expected {expected}",
product[(r, c)]
);
}
}
}
}
#[test]
#[should_panic(expected = "singular")]
fn inverse_of_singular_panics() {
let _ = from_rows(&[&[1.0, 2.0], &[2.0, 4.0]]).inverse();
}
#[test]
fn empty_determinant_is_one() {
assert!((Matrix::new(0, 0).determinant() - 1.0).abs() < 1e-12);
}
#[test]
fn transpose_round_trips() {
let m = from_rows(&[&[1.0, 2.0, 3.0], &[4.0, 5.0, 6.0]]);
let t = m.transpose();
assert_eq!((t.height, t.width), (3, 2));
assert_eq!(t.transpose()[(1, 2)], m[(1, 2)]);
}
}
+85 -2
View File
@@ -14,13 +14,95 @@ pub trait Observer<T: Time>: Send + Sync {
/// Called after each convergence iteration across the whole history.
fn on_iteration_end(&self, _iter: usize, _max_step: (f64, f64)) {}
/// Called after each time slice is processed within an iteration.
fn on_batch_processed(&self, _time: &T, _slice_idx: usize, _n_events: usize) {}
/// Called after each time slice is swept within an iteration.
///
/// A convergence iteration sweeps every slice twice — once travelling
/// backward through the history and once forward — so a multi-slice
/// history fires this twice per slice per iteration. A single-slice
/// history is swept once and fires once.
fn on_slice_processed(&self, _time: &T, _slice_idx: usize, _n_events: usize) {}
/// Called once when convergence completes (or max iters is reached).
fn on_converged(&self, _iters: usize, _final_step: (f64, f64), _converged: bool) {}
}
/// Shared and boxed observers forward to what they point at.
///
/// `History` takes its observer by value, so a caller who wants to *read* what
/// an observer recorded has to keep a handle to it. Without these impls the
/// natural spelling does not compile:
///
/// ```
/// # use std::sync::{Arc, Mutex};
/// # use trueskill_tt::{History, Observer};
/// #[derive(Default)]
/// struct Recorder {
/// iterations: Mutex<Vec<usize>>,
/// }
///
/// impl Observer<i64> for Recorder {
/// fn on_iteration_end(&self, iter: usize, _step: (f64, f64)) {
/// self.iterations.lock().unwrap().push(iter);
/// }
/// }
///
/// let recorder = Arc::new(Recorder::default());
/// let mut h = History::builder().observer(Arc::clone(&recorder)).build();
/// h.record_winner(&"a", &"b", 1).unwrap();
/// h.converge().unwrap();
///
/// // The caller's handle sees what the history's copy recorded.
/// assert!(!recorder.iterations.lock().unwrap().is_empty());
/// ```
///
/// The alternative was for every observer to wrap each of its own fields in an
/// `Arc` and derive `Clone` — one allocation and one lock per field, and a
/// pattern each implementor had to rediscover.
///
/// `?Sized` is deliberate: it makes `Arc<dyn Observer<T>>` and
/// `Box<dyn Observer<T>>` work, so observers can be chosen at runtime.
impl<T: Time, O: Observer<T> + ?Sized> Observer<T> for std::sync::Arc<O> {
fn on_iteration_end(&self, iter: usize, max_step: (f64, f64)) {
(**self).on_iteration_end(iter, max_step);
}
fn on_slice_processed(&self, time: &T, slice_idx: usize, n_events: usize) {
(**self).on_slice_processed(time, slice_idx, n_events);
}
fn on_converged(&self, iters: usize, final_step: (f64, f64), converged: bool) {
(**self).on_converged(iters, final_step, converged);
}
}
impl<T: Time, O: Observer<T> + ?Sized> Observer<T> for Box<O> {
fn on_iteration_end(&self, iter: usize, max_step: (f64, f64)) {
(**self).on_iteration_end(iter, max_step);
}
fn on_slice_processed(&self, time: &T, slice_idx: usize, n_events: usize) {
(**self).on_slice_processed(time, slice_idx, n_events);
}
fn on_converged(&self, iters: usize, final_step: (f64, f64), converged: bool) {
(**self).on_converged(iters, final_step, converged);
}
}
impl<T: Time, O: Observer<T> + ?Sized> Observer<T> for &O {
fn on_iteration_end(&self, iter: usize, max_step: (f64, f64)) {
(**self).on_iteration_end(iter, max_step);
}
fn on_slice_processed(&self, time: &T, slice_idx: usize, n_events: usize) {
(**self).on_slice_processed(time, slice_idx, n_events);
}
fn on_converged(&self, iters: usize, final_step: (f64, f64), converged: bool) {
(**self).on_converged(iters, final_step, converged);
}
}
/// ZST no-op observer; the default when none is configured.
#[derive(Copy, Clone, Debug, Default)]
pub struct NullObserver;
@@ -35,6 +117,7 @@ mod tests {
fn null_observer_compiles_for_i64() {
let o = NullObserver;
<NullObserver as Observer<i64>>::on_iteration_end(&o, 1, (0.0, 0.0));
<NullObserver as Observer<i64>>::on_slice_processed(&o, &7, 0, 3);
<NullObserver as Observer<i64>>::on_converged(&o, 5, (1e-6, 1e-6), true);
}
+53 -8
View File
@@ -29,14 +29,50 @@ pub enum Outcome {
impl Outcome {
/// `n`-team outcome where team `winner` won and everyone else tied for last.
///
/// Panics if `winner >= n`.
/// Note this ties every loser, so for `n >= 3` it needs a positive
/// `p_draw` — see `InferenceError::TieWithoutDrawProbability`.
///
/// # Panics
///
/// Panics if `winner >= n`. Use [`Outcome::try_winner`] when the index
/// comes from data rather than a literal.
///
/// This is the one constructor here that validates, and deliberately so.
/// Its siblings build freely and let ingestion reject what it cannot use,
/// which works because a malformed rank vector stays recognisable. An
/// out-of-range winner does not: `winner(5, 2)` would produce ranks
/// `[1, 1]`, an all-tied draw that ingestion accepts without complaint when
/// `p_draw > 0`. Asking "team 5 won" and silently getting "everyone drew"
/// is exactly the class of quiet wrong answer this crate keeps removing, so
/// the check happens here where the mistake is.
#[must_use]
pub fn winner(winner: u32, n: u32) -> Self {
assert!(winner < n, "winner index {winner} out of range 0..{n}");
Self::try_winner(winner, n)
.unwrap_or_else(|_| panic!("winner index {winner} out of range 0..{n}"))
}
/// `n`-team outcome where team `winner` won, or an error if `winner` is not
/// a valid team index.
///
/// The fallible form of [`Outcome::winner`], for when the index is computed
/// or parsed rather than written literally.
///
/// # Errors
///
/// `InvalidParameter` if `winner >= n`.
pub fn try_winner(winner: u32, n: u32) -> Result<Self, crate::InferenceError> {
if winner >= n {
return Err(crate::InferenceError::InvalidParameter {
name: "winner",
value: f64::from(winner),
});
}
let ranks: SmallVec<[u32; 4]> = (0..n).map(|i| if i == winner { 0 } else { 1 }).collect();
Self::Ranked(ranks)
Ok(Self::Ranked(ranks))
}
/// All `n` teams tied.
#[must_use]
pub fn draw(n: u32) -> Self {
Self::Ranked(SmallVec::from_vec(vec![0; n as usize]))
}
@@ -57,15 +93,18 @@ impl Outcome {
/// Explicit per-team continuous scores with a per-event noise override.
///
/// `sigma` must be `> 0.0`; debug-asserts otherwise.
/// `sigma` must be `> 0.0`. Constructing an `Outcome` with a non-positive
/// or NaN sigma is allowed; the value is rejected with
/// `InferenceError::InvalidParameter` when the event is ingested, so
/// callers get an error rather than a panic.
pub fn scores_with_sigma<I: IntoIterator<Item = f64>>(scores: I, sigma: f64) -> Self {
debug_assert!(sigma > 0.0, "score_sigma must be > 0.0 (got {sigma})");
Self::Scored {
scores: scores.into_iter().collect(),
sigma: Some(sigma),
}
}
#[must_use]
pub fn team_count(&self) -> usize {
match self {
Self::Ranked(r) => r.len(),
@@ -169,9 +208,15 @@ mod tests {
}
}
/// Construction accepts any sigma; the value is validated at ingestion so
/// callers receive an `InferenceError` rather than a panic. See
/// `tests/degenerate_inputs.rs::scored_event_rejects_non_positive_sigma`.
#[test]
#[should_panic(expected = "score_sigma must be > 0.0")]
fn scores_with_sigma_rejects_zero() {
let _ = Outcome::scores_with_sigma([3.0, 1.0], 0.0);
fn scores_with_sigma_defers_validation_to_ingestion() {
let o = Outcome::scores_with_sigma([3.0, 1.0], 0.0);
match o {
Outcome::Scored { sigma, .. } => assert_eq!(sigma, Some(0.0)),
Outcome::Ranked(_) => panic!("expected Scored variant"),
}
}
}
+729
View File
@@ -0,0 +1,729 @@
//! Outcome prediction: who wins, and how likely is a given finishing order.
//!
//! Prediction runs on *performances*, not skills. A competitor's skill is
//! inflated by their performance noise `beta` before any comparison, which is
//! what separates "how good are they" from "how will they do today".
//!
//! Two questions, two algorithms:
//!
//! - **Who finishes first.** Because performances are independent Gaussians,
//! the probability that team `i` beats every other team separates into a
//! *one-dimensional* integral — no multivariate orthant integral is
//! involved. [`quadrature::integrate`] evaluates it to near machine
//! precision for a few hundred `cdf` calls.
//! - **A specific finishing order.** The factor graph only ever constrains
//! rank-*adjacent* teams (see `Game::run_chain`), so the joint probability
//! of a full order is a chain of local constraints rather than a general
//! orthant probability. That chain collapses into a sequential recursion:
//! one cumulative integral per adjacent pair, `O(teams * grid)` overall.
//!
//! Both are deterministic. A sampler would have been easier to write and
//! would have made every `predict_*` call return a slightly different number,
//! which is not a property a rating library should have.
use crate::{Gaussian, quadrature};
/// Teams beyond this count make the outcome enumeration impractical.
///
/// Each realisation sorts into exactly one (permutation, tie-pattern) event,
/// so the space has `n! * 2^(n-1)` members: 24 at 3 teams, 192 at 4, 1_920 at
/// 5, 23_040 at 6. The jump to 322_560 at 7 is where enumerating stops being
/// a reasonable thing to do on a caller's behalf.
pub(crate) const MAX_TEAMS_FOR_DISTRIBUTION: usize = 6;
/// Relative tolerance for the first-place integrals.
///
/// The adaptive integrator reaches the exact two-team closed form to ~1e-15 at
/// this tolerance, which is round-off for a probability. `cdf` is no longer the
/// limit — it went to ~1 ULP when `erfc` moved to `libm` — so this is the
/// integrator's own floor.
const WIN_TOLERANCE: f64 = 1e-8;
/// Nodes for the ranking grid, and the floor below which a grid is pointless.
///
/// The recursion converges as O(h^2), so this trades nodes against accuracy
/// directly. Measured against the exact two-team closed form, 2_048 nodes leave
/// ~1.2e-6 of discretisation error and 8_192 reach ~1e-7.
///
/// Unlike the adaptive path there is no approximation floor underneath this any
/// more — `cdf` is accurate to ~1 ULP since `erfc` moved to `libm` — so the
/// error here is purely the grid, and a caller who needs more can only get it
/// by paying for more nodes. 8_192 is the accuracy/cost point chosen, not a
/// point where refining stops helping.
const MIN_GRID_POINTS: usize = 8_192;
const MAX_GRID_POINTS: usize = 262_144;
/// How many standard deviations of support the grid and integrals cover.
///
/// The normal density is below 1e-18 of its peak past nine sigma, far under
/// the precision of everything else here.
const SUPPORT_SIGMAS: f64 = 9.0;
/// Standard normal CDF at `z`.
fn phi(z: f64) -> f64 {
crate::cdf(z, 0.0, 1.0)
}
/// Normal density of `x` under `g`.
fn density(g: Gaussian, x: f64) -> f64 {
let sigma = g.sigma();
let z = (x - g.mu()) / sigma;
libm::exp(-0.5 * z * z) / (sigma * (2.0 * std::f64::consts::PI).sqrt())
}
/// Per-pair draw margins.
///
/// The margin is *not* a single number for the whole game: inference derives
/// it per rank-adjacent pair from those two teams' betas (`Game::likelihoods`).
/// Prediction has to use the same per-pair values or it answers a question
/// about a different model than the one that will actually be fitted.
pub(crate) struct Margins {
n: usize,
values: Vec<f64>,
}
impl Margins {
/// Build from a per-pair margin function.
pub(crate) fn new<F: Fn(usize, usize) -> f64>(n: usize, f: F) -> Self {
let mut values = vec![0.0; n * n];
for i in 0..n {
for j in 0..n {
if i != j {
values[i * n + j] = f(i, j);
}
}
}
Self { n, values }
}
fn get(&self, i: usize, j: usize) -> f64 {
self.values[i * self.n + j]
}
/// True when no pair can draw, so every tie has probability zero.
fn all_zero(&self) -> bool {
self.values.iter().all(|&v| v == 0.0)
}
}
/// `P(team i finishes strictly first)` for every team.
///
/// Strictly means beating each rival by more than that pair's draw margin, so
/// with a non-zero margin these sum to less than one; the shortfall is the
/// probability that the top place is shared.
pub(crate) fn win_probabilities(perf: &[Gaussian], margins: &Margins) -> Vec<f64> {
(0..perf.len())
.map(|i| {
let (mu, sigma) = (perf[i].mu(), perf[i].sigma());
let (lo, hi) = (mu - SUPPORT_SIGMAS * sigma, mu + SUPPORT_SIGMAS * sigma);
// Each rival's CDF turns over near its own mean plus the margin.
// Seeding there is what keeps a rival with a tiny sigma — a step
// function in disguise — from being stepped over.
let mut seeds = Vec::with_capacity(3 * perf.len());
for (j, rival) in perf.iter().enumerate().filter(|&(j, _)| j != i) {
let centre = rival.mu() + margins.get(i, j);
seeds.extend_from_slice(&[centre - rival.sigma(), centre, centre + rival.sigma()]);
}
quadrature::integrate(
|x| {
let d = density(perf[i], x);
if d == 0.0 {
return 0.0;
}
let beaten: f64 = (0..perf.len())
.filter(|&j| j != i)
.map(|j| phi((x - margins.get(i, j) - perf[j].mu()) / perf[j].sigma()))
.product();
d * beaten
},
lo,
hi,
&seeds,
WIN_TOLERANCE,
)
})
.collect()
}
/// Grid bounds and resolution covering every team's support.
///
/// Resolution is set by the *smallest* feature in play — the narrowest sigma,
/// or a draw margin narrower still — because that is what the recursion has to
/// resolve. A grid sized off the widest team would step over the narrow one.
fn grid_shape(perf: &[Gaussian], margins: &Margins) -> (f64, f64, usize) {
let lo = perf
.iter()
.map(|g| g.mu() - SUPPORT_SIGMAS * g.sigma())
.fold(f64::INFINITY, f64::min);
let hi = perf
.iter()
.map(|g| g.mu() + SUPPORT_SIGMAS * g.sigma())
.fold(f64::NEG_INFINITY, f64::max);
let narrowest = perf
.iter()
.map(Gaussian::sigma)
.fold(f64::INFINITY, f64::min);
let smallest_margin = margins
.values
.iter()
.copied()
.filter(|&m| m > 0.0)
.fold(f64::INFINITY, f64::min);
let feature = narrowest.min(smallest_margin);
let wanted = if feature.is_finite() && feature > 0.0 {
((hi - lo) / (feature / 12.0)).ceil()
} else {
MIN_GRID_POINTS as f64
};
let points = if wanted.is_finite() {
(wanted as usize).clamp(MIN_GRID_POINTS, MAX_GRID_POINTS)
} else {
MIN_GRID_POINTS
};
(lo, hi, points)
}
/// Densities of each team sampled on the shared grid.
struct Sampled {
lo: f64,
step: f64,
points: usize,
density: Vec<Vec<f64>>,
}
impl Sampled {
fn new(perf: &[Gaussian], margins: &Margins) -> Self {
let (lo, hi, points) = grid_shape(perf, margins);
let step = (hi - lo) / (points - 1) as f64;
let density = perf
.iter()
.map(|&g| {
(0..points)
.map(|i| density(g, lo + i as f64 * step))
.collect()
})
.collect();
Self {
lo,
step,
points,
density,
}
}
fn node(&self, i: usize) -> f64 {
self.lo + i as f64 * self.step
}
}
/// `P(order[0] >= order[1] >= ... )` with the given adjacency pattern.
///
/// `tied[k]` says whether `order[k]` and `order[k + 1]` finish within that
/// pair's draw margin. The recursion runs bottom-up: `carry` holds, for each
/// grid node, the probability that everything *below* the current team holds
/// given that team landed on that node. A strict gap reads a cumulative
/// integral; a tie reads a window. Both are O(1) against one prefix array,
/// so each level costs O(grid) and the whole order costs O(teams * grid).
fn order_probability(margins: &Margins, sampled: &Sampled, order: &[usize], tied: &[bool]) -> f64 {
let mut carry = vec![1.0; sampled.points];
for k in (0..order.len() - 1).rev() {
let below = order[k + 1];
let above = order[k];
let margin = margins.get(above, below);
let integrand: Vec<f64> = (0..sampled.points)
.map(|i| sampled.density[below][i] * carry[i])
.collect();
let cumulative = quadrature::Grid::from_values(sampled.lo, sampled.step, integrand);
carry = (0..sampled.points)
.map(|i| {
let x = sampled.node(i);
if tied[k] {
// Sorted order already implies `below <= above`, so the
// tie window is one-sided: [x - margin, x].
cumulative.integral_between(x - margin, x)
} else {
cumulative.integral_to(x - margin)
}
})
.collect();
}
let top = order[0];
let integrand: Vec<f64> = (0..sampled.points)
.map(|i| sampled.density[top][i] * carry[i])
.collect();
quadrature::Grid::from_values(sampled.lo, sampled.step, integrand).total()
}
/// Dense ranks implied by a sorted order and its tie pattern.
fn ranks_of(order: &[usize], tied: &[bool], n: usize) -> Vec<u32> {
let mut ranks = vec![0u32; n];
let mut rank = 0u32;
ranks[order[0]] = 0;
for k in 0..order.len() - 1 {
if !tied[k] {
rank += 1;
}
ranks[order[k + 1]] = rank;
}
ranks
}
/// Every (order, tie-pattern) event, or only the strict ones when no pair can
/// draw — a tie then has probability exactly zero and is not worth integrating.
fn events(n: usize, strict_only: bool) -> Vec<(Vec<usize>, Vec<bool>)> {
fn permute(current: &mut Vec<usize>, k: usize, out: &mut Vec<Vec<usize>>) {
if k == current.len() {
out.push(current.clone());
return;
}
for i in k..current.len() {
current.swap(k, i);
permute(current, k + 1, out);
current.swap(k, i);
}
}
let mut orders = Vec::new();
permute(&mut (0..n).collect(), 0, &mut orders);
let patterns: Vec<Vec<bool>> = if strict_only {
vec![vec![false; n - 1]]
} else {
(0..(1u32 << (n - 1)))
.map(|mask| (0..n - 1).map(|i| mask >> i & 1 == 1).collect())
.collect()
};
let mut out = Vec::with_capacity(orders.len() * patterns.len());
for order in orders {
for pattern in &patterns {
out.push((order.clone(), pattern.clone()));
}
}
out
}
/// The full distribution over finishing orders, aggregated by rank vector.
///
/// Orders that differ only *within* a tied group describe the same finishing
/// order, so their probabilities are summed into one entry.
pub(crate) fn outcome_distribution(perf: &[Gaussian], margins: &Margins) -> Vec<(Vec<u32>, f64)> {
let n = perf.len();
let sampled = Sampled::new(perf, margins);
let mut aggregated: Vec<(Vec<u32>, f64)> = Vec::new();
for (order, tied) in events(n, margins.all_zero()) {
let p = order_probability(margins, &sampled, &order, &tied);
let ranks = ranks_of(&order, &tied, n);
match aggregated.iter_mut().find(|(r, _)| *r == ranks) {
Some((_, acc)) => *acc += p,
None => aggregated.push((ranks, p)),
}
}
aggregated.sort_by(|a, b| b.1.partial_cmp(&a.1).unwrap_or(std::cmp::Ordering::Equal));
aggregated
}
/// All permutations of `items`.
fn permutations(items: &[usize]) -> Vec<Vec<usize>> {
fn go(current: &mut Vec<usize>, k: usize, out: &mut Vec<Vec<usize>>) {
if k == current.len() {
out.push(current.clone());
return;
}
for i in k..current.len() {
current.swap(k, i);
go(current, k + 1, out);
current.swap(k, i);
}
}
let mut out = Vec::new();
go(&mut items.to_vec(), 0, &mut out);
out
}
/// Every (order, tie-pattern) event consistent with a grouping by rank.
///
/// Teams sharing a rank may finish in any internal order, so this is the
/// product of each group's permutations. Adjacencies inside a group are ties;
/// the adjacency joining one group to the next is not.
fn orders_for_groups(groups: &[Vec<usize>]) -> Vec<(Vec<usize>, Vec<bool>)> {
let per_group: Vec<Vec<Vec<usize>>> = groups.iter().map(|g| permutations(g)).collect();
let mut out = Vec::new();
let mut choice = vec![0usize; groups.len()];
loop {
let mut order = Vec::new();
let mut tied = Vec::new();
for (gi, group) in per_group.iter().enumerate() {
for (offset, &member) in group[choice[gi]].iter().enumerate() {
if !order.is_empty() {
tied.push(offset != 0);
}
order.push(member);
}
}
out.push((order, tied));
let mut k = 0;
loop {
if k == choice.len() {
return out;
}
choice[k] += 1;
if choice[k] < per_group[k].len() {
break;
}
choice[k] = 0;
k += 1;
}
}
}
/// Probability of one specific rank vector.
///
/// Ties in `ranks` mean the tied teams may finish in any internal order, so
/// this sums the orders consistent with the requested ranking rather than
/// picking one.
pub(crate) fn ranking_probability(perf: &[Gaussian], margins: &Margins, ranks: &[u32]) -> f64 {
let n = perf.len();
let sampled = Sampled::new(perf, margins);
let mut distinct: Vec<u32> = ranks.to_vec();
distinct.sort_unstable();
distinct.dedup();
let groups: Vec<Vec<usize>> = distinct
.iter()
.map(|&r| (0..n).filter(|&i| ranks[i] == r).collect())
.collect();
orders_for_groups(&groups)
.iter()
.map(|(order, tied)| order_probability(margins, &sampled, order, tied))
.sum()
}
/// A distribution over the ways a contest could finish.
///
/// Each entry pairs a rank vector — the same shape [`crate::Outcome::ranking`]
/// takes, with equal ranks meaning a tie — against its probability. Entries
/// are ordered most likely first, and cover the whole outcome space, so the
/// probabilities sum to one.
///
/// The rank vectors compose directly with inference: feeding one to
/// `Game::ranked` asks "what would we believe if *this* happened", which is
/// what an expected-information-gain calculation needs alongside the weight.
#[derive(Clone, Debug, PartialEq)]
pub struct Prediction {
outcomes: Vec<(Vec<u32>, f64)>,
}
impl Prediction {
pub(crate) fn new(outcomes: Vec<(Vec<u32>, f64)>) -> Self {
Self { outcomes }
}
/// Every possible finishing order and its probability, most likely first.
pub fn outcomes(&self) -> impl ExactSizeIterator<Item = (&[u32], f64)> {
self.outcomes.iter().map(|(r, p)| (r.as_slice(), *p))
}
/// The single most likely finishing order.
#[must_use]
pub fn most_likely(&self) -> Option<(&[u32], f64)> {
self.outcomes.first().map(|(r, p)| (r.as_slice(), *p))
}
/// Probability of one specific finishing order, or zero if it cannot occur.
#[must_use]
pub fn probability_of(&self, ranks: &[u32]) -> f64 {
self.outcomes
.iter()
.find(|(r, _)| r.as_slice() == ranks)
.map_or(0.0, |(_, p)| *p)
}
/// `P(team i finishes strictly first)`, for each team.
///
/// Sums to less than one exactly when the top place can be shared; the
/// shortfall is [`Prediction::shared_first_place`].
#[must_use]
pub fn win_probabilities(&self) -> Vec<f64> {
let n = self.outcomes.first().map_or(0, |(r, _)| r.len());
let mut wins = vec![0.0; n];
for (ranks, p) in &self.outcomes {
let leaders = ranks.iter().filter(|&&r| r == 0).count();
if leaders == 1 {
let winner = ranks.iter().position(|&r| r == 0).expect("a rank-0 team");
wins[winner] += p;
}
}
wins
}
/// Probability that two or more teams share first place.
#[must_use]
pub fn shared_first_place(&self) -> f64 {
self.outcomes
.iter()
.filter(|(r, _)| r.iter().filter(|&&x| x == 0).count() > 1)
.map(|(_, p)| p)
.sum()
}
/// Total probability mass, which should be one.
///
/// Exposed because it is a genuine check on the numerics rather than a
/// formality: the outcome space is exhaustive and disjoint by construction,
/// so any drift from one is integration error and nothing else.
#[must_use]
pub fn total(&self) -> f64 {
self.outcomes.iter().map(|(_, p)| p).sum()
}
}
#[cfg(test)]
mod tests {
use super::*;
fn g(mu: f64, sigma: f64) -> Gaussian {
Gaussian::from_ms(mu, sigma)
}
fn flat(n: usize, eps: f64) -> Margins {
Margins::new(n, |_, _| eps)
}
/// Exact two-team result: `P(a first) = Phi((mu_a - mu_b - eps) / sd)`.
fn closed_form_two(a: Gaussian, b: Gaussian, eps: f64) -> (f64, f64) {
let sd = a.sigma().hypot(b.sigma());
(
phi((a.mu() - b.mu() - eps) / sd),
phi((b.mu() - a.mu() - eps) / sd),
)
}
#[test]
fn two_team_win_probabilities_match_the_closed_form() {
for (ma, sa, mb, sb, eps) in [
(0.0, 6.0, 0.0, 6.0, 0.0),
(3.0, 6.0, -2.0, 1.0, 0.0),
(0.0, 6.0, 0.0, 6.0, 2.0),
(3.0, 6.0, -2.0, 1.0, 1.5),
(40.0, 1.0, 0.0, 1.0, 0.0),
] {
let perf = [g(ma, sa), g(mb, sb)];
let got = win_probabilities(&perf, &flat(2, eps));
let (wa, wb) = closed_form_two(perf[0], perf[1], eps);
assert!(
(got[0] - wa).abs() < 1e-12 && (got[1] - wb).abs() < 1e-12,
"mu=({ma},{mb}) sigma=({sa},{sb}) eps={eps}: got {got:?}, want [{wa}, {wb}]"
);
}
}
/// The identity that a wrong-but-plausible implementation cannot fake:
/// with no draw margin, exactly one team finishes first.
#[test]
fn win_probabilities_sum_to_one_without_a_draw_margin() {
for perf in [
vec![g(0.0, 6.0), g(0.0, 6.0)],
vec![g(5.0, 6.0), g(0.0, 3.0), g(-5.0, 1.0)],
vec![
g(8.0, 2.0),
g(3.0, 6.0),
g(0.0, 1.0),
g(-3.0, 4.0),
g(-8.0, 6.0),
],
] {
let sum: f64 = win_probabilities(&perf, &flat(perf.len(), 0.0))
.iter()
.sum();
assert!(
(sum - 1.0).abs() < 1e-7,
"{} teams: sum = {sum}",
perf.len()
);
}
}
/// A rival with a tiny sigma is a step function in disguise. Fixed-node
/// quadrature steps over it and lands ~1e-2 out while still looking like a
/// probability; this is the case that rules that approach out.
#[test]
fn win_probabilities_survive_a_rival_with_a_tiny_sigma() {
let perf = [g(0.0, 0.001), g(0.5, 6.0), g(-0.5, 6.0)];
let got = win_probabilities(&perf, &flat(3, 0.0));
let sum: f64 = got.iter().sum();
assert!((sum - 1.0).abs() < 1e-6, "sum = {sum}, probs = {got:?}");
}
#[test]
fn a_stronger_team_is_more_likely_to_win() {
let perf = [g(10.0, 3.0), g(0.0, 3.0), g(-10.0, 3.0)];
let p = win_probabilities(&perf, &flat(3, 0.0));
assert!(p[0] > p[1] && p[1] > p[2], "not monotone: {p:?}");
}
#[test]
fn identical_teams_are_equally_likely_to_win() {
let perf = [g(1.0, 4.0), g(1.0, 4.0), g(1.0, 4.0)];
let p = win_probabilities(&perf, &flat(3, 0.0));
for probs in p.windows(2) {
assert!((probs[0] - probs[1]).abs() < 1e-9, "asymmetric: {p:?}");
}
}
/// Every realisation sorts into exactly one finishing order, so the whole
/// distribution must sum to one — with or without a draw margin.
#[test]
fn outcome_distribution_sums_to_one() {
for (perf, eps) in [
(vec![g(0.0, 6.0), g(0.0, 6.0)], 0.0),
(vec![g(0.0, 6.0), g(0.0, 6.0)], 2.0),
(vec![g(0.0, 6.0), g(0.0, 6.0), g(0.0, 6.0)], 0.0),
(vec![g(5.0, 6.0), g(0.0, 3.0), g(-5.0, 1.0)], 1.5),
(vec![g(0.0, 0.05), g(0.5, 6.0), g(-0.5, 6.0)], 1.0),
(
vec![g(6.0, 2.0), g(2.0, 6.0), g(-2.0, 1.0), g(-6.0, 4.0)],
1.0,
),
] {
let n = perf.len();
let dist = outcome_distribution(&perf, &flat(n, eps));
let sum: f64 = dist.iter().map(|(_, p)| p).sum();
assert!(
(sum - 1.0).abs() < 1e-6,
"{n} teams, eps={eps}: sum = {sum} over {} outcomes",
dist.len()
);
assert!(dist.iter().all(|(_, p)| *p >= 0.0), "negative probability");
}
}
/// With two teams the distribution is the exact win/draw/loss triple.
#[test]
fn two_team_distribution_matches_the_closed_form() {
let perf = [g(3.0, 6.0), g(-2.0, 1.0)];
let eps = 1.5;
let dist = outcome_distribution(&perf, &flat(2, eps));
let (wa, wb) = closed_form_two(perf[0], perf[1], eps);
let find = |ranks: &[u32]| {
dist.iter()
.find(|(r, _)| r == ranks)
.map_or(0.0, |(_, p)| *p)
};
assert!(
(find(&[0, 1]) - wa).abs() < 1e-6,
"a wins: {}",
find(&[0, 1])
);
assert!(
(find(&[1, 0]) - wb).abs() < 1e-6,
"b wins: {}",
find(&[1, 0])
);
assert!(
(find(&[0, 0]) - (1.0 - wa - wb)).abs() < 1e-6,
"draw: {}",
find(&[0, 0])
);
}
/// Asking for one ranking must agree with that ranking's entry in the
/// full distribution — the two use different code paths to the same value.
#[test]
fn ranking_probability_agrees_with_the_distribution() {
let perf = [g(5.0, 6.0), g(0.0, 3.0), g(-5.0, 1.0)];
let eps = 1.5;
let margins = flat(3, eps);
let dist = outcome_distribution(&perf, &margins);
for (ranks, expected) in &dist {
let direct = ranking_probability(&perf, &margins, ranks);
assert!(
(direct - expected).abs() < 1e-9,
"ranks {ranks:?}: direct {direct} vs distribution {expected}"
);
}
}
/// Tie mass is controlled by the draw margin. Only the *all-tied* outcome
/// is monotone in it: every one of its constraints is a window that widens
/// with the margin. A partially-tied outcome like `[0, 0, 1]` is not, and
/// must not be asserted to be — widening the margin makes its tie easier
/// but its "and the last team is strictly behind by more than the margin"
/// clause harder, so it peaks and then falls.
#[test]
fn all_tied_probability_grows_with_the_draw_margin() {
let perf = [g(0.0, 4.0), g(0.0, 4.0), g(-8.0, 2.0)];
let mut previous = 0.0;
for eps in [0.0, 0.5, 1.0, 2.0, 4.0, 8.0, 24.0] {
let p = ranking_probability(&perf, &flat(3, eps), &[0, 0, 0]);
assert!(p >= previous, "eps={eps}: {p} < {previous}");
if eps == 0.0 {
assert!(p < 1e-12, "a tie needs a margin, got {p}");
}
previous = p;
}
assert!(
previous > 0.9,
"a very wide margin ties everyone: {previous}"
);
}
/// The converse, stated as the non-property it is: a partially-tied
/// outcome is non-monotone in the margin. Pinning this down stops a future
/// change from "fixing" it into monotonicity and quietly breaking the model.
#[test]
fn a_partially_tied_outcome_peaks_in_the_middle() {
let perf = [g(0.0, 4.0), g(0.0, 4.0), g(-8.0, 2.0)];
let sweep: Vec<f64> = [0.5, 2.0, 4.0, 8.0, 16.0]
.iter()
.map(|&eps| ranking_probability(&perf, &flat(3, eps), &[0, 0, 1]))
.collect();
let peak = sweep
.iter()
.enumerate()
.fold(
(0, 0.0),
|(bi, bv), (i, &v)| if v > bv { (i, v) } else { (bi, bv) },
)
.0;
assert!(
peak > 0 && peak < sweep.len() - 1,
"expected an interior peak: {sweep:?}"
);
}
/// With no draw margin a tie has probability exactly zero, and the
/// enumeration must not waste work pretending otherwise.
#[test]
fn ties_are_impossible_without_a_draw_margin() {
let perf = [g(0.0, 4.0), g(0.0, 4.0), g(0.0, 4.0)];
let dist = outcome_distribution(&perf, &flat(3, 0.0));
assert_eq!(dist.len(), 6, "expected only the 6 strict orders: {dist:?}");
assert!(dist.iter().all(|(r, _)| {
let mut seen = r.clone();
seen.sort_unstable();
seen.dedup();
seen.len() == r.len()
}));
}
}
+322
View File
@@ -0,0 +1,322 @@
//! Deterministic numerical integration for the prediction paths.
//!
//! Prediction asks two questions that have no closed form beyond two teams:
//! "who finishes first" and "how likely is this exact finishing order". Both
//! reduce to integrals over a single performance variable, so neither needs a
//! sampler — and that matters, because a Monte Carlo predictor would make
//! `predict_*` non-reproducible and would answer a slightly different question
//! on every call.
//!
//! Two routines live here:
//!
//! - [`integrate`], adaptive Gauss-Kronrod G7-K15, for the first-place
//! marginals. It carries its own error estimate, so it can refine where the
//! integrand actually bends instead of guessing a node count up front.
//! - [`Grid`], a uniform grid with trapezoid prefix sums, for the ranking
//! chain recursion, where each level needs the *running* integral of the
//! level below at arbitrary points rather than one definite integral.
//!
//! Fixed-node Gauss-Hermite is the obvious tool for the first of these and is
//! a trap: the integrand is a product of normal CDFs, and when one team's
//! sigma is much smaller than the integrating team's, that product turns into
//! a near-step function narrower than the node spacing. The nodes step over
//! it and the result is wrong by ~1e-2 while still looking like a probability.
//! Adaptive refinement is what makes the small-sigma case safe.
/// Kronrod 15-point abscissae, non-negative half, descending.
const XGK: [f64; 8] = [
0.991_455_371_120_813,
0.949_107_912_342_759,
0.864_864_423_359_769,
0.741_531_185_599_394,
0.586_087_235_467_691,
0.405_845_151_377_397,
0.207_784_955_007_898,
0.0,
];
/// Kronrod 15-point weights, matching [`XGK`].
const WGK: [f64; 8] = [
0.022_935_322_010_529,
0.063_092_092_629_979,
0.104_790_010_322_250,
0.140_653_259_715_525,
0.169_004_726_639_267,
0.190_350_578_064_785,
0.204_432_940_075_298,
0.209_482_141_084_728,
];
/// Gauss 7-point weights, applying to the odd-indexed [`XGK`] entries.
const WG: [f64; 4] = [
0.129_484_966_168_870,
0.279_705_391_489_277,
0.381_830_050_505_119,
0.417_959_183_673_469,
];
/// Panels are bisected worst-first; this bounds the work on a pathological
/// integrand rather than letting it spin.
const MAX_SUBDIVISIONS: usize = 200;
/// One G7-K15 panel over `[a, b]`: `(integral, absolute error estimate)`.
///
/// The error estimate is the gap between the embedded 7-point Gauss rule and
/// the 15-point Kronrod extension. It is the only reason this is preferable
/// to a fixed rule: it tells the caller *where* the integrand is hard.
fn gk15<F: Fn(f64) -> f64>(f: &F, a: f64, b: f64) -> (f64, f64) {
let centre = 0.5 * (a + b);
let half = 0.5 * (b - a);
let mut kronrod = 0.0;
let mut gauss = 0.0;
for i in 0..8 {
let offset = XGK[i] * half;
// XGK[7] is the centre node and must not be counted twice.
let sum = if i == 7 {
f(centre)
} else {
f(centre - offset) + f(centre + offset)
};
kronrod += WGK[i] * sum;
if i % 2 == 1 {
gauss += WG[i / 2] * sum;
}
}
(kronrod * half, ((kronrod - gauss) * half).abs())
}
/// Adaptively integrate `f` over `[a, b]` to relative tolerance `tol`.
///
/// `seeds` are interior points where the integrand is known to bend sharply —
/// for a product of normal CDFs, each rival's transition centre. Splitting
/// there up front costs nothing and saves the adaptive loop from having to
/// discover a step by bisection.
///
/// Returns the integral. The error estimate is consumed internally rather
/// than returned: callers here integrate probability densities, where the
/// meaningful check is the sum-to-one identity over a whole outcome space,
/// not a per-integral residual.
pub(crate) fn integrate<F: Fn(f64) -> f64>(f: F, a: f64, b: f64, seeds: &[f64], tol: f64) -> f64 {
// Explicit rather than `!(b > a)`: a NaN bound must fall through to zero
// rather than being read as a valid ordering.
if a.partial_cmp(&b) != Some(std::cmp::Ordering::Less) {
return 0.0;
}
let mut edges: Vec<f64> = Vec::with_capacity(seeds.len() + 2);
edges.push(a);
edges.push(b);
for &s in seeds {
if s > a && s < b {
edges.push(s);
}
}
edges.sort_by(|p, q| p.partial_cmp(q).expect("integration bounds are finite"));
edges.dedup();
// (lo, hi, integral, error)
let mut panels: Vec<(f64, f64, f64, f64)> = edges
.windows(2)
.map(|w| {
let (v, e) = gk15(&f, w[0], w[1]);
(w[0], w[1], v, e)
})
.collect();
for _ in 0..MAX_SUBDIVISIONS {
let total: f64 = panels.iter().map(|p| p.2).sum();
let error: f64 = panels.iter().map(|p| p.3).sum();
// Absolute floor as well as relative: these integrands are
// probabilities, so an absolute 1e-15 is already past the useful
// precision of the underlying `cdf`.
if error <= tol * total.abs().max(1e-12) || error < 1e-15 {
break;
}
let worst = panels
.iter()
.enumerate()
.fold((0usize, f64::NEG_INFINITY), |(bi, be), (i, p)| {
if p.3 > be { (i, p.3) } else { (bi, be) }
})
.0;
let (lo, hi, _, _) = panels[worst];
let mid = 0.5 * (lo + hi);
// Bisection has hit the floating-point floor; refining further would
// loop without reducing the error.
if !(mid > lo && mid < hi) {
break;
}
let (v1, e1) = gk15(&f, lo, mid);
let (v2, e2) = gk15(&f, mid, hi);
panels[worst] = (lo, mid, v1, e1);
panels.push((mid, hi, v2, e2));
}
panels.iter().map(|p| p.2).sum()
}
/// A uniform grid carrying trapezoid prefix sums of one integrand.
///
/// The ranking recursion needs, at every level, the running integral of the
/// level below evaluated at arbitrary points — a cumulative integral, not a
/// definite one. Prefix sums give that in O(1) per query after an O(G) build,
/// which is what keeps a full ranking probability linear in the team count.
pub(crate) struct Grid {
lo: f64,
step: f64,
/// Integrand sampled at each node.
values: Vec<f64>,
/// `prefix[i]` is the integral from `lo` to node `i`.
prefix: Vec<f64>,
}
impl Grid {
/// Build directly from already-sampled values.
///
/// The ranking recursion evaluates every level on the same nodes, so the
/// per-team densities are sampled once and reused; re-evaluating `exp`
/// per level would dominate the cost.
pub(crate) fn from_values(lo: f64, step: f64, values: Vec<f64>) -> Self {
let mut prefix = vec![0.0; values.len()];
for i in 1..values.len() {
prefix[i] = prefix[i - 1] + 0.5 * step * (values[i - 1] + values[i]);
}
Self {
lo,
step,
values,
prefix,
}
}
/// Integral from the grid's lower bound up to `x`.
///
/// Clamped at both ends: the caller sizes the grid to cover the whole
/// support, so a query outside it is asking for a tail that is zero (below)
/// or the whole mass (above).
pub(crate) fn integral_to(&self, x: f64) -> f64 {
let last = self.values.len() - 1;
if x <= self.lo {
return 0.0;
}
if x >= self.lo + last as f64 * self.step {
return self.prefix[last];
}
let scaled = (x - self.lo) / self.step;
let i = scaled.floor() as usize;
let frac = scaled - i as f64;
// Whole cells, plus the trapezoid over the partial cell. The integrand
// is linear within a cell under the trapezoid rule, so the partial
// piece is exact with respect to that same approximation.
self.prefix[i]
+ frac
* self.step
* (self.values[i] + 0.5 * frac * (self.values[i + 1] - self.values[i]))
}
/// Integral over `[from, to]`.
pub(crate) fn integral_between(&self, from: f64, to: f64) -> f64 {
(self.integral_to(to) - self.integral_to(from)).max(0.0)
}
/// Total integral over the whole grid.
pub(crate) fn total(&self) -> f64 {
self.prefix[self.values.len() - 1]
}
}
#[cfg(test)]
mod tests {
use super::*;
const TOL: f64 = 1e-10;
/// Sample `f` over `[lo, hi]` at `points` nodes.
fn sample<F: FnMut(f64) -> f64>(lo: f64, hi: f64, points: usize, mut f: F) -> Grid {
let step = (hi - lo) / (points - 1) as f64;
Grid::from_values(
lo,
step,
(0..points).map(|i| f(lo + i as f64 * step)).collect(),
)
}
#[test]
fn integrates_a_polynomial_exactly() {
// G7-K15 is exact for polynomials well past cubic, so a single panel
// should already be at round-off.
let v = integrate(|x| 3.0 * x * x + 2.0 * x + 1.0, 0.0, 2.0, &[], TOL);
assert!((v - 14.0).abs() < 1e-12, "got {v}");
}
#[test]
fn integrates_a_gaussian_density_to_one() {
let f = |x: f64| (-0.5 * x * x).exp() / (2.0 * std::f64::consts::PI).sqrt();
let v = integrate(f, -10.0, 10.0, &[], TOL);
assert!((v - 1.0).abs() < 1e-12, "got {v}");
}
#[test]
fn resolves_a_step_far_narrower_than_the_initial_panel() {
// The failure mode that rules out fixed-node quadrature: a transition
// 1e-4 wide inside a range of 20. A fixed rule steps over it.
let f = |x: f64| if x < 0.5 { 0.0 } else { 1.0 };
let v = integrate(f, -10.0, 10.0, &[0.5], TOL);
assert!((v - 9.5).abs() < 1e-6, "got {v}");
}
#[test]
fn seeds_do_not_change_the_value_of_a_smooth_integrand() {
let f = |x: f64| (-0.5 * x * x).exp();
let plain = integrate(f, -8.0, 8.0, &[], TOL);
let seeded = integrate(f, -8.0, 8.0, &[-3.0, 0.25, 5.5], TOL);
assert!((plain - seeded).abs() < 1e-12, "{plain} vs {seeded}");
}
#[test]
fn empty_or_inverted_range_integrates_to_zero() {
assert_eq!(integrate(|_| 1.0, 1.0, 1.0, &[], TOL), 0.0);
assert_eq!(integrate(|_| 1.0, 2.0, 1.0, &[], TOL), 0.0);
}
#[test]
fn grid_prefix_matches_a_known_cumulative_integral() {
// f(x) = x over [0, 4]; integral to x is x^2/2.
let g = sample(0.0, 4.0, 4001, |x| x);
for probe in [0.0, 0.5, 1.0, 2.5, 3.75, 4.0] {
let want = probe * probe / 2.0;
let got = g.integral_to(probe);
assert!(
(got - want).abs() < 1e-9,
"at {probe}: got {got}, want {want}"
);
}
assert!((g.total() - 8.0).abs() < 1e-9);
}
#[test]
fn grid_between_is_the_difference_of_two_prefixes() {
let g = sample(-5.0, 5.0, 8001, |x| (-0.5 * x * x).exp());
let whole = g.integral_between(-5.0, 5.0);
let split = g.integral_between(-5.0, 0.3) + g.integral_between(0.3, 5.0);
assert!((whole - split).abs() < 1e-12, "{whole} vs {split}");
}
#[test]
fn grid_clamps_queries_outside_its_support() {
let g = sample(0.0, 1.0, 101, |_| 1.0);
assert_eq!(g.integral_to(-3.0), 0.0);
assert!((g.integral_to(9.0) - 1.0).abs() < 1e-12);
// Reversed bounds must not produce negative probability mass.
assert_eq!(g.integral_between(0.8, 0.2), 0.0);
}
}
+57 -2
View File
@@ -9,13 +9,16 @@ use crate::{
/// Static rating configuration: prior skill, performance noise `beta`, drift.
///
/// Renamed from `Player` in T2; `Rating` better describes the data
/// (a configuration) vs. a person (who's a `Competitor` with state).
/// A configuration rather than a person: the per-history temporal state
/// (messages, last appearance) lives on `Competitor`.
#[derive(Clone, Copy, Debug)]
pub struct Rating<T: Time = i64, D: Drift<T> = ConstantDrift> {
pub(crate) prior: Gaussian,
pub(crate) beta: f64,
pub(crate) drift: D,
/// Multiplier on the drift *variance* this competitor accumulates; 1.0 is
/// the neutral default. Set per competitor via `Member::with_drift_scale`.
pub(crate) drift_scale: f64,
pub(crate) _time: PhantomData<T>,
}
@@ -25,10 +28,61 @@ impl<T: Time, D: Drift<T>> Rating<T, D> {
prior,
beta,
drift,
drift_scale: 1.0,
_time: PhantomData,
}
}
/// Scale how fast this competitor drifts, relative to `drift`.
///
/// Multiplies the drift *variance*, so the scale is in the same units as
/// `gamma`. `0.0` pins the competitor still.
#[must_use]
pub fn with_drift_scale(mut self, drift_scale: f64) -> Self {
self.drift_scale = drift_scale;
self
}
/// The configured prior skill estimate.
#[must_use]
pub fn prior(&self) -> Gaussian {
self.prior
}
/// Performance noise: how much a single showing varies around the skill.
#[must_use]
pub fn beta(&self) -> f64 {
self.beta
}
/// The drift model governing how skill may move between events.
#[must_use]
pub fn drift(&self) -> D {
self.drift
}
/// This competitor's multiplier on the drift variance; 1.0 is neutral.
#[must_use]
pub fn drift_scale(&self) -> f64 {
self.drift_scale
}
/// Drift variance accumulated over `from -> to`, scaled for this competitor.
///
/// The single place the scale is applied for a `Time`-typed span. Callers
/// must go through this rather than `self.drift` directly, so a competitor's
/// scale cannot be silently skipped.
pub(crate) fn drift_variance_delta(&self, from: &T, to: &T) -> f64 {
self.drift.variance_delta(from, to) * self.drift_scale * self.drift_scale
}
/// Drift variance for a cached elapsed count, scaled for this competitor.
///
/// The counterpart of `drift_variance_delta` for the cached-elapsed paths.
pub(crate) fn drift_variance_for_elapsed(&self, elapsed: i64) -> f64 {
self.drift.variance_for_elapsed(elapsed) * self.drift_scale * self.drift_scale
}
pub(crate) fn performance(&self) -> Gaussian {
self.prior.forget(self.beta.powi(2))
}
@@ -40,6 +94,7 @@ impl Default for Rating<i64, ConstantDrift> {
prior: Gaussian::default(),
beta: BETA,
drift: ConstantDrift(GAMMA),
drift_scale: 1.0,
_time: PhantomData,
}
}
-126
View File
@@ -1,126 +0,0 @@
//! Schedule trait and built-in implementations.
//!
//! A schedule drives factor propagation to convergence. The default
//! `EpsilonOrMax` performs one TeamSum sweep (setup) then alternating
//! forward/backward sweeps over the iterating factors until the max
//! delta drops below epsilon or `max` iterations is reached.
use crate::factor::{BuiltinFactor, Factor, VarStore};
/// Result returned by a `Schedule::run` call.
#[derive(Debug, Clone, Copy)]
pub struct ScheduleReport {
pub iterations: usize,
pub final_step: (f64, f64),
pub converged: bool,
}
/// Drives factor propagation to convergence.
pub trait Schedule: Send + Sync {
fn run(&self, factors: &mut [BuiltinFactor], vars: &mut VarStore) -> ScheduleReport;
}
/// Default schedule: sweep forward then backward until step ≤ eps or iter == max.
///
/// Matches the existing `Game::likelihoods` loop bit-for-bit when given the
/// same factor layout (TeamSums first, then alternating RankDiff/Trunc pairs).
#[derive(Debug, Clone, Copy)]
pub struct EpsilonOrMax {
pub eps: f64,
pub max: usize,
}
impl Default for EpsilonOrMax {
fn default() -> Self {
// Matches today's hard-coded tolerance and iteration cap.
Self { eps: 1e-6, max: 10 }
}
}
impl Schedule for EpsilonOrMax {
fn run(&self, factors: &mut [BuiltinFactor], vars: &mut VarStore) -> ScheduleReport {
// Partition: leading run of TeamSum factors run exactly once (setup).
let n_setup = factors
.iter()
.position(|f| !matches!(f, BuiltinFactor::TeamSum(_)))
.unwrap_or(factors.len());
for f in factors[..n_setup].iter_mut() {
f.propagate(vars);
}
let mut iterations = 0;
let mut final_step = (f64::INFINITY, f64::INFINITY);
let mut converged = false;
if n_setup < factors.len() {
for _ in 0..self.max {
let mut step = (0.0_f64, 0.0_f64);
// Forward sweep over iterating factors.
for f in factors[n_setup..].iter_mut() {
let d = f.propagate(vars);
step.0 = step.0.max(d.0);
step.1 = step.1.max(d.1);
}
// Backward sweep.
for f in factors[n_setup..].iter_mut().rev() {
let d = f.propagate(vars);
step.0 = step.0.max(d.0);
step.1 = step.1.max(d.1);
}
iterations += 1;
final_step = step;
if step.0 <= self.eps && step.1 <= self.eps {
converged = true;
break;
}
}
}
ScheduleReport {
iterations,
final_step,
converged,
}
}
}
#[cfg(test)]
mod tests {
use super::*;
use crate::{N_INF, factor::team_sum::TeamSumFactor, gaussian::Gaussian};
#[test]
fn schedule_runs_setup_factors_once() {
// Single TeamSum factor; schedule should propagate it exactly once and report 0 iterations.
let mut vars = VarStore::new();
let out = vars.alloc(N_INF);
let mut factors = vec![BuiltinFactor::TeamSum(TeamSumFactor {
inputs: vec![(Gaussian::from_ms(5.0, 1.0), 1.0)],
out,
})];
let schedule = EpsilonOrMax::default();
let report = schedule.run(&mut factors, &mut vars);
assert_eq!(report.iterations, 0);
// The team-perf var should hold the sum.
let result = vars.get(out);
assert!((result.mu() - 5.0).abs() < 1e-12);
}
#[test]
fn report_marks_converged_when_no_iterating_factors() {
// No iterating factors → 0 iterations, converged stays false (loop never ran).
let mut vars = VarStore::new();
let out = vars.alloc(N_INF);
let mut factors = vec![BuiltinFactor::TeamSum(TeamSumFactor {
inputs: vec![(Gaussian::from_ms(0.0, 1.0), 1.0)],
out,
})];
let report = EpsilonOrMax::default().run(&mut factors, &mut vars);
assert_eq!(report.iterations, 0);
}
}
+6 -1
View File
@@ -2,7 +2,7 @@ use crate::{Index, competitor::Competitor, drift::Drift, time::Time};
/// Dense Vec-backed store for competitor state in History.
///
/// Indexed directly by Index.0, eliminating HashMap hashing in the
/// Indexed directly by Index.0, eliminating `HashMap` hashing in the
/// forward/backward sweep. Uses `Vec<Option<Competitor<T, D>>>` so slots can be
/// absent without an explicit present mask.
#[derive(Debug)]
@@ -21,6 +21,7 @@ impl<T: Time, D: Drift<T>> Default for CompetitorStore<T, D> {
}
impl<T: Time, D: Drift<T>> CompetitorStore<T, D> {
#[must_use]
pub fn new() -> Self {
Self::default()
}
@@ -39,6 +40,7 @@ impl<T: Time, D: Drift<T>> CompetitorStore<T, D> {
self.competitors[idx.0] = Some(competitor);
}
#[must_use]
pub fn get(&self, idx: Index) -> Option<&Competitor<T, D>> {
self.competitors.get(idx.0).and_then(|slot| slot.as_ref())
}
@@ -49,14 +51,17 @@ impl<T: Time, D: Drift<T>> CompetitorStore<T, D> {
.and_then(|slot| slot.as_mut())
}
#[must_use]
pub fn contains(&self, idx: Index) -> bool {
self.get(idx).is_some()
}
#[must_use]
pub fn len(&self) -> usize {
self.n_present
}
#[must_use]
pub fn is_empty(&self) -> bool {
self.n_present == 0
}
+110 -51
View File
@@ -1,15 +1,27 @@
use std::collections::HashMap;
use crate::{Index, time_slice::Skill};
/// Dense Vec-backed store for per-agent skill state within a TimeSlice.
/// Compact per-slice store for skill state, addressed by a slice-local slot.
///
/// Indexed directly by Index.0, eliminating HashMap hashing in the inner
/// convergence loop. Uses a parallel `present` mask so iteration skips
/// absent slots without incurring per-slot Option overhead in the hot path.
/// `skills` holds one entry per competitor **in this slice**, so memory is
/// O(competitors in the slice). It used to be a dense `Vec<Skill>` indexed by
/// the global `Index.0`, which made a slice's footprint O(largest index it
/// touches): a single 1v1 game between competitors 19998 and 19999 reserved
/// 20,000 slots.
///
/// The dense layout existed to keep `HashMap` hashing out of the inner
/// convergence loop, and that property is preserved. `slots` is consulted only
/// while building a slice; every hot-path access goes through
/// [`SkillStore::at`] / [`SkillStore::at_mut`] with a slot resolved once at
/// ingestion and cached on the event's `Item`.
#[derive(Debug, Default)]
pub struct SkillStore {
skills: Vec<Skill>,
present: Vec<bool>,
n_present: usize,
/// Slot -> global index, parallel to `skills`, so iteration can report the
/// global index without a reverse lookup.
indices: Vec<Index>,
slots: HashMap<Index, u32>,
}
impl SkillStore {
@@ -17,76 +29,99 @@ impl SkillStore {
Self::default()
}
fn ensure_capacity(&mut self, idx: usize) {
if idx >= self.skills.len() {
self.skills.resize_with(idx + 1, Skill::default);
self.present.resize(idx + 1, false);
}
/// Resolve a global index to this slice's slot, if the competitor is here.
///
/// This hashes. Call it at ingestion and cache the result; do not call it
/// from the convergence loop.
pub fn slot_of(&self, idx: Index) -> Option<u32> {
self.slots.get(&idx).copied()
}
pub fn insert(&mut self, idx: Index, skill: Skill) {
self.ensure_capacity(idx.0);
if !self.present[idx.0] {
self.n_present += 1;
/// Skill at a slot resolved earlier by [`SkillStore::slot_of`].
///
/// # Panics
///
/// Panics if `slot` is out of range, which means it came from a different
/// slice's store.
pub fn at(&self, slot: u32) -> &Skill {
&self.skills[slot as usize]
}
/// Mutable counterpart to [`SkillStore::at`].
///
/// # Panics
///
/// Panics if `slot` is out of range.
pub fn at_mut(&mut self, slot: u32) -> &mut Skill {
&mut self.skills[slot as usize]
}
/// Insert or overwrite a competitor's skill, returning its slot.
pub fn insert(&mut self, idx: Index, skill: Skill) -> u32 {
match self.slots.get(&idx) {
Some(&slot) => {
self.skills[slot as usize] = skill;
slot
}
None => {
let slot = u32::try_from(self.skills.len())
.expect("a time slice cannot hold more than u32::MAX competitors");
self.skills.push(skill);
self.indices.push(idx);
self.slots.insert(idx, slot);
slot
}
}
self.skills[idx.0] = skill;
self.present[idx.0] = true;
}
pub fn get(&self, idx: Index) -> Option<&Skill> {
if idx.0 < self.present.len() && self.present[idx.0] {
Some(&self.skills[idx.0])
} else {
None
}
self.slot_of(idx).map(|slot| self.at(slot))
}
pub fn get_mut(&mut self, idx: Index) -> Option<&mut Skill> {
if idx.0 < self.present.len() && self.present[idx.0] {
Some(&mut self.skills[idx.0])
} else {
None
}
self.slot_of(idx)
.map(|slot| &mut self.skills[slot as usize])
}
#[allow(dead_code)]
/// Whether a competitor is present in this slice. Test-only.
#[cfg(test)]
pub fn contains(&self, idx: Index) -> bool {
idx.0 < self.present.len() && self.present[idx.0]
self.slots.contains_key(&idx)
}
#[allow(dead_code)]
/// Number of competitors in this slice. Test-only.
#[cfg(test)]
pub fn len(&self) -> usize {
self.n_present
self.skills.len()
}
#[allow(dead_code)]
pub fn is_empty(&self) -> bool {
self.n_present == 0
/// Slots actually allocated — the quantity #17 is about, and NOT the same
/// as `len` for every possible implementation.
///
/// A store indexed by the global `Index` must report `max_index + 1` here
/// while reporting the true competitor count from `len`, which is exactly
/// how the original defect hid. Tests that mean to pin the footprint must
/// assert on this.
#[cfg(test)]
pub fn allocated_slots(&self) -> usize {
self.skills.len()
}
/// Iterate in slot order — the order competitors were first seen in this
/// slice. Deterministic for a given event order, which is what the
/// cross-thread determinism test relies on.
pub fn iter(&self) -> impl Iterator<Item = (Index, &Skill)> {
self.present.iter().enumerate().filter_map(|(i, &p)| {
if p {
Some((Index(i), &self.skills[i]))
} else {
None
}
})
self.indices.iter().copied().zip(self.skills.iter())
}
pub fn iter_mut(&mut self) -> impl Iterator<Item = (Index, &mut Skill)> {
self.skills
.iter_mut()
.zip(self.present.iter())
.enumerate()
.filter_map(|(i, (s, &p))| if p { Some((Index(i), s)) } else { None })
self.indices.iter().copied().zip(self.skills.iter_mut())
}
pub fn keys(&self) -> impl Iterator<Item = Index> + '_ {
self.present
.iter()
.enumerate()
.filter_map(|(i, &p)| if p { Some(Index(i)) } else { None })
self.indices.iter().copied()
}
}
@@ -112,7 +147,7 @@ mod tests {
}
#[test]
fn iter_skips_absent_slots() {
fn iter_reports_global_indices() {
let mut store = SkillStore::new();
store.insert(Index(0), Skill::default());
store.insert(Index(5), Skill::default());
@@ -127,4 +162,28 @@ mod tests {
store.insert(Index(2), Skill::default());
assert_eq!(store.len(), 1);
}
/// The defect in #17: a slice holding two competitors must cost the same
/// whether their indices are small or large.
#[test]
fn footprint_is_independent_of_index_magnitude() {
let mut low = SkillStore::new();
low.insert(Index(0), Skill::default());
low.insert(Index(1), Skill::default());
let mut high = SkillStore::new();
high.insert(Index(19_998), Skill::default());
high.insert(Index(19_999), Skill::default());
assert_eq!(low.len(), high.len());
assert_eq!(low.skills.capacity(), high.skills.capacity());
}
#[test]
fn slot_survives_reinsert() {
let mut store = SkillStore::new();
let first = store.insert(Index(7), Skill::default());
let again = store.insert(Index(7), Skill::default());
assert_eq!(first, again);
}
}
+390 -109
View File
@@ -14,7 +14,6 @@ use crate::{
rating::Rating,
storage::{CompetitorStore, SkillStore},
time::Time,
tuple_gt, tuple_max,
};
#[derive(Debug)]
@@ -23,7 +22,6 @@ pub(crate) struct Skill {
backward: Gaussian,
likelihood: Gaussian,
pub(crate) elapsed: i64,
pub(crate) online: Gaussian,
}
impl Skill {
@@ -39,7 +37,6 @@ impl Default for Skill {
backward: N_INF,
likelihood: N_INF,
elapsed: 0,
online: N_INF,
}
}
}
@@ -51,43 +48,48 @@ pub enum EventKind {
Scored { score_sigma: f64 },
}
#[derive(Debug)]
#[derive(Clone, Debug)]
struct Item {
agent: Index,
/// This competitor's slot in the owning slice's `SkillStore`, resolved
/// once at ingestion.
///
/// The convergence loop reaches skills through this rather than through
/// `agent`, which is what keeps `HashMap` hashing out of the hot path now
/// that the store is compact rather than indexed by the global `Index`.
slot: u32,
likelihood: Gaussian,
}
impl Item {
fn within_prior<T: Time, D: Drift<T>>(
&self,
online: bool,
forward: bool,
skills: &SkillStore,
agents: &CompetitorStore<T, D>,
) -> Rating<T, D> {
let r = &agents[self.agent].rating;
let skill = skills.get(self.agent).unwrap();
let skill = skills.at(self.slot);
if online {
Rating::new(skill.online, r.beta, r.drift)
} else if forward {
Rating::new(skill.forward, r.beta, r.drift)
if forward {
Rating::new(skill.forward, r.beta, r.drift).with_drift_scale(r.drift_scale)
} else {
Rating::new(skill.posterior() / self.likelihood, r.beta, r.drift)
.with_drift_scale(r.drift_scale)
}
}
}
#[derive(Debug)]
#[derive(Clone, Debug)]
struct Team {
items: Vec<Item>,
output: f64,
}
#[derive(Debug)]
#[derive(Clone, Debug)]
pub(crate) struct Event {
teams: Vec<Team>,
evidence: f64,
log_evidence: f64,
weights: Vec<Vec<f64>>,
kind: EventKind,
}
@@ -108,7 +110,6 @@ impl Event {
pub(crate) fn within_priors<T: Time, D: Drift<T>>(
&self,
online: bool,
forward: bool,
skills: &SkillStore,
agents: &CompetitorStore<T, D>,
@@ -118,25 +119,26 @@ impl Event {
.map(|team| {
team.items
.iter()
.map(|item| item.within_prior(online, forward, skills, agents))
.map(|item| item.within_prior(forward, skills, agents))
.collect::<Vec<_>>()
})
.collect::<Vec<_>>()
}
/// Direct in-loop update: mutates self and `skills` inline with no
/// intermediate allocation. Used by both the sequential sweep path and,
/// via unsafe, by the parallel rayon path for events in the same color
/// group (which have disjoint agent sets — see `sweep_color_groups`).
fn iteration_direct<T: Time, D: Drift<T>>(
&mut self,
skills: &mut SkillStore,
/// Run inference for this event and return its per-item likelihoods.
///
/// Reads `skills` immutably and does not touch `self`, so every event in
/// a color group can run concurrently without any aliasing question —
/// the mutation is deferred to `apply`.
fn compute<T: Time, D: Drift<T>>(
&self,
skills: &SkillStore,
agents: &CompetitorStore<T, D>,
p_draw: f64,
convergence: crate::ConvergenceOptions,
arena: &mut ScratchArena,
) {
let teams = self.within_priors(false, false, skills, agents);
) -> EventUpdate {
let teams = self.within_priors(false, skills, agents);
let result = self.outputs();
let g = match self.kind {
EventKind::Ranked => {
@@ -152,17 +154,58 @@ impl Event {
),
};
for (t, team) in self.teams.iter_mut().enumerate() {
for (i, item) in team.items.iter_mut().enumerate() {
let old_likelihood = skills.get(item.agent).unwrap().likelihood;
let new_likelihood = (old_likelihood / item.likelihood) * g.likelihoods[t][i];
skills.get_mut(item.agent).unwrap().likelihood = new_likelihood;
item.likelihood = g.likelihoods[t][i];
EventUpdate {
log_evidence: g.log_evidence,
likelihoods: g.likelihoods,
}
}
self.evidence = g.evidence;
/// Fold a computed update into the skill store and cache it on the items.
fn apply(&mut self, skills: &mut SkillStore, update: EventUpdate) {
for (t, team) in self.teams.iter_mut().enumerate() {
for (i, item) in team.items.iter_mut().enumerate() {
let fresh = update.likelihoods[t][i];
let old_likelihood = skills.at(item.slot).likelihood;
let new_likelihood = (old_likelihood / item.likelihood) * fresh;
skills.at_mut(item.slot).likelihood = new_likelihood;
item.likelihood = fresh;
}
}
self.log_evidence = update.log_evidence;
}
/// Compute and apply in one step — the sequential sweep.
fn iteration_direct<T: Time, D: Drift<T>>(
&mut self,
skills: &mut SkillStore,
agents: &CompetitorStore<T, D>,
p_draw: f64,
convergence: crate::ConvergenceOptions,
arena: &mut ScratchArena,
) {
let update = self.compute(skills, agents, p_draw, convergence, arena);
self.apply(skills, update);
}
}
/// The result of running inference for one event, before it is folded back
/// into the shared skill store.
#[derive(Debug)]
struct EventUpdate {
log_evidence: f64,
likelihoods: Vec<Vec<Gaussian>>,
}
/// One slice's worth of forward-only inference.
///
/// `posteriors` doubles as the outgoing forward message: the scratch sweep
/// never writes `backward`, so it stays `N_INF`, and `Skill::posterior()`
/// and `forward_prior_out` are then the same product.
#[derive(Debug)]
pub(crate) struct FilteredStep {
pub(crate) log_evidence: f64,
pub(crate) posteriors: Vec<(Index, Gaussian)>,
}
#[derive(Debug)]
@@ -174,6 +217,14 @@ pub struct TimeSlice<T: Time = i64> {
pub(crate) convergence: crate::ConvergenceOptions,
arena: ScratchArena,
pub(crate) color_groups: ColorGroups,
/// Whether `color_groups` still reflects `events`.
///
/// Coloring is rebuilt lazily, on the first full sweep after an append,
/// rather than eagerly per append: the partition is thrown away and
/// recomputed wholesale either way, so doing it per append made ingesting
/// n events O(n^2) with no benefit — nothing reads the partition between
/// an append and the next full sweep.
color_groups_dirty: bool,
}
impl<T: Time> TimeSlice<T> {
@@ -186,6 +237,7 @@ impl<T: Time> TimeSlice<T> {
convergence,
arena: ScratchArena::new(),
color_groups: ColorGroups::new(),
color_groups_dirty: false,
}
}
@@ -198,6 +250,7 @@ impl<T: Time> TimeSlice<T> {
let n = self.events.len();
if n == 0 {
self.color_groups = ColorGroups::new();
self.color_groups_dirty = false;
return;
}
@@ -221,13 +274,19 @@ impl<T: Time> TimeSlice<T> {
self.events = reordered;
self.color_groups = ColorGroups { groups: new_groups };
self.color_groups_dirty = false;
debug_assert!(
self.color_groups.groups_are_contiguous(),
"color groups must occupy contiguous event ranges"
);
}
pub fn add_events<D: Drift<T>>(
&mut self,
composition: Vec<Vec<Vec<Index>>>,
results: Vec<Vec<f64>>,
weights: Vec<Vec<Vec<f64>>>,
results: Option<Vec<Vec<f64>>>,
weights: Option<Vec<Vec<Vec<f64>>>>,
kinds: Vec<EventKind>,
agents: &CompetitorStore<T, D>,
) {
@@ -246,21 +305,26 @@ impl<T: Time> TimeSlice<T> {
for idx in this_agent {
let elapsed = compute_elapsed(agents[*idx].last_time.as_ref(), &self.time);
let forward = agents[*idx].receive(&self.time);
if let Some(skill) = self.skills.get_mut(*idx) {
skill.elapsed = elapsed;
skill.forward = agents[*idx].receive(&self.time);
skill.forward = forward;
} else {
self.skills.insert(
*idx,
Skill {
forward: agents[*idx].receive(&self.time),
forward,
backward: N_INF,
likelihood: N_INF,
elapsed,
..Default::default()
},
);
}
}
let skills = &self.skills;
let events = composition.iter().enumerate().map(|(e, event)| {
let teams = event
.iter()
@@ -270,33 +334,37 @@ impl<T: Time> TimeSlice<T> {
.iter()
.map(|&agent| Item {
agent,
// Every participant was inserted into `skills`
// just above, so the slot always resolves.
slot: skills
.slot_of(agent)
.expect("participant must be present in the slice store"),
likelihood: N_INF,
})
.collect::<Vec<_>>();
Team {
items,
output: if results.is_empty() {
(event.len() - (t + 1)) as f64
} else {
results[e][t]
output: match &results {
Some(results) => results[e][t],
// No explicit result: rank by position, first team best.
None => (event.len() - (t + 1)) as f64,
},
}
})
.collect::<Vec<_>>();
let weights = if weights.is_empty() {
teams
let weights = match &weights {
Some(weights) => weights[e].clone(),
None => teams
.iter()
.map(|team| vec![1.0; team.items.len()])
.collect::<Vec<_>>()
} else {
weights[e].clone()
.collect::<Vec<_>>(),
};
Event {
teams,
evidence: 0.0,
log_evidence: 0.0,
weights,
kind: kinds[e],
}
@@ -306,8 +374,9 @@ impl<T: Time> TimeSlice<T> {
self.events.extend(events);
self.color_groups_dirty = true;
self.iteration(from, agents);
self.recompute_color_groups();
}
pub(crate) fn posteriors(&self) -> HashMap<Index, Gaussian> {
@@ -317,11 +386,22 @@ impl<T: Time> TimeSlice<T> {
.collect::<HashMap<_, _>>()
}
/// Sweep this slice's events once, starting at index `from`.
///
/// # Panics
///
/// Panics if an event references a competitor with no entry in this
/// slice's skill store. `add_events` inserts one for every participant, so
/// this cannot happen for slices built through the public API.
pub fn iteration<D: Drift<T>>(&mut self, from: usize, agents: &CompetitorStore<T, D>) {
if from == 0 && self.color_groups_dirty {
self.recompute_color_groups();
}
if from > 0 || self.color_groups.is_empty() {
// Initial pass (add_events) or no color groups yet: simple sequential sweep.
for event in self.events.iter_mut().skip(from) {
let teams = event.within_priors(false, false, &self.skills, agents);
let teams = event.within_priors(false, &self.skills, agents);
let result = event.outputs();
let g = match event.kind {
@@ -345,15 +425,15 @@ impl<T: Time> TimeSlice<T> {
for (t, team) in event.teams.iter_mut().enumerate() {
for (i, item) in team.items.iter_mut().enumerate() {
let old_likelihood = self.skills.get(item.agent).unwrap().likelihood;
let old_likelihood = self.skills.at(item.slot).likelihood;
let new_likelihood =
(old_likelihood / item.likelihood) * g.likelihoods[t][i];
self.skills.get_mut(item.agent).unwrap().likelihood = new_likelihood;
self.skills.at_mut(item.slot).likelihood = new_likelihood;
item.likelihood = g.likelihoods[t][i];
}
}
event.evidence = g.evidence;
event.log_evidence = g.log_evidence;
}
} else {
self.sweep_color_groups(agents);
@@ -363,14 +443,13 @@ impl<T: Time> TimeSlice<T> {
/// Full event sweep using the color-group partition. Colors are processed
/// sequentially; within each color the inner loop is parallel under rayon.
///
/// Events within each color group touch disjoint agent sets (guaranteed by
/// the greedy coloring). This lets each rayon thread write directly to its
/// events' skill likelihoods without a deferred-apply step, matching the
/// sequential path's allocation profile. The unsafe block is sound because:
/// 1. `self.events[range]` and `self.skills` are separate fields → disjoint.
/// 2. Events in the same color group access disjoint `Index` values in
/// `self.skills`, so concurrent writes land on different memory locations.
/// 3. Each event only writes to its own items' likelihoods (no sharing).
/// Events in one color group touch disjoint agent sets, so none of them
/// can observe another's writes. That makes the sweep separable: inference
/// runs concurrently over shared `&self.skills`, and the resulting updates
/// are folded in afterwards in index order. Splitting it this way needs no
/// `unsafe` and no aliasing argument, and it keeps results bit-identical
/// across thread counts because the apply order does not depend on which
/// worker finished first.
#[cfg(feature = "rayon")]
fn sweep_color_groups<D: Drift<T>>(&mut self, agents: &CompetitorStore<T, D>) {
use rayon::prelude::*;
@@ -390,29 +469,28 @@ impl<T: Time> TimeSlice<T> {
if group_len == 0 {
continue;
}
let range = self.color_groups.color_range(color_idx);
let p_draw = self.p_draw;
let convergence = self.convergence;
if group_len >= RAYON_THRESHOLD {
// Obtain a raw pointer from the unique `&mut self.skills` reference.
// Casting back to `&mut` inside the closure is sound because:
// 1. The pointer originates from a `&mut` — no aliasing with shared refs.
// 2. Events in the same color group touch disjoint `Index` slots in the
// underlying Vec, so concurrent writes from different threads land on
// different memory locations — no data race.
// 3. `self.events[range]` and `self.skills` are separate struct fields,
// so the borrow splits cleanly.
let skills_addr: usize = (&mut self.skills as *mut SkillStore) as usize;
self.events[range].par_iter_mut().for_each(move |ev| {
// SAFETY: see above.
let skills: &mut SkillStore = unsafe { &mut *(skills_addr as *mut SkillStore) };
let skills = &self.skills;
let updates: Vec<EventUpdate> = self.events[range.clone()]
.par_iter()
.map(|ev| {
ARENA.with(|cell| {
let mut arena = cell.borrow_mut();
arena.reset();
ev.iteration_direct(skills, agents, p_draw, convergence, &mut arena);
});
});
ev.compute(skills, agents, p_draw, convergence, &mut arena)
})
})
.collect();
for (ev, update) in self.events[range].iter_mut().zip(updates) {
ev.apply(&mut self.skills, update);
}
} else {
for ev in &mut self.events[range] {
ev.iteration_direct(
@@ -454,18 +532,29 @@ impl<T: Time> TimeSlice<T> {
}
}
#[allow(dead_code)]
/// Iterate this slice alone until its posteriors stop moving, returning
/// the number of iterations taken.
///
/// Used by `filtered_step` to drive a scratch copy of the slice, and by
/// tests. Production convergence across slices is driven by
/// `History::converge`, which calls `iteration` directly.
///
/// Honours `self.convergence`; it previously hard-coded an epsilon and a
/// 20-iteration cap that matched neither `ConvergenceOptions` nor the
/// schedule default.
pub(crate) fn iterate_to_convergence<D: Drift<T>>(
&mut self,
agents: &CompetitorStore<T, D>,
) -> usize {
let epsilon = 1e-6;
let iterations = 20;
use crate::{tuple_gt, tuple_max};
let epsilon = self.convergence.epsilon;
let max_iter = self.convergence.max_iter;
let mut step = (f64::INFINITY, f64::INFINITY);
let mut i = 0;
while tuple_gt(step, epsilon) && i < iterations {
while tuple_gt(step, epsilon) && i < max_iter {
let old = self.posteriors();
self.iteration(0, agents);
@@ -477,6 +566,10 @@ impl<T: Time> TimeSlice<T> {
});
i += 1;
if !crate::step_is_finite(step) {
break;
}
}
i
@@ -497,14 +590,13 @@ impl<T: Time> TimeSlice<T> {
n.forget(
agents[*agent]
.rating
.drift
.variance_for_elapsed(skill.elapsed),
.drift_variance_for_elapsed(skill.elapsed),
)
}
pub(crate) fn new_backward_info<D: Drift<T>>(&mut self, agents: &CompetitorStore<T, D>) {
for (agent, skill) in self.skills.iter_mut() {
skill.backward = agents[agent].message;
skill.backward = agents[agent].message.unwrap_or(N_INF);
}
self.iteration(0, agents);
}
@@ -516,21 +608,99 @@ impl<T: Time> TimeSlice<T> {
self.iteration(0, agents);
}
/// Run this slice's events on forward (filtering) information alone.
///
/// `incoming` holds each competitor's forward message out of their
/// previous appearance; a competitor absent from it starts at their
/// configured prior. The sweep runs on a scratch copy, so the real slice
/// is untouched — which is what makes the filtered estimates independent
/// of whether `History::converge` has run.
pub(crate) fn filtered_step<D: Drift<T>>(
&self,
incoming: &HashMap<Index, Gaussian>,
agents: &CompetitorStore<T, D>,
) -> FilteredStep {
let mut scratch = TimeSlice {
events: self.events.clone(),
skills: SkillStore::new(),
time: self.time,
p_draw: self.p_draw,
convergence: self.convergence,
arena: ScratchArena::new(),
color_groups: ColorGroups::new(),
color_groups_dirty: true,
};
for event in &mut scratch.events {
for team in &mut event.teams {
for item in &mut team.items {
item.likelihood = N_INF;
}
}
event.log_evidence = 0.0;
}
for (agent, skill) in self.skills.iter() {
let rating = &agents[agent].rating;
let forward = match incoming.get(&agent) {
Some(message) => message.forget(rating.drift_variance_for_elapsed(skill.elapsed)),
None => rating.prior,
};
let slot = scratch.skills.insert(
agent,
Skill {
forward,
backward: N_INF,
likelihood: N_INF,
elapsed: skill.elapsed,
},
);
// The cloned events carry slots resolved against the REAL store, so
// the scratch must assign the same ones. It does because `iter()`
// yields slot order and `insert` allocates slots in call order —
// but that is a coupling between two types, so pin it here rather
// than leave it to be rediscovered after it breaks.
debug_assert_eq!(
Some(slot),
self.skills.slot_of(agent),
"scratch slot must match the real slice's slot for {agent:?}"
);
}
scratch.iterate_to_convergence(agents);
FilteredStep {
log_evidence: scratch.events.iter().map(|event| event.log_evidence).sum(),
posteriors: scratch
.skills
.iter()
.map(|(agent, skill)| (agent, skill.posterior()))
.collect(),
}
}
pub(crate) fn log_evidence<D: Drift<T>>(
&self,
online: bool,
targets: &[Index],
forward: bool,
agents: &CompetitorStore<T, D>,
) -> f64 {
// Hashed once rather than scanned per player per event, so a
// `log_evidence_for` with many keys is not quadratic.
let target_set: std::collections::HashSet<Index> = targets.iter().copied().collect();
// log_evidence is infrequent; a local arena avoids needing &mut self.
let mut arena = ScratchArena::new();
let run_event = |event: &Event, arena: &mut ScratchArena| -> f64 {
let teams = event.within_priors(online, forward, &self.skills, agents);
let teams = event.within_priors(forward, &self.skills, agents);
let result = event.outputs();
match event.kind {
EventKind::Ranked => Game::ranked_with_arena(
EventKind::Ranked => {
Game::ranked_with_arena(
teams,
&result,
&event.weights,
@@ -538,9 +708,10 @@ impl<T: Time> TimeSlice<T> {
self.convergence,
arena,
)
.evidence
.ln(),
EventKind::Scored { score_sigma } => Game::scored_with_arena(
.log_evidence
}
EventKind::Scored { score_sigma } => {
Game::scored_with_arena(
teams,
&result,
&event.weights,
@@ -548,21 +719,21 @@ impl<T: Time> TimeSlice<T> {
self.convergence,
arena,
)
.evidence
.ln(),
.log_evidence
}
}
};
if targets.is_empty() {
if online || forward {
if forward {
self.events
.iter()
.map(|event| run_event(event, &mut arena))
.sum()
} else {
self.events.iter().map(|event| event.evidence.ln()).sum()
self.events.iter().map(|event| event.log_evidence).sum()
}
} else if online || forward {
} else if forward {
self.events
.iter()
.filter(|event| {
@@ -570,7 +741,7 @@ impl<T: Time> TimeSlice<T> {
.teams
.iter()
.flat_map(|team| &team.items)
.any(|item| targets.contains(&item.agent))
.any(|item| target_set.contains(&item.agent))
})
.map(|event| run_event(event, &mut arena))
.sum()
@@ -582,9 +753,9 @@ impl<T: Time> TimeSlice<T> {
.teams
.iter()
.flat_map(|team| &team.items)
.any(|item| targets.contains(&item.agent))
.any(|item| target_set.contains(&item.agent))
})
.map(|event| event.evidence.ln())
.map(|event| event.log_evidence)
.sum()
}
}
@@ -616,8 +787,113 @@ impl<T: Time> TimeSlice<T> {
}
}
/// Elapsed time from a competitor's previous appearance to `current`.
///
/// A negative elapsed means slices are being visited out of time order, which
/// would make drift *reduce* uncertainty. Release builds clamp to zero so a
/// bad timestamp degrades to "no drift" rather than corrupting the posterior;
/// debug builds trip instead, because reaching here is a bug in slice ordering
/// rather than something callers can cause with ordinary data.
pub(crate) fn compute_elapsed<T: Time>(last: Option<&T>, current: &T) -> i64 {
last.map(|l| l.elapsed_to(current).max(0)).unwrap_or(0)
let Some(last) = last else {
return 0;
};
let elapsed = last.elapsed_to(current);
debug_assert!(
elapsed >= 0,
"negative elapsed ({elapsed}) — slices visited out of time order"
);
elapsed.max(0)
}
impl<T: Time> TimeSlice<T> {
/// Precision matrix of the joint posterior over this slice's competitors.
///
/// Message passing produces per-competitor marginals and throws the
/// correlation away — `Item::likelihood` is already the projection of an
/// event's factor down onto one competitor. So the joint has to be rebuilt
/// from the factor structure rather than recovered from the messages.
///
/// Usefully, a precision matrix depends only on *structure* — who played
/// whom, with what weights and what observation noise — and not on the
/// observed outcomes. The means are already exact (Gaussian belief
/// propagation gets those right even with cycles), so only the second
/// moment needs rebuilding.
///
/// Returns the competitor order and the dense matrix in row-major order.
/// Only scored events contribute their factors exactly; see the caller.
pub(crate) fn joint_precision<D: Drift<T>>(
&self,
agents: &CompetitorStore<T, D>,
) -> (Vec<Index>, Vec<f64>) {
let order: Vec<Index> = self.skills.keys().collect();
let n = order.len();
let mut row_of: HashMap<Index, usize> = HashMap::with_capacity(n);
for (r, idx) in order.iter().enumerate() {
row_of.insert(*idx, r);
}
let mut lambda = vec![0.0; n * n];
// Everything outside this slice enters as each competitor's forward and
// backward messages, which message passing treats as independent.
for (r, idx) in order.iter().enumerate() {
let skill = self.skills.get(*idx).expect("slice key has a skill");
lambda[r * n + r] += (skill.forward * skill.backward).pi();
}
for event in &self.events {
let EventKind::Scored { score_sigma } = event.kind else {
continue;
};
// Teams best-first, matching the diff chain inference builds.
let mut order_idx: Vec<usize> = (0..event.teams.len()).collect();
order_idx.sort_by(|&a, &b| {
event.teams[b]
.output
.partial_cmp(&event.teams[a].output)
.unwrap_or(std::cmp::Ordering::Equal)
});
for pair in order_idx.windows(2) {
let (hi, lo) = (pair[0], pair[1]);
// Contrast vector, and the observation noise that sits on top
// of the skills: per-member performance noise plus the score
// noise itself.
let mut contrast: HashMap<usize, f64> = HashMap::new();
let mut noise = score_sigma * score_sigma;
for (team, sign) in [(hi, 1.0), (lo, -1.0)] {
for (m, item) in event.teams[team].items.iter().enumerate() {
let w = event.weights[team][m];
let beta = agents[item.agent].rating.beta;
noise += w * w * beta * beta;
*contrast.entry(row_of[&item.agent]).or_insert(0.0) += sign * w;
}
}
for (&i, &ci) in &contrast {
for (&j, &cj) in &contrast {
lambda[i * n + j] += ci * cj / noise;
}
}
}
}
(order, lambda)
}
/// True when every event here is scored, so `joint_precision` is exact.
pub(crate) fn all_scored(&self) -> bool {
self.events
.iter()
.all(|e| matches!(e.kind, EventKind::Scored { .. }))
}
}
#[cfg(test)]
@@ -665,8 +941,8 @@ mod tests {
vec![vec![c], vec![d]],
vec![vec![e], vec![f]],
],
vec![vec![1.0, 0.0], vec![0.0, 1.0], vec![1.0, 0.0]],
vec![],
Some(vec![vec![1.0, 0.0], vec![0.0, 1.0], vec![1.0, 0.0]]),
None,
vec![EventKind::Ranked; 3],
&agents,
);
@@ -742,8 +1018,8 @@ mod tests {
vec![vec![a], vec![c]],
vec![vec![b], vec![c]],
],
vec![vec![1.0, 0.0], vec![0.0, 1.0], vec![1.0, 0.0]],
vec![],
Some(vec![vec![1.0, 0.0], vec![0.0, 1.0], vec![1.0, 0.0]]),
None,
vec![EventKind::Ranked; 3],
&agents,
);
@@ -822,8 +1098,8 @@ mod tests {
vec![vec![a], vec![c]],
vec![vec![b], vec![c]],
],
vec![vec![1.0, 0.0], vec![0.0, 1.0], vec![1.0, 0.0]],
vec![],
Some(vec![vec![1.0, 0.0], vec![0.0, 1.0], vec![1.0, 0.0]]),
None,
vec![EventKind::Ranked; 3],
&agents,
);
@@ -854,8 +1130,8 @@ mod tests {
vec![vec![a], vec![c]],
vec![vec![b], vec![c]],
],
vec![vec![1.0, 0.0], vec![0.0, 1.0], vec![1.0, 0.0]],
vec![],
Some(vec![vec![1.0, 0.0], vec![0.0, 1.0], vec![1.0, 0.0]]),
None,
vec![EventKind::Ranked; 3],
&agents,
);
@@ -866,19 +1142,24 @@ mod tests {
let post = time_slice.posteriors();
// These are convergence residuals, not exact values: by symmetry the
// true mean is 25.0 and the iteration approaches it from above. The
// previous expectation of 25.000003 was the residual after the
// hard-coded 20-iteration cap; honouring `ConvergenceOptions` runs to
// 30 and lands nearer the truth.
assert_ulps_eq!(
post[&a],
Gaussian::from_ms(25.000003, 3.880150),
Gaussian::from_ms(25.000001, 3.880150),
epsilon = 1e-6
);
assert_ulps_eq!(
post[&b],
Gaussian::from_ms(25.000003, 3.880150),
Gaussian::from_ms(25.000001, 3.880150),
epsilon = 1e-6
);
assert_ulps_eq!(
post[&c],
Gaussian::from_ms(25.000003, 3.880150),
Gaussian::from_ms(25.000001, 3.880150),
epsilon = 1e-6
);
}
@@ -920,8 +1201,8 @@ mod tests {
vec![vec![c], vec![d]],
vec![vec![a], vec![c]],
],
vec![vec![1.0, 0.0], vec![1.0, 0.0], vec![1.0, 0.0]],
vec![],
Some(vec![vec![1.0, 0.0], vec![1.0, 0.0], vec![1.0, 0.0]]),
None,
vec![EventKind::Ranked; 3],
&agents,
);
+142
View File
@@ -0,0 +1,142 @@
//! What an additive model does to uncertainty, and why "add the marginals" is
//! unsafe in one direction and merely wasteful in the other.
//!
//! Structurally this is the shape a joint player/layout model takes: every
//! observation measures a *sum* of nodes against a reference, so the data pins
//! differences and leaves the overall level to the prior. That is the classic
//! rating-scale indeterminacy, not a defect.
//!
//! The consequence for a consumer is that combining marginals is wrong in
//! opposite directions depending on the combination, which is worth pinning
//! because the unsafe direction is not the one you would guess:
//!
//! - **Differences** (`a - b`): the shared level cancels, so the exact width is
//! small — and adding marginals lands within a couple of percent of it here,
//! because the loopy underestimate offsets the ignored correlation.
//! - **Sums** (`a + b`): the shared level does *not* cancel, so the exact width
//! is large, and adding marginals is roughly five times too narrow. That is
//! overconfident, and it is the direction that publishes a claim the data
//! does not support.
use smallvec::smallvec;
use trueskill_tt::{ConstantDrift, ConvergenceOptions, Event, History, Member, Outcome, Team};
#[test]
fn additive_structure_makes_sums_wide_and_differences_tight() {
// Structurally like ustat: every round is (player + hole) measured against
// a fixed reference. Only SUMS are pinned by the data; the split between
// player and hole is pinned only by the prior.
let players = ["p0", "p1", "p2"];
let holes = ["h0", "h1"];
let mut h: History<i64, _, _, &'static str> = History::builder()
.mu(0.0)
.sigma(6.0)
.beta(1.0)
.score_sigma(2.0)
.drift(ConstantDrift(0.0))
.convergence(ConvergenceOptions {
max_iter: 20_000,
epsilon: 1e-12,
alpha: 1.0,
})
.build();
let mut seed = 3u64;
let mut rnd = move || {
seed ^= seed << 13;
seed ^= seed >> 7;
seed ^= seed << 17;
seed
};
// true skills, so we know what the data encodes
let truth_p = [2.0, 0.0, -2.0];
let truth_h = [1.0, -1.0];
let mut events = Vec::new();
for _ in 0..60 {
let p = (rnd() as usize) % 3;
let q = (rnd() as usize) % 2;
let noise = ((rnd() % 1000) as f64 / 1000.0 - 0.5) * 2.0;
let score = truth_p[p] + truth_h[q] + noise;
events.push(Event {
time: 1,
teams: smallvec![
Team::with_members([Member::new(players[p]), Member::new(holes[q])]),
Team::with_members([Member::new("reference")]),
],
outcome: Outcome::scores([score, 0.0]),
});
}
h.add_events(events).unwrap();
let r = h.converge().unwrap();
assert!(r.converged, "{:?}", r.final_step);
println!("\n== marginals (what current_skill reports) ==");
for k in players.iter().chain(holes.iter()) {
let g = h.current_skill(k).unwrap();
println!(" {k}: mu {:>8.4} sigma {:>8.4}", g.mu(), g.sigma());
}
println!("\n== the same nodes via posterior_of (exact marginal) ==");
for k in players.iter().chain(holes.iter()) {
let g = h.posterior_of(&[(k, 1.0)]).unwrap();
println!(" {k}: mu {:>8.4} sigma {:>8.4}", g.mu(), g.sigma());
}
println!("\n== combinations the data actually pins ==");
for (label, terms) in [
("p0 + h0 (a round)", vec![(&"p0", 1.0), (&"h0", 1.0)]),
(
"p0 - p1 (rank two players)",
vec![(&"p0", 1.0), (&"p1", -1.0)],
),
("p0 - p2", vec![(&"p0", 1.0), (&"p2", -1.0)]),
(
"h0 - h1 (rank two holes)",
vec![(&"h0", 1.0), (&"h1", -1.0)],
),
] {
let joint = h.posterior_of(&terms).unwrap();
// what a consumer gets today by adding marginals
let naive: f64 = terms
.iter()
.map(|(k, c)| c * c * h.current_skill(*k).unwrap().sigma().powi(2))
.sum::<f64>()
.sqrt();
println!(
" {label:<28} exact sigma {:>7.4} adding marginals {:>7.4} {:>5.2}x over",
joint.sigma(),
naive,
naive / joint.sigma()
);
let ratio = naive / joint.sigma();
if label.contains('+') {
assert!(
ratio < 0.5,
"{label}: adding marginals should be badly OVERconfident for a \
sum, got {ratio:.3}x"
);
} else {
assert!(
(0.8..1.25).contains(&ratio),
"{label}: adding marginals happens to be close for a difference, \
got {ratio:.3}x"
);
}
}
// A single node in an additive model is weakly identified: its exact
// posterior is far wider than message passing reports, because the level it
// shares with its partners is pinned only by the prior.
for k in players.iter().chain(holes.iter()) {
let bp = h.current_skill(k).unwrap().sigma();
let exact = h.posterior_of(&[(k, 1.0)]).unwrap().sigma();
assert!(
exact > 3.0 * bp,
"{k}: exact marginal {exact} should be much wider than the reported \
{bp} in an additive model"
);
}
}
+59 -12
View File
@@ -65,7 +65,7 @@ fn add_events_draw() {
outcome: Outcome::draw(2),
}];
h.add_events(events).unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
}
#[test]
@@ -123,7 +123,7 @@ fn fluent_event_builder_winner_convenience() {
.winner(0)
.commit()
.unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
}
#[test]
@@ -141,7 +141,7 @@ fn fluent_event_builder_draw() {
.draw()
.commit()
.unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
}
#[test]
@@ -155,7 +155,7 @@ fn current_skill_and_learning_curve() {
.build();
h.record_winner(&"a", &"b", 1).unwrap();
h.record_winner(&"a", &"b", 2).unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
let a = h.current_skill(&"a").unwrap();
assert!(a.mu() > 25.0);
@@ -201,9 +201,9 @@ fn predict_quality_two_teams() {
.p_draw(0.0)
.build();
h.record_winner(&"a", &"b", 1).unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
let q = h.predict_quality(&[&[&"a"], &[&"b"]]);
let q = h.predict_quality(&[&[&"a"], &[&"b"]]).unwrap();
assert!(q > 0.0 && q <= 1.0);
}
@@ -217,12 +217,16 @@ fn predict_outcome_two_teams_sums_to_one() {
.p_draw(0.0)
.build();
h.record_winner(&"a", &"b", 1).unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
let p = h.predict_outcome(&[&[&"a"], &[&"b"]]);
assert_eq!(p.len(), 2);
assert!((p[0] + p[1] - 1.0).abs() < 1e-9);
assert!(p[0] > p[1]);
let p = h.predict_outcome(&[&[&"a"], &[&"b"]]).unwrap();
let wins = p.win_probabilities();
assert_eq!(wins.len(), 2);
// With p_draw == 0 there is no draw outcome, so the two win
// probabilities are the whole space.
assert!((p.total() - 1.0).abs() < 1e-9, "total = {}", p.total());
assert!((wins[0] + wins[1] - 1.0).abs() < 1e-9);
assert!(wins[0] > wins[1]);
}
#[test]
@@ -241,9 +245,52 @@ fn fluent_event_builder_scores() {
.scores([12.0, 4.0])
.commit()
.unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
let a = h.current_skill(&"alice").unwrap();
let b = h.current_skill(&"bob").unwrap();
assert!(a.mu() > b.mu());
}
/// Every field of `ConvergenceReport` must carry real information.
///
/// `slices_skipped` was public, hardcoded to `0`, and reported a plausible
/// value for a feature that never existed — the same shape as the inert
/// `online` flag in #19. It was removed in #33. This pins the remaining fields
/// so the next always-constant member has to survive an assertion rather than
/// just a reviewer's attention.
#[test]
fn every_convergence_report_field_is_populated() {
let mut h = History::builder().build();
for time in 1..=6i64 {
h.record_winner(&"a", &"b", time).unwrap();
}
let report = h.converge().unwrap();
assert!(
report.iterations > 0,
"iterations is zero on a real converge"
);
assert!(report.converged, "fixture must converge");
assert!(
report.final_step.0.is_finite() && report.final_step.1.is_finite(),
"final_step is not finite: {:?}",
report.final_step
);
assert!(
report.log_evidence.is_finite() && report.log_evidence < 0.0,
"log_evidence is not a finite negative log probability: {}",
report.log_evidence
);
assert_eq!(
report.per_iteration_time.len(),
report.iterations,
"per_iteration_time must carry one duration per iteration"
);
}
+38
View File
@@ -0,0 +1,38 @@
//! Helpers shared across the integration suites.
//!
//! Each integration file is its own binary, so `mod common;` compiles a copy
//! per suite. Anything unused in a given suite would warn, hence the
//! `#![allow(dead_code)]`.
#![allow(dead_code)]
use trueskill_tt::Gaussian;
/// A posterior must be finite with a strictly positive sigma.
///
/// A non-finite posterior is the failure mode this crate is most prone to —
/// EP breaking down produces NaN rather than an error — and a zero or negative
/// sigma means the precision went non-positive, which `Gaussian::sigma` reports
/// as improper rather than trapping.
pub fn assert_finite(g: Gaussian, what: &str) {
assert!(
g.mu().is_finite(),
"{what}: mu is not finite (mu={}, sigma={})",
g.mu(),
g.sigma()
);
assert!(
g.sigma().is_finite() && g.sigma() > 0.0,
"{what}: sigma must be finite and positive (mu={}, sigma={})",
g.mu(),
g.sigma()
);
}
/// Every point on every learning curve must be finite.
pub fn assert_curve_finite(curve: &[(i64, Gaussian)], who: &str) {
for (time, g) in curve {
assert_finite(*g, &format!("{who} at t={time}"));
}
}
+222
View File
@@ -0,0 +1,222 @@
//! `Member::with_prior` / `with_drift_scale` — competitor configuration.
//!
//! Both were previously consumed only on the branch that *creates* a
//! competitor, so configuration supplied for a key the history already knew was
//! dropped with no error. `with_prior` had no coverage in this directory at
//! all, which is how that survived.
use smallvec::smallvec;
use trueskill_tt::{
ConvergenceOptions, Event, Gaussian, History, InferenceError, Member, Outcome, Team,
};
const CONVERGENCE: ConvergenceOptions = ConvergenceOptions {
max_iter: 2_000,
epsilon: 1e-12,
alpha: 1.0,
};
fn history() -> History {
History::builder()
.mu(25.0)
.sigma(25.0 / 3.0)
.beta(25.0 / 6.0)
.p_draw(0.0)
.convergence(CONVERGENCE)
.build()
}
/// One event, optionally configuring `a`.
fn bout(
a: &'static str,
b: &'static str,
time: i64,
prior: Option<Gaussian>,
scale: Option<f64>,
) -> Event<i64, &'static str> {
let mut member = Member::new(a);
if let Some(p) = prior {
member = member.with_prior(p);
}
if let Some(s) = scale {
member = member.with_drift_scale(s);
}
Event {
time,
teams: smallvec![
Team::with_members([member]),
Team::with_members([Member::new(b)]),
],
outcome: Outcome::winner(0, 2),
}
}
fn skill_of(h: &History, key: &str) -> Gaussian {
h.current_skill(&key).expect("key in history")
}
/// Baseline: the mechanism works at all on a competitor's first appearance.
#[test]
fn a_prior_applies_to_a_new_competitor() {
let seeded = Gaussian::from_ms(40.0, 1.0);
let mut with = history();
with.add_events(vec![bout("a", "b", 0, Some(seeded), None)])
.unwrap();
let _ = with.converge().unwrap();
let mut without = history();
without
.add_events(vec![bout("a", "b", 0, None, None)])
.unwrap();
let _ = without.converge().unwrap();
assert!(
(skill_of(&with, "a").mu() - skill_of(&without, "a").mu()).abs() > 1.0,
"a seeded prior should move the fit"
);
}
/// The defect in #10: a prior supplied for a competitor the history already
/// knows was silently discarded, and the caller got output computed from the
/// default prior with no indication anything had been dropped.
#[test]
fn a_prior_applies_to_a_competitor_the_history_already_knows() {
let seeded = Gaussian::from_ms(40.0, 1.0);
let mut late = history();
late.add_events(vec![bout("a", "b", 0, None, None)])
.unwrap();
// "a" now exists. Configuring it here used to do nothing whatsoever.
late.add_events(vec![bout("a", "b", 1, Some(seeded), None)])
.unwrap();
let _ = late.converge().unwrap();
let mut never = history();
never
.add_events(vec![
bout("a", "b", 0, None, None),
bout("a", "b", 1, None, None),
])
.unwrap();
let _ = never.converge().unwrap();
assert!(
(skill_of(&late, "a").mu() - skill_of(&never, "a").mu()).abs() > 1.0,
"a late prior must not be silently dropped: {} vs {}",
skill_of(&late, "a").mu(),
skill_of(&never, "a").mu()
);
}
/// Configuration is competitor-scoped, not event-scoped, and `converge` refits
/// from competitor state — so seeding late reaches the same fit as seeding from
/// the start. This is the documented scope, asserted rather than assumed.
#[test]
fn a_prior_is_whole_history_scoped_not_per_event() {
let seeded = Gaussian::from_ms(40.0, 1.0);
let mut late = history();
late.add_events(vec![bout("a", "b", 0, None, None)])
.unwrap();
late.add_events(vec![bout("a", "b", 1, Some(seeded), None)])
.unwrap();
let _ = late.converge().unwrap();
let mut early = history();
early
.add_events(vec![
bout("a", "b", 0, Some(seeded), None),
bout("a", "b", 1, Some(seeded), None),
])
.unwrap();
let _ = early.converge().unwrap();
let (l, e) = (skill_of(&late, "a"), skill_of(&early, "a"));
assert!(
(l.mu() - e.mu()).abs() < 1e-9 && (l.sigma() - e.sigma()).abs() < 1e-9,
"late seeding should refit the whole history: {l:?} vs {e:?}"
);
}
#[test]
fn repeating_the_same_prior_is_inert() {
let seeded = Gaussian::from_ms(40.0, 1.0);
let mut once = history();
once.add_events(vec![
bout("a", "b", 0, Some(seeded), None),
bout("a", "b", 1, None, None),
])
.unwrap();
let _ = once.converge().unwrap();
let mut every_time = history();
every_time
.add_events(vec![
bout("a", "b", 0, Some(seeded), None),
bout("a", "b", 1, Some(seeded), None),
])
.unwrap();
let _ = every_time.converge().unwrap();
let (o, e) = (skill_of(&once, "a"), skill_of(&every_time, "a"));
assert!(
(o.mu() - e.mu()).abs() < 1e-12 && (o.sigma() - e.sigma()).abs() < 1e-12,
"declaring the same prior repeatedly changed the fit: {o:?} vs {e:?}"
);
}
/// Events within a batch have no order, so two different values for one
/// competitor have no well-defined winner. Rejecting is what keeps the answer
/// independent of iteration order.
#[test]
fn a_batch_declaring_two_different_priors_is_rejected() {
let mut h = history();
let err = h
.add_events(vec![
bout("a", "b", 0, Some(Gaussian::from_ms(40.0, 1.0)), None),
bout("a", "b", 1, Some(Gaussian::from_ms(10.0, 1.0)), None),
])
.expect_err("two different priors for one competitor in one batch");
assert!(
matches!(
err,
InferenceError::ConflictingCompetitorConfig { field: "prior", .. }
),
"got {err:?}"
);
}
/// A member setting only `drift_scale` must not also assert the default prior,
/// or it would silently undo a prior seeded earlier. This is why the collected
/// configuration tracks each field separately rather than a merged `Rating`.
#[test]
fn setting_one_field_late_leaves_the_other_alone() {
let seeded = Gaussian::from_ms(40.0, 1.0);
let mut h = history();
h.add_events(vec![bout("a", "b", 0, Some(seeded), None)])
.unwrap();
// Only the scale this time — the prior above must survive.
h.add_events(vec![bout("a", "b", 1, None, Some(0.5))])
.unwrap();
let _ = h.converge().unwrap();
let mut both_upfront = history();
both_upfront
.add_events(vec![
bout("a", "b", 0, Some(seeded), Some(0.5)),
bout("a", "b", 1, None, None),
])
.unwrap();
let _ = both_upfront.converge().unwrap();
let (a, b) = (skill_of(&h, "a"), skill_of(&both_upfront, "a"));
assert!(
(a.mu() - b.mu()).abs() < 1e-9 && (a.sigma() - b.sigma()).abs() < 1e-9,
"setting drift_scale late clobbered the earlier prior: {a:?} vs {b:?}"
);
}
+423
View File
@@ -0,0 +1,423 @@
//! Degenerate, boundary, and error-path coverage.
//!
//! These run in both debug and release: the defects they pin were all
//! guarded only by `debug_assert!`, so a debug-only suite never saw them.
mod common;
use common::assert_finite;
use trueskill_tt::{
ConstantDrift, ConvergenceOptions, Game, GameOptions, Gaussian, History, InferenceError,
NullObserver, Outcome, Rating,
};
type R = Rating<i64, ConstantDrift>;
fn rating() -> R {
R::new(
Gaussian::from_ms(25.0, 25.0 / 3.0),
25.0 / 6.0,
ConstantDrift(25.0 / 300.0),
)
}
#[test]
fn record_draw_without_draw_probability_is_rejected() {
let mut h = History::default();
let err = h.record_draw(&"a", &"b", 1).unwrap_err();
assert!(matches!(
err,
InferenceError::TieWithoutDrawProbability { .. }
));
}
#[test]
fn builder_draw_without_draw_probability_is_rejected() {
let mut h = History::default();
let err = h
.event(1)
.team(["a"])
.team(["b"])
.draw()
.commit()
.unwrap_err();
assert!(matches!(
err,
InferenceError::TieWithoutDrawProbability { .. }
));
}
#[test]
fn draw_with_positive_draw_probability_is_finite() {
let mut h = History::builder().p_draw(0.25).build();
h.record_draw(&"a", &"b", 1).unwrap();
let report = h.converge().unwrap();
assert_finite(h.current_skill("a").unwrap(), "drawn competitor skill");
assert_finite(h.current_skill("b").unwrap(), "drawn competitor skill");
assert!(report.log_evidence.is_finite());
assert!(report.converged);
}
#[test]
fn game_ranked_rejects_tie_without_draw_probability() {
let a = [rating()];
let b = [rating()];
let teams: Vec<&[R]> = vec![&a, &b];
let err = Game::ranked(&teams, Outcome::draw(2), &GameOptions::default()).unwrap_err();
assert!(matches!(
err,
InferenceError::TieWithoutDrawProbability { .. }
));
}
/// `Outcome::winner(w, n)` ties every loser, so any n >= 3 free-for-all hits
/// the tie path even though the caller never asked for a draw.
#[test]
fn winner_of_three_or_more_requires_draw_probability() {
let a = [rating()];
let b = [rating()];
let c = [rating()];
let teams: Vec<&[R]> = vec![&a, &b, &c];
let err = Game::ranked(&teams, Outcome::winner(0, 3), &GameOptions::default()).unwrap_err();
assert!(matches!(
err,
InferenceError::TieWithoutDrawProbability { .. }
));
let opts = GameOptions {
p_draw: 0.1,
..GameOptions::default()
};
let game = Game::ranked(&teams, Outcome::winner(0, 3), &opts).unwrap();
for team in game.posteriors() {
for skill in team {
assert_finite(skill, "3-team winner posterior");
}
}
}
#[test]
fn full_ranking_without_ties_needs_no_draw_probability() {
let a = [rating()];
let b = [rating()];
let c = [rating()];
let teams: Vec<&[R]> = vec![&a, &b, &c];
let game = Game::ranked(&teams, Outcome::ranking([0, 1, 2]), &GameOptions::default()).unwrap();
for team in game.posteriors() {
for skill in team {
assert_finite(skill, "strict ranking posterior");
}
}
}
#[test]
fn empty_history_converges_trivially() {
let mut h = History::default();
let report = h.converge().unwrap();
assert_eq!(report.iterations, 0);
assert!(report.converged);
}
/// Issue #27's exact reproduction: a non-default key type reaching `converge`
/// with no events at all. The underflow it reported trapped in debug and
/// indexed out of bounds in release, so this must run in both profiles.
#[test]
fn converge_on_an_empty_history_with_owned_keys() {
let mut history: History<i64, ConstantDrift, NullObserver, String> =
History::builder_with_key().score_sigma(5.0).build();
let report = history.converge().unwrap();
assert_eq!(report.iterations, 0);
assert!(report.converged);
}
/// A weights/team length mismatch used to be a `debug_assert!`, so release
/// builds ingested the event with the weights silently unapplied. This file's
/// CI job runs in release too, which is the point of pinning it here.
#[test]
fn event_builder_rejects_a_weights_length_mismatch() {
let mut h = History::default();
let err = h
.event(1)
.team(["a"])
.weights([1.0, 2.0])
.team(["b"])
.winner(0)
.commit()
.unwrap_err();
assert!(
matches!(
err,
InferenceError::MismatchedShape {
kind: "weights",
expected: 1,
got: 2,
}
),
"expected a weights MismatchedShape, got {err:?}"
);
}
/// The mismatch must not be applied even partially — a half-weighted team
/// reaching the history would be worse than the error.
#[test]
fn event_builder_weights_mismatch_leaves_the_history_untouched() {
let mut h = History::default();
// Two teams, so ingestion would otherwise succeed — a one-team event is
// rejected for an unrelated reason and would pass this vacuously.
let _ = h
.event(1)
.team(["a"])
.weights([1.0, 2.0])
.team(["b"])
.winner(0)
.commit();
assert!(h.learning_curve("a").is_empty());
}
#[test]
fn empty_event_stream_then_converge() {
let mut h = History::default();
h.add_events(std::iter::empty()).unwrap();
let report = h.converge().unwrap();
assert_eq!(report.iterations, 0);
}
#[test]
fn empty_history_queries_do_not_panic() {
let h = History::default();
assert!(h.learning_curves().is_empty());
assert!(h.learning_curve("nobody").is_empty());
assert!(h.current_skill("nobody").is_none());
}
#[test]
fn single_event_history_converges() {
let mut h = History::default();
h.record_winner(&"a", &"b", 1).unwrap();
let report = h.converge().unwrap();
assert!(report.converged);
assert_finite(h.current_skill("a").unwrap(), "single-event skill");
}
#[test]
fn scored_event_rejects_non_positive_sigma() {
let mut h = History::builder().score_sigma(2.0).build();
let err = h
.event(1)
.team(["a"])
.team(["b"])
.scores_with_sigma([3.0, 1.0], f64::NAN)
.commit()
.unwrap_err();
assert!(matches!(
err,
InferenceError::InvalidParameter {
name: "score_sigma",
..
}
));
}
#[test]
fn convergence_reports_are_finite_across_many_teams() {
let opts = GameOptions {
p_draw: 0.1,
convergence: ConvergenceOptions::default(),
..GameOptions::default()
};
let holders: Vec<[R; 1]> = (0..12).map(|_| [rating()]).collect();
let teams: Vec<&[R]> = holders.iter().map(|t| t.as_slice()).collect();
let game = Game::ranked(&teams, Outcome::ranking(0..12), &opts).unwrap();
assert!(
game.log_evidence().is_finite(),
"12-team log-evidence must be finite, got {}",
game.log_evidence()
);
for team in game.posteriors() {
for skill in team {
assert_finite(skill, "12-team posterior");
}
}
}
/// A long diff chain underflows a linear evidence product: each link
/// contributes a probability in (0, 1], so ~1000 links flush the product to
/// exactly 0.0 and `ln(0.0)` is `-inf`. Accumulating in log space keeps it
/// finite.
#[test]
fn log_evidence_survives_a_long_diff_chain() {
let holders: Vec<[R; 1]> = (0..1200).map(|_| [rating()]).collect();
let teams: Vec<&[R]> = holders.iter().map(|t| t.as_slice()).collect();
let game = Game::ranked(
&teams,
Outcome::ranking(0..holders.len() as u32),
&GameOptions::default(),
)
.unwrap();
let log_evidence = game.log_evidence();
assert!(
log_evidence.is_finite(),
"1200-team log-evidence must be finite, got {log_evidence}"
);
assert!(
log_evidence < 0.0,
"log-evidence of a probability must be negative, got {log_evidence}"
);
}
/// A near-certain outcome rounds the losing tail to exactly zero in the
/// `erfc` approximation; the evidence floor keeps `ln` finite.
#[test]
fn log_evidence_finite_for_near_certain_outcome() {
let overwhelming = R::new(Gaussian::from_ms(5_000.0, 0.5), 1.0, ConstantDrift(0.0));
let hopeless = R::new(Gaussian::from_ms(-5_000.0, 0.5), 1.0, ConstantDrift(0.0));
let a = [overwhelming];
let b = [hopeless];
let teams: Vec<&[R]> = vec![&a, &b];
let game = Game::ranked(&teams, Outcome::winner(0, 2), &GameOptions::default()).unwrap();
assert!(
game.log_evidence().is_finite(),
"got {}",
game.log_evidence()
);
// And the reverse — a colossal upset — must also stay finite.
let upset = Game::ranked(&teams, Outcome::winner(1, 2), &GameOptions::default()).unwrap();
assert!(
upset.log_evidence().is_finite(),
"upset log-evidence must be finite, got {}",
upset.log_evidence()
);
}
#[test]
fn empty_history_has_no_filtered_estimates() {
let history: History = History::builder().build();
assert_eq!(history.filtered_log_evidence(), 0.0);
assert!(history.filtered_learning_curves().is_empty());
assert!(history.filtered_learning_curve("nobody").is_empty());
}
// --- Boundary inputs (#26) ----------------------------------------------
fn tight() -> ConvergenceOptions {
ConvergenceOptions {
max_iter: 2_000,
epsilon: 1e-12,
..ConvergenceOptions::default()
}
}
fn assert_curve_finite(h: &History, keys: &[&str], what: &str) {
for key in keys {
for (time, g) in h.learning_curve(*key) {
assert!(
g.mu().is_finite() && g.sigma().is_finite(),
"{what}: non-finite posterior for {key} at t={time} (mu={} sigma={})",
g.mu(),
g.sigma()
);
}
}
}
/// A zero weight reaches `(m - performance.exclude(..)) * (1.0 / w)`, i.e. a
/// division by zero. The commit is accepted today, so this pins that the
/// resulting posterior is still finite rather than quietly NaN.
#[test]
fn zero_weight_does_not_produce_a_non_finite_posterior() {
let mut h = History::builder().build();
h.event(1)
.team(["a"])
.weights([0.0])
.team(["b"])
.winner(0)
.commit()
.expect("a zero weight is accepted today; update this test if that changes");
let _ = h.converge().unwrap();
assert_curve_finite(&h, &["a", "b"], "zero weight");
}
#[test]
fn negative_weight_does_not_produce_a_non_finite_posterior() {
let mut h = History::builder().build();
h.event(1)
.team(["a"])
.weights([-1.0])
.team(["b"])
.winner(0)
.commit()
.expect("a negative weight is accepted today; update this test if that changes");
let _ = h.converge().unwrap();
assert_curve_finite(&h, &["a", "b"], "negative weight");
}
/// Events supplied newest-first must land in the same slices as oldest-first:
/// ingestion sorts by time rather than trusting arrival order.
#[test]
fn out_of_order_timestamps_converge_to_the_same_answer() {
fn build(descending: bool) -> History {
let mut h = History::builder().convergence(tight()).build();
let mut times: Vec<i64> = (1..=6).collect();
if descending {
times.reverse();
}
for time in times {
h.record_winner(&"a", &"b", time).unwrap();
}
let _ = h.converge().unwrap();
h
}
let ascending = build(false);
let descending = build(true);
let one = ascending.current_skill("a").unwrap();
let other = descending.current_skill("a").unwrap();
assert!(
(one.mu() - other.mu()).abs() < 1e-8 && (one.sigma() - other.sigma()).abs() < 1e-8,
"arrival order changed the answer: ascending mu={} sigma={}, descending mu={} sigma={}",
one.mu(),
one.sigma(),
other.mu(),
other.sigma()
);
}
#[test]
fn extreme_beta_and_sigma_stay_finite() {
for (beta, sigma) in [(1e-6, 1e-6), (1e6, 1e6), (1e-6, 1e6), (1e6, 1e-6)] {
let mut h = History::builder().beta(beta).sigma(sigma).build();
h.record_winner(&"a", &"b", 1).unwrap();
h.record_winner(&"a", &"b", 2).unwrap();
let _ = h.converge().unwrap();
assert_curve_finite(&h, &["a", "b"], &format!("beta={beta} sigma={sigma}"));
}
}
+1 -1
View File
@@ -47,7 +47,7 @@ fn build_and_converge(seed: u64) -> Vec<(i64, trueskill_tt::Gaussian)> {
});
}
h.add_events(events).unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
// Sample one competitor's curve for the comparison.
h.learning_curve("p0")
}
+499
View File
@@ -0,0 +1,499 @@
//! Per-competitor drift scaling via `Member::with_drift_scale`.
//!
//! The scale multiplies the *variance* the history's `Drift` contributes for
//! that competitor, so `scale` is in the same units as `gamma`:
//! `ConstantDrift(g)` at `scale = s` behaves as `ConstantDrift(g * s)` would.
//! `scale = 0.0` pins a competitor still — an anchor, a rating floor, a course
//! difficulty — while everyone around them keeps drifting.
use smallvec::smallvec;
use trueskill_tt::{
ConstantDrift, ConvergenceOptions, Event, Gaussian, History, InferenceError, Member,
NullObserver, Outcome, Team,
};
type Fit = History<i64, ConstantDrift, NullObserver, &'static str>;
const CONVERGENCE: ConvergenceOptions = ConvergenceOptions {
max_iter: 64,
epsilon: 1e-9,
alpha: 1.0,
};
/// Two events separated by a long gap, so drift has room to matter.
fn distant_pair(anchor_scale: Option<f64>) -> Vec<Event<i64, &'static str>> {
let anchor = |s: Option<f64>| match s {
Some(scale) => Member::new("anchor").with_drift_scale(scale),
None => Member::new("anchor"),
};
vec![
Event {
time: 0,
teams: smallvec![
Team::with_members([anchor(anchor_scale)]),
Team::with_members([Member::new("player")]),
],
outcome: Outcome::winner(0, 2),
},
Event {
time: 1000,
teams: smallvec![
Team::with_members([anchor(anchor_scale)]),
Team::with_members([Member::new("player")]),
],
outcome: Outcome::winner(1, 2),
},
]
}
fn fit(events: Vec<Event<i64, &'static str>>, gamma: f64) -> Fit {
let mut h = History::builder()
.mu(25.0)
.sigma(25.0 / 3.0)
.beta(25.0 / 6.0)
.p_draw(0.0)
.drift(ConstantDrift(gamma))
.convergence(CONVERGENCE)
.build();
h.add_events(events).unwrap();
let _ = h.converge().unwrap();
h
}
fn curve(h: &Fit, key: &str) -> Vec<(i64, Gaussian)> {
let mut c = h.learning_curves().remove(key).expect("key in curves");
c.sort_by_key(|(t, _)| *t);
c
}
/// A competitor at `scale = 0.0` is one latent skill observed twice, so the
/// posterior is the same distribution at both times — and strictly tighter
/// than the same competitor left to drift.
#[test]
fn zero_scale_pins_a_competitor_still() {
let pinned = fit(distant_pair(Some(0.0)), 25.0 / 300.0);
let drifting = fit(distant_pair(None), 25.0 / 300.0);
let pinned_curve = curve(&pinned, "anchor");
assert_eq!(pinned_curve.len(), 2);
let (t0, first) = pinned_curve[0];
let (t1, second) = pinned_curve[1];
assert_eq!((t0, t1), (0, 1000));
assert!(
(first.sigma() - second.sigma()).abs() < 1e-9,
"a pinned competitor's uncertainty must not move between t=0 and t=1000: \
{} vs {}",
first.sigma(),
second.sigma()
);
assert!(
(first.mu() - second.mu()).abs() < 1e-9,
"a pinned competitor's mean must not move: {} vs {}",
first.mu(),
second.mu()
);
let drifting_curve = curve(&drifting, "anchor");
assert!(
drifting_curve[0].1.sigma() > first.sigma() + 1e-6,
"drift must leave the anchor less certain than pinning does: {} vs {}",
drifting_curve[0].1.sigma(),
first.sigma()
);
}
/// The scale is composable with `gamma`: scaling every competitor by `s` is
/// exactly the same fit as scaling the history's drift by `s`.
#[test]
fn scale_is_equivalent_to_scaling_gamma() {
let scaled: Vec<Event<i64, &'static str>> = vec![
Event {
time: 0,
teams: smallvec![
Team::with_members([Member::new("a").with_drift_scale(0.5)]),
Team::with_members([Member::new("b").with_drift_scale(0.5)]),
],
outcome: Outcome::winner(0, 2),
},
Event {
time: 400,
teams: smallvec![
Team::with_members([Member::new("b").with_drift_scale(0.5)]),
Team::with_members([Member::new("a").with_drift_scale(0.5)]),
],
outcome: Outcome::winner(0, 2),
},
];
let plain: Vec<Event<i64, &'static str>> = vec![
Event {
time: 0,
teams: smallvec![
Team::with_members([Member::new("a")]),
Team::with_members([Member::new("b")]),
],
outcome: Outcome::winner(0, 2),
},
Event {
time: 400,
teams: smallvec![
Team::with_members([Member::new("b")]),
Team::with_members([Member::new("a")]),
],
outcome: Outcome::winner(0, 2),
},
];
let by_scale = fit(scaled, 0.3);
let by_gamma = fit(plain, 0.15);
for key in ["a", "b"] {
let lhs = curve(&by_scale, key);
let rhs = curve(&by_gamma, key);
assert_eq!(lhs.len(), rhs.len());
for ((t_l, g_l), (t_r, g_r)) in lhs.iter().zip(rhs.iter()) {
assert_eq!(t_l, t_r);
assert!(
(g_l.mu() - g_r.mu()).abs() < 1e-9 && (g_l.sigma() - g_r.sigma()).abs() < 1e-9,
"ConstantDrift(0.3) at scale 0.5 must equal ConstantDrift(0.15) for {key} at \
t={t_l}: ({}, {}) vs ({}, {})",
g_l.mu(),
g_l.sigma(),
g_r.mu(),
g_r.sigma()
);
}
}
}
/// `None` means 1.0: an explicit unit scale changes nothing.
#[test]
fn unset_scale_matches_an_explicit_unit_scale() {
let implicit = fit(distant_pair(None), 25.0 / 300.0);
let explicit = fit(distant_pair(Some(1.0)), 25.0 / 300.0);
for key in ["anchor", "player"] {
let lhs = curve(&implicit, key);
let rhs = curve(&explicit, key);
assert_eq!(lhs.len(), rhs.len());
for ((t_l, g_l), (t_r, g_r)) in lhs.iter().zip(rhs.iter()) {
assert_eq!(t_l, t_r);
assert_eq!(
(g_l.mu(), g_l.sigma()),
(g_r.mu(), g_r.sigma()),
"an explicit scale of 1.0 must be bit-identical to leaving it unset, \
for {key} at t={t_l}"
);
}
}
}
/// The use case from the issue: a static difficulty alongside drifting players,
/// in one graph. The anchor must hold still without absorbing drift through its
/// neighbours, and everything must stay finite.
#[test]
fn mixed_static_and_drifting_graph_converges() {
let mut events: Vec<Event<i64, &'static str>> = Vec::new();
let players = ["p0", "p1", "p2"];
for (i, p) in players.iter().cycle().take(9).enumerate() {
events.push(Event {
time: (i as i64) * 100,
teams: smallvec![
Team::with_members([Member::new(*p)]),
Team::with_members([Member::new("layout").with_drift_scale(0.0)]),
],
outcome: Outcome::winner((i % 2) as u32, 2),
});
}
let mut h = History::builder()
.mu(25.0)
.sigma(25.0 / 3.0)
.beta(25.0 / 6.0)
.p_draw(0.0)
.drift(ConstantDrift(25.0 / 300.0))
.convergence(CONVERGENCE)
.build();
h.add_events(events).unwrap();
let report = h.converge().unwrap();
assert!(report.converged, "mixed graph must converge: {report:?}");
let curves = h.learning_curves();
for (key, points) in &curves {
for (t, g) in points {
assert!(
g.mu().is_finite() && g.sigma().is_finite() && g.sigma() > 0.0,
"{key} at t={t} is not a usable posterior: mu={}, sigma={}",
g.mu(),
g.sigma()
);
}
}
let layout = curve(&h, "layout");
assert_eq!(layout.len(), 9);
let (_, first) = layout[0];
for (t, g) in &layout {
assert!(
(g.sigma() - first.sigma()).abs() < 1e-9,
"a static layout must not accumulate uncertainty; t={t} has sigma {} vs {}",
g.sigma(),
first.sigma()
);
}
let p0 = curve(&h, "p0");
assert!(
p0.last().unwrap().1.sigma() > 0.0,
"a drifting player should still have a proper posterior"
);
}
fn reject(scale: f64) -> InferenceError {
let mut h = History::builder()
.drift(ConstantDrift(25.0 / 300.0))
.build();
let events: Vec<Event<i64, &'static str>> = vec![Event {
time: 0,
teams: smallvec![
Team::with_members([Member::new("a").with_drift_scale(scale)]),
Team::with_members([Member::new("b")]),
],
outcome: Outcome::winner(0, 2),
}];
h.add_events(events)
.expect_err("an out-of-range drift_scale must be rejected")
}
#[test]
fn negative_scale_is_rejected() {
assert_eq!(
reject(-1.0),
InferenceError::InvalidParameter {
name: "drift_scale",
value: -1.0
}
);
}
#[test]
fn non_finite_scale_is_rejected() {
for scale in [f64::NAN, f64::INFINITY, f64::NEG_INFINITY] {
assert!(
matches!(
reject(scale),
InferenceError::InvalidParameter {
name: "drift_scale",
..
}
),
"a drift_scale of {scale} must be rejected as an invalid parameter"
);
}
}
/// The scale must reach the filtering pass too, not just `converge()`.
/// `filtered_learning_curves` runs its own drift application, so a pinned
/// competitor has to stay pinned there as well.
#[test]
fn zero_scale_pins_a_competitor_in_the_filtered_pass() {
let pinned = fit(distant_pair(Some(0.0)), 25.0 / 300.0);
let drifting = fit(distant_pair(None), 25.0 / 300.0);
let filtered = |h: &Fit| -> Vec<(i64, Gaussian)> {
let mut c = h
.filtered_learning_curves()
.remove("anchor")
.expect("anchor in filtered curves");
c.sort_by_key(|(t, _)| *t);
c
};
let pinned_curve = filtered(&pinned);
let drifting_curve = filtered(&drifting);
assert_eq!(pinned_curve.len(), 2);
assert_eq!(drifting_curve.len(), 2);
assert!(
pinned_curve[1].1.sigma() < pinned_curve[0].1.sigma(),
"a pinned competitor's filtered uncertainty must shrink with a second \
observation, not be re-inflated by drift: {} then {}",
pinned_curve[0].1.sigma(),
pinned_curve[1].1.sigma()
);
assert!(
pinned_curve[1].1.sigma() < drifting_curve[1].1.sigma() - 1e-6,
"pinning must leave the filtered estimate tighter than drifting does: \
{} vs {}",
pinned_curve[1].1.sigma(),
drifting_curve[1].1.sigma()
);
}
/// `drift_scale` is competitor configuration, and configuration supplied for a
/// competitor the history already knows is now *applied* rather than dropped.
///
/// This test previously asserted the opposite. It was written as a deliberate
/// change-detector — "moving the capture would be a visible break, not a silent
/// one" — and that is exactly what happened: the capture moved, and the
/// assertion inverted rather than being deleted.
///
/// Because configuration lives on the competitor and `converge` refits from
/// competitor state, a late pin applies to the *whole* history, not just to
/// events after it. So a scale set on the second batch must reach the same fit
/// as one set from the very first event.
#[test]
fn drift_scale_applies_when_set_after_first_appearance() {
let mut late = History::builder()
.mu(25.0)
.sigma(25.0 / 3.0)
.beta(25.0 / 6.0)
.p_draw(0.0)
.drift(ConstantDrift(25.0 / 300.0))
.convergence(CONVERGENCE)
.build();
// First batch creates "anchor" with the default scale.
late.add_events(vec![Event {
time: 0,
teams: smallvec![
Team::with_members([Member::new("anchor")]),
Team::with_members([Member::new("player")]),
],
outcome: Outcome::winner(0, 2),
}])
.unwrap();
// Second batch asks for a pin. No longer too late.
late.add_events(vec![Event {
time: 1000,
teams: smallvec![
Team::with_members([Member::new("anchor").with_drift_scale(0.0)]),
Team::with_members([Member::new("player")]),
],
outcome: Outcome::winner(1, 2),
}])
.unwrap();
let _ = late.converge().unwrap();
let applied = curve(&late, "anchor");
let pinned_from_the_start = curve(&fit(distant_pair(Some(0.0)), 25.0 / 300.0), "anchor");
let never_pinned = curve(&fit(distant_pair(None), 25.0 / 300.0), "anchor");
for ((t_l, g_l), (t_r, g_r)) in applied.iter().zip(pinned_from_the_start.iter()) {
assert_eq!(t_l, t_r);
assert!(
(g_l.sigma() - g_r.sigma()).abs() < 1e-9,
"a late pin should refit the whole history: t={t_l}, {} vs {}",
g_l.sigma(),
g_r.sigma()
);
}
// And it must actually have done something.
assert!(
applied
.iter()
.zip(never_pinned.iter())
.any(|((_, a), (_, b))| (a.sigma() - b.sigma()).abs() > 1e-9),
"the pin had no effect at all — the silent drop is back"
);
}
/// Re-declaring the same configuration must be inert. This is the shape a
/// caller gets when the configuration is a property of the domain — "layouts
/// are static" — so every ingestion path repeats it on every event.
///
/// Both histories see exactly the same events; only how many times the scale
/// is declared differs.
#[test]
fn repeating_the_same_configuration_changes_nothing() {
let events = |declare_every_time: bool| {
let anchor = |first: bool| {
if first || declare_every_time {
Member::new("anchor").with_drift_scale(0.0)
} else {
Member::new("anchor")
}
};
vec![
Event {
time: 0,
teams: smallvec![
Team::with_members([anchor(true)]),
Team::with_members([Member::new("player")]),
],
outcome: Outcome::winner(0, 2),
},
Event {
time: 1000,
teams: smallvec![
Team::with_members([anchor(false)]),
Team::with_members([Member::new("player")]),
],
outcome: Outcome::winner(1, 2),
},
]
};
let once = curve(&fit(events(false), 25.0 / 300.0), "anchor");
let every_time = curve(&fit(events(true), 25.0 / 300.0), "anchor");
for ((t_l, a), (t_r, b)) in once.iter().zip(every_time.iter()) {
assert_eq!(t_l, t_r);
assert!(
(a.sigma() - b.sigma()).abs() < 1e-12,
"t={t_l}: declaring the same scale repeatedly changed the fit, {} vs {}",
a.sigma(),
b.sigma()
);
}
}
#[test]
fn a_batch_that_contradicts_itself_is_rejected() {
let mut h = History::builder().convergence(CONVERGENCE).build();
let err = h
.add_events(vec![
Event {
time: 0,
teams: smallvec![
Team::with_members([Member::new("anchor").with_drift_scale(0.0)]),
Team::with_members([Member::new("player")]),
],
outcome: Outcome::winner(0, 2),
},
Event {
time: 1,
teams: smallvec![
Team::with_members([Member::new("anchor").with_drift_scale(1.0)]),
Team::with_members([Member::new("player")]),
],
outcome: Outcome::winner(0, 2),
},
])
.expect_err("two different scales for one competitor in one batch");
assert!(
matches!(
err,
InferenceError::ConflictingCompetitorConfig {
field: "drift_scale",
..
}
),
"got {err:?}"
);
}
+7 -12
View File
@@ -19,7 +19,8 @@ fn ts_rating(mu: f64, sigma: f64, beta: f64, gamma: f64) -> R {
fn game_1v1_golden_matches_historical() {
let a = ts_rating(25.0, 25.0 / 3.0, 25.0 / 6.0, 25.0 / 300.0);
let b = ts_rating(25.0, 25.0 / 3.0, 25.0 / 6.0, 25.0 / 300.0);
let (a_post, b_post) = Game::<i64, _>::one_v_one(&a, &b, Outcome::winner(0, 2)).unwrap();
let (a_post, b_post) =
Game::<i64, _>::one_v_one(&a, &b, Outcome::winner(0, 2), &GameOptions::default()).unwrap();
// Historical golden from pre-T2 test_1vs1 (team 0 wins):
assert_ulps_eq!(
a_post,
@@ -48,15 +49,9 @@ fn game_1v1_draw_golden() {
)
.unwrap();
let p = g.posteriors();
// Historical golden from pre-T2 test_1vs1_draw:
assert_ulps_eq!(
p[0][0],
Gaussian::from_ms(24.999999, 6.469480),
epsilon = 1e-6
);
assert_ulps_eq!(
p[1][0],
Gaussian::from_ms(24.999999, 6.469480),
epsilon = 1e-6
);
// Historical golden from pre-T2 test_1vs1_draw. The mean is 25.0 exactly
// by symmetry — two identical competitors drawing cannot move apart — and
// the reference's 24.999999 is that value transcribed to six decimals.
assert_ulps_eq!(p[0][0], Gaussian::from_ms(25.0, 6.469480), epsilon = 1e-6);
assert_ulps_eq!(p[1][0], Gaussian::from_ms(25.0, 6.469480), epsilon = 1e-6);
}
+254
View File
@@ -0,0 +1,254 @@
//! Forward-only (filtering) estimates: what the model knew at the time,
//! as opposed to the smoothed posteriors `learning_curve` reports.
use smallvec::smallvec;
use trueskill_tt::{ConvergenceOptions, Event, History, Member, Outcome, Team};
/// `games` one-on-one matches at successive times, won by "a" every time,
/// built with the given convergence options.
fn repeated_winner_with(games: i64, convergence: ConvergenceOptions) -> History {
let mut history = History::builder().convergence(convergence).build();
for time in 1..=games {
history
.add_events([Event {
time,
teams: smallvec![
Team::with_members([Member::new("a")]),
Team::with_members([Member::new("b")]),
],
outcome: Outcome::winner(0, 2),
}])
.unwrap();
}
history
}
/// `games` one-on-one matches at successive times, won by "a" every time.
///
/// This is the fixture from issue #19, where `online(true)` reported
/// `games * ln(0.5)`.
fn repeated_winner(games: i64) -> History {
repeated_winner_with(games, ConvergenceOptions::default())
}
/// The default 30-iteration cap leaves a residual around 1e-6, which would
/// swamp these comparisons. Drive both sides well past the fixed point.
fn tight() -> ConvergenceOptions {
ConvergenceOptions {
max_iter: 2_000,
epsilon: 1e-12,
..ConvergenceOptions::default()
}
}
#[test]
fn filtered_evidence_sits_between_coin_flip_and_batch() {
let mut history = repeated_winner(5);
let _ = history.converge().unwrap();
let coin_flip = 5.0 * 0.5f64.ln();
let batch = history.log_evidence();
let filtered = history.filtered_log_evidence();
assert!(
filtered > coin_flip,
"filtered evidence {filtered} is at or below {coin_flip}, the all-coin-flip \
value the inert online flag reported; game one is a coin flip but games two \
through five are not"
);
assert!(
filtered < batch,
"filtered evidence {filtered} is not below the smoothed {batch}; filtering \
scores each game on strictly less information than smoothing does"
);
}
#[test]
fn filtered_first_point_is_less_certain_than_smoothed() {
let mut history = repeated_winner(12);
let _ = history.converge().unwrap();
let smoothed = history.learning_curve("a");
let filtered = history.filtered_learning_curve("a");
assert_eq!(
smoothed.len(),
filtered.len(),
"both curves must cover the same time points"
);
let (smoothed_time, first_smoothed) = smoothed[0];
let (filtered_time, first_filtered) = filtered[0];
assert_eq!(smoothed_time, filtered_time);
assert!(
first_filtered.sigma() > first_smoothed.sigma(),
"filtered sigma {} at the first point is not above smoothed {}; the smoother \
collapses uncertainty before the first round is drawn, which is the whole \
reason this method exists",
first_filtered.sigma(),
first_smoothed.sigma()
);
assert!(
first_filtered.sigma() < trueskill_tt::SIGMA,
"filtered sigma {} at the first point is not below the prior {}; one game was \
played, so some uncertainty must have been resolved",
first_filtered.sigma(),
trueskill_tt::SIGMA
);
for pair in filtered.windows(2) {
assert!(
pair[1].1.mu() > pair[0].1.mu(),
"filtered mu must climb at every step for a competitor who wins every \
game: t={} mu={} then t={} mu={}",
pair[0].0,
pair[0].1.mu(),
pair[1].0,
pair[1].1.mu()
);
}
}
#[test]
fn filtered_curves_plural_agrees_with_singular() {
let mut history = repeated_winner(4);
let _ = history.converge().unwrap();
let curves = history.filtered_learning_curves();
assert_eq!(
curves["b"],
history.filtered_learning_curve("b"),
"the plural form must agree with the singular for the same key"
);
}
#[test]
fn filtered_evidence_is_invariant_to_convergence() {
let mut history = repeated_winner_with(6, tight());
let before = history.filtered_log_evidence();
let report = history.converge().unwrap();
assert!(
report.converged,
"fixture must converge: {:?}",
report.final_step
);
let after = history.filtered_log_evidence();
assert!(
(before - after).abs() < 1e-8,
"filtered evidence moved across converge(): {before} -> {after}. The pass must \
carry its own forward messages; anything reading skill.forward shows exactly \
this drift, because converge() contaminates it with backward information."
);
}
#[test]
fn single_slice_filtered_matches_smoothed() {
let mut history = History::builder().convergence(tight()).build();
history
.add_events([
Event {
time: 1,
teams: smallvec![
Team::with_members([Member::new("a")]),
Team::with_members([Member::new("b")]),
],
outcome: Outcome::winner(0, 2),
},
Event {
time: 1,
teams: smallvec![
Team::with_members([Member::new("c")]),
Team::with_members([Member::new("d")]),
],
outcome: Outcome::winner(0, 2),
},
])
.unwrap();
let _ = history.converge().unwrap();
let smoothed = history.learning_curve("a");
let filtered = history.filtered_learning_curve("a");
assert_eq!(smoothed.len(), 1);
assert_eq!(filtered.len(), 1);
assert!(
(smoothed[0].1.mu() - filtered[0].1.mu()).abs() < 1e-8
&& (smoothed[0].1.sigma() - filtered[0].1.sigma()).abs() < 1e-8,
"one slice has no future to propagate back, so filtered and smoothed must \
agree: smoothed mu={} sigma={}, filtered mu={} sigma={}",
smoothed[0].1.mu(),
smoothed[0].1.sigma(),
filtered[0].1.mu(),
filtered[0].1.sigma()
);
}
#[test]
fn filtered_curves_do_not_depend_on_ingestion_order() {
let events = |time: i64, winner: &'static str, loser: &'static str| Event {
time,
teams: smallvec![
Team::with_members([Member::new(winner)]),
Team::with_members([Member::new(loser)]),
],
outcome: Outcome::winner(0, 2),
};
let all = vec![
events(1, "a", "b"),
events(1, "c", "d"),
events(1, "a", "c"),
events(1, "b", "d"),
events(2, "a", "d"),
events(2, "b", "c"),
events(2, "a", "b"),
];
let mut batched = History::builder().convergence(tight()).build();
batched.add_events(all.clone()).unwrap();
let _ = batched.converge().unwrap();
let mut incremental = History::builder().convergence(tight()).build();
for event in all {
incremental.add_events([event]).unwrap();
}
let _ = incremental.converge().unwrap();
let from_batched = batched.filtered_learning_curve("a");
let from_incremental = incremental.filtered_learning_curve("a");
assert_eq!(from_batched.len(), from_incremental.len());
for ((time_b, gaussian_b), (time_i, gaussian_i)) in
from_batched.iter().zip(from_incremental.iter())
{
assert_eq!(time_b, time_i);
assert!(
(gaussian_b.mu() - gaussian_i.mu()).abs() < 1e-8
&& (gaussian_b.sigma() - gaussian_i.sigma()).abs() < 1e-8,
"at t={time_b}: batched mu={} sigma={}, incremental mu={} sigma={}",
gaussian_b.mu(),
gaussian_b.sigma(),
gaussian_i.mu(),
gaussian_i.sigma()
);
}
}
+44 -1
View File
@@ -32,7 +32,8 @@ fn game_ranked_1v1_golden() {
fn game_one_v_one_shortcut() {
let a = default_rating();
let b = default_rating();
let (a_post, b_post) = Game::<i64, _>::one_v_one(&a, &b, Outcome::winner(0, 2)).unwrap();
let (a_post, b_post) =
Game::<i64, _>::one_v_one(&a, &b, Outcome::winner(0, 2), &GameOptions::default()).unwrap();
assert!(a_post.mu() > 25.0);
assert!(b_post.mu() < 25.0);
}
@@ -95,3 +96,45 @@ fn game_log_evidence_is_finite() {
assert!(g.log_evidence().is_finite());
assert!(g.log_evidence() < 0.0);
}
/// `one_v_one` used to hardcode `GameOptions::default()`, so a 1v1 could
/// never set `p_draw` and a drawn 1v1 was unreachable through it.
#[test]
fn one_v_one_honours_the_draw_probability_it_is_given() {
let a = default_rating();
let b = default_rating();
// Default options still reject a draw, because the default p_draw is zero.
let err = Game::<i64, _>::one_v_one(&a, &b, Outcome::draw(2), &GameOptions::default())
.expect_err("a draw needs a positive p_draw");
assert!(matches!(
err,
InferenceError::TieWithoutDrawProbability { .. }
));
// With a draw probability supplied it succeeds — which was impossible
// before the signature took options.
let options = GameOptions {
p_draw: 0.25,
..GameOptions::default()
};
let (a_post, b_post) = Game::<i64, _>::one_v_one(&a, &b, Outcome::draw(2), &options)
.expect("a draw is representable once p_draw is positive");
// A symmetric draw leaves the means alone and sharpens both sides.
assert!((a_post.mu() - b_post.mu()).abs() < 1e-9);
assert!(a_post.sigma() < 25.0 / 3.0);
}
/// Convergence options reach the 1v1 path too, not just `p_draw`.
#[test]
fn one_v_one_honours_convergence_options() {
let a = default_rating();
let b = default_rating();
let options = GameOptions {
convergence: ConvergenceOptions::default(),
..GameOptions::default()
};
let (a_post, _) = Game::<i64, _>::one_v_one(&a, &b, Outcome::winner(0, 2), &options).unwrap();
assert!(a_post.mu() > 25.0);
}
+225
View File
@@ -0,0 +1,225 @@
//! Ingesting the same events must give the same answer however they were
//! batched.
//!
//! The numerical goldens all ingest in a single call with one slice per
//! timestamp, so they never exercise the "append to an existing slice" path.
//! These do.
use smallvec::smallvec;
use trueskill_tt::{ConvergenceOptions, Event, Gaussian, History, Member, Outcome, Team};
/// Converge tightly: the default cap of 30 iterations leaves a residual around
/// 1e-6, which would swamp the comparison. Both paths must reach the same
/// fixed point, so drive both well past it.
fn tight() -> ConvergenceOptions {
ConvergenceOptions {
max_iter: 2_000,
epsilon: 1e-12,
..ConvergenceOptions::default()
}
}
fn event(a: &str, b: &str, time: i64) -> Event<i64, String> {
Event {
time,
teams: smallvec![
Team::with_members([Member::new(a.to_string())]),
Team::with_members([Member::new(b.to_string())]),
],
outcome: Outcome::winner(0, 2),
}
}
/// Like [`event`], but `a` carries competitor configuration.
///
/// `prior` and `drift_scale` configure the competitor rather than the event, so
/// they are the part of ingestion most exposed to order: they are consumed once,
/// where the competitor's state is written.
fn configured_event(a: &str, b: &str, time: i64, scale: f64) -> Event<i64, String> {
Event {
time,
teams: smallvec![
Team::with_members([Member::new(a.to_string()).with_drift_scale(scale)]),
Team::with_members([Member::new(b.to_string())]),
],
outcome: Outcome::winner(0, 2),
}
}
fn converged_skills(events: Vec<Event<i64, String>>, batched: bool) -> Vec<(String, Gaussian)> {
let mut h: History<i64, _, _, String> =
History::builder_with_key().convergence(tight()).build();
if batched {
h.add_events(events).unwrap();
} else {
for ev in events {
h.add_events(std::iter::once(ev)).unwrap();
}
}
let report = h.converge().unwrap();
assert!(
report.converged,
"fixture must converge before results can be compared; final step {:?}",
report.final_step
);
let mut skills: Vec<(String, Gaussian)> = h
.learning_curves()
.into_iter()
.map(|(key, curve)| (key, curve.last().unwrap().1))
.collect();
skills.sort_by(|a, b| a.0.cmp(&b.0));
skills
}
fn assert_same(batched: &[(String, Gaussian)], incremental: &[(String, Gaussian)], what: &str) {
assert_eq!(
batched.len(),
incremental.len(),
"{what}: competitor count differs"
);
for ((kb, gb), (ki, gi)) in batched.iter().zip(incremental.iter()) {
assert_eq!(kb, ki, "{what}: key order differs");
assert!(
(gb.mu() - gi.mu()).abs() < 1e-8 && (gb.sigma() - gi.sigma()).abs() < 1e-8,
"{what}: {kb} differs — batched mu={} sigma={}, incremental mu={} sigma={}",
gb.mu(),
gb.sigma(),
gi.mu(),
gi.sigma()
);
}
}
/// All events share one timestamp, so incremental ingestion repeatedly appends
/// to an existing slice.
#[test]
fn same_slice_incremental_matches_batched() {
let events = vec![
event("a", "b", 1),
event("c", "d", 1),
event("e", "f", 1),
event("a", "c", 1),
event("b", "e", 1),
];
let batched = converged_skills(events.clone(), true);
let incremental = converged_skills(events, false);
assert_same(&batched, &incremental, "single shared slice");
}
/// Distinct timestamps, so each append lands in a fresh slice appended after
/// the existing ones.
#[test]
fn distinct_slices_incremental_matches_batched() {
let events = vec![
event("a", "b", 1),
event("b", "c", 2),
event("c", "a", 3),
event("a", "c", 4),
];
let batched = converged_skills(events.clone(), true);
let incremental = converged_skills(events, false);
assert_same(&batched, &incremental, "distinct slices");
}
/// Several events per timestamp across several timestamps — appends to
/// existing slices interleaved with new ones.
#[test]
fn mixed_slices_incremental_matches_batched() {
let events = vec![
event("a", "b", 1),
event("c", "d", 1),
event("a", "c", 2),
event("b", "d", 2),
event("a", "d", 3),
event("b", "c", 3),
];
let batched = converged_skills(events.clone(), true);
let incremental = converged_skills(events, false);
assert_same(&batched, &incremental, "mixed slices");
}
/// Appending an event to a slice that is *not* the most recent one exercises
/// the forward refresh of every later slice.
#[test]
fn back_dated_event_matches_batched() {
let events = vec![
event("a", "b", 1),
event("b", "c", 5),
event("c", "a", 9),
// arrives last, but belongs to the middle slice
event("a", "c", 5),
];
let batched = converged_skills(events.clone(), true);
let incremental = converged_skills(events, false);
assert_same(&batched, &incremental, "back-dated event");
}
/// The invariant this file protects was only ever checked for *unconfigured*
/// competitors — every helper above built members with `Member::new`.
///
/// Configuration is the part most exposed to ordering, because it is consumed
/// once at the point the competitor's state is written rather than replayed per
/// event. These cover it.
#[test]
fn configured_competitors_are_order_independent() {
let events = vec![
configured_event("a", "b", 0, 0.0),
configured_event("a", "c", 1, 0.0),
configured_event("a", "b", 2, 0.0),
event("b", "c", 3),
];
assert_same(
&converged_skills(events.clone(), true),
&converged_skills(events, false),
"configuration repeated on every appearance",
);
}
/// Configuration supplied only on a *later* event is the case that used to be
/// silently dropped. It must now reach the same fit either way it is ingested.
#[test]
fn late_configuration_is_order_independent() {
let events = vec![
event("a", "b", 0),
configured_event("a", "c", 1, 0.0),
event("a", "b", 2),
];
assert_same(
&converged_skills(events.clone(), true),
&converged_skills(events, false),
"configuration supplied after first appearance",
);
}
/// And it must actually be doing something — an implementation that dropped
/// configuration entirely would pass both tests above.
#[test]
fn configuration_changes_the_fit_however_it_is_ingested() {
let configured = vec![
event("a", "b", 0),
configured_event("a", "c", 1, 0.0),
event("a", "b", 2),
];
let plain = vec![event("a", "b", 0), event("a", "c", 1), event("a", "b", 2)];
for batched in [true, false] {
let with = converged_skills(configured.clone(), batched);
let without = converged_skills(plain.clone(), batched);
assert!(
with.iter()
.zip(&without)
.any(|((_, x), (_, y))| (x.sigma() - y.sigma()).abs() > 1e-9),
"batched={batched}: configuration had no effect, so the order tests are vacuous"
);
}
}
+1 -1
View File
@@ -46,7 +46,7 @@ fn nan_after_fit(players: usize) -> usize {
let (w, l) = if rng.coin() { (a, b) } else { (b, a) };
h.record_winner(&ids[w], &ids[l], 0).unwrap();
}
h.converge().unwrap();
let _ = h.converge().unwrap();
ids.iter()
.filter(|id| {
+367
View File
@@ -0,0 +1,367 @@
//! Calibration of the crate's marginals against the EXACT posterior.
//!
//! A scored history is linear-Gaussian — `MarginFactor` encodes
//! `score_a - score_b ~ N(perf_a - perf_b, score_sigma^2)` — so the true joint
//! posterior has a closed form and the crate can be checked against ground
//! truth rather than against intuition. That is not possible for ranked
//! outcomes, whose truncation likelihood EP genuinely approximates.
//!
//! Two things are pinned here, and one is deliberately only recorded.
//!
//! **Pinned: on a tree the crate is exact**, means and variances both. Message
//! passing has no approximation to make when the factor graph has no cycles, so
//! any drift here would be a real defect.
//!
//! **Pinned: means are exact even with cycles.** This is the standard result
//! for Gaussian belief propagation (Weiss & Freeman 2001) and it is what makes
//! ratings trustworthy.
//!
//! **Recorded, not asserted: with cycles, marginal variances are too narrow.**
//! Measured on the round-robin fixture below, the crate reports sigma 1.430
//! where the exact posterior is 2.851 — a ratio of 0.502. That is the known
//! behaviour of loopy Gaussian BP, not a bug in this crate, and it is left
//! unasserted because fixing it is exactly what #46 proposes.
//!
//! Why that matters for a consumer, and why #46 cannot be implemented as "add
//! a covariance accessor": the exact correlation between two nodes here is
//! +0.857, so a consumer computing `sqrt(sa^2 + sb^2)` for a difference
//! overstates its width. But the too-narrow marginals partially cancel that,
//! leaving 1.327x rather than 2.646x. Adding true correlations to these
//! marginals without also correcting them would give 0.765 against a true
//! 1.524 — *overconfident*, which is the unsafe direction.
use smallvec::smallvec;
use trueskill_tt::{ConstantDrift, ConvergenceOptions, Event, History, Member, Outcome, Team};
const N: usize = 5;
const MU0: f64 = 0.0;
const SIGMA0: f64 = 6.0;
const BETA: f64 = 1.0;
const SCORE_SIGMA: f64 = 2.0;
/// A STAR: every event touches c0, so the node-event graph is a tree and
/// Gaussian BP is exact. Any discrepancy here is not caused by loops.
fn tree_fixture() -> Vec<(usize, usize, f64)> {
vec![(0, 1, 3.0), (0, 2, 5.0), (0, 3, 4.0), (0, 4, 6.0)]
}
/// (winner, loser, score_diff)
fn fixture() -> Vec<(usize, usize, f64)> {
vec![
(0, 1, 3.0),
(0, 2, 5.0),
(1, 2, 2.0),
(3, 4, 1.0),
(0, 3, 4.0),
(1, 4, 2.5),
(2, 3, 0.5),
(0, 4, 6.0),
(1, 3, 1.5),
(2, 4, 3.0),
]
}
/// Invert a small symmetric positive-definite matrix by Gauss-Jordan.
fn inverse(mut a: Vec<Vec<f64>>) -> Vec<Vec<f64>> {
let n = a.len();
let mut inv: Vec<Vec<f64>> = (0..n)
.map(|i| (0..n).map(|j| if i == j { 1.0 } else { 0.0 }).collect())
.collect();
for col in 0..n {
// partial pivot
let mut piv = col;
for r in col + 1..n {
if a[r][col].abs() > a[piv][col].abs() {
piv = r;
}
}
a.swap(col, piv);
inv.swap(col, piv);
let d = a[col][col];
for j in 0..n {
a[col][j] /= d;
inv[col][j] /= d;
}
for r in 0..n {
if r == col {
continue;
}
let f = a[r][col];
for j in 0..n {
a[r][j] -= f * a[col][j];
inv[r][j] -= f * inv[col][j];
}
}
}
inv
}
/// The exact posterior of a linear-Gaussian model:
/// precision = prior precision + sum of a_k a_k^T / v_k.
fn exact_for(obs: &[(usize, usize, f64)]) -> (Vec<f64>, Vec<Vec<f64>>) {
let mut lambda = vec![vec![0.0; N]; N];
let mut eta = [0.0; N];
for (i, row) in lambda.iter_mut().enumerate() {
row[i] = 1.0 / (SIGMA0 * SIGMA0);
eta[i] = MU0 / (SIGMA0 * SIGMA0);
}
// Each 1v1 observation: d ~ N(x_a - x_b, score_sigma^2 + 2 beta^2)
let v = SCORE_SIGMA * SCORE_SIGMA + 2.0 * BETA * BETA;
for &(a, b, d) in obs {
let mut vec_a = [0.0; N];
vec_a[a] = 1.0;
vec_a[b] = -1.0;
for i in 0..N {
for j in 0..N {
lambda[i][j] += vec_a[i] * vec_a[j] / v;
}
eta[i] += vec_a[i] * d / v;
}
}
let cov = inverse(lambda);
let mean: Vec<f64> = (0..N)
.map(|i| (0..N).map(|j| cov[i][j] * eta[j]).sum())
.collect();
(mean, cov)
}
fn key(i: usize) -> &'static str {
["c0", "c1", "c2", "c3", "c4"][i]
}
/// Returns (worst mean error, worst sd ratio).
fn fitted(
obs: &[(usize, usize, f64)],
) -> History<i64, ConstantDrift, trueskill_tt::NullObserver, &'static str> {
let mut h: History<i64, _, _, &'static str> = History::builder()
.mu(MU0)
.sigma(SIGMA0)
.beta(BETA)
.score_sigma(SCORE_SIGMA)
.drift(ConstantDrift(0.0))
.convergence(ConvergenceOptions {
max_iter: 20_000,
epsilon: 1e-13,
alpha: 1.0,
})
.build();
let events: Vec<Event<i64, &'static str>> = obs
.iter()
.copied()
.map(|(a, b, d)| Event {
time: 1,
teams: smallvec![
Team::with_members([Member::new(key(a))]),
Team::with_members([Member::new(key(b))]),
],
outcome: Outcome::scores([d, 0.0]),
})
.collect();
h.add_events(events).unwrap();
let report = h.converge().unwrap();
assert!(
report.converged,
"fixture must converge: {:?}",
report.final_step
);
h
}
/// Returns (worst mean error, worst sd ratio gap).
fn run(name: &str, obs: Vec<(usize, usize, f64)>) -> (f64, f64) {
println!("\n########## {name} ##########");
let h = fitted(&obs);
let (mean, cov) = exact_for(&obs);
println!("\n== marginals: crate vs the exact linear-Gaussian posterior ==");
println!(
"{:>4} {:>12} {:>12} {:>12} {:>12} {:>8}",
"node", "crate mu", "exact mu", "crate sd", "exact sd", "sd ratio"
);
for i in 0..N {
let g = h.current_skill(&key(i)).unwrap();
let exact_sd = cov[i][i].sqrt();
println!(
"{:>4} {:>12.6} {:>12.6} {:>12.6} {:>12.6} {:>8.3}",
key(i),
g.mu(),
mean[i],
g.sigma(),
exact_sd,
g.sigma() / exact_sd
);
}
let mut worst_mean = 0.0f64;
let mut worst_ratio_gap = 0.0f64;
for i in 0..N {
let g = h.current_skill(&key(i)).unwrap();
worst_mean = worst_mean.max((g.mu() - mean[i]).abs());
worst_ratio_gap = worst_ratio_gap.max((g.sigma() / cov[i][i].sqrt() - 1.0).abs());
}
println!("\n== what a consumer actually computes for a DIFFERENCE ==");
println!(
"{:>8} {:>12} {:>14} {:>14} {:>12}",
"pair", "exact", "naive(exact)", "naive(crate)", "crate err"
);
for i in 0..N {
for j in i + 1..N {
if i != 0 && j != 1 {
continue;
}
let gi = h.current_skill(&key(i)).unwrap();
let gj = h.current_skill(&key(j)).unwrap();
let exact_sd = (cov[i][i] + cov[j][j] - 2.0 * cov[i][j]).sqrt();
let naive_exact = (cov[i][i] + cov[j][j]).sqrt();
let naive_crate = (gi.sigma().powi(2) + gj.sigma().powi(2)).sqrt();
let corr = cov[i][j] / (cov[i][i].sqrt() * cov[j][j].sqrt());
println!(
"{:>8} {:>12.6} {:>14.6} {:>14.6} {:>11.3}x (corr {corr:.4})",
format!("{}-{}", key(i), key(j)),
exact_sd,
naive_exact,
naive_crate,
naive_crate / exact_sd
);
}
}
(worst_mean, worst_ratio_gap)
}
/// With no cycles there is nothing for message passing to approximate.
#[test]
fn on_a_tree_the_marginals_are_exact() {
let (mean_err, sd_gap) = run("TREE (star: no loops, BP is exact)", tree_fixture());
assert!(
mean_err < 1e-9,
"tree means should be exact, worst error {mean_err}"
);
assert!(
sd_gap < 1e-9,
"tree sigmas should be exact, worst ratio gap {sd_gap}"
);
}
/// With cycles the means stay exact — the property ratings depend on — while
/// the variances do not. The variance gap is measured and reported rather than
/// asserted; see the module docs.
#[test]
fn with_cycles_the_means_stay_exact_but_the_variances_shrink() {
let (mean_err, sd_gap) = run("LOOPY (round robin)", fixture());
assert!(
mean_err < 1e-9,
"loopy means must still be exact, worst error {mean_err}"
);
assert!(
sd_gap > 0.1,
"the loopy variance gap is the premise of #46; if it has closed, that \
issue and these docs need revisiting (worst ratio gap {sd_gap})"
);
}
/// The point of #46: `posterior_of` must reproduce the exact joint, including
/// the correlation that marginals cannot express.
#[test]
fn posterior_of_matches_the_exact_joint() {
for (name, obs) in [("tree", tree_fixture()), ("loopy", fixture())] {
let h = fitted(&obs);
let (_, cov) = exact_for(&obs);
println!("\n== posterior_of vs exact ({name}) ==");
println!(
"{:>12} {:>14} {:>14} {:>10}",
"functional", "posterior_of", "exact", "ratio"
);
for (i, j) in [(0usize, 1usize), (0, 2), (1, 3), (2, 4)] {
let got = h
.posterior_of(&[(&key(i), 1.0), (&key(j), -1.0)])
.expect("scored slice should have a joint");
let exact_sd = (cov[i][i] + cov[j][j] - 2.0 * cov[i][j]).sqrt();
println!(
"{:>12} {:>14.6} {:>14.6} {:>10.4}",
format!("{}-{}", key(i), key(j)),
got.sigma(),
exact_sd,
got.sigma() / exact_sd
);
assert!(
(got.sigma() - exact_sd).abs() / exact_sd < 1e-9,
"{name} {}-{}: posterior_of gave {} where the exact joint is {exact_sd}",
key(i),
key(j),
got.sigma()
);
}
// A single competitor: this is where the loopy marginal was 2x narrow.
for (i, row) in cov.iter().enumerate() {
let got = h.posterior_of(&[(&key(i), 1.0)]).unwrap();
let exact_sd = row[i].sqrt();
assert!(
(got.sigma() - exact_sd).abs() / exact_sd < 1e-9,
"{name} {}: posterior_of gave {} where exact is {exact_sd}",
key(i),
got.sigma()
);
}
println!(" single-competitor marginals also exact");
}
}
/// Cost of the dense solve as the slice grows. Recorded, not asserted.
#[test]
#[ignore = "timing probe, run explicitly"]
fn cost_scaling() {
use std::time::Instant;
for n in [50usize, 100, 200, 400, 800] {
let names: Vec<String> = (0..n).map(|i| format!("c{i}")).collect();
let mut h: History<i64, _, _, String> = History::builder_with_key()
.score_sigma(2.0)
.drift(ConstantDrift(0.0))
.convergence(ConvergenceOptions {
max_iter: 200,
epsilon: 1e-8,
alpha: 1.0,
})
.build();
let mut seed = 5u64;
let mut rnd = move || {
seed ^= seed << 13;
seed ^= seed >> 7;
seed ^= seed << 17;
seed
};
let events: Vec<Event<i64, String>> = (0..n * 4)
.map(|_| {
let a = (rnd() as usize) % n;
let mut b = (rnd() as usize) % n;
if b == a {
b = (b + 1) % n;
}
Event {
time: 1,
teams: smallvec![
Team::with_members([Member::new(names[a].clone())]),
Team::with_members([Member::new(names[b].clone())]),
],
outcome: Outcome::scores([1.0, 0.0]),
}
})
.collect();
h.add_events(events).unwrap();
let _ = h.converge().unwrap();
let t = Instant::now();
let g = h
.posterior_of(&[(&names[0], 1.0), (&names[1], -1.0)])
.unwrap();
println!(" n={n:>4}: {:>10.2?} sigma {:.6}", t.elapsed(), g.sigma());
}
}
+161
View File
@@ -0,0 +1,161 @@
//! `Observer` callbacks must actually fire.
//!
//! `on_slice_processed` (formerly `on_batch_processed`) was declared on the
//! trait and never called from anywhere, so implementors wired up a callback
//! that could not run. These tests exist so that cannot silently recur.
use std::sync::{Arc, Mutex};
use trueskill_tt::{History, Observer};
/// Plain fields. `Arc<O>` implements `Observer`, so the caller shares the
/// observer itself rather than wrapping each field in its own `Arc`.
#[derive(Default)]
struct Recorder {
iterations: Mutex<Vec<usize>>,
slices: Mutex<Vec<(i64, usize, usize)>>,
converged: Mutex<Vec<(usize, bool)>>,
}
impl Observer<i64> for Recorder {
fn on_iteration_end(&self, iter: usize, _max_step: (f64, f64)) {
self.iterations.lock().unwrap().push(iter);
}
fn on_slice_processed(&self, time: &i64, slice_idx: usize, n_events: usize) {
self.slices
.lock()
.unwrap()
.push((*time, slice_idx, n_events));
}
fn on_converged(&self, iters: usize, _final_step: (f64, f64), converged: bool) {
self.converged.lock().unwrap().push((iters, converged));
}
}
#[test]
fn every_observer_callback_fires() {
let recorder = Arc::new(Recorder::default());
let mut h = History::builder().observer(Arc::clone(&recorder)).build();
h.record_winner(&"a", &"b", 1).unwrap();
h.record_winner(&"b", &"c", 2).unwrap();
h.record_winner(&"c", &"a", 3).unwrap();
let _ = h.converge().unwrap();
assert!(
!recorder.iterations.lock().unwrap().is_empty(),
"on_iteration_end never fired"
);
assert!(
!recorder.converged.lock().unwrap().is_empty(),
"on_converged never fired"
);
assert!(
!recorder.slices.lock().unwrap().is_empty(),
"on_slice_processed never fired — the defect this test exists for"
);
}
#[test]
fn slice_callbacks_report_the_slice_they_swept() {
let recorder = Arc::new(Recorder::default());
let mut h = History::builder().observer(Arc::clone(&recorder)).build();
h.record_winner(&"a", &"b", 10).unwrap();
h.record_winner(&"a", &"b", 20).unwrap();
let _ = h.converge().unwrap();
let slices = recorder.slices.lock().unwrap();
// Only the times actually in the history, and each with its own events.
for &(time, idx, events) in slices.iter() {
assert!(time == 10 || time == 20, "unexpected slice time {time}");
assert!(idx < 2, "slice index {idx} out of range");
assert_eq!(events, 1, "each slice holds exactly one event");
}
// Both slices must be reported, not just one end of the sweep.
assert!(
slices.iter().any(|&(t, ..)| t == 10),
"slice 10 never reported"
);
assert!(
slices.iter().any(|&(t, ..)| t == 20),
"slice 20 never reported"
);
}
#[test]
fn a_single_slice_history_still_reports_its_sweep() {
let recorder = Arc::new(Recorder::default());
let mut h = History::builder().observer(Arc::clone(&recorder)).build();
h.record_winner(&"a", &"b", 1).unwrap();
let _ = h.converge().unwrap();
let slices = recorder.slices.lock().unwrap();
assert!(
!slices.is_empty(),
"the single-slice path must report its sweep too"
);
assert!(slices.iter().all(|&(t, idx, _)| t == 1 && idx == 0));
}
/// The gap #40 closed: without `impl Observer for Arc<O>`, an observer that
/// accumulates anything had to wrap every field in its own `Arc` and derive
/// `Clone`, because `History` consumes the observer and never hands it back.
#[test]
fn a_shared_observer_reaches_the_callers_handle() {
let recorder = Arc::new(Recorder::default());
let mut h = History::builder().observer(Arc::clone(&recorder)).build();
h.record_winner(&"a", &"b", 1).unwrap();
let _ = h.converge().unwrap();
assert!(!recorder.iterations.lock().unwrap().is_empty());
assert!(!recorder.slices.lock().unwrap().is_empty());
assert!(!recorder.converged.lock().unwrap().is_empty());
}
/// `?Sized` on the blanket impls means the observer can be chosen at runtime.
#[test]
fn a_trait_object_observer_works() {
let boxed: Box<dyn Observer<i64>> = Box::new(Recorder::default());
let mut h = History::builder().observer(boxed).build();
h.record_winner(&"a", &"b", 1).unwrap();
let _ = h.converge().unwrap();
let shared: Arc<dyn Observer<i64>> = Arc::new(Recorder::default());
let mut h = History::builder().observer(Arc::clone(&shared)).build();
h.record_winner(&"a", &"b", 1).unwrap();
let _ = h.converge().unwrap();
}
/// A non-shared observer can be reclaimed after convergence instead.
#[test]
fn into_observer_returns_the_accumulated_state() {
let mut h = History::builder().observer(Recorder::default()).build();
h.record_winner(&"a", &"b", 1).unwrap();
let _ = h.converge().unwrap();
// Readable in place...
assert!(!h.observer().iterations.lock().unwrap().is_empty());
// ...and reclaimable by value.
let recorder = h.into_observer();
assert!(!recorder.slices.lock().unwrap().is_empty());
}
/// Borrowing works too, for an observer that outlives the history.
#[test]
fn a_borrowed_observer_works() {
let recorder = Recorder::default();
{
let mut h = History::builder().observer(&recorder).build();
h.record_winner(&"a", &"b", 1).unwrap();
let _ = h.converge().unwrap();
}
assert!(!recorder.iterations.lock().unwrap().is_empty());
}
+153
View File
@@ -0,0 +1,153 @@
//! `predict_margin`: the predictive distribution of a scored matchup.
use smallvec::smallvec;
use trueskill_tt::{
ConstantDrift, ConvergenceOptions, Event, History, InferenceError, Member, Outcome, Team,
UnknownKeys,
};
fn builder(
policy: UnknownKeys,
) -> History<i64, ConstantDrift, trueskill_tt::NullObserver, &'static str> {
History::builder()
.mu(0.0)
.sigma(6.0)
.beta(1.0)
.score_sigma(2.0)
.drift(ConstantDrift(0.0))
.unknown_keys(policy)
.convergence(ConvergenceOptions {
max_iter: 5_000,
epsilon: 1e-12,
alpha: 1.0,
})
.build()
}
fn round(a: &'static str, b: &'static str, sa: f64, sb: f64) -> Event<i64, &'static str> {
Event {
time: 1,
teams: smallvec![
Team::with_members([Member::new(a)]),
Team::with_members([Member::new(b)]),
],
outcome: Outcome::scores([sa, sb]),
}
}
/// A history where "veteran" and "regular" are well observed and "novice"
/// appears once.
fn fitted(
policy: UnknownKeys,
) -> History<i64, ConstantDrift, trueskill_tt::NullObserver, &'static str> {
let mut h = builder(policy);
let mut events: Vec<_> = (0..40)
.map(|t| round("veteran", "regular", 10.0 + f64::from(t % 3), 5.0))
.collect();
events.push(round("veteran", "novice", 10.0, 6.0));
h.add_events(events).unwrap();
let _ = h.converge().unwrap();
h
}
/// The property #48 exists for: the interval must widen when the model knows
/// less. Their hand-fitted noise law quoted the same sigma for a competitor
/// with forty rounds and one with none.
#[test]
fn the_interval_widens_as_the_model_knows_less() {
let h = fitted(UnknownKeys::Prior);
let well_known = h
.predict_margin(&[&[&"veteran"], &[&"regular"]])
.unwrap()
.sigma();
let thin = h
.predict_margin(&[&[&"veteran"], &[&"novice"]])
.unwrap()
.sigma();
let unseen = h
.predict_margin(&[&[&"veteran"], &[&"stranger"]])
.unwrap()
.sigma();
assert!(
well_known < thin && thin < unseen,
"margin width should grow as evidence thins: {well_known} < {thin} < {unseen}"
);
}
/// #48's second requirement: an unseen competitor is a legitimate question, not
/// an error, and the answer should come from the prior rather than be faked.
#[test]
fn an_unseen_competitor_is_answered_from_the_prior() {
let h = fitted(UnknownKeys::Prior);
let g = h.predict_margin(&[&[&"nobody"], &[&"no_one"]]).unwrap();
// Two unknowns: the gap is centred on zero and carries both priors plus
// both performance noises plus the observation noise.
assert!(g.mu().abs() < 1e-9, "mu {}", g.mu());
let expected = (2.0 * 36.0 + 2.0 * 1.0 + 4.0f64).sqrt();
assert!(
(g.sigma() - expected).abs() < 1e-9,
"sigma {} vs expected {expected}",
g.sigma()
);
}
#[test]
fn reject_still_rejects() {
let h = fitted(UnknownKeys::Reject);
assert!(matches!(
h.predict_margin(&[&[&"veteran"], &[&"stranger"]]),
Err(InferenceError::UnknownKey { .. })
));
}
/// The margin is the *difference*, so it must be antisymmetric in the teams.
#[test]
fn swapping_the_teams_negates_the_margin() {
let h = fitted(UnknownKeys::Prior);
let forward = h.predict_margin(&[&[&"veteran"], &[&"regular"]]).unwrap();
let reverse = h.predict_margin(&[&[&"regular"], &[&"veteran"]]).unwrap();
assert!((forward.mu() + reverse.mu()).abs() < 1e-9);
assert!((forward.sigma() - reverse.sigma()).abs() < 1e-12);
}
/// The predictive interval must be wider than the skill gap alone: it also
/// carries per-event performance noise and the observation noise.
#[test]
fn the_predictive_interval_exceeds_the_skill_uncertainty() {
let h = fitted(UnknownKeys::Prior);
let skill_gap = h
.posterior_of(&[(&"veteran", 1.0), (&"regular", -1.0)])
.unwrap();
let predictive = h.predict_margin(&[&[&"veteran"], &[&"regular"]]).unwrap();
assert!(
(predictive.mu() - skill_gap.mu()).abs() < 1e-12,
"means agree"
);
// beta^2 twice plus score_sigma^2 = 2 + 4.
let expected = (skill_gap.sigma().powi(2) + 6.0).sqrt();
assert!((predictive.sigma() - expected).abs() < 1e-12);
assert!(predictive.sigma() > skill_gap.sigma());
}
#[test]
fn shape_errors_are_reported() {
let h = fitted(UnknownKeys::Prior);
assert!(matches!(
h.predict_margin(&[&[&"veteran"]]),
Err(InferenceError::MismatchedShape {
expected: 2,
got: 1,
..
})
));
let empty: [&&str; 0] = [];
assert!(matches!(
h.predict_margin(&[&[&"veteran"], &empty]),
Err(InferenceError::EmptyTeam { team: 1 })
));
}
+417
View File
@@ -0,0 +1,417 @@
//! Prediction API: N-team outcomes, draw mass, and the error paths that used
//! to be panics or silent wrong answers.
use trueskill_tt::{History, InferenceError, MAX_PREDICTED_TEAMS};
fn history_with(names: &[&'static str], p_draw: f64) -> History {
let mut h = History::builder().p_draw(p_draw).build();
// Give every competitor a recorded skill by playing a small round robin.
for pair in names.windows(2) {
h.record_winner(&pair[0], &pair[1], 1).unwrap();
}
let _ = h.converge().unwrap();
h
}
#[test]
fn unknown_keys_are_reported_not_silently_dropped() {
let h = history_with(&["a", "b"], 0.0);
let err = h
.predict_outcome(&[&[&"a"], &[&"ghost"]])
.expect_err("an unknown key must not yield a confident prediction");
assert_eq!(
err,
InferenceError::UnknownKey {
team: 1,
member: 0,
key: "\"ghost\"".to_owned(),
}
);
// Every prediction entry point, not just one.
assert!(
h.predict_win_probabilities(&[&[&"a"], &[&"ghost"]])
.is_err()
);
assert!(h.predict_quality(&[&[&"a"], &[&"ghost"]]).is_err());
assert!(h.predict_ranking(&[&[&"a"], &[&"ghost"]], &[0, 1]).is_err());
}
#[test]
fn an_entirely_unknown_team_is_an_error() {
let h = history_with(&["a", "b"], 0.0);
let err = h.predict_outcome(&[&[&"a"], &[&"x", &"y"]]).unwrap_err();
assert_eq!(
err,
InferenceError::UnknownKey {
team: 1,
member: 0,
key: "\"x\"".to_owned(),
}
);
}
#[test]
fn degenerate_team_shapes_are_errors_rather_than_panics() {
let h = history_with(&["a", "b"], 0.0);
assert_eq!(
h.predict_outcome(&[&[&"a"]]).unwrap_err(),
InferenceError::NotEnoughTeams { got: 1 }
);
assert_eq!(
h.predict_outcome(&[]).unwrap_err(),
InferenceError::NotEnoughTeams { got: 0 }
);
assert_eq!(
h.predict_outcome(&[&[&"a"], &[]]).unwrap_err(),
InferenceError::EmptyTeam { team: 1 }
);
}
#[test]
fn more_than_two_teams_no_longer_panics() {
let h = history_with(&["a", "b", "c"], 0.0);
let p = h
.predict_outcome(&[&[&"a"], &[&"b"], &[&"c"]])
.expect("three teams must be supported");
assert!((p.total() - 1.0).abs() < 1e-6, "total = {}", p.total());
// Three teams, no draws possible: exactly the six strict orderings.
assert_eq!(p.outcomes().len(), 6);
}
#[test]
fn the_outcome_space_is_capped_rather_than_hanging() {
let names: Vec<&'static str> = vec!["a", "b", "c", "d", "e", "f", "g", "h"];
let h = history_with(&names, 0.0);
let teams: Vec<&[&&'static str]> = Vec::new();
let _ = teams;
let too_many: Vec<Vec<&&str>> = names.iter().map(|n| vec![n]).collect();
let refs: Vec<&[&&str]> = too_many.iter().map(Vec::as_slice).collect();
let err = h.predict_outcome(&refs).unwrap_err();
assert_eq!(
err,
InferenceError::TooManyTeams {
got: 8,
max: MAX_PREDICTED_TEAMS
}
);
// The cheap paths stay available at any size.
let wins = h.predict_win_probabilities(&refs).unwrap();
assert_eq!(wins.len(), 8);
assert!(
(wins.iter().sum::<f64>() - 1.0).abs() < 1e-6,
"win probabilities must still sum to one: {wins:?}"
);
}
/// The defect that made every draw-enabled prediction wrong: `[p, 1 - p]`
/// allocated no mass to a draw even with `p_draw > 0`.
#[test]
fn a_draw_carries_probability_mass_when_p_draw_is_positive() {
let h = history_with(&["a", "b"], 0.25);
let p = h.predict_outcome(&[&[&"a"], &[&"b"]]).unwrap();
let draw = p.probability_of(&[0, 0]);
assert!(draw > 0.0, "a draw-enabled model must give draws mass");
assert!((p.total() - 1.0).abs() < 1e-6, "total = {}", p.total());
let wins = p.win_probabilities();
assert!(
(wins.iter().sum::<f64>() + draw - 1.0).abs() < 1e-6,
"wins {wins:?} plus draw {draw} must be the whole space"
);
assert!(
(p.shared_first_place() - draw).abs() < 1e-12,
"a two-team draw is a shared first place"
);
}
#[test]
fn a_zero_draw_probability_admits_no_ties() {
let h = history_with(&["a", "b"], 0.0);
let p = h.predict_outcome(&[&[&"a"], &[&"b"]]).unwrap();
assert_eq!(p.probability_of(&[0, 0]), 0.0);
assert!(p.shared_first_place() < 1e-12);
}
/// The two routes to a win probability run through entirely different
/// algorithms — adaptive quadrature versus the enumerated chain recursion —
/// so agreement between them is a real cross-check, not a tautology.
#[test]
fn the_cheap_and_exhaustive_paths_agree() {
for p_draw in [0.0, 0.1] {
let h = history_with(&["a", "b", "c"], p_draw);
let teams: &[&[&&str]] = &[&[&"a"], &[&"b"], &[&"c"]];
let cheap = h.predict_win_probabilities(teams).unwrap();
let exhaustive = h.predict_outcome(teams).unwrap().win_probabilities();
for (i, (a, b)) in cheap.iter().zip(&exhaustive).enumerate() {
assert!(
(a - b).abs() < 1e-6,
"p_draw={p_draw} team {i}: quadrature {a} vs enumeration {b}"
);
}
}
}
#[test]
fn predict_ranking_agrees_with_the_distribution() {
let h = history_with(&["a", "b", "c"], 0.1);
let teams: &[&[&&str]] = &[&[&"a"], &[&"b"], &[&"c"]];
let dist = h.predict_outcome(teams).unwrap();
for (ranks, expected) in dist.outcomes() {
let direct = h.predict_ranking(teams, ranks).unwrap();
assert!(
(direct - expected).abs() < 1e-9,
"ranks {ranks:?}: {direct} vs {expected}"
);
}
}
#[test]
fn predict_ranking_checks_its_shape() {
let h = history_with(&["a", "b"], 0.0);
let err = h
.predict_ranking(&[&[&"a"], &[&"b"]], &[0, 1, 2])
.unwrap_err();
assert!(matches!(
err,
InferenceError::MismatchedShape {
expected: 2,
got: 3,
..
}
));
}
#[test]
fn the_stronger_competitor_is_favoured() {
let mut h = History::builder().build();
for t in 1..=10 {
h.record_winner(&"strong", &"weak", t).unwrap();
}
let _ = h.converge().unwrap();
let p = h.predict_outcome(&[&[&"strong"], &[&"weak"]]).unwrap();
let (best, _) = p.most_likely().expect("a most likely outcome");
assert_eq!(best, &[0, 1], "the winner should be favoured");
let wins = p.win_probabilities();
assert!(wins[0] > wins[1], "{wins:?}");
}
/// Unequal team sizes change the draw margin, because inference derives it
/// from the teams' betas. Prediction has to follow, or it describes a
/// different model than the one that will be fitted.
#[test]
fn team_size_affects_the_prediction() {
let mut h = History::builder().p_draw(0.2).build();
h.event(1)
.team(["a", "b"])
.team(["c"])
.winner(0)
.commit()
.unwrap();
let _ = h.converge().unwrap();
let p = h.predict_outcome(&[&[&"a", &"b"], &[&"c"]]).unwrap();
assert!((p.total() - 1.0).abs() < 1e-6, "total = {}", p.total());
assert!(p.probability_of(&[0, 0]) > 0.0);
}
// ---------------------------------------------------------------------------
// Expected information gain
// ---------------------------------------------------------------------------
/// The whole point of #39: "which comparison should I run next?" is a
/// different question from "who will win?" or "is this fair?".
#[test]
fn information_gain_prefers_the_uncertain_pairing() {
let mut h = History::builder().build();
// "known" and "rival" have played a lot; "newcomer" has played once.
for t in 1..=15 {
h.record_winner(&"known", &"rival", t).unwrap();
h.record_winner(&"rival", &"known", t + 100).unwrap();
}
h.record_winner(&"known", &"newcomer", 500).unwrap();
let _ = h.converge().unwrap();
let settled = h
.expected_information_gain(&[&[&"known"], &[&"rival"]])
.unwrap();
let unknown = h
.expected_information_gain(&[&[&"known"], &[&"newcomer"]])
.unwrap();
assert!(
unknown > settled,
"pairing against the newcomer should teach more: {unknown} vs {settled}"
);
}
/// The analytic ceiling, through the `History` entry point rather than the
/// standalone one.
#[test]
fn information_gain_respects_the_entropy_ceiling() {
let h = history_with(&["a", "b", "c"], 0.0);
let two = h.expected_information_gain(&[&[&"a"], &[&"b"]]).unwrap();
assert!(
(0.0..=std::f64::consts::LN_2).contains(&two),
"two-team EIG {two} outside [0, ln 2]"
);
let three = h
.expected_information_gain(&[&[&"a"], &[&"b"], &[&"c"]])
.unwrap();
assert!(
(0.0..=6.0f64.ln()).contains(&three),
"three-team EIG {three} outside [0, ln 6]"
);
}
#[test]
fn information_gain_reports_unknown_keys() {
let h = history_with(&["a", "b"], 0.0);
assert_eq!(
h.expected_information_gain(&[&[&"a"], &[&"ghost"]])
.unwrap_err(),
InferenceError::UnknownKey {
team: 1,
member: 0,
key: "\"ghost\"".to_owned(),
}
);
}
/// A draw-enabled history has three outcomes to weigh rather than two, so the
/// draw branch must actually be reachable through this path.
#[test]
fn information_gain_accounts_for_draws() {
let with_draws = history_with(&["a", "b"], 0.25);
let g = with_draws
.expected_information_gain(&[&[&"a"], &[&"b"]])
.unwrap();
assert!(g > 0.0 && g <= 3.0f64.ln(), "{g}");
// The draw outcome carries mass, so it is genuinely being weighed.
let dist = with_draws.predict_outcome(&[&[&"a"], &[&"b"]]).unwrap();
assert!(dist.probability_of(&[0, 0]) > 0.0);
}
/// The defect that cost a consumer a day: `UnknownKey { team: 0, member: 0 }`
/// says nothing about *which* key is unknown, so the natural handling — log it,
/// fall back to a neutral value — converts a total miss into a plausible
/// constant. The key has to be in the error, and in its `Display`.
#[test]
fn unknown_key_names_the_key_it_could_not_find() {
let h = history_with(&["a", "b"], 0.0);
let err = h.predict_outcome(&[&[&"a"], &[&"never_seen"]]).unwrap_err();
match &err {
InferenceError::UnknownKey { key, .. } => {
assert!(
key.contains("never_seen"),
"the error should name the key, got {key}"
);
}
other => panic!("expected UnknownKey, got {other:?}"),
}
let rendered = err.to_string();
assert!(
rendered.contains("never_seen"),
"Display should name the key: {rendered}"
);
assert!(
rendered.contains("pre-filter"),
"Display should say what to do about it: {rendered}"
);
}
// ---------------------------------------------------------------------------
// UnknownKeys policy
// ---------------------------------------------------------------------------
fn history_with_policy(names: &[&'static str], policy: trueskill_tt::UnknownKeys) -> History {
let mut h = History::builder().unknown_keys(policy).build();
for pair in names.windows(2) {
h.record_winner(&pair[0], &pair[1], 1).unwrap();
}
let _ = h.converge().unwrap();
h
}
#[test]
fn reject_is_the_default() {
let h = history_with(&["a", "b"], 0.0);
assert!(matches!(
h.predict_outcome(&[&[&"a"], &[&"ghost"]]),
Err(InferenceError::UnknownKey { .. })
));
}
#[test]
fn prior_answers_instead_of_erroring() {
let h = history_with_policy(&["a", "b"], trueskill_tt::UnknownKeys::Prior);
let p = h
.predict_outcome(&[&[&"a"], &[&"ghost"]])
.expect("Prior should answer rather than reject");
assert!((p.total() - 1.0).abs() < 1e-6);
}
/// Two competitors the model has never seen are genuinely a coin flip. The
/// point is that this is now *derived* rather than a constant a caller
/// substitutes after swallowing an error.
#[test]
fn two_unknown_competitors_are_an_honest_coin_flip() {
let h = history_with_policy(&["a", "b"], trueskill_tt::UnknownKeys::Prior);
let wins = h
.predict_win_probabilities(&[&[&"nobody"], &[&"no_one"]])
.unwrap();
assert!((wins[0] - 0.5).abs() < 1e-9, "{wins:?}");
assert!((wins[1] - 0.5).abs() < 1e-9, "{wins:?}");
}
/// The property that rules out a `Skip` mode: an unknown member must make a
/// team *less* certain, never more. Skipping would drop the member's variance
/// from the sum and narrow the team, which is backwards.
#[test]
fn an_unknown_member_widens_its_team_rather_than_narrowing_it() {
let h = history_with_policy(&["a", "b", "c"], trueskill_tt::UnknownKeys::Prior);
// "a" alone against "b" — then "a" plus an unknown partner against "b".
let solo = h.predict_win_probabilities(&[&[&"a"], &[&"b"]]).unwrap();
let with_unknown = h
.predict_win_probabilities(&[&[&"a", &"stranger"], &[&"b"]])
.unwrap();
// Adding an unknown partner pulls the outcome toward even, because the
// team's performance spread grew.
assert!(
(with_unknown[0] - 0.5).abs() < (solo[0] - 0.5).abs(),
"an unknown partner should make the result less certain: solo {solo:?}, \
with unknown {with_unknown:?}"
);
}
#[test]
fn prior_reaches_every_prediction_entry_point() {
let h = history_with_policy(&["a", "b"], trueskill_tt::UnknownKeys::Prior);
let teams: &[&[&&str]] = &[&[&"a"], &[&"ghost"]];
assert!(h.predict_quality(teams).is_ok());
assert!(h.predict_win_probabilities(teams).is_ok());
assert!(h.predict_outcome(teams).is_ok());
assert!(h.predict_ranking(teams, &[0, 1]).is_ok());
assert!(h.expected_information_gain(teams).is_ok());
}
+7
View File
@@ -0,0 +1,7 @@
# Seeds for failure cases proptest has generated in the past. It is
# automatically read and these particular cases re-run before any
# novel cases are generated.
#
# It is recommended to check this file in to source control so that
# everyone who runs the test benefits from these saved cases.
cc 8859be600e638573980f78622b8fcd8b4553ca34a9a041746c417f7e4293f89c # shrinks to games = [(0, 1), (0, 1), (4, 6), (2, 0), (0, 1), (0, 6), (0, 1), (2, 0), (6, 4), (0, 1), (0, 2), (6, 4), (1, 0), (4, 0), (0, 2)]
+183
View File
@@ -0,0 +1,183 @@
//! Property-based tests over generated histories.
//!
//! The golden suite pins exact values against the Python/Julia reference on a
//! handful of fixtures. These pin *invariants* over inputs nobody wrote by
//! hand, which is where the defects this crate has actually shipped were
//! hiding: a linear evidence product that underflowed only past ~1000 teams,
//! and a batching path no golden exercised because every golden ingests in one
//! call.
mod common;
use common::assert_finite;
use proptest::prelude::*;
use smallvec::smallvec;
use trueskill_tt::{ConvergenceOptions, Event, History, Member, Outcome, Team};
/// Distinct competitors, so no event pits someone against themselves.
fn pairs() -> impl Strategy<Value = Vec<(usize, usize)>> {
prop::collection::vec((0usize..8, 0usize..8), 1..24)
.prop_map(|v| v.into_iter().filter(|(a, b)| a != b).collect::<Vec<_>>())
.prop_filter("needs at least one valid pair", |v| !v.is_empty())
}
const KEYS: [&str; 8] = ["a", "b", "c", "d", "e", "f", "g", "h"];
fn history_from(games: &[(usize, usize)]) -> History {
let mut h = History::builder()
.convergence(ConvergenceOptions {
// 200 was not enough: the batched side stopped at the cap with a
// step of 3.4e-9, so this test was comparing two truncated fits and
// attributing the gap to ingestion order.
max_iter: 20_000,
epsilon: 1e-10,
..ConvergenceOptions::default()
})
.build();
let events: Vec<Event<i64, &'static str>> = games
.iter()
.enumerate()
.map(|(i, &(a, b))| Event {
time: i as i64 + 1,
teams: smallvec![
Team::with_members([Member::new(KEYS[a])]),
Team::with_members([Member::new(KEYS[b])]),
],
outcome: Outcome::winner(0, 2),
})
.collect();
h.add_events(events).unwrap();
h
}
proptest! {
#![proptest_config(ProptestConfig::with_cases(48))]
/// Whatever the schedule of games, convergence must not produce NaN or an
/// improper posterior. `converge` returns `NonFiniteResult` rather than
/// silently reporting a NaN step as converged, so a break shows up here as
/// either an Err or a non-finite curve point.
#[test]
fn converged_posteriors_are_always_finite(games in pairs()) {
let mut h = history_from(&games);
let _ = h.converge().unwrap();
for key in KEYS {
for (time, g) in h.learning_curve(key) {
assert_finite(g, &format!("{key} at t={time}"));
}
}
}
/// Log-evidence is a log probability: finite, and never above zero.
///
/// The linear-product implementation this replaced underflowed to zero on
/// long chains, making `ln(0)` = -inf — finite-ness is the property that
/// would have caught it.
#[test]
fn log_evidence_is_a_finite_log_probability(games in pairs()) {
let mut h = history_from(&games);
let _ = h.converge().unwrap();
let batch = h.log_evidence();
let filtered = h.filtered_log_evidence();
prop_assert!(batch.is_finite(), "batch log-evidence {batch} is not finite");
prop_assert!(batch <= 0.0, "batch log-evidence {batch} exceeds zero");
prop_assert!(filtered.is_finite(), "filtered log-evidence {filtered} is not finite");
prop_assert!(filtered <= 0.0, "filtered log-evidence {filtered} exceeds zero");
}
/// Filtered estimates must not depend on whether `converge` has run — the
/// property the whole forward-only design rests on.
#[test]
fn filtered_evidence_is_invariant_to_convergence(games in pairs()) {
let mut h = history_from(&games);
let before = h.filtered_log_evidence();
let _ = h.converge().unwrap();
let after = h.filtered_log_evidence();
prop_assert!(
(before - after).abs() < 1e-8,
"filtered evidence moved across converge(): {before} -> {after}"
);
}
/// Ingesting the same games one at a time must reach the same fixed point
/// as ingesting them in one call.
#[test]
fn ingestion_order_does_not_change_the_answer(games in pairs()) {
let batched = {
let mut h = history_from(&games);
let report = h.converge().unwrap();
prop_assert!(
report.converged,
"batched side stopped at {} iterations with step {:?}; comparing \
two fits that have not converged measures truncation, not order",
report.iterations,
report.final_step
);
h
};
let incremental = {
let mut h = History::builder()
.convergence(ConvergenceOptions {
max_iter: 20_000,
epsilon: 1e-10,
..ConvergenceOptions::default()
})
.build();
for (i, &(a, b)) in games.iter().enumerate() {
h.add_events([Event {
time: i as i64 + 1,
teams: smallvec![
Team::with_members([Member::new(KEYS[a])]),
Team::with_members([Member::new(KEYS[b])]),
],
outcome: Outcome::winner(0, 2),
}])
.unwrap();
}
let report = h.converge().unwrap();
prop_assert!(
report.converged,
"incremental side stopped at {} iterations with step {:?}",
report.iterations,
report.final_step
);
h
};
for key in KEYS {
let one = batched.current_skill(key);
let other = incremental.current_skill(key);
match (one, other) {
(Some(one), Some(other)) => {
prop_assert!(
(one.mu() - other.mu()).abs() < 1e-6
&& (one.sigma() - other.sigma()).abs() < 1e-6,
"{key}: batched mu={} sigma={}, incremental mu={} sigma={}",
one.mu(),
one.sigma(),
other.mu(),
other.sigma()
);
}
(None, None) => {}
_ => prop_assert!(false, "{key} present in only one history"),
}
}
}
}
+166
View File
@@ -0,0 +1,166 @@
//! `quality()` beyond two rating groups.
//!
//! The historical golden (two equal singletons) is asserted in
//! `src/lib.rs::tests::test_quality`. These cover the N-group generalisation,
//! which previously panicked with an out-of-bounds index at 3+ groups.
use trueskill_tt::{Gaussian, quality};
const BETA: f64 = 25.0 / 3.0 / 2.0;
fn rating(mu: f64, sigma: f64) -> Gaussian {
Gaussian::from_ms(mu, sigma)
}
#[test]
fn three_equal_groups_is_finite_and_in_range() {
let r = rating(25.0, 3.0);
let q = quality(&[&[r], &[r], &[r]], BETA);
assert!(q.is_finite(), "quality must be finite, got {q}");
assert!((0.0..=1.0).contains(&q), "quality out of range: {q}");
}
#[test]
fn quality_supports_many_groups() {
let r = rating(25.0, 3.0);
for n in 2..=8 {
let holders: Vec<[Gaussian; 1]> = (0..n).map(|_| [r]).collect();
let groups: Vec<&[Gaussian]> = holders.iter().map(|g| g.as_slice()).collect();
let q = quality(&groups, BETA);
assert!(q.is_finite(), "n={n}: quality must be finite, got {q}");
assert!((0.0..=1.0).contains(&q), "n={n}: out of range: {q}");
}
}
/// Equal-strength groups are the best-matched case: introducing a skill gap
/// must lower quality.
#[test]
fn imbalance_lowers_quality() {
let strong = rating(40.0, 3.0);
let average = rating(25.0, 3.0);
let balanced = quality(&[&[average], &[average], &[average]], BETA);
let lopsided = quality(&[&[strong], &[average], &[average]], BETA);
assert!(
lopsided < balanced,
"expected imbalanced quality {lopsided} < balanced {balanced}"
);
}
/// Quality is a property of the multiset of groups, not their order.
#[test]
fn quality_is_permutation_invariant() {
let a = rating(30.0, 2.0);
let b = rating(25.0, 3.0);
let c = rating(20.0, 4.0);
let forward = quality(&[&[a], &[b], &[c]], BETA);
let reversed = quality(&[&[c], &[b], &[a]], BETA);
assert!(
(forward - reversed).abs() < 1e-9,
"permutation changed quality: {forward} vs {reversed}"
);
}
#[test]
fn multi_player_groups_work() {
let r = rating(25.0, 3.0);
let q = quality(&[&[r, r], &[r, r], &[r, r]], BETA);
assert!(q.is_finite());
assert!((0.0..=1.0).contains(&q));
}
#[test]
fn uneven_group_sizes_work() {
let r = rating(25.0, 3.0);
let q = quality(&[&[r, r], &[r], &[r, r, r]], BETA);
assert!(q.is_finite(), "got {q}");
assert!((0.0..=1.0).contains(&q), "got {q}");
}
#[test]
#[should_panic(expected = "at least 2 rating groups")]
fn single_group_panics_with_clear_message() {
let r = rating(25.0, 3.0);
let _ = quality(&[&[r]], BETA);
}
#[test]
#[should_panic(expected = "at least 2 rating groups")]
fn zero_groups_panics_with_clear_message() {
let _ = quality(&[], BETA);
}
#[test]
#[should_panic(expected = "non-empty")]
fn empty_group_panics_with_clear_message() {
let r = rating(25.0, 3.0);
let _ = quality(&[&[r], &[]], BETA);
}
#[test]
fn history_predict_quality_supports_three_teams() {
use trueskill_tt::History;
let mut h = History::default();
h.record_winner(&"a", &"b", 1).unwrap();
h.record_winner(&"b", &"c", 2).unwrap();
let _ = h.converge().unwrap();
let q = h.predict_quality(&[&[&"a"], &[&"b"], &[&"c"]]).unwrap();
assert!(
q.is_finite(),
"3-team predict_quality must be finite, got {q}"
);
assert!((0.0..=1.0).contains(&q), "out of range: {q}");
}
/// `quality()` for N identical teams has a closed form, which pins the N-group
/// determinant path across the whole range rather than at a single golden.
///
/// For two identical single-player teams the standard result is
/// `sqrt(2b^2 / (2b^2 + s1^2 + s2^2))`. With the conventional parameters
/// (`sigma = 25/3`, `beta = 25/6`) that ratio is exactly `1/5`, and the N-group
/// generalisation is `(1/5)^((n-1)/2)` — one factor per adjacent pair.
///
/// The n=3 and n=5 values this produces (0.200 and 0.040) are also what the
/// `trueskill` Python package returns for the same configuration, so this
/// doubles as the cross-implementation check the README asked for.
#[test]
fn quality_of_identical_teams_follows_its_closed_form() {
let g = Gaussian::from_ms(25.0, 25.0 / 3.0);
let beta = 25.0 / 6.0;
for n in 2..=10usize {
let groups: Vec<Vec<Gaussian>> = (0..n).map(|_| vec![g]).collect();
let refs: Vec<&[Gaussian]> = groups.iter().map(Vec::as_slice).collect();
let got = quality(&refs, beta);
let expected = 0.2f64.powf((n - 1) as f64 / 2.0);
assert!(
(got - expected).abs() / expected < 1e-9,
"n={n}: quality {got}, closed form {expected}"
);
}
}
/// Spot-check against the two values the `trueskill` Python package is known
/// to produce for this configuration, stated as literals so a future change to
/// the closed-form reasoning above cannot quietly take these with it.
#[test]
fn quality_matches_the_reference_implementation() {
let g = Gaussian::from_ms(25.0, 25.0 / 3.0);
let beta = 25.0 / 6.0;
let three: Vec<Vec<Gaussian>> = (0..3).map(|_| vec![g]).collect();
let refs: Vec<&[Gaussian]> = three.iter().map(Vec::as_slice).collect();
assert!((quality(&refs, beta) - 0.200).abs() < 1e-9);
let five: Vec<Vec<Gaussian>> = (0..5).map(|_| vec![g]).collect();
let refs: Vec<&[Gaussian]> = five.iter().map(Vec::as_slice).collect();
assert!((quality(&refs, beta) - 0.040).abs() < 1e-9);
}
+165
View File
@@ -0,0 +1,165 @@
//! Converging, appending, and converging again must reach the same fixed point
//! as converging once over the whole event set.
//!
//! `tests/ingestion_equivalence.rs` covers a different question: it varies how
//! events are *batched* but converges only at the end. This file converges
//! between batches, which is the path a caller takes when it fits, serves for a
//! while, then ingests more.
//!
//! The property matters beyond ergonomics. It says `converge` reaches a fixed
//! point determined by the events, ratings and configuration alone — not by the
//! message state it started from. That is what makes a restored snapshot safe:
//! an inexact one cannot corrupt the answer, only cost an extra sweep. See #45.
use smallvec::smallvec;
use trueskill_tt::{ConvergenceOptions, Event, Gaussian, History, Member, Outcome, Team};
fn tight() -> ConvergenceOptions {
ConvergenceOptions {
max_iter: 5_000,
epsilon: 1e-12,
alpha: 1.0,
}
}
fn ev(a: &str, b: &str, time: i64) -> Event<i64, String> {
Event {
time,
teams: smallvec![
Team::with_members([Member::new(a.to_string())]),
Team::with_members([Member::new(b.to_string())]),
],
outcome: Outcome::winner(0, 2),
}
}
/// Ingest each chunk in turn, converging fully after every one.
fn fit_in_chunks(chunks: Vec<Events>) -> Vec<(String, Gaussian)> {
let mut h: History<i64, _, _, String> =
History::builder_with_key().convergence(tight()).build();
for chunk in chunks {
h.add_events(chunk).unwrap();
let report = h.converge().unwrap();
assert!(
report.converged,
"a chunk failed to converge, so any comparison would be measuring \
truncation rather than the fixed point; final step {:?}",
report.final_step
);
}
let mut skills: Vec<(String, Gaussian)> = h
.learning_curves()
.into_iter()
.map(|(k, curve)| (k, curve.last().unwrap().1))
.collect();
skills.sort_by(|a, b| a.0.cmp(&b.0));
skills
}
fn assert_same(a: &[(String, Gaussian)], b: &[(String, Gaussian)], what: &str) {
assert_eq!(a.len(), b.len(), "{what}: competitor count differs");
for ((ka, ga), (kb, gb)) in a.iter().zip(b) {
assert_eq!(ka, kb, "{what}: key order differs");
// Measured: 6.2e-13 for a later append, 8.9e-11 for an interleaved one.
// The bar is well clear of both but far under anything that would let a
// genuine divergence through.
assert!(
(ga.mu() - gb.mu()).abs() < 1e-8 && (ga.sigma() - gb.sigma()).abs() < 1e-8,
"{what}: {ka} differs — one-shot mu={} sigma={}, chunked mu={} sigma={}",
ga.mu(),
ga.sigma(),
gb.mu(),
gb.sigma()
);
}
}
type Events = Vec<Event<i64, String>>;
/// Two chunks of events: the first at times 0..20, the second at 100..120.
fn fixture() -> (Events, Events) {
let names = ["a", "b", "c", "d", "e"];
let mut seed = 7u64;
let mut rnd = move || {
seed ^= seed << 13;
seed ^= seed >> 7;
seed ^= seed << 17;
seed
};
let (mut early, mut late) = (Vec::new(), Vec::new());
for t in 0..40i64 {
let i = (rnd() % 5) as usize;
let mut j = (rnd() % 5) as usize;
if j == i {
j = (j + 1) % 5;
}
if t < 20 {
early.push(ev(names[i], names[j], t));
} else {
late.push(ev(names[i], names[j], 100 + t));
}
}
(early, late)
}
/// The ordinary case: new events are strictly later than everything fitted.
#[test]
fn appending_later_events_matches_a_single_fit() {
let (early, late) = fixture();
let all: Vec<_> = early.iter().cloned().chain(late.iter().cloned()).collect();
assert_same(
&fit_in_chunks(vec![all]),
&fit_in_chunks(vec![early, late]),
"append strictly later",
);
}
/// The case the design question suspected might be weaker: appended events
/// interleave with slices that are already fitted, so the append legitimately
/// revises the past. It is not weaker — Through Time revises the past on every
/// converge regardless, so there is nothing special about doing it in two steps.
#[test]
fn appending_interleaved_events_matches_a_single_fit() {
let (early, late) = fixture();
let all: Vec<_> = early.iter().cloned().chain(late.iter().cloned()).collect();
// Split by parity so the second chunk is back-dated into the first's range.
let first: Vec<_> = all.iter().step_by(2).cloned().collect();
let second: Vec<_> = all.iter().skip(1).step_by(2).cloned().collect();
let together: Vec<_> = first
.iter()
.cloned()
.chain(second.iter().cloned())
.collect();
assert_same(
&fit_in_chunks(vec![together]),
&fit_in_chunks(vec![first, second]),
"append interleaved",
);
}
/// Converging an already-converged history is a no-op, which is what makes a
/// restored snapshot worth having: the work is skipped rather than redone.
#[test]
fn re_converging_an_unchanged_history_costs_one_iteration() {
let (early, late) = fixture();
let all: Vec<_> = early.into_iter().chain(late).collect();
let mut h: History<i64, _, _, String> =
History::builder_with_key().convergence(tight()).build();
h.add_events(all).unwrap();
let first = h.converge().unwrap();
assert!(first.converged);
let again = h.converge().unwrap();
assert_eq!(
again.iterations, 1,
"a converged history should settle immediately, not re-grind"
);
assert!(again.converged);
}
+2 -2
View File
@@ -15,7 +15,7 @@ fn record_winner_builds_history() {
.build();
h.record_winner(&"alice", &"bob", 1).unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
let a_idx = h.lookup(&"alice").unwrap();
let b_idx = h.lookup(&"bob").unwrap();
@@ -48,7 +48,7 @@ fn record_draw_with_p_draw_set() {
.build();
h.record_draw(&"alice", &"bob", 1).unwrap();
h.converge().unwrap();
let _ = h.converge().unwrap();
assert!(h.lookup(&"alice").is_some());
assert!(h.lookup(&"bob").is_some());
+186
View File
@@ -0,0 +1,186 @@
//! Input validation must hold in **release**, where `debug_assert!` is gone.
//!
//! The engine guards itself with `debug_assert!`, which documents invariants
//! but vanishes in the profile users actually ship. Anything reachable from the
//! public API has to be rejected with an `InferenceError` instead, at the
//! boundary, rather than becoming NaN or an out-of-bounds panic deep inside
//! `run_chain`.
//!
//! `GameOptions` and `ConvergenceOptions` both have public fields, so the
//! eager asserts on `HistoryBuilder` do not cover the `Game` constructors —
//! a caller can build the options struct directly.
use smallvec::smallvec;
use trueskill_tt::{
ConstantDrift, ConvergenceOptions, Event, Game, GameOptions, Gaussian, History, InferenceError,
Member, Outcome, Rating, Team,
};
type R = Rating<i64, ConstantDrift>;
fn rating() -> R {
R::new(
Gaussian::from_ms(25.0, 25.0 / 3.0),
25.0 / 6.0,
ConstantDrift(0.0),
)
}
fn options_with_alpha(alpha: f64) -> GameOptions {
GameOptions {
convergence: ConvergenceOptions {
alpha,
..ConvergenceOptions::default()
},
..GameOptions::default()
}
}
/// `alpha == 0.0` leaves every EP update unapplied, so inference silently
/// returns the priors — the worst possible failure, since the output looks
/// entirely reasonable.
#[test]
fn ranked_rejects_a_zero_damping_factor() {
let (a, b) = (rating(), rating());
let err = Game::<i64, _>::ranked(
&[&[a], &[b]],
Outcome::winner(0, 2),
&options_with_alpha(0.0),
)
.expect_err("alpha = 0 must be rejected");
assert!(
matches!(err, InferenceError::InvalidParameter { name: "alpha", .. }),
"got {err:?}"
);
}
#[test]
fn ranked_rejects_an_out_of_range_damping_factor() {
let (a, b) = (rating(), rating());
for alpha in [-0.5, 1.5, f64::NAN] {
let err = Game::<i64, _>::ranked(
&[&[a], &[b]],
Outcome::winner(0, 2),
&options_with_alpha(alpha),
)
.expect_err("alpha out of (0, 1] must be rejected");
assert!(
matches!(err, InferenceError::InvalidParameter { name: "alpha", .. }),
"alpha={alpha}: got {err:?}"
);
}
}
#[test]
fn scored_rejects_a_bad_damping_factor() {
let (a, b) = (rating(), rating());
let err = Game::<i64, _>::scored(
&[&[a], &[b]],
Outcome::scores([21.0, 9.0]),
&options_with_alpha(0.0),
)
.expect_err("alpha = 0 must be rejected");
assert!(
matches!(err, InferenceError::InvalidParameter { name: "alpha", .. }),
"got {err:?}"
);
}
/// Already covered by `Game::ranked`, asserted here so the release-mode
/// guarantee is stated in one place.
#[test]
fn ranked_rejects_an_out_of_range_draw_probability() {
let (a, b) = (rating(), rating());
for p_draw in [-0.5, 1.0, 1.5] {
let options = GameOptions {
p_draw,
..GameOptions::default()
};
assert!(
Game::<i64, _>::ranked(&[&[a], &[b]], Outcome::winner(0, 2), &options).is_err(),
"p_draw={p_draw} must be rejected"
);
}
}
#[test]
fn scored_rejects_a_non_positive_noise() {
let (a, b) = (rating(), rating());
for score_sigma in [0.0, -1.0, f64::NAN] {
let options = GameOptions {
score_sigma,
..GameOptions::default()
};
assert!(
Game::<i64, _>::scored(&[&[a], &[b]], Outcome::scores([21.0, 9.0]), &options).is_err(),
"score_sigma={score_sigma} must be rejected"
);
}
}
/// A tie with no draw probability makes the truncation margin zero and the
/// two-sided update evaluate 0/0. Ingestion must refuse it.
#[test]
fn ingestion_rejects_a_tie_without_a_draw_probability() {
let mut h = History::builder().p_draw(0.0).build();
let err = h
.add_events(vec![Event {
time: 0,
teams: smallvec![
Team::with_members([Member::new("a")]),
Team::with_members([Member::new("b")]),
],
outcome: Outcome::draw(2),
}])
.expect_err("a tie with p_draw = 0 must be rejected");
assert!(
matches!(err, InferenceError::TieWithoutDrawProbability { .. }),
"got {err:?}"
);
}
/// `Outcome::scores_with_sigma` documents that a non-positive sigma is
/// accepted at construction and rejected at ingestion.
#[test]
fn ingestion_rejects_a_non_positive_per_event_score_sigma() {
for sigma in [0.0, -1.0, f64::NAN] {
let mut h = History::builder().build();
let err = h
.add_events(vec![Event {
time: 0,
teams: smallvec![
Team::with_members([Member::new("a")]),
Team::with_members([Member::new("b")]),
],
outcome: Outcome::scores_with_sigma([21.0, 9.0], sigma),
}])
.expect_err("a non-positive per-event sigma must be rejected");
assert!(
matches!(err, InferenceError::InvalidParameter { .. }),
"sigma={sigma}: got {err:?}"
);
}
}
/// Per-team weights must match that team's membership. The top-level length
/// checks in ingestion do not cover the inner dimension.
#[test]
fn ingestion_rejects_weights_that_do_not_match_their_team() {
let mut h = History::builder().build();
let mut team = Team::with_members([Member::new("a"), Member::new("b")]);
team.members[0].weight = 1.0;
let err = h
.event(0)
.team(["a", "b"])
.team(["c"])
// Three weights for a two-member team.
.weights([1.0, 1.0, 1.0])
.winner(0)
.commit()
.expect_err("a weight/member length mismatch must be rejected");
assert!(
matches!(err, InferenceError::MismatchedShape { .. }),
"got {err:?}"
);
}
+156
View File
@@ -0,0 +1,156 @@
//! `expected_variance_reduction`: which matchup best sharpens a given question.
use smallvec::smallvec;
use trueskill_tt::{
ConstantDrift, ConvergenceOptions, Event, History, InferenceError, Member, Outcome, Team,
UnknownKeys,
};
type H = History<i64, ConstantDrift, trueskill_tt::NullObserver, &'static str>;
fn round(a: &'static str, b: &'static str, sa: f64, sb: f64) -> Event<i64, &'static str> {
Event {
time: 1,
teams: smallvec![
Team::with_members([Member::new(a)]),
Team::with_members([Member::new(b)]),
],
outcome: Outcome::scores([sa, sb]),
}
}
fn base() -> Vec<Event<i64, &'static str>> {
vec![
round("a", "b", 5.0, 2.0),
round("a", "c", 6.0, 1.0),
round("b", "c", 4.0, 3.0),
round("c", "d", 2.0, 1.0),
round("a", "d", 7.0, 2.0),
]
}
fn fit(extra: Option<Event<i64, &'static str>>, policy: UnknownKeys) -> H {
let mut h: History<i64, _, _, &'static str> = History::builder()
.mu(0.0)
.sigma(6.0)
.beta(1.0)
.score_sigma(2.0)
.drift(ConstantDrift(0.0))
.unknown_keys(policy)
.convergence(ConvergenceOptions {
max_iter: 20_000,
epsilon: 1e-13,
alpha: 1.0,
})
.build();
let mut ev = base();
if let Some(e) = extra {
ev.push(e);
}
h.add_events(ev).unwrap();
let _ = h.converge().unwrap();
h
}
/// The closed form must equal what actually happens if the matchup is played.
/// This is the assertion that makes the whole call trustworthy: a wrong
/// acquisition function returns plausible numbers and quietly picks worse
/// matchups forever.
#[test]
fn the_closed_form_matches_an_actual_refit() {
let h = fit(None, UnknownKeys::Reject);
let target: Vec<(&&str, f64)> = vec![(&"a", 1.0), (&"b", -1.0)];
let before = h.posterior_of(&target).unwrap().sigma().powi(2);
for (x, y) in [("a", "b"), ("c", "d"), ("a", "c"), ("b", "d")] {
let predicted = h
.expected_variance_reduction(&[&[&x], &[&y]], &target)
.unwrap();
let after = fit(Some(round(x, y, 3.0, 1.0)), UnknownKeys::Reject);
let actual = before - after.posterior_of(&target).unwrap().sigma().powi(2);
assert!(
(predicted - actual).abs() / actual.abs() < 1e-9,
"{x} vs {y}: predicted {predicted}, actual {actual}"
);
}
}
/// The reduction cannot depend on the score, because for a Gaussian likelihood
/// the posterior variance update is data-independent. This is why the call
/// needs no expectation despite its name.
#[test]
fn the_outcome_does_not_change_the_reduction() {
let target: Vec<(&&str, f64)> = vec![(&"a", 1.0), (&"b", -1.0)];
let h = fit(None, UnknownKeys::Reject);
let before = h.posterior_of(&target).unwrap().sigma().powi(2);
let mut seen = Vec::new();
for (sa, sb) in [(3.0, 1.0), (100.0, -50.0), (0.0, 0.0)] {
let after = fit(Some(round("c", "d", sa, sb)), UnknownKeys::Reject);
seen.push(before - after.posterior_of(&target).unwrap().sigma().powi(2));
}
for w in seen.windows(2) {
assert!(
(w[0] - w[1]).abs() < 1e-12,
"variance reduction moved with the observed score: {seen:?}"
);
}
}
/// The point of the call: it must rank candidate matchups usefully. Playing the
/// pair you are trying to separate helps most; an unrelated pair helps least.
#[test]
fn it_ranks_candidates_by_how_much_they_answer_the_question() {
let h = fit(None, UnknownKeys::Reject);
let target: Vec<(&&str, f64)> = vec![(&"a", 1.0), (&"b", -1.0)];
let direct = h
.expected_variance_reduction(&[&[&"a"], &[&"b"]], &target)
.unwrap();
let unrelated = h
.expected_variance_reduction(&[&[&"c"], &[&"d"]], &target)
.unwrap();
assert!(direct > 0.0 && unrelated > 0.0);
assert!(
direct > 5.0 * unrelated,
"playing the target pair should dominate: {direct} vs {unrelated}"
);
}
/// A matchup between two competitors nobody has seen still teaches something
/// about them, but nothing about a target that does not involve them.
#[test]
fn an_unrelated_unseen_matchup_teaches_nothing_about_the_target() {
let h = fit(None, UnknownKeys::Prior);
let target: Vec<(&&str, f64)> = vec![(&"a", 1.0), (&"b", -1.0)];
let reduction = h
.expected_variance_reduction(&[&[&"stranger"], &[&"nobody"]], &target)
.unwrap();
assert!(
reduction.abs() < 1e-12,
"an unseen pair shares nothing with the target: {reduction}"
);
}
#[test]
fn shape_errors_are_reported() {
let h = fit(None, UnknownKeys::Reject);
let target: Vec<(&&str, f64)> = vec![(&"a", 1.0), (&"b", -1.0)];
assert!(matches!(
h.expected_variance_reduction(&[&[&"a"]], &target),
Err(InferenceError::MismatchedShape {
expected: 2,
got: 1,
..
})
));
assert!(matches!(
h.expected_variance_reduction(&[&[&"a"], &[&"ghost"]], &target),
Err(InferenceError::UnknownKey { .. })
));
}