fix: correct erfc_inv's sign error and keep evidence in log space
A systematic scan for precision defects, following the tail-precision work in7341669. Three findings; the first is a correctness bug in a released version. 1. `erfc_inv`'s initial guess had the wrong sign. Numerical Recipes' `inverfc` uses -0.70711 as the leading coefficient; this used +FRAC_1_SQRT_2. Since `rational - t` is negative, that put Newton on the mirror image of the root, and three fixed iterations could not cross back. Measured against exact standard-normal quantiles: p_draw old rel err new rel err 0.50 1.46e-7 8.40e-8 0.90 1.02e-1 8.63e-9 0.95 3.06e-1 1.91e-8 0.99 8.05e-1 5.89e-9 `compute_margin` inherited it, so the draw margin was wrong for any `p_draw` above about 0.6 and *non-monotone* above 0.9 — it ran 0.674, 1.476, 0.503, 0.982 as p_draw went 0.5, 0.9, 0.99, 0.999. A history configured for a 0.99 draw rate was being fitted at 0.385. Note it was slightly wrong everywhere, not only in the tail. 2. `MarginFactor` computed a density and clamped it. `pdf` underflows past ~38 sigma, so `ln` of the clamped zero reported -708 nats however far out the score actually was: 4292 nats adrift at 100 sigma, and unbounded beyond. This is the same defect as the one fixed in `TruncFactor`, one file over, on the scored-outcome path. 3. `TruncFactor` still bottomed out past ~38 sigma even after7341669removed the cancellation, because the linear probability itself underflows there. 2 and 3 are fixed the same way: factors cache a *log* evidence, built from new `ln_pdf`, `ln_sf` and `ln_interval` helpers that factor the shared exponential out analytically via the `erfcx` added earlier. Nothing underflows, at any separation. One golden moved. `test_1vs1vs1` runs at `p_draw = 0.5`, so it goes through `compute_margin`; its 1e-6-place values shifted. Verified as movement *toward* analytic truth by comparing both the old and new inverse against exact quantiles, per the goldens policy in CLAUDE.md — not re-baselined on faith. Two test tolerances are asserted at 1e-6 rather than tighter because above x = 2 `erfcx` uses a continued fraction accurate to ~1e-15 while `erfc` carries ~1e-7, so the log path is the more accurate of the two and they part company at `erfc`'s error. That floor is tracked in #41. Refs #41 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011hcFjNDmHXZF8URGLku5zZ
This commit is contained in:
+10
-4
@@ -46,8 +46,8 @@ impl DiffFactor {
|
||||
/// reaches.
|
||||
pub(crate) fn log_evidence(&self) -> f64 {
|
||||
match self {
|
||||
Self::Trunc(f) => f.evidence_cached.unwrap_or(1.0).ln(),
|
||||
Self::Margin(f) => f.evidence_cached.unwrap_or(1.0).ln(),
|
||||
Self::Trunc(f) => f.log_evidence_cached.unwrap_or(0.0),
|
||||
Self::Margin(f) => f.log_evidence_cached.unwrap_or(0.0),
|
||||
}
|
||||
}
|
||||
|
||||
@@ -733,9 +733,15 @@ mod tests {
|
||||
let c = p[2][0];
|
||||
|
||||
// T1 ULP shift: mu rounds to 25.0 (was 24.999999) under natural-parameter storage.
|
||||
//
|
||||
// The 1e-6-place values moved when `erfc_inv`'s sign error was fixed:
|
||||
// this case runs at `p_draw = 0.5`, so it goes through `compute_margin`,
|
||||
// and the margin is now 8.4e-8 from the exact quantile where it was
|
||||
// 1.46e-7. Verified as movement *toward* analytic truth, not a
|
||||
// regression — see `erfc_inv_matches_known_quantiles`.
|
||||
assert_ulps_eq!(a, Gaussian::from_ms(25.0, 6.092561), epsilon = 1e-6);
|
||||
assert_ulps_eq!(b, Gaussian::from_ms(33.379314, 6.483575), epsilon = 1e-6);
|
||||
assert_ulps_eq!(c, Gaussian::from_ms(16.620685, 6.483575), epsilon = 1e-6);
|
||||
assert_ulps_eq!(b, Gaussian::from_ms(33.379315, 6.483576), epsilon = 1e-6);
|
||||
assert_ulps_eq!(c, Gaussian::from_ms(16.620685, 6.483576), epsilon = 1e-6);
|
||||
}
|
||||
|
||||
#[test]
|
||||
|
||||
Reference in New Issue
Block a user