Skip to content

Add supplied frozen coordinate chart with executed validation gates - #263

Merged
junpenglao merged 13 commits into
mainfrom
geodesic-frozen-transport
Sep 7, 2026
Merged

junpenglao merged 13 commits into
mainfrom
geodesic-frozen-transport

Conversation

@junpenglao

@junpenglao junpenglao commented Sep 7, 2026 •

Copy link
Copy Markdown
Member

Summary

Adds a supplied (frozen) affine-flow coordinate chart to tuningfork, together with the
algebra needed to validate it. The chart is handed complete parameters and evaluates them; it
learns nothing.

Not in this PR: no fitter, no chart-selection rule, no warmup descriptor, no execution
recipe, no registry entry, no sampler integration, and no changes to BlackJAX or to any existing
tuningfork file. Every file here is new. No sampling, mixing, efficiency or posterior-accuracy
claim is made anywhere, and no chain is run.

The chart

In preconditioned coordinates z = L^{-1}(q - center), a constant affine field

V(z) = alpha * z + a * (h . z) + c

is constrained at construction to ||h|| = 1, h . c = 1, h . a = -alpha. Those three
constraints make the clock rate identically one, d/ds (h . z) = 1, which yields three
consequences exact in real arithmetic:

  1. the last chart coordinate (the clock) advances at unit rate, so it is readable as t = h . z;
  2. the inverse is closed-form — flowing back by t lands on the section h . z = 0, with no
    iteration and no root find;
  3. log|det J| = alpha (d-1) t + log|det L| is the exact log-Jacobian.

These are identities of the map, not statements about its floating-point evaluation. The
implementation's finite-precision behaviour is a separate matter and is treated as such below.

CONVENTION_VERSION identifies the coordinate and series conventions the module fixes: clock-last
ordering, the reflector orientation, and the per-dtype series crossover.

Delivered surface

The chart implementation is private; only the convention identifier is exported. The private
constructor does two distinct things and does not conflate them:

  • Projects — h is normalised, and c and a are projected so that h . c = 1 and
    h . a = -alpha hold by construction.
  • Refuses — representation preconditions are enforced rather than repaired: non-zero h;
    matching, non-zero scale; declared shapes for a, c, center and a scalar alpha; finite
    a, c, alpha, center; finite low-rank basis and eigenvalues, checked before any spectral
    work; strictly positive eigenvalues; and orthonormality of the spectrally active basis
    columns.

Orthonormality is scoped to active columns deliberately. Columns with unit eigenvalue contribute
exactly zero to the preconditioner at every power and zero to the log-determinant, so requiring
orthonormality of them would refuse legitimate inputs carrying a neutral direction.

log_det is a log-absolute determinant, so a negative scale entry is supported: it flips
orientation without changing the volume element.

phi serves float32 and float64 only. Half precisions raise rather than falling back. The
reason is directional: a coarser dtype wants a larger crossover, so falling back to the float64
constant moves the wrong way, and a silent fallback would be worse than no support.

What is checked

  • Algebraic identities — the structural constraints; the closed-form log-Jacobian against
    slogdet of the AD Jacobian, near and far from the origin; the two-sided inverse; the supplied
    score against a reference; and invertibility of the score pullback. Diagonal and low-rank
    preconditioners, including an active factor and a neutral column.
  • Gates — a valid chart, a wrong log-determinant, a wrong score, non-finite providers, and a
    chart built without normalising h. Gate success requires finiteness as well as agreement:
    a reject-if-error-exceeds comparison does not reject NaN, because NaN > tolerance is false.
    Explicit finiteness checks handle non-finite values.
  • Special cases — a pure reflection, an exercised low-rank factor, a non-normal generator, and
    a table of representation violations against a baseline that is separately asserted to
    construct, so a mistake in the fixture cannot make every row vacuously pass.
  • phi — values and derivatives against an independent 60-digit decimal oracle at named
    points: zero, small arguments, either side of each crossover, and a large-argument case where
    the direct form cancels. These are agreements at those points. No accuracy bound, floor or
    optimality is claimed for the crossovers or for either branch, and no global property is
    inferred from the sampled points.
  • The clock residual — recorded as a runtime diagnostic and asserted only to be available and
    clean where the chart is well conditioned.

One methodological point the tests encode:

  • The score reference is not hooked. Differentiating a custom_jvp density to test the score
    it supplies returns the rule under test, so the reference is the plain composed expression. For
    the same reason the inverse is tested as a supplied map: AD cannot manufacture a global inverse
    from the forward map alone.

Scope of the claims

The funnel case checks that, for that target's own scaling map, the transformed clock score is
affine in the clock with no section dependence, so the density factorises into a Gaussian clock
and a section marginal. The premise for that conclusion is stated in full in the design note: it
requires the identity to hold on the full real line, with a constant positive precision,
on a global Cartesian chart with global support. For a supplied chart those premises are
discharged by the algebra and the evaluations corroborate them; for a fitted chart they would
remain open. A residual measured on any finite set of points establishes none of them.

The clock residual is a component-specific detector, not a certificate. It observes error in
the clock coordinate and is blind to anything leaving that coordinate intact, so it certifies
neither the score nor the Jacobian; those are covered by the algebraic checks. A projection that
removes the residual is described in the design note and is not shipped — it would change the
implemented map and needs its own verification.

The overflowing-basis case establishes that a basis with finite entries and a non-finite derived
Gram is refused. It does not establish coverage of the non-finite-comparison path specifically:
on this backend that Gram reduces to an infinity, which a comparison-only gate would also reject.
The guard rests on the reasoning recorded beside it, not on that test.

Verification

Final-head CI at fa1f3fb62711726f8273a25b16cb414c41c5ba71:

job result
pre-commit success
test-fast (py3.13) 1977 passed, 6 skipped in 383.99 s
test-slow 142 passed, 2 skipped in 870.07 s

The modules added here trace JAX and run in seconds, so they are marked slow and run in the
slow job. Earlier heads of this branch carried different counts and, in two cases, failures; those
results belong to those commits and are not restated here. A local ordered single-worker check of
the added modules together with two previously affected cases was run at an earlier commit, and a
single-test check was run against a working tree that was not this commit; neither is offered as
evidence about this head.

junpenglao and others added 13 commits September 7, 2026 13:55
Validation-first artifact for a frozen affine-flow coordinate chart: the
supplied chart plus the algebra needed to check it. No fitter, no selection
rule, no warmup descriptor, no execution recipe, no registry wiring.

The chart family is constrained at construction to |h|=1, h.c=1, h.a=-alpha,
which makes the clock rate identically one and therefore makes the inverse and
the log-Jacobian exact in real arithmetic rather than fitted surrogates.

Finding: a non-zero Stein residual is the USEFUL case, not a defect.
grad(log pi).V + div V is exactly d/dt log pi_chart. If that equals
-kappa*t + beta globally with kappa>0 and no section dependence, integrating
gives a Gaussian clock independent of the section. An exactly zero residual
would instead mean a flat, improper clock. Earlier reasoning here treated zero
as the target and conflated the fitter's regression residual (which carries
affine-in-clock slack) with the deployed field's residual.
Fix: test_chart_refuters.py pins all three quantities separately on the exact
funnel -- raw residual -t/9, slack-absorbed regression residual ~0, and the
resulting structure via the decisive discriminator d_s d_t log pi_chart == 0
with constant clock curvature -1/9.

Finding: differentiating a custom_jvp density to test its own supplied score is
circular, and AD cannot manufacture a global inverse from forward alone.
Fix: every score gate references the plain, non-hooked
native_logdensity(forward(y)) + log_det(y); the inverse is tested as a supplied
map.

Finding: predicting which gate catches a mutant is not the same as executing it.
Rescaling h was expected to fail the determinant gate and does not -- the
Householder reflector is built from the normalised direction, so the section
stays orthogonal to h and forward is unchanged to rounding.
Fix: all seven mutants are executed through the real helper and the expected
gate pattern is asserted exactly. M4 and M5 pass the log-determinant and
round-trip gates and are caught only by the independent score gate, which is
why that gate is mandatory.

Finding: no single phi series crossover serves both supported dtypes, and two
measurement traps hide this -- a float128 closed-form oracle cancels worse than
the code under test at |x|~1e-11, and a numpy stand-in disagrees with the JAX
path in float32.
Fix: hybrid extended-precision oracle, comparison against the shipped code
path, and a per-dtype threshold (float64 0.1 -> 3.8e-15; float32 0.3 ->
8.8e-07). Both value and derivative error are measured.

Finding: the chart fails silently. The score is wrong by O(1) relative at
alpha*clock ~ 20 while every component is still finite, and only goes non-finite
around alpha*clock ~ 40, so a reject-on-NaN guard cannot catch it.
Fix: test_chart_conditioning.py records the band with a float32-vs-float64
ladder on the shipped code path, and labels exp(alpha*clock) an empirical
indicator explicitly, not a certified bound.

Also fixed during development: push_score used the wrong preconditioner power
(+0.5 instead of -0.5), which the diagonal case cannot detect and the low-rank
algebra test caught.

Verification: 40 tests pass (pytest tests/transport, CPU, ~15 s); pre-commit
clean on the ten added files. No existing tuningfork or BlackJAX file changed.
No sampling, no campaign, and no mixing, efficiency or posterior-accuracy claim.

Refs: local Issue#1200, Experiment#925, Experiment#926, Belief#1371, Belief#1372.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FBFCNAvKSxsz5gb1TWcyrE
Finding: the Gaussian-clock statement was written as though finite evaluations
established it. They do not. The implication needs three conditions stated
together -- the identity d/dt log pi_chart = -kappa*t + beta holding on the FULL
real line in t, kappa CONSTANT and positive, and a global Cartesian chart with
global support -- and only under those does kappa <= 0 fail to normalise. The
previous wording also read as a universal claim that a non-positive precision is
always improper, which is false on a bounded or otherwise restricted clock
domain.
Fix: state the analytic implication with its premise explicit, note that for the
supplied chart tested here the premise is discharged by the algebra while the
evaluations only corroborate it, and record that for a LEARNED chart the premise
is open -- a training-support fit says nothing about the full line, nothing about
whether kappa is constant off that support, and does not guarantee kappa > 0.
Added "constant" to the list of properties a small training residual fails to
establish; it was missing alongside positivity.

No behaviour change: documentation and test docstrings only. 40 tests still pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FBFCNAvKSxsz5gb1TWcyrE
Finding: the shipped docstrings and design doc carried internal review
vocabulary -- numbered tiers, and labels naming the two earlier exploratory
implementations as investigation arms. These are meaningful only inside the
private review process and are noise, or worse, misleading provenance, to a
reader of the repository.
Fix: name each suite by what it establishes rather than by a tier number; refer
to "earlier exploratory implementations" instead of arms; rename the comparison
constant to PRIOR_CROSSOVERS. No test logic, tolerance, constant or assertion
changed -- the phi selection still compares the shipped path against both prior
crossovers and still asserts the chosen per-dtype threshold wins.

40 tests still pass; pre-commit clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FBFCNAvKSxsz5gb1TWcyrE
…uite

Independent paired review found the chart mathematics clean and the validation
around it defective. This repairs the validation and the two real code defects,
and removes machinery that was re-litigating past decisions instead of
protecting anything.

Finding: the gates compared `float(error) > ATOL`, which is False for NaN, so a
provider returning NaN everywhere passed all four gates. The suite's central
claim -- that these gates catch wrong charts -- had a hole big enough to admit
a chart that computes nothing.
Fix: finiteness is part of gate success. Four providers replace the seven-mutant
table: a valid chart, a wrong log-determinant, a wrong score, and non-finite
controls for NaN/+Inf/-Inf. The pass/fail vector over a mutant family was never
a product contract -- which gate a mutation perturbs is a property of how it was
injected -- so it is gone rather than rebuilt.

Finding: all five test modules mutated jax_enable_x64 at MODULE level, i.e. at
collection, before the shared conftest guard snapshots it. The guard then
restored the already-corrupted value forever after, leaking float64 into the
whole session and failing two groundtruth tests in files this branch never
touched.
Fix: a module-scoped fixture that enables x64 at setup and restores the prior
value in a finally. No collection-time mutation and no conftest change. The
modules trace JAX and run in seconds, so they are marked slow, not fast.

Finding: `log_det` used `log(scale)` while claiming log|det|, returning NaN for a
chart whose forward, inverse and score were all exact.
Fix: log-absolute determinant, so signed scale is supported consistently rather
than left implicit.

Finding: "correct by construction" covered only what make_chart projects. Zero h,
mismatched low-rank arguments, non-positive eigenvalues and non-orthonormal
factors all still produced silently wrong charts -- a non-orthonormal active
basis broke the inverse by 2.3 with no error raised.
Fix: a stated, enforced input contract that refuses what it cannot repair.
Orthonormality is required only of SPECTRALLY ACTIVE columns: neutral columns
contribute exactly zero at every power, so demanding it of them would reject
legitimate inputs carrying a neutral direction.

Finding: the phi "measured floors" were maxima over a sampled grid, and review
found points exceeding two of them. A grid maximum is a lower bound on the worst
case, so promoting one to a "floor" was wrong in kind. Worse, the f32 crossover
was selected because the table recorded 0.3 and 1.0 as tied -- and that tie was
itself a sampling artifact.
Fix: no floor, bound or accuracy is claimed for phi at any dtype. The competition
tests and the shipped-implementation copy they compared against are removed; a
regression suite should not re-run a past threshold decision on every invocation.
What remains checks values AND derivatives against an independent 60-digit
decimal oracle at points chosen to fail differently: zero, small arguments, both
sides of each crossover, and a large-argument cancellation.

Finding: the series branch received arbitrarily large x under `where`, so its
Horner recurrence overflowed and an inf-inf tangent in the unselected branch
poisoned the selected one -- jax.grad returned NaN where the value was exact.
Fix: clamp the series argument as the direct denominator already was, and assert
gradient finiteness where the value is finite. Unsupported half precisions now
raise instead of silently taking the float64 threshold, which measurement showed
is the worst available choice for them.

Finding: every gate tolerance was absolute and calibrated at |y| ~ 1, so the
UNMUTATED chart failed its own log-det gate at |y| ~ 100 (2.669e-11 against
1e-11) at a relative error of 1.08e-13. The suite would reject a correct
implementation -- the failure direction that costs most later.
Fix: relative tolerances throughout, and a far-field case that an absolute bound
would fail.

Also: the "identity" refuter's forward was the identity map and the "rotated"
refuter's low-rank factor was exactly I, so neither exercised what it named; both
are replaced with cases that do. The conditioning module no longer asserts that
the implementation degrades -- pinning today's rounding behaviour would block
tomorrow's repair -- and instead records the clock residual |h.z - t| as an O(d)
runtime diagnostic. The claim that a guard could not be inferred from returned
values is withdrawn; that residual is exactly such a guard. The clock projection
that repairs it is documented, NOT shipped: it changes the implemented map and
needs its own verification. CONVENTION_VERSION stays v1 -- these repairs change
results only where the old code returned NaN or served an unsupported dtype.

NOT VERIFIED LOCALLY: the machine-wide numerical lock was held throughout
(flock exit 75, twice, no retry loop). This batch is lint-clean but unexecuted;
remote CI is the first execution, and its test-slow job runs the new tests
followed by the two previously failing groundtruth tests in one process, which
is the isolation check the x64 fix needs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FBFCNAvKSxsz5gb1TWcyrE
Finding: the independent phi oracle built its series with `xd ** k`, and
`Decimal(0) ** 0` raises InvalidOperation. Since zero is one of the named test
points -- and the most worth testing, being where the direct form cancels
hardest -- the oracle would have raised rather than compared, erroring the test
instead of failing it. The local numerical lock was held, so this was found by
reading the new code rather than by running it.
Fix: accumulate powers iteratively; no exponentiation anywhere in the oracle.

Validated in pure Python, no JAX: phi1(0)=1, phi2(0)=0.5, phi1'(0)=0.5,
phi2'(0)=1/6 exactly, and agreement with float references at 0.3, +/-12. At
x=1e-8 the oracle gives phi2 = 0.5000000016667 against the float form's
0.500000008431 -- the oracle is correct and the float form is the one cancelling,
which is the reason it exists.

Still unexecuted under JAX: the numerical lock remains held by another station.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FBFCNAvKSxsz5gb1TWcyrE
Finding: the design doc and the conditioning module presented |h.z - t| as if a
clean reading meant a healthy chart. It does not. The residual observes error in
the clock component and is blind to anything that leaves h.z intact, so it
certifies neither the score nor the log-Jacobian.
Fix: both files now say so. The residual detects one specific failure; accuracy
of the score and Jacobian is covered by the algebra tests against an independent
reference, not by this diagnostic.

Also retracted, source-refuted: the previous commit message called CI's test-slow
the ordered isolation check for the x64 fix. .github/workflows/test-slow.yml
runs `pytest tests -m slow -n auto`, so there is no guaranteed process or
ordering. CI covers collection-time behaviour and distributed regression; the
ordered check is a separate scoped single-worker run, still queued behind the
shared numerical lock.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FBFCNAvKSxsz5gb1TWcyrE
A source-only delta review of the repair itself found eight items. The first
would have failed the suite deterministically, and my ordered isolation check
did not cover that file -- so the batch I reported as lint-clean was not green.

Finding: test_gradients_stay_finite_where_the_value_is_finite asserted the VALUE
is finite at large POSITIVE x, where it provably is not. phi1(x) = (e^x-1)/x is
unbounded, so expm1 overflows above 88.7 (float32) / 709.8 (float64) and the
assertion fails at the first positive entry in both dtypes. The test's own name
stated the correct scope; its body asserted more.
Fix: probe only large negative arguments, where the value is small and exact and
where the repaired gradient bug actually lived, and pin the positive overflow
separately as a property of the function so it is not later mistaken for a bug.

Finding: the orthonormality gate used an absolute 1e-8 Gram tolerance evaluated
in the input dtype. A genuinely orthonormal float32 basis carries Gram error
around sqrt(d)*eps32 ~ 3e-7, so the contract rejected VALID float32 input. No
test built a float32 low-rank chart, which is why it went unseen.
Fix: scale the tolerance with dtype eps and rank; real violations are O(0.1),
far above it. Added the missing float32 low-rank case.

Finding: two gaps in the enumerated input contract -- center size was never
checked against h, and finiteness was checked for h and scale but not for a, c,
alpha or center, so a NaN in c propagated through the projection and silently
built a NaN chart.
Fix: both checked and refused.

Finding: test_projection_is_a_no_op_on_a_well_conditioned_point could not fail.
The displacement is h(t - h.z), whose largest component is the residual times
max|h_i| <= residual because ||h|| = 1, so the assertion restates an identity of
the expression and holds for any chart including a broken one -- the same defect
class this branch removed from the mutant tests two commits ago.
Fix: deleted, with a comment recording why. The projection is documented, not
shipped; testing it here tested nothing.

Finding: test_omitted_normalisation_is_rejected described itself as bypassing
make_chart's projection but actually rescaled h on a finished chart, leaving u,
a and c consistent with the normalised direction -- precisely the criticism this
branch made of the previous M6 mutant, reintroduced while fixing it.
Fix: the chart is now built the way forgetting the normalisation would build it,
with the projection and the Householder reflector both derived from the raw h.

Finding: the gradient-poisoning mechanism was named wrong in two places. Horner's
s*x + c produces inf, never inf - inf; the poisoning is 0 * inf in the reverse
pass, where the unselected branch's zero cotangent multiplies a saved infinite
intermediate. That sentence is the one explaining why the clamp works, so being
wrong in it matters more than its length suggests.
Fix: corrected in the module and the test.

Also: make_chart validates with Python control flow and is host-side only, not
jit/vmap-traceable -- now stated, since the previous implementation was traceable
and Chart is a pytree that does cross those boundaries. Removed a docstring
pointer to a table this branch had already deleted.

Still unexecuted at the time of writing: the transport suite has never run under
JAX. The ordered isolation check at 8e0adc7 covered five tests and verified that
sequence, not this.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FBFCNAvKSxsz5gb1TWcyrE
Finding: the previous commit claimed to scale the Gram tolerance by dtype and
rank. It did not. The edit was applied with a string replacement whose target
had already been reformatted onto multiple lines by black, so the replacement
matched nothing and wrote a silent no-op while the commit message described the
fix as done. The hardcoded 1e-8 was still in the file.
Fix: applied for real, and asserted the match before writing plus the result
after writing, so a non-matching patch fails loudly instead of passing quietly.

This was caught by the first execution of the transport suite: a valid float32
low-rank chart was rejected with max |U^T U - I| = 1.192e-07, which is the
sqrt(d)*eps32 Gram error an orthonormal float32 basis legitimately carries. The
test that caught it was added in the same commit as the fix that failed to
apply, from the delta review -- so the review's finding was confirmed by its own
regression test catching the repair's failure to take effect.

Suite status at the previous commit: 1 failed, 38 passed. The single failure is
this one.

Process note recorded because it nearly shipped: a wrapper script around pytest
reported exit 0 while pytest itself exited 1, because the wrapper's last
statement was an echo. The real status was only visible because the script
printed it immediately after the pytest line. An unpiped wrapper exit code is
not automatically the command's exit code.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FBFCNAvKSxsz5gb1TWcyrE
Finding: make_chart still had a NaN-blind path, in the code written to enforce
the contract. lr_basis had no finiteness check, so a NaN in an active column
makes the Gram residual NaN, and `if off > tol` is False for NaN -- the gate
accepts a basis it cannot evaluate. This is the same comparison defect this
branch repaired in its own test gates two commits ago, reproduced in the fix for
it. A positive infinity in lr_eigenvalues likewise satisfies a bare `> 0` test.
Fix: require finite lr_basis and lr_eigenvalues BEFORE any Gram or spectral
work, for every column including neutral ones -- 0 * NaN is NaN, so a neutral
column is not numerically inert. One NaN-basis refusal and one infinite-
eigenvalue refusal cover it; the valid float32 and finite-neutral controls are
kept.

Finding: a, c and alpha shapes were never checked, so broadcasting could
construct a chart outside the declared family -- a scalar `a`, or an `alpha`
carrying a trailing axis, both produce well-formed arrays.
Fix: shapes checked explicitly against h, and alpha required to be a scalar.

Finding: the tolerance comment claimed "real violations are O(0.1), orders above
this". That is not something the code knows. A violation can be arbitrarily
small, and one below the rounding allowance is simply not detected.
Fix: the comment now says what the tolerance is -- a rounding allowance, not a
separation guarantee.

Finding: the phi tests described expm1's overflow at ~88.7 / ~709.8 as the point
where "the function itself is unbounded", and treated it as a boundary the test
proved. Both are wrong. phi1(x) = (e^x - 1)/x is finite at every finite real x,
and because the division follows the exponential there are arguments where the
exact result is representable while the intermediate expm1 has already
overflowed. The limit observed belongs to the evaluation order, not to the
function.
Fix: renamed and rewritten to record where THIS IMPLEMENTATION stops returning
finite values, stating explicitly that it is not the representability boundary
and that no boundary is located or claimed. The 2*limit arguments stay; they are
comfortably past the limit rather than evidence of where it is.

Finding: the oracle set decimal precision with getcontext().prec = 60 at module
scope, mutating global decimal state during collection and leaking it into every
later test in the session. That is the same defect class as the x64 collection
leak this branch already repaired, in a different global, introduced while
repairing it.
Fix: local context inside the oracle.

Process note: the previous commit's fix silently did not apply because a string
replacement matched nothing after black had reformatted its target. Every edit
in this commit asserts its match before writing and verifies the result after.

Not verified locally: no compute slot was granted for this commit; final-head CI
executes it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FBFCNAvKSxsz5gb1TWcyrE
Finding: the orthonormality gate was still NaN-blind, and this is the fourth
instance of that defect on this branch. The previous three were closed by
checking finiteness before the comparison. That discipline does not reach here,
because `off` is not an input: it is DERIVED from the inputs after they have all
been validated. An off-diagonal Gram entry is a signed sum, so a basis with
every entry finite -- around sqrt(dtype max), 1.8e19 in float32 or 1.3e154 in
float64, with at least two active columns of mixed sign -- yields
(+inf) + (-inf) = NaN, which jnp.max propagates. Written as `off > tol` the gate
then accepts a basis it could not evaluate, since NaN > tol is False. Diagonal
entries are sums of squares and can only overflow to +inf, which raises
correctly; the reachable path is specifically the off-diagonal.
Fix: `if not (off <= tol):`. The structural lesson is recorded in the code
because it generalises past this case -- finiteness-before-comparison protects
inputs and cannot protect a quantity computed after the inputs are cleared, so
the only durable guard on a derived value is the comparison's polarity.

The regression test asserts the basis entries are finite and the Gram is NaN
BEFORE asserting the refusal, so it cannot silently stop exercising the path if
the arithmetic underneath it changes.

Finding: the SERIES_THRESHOLD docstring still asserted an optimality verdict --
"the optimum is dtype-dependent", "a single constant is measurably wrong for one
of them" -- which is the deleted "measured floor" claim in another costume. It
had no test behind it, since the threshold comparison was deliberately removed
for resting on a sampled grid, and it sat two screens below a module docstring
disclaiming exactly that kind of statement. An outstanding measurement also
contradicts it.
Fix: keep the mechanism, drop the verdict. The two branches' errors move in
opposite directions in |x|, so a useful crossover depends on the dtype's
epsilon; these are the shipped constants and no optimality is claimed for
either.

Finding: the half-precision refusal was justified by quoting specific sampled
bfloat16 error figures, which is the same sampling caveat this module applies to
everything else.
Fix: the justification is now directional -- a coarser dtype wants a larger
crossover, so falling back to the float64 constant moves the wrong way and a
silent fallback is worse than no support. The measurements pointed the same way
but are no longer quoted as bounds.

Not verified locally: no compute slot for this commit; CI executes it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FBFCNAvKSxsz5gb1TWcyrE
The delta review's closeout: three prose sites carrying one claim, and two
duplications with structural rather than cosmetic value.

Finding: rewriting the half-precision justification in _phi.py left the same
claim standing in two other files -- a test docstring and the design doc -- and
the test docstring was the worse of them, attributing a measurement to a test
that performs none. It only asserts that a TypeError is raised.
Fix: all three now give the directional reason. Also dropped the word "worst":
directional reasoning supports "the float64 constant moves the wrong way for a
coarser dtype, so a silent fallback is worse than refusing", but not "worst of
all available", which needs the ranked comparison this suite deliberately
removed. The refusal stands on the directional argument alone.

Finding: `_rel` was duplicated verbatim in two test modules and `_agrees` in a
third was the same function plus a finiteness guard -- three near-copies of one
comparison with the guard in only one of them. The NaN-comparison defect has now
recurred four times in this artifact, and a helper that must be re-remembered at
each site is the version of "get the polarity right" most likely to fail again.
Fix: `rel_error` and `agrees` live once in tests/transport/__init__.py, with the
invariant recorded there. The sharper form of it, from the review: in ASSERT
form the dangerous polarity does not exist, because assert fires on falsity and
every comparison with NaN is false, so `assert x < tol` and `assert x > tol` are
both NaN-safe. Only the PREDICATE form `if err > tol: reject` is unsafe, where
NaN means "do not reject". This artifact contains exactly two such constructs --
make_chart's Gram check and the gate functions -- and both are now guarded. That
is a structural argument rather than an enumeration, which is why it can be
relied on going forward.

Finding: five separate raises-tests against make_chart each rebuilt their own
fixtures, and their differing keyword sets made it non-obvious which check fires
for a given bad input -- for a contract that has now grown three times.
Fix: one parametrized table of eleven violations against a single valid
baseline, plus a control asserting that baseline actually constructs, so a
mistake in the fixture cannot make every row vacuously pass. The Gram-overflow
case stays standalone: it asserts properties of its own inputs (entries finite,
Gram NaN) before asserting the refusal, and collapsing it would lose that.

Not applied: the review also suggested collapsing a duplicated funnel
log-density. Checked -- it is defined once, in the refuters module only; the
conditioning module builds a funnel chart but no log-density. No change made.

Also: fixed an over-indented line in an attribute docstring that reST rendered
as a block quote.

Not verified locally: no compute slot for this commit; CI executes it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FBFCNAvKSxsz5gb1TWcyrE
Finding: test_finite_basis_with_overflowing_gram_is_refused failed slow CI at
28e143f. The failure is in the test's own setup, not in the product: the line
`gram = basis.T @ basis` overflows by design -- that overflow is how the NaN
Gram entry is constructed -- and pytest.ini turns warnings into errors, so the
RuntimeWarning aborted the test before make_chart was ever called. The assertion
under test never ran.
Fix: np.errstate(over="ignore", invalid="ignore") around that single NumPy
operation and nothing else. The finite-input premise, the NaN-Gram premise and
the ValueError assertion are all unchanged, so the test still fails if the
basis stops being finite, if the Gram stops being NaN, or if make_chart stops
refusing it.

Deliberately not done: no global warning policy change, no tolerance change, no
new test family or helper. The narrow suppression is correct here precisely
because the overflow is the premise; suppressing it anywhere wider would hide
overflows that are not.

CI state being corrected: 141 passed, 1 failed, 2 skipped in 611.24s at 28e143f.
That run does not establish product correctness for the 141 either -- it is the
same head, and the single failure was a harness collision rather than evidence
about the chart.

Not verified locally: no compute slot for this commit; CI executes it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FBFCNAvKSxsz5gb1TWcyrE
Finding: the previous repair moved the failure one line down rather than fixing
it. Suppressing the NumPy warning let the test reach its next premise, which
then failed: the NumPy Gram is [[inf,inf],[inf,inf]], not NaN. make_chart was
still never reached, so the assertion under test still never ran. Two things
were wrong with that premise. Overflow reduction behaviour is not portable, so
requiring NaN specifically was requiring an implementation detail. And NumPy is
not the path the factory takes -- asserting a property of a NumPy matmul says
nothing about what the JAX code under test computes.
Fix: the premise is now evaluated on the same dtype and backend the factory
uses, and requires only NON-FINITENESS, accepting inf or NaN. The np.errstate
suppression is gone with it, since nothing NumPy overflows any more. The
finite-input premise and the ValueError refusal are unchanged.

Renamed to test_finite_but_overflowing_basis_is_refused, because that is what it
covers. The docstring now states the limit rather than leaving it to be
inferred: this does NOT cover the NaN-only polarity path. On this backend the
derived Gram reduces to +inf, and `inf > tol` is already True, so a gate written
with the unsafe polarity would reject this input too. The polarity guard in
make_chart is unchanged and rests on the reasoning recorded beside it, not on
this test.

Preflight (one existing test, inside the lock, 300s bound): REAL_EXIT=0,
1 passed in 1.09s, 2s wall. Import resolution verified inside the lease --
tuningfork resolved to the worktree, not an installed sibling. Run against
afeefc4 plus this uncommitted change; the tree was dirty by design and the run
output records it.

CI state this corrects: 141 passed, 1 failed, 2 skipped in 621.71s at afeefc4.
The 141 remain evidence about the properties they assert; they do not certify
the product, and the single failure was a broken premise rather than a finding
about the chart.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FBFCNAvKSxsz5gb1TWcyrE
@junpenglao
junpenglao marked this pull request as ready for review September 7, 2026 18:53
@junpenglao
junpenglao merged commit e0030f1 into main Sep 7, 2026
3 checks passed
@junpenglao
junpenglao deleted the geodesic-frozen-transport branch September 7, 2026 19:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant