Repository navigation
Add supplied frozen coordinate chart with executed validation gates - #263
Merged
Merged
Conversation
Validation-first artifact for a frozen affine-flow coordinate chart: the supplied chart plus the algebra needed to check it. No fitter, no selection rule, no warmup descriptor, no execution recipe, no registry wiring. The chart family is constrained at construction to |h|=1, h.c=1, h.a=-alpha, which makes the clock rate identically one and therefore makes the inverse and the log-Jacobian exact in real arithmetic rather than fitted surrogates. Finding: a non-zero Stein residual is the USEFUL case, not a defect. grad(log pi).V + div V is exactly d/dt log pi_chart. If that equals -kappa*t + beta globally with kappa>0 and no section dependence, integrating gives a Gaussian clock independent of the section. An exactly zero residual would instead mean a flat, improper clock. Earlier reasoning here treated zero as the target and conflated the fitter's regression residual (which carries affine-in-clock slack) with the deployed field's residual. Fix: test_chart_refuters.py pins all three quantities separately on the exact funnel -- raw residual -t/9, slack-absorbed regression residual ~0, and the resulting structure via the decisive discriminator d_s d_t log pi_chart == 0 with constant clock curvature -1/9. Finding: differentiating a custom_jvp density to test its own supplied score is circular, and AD cannot manufacture a global inverse from forward alone. Fix: every score gate references the plain, non-hooked native_logdensity(forward(y)) + log_det(y); the inverse is tested as a supplied map. Finding: predicting which gate catches a mutant is not the same as executing it. Rescaling h was expected to fail the determinant gate and does not -- the Householder reflector is built from the normalised direction, so the section stays orthogonal to h and forward is unchanged to rounding. Fix: all seven mutants are executed through the real helper and the expected gate pattern is asserted exactly. M4 and M5 pass the log-determinant and round-trip gates and are caught only by the independent score gate, which is why that gate is mandatory. Finding: no single phi series crossover serves both supported dtypes, and two measurement traps hide this -- a float128 closed-form oracle cancels worse than the code under test at |x|~1e-11, and a numpy stand-in disagrees with the JAX path in float32. Fix: hybrid extended-precision oracle, comparison against the shipped code path, and a per-dtype threshold (float64 0.1 -> 3.8e-15; float32 0.3 -> 8.8e-07). Both value and derivative error are measured. Finding: the chart fails silently. The score is wrong by O(1) relative at alpha*clock ~ 20 while every component is still finite, and only goes non-finite around alpha*clock ~ 40, so a reject-on-NaN guard cannot catch it. Fix: test_chart_conditioning.py records the band with a float32-vs-float64 ladder on the shipped code path, and labels exp(alpha*clock) an empirical indicator explicitly, not a certified bound. Also fixed during development: push_score used the wrong preconditioner power (+0.5 instead of -0.5), which the diagonal case cannot detect and the low-rank algebra test caught. Verification: 40 tests pass (pytest tests/transport, CPU, ~15 s); pre-commit clean on the ten added files. No existing tuningfork or BlackJAX file changed. No sampling, no campaign, and no mixing, efficiency or posterior-accuracy claim. Refs: local Issue#1200, Experiment#925, Experiment#926, Belief#1371, Belief#1372. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FBFCNAvKSxsz5gb1TWcyrE
Finding: the Gaussian-clock statement was written as though finite evaluations established it. They do not. The implication needs three conditions stated together -- the identity d/dt log pi_chart = -kappa*t + beta holding on the FULL real line in t, kappa CONSTANT and positive, and a global Cartesian chart with global support -- and only under those does kappa <= 0 fail to normalise. The previous wording also read as a universal claim that a non-positive precision is always improper, which is false on a bounded or otherwise restricted clock domain. Fix: state the analytic implication with its premise explicit, note that for the supplied chart tested here the premise is discharged by the algebra while the evaluations only corroborate it, and record that for a LEARNED chart the premise is open -- a training-support fit says nothing about the full line, nothing about whether kappa is constant off that support, and does not guarantee kappa > 0. Added "constant" to the list of properties a small training residual fails to establish; it was missing alongside positivity. No behaviour change: documentation and test docstrings only. 40 tests still pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FBFCNAvKSxsz5gb1TWcyrE
Finding: the shipped docstrings and design doc carried internal review vocabulary -- numbered tiers, and labels naming the two earlier exploratory implementations as investigation arms. These are meaningful only inside the private review process and are noise, or worse, misleading provenance, to a reader of the repository. Fix: name each suite by what it establishes rather than by a tier number; refer to "earlier exploratory implementations" instead of arms; rename the comparison constant to PRIOR_CROSSOVERS. No test logic, tolerance, constant or assertion changed -- the phi selection still compares the shipped path against both prior crossovers and still asserts the chosen per-dtype threshold wins. 40 tests still pass; pre-commit clean. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FBFCNAvKSxsz5gb1TWcyrE
…uite Independent paired review found the chart mathematics clean and the validation around it defective. This repairs the validation and the two real code defects, and removes machinery that was re-litigating past decisions instead of protecting anything. Finding: the gates compared `float(error) > ATOL`, which is False for NaN, so a provider returning NaN everywhere passed all four gates. The suite's central claim -- that these gates catch wrong charts -- had a hole big enough to admit a chart that computes nothing. Fix: finiteness is part of gate success. Four providers replace the seven-mutant table: a valid chart, a wrong log-determinant, a wrong score, and non-finite controls for NaN/+Inf/-Inf. The pass/fail vector over a mutant family was never a product contract -- which gate a mutation perturbs is a property of how it was injected -- so it is gone rather than rebuilt. Finding: all five test modules mutated jax_enable_x64 at MODULE level, i.e. at collection, before the shared conftest guard snapshots it. The guard then restored the already-corrupted value forever after, leaking float64 into the whole session and failing two groundtruth tests in files this branch never touched. Fix: a module-scoped fixture that enables x64 at setup and restores the prior value in a finally. No collection-time mutation and no conftest change. The modules trace JAX and run in seconds, so they are marked slow, not fast. Finding: `log_det` used `log(scale)` while claiming log|det|, returning NaN for a chart whose forward, inverse and score were all exact. Fix: log-absolute determinant, so signed scale is supported consistently rather than left implicit. Finding: "correct by construction" covered only what make_chart projects. Zero h, mismatched low-rank arguments, non-positive eigenvalues and non-orthonormal factors all still produced silently wrong charts -- a non-orthonormal active basis broke the inverse by 2.3 with no error raised. Fix: a stated, enforced input contract that refuses what it cannot repair. Orthonormality is required only of SPECTRALLY ACTIVE columns: neutral columns contribute exactly zero at every power, so demanding it of them would reject legitimate inputs carrying a neutral direction. Finding: the phi "measured floors" were maxima over a sampled grid, and review found points exceeding two of them. A grid maximum is a lower bound on the worst case, so promoting one to a "floor" was wrong in kind. Worse, the f32 crossover was selected because the table recorded 0.3 and 1.0 as tied -- and that tie was itself a sampling artifact. Fix: no floor, bound or accuracy is claimed for phi at any dtype. The competition tests and the shipped-implementation copy they compared against are removed; a regression suite should not re-run a past threshold decision on every invocation. What remains checks values AND derivatives against an independent 60-digit decimal oracle at points chosen to fail differently: zero, small arguments, both sides of each crossover, and a large-argument cancellation. Finding: the series branch received arbitrarily large x under `where`, so its Horner recurrence overflowed and an inf-inf tangent in the unselected branch poisoned the selected one -- jax.grad returned NaN where the value was exact. Fix: clamp the series argument as the direct denominator already was, and assert gradient finiteness where the value is finite. Unsupported half precisions now raise instead of silently taking the float64 threshold, which measurement showed is the worst available choice for them. Finding: every gate tolerance was absolute and calibrated at |y| ~ 1, so the UNMUTATED chart failed its own log-det gate at |y| ~ 100 (2.669e-11 against 1e-11) at a relative error of 1.08e-13. The suite would reject a correct implementation -- the failure direction that costs most later. Fix: relative tolerances throughout, and a far-field case that an absolute bound would fail. Also: the "identity" refuter's forward was the identity map and the "rotated" refuter's low-rank factor was exactly I, so neither exercised what it named; both are replaced with cases that do. The conditioning module no longer asserts that the implementation degrades -- pinning today's rounding behaviour would block tomorrow's repair -- and instead records the clock residual |h.z - t| as an O(d) runtime diagnostic. The claim that a guard could not be inferred from returned values is withdrawn; that residual is exactly such a guard. The clock projection that repairs it is documented, NOT shipped: it changes the implemented map and needs its own verification. CONVENTION_VERSION stays v1 -- these repairs change results only where the old code returned NaN or served an unsupported dtype. NOT VERIFIED LOCALLY: the machine-wide numerical lock was held throughout (flock exit 75, twice, no retry loop). This batch is lint-clean but unexecuted; remote CI is the first execution, and its test-slow job runs the new tests followed by the two previously failing groundtruth tests in one process, which is the isolation check the x64 fix needs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FBFCNAvKSxsz5gb1TWcyrE
Finding: the independent phi oracle built its series with `xd ** k`, and `Decimal(0) ** 0` raises InvalidOperation. Since zero is one of the named test points -- and the most worth testing, being where the direct form cancels hardest -- the oracle would have raised rather than compared, erroring the test instead of failing it. The local numerical lock was held, so this was found by reading the new code rather than by running it. Fix: accumulate powers iteratively; no exponentiation anywhere in the oracle. Validated in pure Python, no JAX: phi1(0)=1, phi2(0)=0.5, phi1'(0)=0.5, phi2'(0)=1/6 exactly, and agreement with float references at 0.3, +/-12. At x=1e-8 the oracle gives phi2 = 0.5000000016667 against the float form's 0.500000008431 -- the oracle is correct and the float form is the one cancelling, which is the reason it exists. Still unexecuted under JAX: the numerical lock remains held by another station. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FBFCNAvKSxsz5gb1TWcyrE
Finding: the design doc and the conditioning module presented |h.z - t| as if a clean reading meant a healthy chart. It does not. The residual observes error in the clock component and is blind to anything that leaves h.z intact, so it certifies neither the score nor the log-Jacobian. Fix: both files now say so. The residual detects one specific failure; accuracy of the score and Jacobian is covered by the algebra tests against an independent reference, not by this diagnostic. Also retracted, source-refuted: the previous commit message called CI's test-slow the ordered isolation check for the x64 fix. .github/workflows/test-slow.yml runs `pytest tests -m slow -n auto`, so there is no guaranteed process or ordering. CI covers collection-time behaviour and distributed regression; the ordered check is a separate scoped single-worker run, still queued behind the shared numerical lock. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FBFCNAvKSxsz5gb1TWcyrE
A source-only delta review of the repair itself found eight items. The first would have failed the suite deterministically, and my ordered isolation check did not cover that file -- so the batch I reported as lint-clean was not green. Finding: test_gradients_stay_finite_where_the_value_is_finite asserted the VALUE is finite at large POSITIVE x, where it provably is not. phi1(x) = (e^x-1)/x is unbounded, so expm1 overflows above 88.7 (float32) / 709.8 (float64) and the assertion fails at the first positive entry in both dtypes. The test's own name stated the correct scope; its body asserted more. Fix: probe only large negative arguments, where the value is small and exact and where the repaired gradient bug actually lived, and pin the positive overflow separately as a property of the function so it is not later mistaken for a bug. Finding: the orthonormality gate used an absolute 1e-8 Gram tolerance evaluated in the input dtype. A genuinely orthonormal float32 basis carries Gram error around sqrt(d)*eps32 ~ 3e-7, so the contract rejected VALID float32 input. No test built a float32 low-rank chart, which is why it went unseen. Fix: scale the tolerance with dtype eps and rank; real violations are O(0.1), far above it. Added the missing float32 low-rank case. Finding: two gaps in the enumerated input contract -- center size was never checked against h, and finiteness was checked for h and scale but not for a, c, alpha or center, so a NaN in c propagated through the projection and silently built a NaN chart. Fix: both checked and refused. Finding: test_projection_is_a_no_op_on_a_well_conditioned_point could not fail. The displacement is h(t - h.z), whose largest component is the residual times max|h_i| <= residual because ||h|| = 1, so the assertion restates an identity of the expression and holds for any chart including a broken one -- the same defect class this branch removed from the mutant tests two commits ago. Fix: deleted, with a comment recording why. The projection is documented, not shipped; testing it here tested nothing. Finding: test_omitted_normalisation_is_rejected described itself as bypassing make_chart's projection but actually rescaled h on a finished chart, leaving u, a and c consistent with the normalised direction -- precisely the criticism this branch made of the previous M6 mutant, reintroduced while fixing it. Fix: the chart is now built the way forgetting the normalisation would build it, with the projection and the Householder reflector both derived from the raw h. Finding: the gradient-poisoning mechanism was named wrong in two places. Horner's s*x + c produces inf, never inf - inf; the poisoning is 0 * inf in the reverse pass, where the unselected branch's zero cotangent multiplies a saved infinite intermediate. That sentence is the one explaining why the clamp works, so being wrong in it matters more than its length suggests. Fix: corrected in the module and the test. Also: make_chart validates with Python control flow and is host-side only, not jit/vmap-traceable -- now stated, since the previous implementation was traceable and Chart is a pytree that does cross those boundaries. Removed a docstring pointer to a table this branch had already deleted. Still unexecuted at the time of writing: the transport suite has never run under JAX. The ordered isolation check at 8e0adc7 covered five tests and verified that sequence, not this. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FBFCNAvKSxsz5gb1TWcyrE
Finding: the previous commit claimed to scale the Gram tolerance by dtype and rank. It did not. The edit was applied with a string replacement whose target had already been reformatted onto multiple lines by black, so the replacement matched nothing and wrote a silent no-op while the commit message described the fix as done. The hardcoded 1e-8 was still in the file. Fix: applied for real, and asserted the match before writing plus the result after writing, so a non-matching patch fails loudly instead of passing quietly. This was caught by the first execution of the transport suite: a valid float32 low-rank chart was rejected with max |U^T U - I| = 1.192e-07, which is the sqrt(d)*eps32 Gram error an orthonormal float32 basis legitimately carries. The test that caught it was added in the same commit as the fix that failed to apply, from the delta review -- so the review's finding was confirmed by its own regression test catching the repair's failure to take effect. Suite status at the previous commit: 1 failed, 38 passed. The single failure is this one. Process note recorded because it nearly shipped: a wrapper script around pytest reported exit 0 while pytest itself exited 1, because the wrapper's last statement was an echo. The real status was only visible because the script printed it immediately after the pytest line. An unpiped wrapper exit code is not automatically the command's exit code. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FBFCNAvKSxsz5gb1TWcyrE
Finding: make_chart still had a NaN-blind path, in the code written to enforce the contract. lr_basis had no finiteness check, so a NaN in an active column makes the Gram residual NaN, and `if off > tol` is False for NaN -- the gate accepts a basis it cannot evaluate. This is the same comparison defect this branch repaired in its own test gates two commits ago, reproduced in the fix for it. A positive infinity in lr_eigenvalues likewise satisfies a bare `> 0` test. Fix: require finite lr_basis and lr_eigenvalues BEFORE any Gram or spectral work, for every column including neutral ones -- 0 * NaN is NaN, so a neutral column is not numerically inert. One NaN-basis refusal and one infinite- eigenvalue refusal cover it; the valid float32 and finite-neutral controls are kept. Finding: a, c and alpha shapes were never checked, so broadcasting could construct a chart outside the declared family -- a scalar `a`, or an `alpha` carrying a trailing axis, both produce well-formed arrays. Fix: shapes checked explicitly against h, and alpha required to be a scalar. Finding: the tolerance comment claimed "real violations are O(0.1), orders above this". That is not something the code knows. A violation can be arbitrarily small, and one below the rounding allowance is simply not detected. Fix: the comment now says what the tolerance is -- a rounding allowance, not a separation guarantee. Finding: the phi tests described expm1's overflow at ~88.7 / ~709.8 as the point where "the function itself is unbounded", and treated it as a boundary the test proved. Both are wrong. phi1(x) = (e^x - 1)/x is finite at every finite real x, and because the division follows the exponential there are arguments where the exact result is representable while the intermediate expm1 has already overflowed. The limit observed belongs to the evaluation order, not to the function. Fix: renamed and rewritten to record where THIS IMPLEMENTATION stops returning finite values, stating explicitly that it is not the representability boundary and that no boundary is located or claimed. The 2*limit arguments stay; they are comfortably past the limit rather than evidence of where it is. Finding: the oracle set decimal precision with getcontext().prec = 60 at module scope, mutating global decimal state during collection and leaking it into every later test in the session. That is the same defect class as the x64 collection leak this branch already repaired, in a different global, introduced while repairing it. Fix: local context inside the oracle. Process note: the previous commit's fix silently did not apply because a string replacement matched nothing after black had reformatted its target. Every edit in this commit asserts its match before writing and verifies the result after. Not verified locally: no compute slot was granted for this commit; final-head CI executes it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FBFCNAvKSxsz5gb1TWcyrE
Finding: the orthonormality gate was still NaN-blind, and this is the fourth instance of that defect on this branch. The previous three were closed by checking finiteness before the comparison. That discipline does not reach here, because `off` is not an input: it is DERIVED from the inputs after they have all been validated. An off-diagonal Gram entry is a signed sum, so a basis with every entry finite -- around sqrt(dtype max), 1.8e19 in float32 or 1.3e154 in float64, with at least two active columns of mixed sign -- yields (+inf) + (-inf) = NaN, which jnp.max propagates. Written as `off > tol` the gate then accepts a basis it could not evaluate, since NaN > tol is False. Diagonal entries are sums of squares and can only overflow to +inf, which raises correctly; the reachable path is specifically the off-diagonal. Fix: `if not (off <= tol):`. The structural lesson is recorded in the code because it generalises past this case -- finiteness-before-comparison protects inputs and cannot protect a quantity computed after the inputs are cleared, so the only durable guard on a derived value is the comparison's polarity. The regression test asserts the basis entries are finite and the Gram is NaN BEFORE asserting the refusal, so it cannot silently stop exercising the path if the arithmetic underneath it changes. Finding: the SERIES_THRESHOLD docstring still asserted an optimality verdict -- "the optimum is dtype-dependent", "a single constant is measurably wrong for one of them" -- which is the deleted "measured floor" claim in another costume. It had no test behind it, since the threshold comparison was deliberately removed for resting on a sampled grid, and it sat two screens below a module docstring disclaiming exactly that kind of statement. An outstanding measurement also contradicts it. Fix: keep the mechanism, drop the verdict. The two branches' errors move in opposite directions in |x|, so a useful crossover depends on the dtype's epsilon; these are the shipped constants and no optimality is claimed for either. Finding: the half-precision refusal was justified by quoting specific sampled bfloat16 error figures, which is the same sampling caveat this module applies to everything else. Fix: the justification is now directional -- a coarser dtype wants a larger crossover, so falling back to the float64 constant moves the wrong way and a silent fallback is worse than no support. The measurements pointed the same way but are no longer quoted as bounds. Not verified locally: no compute slot for this commit; CI executes it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FBFCNAvKSxsz5gb1TWcyrE
The delta review's closeout: three prose sites carrying one claim, and two duplications with structural rather than cosmetic value. Finding: rewriting the half-precision justification in _phi.py left the same claim standing in two other files -- a test docstring and the design doc -- and the test docstring was the worse of them, attributing a measurement to a test that performs none. It only asserts that a TypeError is raised. Fix: all three now give the directional reason. Also dropped the word "worst": directional reasoning supports "the float64 constant moves the wrong way for a coarser dtype, so a silent fallback is worse than refusing", but not "worst of all available", which needs the ranked comparison this suite deliberately removed. The refusal stands on the directional argument alone. Finding: `_rel` was duplicated verbatim in two test modules and `_agrees` in a third was the same function plus a finiteness guard -- three near-copies of one comparison with the guard in only one of them. The NaN-comparison defect has now recurred four times in this artifact, and a helper that must be re-remembered at each site is the version of "get the polarity right" most likely to fail again. Fix: `rel_error` and `agrees` live once in tests/transport/__init__.py, with the invariant recorded there. The sharper form of it, from the review: in ASSERT form the dangerous polarity does not exist, because assert fires on falsity and every comparison with NaN is false, so `assert x < tol` and `assert x > tol` are both NaN-safe. Only the PREDICATE form `if err > tol: reject` is unsafe, where NaN means "do not reject". This artifact contains exactly two such constructs -- make_chart's Gram check and the gate functions -- and both are now guarded. That is a structural argument rather than an enumeration, which is why it can be relied on going forward. Finding: five separate raises-tests against make_chart each rebuilt their own fixtures, and their differing keyword sets made it non-obvious which check fires for a given bad input -- for a contract that has now grown three times. Fix: one parametrized table of eleven violations against a single valid baseline, plus a control asserting that baseline actually constructs, so a mistake in the fixture cannot make every row vacuously pass. The Gram-overflow case stays standalone: it asserts properties of its own inputs (entries finite, Gram NaN) before asserting the refusal, and collapsing it would lose that. Not applied: the review also suggested collapsing a duplicated funnel log-density. Checked -- it is defined once, in the refuters module only; the conditioning module builds a funnel chart but no log-density. No change made. Also: fixed an over-indented line in an attribute docstring that reST rendered as a block quote. Not verified locally: no compute slot for this commit; CI executes it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FBFCNAvKSxsz5gb1TWcyrE
Finding: test_finite_basis_with_overflowing_gram_is_refused failed slow CI at 28e143f. The failure is in the test's own setup, not in the product: the line `gram = basis.T @ basis` overflows by design -- that overflow is how the NaN Gram entry is constructed -- and pytest.ini turns warnings into errors, so the RuntimeWarning aborted the test before make_chart was ever called. The assertion under test never ran. Fix: np.errstate(over="ignore", invalid="ignore") around that single NumPy operation and nothing else. The finite-input premise, the NaN-Gram premise and the ValueError assertion are all unchanged, so the test still fails if the basis stops being finite, if the Gram stops being NaN, or if make_chart stops refusing it. Deliberately not done: no global warning policy change, no tolerance change, no new test family or helper. The narrow suppression is correct here precisely because the overflow is the premise; suppressing it anywhere wider would hide overflows that are not. CI state being corrected: 141 passed, 1 failed, 2 skipped in 611.24s at 28e143f. That run does not establish product correctness for the 141 either -- it is the same head, and the single failure was a harness collision rather than evidence about the chart. Not verified locally: no compute slot for this commit; CI executes it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FBFCNAvKSxsz5gb1TWcyrE
Finding: the previous repair moved the failure one line down rather than fixing it. Suppressing the NumPy warning let the test reach its next premise, which then failed: the NumPy Gram is [[inf,inf],[inf,inf]], not NaN. make_chart was still never reached, so the assertion under test still never ran. Two things were wrong with that premise. Overflow reduction behaviour is not portable, so requiring NaN specifically was requiring an implementation detail. And NumPy is not the path the factory takes -- asserting a property of a NumPy matmul says nothing about what the JAX code under test computes. Fix: the premise is now evaluated on the same dtype and backend the factory uses, and requires only NON-FINITENESS, accepting inf or NaN. The np.errstate suppression is gone with it, since nothing NumPy overflows any more. The finite-input premise and the ValueError refusal are unchanged. Renamed to test_finite_but_overflowing_basis_is_refused, because that is what it covers. The docstring now states the limit rather than leaving it to be inferred: this does NOT cover the NaN-only polarity path. On this backend the derived Gram reduces to +inf, and `inf > tol` is already True, so a gate written with the unsafe polarity would reject this input too. The polarity guard in make_chart is unchanged and rests on the reasoning recorded beside it, not on this test. Preflight (one existing test, inside the lock, 300s bound): REAL_EXIT=0, 1 passed in 1.09s, 2s wall. Import resolution verified inside the lease -- tuningfork resolved to the worktree, not an installed sibling. Run against afeefc4 plus this uncommitted change; the tree was dirty by design and the run output records it. CI state this corrects: 141 passed, 1 failed, 2 skipped in 621.71s at afeefc4. The 141 remain evidence about the properties they assert; they do not certify the product, and the single failure was a broken premise rather than a finding about the chart. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FBFCNAvKSxsz5gb1TWcyrE
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds a supplied (frozen) affine-flow coordinate chart to tuningfork, together with the
algebra needed to validate it. The chart is handed complete parameters and evaluates them; it
learns nothing.
Not in this PR: no fitter, no chart-selection rule, no warmup descriptor, no execution
recipe, no registry entry, no sampler integration, and no changes to BlackJAX or to any existing
tuningfork file. Every file here is new. No sampling, mixing, efficiency or posterior-accuracy
claim is made anywhere, and no chain is run.
The chart
In preconditioned coordinates
z = L^{-1}(q - center), a constant affine fieldis constrained at construction to
||h|| = 1,h . c = 1,h . a = -alpha. Those threeconstraints make the clock rate identically one,
d/ds (h . z) = 1, which yields threeconsequences exact in real arithmetic:
t = h . z;tlands on the sectionh . z = 0, with noiteration and no root find;
log|det J| = alpha (d-1) t + log|det L|is the exact log-Jacobian.These are identities of the map, not statements about its floating-point evaluation. The
implementation's finite-precision behaviour is a separate matter and is treated as such below.
CONVENTION_VERSIONidentifies the coordinate and series conventions the module fixes: clock-lastordering, the reflector orientation, and the per-dtype series crossover.
Delivered surface
The chart implementation is private; only the convention identifier is exported. The private
constructor does two distinct things and does not conflate them:
his normalised, andcandaare projected so thath . c = 1andh . a = -alphahold by construction.h;matching, non-zero
scale; declared shapes fora,c,centerand a scalaralpha; finitea,c,alpha,center; finite low-rank basis and eigenvalues, checked before any spectralwork; strictly positive eigenvalues; and orthonormality of the spectrally active basis
columns.
Orthonormality is scoped to active columns deliberately. Columns with unit eigenvalue contribute
exactly zero to the preconditioner at every power and zero to the log-determinant, so requiring
orthonormality of them would refuse legitimate inputs carrying a neutral direction.
log_detis a log-absolute determinant, so a negativescaleentry is supported: it flipsorientation without changing the volume element.
phiserves float32 and float64 only. Half precisions raise rather than falling back. Thereason is directional: a coarser dtype wants a larger crossover, so falling back to the float64
constant moves the wrong way, and a silent fallback would be worse than no support.
What is checked
slogdetof the AD Jacobian, near and far from the origin; the two-sided inverse; the suppliedscore against a reference; and invertibility of the score pullback. Diagonal and low-rank
preconditioners, including an active factor and a neutral column.
chart built without normalising
h. Gate success requires finiteness as well as agreement:a reject-if-error-exceeds comparison does not reject NaN, because NaN > tolerance is false.
Explicit finiteness checks handle non-finite values.
a table of representation violations against a baseline that is separately asserted to
construct, so a mistake in the fixture cannot make every row vacuously pass.
phi— values and derivatives against an independent 60-digit decimal oracle at namedpoints: zero, small arguments, either side of each crossover, and a large-argument case where
the direct form cancels. These are agreements at those points. No accuracy bound, floor or
optimality is claimed for the crossovers or for either branch, and no global property is
inferred from the sampled points.
clean where the chart is well conditioned.
One methodological point the tests encode:
custom_jvpdensity to test the scoreit supplies returns the rule under test, so the reference is the plain composed expression. For
the same reason the inverse is tested as a supplied map: AD cannot manufacture a global inverse
from the forward map alone.
Scope of the claims
The funnel case checks that, for that target's own scaling map, the transformed clock score is
affine in the clock with no section dependence, so the density factorises into a Gaussian clock
and a section marginal. The premise for that conclusion is stated in full in the design note: it
requires the identity to hold on the full real line, with a constant positive precision,
on a global Cartesian chart with global support. For a supplied chart those premises are
discharged by the algebra and the evaluations corroborate them; for a fitted chart they would
remain open. A residual measured on any finite set of points establishes none of them.
The clock residual is a component-specific detector, not a certificate. It observes error in
the clock coordinate and is blind to anything leaving that coordinate intact, so it certifies
neither the score nor the Jacobian; those are covered by the algebraic checks. A projection that
removes the residual is described in the design note and is not shipped — it would change the
implemented map and needs its own verification.
The overflowing-basis case establishes that a basis with finite entries and a non-finite derived
Gram is refused. It does not establish coverage of the non-finite-comparison path specifically:
on this backend that Gram reduces to an infinity, which a comparison-only gate would also reject.
The guard rests on the reasoning recorded beside it, not on that test.
Verification
Final-head CI at
fa1f3fb62711726f8273a25b16cb414c41c5ba71:The modules added here trace JAX and run in seconds, so they are marked slow and run in the
slow job. Earlier heads of this branch carried different counts and, in two cases, failures; those
results belong to those commits and are not restated here. A local ordered single-worker check of
the added modules together with two previously affected cases was run at an earlier commit, and a
single-test check was run against a working tree that was not this commit; neither is offered as
evidence about this head.