Status: bounded local implementation, real cohorts, memory workloads and
repository gates verified. Requirements REQ-CTX-009 and REQ-CTX-010;
decisions ADR-025 and amended ADR-017. Broad tasks stay open.
A supported source is converted to .synapse before execution. The real
Qwen2.5-0.5B Q8_0 lossless artifact has file SHA-256
b67620a926cd299931b0a7cbfbd31f112c893832658cefa64e67f4d4a21f514d and
source SHA-256
ca59ca7f13d0e15a8cfa77bd17e65d24f6844b554a7b6c12e07a5f89ff76844e.
Every measured profile binds package, graph, runtime and exact prompt IDs.
The natural haystack concatenates 49 Markdown documents from docs,
README.md and synapse.plan.md at 2bcd73c: 289,473 UTF-8 bytes,
SHA-256 b71f38b349b40b544deab91359f9f18c8a3ce160ade74d607d7cc83a61d0f175.
The tree was exported with git archive; a C# tool normalizes and concatenates
sorted paths. Ignored artifacts/optimization-ablation holds plans, snapshots,
binary hashes, raw JSON, corpus and logs. Model/generated bytes are not in Git.
The FP32/FP16 Metal plan requests 4,000/8,000/16,000 prompt tokens at three depths for needle and multi-key tasks, plus one variable-tracking case per length: 21 cases, two paired rounds, 84 fresh children. Factory paragraph cuts make actual prompt counts differ from targets; both are recorded. Repeated requests with/without exact prefix reuse form a separate cohort. Profiles rotate order. The same-prompt tail NLL and greedy agreement diagnose numerical drift; exact answers diagnose these bounded tasks. Neither proves general held-out language quality.
Hardware: M2 Pro, 12 CPU cores (8 performance/4 efficiency), 19 GPU cores,
32 GB, macOS 27.0.1 ARM64; .NET SDK 10.0.401/runtime 10.0.12. Real Metal runs
use narrow sandbox escalation because the sandbox hides the device. The host
also runs unrelated development work; timing is local diagnostic evidence.
Independent builds avoid the shared compiler queue using
-p:UseSharedCompilation=false -m:1 /nr:false.
The real 3,754-token needle at depth 0.5 produced these cold results from five fresh workers (one pair per candidate, not statistical qualification):
| Profile | TTFT ms | Generation ms | Output tokens | Correct | Allocated KV MiB |
|---|---|---|---|---|---|
| dense FP32 | 1518.920 | 1568.813 | 8 | yes | 96 |
| FP16 KV | 1391.515 | 1439.357 | 8 | yes | 48 |
| mean2 | 1635.946 | 1953.731 | 16 | no | 96 |
| mean4 | 1501.895 | 1631.828 | 16 | no | 96 |
| mean8 | 1556.896 | 1719.937 | 16 | no | 96 |
FP16 tail mean absolute NLL drift was 0.002825 with greedy agreement 1. Mean2/4/8 greedy agreement was zero and tail perplexity ratios were 3346.35/8257.71/27187.02. Averaging retained storage/arithmetic and lost answers. The owner rejected it: all averaging code/options and generated mean models are removed. Its earlier red/green tests and focused coverage describe a removed experiment, not current product verification.
Raw: pilot-initial.json; derived: pilot-initial-summary.md; progress:
pilot-initial.log under the ignored artifact folder. A single-pair 1.092×
FP16 TTFT ratio is not a qualified speedup. The larger cohort below decides
the observed speed/quality effect.
CLI mode tests first failed 21/21 because old commands ignored the mode. Runner shell plan/worker cases first failed against real processes. Subsequent source-generated JSON tests exposed omitted init-only defaults becoming zero; internal mutable plan/profile/KV DTOs fix that without changing public runtime options. Real child tests check provenance forgeries, pairing, repeated controls, failure retention, cancellation/PID exit and output sentinel preservation. Final controls/capacity suites contain 33 cases. The runner has 34 cases after the memory workload, backend guard and finite-log score regressions were added.
Growing-only KV and retained disposed CPU arrays are the memory defects under test. ADR-017 adds fourfold contraction hysteresis, preserves surviving bytes, reserves newly assigned batch slots and releases arrays under the execution gate. Real same-backend memory/lifecycle tests pass 49/49, including Metal FP32/FP16, ordered bitwise cold logits, retained-prefix outputs, independent slots and zero owned KV after disposal. Owner guards reject nested direct math and callback disposal before changing state; asynchronous admission is fenced against closing. These defects first failed real behavior tests. The final memory-focused isolated coverage union is 574/644 executable lines (89.13%) across seven runtime files, not whole-tree or changed-line qualification.
The paired runner now rejects another backend before measurement. Real prepared fixtures with large finite output norms previously made worker and memory JSON publication exit 134: finite NLL overflowed its exponential, or the actual score was nonfinite. Seven regressions now pass: finite log evidence survives exponent overflow with nullable perplexity/status, while nonfinite scores fail cleanly and retain preceding generation/memory phases. Ratios use log differences. The helper's isolated file coverage is 14/16 lines (87.5%).
no-mean-snapshot/Synapse.ReferenceBenchmarks.dll optimization-eval --plan plans/weights.json --output fp16-long-context.json (paths under the artifact
folder) exited 0 with completed status: 84 fresh Metal workers, 42 pairs,
21 unique cases repeated twice. All generated continuations are identical;
there are zero candidate answer regressions or improvements.
| Context | FP32 / FP16 TTFT ms (medians) | Median paired TTFT ratio | FP32 / FP16 generation ms | FP32 / FP16 allocated KV MiB | Correct answers each |
|---|---|---|---|---|---|
| 4K | 1578.471 / 1523.101 | 1.026 | 1678.103 / 1601.854 | 96 / 48 | 8/14 |
| 8K | 3679.897 / 3415.704 | 1.067 | 3757.568 / 3488.857 | 192 / 96 | 10/14 |
| 16K | 9801.376 / 8585.593 | 1.137 | 9889.396 / 8725.290 | 384 / 192 | 8/14 |
Median mean absolute tail NLL drift is 0.001998 / 0.003250 / 0.002735. Minimum tail greedy agreement is 1 / 0.96875 / 1; maximum mean absolute NLL drift is 0.002510 / 0.003856 / 0.006019. Numerical drift exists even where generated answers are identical. The original 0.5B model itself solves only 26/42 measured requests; this result preserves that bounded behavior.
Timings are descriptive local diagnostics, not a release/statistical speed
qualification. Unrelated Prostir work and own independent builds/coverage
overlapped parts of the frozen-binary cohort. Candidate order rotates, and raw
paired rows retain every output length and failure. The immutable managed
DLL hashes before/after match literally (cmp, exit 0); no benchmark snapshot
was coverage-instrumented. This snapshot precedes the allocator/lifecycle
repair, which changes storage ownership rather than the numerical kernels.
The post-repair real memory workload and same-backend bitwise tests validate
that separate change.
Raw: fp16-long-context.json; progress: fp16-long-context.progress.log;
derived summary: fp16-long-context-summary.md; binary hashes:
no-mean-binary-hashes-{before,after}.txt. Generation cap is the lower of the
plan and task factory cap (needle/multikey 16; variable tracking 40), recorded
in each row. Requested 64 is not silently reported as 64 actual outputs.
The separate plans/reuse.json cohort is completed: 24 fresh children, one
warmup and three measured paired rounds per context, 18 measured samples.
Both profiles use original weights, Metal and FP32 KV; only exact prompt-prefix
reuse differs. Each child records cold and repeated generation separately.
| Context | Dense / reuse cold TTFT ms | Dense / reuse repeated TTFT ms | Dense / reuse repeated generation ms |
|---|---|---|---|
| 4K | 1345.045 / 1345.284 | 1295.466 / 6.025 | 1336.635 / 49.044 |
| 8K | 3392.764 / 3381.221 | 3316.884 / 7.159 | 3364.065 / 53.693 |
| 16K | 9519.432 / 9664.538 | 9224.563 / 15.981 | 9288.005 / 84.825 |
All nine measured pairs have identical generated token sequences, exact
tail NLL, greedy agreement 1 and correct answers in both turns. This cohort
contains only three unique needle cases at depth 0.5; it establishes exact
prefix reuse on these cases, not general task quality. It does not accelerate
a new prompt. KV capacity stays 96/192/384 MiB: reuse saves repeated prefill
work, not the stored attention state. Both managed and native hashes before
and after this cohort match (cmp, exit 0).
Raw: reuse-long-context.json; derived: reuse-long-context-summary.md;
progress: reuse-long-context.progress.log; hashes:
no-mean-binary-hashes-{before,after-reuse}.txt and
no-mean-native-hashes-{before-reuse,after-reuse}.txt. The completed raw report
and every child's exit are retained; the report utility exits 0. The original
parent shell's completion exit was not retained by its terminal session, so
it is not separately claimed here.
current-final-snapshot comes from the post-repair whole Release build and
is never coverage-instrumented. Each memory-eval loads the separately
prepared original package, runs short→long→short, disposes it, then runs
independent fresh-short/fresh-long numerical oracles on the same backend.
Generation and subsequent tail-score checkpoints are separate. There is no
forced GC, RSS promise, cache flush or timing comparison across these phases.
| Backend / KV | Prompt tokens short / long | Short-first KV MiB | Long KV MiB | Short-after-long KV MiB | Disposed KV MiB |
|---|---|---|---|---|---|
| Metal / FP32 | 456 / 15921 | 24 | 384 | 24 | 0 |
| Metal / FP16 | 456 / 15921 | 12 | 192 | 12 | 0 |
| Native / FP32 | 456 / 3754 | 24 | 96 | 24 | 0 |
All three completed commands exited 0. All nine retained-vs-fresh comparisons
independently have identical ordered generated IDs, tail NLL double bits and
greedy IDs; absolute NLL difference is 0. This is exact same-backend evidence
for contraction; FP16-vs-FP32 drift is separately measured above. Physical
footprint, working set and managed heap remain separate metrics in each raw
phase; mapped model bytes and other allocations do not disappear with KV.
The complete derived memory-real-summary.md preserves all 16 phases per
profile and raw SHA-256 identities. Long→short owned KV contracts by 93.75%
for the two Metal runs and 75% for the shorter native run. Native context
differs, so its timings do not establish a backend speed comparison.
Peak observed footprint is 529.142 / 330.767 / 191.861 MiB for Metal FP32 / Metal FP16 / native FP32; peak working set is 587.953 / 599.125 / 684.625 MiB. These scope differences are real. For FP16, the immediate short-after-long footprint is still 325.329 MiB despite 12 MiB of owned KV; the separately recorded short-score phase observes 133.361 MiB. No immediate RSS/footprint release is promised. Native managed heap at disposal still observes 130.111 MiB without GC while owned KV is zero.
Exact commands, from the repository root:
dotnet artifacts/optimization-ablation/current-final-snapshot/Synapse.ReferenceBenchmarks.dll memory-eval --request artifacts/optimization-ablation/plans/memory/memory-metal-f16-natural.json --output artifacts/optimization-ablation/memory-metal-f16-natural.json
dotnet artifacts/optimization-ablation/current-final-snapshot/Synapse.ReferenceBenchmarks.dll memory-eval --request artifacts/optimization-ablation/plans/memory/memory-metal-f32-natural.json --output artifacts/optimization-ablation/memory-metal-f32-natural.json
dotnet artifacts/optimization-ablation/current-final-snapshot/Synapse.ReferenceBenchmarks.dll memory-eval --request artifacts/optimization-ablation/plans/memory/memory-native-f32-natural.json --output artifacts/optimization-ablation/memory-native-f32-natural.json
Raw/progress and exact prompt manifest live in the artifact folder. The fourfold contraction threshold avoids repeated small reallocations. CPU release removes owned arrays; managed GC can retain former memory temporarily. CUDA contraction is unimplemented/unqualified, and full-context RoPE tables remain a separate profiling opportunity. Neither is claimed solved.
- Locked restore: exit 0 (
restore.log). - Post-removal Release build/analyzers with independent compiler: exit 0
(
build-no-averaging.log); earlier post-JSON-fix builds had zero warnings. - Final Rust fmt/clippy both exit 0 (
cargo-fmt-final.log,cargo-clippy-final.log). - Post-native-change
PATH=/opt/homebrew/opt/rustup/bin:$PATH cargo test --manifest-path native/Cargo.toml --locked: exit 0, 33 actual tests. Calling cargo by absolute path without putting rustc on PATH first exited 101 before execution; the corrected run iscargo-test-final.log. - Dependency audit
dotnet list Synapse.slnx package --vulnerable --include-transitive --format json: exit 0, NuGet.org, six projects and no reported vulnerable packages (security-vulnerabilities.json). - Final whole Release build/analyzers with independent compiler exits 0,
zero warnings/errors, 5.62 s (
build-current-final-green2.log). Two prior combined builds rejected redundant tuple casts/explicit types through IDE0004/IDE0007; corrected code passes without suppressions. - Final
dotnet format Synapse.slnx --verify-no-changes --no-restore --verbosity minimal: exit 0 (format-current-final-host.log). Two sandbox attempts exit 1 because Roslyn's build-host IPC bind is denied; the narrow host retry passes. It verifies source without changing benchmark binaries. - Both doctor entry points exit 0; the C# entry performs a real ZoneTree
durable round trip (
doctor-final.json,doctor-native-final.json). - Final independent read-only architecture/security review finds all four
earlier concrete defects fixed: nested direct owner math, scheduler callback
self-join, cross-backend pairing and exponent overflow. Prepared provenance,
bounded strict JSON and owned child cancellation remain explicit. Source
searches leave averaging flags only in removed-option rejection tests.
The planned
tools/Synapse.Buildautomated architecture/changed-line gates do not exist yet and are not claimed passed. - Final focused controls/capacity: 33/33; runtime memory/lifecycle: 49/49;
runner: 34/34, all exit 0 and zero skips. Final 34-case runner coverage is
611/655 deduplicated executable lines (93.28%) across 14 current-source
files; final CLI coverage is 511/599 (85.31%) across seven files.
runner-focused-coverage-current.json,memory-focused-coverage.jsonand../optimization-controls/final-cli-coverage-union.jsonretain raw XML inputs and source identities. All collectors instrumented separate copied binaries. A historical runner report and a failed broad instrumentation attempt remain explicitly distinct; neither substitutes for the current module-only green run. These are focused full-file coverage observations, not whole-tree or changed-line qualification. - Canonical whole-solution .NET tests exit 0: 558/558, zero failed/skipped,
6m36.758s. The log has no unhandled/background exception. HTML/TRX are in
full-test-results; the realtest-reportexporter exits 0 and retains reconciled counts infull-test-results.json. The host runs the available dotLLM/llama.cpp processes and prepared model/Foundry cache prerequisites. - Current final managed/native snapshot hashes before/after the real memory
runs and tests match (
cmp, exit 0). Raw snapshots are not collectors.current-final-source.patch, the untracked-source archive and their hashes preserve the working source relative to HEAD; no commit/release is claimed.
Final gate commands (run from repository root; each exits 0):
dotnet restore Synapse.slnx --locked-mode
dotnet build Synapse.slnx --configuration Release --no-restore -p:UseSharedCompilation=false -m:1 /nr:false
dotnet format Synapse.slnx --verify-no-changes --no-restore --verbosity minimal
SYNAPSE_DOTLLM_EXECUTABLE=/Users/ksemenenko/Developer/Synapse/_external/dotLLM/src/DotLLM.Cli/bin/Release/net10.0/DotLLM.Cli SYNAPSE_DOTLLM_VERSION=d88040451d7db56e5dfef9d5754ad0955b0f7fe5 SYNAPSE_LLAMACPP_EXECUTABLE=/opt/homebrew/bin/llama-completion SYNAPSE_LLAMACPP_VERSION=b29c606e2 SYNAPSE_MODEL_ROOT=/Users/ksemenenko/Developer/Synapse/artifacts/models SYNAPSE_FOUNDRY_CACHE=/Users/ksemenenko/Developer/Synapse/artifacts/foundry-local dotnet test Synapse.slnx --configuration Release --no-build --output Detailed --timeout 35m --minimum-expected-tests 1 --zero-tests-policy strict --report-trx --results-directory /Users/ksemenenko/Developer/Synapse/artifacts/optimization-ablation/full-test-results
PATH=/opt/homebrew/opt/rustup/bin:$PATH cargo test --manifest-path native/Cargo.toml --locked
PATH=/opt/homebrew/opt/rustup/bin:$PATH cargo fmt --manifest-path native/Cargo.toml --all --check
PATH=/opt/homebrew/opt/rustup/bin:$PATH cargo clippy --manifest-path native/Cargo.toml --workspace --all-targets -- -D warnings
dotnet list Synapse.slnx package --vulnerable --include-transitive --format json
git diff --check
The original locked restore/audit remain applicable: no package references or locks changed during this slice. Build/style failure logs and sandbox IPC limitations are retained rather than counted as passing attempts. CUDA qualification, held-out/model-family breadth, hosted CI and planned automatic architecture/changed-line gates remain open.
No preparation microbenchmark is claimed as token generation acceleration.