Skip to content

Latest commit

 

History

History
279 lines (239 loc) · 17.3 KB

File metadata and controls

279 lines (239 loc) · 17.3 KB

Optimization and dynamic memory evidence, 2026-10-03

Status: bounded local implementation, real cohorts, memory workloads and repository gates verified. Requirements REQ-CTX-009 and REQ-CTX-010; decisions ADR-025 and amended ADR-017. Broad tasks stay open.

Scope and reproducibility

A supported source is converted to .synapse before execution. The real Qwen2.5-0.5B Q8_0 lossless artifact has file SHA-256 b67620a926cd299931b0a7cbfbd31f112c893832658cefa64e67f4d4a21f514d and source SHA-256 ca59ca7f13d0e15a8cfa77bd17e65d24f6844b554a7b6c12e07a5f89ff76844e. Every measured profile binds package, graph, runtime and exact prompt IDs.

The natural haystack concatenates 49 Markdown documents from docs, README.md and synapse.plan.md at 2bcd73c: 289,473 UTF-8 bytes, SHA-256 b71f38b349b40b544deab91359f9f18c8a3ce160ade74d607d7cc83a61d0f175. The tree was exported with git archive; a C# tool normalizes and concatenates sorted paths. Ignored artifacts/optimization-ablation holds plans, snapshots, binary hashes, raw JSON, corpus and logs. Model/generated bytes are not in Git.

The FP32/FP16 Metal plan requests 4,000/8,000/16,000 prompt tokens at three depths for needle and multi-key tasks, plus one variable-tracking case per length: 21 cases, two paired rounds, 84 fresh children. Factory paragraph cuts make actual prompt counts differ from targets; both are recorded. Repeated requests with/without exact prefix reuse form a separate cohort. Profiles rotate order. The same-prompt tail NLL and greedy agreement diagnose numerical drift; exact answers diagnose these bounded tasks. Neither proves general held-out language quality.

Hardware: M2 Pro, 12 CPU cores (8 performance/4 efficiency), 19 GPU cores, 32 GB, macOS 27.0.1 ARM64; .NET SDK 10.0.401/runtime 10.0.12. Real Metal runs use narrow sandbox escalation because the sandbox hides the device. The host also runs unrelated development work; timing is local diagnostic evidence. Independent builds avoid the shared compiler queue using -p:UseSharedCompilation=false -m:1 /nr:false.

Rejected averaging pilot

The real 3,754-token needle at depth 0.5 produced these cold results from five fresh workers (one pair per candidate, not statistical qualification):

Profile TTFT ms Generation ms Output tokens Correct Allocated KV MiB
dense FP32 1518.920 1568.813 8 yes 96
FP16 KV 1391.515 1439.357 8 yes 48
mean2 1635.946 1953.731 16 no 96
mean4 1501.895 1631.828 16 no 96
mean8 1556.896 1719.937 16 no 96

FP16 tail mean absolute NLL drift was 0.002825 with greedy agreement 1. Mean2/4/8 greedy agreement was zero and tail perplexity ratios were 3346.35/8257.71/27187.02. Averaging retained storage/arithmetic and lost answers. The owner rejected it: all averaging code/options and generated mean models are removed. Its earlier red/green tests and focused coverage describe a removed experiment, not current product verification.

Raw: pilot-initial.json; derived: pilot-initial-summary.md; progress: pilot-initial.log under the ignored artifact folder. A single-pair 1.092× FP16 TTFT ratio is not a qualified speedup. The larger cohort below decides the observed speed/quality effect.

Behavior-first verification

CLI mode tests first failed 21/21 because old commands ignored the mode. Runner shell plan/worker cases first failed against real processes. Subsequent source-generated JSON tests exposed omitted init-only defaults becoming zero; internal mutable plan/profile/KV DTOs fix that without changing public runtime options. Real child tests check provenance forgeries, pairing, repeated controls, failure retention, cancellation/PID exit and output sentinel preservation. Final controls/capacity suites contain 33 cases. The runner has 34 cases after the memory workload, backend guard and finite-log score regressions were added.

Growing-only KV and retained disposed CPU arrays are the memory defects under test. ADR-017 adds fourfold contraction hysteresis, preserves surviving bytes, reserves newly assigned batch slots and releases arrays under the execution gate. Real same-backend memory/lifecycle tests pass 49/49, including Metal FP32/FP16, ordered bitwise cold logits, retained-prefix outputs, independent slots and zero owned KV after disposal. Owner guards reject nested direct math and callback disposal before changing state; asynchronous admission is fenced against closing. These defects first failed real behavior tests. The final memory-focused isolated coverage union is 574/644 executable lines (89.13%) across seven runtime files, not whole-tree or changed-line qualification.

The paired runner now rejects another backend before measurement. Real prepared fixtures with large finite output norms previously made worker and memory JSON publication exit 134: finite NLL overflowed its exponential, or the actual score was nonfinite. Seven regressions now pass: finite log evidence survives exponent overflow with nullable perplexity/status, while nonfinite scores fail cleanly and retain preceding generation/memory phases. Ratios use log differences. The helper's isolated file coverage is 14/16 lines (87.5%).

Executed commands and remaining gates

Completed cold FP16 cohort

no-mean-snapshot/Synapse.ReferenceBenchmarks.dll optimization-eval --plan plans/weights.json --output fp16-long-context.json (paths under the artifact folder) exited 0 with completed status: 84 fresh Metal workers, 42 pairs, 21 unique cases repeated twice. All generated continuations are identical; there are zero candidate answer regressions or improvements.

Context FP32 / FP16 TTFT ms (medians) Median paired TTFT ratio FP32 / FP16 generation ms FP32 / FP16 allocated KV MiB Correct answers each
4K 1578.471 / 1523.101 1.026 1678.103 / 1601.854 96 / 48 8/14
8K 3679.897 / 3415.704 1.067 3757.568 / 3488.857 192 / 96 10/14
16K 9801.376 / 8585.593 1.137 9889.396 / 8725.290 384 / 192 8/14

Median mean absolute tail NLL drift is 0.001998 / 0.003250 / 0.002735. Minimum tail greedy agreement is 1 / 0.96875 / 1; maximum mean absolute NLL drift is 0.002510 / 0.003856 / 0.006019. Numerical drift exists even where generated answers are identical. The original 0.5B model itself solves only 26/42 measured requests; this result preserves that bounded behavior.

Timings are descriptive local diagnostics, not a release/statistical speed qualification. Unrelated Prostir work and own independent builds/coverage overlapped parts of the frozen-binary cohort. Candidate order rotates, and raw paired rows retain every output length and failure. The immutable managed DLL hashes before/after match literally (cmp, exit 0); no benchmark snapshot was coverage-instrumented. This snapshot precedes the allocator/lifecycle repair, which changes storage ownership rather than the numerical kernels. The post-repair real memory workload and same-backend bitwise tests validate that separate change.

Raw: fp16-long-context.json; progress: fp16-long-context.progress.log; derived summary: fp16-long-context-summary.md; binary hashes: no-mean-binary-hashes-{before,after}.txt. Generation cap is the lower of the plan and task factory cap (needle/multikey 16; variable tracking 40), recorded in each row. Requested 64 is not silently reported as 64 actual outputs.

Completed repeated-prefix cohort

The separate plans/reuse.json cohort is completed: 24 fresh children, one warmup and three measured paired rounds per context, 18 measured samples. Both profiles use original weights, Metal and FP32 KV; only exact prompt-prefix reuse differs. Each child records cold and repeated generation separately.

Context Dense / reuse cold TTFT ms Dense / reuse repeated TTFT ms Dense / reuse repeated generation ms
4K 1345.045 / 1345.284 1295.466 / 6.025 1336.635 / 49.044
8K 3392.764 / 3381.221 3316.884 / 7.159 3364.065 / 53.693
16K 9519.432 / 9664.538 9224.563 / 15.981 9288.005 / 84.825

All nine measured pairs have identical generated token sequences, exact tail NLL, greedy agreement 1 and correct answers in both turns. This cohort contains only three unique needle cases at depth 0.5; it establishes exact prefix reuse on these cases, not general task quality. It does not accelerate a new prompt. KV capacity stays 96/192/384 MiB: reuse saves repeated prefill work, not the stored attention state. Both managed and native hashes before and after this cohort match (cmp, exit 0).

Raw: reuse-long-context.json; derived: reuse-long-context-summary.md; progress: reuse-long-context.progress.log; hashes: no-mean-binary-hashes-{before,after-reuse}.txt and no-mean-native-hashes-{before-reuse,after-reuse}.txt. The completed raw report and every child's exit are retained; the report utility exits 0. The original parent shell's completion exit was not retained by its terminal session, so it is not separately claimed here.

Dynamic owned memory on the current implementation

current-final-snapshot comes from the post-repair whole Release build and is never coverage-instrumented. Each memory-eval loads the separately prepared original package, runs short→long→short, disposes it, then runs independent fresh-short/fresh-long numerical oracles on the same backend. Generation and subsequent tail-score checkpoints are separate. There is no forced GC, RSS promise, cache flush or timing comparison across these phases.

Backend / KV Prompt tokens short / long Short-first KV MiB Long KV MiB Short-after-long KV MiB Disposed KV MiB
Metal / FP32 456 / 15921 24 384 24 0
Metal / FP16 456 / 15921 12 192 12 0
Native / FP32 456 / 3754 24 96 24 0

All three completed commands exited 0. All nine retained-vs-fresh comparisons independently have identical ordered generated IDs, tail NLL double bits and greedy IDs; absolute NLL difference is 0. This is exact same-backend evidence for contraction; FP16-vs-FP32 drift is separately measured above. Physical footprint, working set and managed heap remain separate metrics in each raw phase; mapped model bytes and other allocations do not disappear with KV. The complete derived memory-real-summary.md preserves all 16 phases per profile and raw SHA-256 identities. Long→short owned KV contracts by 93.75% for the two Metal runs and 75% for the shorter native run. Native context differs, so its timings do not establish a backend speed comparison.

Peak observed footprint is 529.142 / 330.767 / 191.861 MiB for Metal FP32 / Metal FP16 / native FP32; peak working set is 587.953 / 599.125 / 684.625 MiB. These scope differences are real. For FP16, the immediate short-after-long footprint is still 325.329 MiB despite 12 MiB of owned KV; the separately recorded short-score phase observes 133.361 MiB. No immediate RSS/footprint release is promised. Native managed heap at disposal still observes 130.111 MiB without GC while owned KV is zero.

Exact commands, from the repository root:

dotnet artifacts/optimization-ablation/current-final-snapshot/Synapse.ReferenceBenchmarks.dll memory-eval --request artifacts/optimization-ablation/plans/memory/memory-metal-f16-natural.json --output artifacts/optimization-ablation/memory-metal-f16-natural.json
dotnet artifacts/optimization-ablation/current-final-snapshot/Synapse.ReferenceBenchmarks.dll memory-eval --request artifacts/optimization-ablation/plans/memory/memory-metal-f32-natural.json --output artifacts/optimization-ablation/memory-metal-f32-natural.json
dotnet artifacts/optimization-ablation/current-final-snapshot/Synapse.ReferenceBenchmarks.dll memory-eval --request artifacts/optimization-ablation/plans/memory/memory-native-f32-natural.json --output artifacts/optimization-ablation/memory-native-f32-natural.json

Raw/progress and exact prompt manifest live in the artifact folder. The fourfold contraction threshold avoids repeated small reallocations. CPU release removes owned arrays; managed GC can retain former memory temporarily. CUDA contraction is unimplemented/unqualified, and full-context RoPE tables remain a separate profiling opportunity. Neither is claimed solved.

  • Locked restore: exit 0 (restore.log).
  • Post-removal Release build/analyzers with independent compiler: exit 0 (build-no-averaging.log); earlier post-JSON-fix builds had zero warnings.
  • Final Rust fmt/clippy both exit 0 (cargo-fmt-final.log, cargo-clippy-final.log).
  • Post-native-change PATH=/opt/homebrew/opt/rustup/bin:$PATH cargo test --manifest-path native/Cargo.toml --locked: exit 0, 33 actual tests. Calling cargo by absolute path without putting rustc on PATH first exited 101 before execution; the corrected run is cargo-test-final.log.
  • Dependency audit dotnet list Synapse.slnx package --vulnerable --include-transitive --format json: exit 0, NuGet.org, six projects and no reported vulnerable packages (security-vulnerabilities.json).
  • Final whole Release build/analyzers with independent compiler exits 0, zero warnings/errors, 5.62 s (build-current-final-green2.log). Two prior combined builds rejected redundant tuple casts/explicit types through IDE0004/IDE0007; corrected code passes without suppressions.
  • Final dotnet format Synapse.slnx --verify-no-changes --no-restore --verbosity minimal: exit 0 (format-current-final-host.log). Two sandbox attempts exit 1 because Roslyn's build-host IPC bind is denied; the narrow host retry passes. It verifies source without changing benchmark binaries.
  • Both doctor entry points exit 0; the C# entry performs a real ZoneTree durable round trip (doctor-final.json, doctor-native-final.json).
  • Final independent read-only architecture/security review finds all four earlier concrete defects fixed: nested direct owner math, scheduler callback self-join, cross-backend pairing and exponent overflow. Prepared provenance, bounded strict JSON and owned child cancellation remain explicit. Source searches leave averaging flags only in removed-option rejection tests. The planned tools/Synapse.Build automated architecture/changed-line gates do not exist yet and are not claimed passed.
  • Final focused controls/capacity: 33/33; runtime memory/lifecycle: 49/49; runner: 34/34, all exit 0 and zero skips. Final 34-case runner coverage is 611/655 deduplicated executable lines (93.28%) across 14 current-source files; final CLI coverage is 511/599 (85.31%) across seven files. runner-focused-coverage-current.json, memory-focused-coverage.json and ../optimization-controls/final-cli-coverage-union.json retain raw XML inputs and source identities. All collectors instrumented separate copied binaries. A historical runner report and a failed broad instrumentation attempt remain explicitly distinct; neither substitutes for the current module-only green run. These are focused full-file coverage observations, not whole-tree or changed-line qualification.
  • Canonical whole-solution .NET tests exit 0: 558/558, zero failed/skipped, 6m36.758s. The log has no unhandled/background exception. HTML/TRX are in full-test-results; the real test-report exporter exits 0 and retains reconciled counts in full-test-results.json. The host runs the available dotLLM/llama.cpp processes and prepared model/Foundry cache prerequisites.
  • Current final managed/native snapshot hashes before/after the real memory runs and tests match (cmp, exit 0). Raw snapshots are not collectors. current-final-source.patch, the untracked-source archive and their hashes preserve the working source relative to HEAD; no commit/release is claimed.

Final gate commands (run from repository root; each exits 0):

dotnet restore Synapse.slnx --locked-mode
dotnet build Synapse.slnx --configuration Release --no-restore -p:UseSharedCompilation=false -m:1 /nr:false
dotnet format Synapse.slnx --verify-no-changes --no-restore --verbosity minimal
SYNAPSE_DOTLLM_EXECUTABLE=/Users/ksemenenko/Developer/Synapse/_external/dotLLM/src/DotLLM.Cli/bin/Release/net10.0/DotLLM.Cli SYNAPSE_DOTLLM_VERSION=d88040451d7db56e5dfef9d5754ad0955b0f7fe5 SYNAPSE_LLAMACPP_EXECUTABLE=/opt/homebrew/bin/llama-completion SYNAPSE_LLAMACPP_VERSION=b29c606e2 SYNAPSE_MODEL_ROOT=/Users/ksemenenko/Developer/Synapse/artifacts/models SYNAPSE_FOUNDRY_CACHE=/Users/ksemenenko/Developer/Synapse/artifacts/foundry-local dotnet test Synapse.slnx --configuration Release --no-build --output Detailed --timeout 35m --minimum-expected-tests 1 --zero-tests-policy strict --report-trx --results-directory /Users/ksemenenko/Developer/Synapse/artifacts/optimization-ablation/full-test-results
PATH=/opt/homebrew/opt/rustup/bin:$PATH cargo test --manifest-path native/Cargo.toml --locked
PATH=/opt/homebrew/opt/rustup/bin:$PATH cargo fmt --manifest-path native/Cargo.toml --all --check
PATH=/opt/homebrew/opt/rustup/bin:$PATH cargo clippy --manifest-path native/Cargo.toml --workspace --all-targets -- -D warnings
dotnet list Synapse.slnx package --vulnerable --include-transitive --format json
git diff --check

The original locked restore/audit remain applicable: no package references or locks changed during this slice. Build/style failure logs and sandbox IPC limitations are retained rather than counted as passing attempts. CUDA qualification, held-out/model-family breadth, hosted CI and planned automatic architecture/changed-line gates remain open.

No preparation microbenchmark is claimed as token generation acceleration.