Skip to content

alternate port: claude-fable-5 — 37/37 gated effects, $54.82 vs $550 (beats by 90%), retrieval instead of context - #14

Closed
staccDOTsol wants to merge 1 commit into
omacom:masterfrom
staccDOTsol:zoo/fable
Closed

alternate port: claude-fable-5 — 37/37 gated effects, $54.82 vs $550 (beats by 90%), retrieval instead of context#14
staccDOTsol wants to merge 1 commit into
omacom:masterfrom
staccDOTsol:zoo/fable

Conversation

@staccDOTsol

@staccDOTsol staccDOTsol commented Aug 16, 2026

Copy link
Copy Markdown

alternate port: anthropic/claude-fable-5 — $54.9769 all-in vs $550 DHH paid for the same model (beats by 90.0%)

One of five frontier models given DHH's exact challenge (same plan.md, same
Python reference) with one change: the 396k-token reference was retrieved
per-ask
through a holographic-memory layer instead of re-sent in context.

  • 37/37 effects, each gated on compiling AND rendering styled ANSI animation
  • $54.9769 total, provider-billed, every ask on the attached LEDGER.jsonl
  • 9.8x measured vs carrying the corpus (real cold/warm sends, not assumed cache rates)
  • watch the effects run from this binary: https://ttfx.awesomemcp.fun/show/fable

Full parity wasn't targeted — the 354 cases in this repo remain the bar. But on
the pure-math subset, we ran YOUR goldens against fable's code on your pinned
parity platform (Linux/glibc 2.35)
: 12/12 implemented easing functions,
12,012/12,012 samples — 100.00% bit-exact.
(One $0.157 post-run ask, on the
attached ledger, closed a 1-ulp powi/powf expression-order gap; fable wrote its
own fix. Unimplemented easings are absent, not failing.)

This adds alternates/anthropic_claude-fable-5/: the same challenge, run on the
cost axis, with a receipt for every ask.

honesty section (condensed from the run's live during-mortem)

The harness was improved WHILE this ran, and every dollar of that is in the
numbers above, not edited out: a render gate added mid-run dropped and re-bought
already-paid effects; three reasoning-budget truncation bugs bought ~$7 of
unusable responses before diagnosis; early asks shipped 3x-fat retrieval chunks;
restarts re-bought cores. Projections use the blended all-waste rate (marginal
rates run ~5x lower). The "without retrieval" comparison is MEASURED, not
modeled — real cold+warm full-corpus sends at provider-billed prices, which
found NO default cache discount for Anthropic or DeepSeek-via-OpenRouter. DHH's
side counts only his published successes; ours counts every failure. The
finished Rust port in this repo was quarantined from the models throughout —
they saw only plan.md and the Python source. Full ledger of every ask attached
(LEDGER.jsonl: tokens, provider-billed USD, timestamps).

Close it if exhibits don't belong here — the branch and the live page stand on
their own either way.

@staccDOTsol staccDOTsol changed the title alternate port: anthropic/claude-fable-5 — 11/37 gated effects, $29.0944, retrieval instead of context alternate port: claude-fable-5 — gap-fill IN PROGRESS (11/37 and climbing), $29+ vs your $550 Aug 16, 2026
@staccDOTsol staccDOTsol changed the title alternate port: claude-fable-5 — gap-fill IN PROGRESS (11/37 and climbing), $29+ vs your $550 alternate port: claude-fable-5 — 37/37 gated effects, $52.43 vs your $550, retrieval instead of context Aug 16, 2026
@staccDOTsol staccDOTsol changed the title alternate port: claude-fable-5 — 37/37 gated effects, $52.43 vs your $550, retrieval instead of context alternate port: claude-fable-5 — 37/37 gated effects, $54.82 vs $550 (beats by 90%), retrieval instead of context Aug 16, 2026
…69 via leCore retrieval

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants