Skip to content

Optimize: representative benchmarks, cold_path hints, fair parse benchmark - #2

Merged
gfx merged 4 commits into
mainfrom
claude/fpfmt-crate-optimize-o8fq3y
Jul 23, 2026
Merged

gfx merged 4 commits into
mainfrom
claude/fpfmt-crate-optimize-o8fq3y

Conversation

@gfx

@gfx gfx commented Jul 23, 2026

Copy link
Copy Markdown
Member

Summary

Performance/benchmark work on the fpfmt crate, keeping correctness intact throughout (full comprehensive suite passes on both feature sets; the parse-scan change was additionally cross-checked with a 5.3M-string differential test before being reverted — see below).

Measured on Intel Xeon (2.10 GHz), Ubuntu 24.04, x86_64, rustc 1.95.

Changes

1. Benchmark at opt-level 3 ([profile.bench]). The workspace [profile.release] uses opt-level = "z" to keep the wasm binary small, and bench inherited it — so benchmarks were measured size-optimized, which no consumer does (a dependency's profile never propagates; downstream builds fpfmt at their own opt-level). A dedicated [profile.bench] opt-level = 3 fixes this while release stays "z" for the wasm story.

  • format 235 ns → 197 ns in the reported benchmark; ryu now measured fairly (was penalized ~5× by size-opt). Wasm binaries unchanged.

2. core::hint::cold_path on rarely-taken branches (subnormal/zero in unpack64, exact-power-of-two and subnormal in short). Zero-cost, changes no results. Runtime effect is within noise for realistic input (those branches are already never taken for normal doubles), but the hint still guides code layout at no cost to the hot path.

  • Requires the toolchain pinned to exactly 1.95 (the release that stabilized cold_path). 1.96/1.97 regress the wasm-size build ~4× on small for this crate, so they are avoided. rust-version = "1.95" declares the new minimum.

3. Fair parse benchmark. The parse row fed each value's f64::to_string(), which expands 5e-324/f64::MAX to 300+ digit decimals that fpfmt's 19-significant-digit parser rejects while stdlib parses — timing rejection against parsing. It now parses each value's shortest round-trip string (valid input both accept).

  • Honest parse numbers: fpfmt 67 ns / small 80 ns / stdlib 113 ns.
  • Note: I prototyped a digit-count early-exit in parse_text (huge speedup on the old benchmark) but an A/B showed it slows valid parsing ~35%, so it was not kept — the per-digit branch pessimizes the common path to speed up rare adversarial rejection.

Benchmarks (8 representative f64 values)

Task fpfmt fpfmt small ryu stdlib
format 150 ns 237 ns 308 ns 927 ns
parse 67 ns 80 ns — 113 ns

Wasm size unchanged from the 1.93 baseline (14,291 / 4,078 bytes on 1.95).

Validation

  • cargo test + cargo test --features small — pass
  • cargo clippy + --features small — clean
  • cargo fmt --check — clean
  • wasm-size builds — 14,291 / 4,078 bytes

🤖 Generated with Claude Code


Generated by Claude Code

claude added 4 commits July 23, 2026 05:15
The workspace `[profile.release]` sets `opt-level = "z"` to keep the shipped
`wasm32` binary small. Because the `bench` profile inherits from `release`,
the benchmarks were being compiled size-optimized — which is not how any
consumer builds this crate. A dependency's own profile never propagates to
the crates that depend on it; downstream users compile fpfmt at their own
opt-level, almost always 3.

Add a dedicated `[profile.bench]` with `opt-level = 3` so the benchmarks
measure what users actually get, while `release` stays at `"z"` for the wasm
size story (wasm binaries are unchanged: 14,379 / 4,175 bytes).

Measured effect (8 representative f64 values, x86_64):
  format  235 ns -> 197 ns  (fpfmt)
  ryu    1978 ns -> 376 ns  (size-opt had penalized ryu ~5x; now fair)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BqVZdaF64NvBgk8LJQhUeD
Annotate the genuinely-rare branches in the format path — subnormal/zero in
unpack64, and the exact-power-of-two and subnormal cases in short — with
core::hint::cold_path() so the optimizer keeps them off the hot path's code
layout. The intrinsic (stabilized in 1.95.0) is zero-cost and does not change
any result; all tests pass on both feature sets.

On careful interleaved A/B measurement the runtime effect is within noise for
realistic input (those branches are already never taken when formatting normal
doubles), but the hint still influences code layout, at no cost to the common
path.

Toolchain is pinned to exactly 1.95 (the release that stabilized cold_path):
1.96 and 1.97 regress the wasm-size build badly for this crate (default
14,291 -> 26,496 bytes, small 4,078 -> 16,283 bytes), which would undermine
the crate's size story. On 1.95 the wasm is unchanged/slightly smaller than
the previous 1.93 pin (14,291 / 4,078 vs 14,379 / 4,175). rust-version is set
to 1.95 to declare the new minimum.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BqVZdaF64NvBgk8LJQhUeD
Re-measured on this machine (Intel Xeon 2.10 GHz, Ubuntu 24.04, x86_64,
rustc 1.95) with the benchmarks now built at opt-level 3. Refresh the
benchmark table and the wasm-size table (14,291 / 4,078 bytes on 1.95),
and adjust the `small`-vs-ryu formatting ratio to match this machine.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BqVZdaF64NvBgk8LJQhUeD
The parse benchmark fed each value's `f64::to_string()` to the parsers. For
`5e-324` and `f64::MAX` that expands to a 300+ digit decimal, which fpfmt's
parser (capped at 19 significant digits) rejects outright while stdlib parses
in full — so the row was timing fpfmt's rejection against stdlib's parsing,
not a like-for-like comparison.

Parse the shortest round-trip string each value formats to instead (valid
input both parsers accept). The parse row now reflects real parsing:

  parse  fpfmt 67 ns  |  small 80 ns  |  stdlib 113 ns   (was fpfmt 422 ns
  against stdlib 1109 ns, dominated by rejecting two over-long strings).

Format numbers are re-measured in the same run for consistency, and the
`small` note now mentions its ~1.2x parsing slowdown (prescale computes the
factor instead of a table lookup) rather than claiming parsing is unaffected.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BqVZdaF64NvBgk8LJQhUeD
@gfx
gfx merged commit ae18700 into main Jul 23, 2026
@gfx
gfx deleted the claude/fpfmt-crate-optimize-o8fq3y branch July 23, 2026 09:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants