Skip to content

Perf regression: vlt/hosted wall +127% (1169ae68, #472) #579

Description

[agent] Bench: vlt/hosted and vlt/rescan take about 2.2–2.4x as long since #472 (1169ae6, "Fix Bun/vlt bundled copies left unpatched"). Wall time is up 118–144% and CPU time 91–106%. The request count is unchanged (94). vlt was already the slowest package manager per package before this change, and both vlt scenarios now take over 750 ms.

Same-machine interleaved compare (socket-patch-bench, PR #485 suite)

round base scenario base wall head wall Δ wall [95% CI] Δ CPU Δ RSS requests
daily, 15 pairs + 10 confirm 6e7ef748 (main −24h) vlt/hosted 371.6 ms 802.7 ms +117.8% [+109.9, +125.1] +91.4% +0.0% 94
daily, 15 pairs + 10 confirm 6e7ef748 vlt/rescan 345.1 ms 765.1 ms +120.6% [+115.3, +137.7] +99.5% +0.0% 94
confirm, 25 pairs + 20 6e7ef748 vlt/hosted 381.6 ms 861.2 ms +126.6% [+118.0, +133.1] +100.3% 94
confirm, 25 pairs + 20 6e7ef748 vlt/rescan 340.5 ms 837.7 ms +144.0% [+136.4, +152.3] +106.0% 94
bisect, 11 pairs cbf1f748 (parent of #472) vlt/hosted 397.0 ms 872.8 ms +127.8% [+119.3, +140.3] +99.5% 94

An A/A check on the same runner (including vlt/hosted) found no regression: −1.1% [−3.6, +5.5].

Hot spot

The instruction count barely moves (callgrind: 2.46G → 2.59G, +5%). The time goes into syscalls and thread handoffs. strace -f -c on one vlt/hosted scan:

parent cbf1f748 #472
total syscalls 27,433 144,713
readlink (all EINVAL) 0 72,243
futex 1,148 37,200
openat 5,022 9,434

The cause is the new vendor::vlt_bundled::store_bundled_copies walk. For each of the 1500 lock nodes, it runs tokio::fs::canonicalize on .vlt/<key>/node_modules/<name>, which issues one readlink per path component, and then a tokio::fs::read_dir/file_type/read for every entry. Every one of these is a separate spawn_blocking round trip. The walk runs twice per scan: once from vex::discover::vlt (lockfile discovery) and once from heal_after_rewrite in scan/hosted/vlt.rs.

Proposed fix (not applied):

  • run the whole walk in one spawn_blocking with std::fs;
  • canonicalize the root once and check each store entry with symlink_metadata or read_dir file types rather than a full canonicalize;
  • compute the copies once per scan and share them between discovery and heal.

The rest of the vlt cost predates #472: redirect::vlt::every_instance_pinned / partition_instances / vlt_preflight::preflight_scope take about 35% of instructions in vlt_lock_text::split_dep_id / DepId::registry_identity, which suggests a per-patch re-parse of every DepId.

Repro

export CARGO_PROFILE_PERF_INHERITS=release CARGO_PROFILE_PERF_LTO=thin CARGO_PROFILE_PERF_STRIP=none
git worktree add /tmp/base cbf1f748 && (cd /tmp/base && CARGO_TARGET_DIR=/tmp/tb cargo build --locked --profile perf -p socket-patch-cli)
# on a checkout of #485 merged with main:
cargo build --locked --profile perf -p socket-patch-cli -p socket-patch-bench
target/perf/socket-patch-bench compare --base /tmp/tb/perf/socket-patch --head target/perf/socket-patch -f '^vlt/'
target/perf/socket-patch-bench serve vlt/hosted --bin target/perf/socket-patch   # then strace -f -c the printed command

Runner: 4 vCPU (nproc = 4), Intel(R) Xeon(R) Processor @ 2.10GHz, cloud sandbox.


Generated by Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions