[agent] Bench: vlt/hosted and vlt/rescan take about 2.2–2.4x as long since #472 (1169ae6, "Fix Bun/vlt bundled copies left unpatched"). Wall time is up 118–144% and CPU time 91–106%. The request count is unchanged (94). vlt was already the slowest package manager per package before this change, and both vlt scenarios now take over 750 ms.
Same-machine interleaved compare (socket-patch-bench, PR #485 suite)
| round |
base |
scenario |
base wall |
head wall |
Δ wall [95% CI] |
Δ CPU |
Δ RSS |
requests |
| daily, 15 pairs + 10 confirm |
6e7ef748 (main −24h) |
vlt/hosted |
371.6 ms |
802.7 ms |
+117.8% [+109.9, +125.1] |
+91.4% |
+0.0% |
94 |
| daily, 15 pairs + 10 confirm |
6e7ef748 |
vlt/rescan |
345.1 ms |
765.1 ms |
+120.6% [+115.3, +137.7] |
+99.5% |
+0.0% |
94 |
| confirm, 25 pairs + 20 |
6e7ef748 |
vlt/hosted |
381.6 ms |
861.2 ms |
+126.6% [+118.0, +133.1] |
+100.3% |
|
94 |
| confirm, 25 pairs + 20 |
6e7ef748 |
vlt/rescan |
340.5 ms |
837.7 ms |
+144.0% [+136.4, +152.3] |
+106.0% |
|
94 |
| bisect, 11 pairs |
cbf1f748 (parent of #472) |
vlt/hosted |
397.0 ms |
872.8 ms |
+127.8% [+119.3, +140.3] |
+99.5% |
|
94 |
An A/A check on the same runner (including vlt/hosted) found no regression: −1.1% [−3.6, +5.5].
Hot spot
The instruction count barely moves (callgrind: 2.46G → 2.59G, +5%). The time goes into syscalls and thread handoffs. strace -f -c on one vlt/hosted scan:
|
parent cbf1f748 |
#472 |
| total syscalls |
27,433 |
144,713 |
readlink (all EINVAL) |
0 |
72,243 |
futex |
1,148 |
37,200 |
openat |
5,022 |
9,434 |
The cause is the new vendor::vlt_bundled::store_bundled_copies walk. For each of the 1500 lock nodes, it runs tokio::fs::canonicalize on .vlt/<key>/node_modules/<name>, which issues one readlink per path component, and then a tokio::fs::read_dir/file_type/read for every entry. Every one of these is a separate spawn_blocking round trip. The walk runs twice per scan: once from vex::discover::vlt (lockfile discovery) and once from heal_after_rewrite in scan/hosted/vlt.rs.
Proposed fix (not applied):
- run the whole walk in one
spawn_blocking with std::fs;
- canonicalize the root once and check each store entry with
symlink_metadata or read_dir file types rather than a full canonicalize;
- compute the copies once per scan and share them between discovery and heal.
The rest of the vlt cost predates #472: redirect::vlt::every_instance_pinned / partition_instances / vlt_preflight::preflight_scope take about 35% of instructions in vlt_lock_text::split_dep_id / DepId::registry_identity, which suggests a per-patch re-parse of every DepId.
Repro
export CARGO_PROFILE_PERF_INHERITS=release CARGO_PROFILE_PERF_LTO=thin CARGO_PROFILE_PERF_STRIP=none
git worktree add /tmp/base cbf1f748 && (cd /tmp/base && CARGO_TARGET_DIR=/tmp/tb cargo build --locked --profile perf -p socket-patch-cli)
# on a checkout of #485 merged with main:
cargo build --locked --profile perf -p socket-patch-cli -p socket-patch-bench
target/perf/socket-patch-bench compare --base /tmp/tb/perf/socket-patch --head target/perf/socket-patch -f '^vlt/'
target/perf/socket-patch-bench serve vlt/hosted --bin target/perf/socket-patch # then strace -f -c the printed command
Runner: 4 vCPU (nproc = 4), Intel(R) Xeon(R) Processor @ 2.10GHz, cloud sandbox.
Generated by Claude Code
[agent] Bench:
vlt/hostedandvlt/rescantake about 2.2–2.4x as long since #472 (1169ae6, "Fix Bun/vlt bundled copies left unpatched"). Wall time is up 118–144% and CPU time 91–106%. The request count is unchanged (94). vlt was already the slowest package manager per package before this change, and both vlt scenarios now take over 750 ms.Same-machine interleaved
compare(socket-patch-bench, PR #485 suite)6e7ef748(main −24h)vlt/hosted6e7ef748vlt/rescan6e7ef748vlt/hosted6e7ef748vlt/rescancbf1f748(parent of #472)vlt/hostedAn A/A check on the same runner (including
vlt/hosted) found no regression: −1.1% [−3.6, +5.5].1169ae68(main), with the bench suite from Add ascanbenchmark suite and a CI performance gate #485 merged locallyHot spot
The instruction count barely moves (callgrind: 2.46G → 2.59G, +5%). The time goes into syscalls and thread handoffs.
strace -f -con onevlt/hostedscan:cbf1f748readlink(all EINVAL)futexopenatThe cause is the new
vendor::vlt_bundled::store_bundled_copieswalk. For each of the 1500 lock nodes, it runstokio::fs::canonicalizeon.vlt/<key>/node_modules/<name>, which issues onereadlinkper path component, and then atokio::fs::read_dir/file_type/readfor every entry. Every one of these is a separatespawn_blockinground trip. The walk runs twice per scan: once fromvex::discover::vlt(lockfile discovery) and once fromheal_after_rewriteinscan/hosted/vlt.rs.Proposed fix (not applied):
spawn_blockingwithstd::fs;symlink_metadataorread_dirfile types rather than a fullcanonicalize;The rest of the vlt cost predates #472:
redirect::vlt::every_instance_pinned/partition_instances/vlt_preflight::preflight_scopetake about 35% of instructions invlt_lock_text::split_dep_id/DepId::registry_identity, which suggests a per-patch re-parse of every DepId.Repro
Runner: 4 vCPU (
nproc= 4), Intel(R) Xeon(R) Processor @ 2.10GHz, cloud sandbox.Generated by Claude Code