Repository navigation
Add time slicing benchmark - #203
Conversation
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Result-reduction errors and missing incremental benchmark implementations block approval.
Review effort: Balanced
Findings: 2
Open (3)
What changed in this PR
Prepares opt-in incremental slicing benchmarks for NWB data and expands ASV result processing.
Changes:
- Adds modality selection and incremental slicing parameters.
- Extends result normalization and exported slicing metadata.
- Documents runtime environment variables and benchmark commands.
| File | Description |
|---|---|
| src/nwb_benchmarks/setup/_reduce_results.py | Normalizes ASV layouts and skipped results. |
| src/nwb_benchmarks/database/_models.py | Exports slicing template and strategy metadata. |
| src/nwb_benchmarks/benchmarks/params.py | Defines modality-filtered incremental slicing parameters. |
| src/nwb_benchmarks/__init__.py | Parses incremental benchmark opt-in settings. |
| docs/running_benchmarks.rst | Documents environment settings and usage. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
@CodyCBakerPhD the changes to |
Use sphinx-tabs group tabs (as in the nwb2bids docs) so each environment variable example shows a macOS / Linux and a Windows (PowerShell) tab; selecting a platform once switches every tab group on the page. Also adds PowerShell equivalents for the RUN_* examples and fixes a heading underline that was one character short. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015K2S7gEkSTfhpnfbXVD5YG
The incremental slicing benchmarks are a new benchmark family, which AGENTS.md treats as a minor bump. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015K2S7gEkSTfhpnfbXVD5YG
|
Two small points:
|
`track_cumulative_slice_times` now returns
`{"cumulative_time_in_seconds": [...]}` instead of one
`cumulative_slice_NNN` key per step. The keys stopped sorting correctly
past 999 steps (zero-padded to 3 digits), and each one became its own
variable in the database. The list keeps the steps in read order: the
first entry is the open time and entry N includes reading the first N
slices.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015K2S7gEkSTfhpnfbXVD5YG
- `RUN_INCREMENTAL_SLICING_BENCHMARKS` parsing and the modality filter of the incremental slicing parameters. - The incremental slicing helpers and `_track_cumulative_slice_times` on an in-memory NWB file, for both slice strategies. - `_extract_successful_results`, `_serialize_parameter_cases` and `reduce_results` on a raw results file written by asv 0.6.1 with `--record-samples`, holding an incremental, a network and a time benchmark, including a failed and an unselected parameter set; and the database reader on the reduced output. The DANDI URL lookup that importing the benchmark parameters does is stubbed in `conftest.py`, so the tests run offline. A `test` dependency group installs pytest with the database and figure dependencies, and the CI runs the tests before the benchmark smoke test. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015K2S7gEkSTfhpnfbXVD5YG
|
Fixed those small points as well as:
|
…rams - `number` only applies to `time_` benchmarks; ASV calls a `track_` benchmark once per sample regardless, so the override on `IncrementalSliceBenchmark` did nothing. - The `incremental_hdf5_*_params` dicts are shared by the HDF5, Zarr and LINDI parameter sets, so they are now `incremental_*_params`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015K2S7gEkSTfhpnfbXVD5YG
… reducer `track_cumulative_slice_times` now returns `dict(samples=..., number=None)`, like the network tracking benchmarks. With the `--record-samples` flag `nwb_benchmarks run` always passes, ASV then writes the cumulative times to the samples column, which the existing `reduce_results` reads, so the reducer rewrite is no longer needed: - `_reduce_results.py` is back to its version on `main`, with one fix: a parameter set that failed (`null` samples) no longer drops the successful parameter sets of the same benchmark through the length-mismatch warning. The warning now covers only a structural mismatch between the parameter, result and samples lists. - `parse_parameter_case` no longer special-cases `()`, which only zero-parameter benchmarks produce and the suite has none of. - The guard in `normalize_time_and_network_results` stays: the database reader still treats every dict result as network statistics otherwise. The reducer tests now run `reduce_results` on a raw results file regenerated with asv 0.6.1, holding the wrapped incremental result, a network, a time, and an unwrapped `track_` benchmark. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015K2S7gEkSTfhpnfbXVD5YG
…a-rzyanr Wrap incremental slicing results as ASV samples and restore the existing reducer
Follow-up to #203 (targets its branch). The raw results fixture behind `tests/test_reduce_results.py` was written with asv 0.6.1. Main now pins asv 0.6.6 and asv-runner 0.3.1 (#205), so this PR regenerates the fixture with those versions. ## Changes - **Merges `main` into the branch.** It merges cleanly and brings in the #205 pins, so the tests in this PR run against the asv version that wrote the fixture. - **Regenerates the fixture.** It comes from the same toy suite as before: `Incremental` (B raises), `Network`, `Timed` (`--bench` selects only parameter set A) and `Unwrapped`. The suite was run with `asv run --python=same --record-samples` and trimmed to the keys the reducer reads, with the same fixed commit hash. - **Renames the fixture** `tests/data/asv_0.6.1_raw_results.json` → `asv_0.6.6_raw_results.json`, and updates the path and comment in the test. ## What changed in the fixture The structure is the same as under 0.6.1: same `result_columns`, same row lengths (12 for the wrapped benchmarks, 5 for `Unwrapped`), NaN result for the unselected `Timed` B, `null` samples for the failed `Incremental` B, and the same `true` results for the wrapped `track_` benchmarks. Only the `Timed` timing, the benchmark version hashes, `started_at` and `duration` differ. No version bump: nothing under `src/nwb_benchmarks/` changes. ## Testing - `python -m pytest tests`: 40 passed, none skipped (polars and seaborn installed, so the database-reader test ran). - `pre-commit run` on the changed files passes. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01At337MLAynHQXz3TZBF3Vv --- _Generated by [Claude Code](https://claude.ai/code/session_01At337MLAynHQXz3TZBF3Vv)_ Co-authored-by: Claude <noreply@anthropic.com>
…arams Only the benchmark suite reads it (`params.py` filters the parameter sets by modality, `track_incremental_slicing.py` gates the benchmark), but the parser and its warning ran in the package `__init__`, so on every import of `nwb_benchmarks`: an invalid value broke every CLI command and the warning printed on all of them. They now live in `benchmarks/params.py`, next to the modality filter, and `__init__.py` is back to its version on `main`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015K2S7gEkSTfhpnfbXVD5YG
|
@oruebel I also folded the doc snippets into sphinx-tabs (for cross platform) Then moved some code from If you think it still looks good / does what you want go ahead and merge |
Thanks @CodyCBakerPhD for the fixes for the fixes. Looks good to me. |


Benchmarks:
Added opt-in incremental slicing benchmark controls via
RUN_INCREMENTAL_SLICING_BENCHMARKSenvironment variable, including support for running all modalities or selectedecephys,ophys, andicephyscombinations. E.g., to run the incremental slicing tests for just icephys one can call:Documented relevant runtime environment variables for network tracking, download benchmarks, and incremental slicing benchmark selection.
Extended database result export with
slice_templateandslice_strategyparameter metadata.slice_strategytells the benchmark how to perform the incremental read sequence.iterate_time_axis: used for ecephys/ophys array-like datasets. It repeatedly reads adjacent chunks along axis 0.iterate_icephys_timeseries: used for icephys files. It iterates over TimeSeries-like objects in acquisition/stimulus and reads each one.slice_templatedefines the shape of each repeated slice foriterate_time_axis.slice_template, because it reads whole TimeSeries objects instead.Result Parsing:
Note:
LINDI,fsspec HTTPS cached,remfile cache,ROS3, andZarr with consolidated metadatafor remote slicing but we can add the others ones too if you want