End-user CLIs for preparing scene tests and analyzing profiling data or args dumps.
All are invokable as Python modules once the simpler wheel is installed —
no repo checkout required.
Dev-only scripts (
benchmark_rounds.sh,verify_packaging.sh) live in the repo-leveltools/directory and are not shipped.
- scene_test_compile — collect and compile selected Scene Test callables without an NPU
- swimlane_converter — perf JSON → Chrome Trace Event (Perfetto)
- sched_overhead_analysis — scheduler overhead / Tail OH breakdown
- critical_path — chip swimlane critical-path compute/stall analysis
- strace_timing — per-stage
chip.runbreakdown (host + AICPU phases) from[STRACE]log markers → TPOT table, per-round table (--rounds-table), nested tree (--tree), or Perfetto JSON - dump_viewer — inspect / export args dumps (see docs/args-dump.md for full workflow)
- deps_viewer —
deps.json(dep_gen) → text or pan/zoom HTML dependency graph
For CLIs that allow an omitted input, auto-detection paths
(outputs/*/chip_swimlane_records.json, outputs/*/args_dump/) are resolved
relative to the current working directory — run these from the directory
that holds your outputs/. Each test case writes into its own
outputs/<case>_<ts>/ directory; those tools auto-pick the latest by mtime.
Populate the persistent Scene Test kernel cache without creating a Worker or
accessing an NPU. Arguments after the module name are passed to pytest
collection, so the warm-up can use the same paths and selection filters as the
later device run.
python -m simpler_setup.tools.scene_test_compile examples tests/st \
-m "not sdma" --platform a2a3 --require-pto-isa --compile-workers 8Compiled ChipCallable blobs are stored under build/cache/kernels/, with
independent incore artifacts under build/cache/kernels/incore/. A normal
pytest or standalone scene-test run first loads a matching callable blob. On a
callable miss, unchanged incore artifacts are reused and only missing kernels
are compiled before assembly. Source, transitive-include, compiler,
compilation-logic, or compiler-visible path changes produce the corresponding
new content key. The compiler runs from the checkout root, so checkout-local
paths are stable relative paths: path-sensitive macros remain correct and cache
entries can move between CI runners. Paths outside the checkout remain absolute.
Entries unused for 14 days are pruned. Cache misses compile with an automatic
process-wide budget: two logical CPUs remain available to pytest/Python and at
most eight compiler processes run across all test classes and callable artifacts.
--compile-workers N overrides that budget. Class-level and per-callable
parallelism share it, so their worker counts never multiply into additional
compiler processes in one process. A class that fails to compile is reported
without aborting the rest of the pass. The warm-up does not inspect or access
NPU devices.
Post-processing analysis over an chip-swimlane run. Given a run directory, it recursively discovers every directory containing all three required artifacts:
chip_swimlane_records.json(or legacyl2_perf_records.json)deps.jsonname_map*.json(the newest matching sibling is used when several exist)
When a sibling merged_swimlane*.json is present, the newest matching file is
used as the source for two additional Perfetto-compatible traces:
CPM_static.json— highlights the Static CPM task set.CPM_observed.json— highlights the Observed path task set.
Both files retain every view, metadata event, bar, and flow from the merged
trace. Only AIC/AIV task bars in Worker View (pid=4) outside the selected path
are renamed to ·(rXtY) (or ·(tY) for ring 0); path bars and every slice
in other views keep their original names. With Perfetto's default name-based
coloring, the middle dot maps to a light blue-purple while digits are removed
before hashing, so anonymous Worker View bars share a subdued color and path
bars remain grouped by function. dummy(...) and alloc(...) bars keep their
original names because the AICore critical-path model does not classify them.
This supports both simpler directories such as outputs/<case>_<ts>/ and
PyPTO directories such as build_output/<case>/dfx_outputs/, including nested
rank/device layouts. Each discovered artifact directory is analyzed separately,
and its critical_path_report.md is written beside
chip_swimlane_records.json. Pointing the command at a whole run therefore
creates one local report per rank/device rather than one combined report at the
scan root.
For each rank/device, the tool builds a happens-before DAG from the dependency graph oriented by observed timestamps, then computes two critical paths:
- Static CPM — the longest duration-weighted path, i.e. the dependency-limited latency floor with unlimited cores.
- Observed — the as-executed backward blame walk from the last-finishing
task. Each task's compute plus its preceding scheduling stall (
data-wait,core-wait, orfront-gap) tiles the makespan exactly.
The report contains a makespan/CPM/compute/stall overview, a per-kernel-family
table, and a full per-task listing. It is pure post-processing: no C++ or device
is required. If no sibling merged_swimlane*.json exists, Markdown analysis
still succeeds and the tool prints a warning that the two Perfetto traces were
skipped.
# Analyze one simpler output directory or an entire PyPTO run tree
python -m simpler_setup.tools.critical_path outputs/<case>_<ts>
python -m simpler_setup.tools.critical_path build_output/<case>
# Customize each local report filename/table size and print a combined stdout view
python -m simpler_setup.tools.critical_path <run-dir> \
--report critical_path_report.md --top 25 --stdout--report accepts a filename, not a path, so the report cannot be redirected
away from the directory containing its source chip_swimlane_records.json.
Convert performance profiling JSON files into Chrome Trace Event format for visualization in Perfetto.
Converts simpler profiling data (chip_swimlane_records_*.json) into the format used by the Perfetto trace viewer (https://ui.perfetto.dev/) and prints a per-function task-execution summary. With --overhead (needs deps.json) it also adds an Overhead Analysis counter group under the AICPU Scheduler track — 8 lines (oh_{aic,aiv}_{idle,ready,overhead} + oh_all_overhead / oh_has_overhead) you can overlay on the task bars. See docs/dfx/sched-overhead-model.md for the model.
# Auto-detect the latest profiling file under ./outputs/
python -m simpler_setup.tools.swimlane_converter
# Specify an input file
python -m simpler_setup.tools.swimlane_converter outputs/<case>_<ts>/chip_swimlane_records.json
# A unique sibling name_map*.json is loaded automatically.
# Override it explicitly when needed:
python -m simpler_setup.tools.swimlane_converter outputs/<case>_<ts>/chip_swimlane_records.json \
--func-names outputs/<case>_<ts>/name_map_<case>.json
# Specify an output file
python -m simpler_setup.tools.swimlane_converter outputs/<case>_<ts>/chip_swimlane_records.json -o custom_output.json
# Load function name mapping from kernel_config.py
python -m simpler_setup.tools.swimlane_converter outputs/<case>_<ts>/chip_swimlane_records.json \
-k examples/host_build_graph/paged_attention/kernels/kernel_config.py
# Verbose mode (for debugging)
python -m simpler_setup.tools.swimlane_converter outputs/<case>_<ts>/chip_swimlane_records.json -v
# Reuse a deps.json captured in an earlier dep_gen run (different output dir)
python -m simpler_setup.tools.swimlane_converter outputs/<case>_<ts>/chip_swimlane_records.json \
--deps-json outputs/<case>_<earlier_ts>/deps.jsonDependency arrows in the Perfetto trace come from
deps.json(dep_gen replay). The device hot path no longer records fanout, so the typical workflow is two runs: a one-time--enable-dep-gencapture per topology to producedeps.json, then any number of--enable-chip-swimlaneruns that consume it. If nodeps.jsonis found alongside the perf JSON (and--deps-jsonisn't passed), the trace still renders but has no arrows; the converter prints a warning.
When neither --func-names nor --kernel-config is specified, the converter
loads a unique name_map*.json next to the input file. If that directory
contains multiple matching files, it prints a warning and uses default function
labels until one is selected explicitly with --func-names.
For SPMD logical tasks (block_num > 1 in deps.json), dependency
arrows anchor on representative subtask rows on physical core lanes
(not a dedicated block-level track). SPMD tasks use the minimum-core_id
subtask row per core_type as the dependency anchor; MIX-type SPMD
tasks pick the minimum separately for AIC and AIV. See
docs/dfx/chip-swimlane-profiling.md §3.5.
Each logical (pred, succ) edge emits flows for the Cartesian product
of pred/succ anchor rows (|pred_anchors| × |succ_anchors|), not a
per-subtask crossbar.
SPMD lane labels append _spmd before (rXtY) unless the function
name already contains spmd (case-insensitive), e.g.
v_proj_spmd(r2t10) vs SPMD_WRITE_AIV(t0).
With -v, the converter prints
dependency arrows anchor on min core_id subtask per core_type when
SPMD tasks are present.
| Option | Short | Description |
|---|---|---|
input |
Input JSON file (chip_swimlane_records_*.json). If omitted, the latest file in outputs/ is used | |
--output |
-o |
Output JSON file (default: outputs/merged_swimlane_<timestamp>.json) |
--kernel-config |
-k |
Path to kernel_config.py, used for function name mapping |
--func-names |
Path to name_map*.json (SceneTest format) for function name mapping | |
--deps-json |
Path to a dep_gen deps.json (defaults to sibling of input). Without one, no dependency arrows are drawn. |
|
--overhead |
Add the 8-line Overhead Analysis counter group (needs deps.json). See sched-overhead-model. |
|
--verbose |
-v |
Enable verbose output |
The tool produces three kinds of output:
A Chrome Trace Event format JSON file that can be visualized in Perfetto:
- File location:
outputs/merged_swimlane_<timestamp>.json - Open https://ui.perfetto.dev/ and drag-and-drop the file to visualize
A statistics summary grouped by function (printed to the console), including Exec/Latency comparison and scheduling overhead analysis:
- Exec: kernel execution time on AICore (end_time - start_time)
- Latency: end-to-end latency from the AICPU perspective (finish_time - dispatch_time, including head OH + Exec + tail OH)
- Head/Tail OH: scheduling head/tail overhead
- Exec_%: Exec / Latency percentage (kernel utilization)
The table prints the source chip_swimlane_level recorded in
chip_swimlane_records.json. At level 1, only AICore timing is captured, so
Latency, Exec%, Head/Tail OH, and Propagation render as -, including total
latency in the TOTAL row. Count, Exec, and Local Setup remain available. The
Total Test Time line is omitted and replaced by an AICore Observed Span
summary. Level 2 and above retain the full latency summary.
swimlane_converter no longer runs the deep-dive inline — it needs the task DAG
(deps.json) from a separate --enable-dep-gen run, which can't be produced
accurately alongside the swimlane capture. Run
sched_overhead_analysis manually with both
artifacts to get the scheduler-starvation / critical-path report.
When running a test with profiling enabled, the converter is invoked automatically:
# Run the test with profiling enabled - merged_swimlane.json is generated automatically after the test passes
python examples/scripts/run_example.py \
-k examples/host_build_graph/vector_example/kernels \
-g examples/host_build_graph/vector_example/golden.py \
--enable-chip-swimlaneAfter the test passes, the tool will:
- Auto-detect the latest
chip_swimlane_records_*.jsonin outputs/ - Load function names from the kernel_config.py specified via
-k - Produce
merged_swimlane_*.jsonfor visualization - Print the task statistics and scheduler overhead deep-dive report to the console
Answer "is the AICPU scheduler the bottleneck, or is it starved?" by measuring, dependency- and MIX-aware, how much of the makespan a free core has ready, undispatched work — vs. legitimately busy or dependency-limited. Full model: docs/dfx/sched-overhead-model.md.
sched_overhead_analysis needs two artifacts, captured in SEPARATE runs
(co-running the flags perturbs timing — dep_gen adds per-submit overhead):
- Perf profiling data (
chip_swimlane_records_*.json, level >= 3) from a--enable-chip-swimlanerun — per-task dispatch/start/end/finish +aicpu_scheduler_phases. deps.json(the task DAG) from a separate--enable-dep-genrun. It drivesready(C) = max(producer.end), which is what separates scheduler bubbles from dependency stalls. Required — the tool errors without it.
# Capture once (two separate runs of the same case):
pytest <case> --platform a2a3 --device N --enable-dep-gen # -> deps.json
pytest <case> --platform a2a3 --device N --enable-chip-swimlane # -> chip_swimlane_records.json (clean timing)
# Analyze:
python -m simpler_setup.tools.sched_overhead_analysis \
--chip-swimlane-records-json outputs/<swimlane case>/chip_swimlane_records.json \
--deps-json outputs/<dep_gen case>/deps.json
deps.jsonis topology-invariant — capture it once per graph and reuse it for any number of swimlane runs. For Host / Device / Effective / Orch / Sched timing from a plain run, usestrace_timing --rounds-tableinstead.
| Option | Description |
|---|---|
--chip-swimlane-records-json |
Path to the chip_swimlane_records_*.json file (level >= 3). If omitted, the latest under outputs/ is auto-selected. |
--deps-json |
Path to deps.json from a --enable-dep-gen run. Required. Falls back to a deps.json sibling of the perf JSON if present. |
Emitted in six parts:
- Part 1: Overhead verdict — per-engine overhead (idle T-core and a ready, undispatched T-task, MIX-aware) + system
all_overhead/has_overhead, all as % of makespan. An engine with no ready work is not overhead (dependency-mandated idle, not waste). - Part 2: aicore switch — the pre-dispatched pickup gap (
dispatch < prev_end), reported per core (min/mean/max, ~0.8 µs each), the overhead-vs-independent split, and the makespan switch bound[min over cores, sum of per-engine minima]. - Part 3 / 4: Head / Tail OH distributions — P10–P99 + mean + total (per-task pickup and detect-latency magnitude).
- Part 5: AICPU scheduler loop breakdown — per-thread loops, ns/loop, complete/dispatch/idle phase ratios, pop_hit / pop_miss, fanout / fanin, + the tail-vs-loop cause analysis.
- Part 6: Critical-path latency attribution — along the makespan path, scheduler-injected µs vs compute µs ("scheduler adds X% to the critical path").
The perf JSON must be captured at chip_swimlane_level >= 3 so that aicpu_scheduler_phases is non-empty (rerun the case with --enable-chip-swimlane if the tool reports the field is missing).
Per-stage breakdown of every simpler_run() from [STRACE] host-trace
markers in a log (host stderr or CANN device log). The runtime emits one
[STRACE] line per span on scope exit (RAII, gated on SIMPLER_HOST_STRACE,
LOG_TIMING), including the AICPU device-phase subdivision (clk=dev). See
docs/dfx/host-trace.md for the marker grammar.
# Per-callable TPOT table (decode = most-invoked hid bucket; prefill = once-seen)
python -m simpler_setup.tools.strace_timing path/to/log
# Per-round Host/Device/Orch/Sched table (the benchmark/--rounds N view)
python -m simpler_setup.tools.strace_timing path/to/log --rounds-table
# Indented nested span tree per callable (chip.run → bind / runner_run →
# device_wall → preamble/config_validate/arena_wire/sm_reset/orch/sched/post_orch)
python -m simpler_setup.tools.strace_timing path/to/log --tree
# Also emit a Chrome-trace / Perfetto JSON (one named lane per invocation, with
# separate host and device(clk=dev) tracks; nested by span containment)
python -m simpler_setup.tools.strace_timing path/to/log --trace-out strace.json
# L3/L4 host scheduler timeline (real OS pid/tid lanes + cross-thread flows)
python -m simpler_setup.tools.strace_timing path/to/log --swimlane host_swimlane.jsonGroups spans by (pid, inv), rebuilds each invocation's tree from depth,
buckets by callable hash hid, and reports each callable's mean chip.run
plus per-stage means. It reads the host-emitted [STRACE] lines and shows the
host stages (bind/runner_run/validate) alongside the AICPU phases.
--tree renders one nested span tree per callable; each node's duration is the
median across every invocation of that callable (not one invocation's
value). This matters for a callable whose invocations differ in cost — e.g.
qwen3 decode, where the pypto-serving profile warmup dispatches a tiny-KV step
(seq_len≈257, ~28 ms) before the real 3.5k-context steps (~40 ms); a
single-invocation tree would report the warmup value.
--rounds-table renders one row per invocation of the busiest hid —
Host always, plus every device column whose marker is present, in the format
tools/benchmark_rounds.sh parses. TMR normally supplies Device / Effective /
Orch / Sched. HBG supplies Device but no device-side orch/sched windows, so its
table contains Host / Device only. Effective is the TMR orch∪sched merged
window (max(orch_end,sched_end) − min(orch_start,sched_start), the old
device-log "Total"), recomputed from the orch/sched markers' ts+dur — no
device log needed. The scene test only emits the markers to stderr; tee a run
to a file (python test_*.py … --rounds N > run.log 2>&1) and pass run.log
here. Because grouping is per (pid, inv), this captures L3 multi-round
(every chip-child invocation), not just round 0.
--swimlane consumes the <level>.* host-scheduler markers (host.,
network1., network2., network3.) and child chip.run markers, plus any
ext.<producer>.* spans a producer outside simpler emitted. Host lanes retain
their OS pid/tid. Because Chrome Trace
JSON has one visible timestamp axis, raw device-domain clk=dev slices are
stored in the top-level unalignedDeviceSpans array rather than placed beside
the unrelated host clock and stretching Perfetto into an empty-looking
multi-day viewport. Their ns timestamps remain unchanged; no clock offset is
invented. This does not alter the established per-invocation --trace-out
view.
The swimlane is the only view that renders ext. spans: every table and
--trace-out keys on (pid, inv), which no external producer has. See
docs/dfx/host-trace.md for that contract.
Render the dep_gen deps.json task graph as either grep-friendly text
(default) or a self-contained pan/zoom HTML page. Pairs naturally with
swimlane_converter: swimlane is the timing view,
this is the structural view.
deps_viewer reads deps.json produced by the dep_gen replay (see
docs/dfx/dep-gen.md) and supports two modes:
- Default text mode — emits
deps_viewer.txtwith:SUMMARY(input path plus task / edge / tensor counts)tasks: number of rendered task idsunique_task_edges: number of unique(pred, succ)pairsannotated_edges: total number of annotated edge rowsperf_sidecar:yeswhenchip_swimlane_records.jsonwas successfully loadedfunc_name_map:yeswhen at least one task name resolved to a namedfunc_namefrom--func-namesor an auto-discoveredname_map*.json.func_name_mapstaysnounless a real human-readable name was resolved.
TASK INDEX(one line per task for grep)kind=distinguishessubmit/dummy/alloc/unknownfunc_id=is taken only fromtasks[].kernel_idsand shows the aligned three-slot[aic,aiv0,aiv1]array forsubmitkind=alloc/kind=dummyrender asfunc_id=none
TASK DETAILS(per-taskFANIN/FANOUTblocks showing peer task references only) Best for "what does task X depend on?" and large-graph debugging.
--format html— renders the task graph as Graphviz SVG wrapped in a self-contained HTML file viewable in any modern browser.- Add
--show-tensor-infoto restore per-task tensor rows and edge routing to specific arg ports in the HTML view.
- Add
# Auto-pick the newest deps.json under ./outputs/ -> deps_viewer.txt
python -m simpler_setup.tools.deps_viewer
# Specific path -> deps_viewer.txt next to deps.json
python -m simpler_setup.tools.deps_viewer outputs/<case>_<ts>/deps.json
# Explicit text output path
python -m simpler_setup.tools.deps_viewer outputs/<case>_<ts>/deps.json -o graph.txt
# HTML output
python -m simpler_setup.tools.deps_viewer outputs/<case>_<ts>/deps.json \
--format html -o graph.html
# HTML output with per-task tensor details and arg-port routing
python -m simpler_setup.tools.deps_viewer outputs/<case>_<ts>/deps.json \
--format html --show-tensor-info -o graph.html
# Force-directed HTML layout for large graphs (>~1000 nodes)
python -m simpler_setup.tools.deps_viewer outputs/<case>_<ts>/deps.json \
--format html --engine sfdp
# Override task labels with a func_id -> name mapping
python -m simpler_setup.tools.deps_viewer outputs/<case>_<ts>/deps.json \
--func-names outputs/<case>_<ts>/name_map_TestPA_basic.json
# Transitive reduction: select non-redundant edges, print what was removed
python -m simpler_setup.tools.deps_viewer outputs/<case>_<ts>/deps.json \
--edge-mode reduced
# Redundant-only: select the transitively-implied edges reduced would drop
python -m simpler_setup.tools.deps_viewer outputs/<case>_<ts>/deps.json \
--edge-mode omitted
# Dataflow-verified view: preserve OUTPUT_EXISTING reuse boundaries and require
# direct TensorMap dataflow around every byte of an omitted INOUT
python -m simpler_setup.tools.deps_viewer outputs/<case>_<ts>/deps.json \
--edge-mode omitted_dataflow--edge-mode selects which structural (pred, succ) edges are visible:
full(default) — every dependency edge.reduced— the transitively-reduced scheduling edge set: anexplicitortensormapedge already implied by a longer path is dropped, e.g.A->CwhenA->B->Cexists. Acreatoredge is always retained because it keeps the task that owns a tensor referenced by the consumer alive; execution order alone cannot replace that lifetime relationship.omitted— only the redundant edgesreducedwould drop (its complement), for auditing exactly which dependencies are transitively covered.reduced_dataflow— structural reduction only selects candidate edges; it does not reduce the annotations used for proof. Every candidate is checked against the complete originalcreatorandtensormapannotations before display filtering. AnOUTPUT_EXISTINGcreator edge is always preserved as a possible reuse-generation boundary. AnINOUTcreator edge is omitted only when directtensormapannotations prove that every occupied byte flows from an earlier Output and continues to a laterINOUTowned by the same creator. Regions are derived from the underlyingbuffer_addr, dtype, shape, start offset, and strides. Missing, ambiguous, or excessively complex metadata is preserved conservatively.omitted_dataflow— only structurally redundant edges that pass the dataflow proof; the complement ofreduced_dataflow.
All reduction modes print the redundant edges to stdout as a
<task> -> <task> list, where each task uses the same label as the rendered
graph — the bare local counter when every task is in ring 0, or the explicit
(ring, local) tuple once any task lives in ring >= 1. Text output emits only
the selected edge set. HTML output keeps every edge in the Graphviz layout and
colors unselected edges like the page background, so reduced / omitted
preserve the full-graph node placement and routing while showing only the
selected edge set. Selected edges are drawn above background-colored edges so
they stay visible where routes overlap. When -o is omitted the graph is
written to a mode-specific stem (deps_viewer_reduced.* /
deps_viewer_omitted.*) rather than deps_viewer.* so it never clobbers a
full-graph render in the same directory. In reduced / omitted, any
annotation with source=creator protects that (pred, succ) pair from
reduction. The dataflow modes use the complete original annotations to prove
the narrow exception described above. Reduction is skipped with a warning if
the graph contains a cycle.
| Option | Short | Description |
|---|---|---|
input |
Path to deps.json (default: newest under ./outputs/) |
|
--output |
-o |
Output path; default stem is deps_viewer, or deps_viewer_{mode} for any reduction mode |
--format |
Output format: text (default) or html |
|
--edge-mode |
Select visible edges: full, reduced, omitted, reduced_dataflow, or omitted_dataflow; HTML preserves full layout. |
|
--engine |
HTML-only Graphviz layout engine: dot (default), sfdp, neato, fdp, circo, twopi |
|
--direction |
HTML-only flow direction for hierarchical layouts: LR (default) / TB / BT / RL |
|
--show-tensor-info |
HTML-only: render per-task tensor rows and route edges to specific arg ports | |
--func-names |
JSON file with callable_id_to_name (or flat {func_id: name}) for task-label enrichment |
Text output has no extra dependencies. HTML output requires Graphviz on PATH:
brew install graphviz # macOS
apt install graphviz # Debian/UbuntuThe HTML viewer is self-contained — no JavaScript or fonts are downloaded at view time.
- drag → pan
- scroll / two-finger swipe → pan
- Ctrl+scroll / trackpad pinch → zoom about cursor
- f → fit to view
- r → reset to 1:1
Inspect and export args captured by the runtime args-dump feature. See docs/args-dump.md for the full capture workflow; this section only documents CLI invocation.
# List all args (auto-picks latest outputs/*/args_dump dir)
python -m simpler_setup.tools.dump_viewer
# Filter by task/stage/role
python -m simpler_setup.tools.dump_viewer --task 0x0000000200000a00 --stage before --role input
# Export the current selection to txt
python -m simpler_setup.tools.dump_viewer --task 0x0000000200000a00 --stage before --role input --export
# Export a specific arg by index (always exports)
python -m simpler_setup.tools.dump_viewer outputs/<case>_<ts>/args_dump/ --index 42The analysis tools share the same input format - the chip_swimlane_records_*.json files generated by the simpler runtime:
{
"chip_swimlane_level": 4,
"tasks": [
{
"task_id": 0,
"func_id": 0,
"core_id": 7,
"core_type": "aiv",
"ring_id": 0,
"start_time_us": 47.46,
"end_time_us": 55.9,
"duration_us": 8.44,
"dispatch_time_us": 45.94,
"finish_time_us": 60.52
},
{
"task_id": 4294967296,
"func_id": 1,
"core_id": 7,
"core_type": "aiv",
"ring_id": 1,
"start_time_us": 68.68,
"end_time_us": 70.42,
"duration_us": 1.74,
"dispatch_time_us": 68.24,
"finish_time_us": 71.2
}
]
}Dependency edges come from deps.json (dep_gen replay) at post-process time —
not from the perf JSON. See swimlane_converter --deps-json.
Top-level layout depends on chip_swimlane_level:
- All levels:
chip_swimlane_level,tasks[](per-task fields above). >= 3: alsoaicpu_scheduler_phases[](per-thread phase records: scan / complete / dispatch / idle) andcore_to_thread[](core_id → scheduler thread index).>= 4: alsoaicpu_orchestrator_phases[](per-task orchestrator phase records).
To display meaningful function names in the output, provide a kernel_config.py file:
KERNELS = [
{
"func_id": 0,
"name": "QK",
# ... other fields
},
{
"func_id": 1,
"name": "SF",
# ... other fields
},
]The tools extract the func_id to name mapping from the KERNELS list.
- A detailed timeline execution view
- To analyze task scheduling across different cores
- To see precise execution times and intervals
- Task execution statistics
- Professional performance analysis and optimization
- A structural view of task dependencies (who feeds whom)
- Fast grep-friendly inspection via the default text output
- A single-file HTML you can open offline and pan by dragging or scrolling; use Ctrl+scroll or trackpad pinch to zoom
- Optional per-task tensor rows and arg-port routing in HTML via
--show-tensor-info - A graph that survives without an associated timing run (deps.json is produced by structural replay, not by hardware profiling)
# 1. Run the test to produce both timing + structural data
pytest tests/st/... --enable-chip-swimlane --enable-dep-gen
# 2. Perfetto timeline (automatic via SceneTest)
# -> outputs/<case>_<ts>/merged_swimlane.json
# open at https://ui.perfetto.dev/
# 3. Structural dependency graph (manual, default text output)
python -m simpler_setup.tools.deps_viewer outputs/<case>_<ts>/deps.json
# -> outputs/<case>_<ts>/deps_viewer.txt
# 4. Same graph as HTML
python -m simpler_setup.tools.deps_viewer outputs/<case>_<ts>/deps.json \
--format html -o outputs/<case>_<ts>/deps_viewer.html
For batch-run hardware regression, see the dev-only script
tools/benchmark_rounds.sh.
- Make sure the test was run with the
--enable-chip-swimlaneflag - Check that the outputs/ directory exists and contains profiling data
- Check the kernel_config.py file format
- Make sure every KERNELS entry has a 'func_id' and 'name' field
- The tools accept chip_swimlane_level 1–4 (the integer captured at runtime
via
--enable-chip-swimlane <N>) - Regenerate the profiling data with a supported level
- This error means the input
chip_swimlane_records_*.jsonlacks fields required by the deep-dive analysis (typicallydispatch_time_us/finish_time_us) - The basic conversion in
swimlane_convertercan still succeed, but the deep-dive will be skipped or fail - Remediation:
- Re-run with
--enable-chip-swimlaneto produce a newoutputs/*/chip_swimlane_records.json - Re-run
swimlane_converterorsched_overhead_analysis - Verify that each task in the JSON contains
dispatch_time_usandfinish_time_us
- Re-run with
- This only affects
--format html - Install graphviz:
brew install graphviz(macOS) orapt install graphviz(Debian/Ubuntu) - Verify with
which dot; should print a path - Use a different layout engine with
--engine sfdpfor very large graphs
| File | Tool | Purpose | Format |
|---|---|---|---|
chip_swimlane_records_*.json |
Runtime | Raw timing profiling data | JSON |
merged_swimlane_*.json |
swimlane_converter | Perfetto visualization | Chrome Trace Event JSON |
deps.json |
Runtime (dep_gen replay) | Structural task dependency graph + per-edge tensor info | JSON |
deps_viewer.txt |
deps_viewer | Grep-friendly dependency graph view | Plain text |
deps_viewer.html |
deps_viewer | Pan/zoom dependency graph viewer | HTML (self-contained) |