wasm: support SIMD in the WAMR classic interpreter - #2723
zhouguangyuan0718 wants to merge 2 commits into
Conversation
There was a problem hiding this comment.
Review: classic-interpreter SIMD support
This PR adds a WAMR 2.4.5 patch that runs WebAssembly SIMD (v128) in the classic interpreter so the legacy-EH / pthreads profile can execute SIMD, plus build wiring, a .wat regression fixture, a Python boundary test, and validation docs. The port of the fast-interpreter SIMD semantics onto the classic value stack looks faithful and I found no merge-blockers. Findings below are minor/optional.
Verified correct (no change needed):
v128.storepops data then address;CHECK_MEMORY_OVERFLOWguards the rightmaddrwith correct widths (16 for full load/store, 8 for extend-loads,width/8for splat/lane/zero). Loads/stores check before dereferencing.- 4-cell (
v128= 4×32-bit) stack discipline is consistent acrossPOP_V128/PUSH_V128,DROP_128/SELECT_128, typedselect, and local/global get/set/tee. - Lane-index immediates (
extract/replace/load_lane/store_lane/shuffle) are not re-checked at runtime, which is safe only because the loader validation (WASM_OP_SIMD_PREFIXcase) is enabled under the new|| WASM_ENABLE_SIMDEguard. Keep the execution and loader guards in lock-step — if a future refactor drops|| WASM_ENABLE_SIMDEfrom the loader's SIMD validation, these handlers become OOB read/write primitives driven by untrusted bytecode. - EH changes: catch now copies from the saved payload (
frame_sp - cell_num_to_copy) avoiding source/dest overlap for a multi-cellv128; rethrow resolves the tag type viamodule->e->tags[...]with theis_import_tagbranch (correctly handling imported tags) and reads the payload after the tag word. V128/GET_V128_FROM_ADDR/PUT_V128_TO_ADDR/LOAD_V128/STORE_V128resolve from shared headers (wasm.h,wasm_runtime_common.h) that the classic TU already includes, so the patch compiles without adding them locally.
See inline comments for the minor items.
| + SIMD_CASE(SIMD_i64x2_extmul_high_i32x4_u, | ||
| + SIMD_DOUBLE_OP(simde_wasm_u64x2_extmul_high_u32x4)); | ||
| + | ||
| + /* f32x4 opertions */ |
There was a problem hiding this comment.
Typo: f32x4 opertions → operations (the i64x2/f64x2 sections nearby spell it correctly).
| + V128 v2 = POP_V128(); | ||
| + V128 v3 = POP_V128(); | ||
| + | ||
| + simde_v128_t simde_result = simde_wasm_v128_bitselect( |
There was a problem hiding this comment.
v128.bitselect operand naming is confusing although the behavior is correct. Here v1 is popped first (top of stack = the mask/condition) and passed as SIMDe's 3rd arg, while v3 is the first spec operand. Naming the condition v1 and operand-1 v3 inverts the Wasm spec's conventional names and makes this hard to audit. Consider renaming to c / v2 / v1 matching the spec so the order is self-evident. Non-blocking.
| +#define SIMD_SPLAT_OP_F64(simde_func) \ | ||
| + SIMD_SPLAT_OP(simde_func, POP_F64, float64) | ||
| + | ||
| + case SIMD_i8x16_splat: |
There was a problem hiding this comment.
SIMD_i8x16_splat is hand-expanded while every other splat uses the SIMD_SPLAT_OP_* macros defined just above. It also reuses the shared val variable (declared as tbl_elem_idx_t, not uint32) rather than a local. SIMD_CASE(SIMD_i8x16_splat, SIMD_SPLAT_OP_I32(simde_wasm_i8x16_splat)); would be consistent and avoid touching val. Non-blocking.
| + SIMD_SPLAT_OP_F32(simde_wasm_f32x4_splat)); | ||
| + SIMD_CASE(SIMD_f64x2_splat, | ||
| + SIMD_SPLAT_OP_F64(simde_wasm_f64x2_splat)); | ||
| +#define SIMD_LANE_HANDLE_UNALIGNED_ACCESS() |
There was a problem hiding this comment.
SIMD_LANE_HANDLE_UNALIGNED_ACCESS() expands to nothing and is invoked in five places (carried over from the fast interpreter). If it's an intentional hook for future unaligned handling, a one-line comment would clarify; otherwise it's dead code. Non-blocking.
| fast interpreter, with the classic value stack, original Wasm immediates, | ||
| and memory bounds checks. Preserve all four cells through locals, globals, | ||
| select/drop, calls, and exception payloads. Copy catch values from the | ||
| saved payload to avoid overlap; rethrow reads that payload after the tag. |
There was a problem hiding this comment.
Doc completeness: the rethrow hunk makes two changes — (a) reading the payload after the tag, and (b) a new tag-type resolution that handles imported tags via module->e->tags[...].is_import_tag. The patch header and wasm-wasi-validation.md describe only (a). Consider mentioning the imported-tag fix in the header so it fully reflects the hunk. (Only the inline code comment currently notes it.) Optional.
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
LLGo baseline benchmarks
Program measurements
Core language and compiler benchmarks
Timer runtime benchmarks
Compared with |
LLGo WebAssembly build benchmarks
WebAssembly output sizes
LLGo WebAssembly build measurements
Compared with |
The main-branch WASI SIMD suite fails before execution with
unknown value type: WAMR 2.4.5's classic interpreter rejectsv128, while the fast interpreter cannot execute the legacy exception handling used by W32.Add a WAMR patch that reuses its existing SIMDe-backed SIMD operations in the classic interpreter. The adapter reads original Wasm immediates, uses the classic value stack and memory checks, preserves all four vector cells through locals/globals/select/drop/calls, and extends block scanning for the rewritten opcodes. Correct overlapping catch-payload copies and the saved-payload offset used by rethrow, with vector regressions for both.
The build keeps the existing classic/EH/pthread configuration, enables SIMD explicitly, includes the new patch in the local build identity, and handles an empty generator-argument array under macOS Bash 3.2. CI's existing WAMR patch hash invalidates its cache automatically.
A raw Wasm regression now runs at the start of the existing WASI acceptance driver. It covers high vector lanes, unaligned memory, indirect calls, select/drop, block scanning, cross-call catch/rethrow, and a trapping out-of-bounds vector load.
Validation on macOS:
bash dev/build_iwasm.shbuild and cache reuse passed.2e44bf79e68cc4fe689f43eea8e7d25b6fdfbce1. Multi-module linking and multi-memory cases are outside this runner configuration. Two input files required whitespace-only normalization for WAMR's Python harness; their assertions were retained.Unblocks WASI execution for #2722. Linux and Windows validation will run in CI.