Standard collections in Julia (Dicts, Arrays of Arrays, structs) often scatter data across memory, causing frequent cache misses. MemoryLayouts.jl packs this data into contiguous blocks.
The advantage of contiguity is that it reduces cache misses and should be expected to improve performance.
| Function | Description | Analogy |
|---|---|---|
layout( x ) |
Aligns immediate fields of x |
Like copy( x ) but packed |
deeplayout( x ) |
Recursively aligns nested structures | Like deepcopy( x ) but packed |
layout!( x ) |
In-place alignment (e.g. for Dicts) | Like layout( x ) but in-place |
withlayout( f ) |
Runs f with a scoped layout handle |
Automatic memory management |
layoutstats( x ) |
Dry run statistics for layout( x ) |
|
deeplayoutstats( x ) |
Dry run statistics for deeplayout( x ) |
|
visualizelayout( x ) |
Visualizes memory layout using terminal graphics | |
deepvisualizelayout( x ) |
Recursively visualizes memory layout |
Both functions accept an optional alignment keyword argument (default 1).
This allows aligning data to specific byte boundaries (e.g., 32 or 64 bytes), which is relevant for maximizing performance with SIMD instructions (AVX2, AVX-512).
aligneddata = layout( data; alignment = 64 )By default, nothing needs to be managed: the contiguous block behind the arrays returned by layout and deeplayout is kept alive by the garbage collector for as long as any of those arrays is alive, and freed afterwards.
For deterministic release of a large block, use withlayout. All calls to layout, deeplayout, and layout! inside the block use a temporary handle that is released when the block exits:
result = withlayout() do
x = deeplayout( a )
y = deeplayout( b )
compute( x, y )
endWarning
Arrays created inside a withlayout block point into freed memory once the block exits. Do not let them escape the block; return plain results instead.
In scientific computing, memory locality is everything.
Benchmark Result:
original: 159.177 μslayout: 111.251 μs (🚀 43% Faster)
Click to see the benchmark code
using MemoryLayouts, BenchmarkTools, StyledStrings
function original( A = 10_000, L = 100, S = 5000)
x = Vector{Vector{Float64}}( undef, A )
s = Vector{Vector{Float64}}( undef, A )
for i ∈ 1:A
x[i] = rand( L )
s[i] = rand( S )
end
return x
end
function computeme( X )
Σ = 0.0
for x ∈ X
Σ += x[5]
end
return Σ
end
print( styled"{red:original}: " ); @btime computeme( X ) setup=(X = original())
print( styled"{green:layout}: " ); @btime computeme( X ) setup=(X = layout( original()))- The above example is included as
example1.jlin theexamplesfolder.
You can inspect the potential improvements in memory contiguity without performing the actual allocation using layoutstats and deeplayoutstats. You can also visualize the memory layout using visualizelayout and deepvisualizelayout.
julia> using MemoryLayouts
julia> data = [rand(10) for _ in 1:5];
julia> layoutstats(data)
LayoutStats(packed=400, blocks=5, span=2304, reduction=1904 (82.6%))
julia> visualizelayout(data)
Memory Layout Visualization
Span: 2 kb
Min : 0x7f0a1c000f60 (leftmost point)
Max : 0x7f0a1c001860 (rightmost point)
Scale: 28 b / char
█░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░█The output indicates:
- packed: The total size (in bytes) of the data if packed.
- blocks: The number of individual arrays identified.
- span: The current distance between the minimum and maximum memory addresses of the data.
- reduction: The potential reduction in memory span.
Important
- Use the statistics to evaluate whether packing is worth it.
- Only plain
Arrays (and the wrappers listed below) withisbitselement types are moved; ranges, views,BitArrays, etc. are passed through untouched. - Shared arrays stay shared, as with
deepcopy. - There are some issues to pay attention to: please read the documentation carefully.
- This is the first version of this package: comments are welcome.
MemoryLayouts is further compatible with:
- 🔑
AxisKeys - 🏷️
NamedDims.jl - 📏
OffsetArrays
(Assumes these packages are loaded by the user.) Arrays of InlineStrings are packed like arrays of numbers; no extension is needed.
There are several other Julia packages that address memory layout and array storage, though with a different focus:
- RaggedArrays.jl: Provides contiguous memory storage specifically for arrays of arrays (jagged/ragged arrays).
- BlockArrays.jl: Focuses on partitioning arrays into blocks. The
BlockedArraytype stores the full array contiguously with a block structure overlaid. - Strided.jl: Specialized for efficient strided array views and operations.
- UnsafeArrays.jl: Provides stack-allocated pointer-based array views.
- Buffers.jl: Manages buffer allocation/deallocation for multidimensional arrays.
MemoryLayouts.jl differs by focusing specifically on physically aligning multiple independent arrays (which may be fields in a struct) into a single contiguous memory block to optimize cache usage, while using unsafe_wrap to present them as standard Julia arrays.
This package contains a dedicated guide for AI agents to help them understand and assist with MemoryLayouts.jl.
See Agents.md for detailed instructions on architecture, safety, and best practices.