Skip to content

Repository files navigation

MemoryLayouts.jl 🧠⚡

Optimize your memory layout for maximum cache efficiency.

Stable Dev Build Status Aqua JET


🚀 The Problem vs. The Solution

Standard collections in Julia (Dicts, Arrays of Arrays, structs) often scatter data across memory, causing frequent cache misses. MemoryLayouts.jl packs this data into contiguous blocks.

The advantage of contiguity is that it reduces cache misses and should be expected to improve performance.

🔮 How it works

Function Description Analogy
layout( x ) Aligns immediate fields of x Like copy( x ) but packed
deeplayout( x ) Recursively aligns nested structures Like deepcopy( x ) but packed
layout!( x ) In-place alignment (e.g. for Dicts) Like layout( x ) but in-place
withlayout( f ) Runs f with a scoped layout handle Automatic memory management
layoutstats( x ) Dry run statistics for layout( x )
deeplayoutstats( x ) Dry run statistics for deeplayout( x )
visualizelayout( x ) Visualizes memory layout using terminal graphics
deepvisualizelayout( x ) Recursively visualizes memory layout

🚀 SIMD Optimization

Both functions accept an optional alignment keyword argument (default 1). This allows aligning data to specific byte boundaries (e.g., 32 or 64 bytes), which is relevant for maximizing performance with SIMD instructions (AVX2, AVX-512).

aligneddata = layout( data; alignment = 64 )

🔒 Memory Management

By default, nothing needs to be managed: the contiguous block behind the arrays returned by layout and deeplayout is kept alive by the garbage collector for as long as any of those arrays is alive, and freed afterwards.

For deterministic release of a large block, use withlayout. All calls to layout, deeplayout, and layout! inside the block use a temporary handle that is released when the block exits:

result = withlayout() do
    x = deeplayout( a )
    y = deeplayout( b )
    compute( x, y )
end

Warning

Arrays created inside a withlayout block point into freed memory once the block exits. Do not let them escape the block; return plain results instead.


⚡ Performance Example

In scientific computing, memory locality is everything.

Benchmark Result: original: 159.177 μs layout: 111.251 μs (🚀 43% Faster)

Click to see the benchmark code
using MemoryLayouts, BenchmarkTools, StyledStrings

function original( A = 10_000, L = 100, S = 5000)
    x = Vector{Vector{Float64}}( undef, A )
    s = Vector{Vector{Float64}}( undef, A )
    for i ∈ 1:A
        x[i] = rand( L )
        s[i] = rand( S )
    end
    return x
end

function computeme( X )
    Σ = 0.0
    for x ∈ X 
        Σ += x[5] 
    end
    return Σ
end

print( styled"{red:original}: " ); @btime computeme( X ) setup=(X = original())
print( styled"{green:layout}: " ); @btime computeme( X ) setup=(X = layout( original()))
  • The above example is included as example1.jl in the examples folder.

📊 Dry Run / Statistics / Visualization

You can inspect the potential improvements in memory contiguity without performing the actual allocation using layoutstats and deeplayoutstats. You can also visualize the memory layout using visualizelayout and deepvisualizelayout.

julia> using MemoryLayouts

julia> data = [rand(10) for _ in 1:5];

julia> layoutstats(data)
LayoutStats(packed=400, blocks=5, span=2304, reduction=1904 (82.6%))

julia> visualizelayout(data)
Memory Layout Visualization
  Span: 2 kb
  Min : 0x7f0a1c000f60 (leftmost point)
  Max : 0x7f0a1c001860 (rightmost point)
  Scale: 28 b / char
█░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░█

The output indicates:

  • packed: The total size (in bytes) of the data if packed.
  • blocks: The number of individual arrays identified.
  • span: The current distance between the minimum and maximum memory addresses of the data.
  • reduction: The potential reduction in memory span.

⚠️ Usage Notes

Important

  1. Use the statistics to evaluate whether packing is worth it.
  2. Only plain Arrays (and the wrappers listed below) with isbits element types are moved; ranges, views, BitArrays, etc. are passed through untouched.
  3. Shared arrays stay shared, as with deepcopy.
  4. There are some issues to pay attention to: please read the documentation carefully.
  5. This is the first version of this package: comments are welcome.

🔌 Compatibility & Extensions

MemoryLayouts is further compatible with:

(Assumes these packages are loaded by the user.) Arrays of InlineStrings are packed like arrays of numbers; no extension is needed.


📚 Related Packages

There are several other Julia packages that address memory layout and array storage, though with a different focus:

  • RaggedArrays.jl: Provides contiguous memory storage specifically for arrays of arrays (jagged/ragged arrays).
  • BlockArrays.jl: Focuses on partitioning arrays into blocks. The BlockedArray type stores the full array contiguously with a block structure overlaid.
  • Strided.jl: Specialized for efficient strided array views and operations.
  • UnsafeArrays.jl: Provides stack-allocated pointer-based array views.
  • Buffers.jl: Manages buffer allocation/deallocation for multidimensional arrays.

MemoryLayouts.jl differs by focusing specifically on physically aligning multiple independent arrays (which may be fields in a struct) into a single contiguous memory block to optimize cache usage, while using unsafe_wrap to present them as standard Julia arrays.


🤖 For AI Agents

This package contains a dedicated guide for AI agents to help them understand and assist with MemoryLayouts.jl. See Agents.md for detailed instructions on architecture, safety, and best practices.

About

aligns memory within structs, dicts, and arrays

Resources

Contributing

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages