pip install jammaThat's it. macOS Accelerate BLAS handles large matrices natively.
For small datasets (<46k samples), the standard install works:
pip install jammaFor large-scale GWAS (>46k samples), install numpy-mkl first — standard numpy uses 32-bit BLAS integers which overflow at ~46k samples. Pre-built ILP64 wheels are available for Python 3.11–3.14:
pip install psutil loguru threadpoolctl click progressbar2 bed-reader
pip install numpy \
--index-url https://michael-denyer.github.io/numpy-mkl \
--force-reinstall --upgrade
pip install jamma --no-depsFrom Git (latest development version):
pip install psutil loguru threadpoolctl click progressbar2 bed-reader
pip install numpy \
--index-url https://michael-denyer.github.io/numpy-mkl \
--force-reinstall --upgrade
pip install git+https://github.com/michael-denyer/jamma.git --no-depsWhy
--no-deps? JAMMA depends onnumpy>=2.4.6, so a normal install will pull in standard numpy and overwrite the ILP64 build.--no-depsprevents this; you install the runtime dependencies manually instead.
git clone https://github.com/michael-denyer/jamma.git
cd jamma
uv sync| Platform | pip install jamma |
Notes |
|---|---|---|
| Linux x86_64 | Full support | ILP64 for >46k samples |
| ARM Mac (M1+) | Full support | Accelerate BLAS |
| ARM Linux | Full support | OpenBLAS |
| Intel Mac (macOS 13.3+) | Full support | Accelerate BLAS |
| Windows (10+) | Full support | ILP64 for >46k samples |
| Windows Server (2016+) | Full support | ILP64 for >46k samples |
JAMMA's heavy computation (eigendecomposition, matrix multiplication, REML optimization) is BLAS-bound. Intel MKL delivers the best throughput, particularly at scale. Apple Accelerate is a close second on Apple Silicon. OpenBLAS works correctly everywhere but is less tuned for these workloads.
JAMMA auto-detects the best available backend. Force a specific backend with:
# CLI flag
jamma -lmm 1 --backend numpy -bfile data/my_study -k kinship.cXX.npy
# Environment variable (overrides CLI flag)
export JAMMA_BACKEND=numpyPriority: JAMMA_BACKEND env var > --backend flag > auto-detect (C+NumPy if C extension available, else NumPy fallback).
flowchart TD
subgraph INPUT["INPUT"]
direction LR
BED[".bed/.bim/.fam"]
COV["Covariates"]
KIN_IN["Kinship (optional)"]
end
subgraph COMPUTE["COMPUTE PIPELINE"]
direction TB
subgraph PHASE1["Phase 1 — Kinship"]
GK["Kinship Accumulation<br/>DGEMM (chunked)"]
end
subgraph PHASE2["Phase 2 — Eigendecomposition"]
EIG["jlinalg.eigh<br/>DSYEVD / DSYEVR"]
end
subgraph PHASE3["Phase 3 — Association"]
direction TB
MEM{"Memory<br/>budget?"}
BATCH["Batch Runner<br/>(genotypes in RAM)"]
STREAM["Streaming Runner<br/>(two-pass disk I/O)"]
CEXT{"C extension<br/>available?"}
OMP["OpenMP + SIMD<br/>accelerated"]
PYFALL["Pure Python<br/>fallback"]
end
end
subgraph OUTPUT["OUTPUT"]
ASSOC[".assoc.txt"]
LOG[".log.txt"]
end
INPUT --> PHASE1
KIN_IN -.->|skip if pre-computed| PHASE2
PHASE1 --> PHASE2
PHASE2 --> PHASE3
MEM -->|fits| BATCH
MEM -->|large| STREAM
BATCH --> CEXT
STREAM --> CEXT
CEXT -->|yes| OMP
CEXT -->|no| PYFALL
OMP --> OUTPUT
PYFALL --> OUTPUT
style INPUT fill:#1a1a2e,stroke:#e94560,color:#eee,stroke-width:2px
style COMPUTE fill:#0f3460,stroke:#e94560,color:#eee,stroke-width:2px
style OUTPUT fill:#1a1a2e,stroke:#e94560,color:#eee,stroke-width:2px
style PHASE1 fill:#16213e,stroke:#53a8b6,color:#eee,stroke-width:2px
style PHASE2 fill:#16213e,stroke:#f5b461,color:#eee,stroke-width:2px
style PHASE3 fill:#16213e,stroke:#e94560,color:#eee,stroke-width:2px
style BED fill:#53a8b6,stroke:#3d8a96,color:#fff
style COV fill:#53a8b6,stroke:#3d8a96,color:#fff
style KIN_IN fill:#53a8b6,stroke:#3d8a96,color:#fff
style GK fill:#53a8b6,stroke:#3d8a96,color:#fff
style EIG fill:#f5b461,stroke:#d4943f,color:#1a1a2e
style MEM fill:#e94560,stroke:#c73550,color:#fff
style BATCH fill:#7b68ae,stroke:#5a4d8a,color:#fff
style STREAM fill:#7b68ae,stroke:#5a4d8a,color:#fff
style CEXT fill:#e94560,stroke:#c73550,color:#fff
style OMP fill:#2ecc71,stroke:#27ae60,color:#1a1a2e
style PYFALL fill:#95a5a6,stroke:#7f8c8d,color:#1a1a2e
style ASSOC fill:#2ecc71,stroke:#27ae60,color:#1a1a2e
style LOG fill:#2ecc71,stroke:#27ae60,color:#1a1a2e
JAMMA uses PLINK binary format (.bed, .bim, .fam files):
my_study.bed # Binary genotype data
my_study.bim # SNP information
my_study.fam # Sample information
Compute genetic relatedness matrix from genotype data:
jamma -gk 1 -bfile data/my_study -o kinship -outdir outputOptions:
-bfile PATH— PLINK binary file prefix (required)-gk MODE— Kinship type: 1 = centered, 2 = standardized-ksnps PATH— SNP list file to restrict kinship computation (one RS ID per line)-n INT— Phenotype column in .fam file (1-based, default: 1)-maf FLOAT— MAF threshold (default: 0.0, no filter for gk mode)-miss FLOAT— Missing rate threshold (default: 1.0, no filter for gk mode)--legacy-text— Write kinship files in GEMMA text format (.cXX.txt) instead of binary.npy-o PREFIX— Output file prefix-outdir DIR— Output directory
Note: Monomorphic SNPs (variance = 0) are always filtered to match GEMMA behavior.
Note: -gk 2 (standardized kinship) cannot be used with -loco mode.
Output:
output/kinship.cXX.npy— Kinship matrix (binary NumPy format, default)output/kinship.log.txt— Run log
Binary .npy is the default output format (10-100x faster I/O at scale). Use
--legacy-text for GEMMA-compatible text format (.cXX.txt):
# Default: binary .npy output (fast)
jamma -gk 1 -bfile data/my_study -o kinship -outdir output
# GEMMA-compatible text output
jamma -gk 1 -bfile data/my_study -o kinship -outdir output --legacy-textThe reader auto-detects format, so existing .cXX.txt files still work as -k input.
Using GEMMA-generated files with JAMMA:
# GEMMA produced kinship.cXX.txt — pass it directly to JAMMA
jamma -lmm 1 -bfile data/my_study -k output/kinship.cXX.txt -o assoc -outdir output
# GEMMA produced eigenvalue/eigenvector files — use -d/-u
jamma -lmm 1 -bfile data/my_study \
-d output/result.eigenD.txt -u output/result.eigenU.txt \
-o assoc -outdir outputJAMMA reads GEMMA's text formats natively (space- or tab-separated). No conversion needed.
Run univariate linear mixed model association tests:
jamma -lmm 1 -bfile data/my_study -k output/kinship.cXX.npy -o assoc -outdir outputWith covariates:
jamma -lmm 1 -bfile data/my_study -k output/kinship.cXX.npy \
-c covariates.txt -o assoc -outdir outputOptions:
-bfile PATH— PLINK binary file prefix (required)-k PATH— Kinship matrix file (required unless-locoor-d/-uare used)-lmm MODE— Test type: 1 = Wald (default), 2 = LRT, 3 = Score, 4 = All-c PATH— Covariate file (GEMMA format: whitespace-delimited, first column should be intercept)-loco— Enable leave-one-chromosome-out analysis (mutually exclusive with-k)-d PATH— Pre-computed eigenvalue file (.eigenD.npyor.eigenD.txt)-u PATH— Pre-computed eigenvector file (.eigenU.npyor.eigenU.txt)-eigen— Write eigendecomposition files (.eigenD.npy,.eigenU.npy; text with--legacy-text)-n INT|"INT INT ..."— Phenotype column(s) in .fam file (1-based, default: 1). Multiple columns can be space- or comma-separated (e.g.,-n "1 2 3"or-n "1,2,3")-snps PATH— SNP list file to restrict association testing (one RS ID per line)-ksnps PATH— SNP list file to restrict kinship computation (one RS ID per line)-hwe FLOAT— HWE p-value threshold; exclude SNPs below this value (default: 0.0, disabled)-lmin FLOAT— Minimum lambda for optimization (default: 1e-5)-lmax FLOAT— Maximum lambda for optimization (default: 1e5)-widv PATH— Individual weights file for kinship pre-transformation (one weight per line)-cat INT [INT ...]— Covariate column indices to one-hot encode as categorical (1-based)-maf FLOAT— MAF threshold (default: 0.01)-miss FLOAT— Missing rate threshold (default: 0.05)--mem-budget GB— Memory budget in GB (default: available - 10%)--no-check-memory— Disable pre-flight memory checks--legacy-text— Write kinship and eigen files in GEMMA text format instead of binary.npy--backend auto|numpy|numpy-streaming— Force compute backend (default: auto)-v/--verbose— Verbose output--version— Show version and exit
Note: Monomorphic SNPs (variance = 0) are always filtered to match GEMMA behavior.
Output:
output/assoc.assoc.txt— Association resultsoutput/assoc.log.txt— Run log
Tab-separated file. The first 7 columns are always present; stat columns depend on -lmm mode:
Common columns (all modes):
| Column | Description |
|---|---|
chr |
Chromosome |
rs |
SNP identifier |
ps |
Position |
n_miss |
Number of missing genotypes |
allele1 |
Effect allele |
allele0 |
Reference allele |
af |
Allele frequency |
Mode-specific stat columns:
| Column | -lmm 1 (Wald) |
-lmm 2 (LRT) |
-lmm 3 (Score) |
-lmm 4 (All) |
|---|---|---|---|---|
beta |
yes | — | yes | yes |
se |
yes | — | yes | yes |
logl_H1 |
yes | — | — | yes |
l_remle |
yes | — | — | yes |
l_mle |
— | yes | — | yes |
p_wald |
yes | — | — | yes |
p_lrt |
— | yes | — | yes |
p_score |
— | — | yes | yes |
Binary format (.cXX.npy, default): NumPy binary array. 10-100x faster I/O than
text at scale. Read with numpy.load("kinship.cXX.npy").
Text format (.cXX.txt, with --legacy-text): Space-separated N×N matrix where
N is the number of samples. Compatible with GEMMA format.
Format auto-detection: When you pass a .txt path (e.g., -k kinship.cXX.txt),
JAMMA checks for a .npy sibling file with the same stem. If a .npy file exists
and its modification time is at least as recent as the .txt file, the .npy is
loaded instead (much faster at scale). If the .txt file is newer — for example,
because you regenerated it with GEMMA — JAMMA ignores the stale .npy and re-parses
the text file. Passing a .npy path directly always loads the binary file.
Important: If you regenerate a .txt file externally (e.g., with GEMMA) and an
older .npy sibling exists from a previous JAMMA run, the .npy is automatically
skipped because the .txt is newer. To be safe, delete stale .npy files when
regenerating kinship or eigen files outside JAMMA.
Leave-one-chromosome-out (LOCO) analysis eliminates proximal contamination by excluding the test chromosome's SNPs from the kinship matrix. JAMMA computes per-chromosome LOCO kinship via streaming subtraction from a full kinship matrix, processing one chromosome at a time for memory efficiency.
# LOCO association (kinship computed internally per chromosome)
jamma -lmm 1 -bfile data/my_study -loco -o loco_results -outdir outputKey constraints:
-locois mutually exclusive with-k(kinship is computed internally)-locois mutually exclusive with-gk 2(standardized kinship not supported in LOCO mode)-locodoes not support multi-phenotype (-n "1 2 3"). Run each phenotype separately when using-loco-hweis not supported with-loco(HWE filtering requires a single-pass architecture)
For multi-phenotype workflows, eigendecomposition (O(n^3)) dominates runtime.
Automatic reuse (recommended): Pass multiple phenotype columns with -n. JAMMA
computes eigendecomposition once and reuses it across all phenotypes:
# All three phenotypes in one invocation — eigendecomp computed once
jamma -lmm 1 -bfile data/my_study -k kinship.cXX.npy \
-n "1 2 3" -o results -outdir outputManual reuse: Save eigendecomposition files and reload them in subsequent runs:
# First phenotype: compute kinship + eigen, save both
jamma -lmm 1 -bfile data/my_study -k kinship.cXX.npy \
-eigen -n 1 -o pheno1 -outdir output
# Second phenotype: reuse eigendecomposition (skips kinship + eigen entirely)
jamma -lmm 1 -bfile data/my_study \
-d output/pheno1.eigenD.npy -u output/pheno1.eigenU.npy \
-n 2 -o pheno2 -outdir outputOutput files when -eigen is used:
output/pheno1.eigenD.npy— Eigenvalues (binary NumPy format)output/pheno1.eigenU.npy— Eigenvectors (binary NumPy format)
Use --legacy-text to write GEMMA-compatible text files (.eigenD.txt, .eigenU.txt)
instead. When --legacy-text is used with -eigen, JAMMA also writes .npy sidecar
files alongside the text files for faster subsequent reads.
The reader auto-detects format using the same sibling rule as kinship: passing a .txt
path checks for a newer .npy sibling first. See Kinship Matrix for
details on the auto-detection logic.
Note: --legacy-text only affects kinship and eigen file writes. It has no effect
on association output (.assoc.txt is always text format).
Restrict which SNPs are used for kinship computation and/or association testing:
# Restrict association to specific SNPs
jamma -lmm 1 -bfile data/my_study -k kinship.cXX.npy \
-snps snp_list.txt -o filtered -outdir output
# Restrict kinship computation to specific SNPs
jamma -gk 1 -bfile data/my_study -ksnps kinship_snps.txt \
-o kinship -outdir output
# HWE quality control
jamma -lmm 1 -bfile data/my_study -k kinship.cXX.npy \
-hwe 0.001 -o qc -outdir outputSNP list file format: One SNP RS ID per line (first whitespace-delimited token used).
HWE filtering: JAMMA uses a chi-squared goodness-of-fit test (df=1) via pure NumPy.
SNPs with p-value below the threshold are excluded from association testing.
HWE filtering is supported on the streaming backend (numpy-streaming) only; it is
not available on the batch backend.
See GEMMA_DIVERGENCES.md for differences from GEMMA's
Wigginton exact test.
For .fam files with multiple phenotype columns, select which to use:
# Single phenotype (default)
jamma -lmm 1 -bfile data/my_study -k kinship.cXX.npy -n 1
# Multiple phenotypes (space-separated)
jamma -lmm 1 -bfile data/my_study -k kinship.cXX.npy -n "1 2 3"
# Multiple phenotypes (comma-separated)
jamma -lmm 1 -bfile data/my_study -k kinship.cXX.npy -n "1,2,3"The -n flag uses 1-based indexing matching GEMMA: -n 1 selects column 6
(standard phenotype), -n 2 selects column 7, etc.
Multi-phenotype mode computes eigendecomposition once and reuses it across all
phenotypes. Each phenotype produces a separate output file with .phenoN. suffix
(e.g., result.pheno1.assoc.txt, result.pheno2.assoc.txt). Single-phenotype runs
produce output without the suffix (e.g., result.assoc.txt).
Samples with missing values (NA/-9) in any selected phenotype column are excluded from the analysis. The sample mask is the intersection across all phenotype columns, ensuring consistent results.
Early sample filtering: When samples are excluded due to missing phenotype or covariate data, JAMMA filters them before kinship computation (not after). The kinship matrix is accumulated at (n_valid x n_valid) size directly, avoiding a full (n_samples x n_samples) allocation. This reduces peak memory proportionally to the number of excluded samples.
Note: LOCO mode (-loco) does not support multi-phenotype. Run each phenotype
separately when using -loco.
The simplest way to run a complete GWAS from Python:
from jamma import gwas
# With pre-computed kinship
result = gwas("data/my_study", kinship_file="data/kinship.cXX.txt")
print(f"Tested {result.n_snps_tested} SNPs in {result.timing['total_s']:.1f}s")
# Compute kinship from scratch, save it for reuse
result = gwas("data/my_study", save_kinship=True, output_dir="output")
# With covariates
result = gwas(
"data/my_study",
kinship_file="k.txt",
covariate_file="covars.txt",
lmm_mode=2, # LRT test
)
# LOCO analysis (leave-one-chromosome-out)
result = gwas("data/my_study", loco=True)
# Multi-phenotype with eigendecomp reuse (Python API)
result = gwas("data/my_study", write_eigen=True, phenotype_column=1)
result = gwas(
"data/my_study",
eigenvalue_file="output/result.eigenD.npy",
eigenvector_file="output/result.eigenU.npy",
phenotype_column=2,
)
# Or use the CLI for automatic multi-phenotype: jamma -lmm 1 ... -n "1 2 3"
# SNP filtering and HWE QC
result = gwas(
"data/my_study",
kinship_file="k.txt",
snps_file="snps.txt",
hwe=0.001,
)gwas() handles the full pipeline: load data, compute or load kinship,
eigendecompose, run LMM association, and write results. Returns a GWASResult
with timing breakdown and summary stats. Access result.pve_estimate for the PVE
(proportion of variance explained) estimate and result.pve_se for the
standard error of PVE computed via the delta method from the REML second
derivative at the null model optimum.
For more control, use the component functions directly:
from jamma.io import load_plink_binary
from jamma.kinship import compute_centered_kinship
# Load genotypes
data = load_plink_binary("data/my_study")
# Compute kinship
K = compute_centered_kinship(data.genotypes)from jamma.lmm import run_lmm_association_numpy
from jamma.lmm.eigen import eigendecompose_kinship
from jamma.lmm.schema import LmmConfig
eigenvalues, eigenvectors = eigendecompose_kinship(K)
# NumPy runner (loads full genotype matrix)
run_result = run_lmm_association_numpy(
genotypes=data.genotypes,
phenotypes=phenotypes,
kinship=None, # Not needed when eigenvalues/eigenvectors provided
snp_info=snp_info, # list of dicts with chr, rs, pos, a1, a0
eigenvalues=eigenvalues,
eigenvectors=eigenvectors,
config=LmmConfig(lmm_mode=1), # 1=Wald, 2=LRT, 3=Score, 4=All
)
results = run_result.associations # list[AssocResult]
pve = run_result.pve # heritability estimate
pve_se = run_result.pve_se # SE of PVE via delta method (None if flat likelihood)The NumPy backend supports Wald, LRT, Score, all-tests modes, and LOCO. HWE filtering (-hwe) is supported on the streaming backend only (numpy-streaming).
JAMMA's LMM requires eigendecomposition of the N×N kinship matrix. The default numpy stack uses LP64 BLAS (32-bit integers), which overflows at ~46k samples (46k × 46k = 2.1 billion elements > int32 max).
flowchart TD
subgraph DETECT["BLAS DETECTION (jlinalg)"]
direction TB
PROBE["Probe process BLAS symbols<br/>dlsym / numpy scan"]
PROBE --> ILP64{"ILP64<br/>backend?"}
end
subgraph DISPATCH["EIGENDECOMP DISPATCH"]
direction TB
ILP64 -->|MKL-ILP64 / Accelerate-ILP64| VENDOR["Vendor LAPACK<br/>(DSYEVD / DSYEVR)"]
ILP64 -->|LP64 only or none| NPFALL["NumPy fallback<br/>(np.linalg.eigh)"]
VENDOR --> DSYEVD{"DSYEVD<br/>fits in RAM?"}
DSYEVD -->|yes| FAST["DSYEVD in-place<br/>O(N²) workspace — fast"]
DSYEVD -->|no| DSYEVR["DSYEVR<br/>O(N) workspace — slower"]
end
subgraph LIMITS["SAMPLE LIMITS"]
direction LR
L1["LP64: ~46k max<br/>(int32 overflow)"]
L2["ILP64: 200k+<br/>(memory-bound)"]
end
FAST --> LIMITS
DSYEVR --> LIMITS
NPFALL --> LIMITS
style DETECT fill:#1a1a2e,stroke:#f5b461,color:#eee,stroke-width:2px
style DISPATCH fill:#0f3460,stroke:#e94560,color:#eee,stroke-width:2px
style LIMITS fill:#1a1a2e,stroke:#53a8b6,color:#eee,stroke-width:2px
style PROBE fill:#f5b461,stroke:#d4943f,color:#1a1a2e
style ILP64 fill:#e94560,stroke:#c73550,color:#fff
style VENDOR fill:#2ecc71,stroke:#27ae60,color:#1a1a2e
style NPFALL fill:#95a5a6,stroke:#7f8c8d,color:#1a1a2e
style DSYEVD fill:#e94560,stroke:#c73550,color:#fff
style FAST fill:#2ecc71,stroke:#27ae60,color:#1a1a2e
style DSYEVR fill:#f5b461,stroke:#d4943f,color:#1a1a2e
style L1 fill:#e74c3c,stroke:#c0392b,color:#fff
style L2 fill:#2ecc71,stroke:#27ae60,color:#1a1a2e
Install numpy-mkl using the commands in Linux / Windows above. Pre-built ILP64 wheels are available for numpy 2.4.6 (Python 3.11–3.14) and 2.5.1 (Python 3.12–3.14), Linux and Windows.
Note: scipy does not support ILP64 — it hardcodes
ilp64=Falseinget_lapack_funcs()(scipy#23351). JAMMA does not use scipy at runtime (it is a dev-only dependency for tests). JAMMA usesjlinalg.eighwhich dispatches to vendor DSYEVD/DSYEVR via the jlinalg C layer, correctly using ILP64 when an ILP64 BLAS backend is available.
Verify ILP64 is active:
import numpy as np
cfg = np.show_config(mode="dicts")
blas = cfg["Build Dependencies"]["blas"]
print(f"BLAS: {blas['name']}") # Should show: mkl
print(f"Symbol suffix: {blas.get('symbol suffix', 'none')}") # Should show: _64Testing the ILP64 build:
# Run JAMMA's validation suite to confirm equivalence
uv run pytest tests/test_kinship_validation.py tests/test_validation.py \
tests/test_validation_assoc.py -v
# Quick eigendecomposition sanity check
python -c "
import numpy as np
n = 50000 # Exceeds LP64 limit
K = np.random.randn(n, 100) @ np.random.randn(100, n)
K = (K + K.T) / 2
vals, vecs = np.linalg.eigh(K)
print(f'Eigendecomposition of {n}x{n} matrix: OK')
print(f'Top eigenvalue: {vals[-1]:.2f}')
"MKL is distributed under the Intel Simplified Software License (ISSL), which permits free redistribution with no royalty fees. However, the ISSL is not an open source license — it restricts reverse engineering and decompilation, and is not GPL-compatible.
This does not affect JAMMA itself (GPL-3.0). JAMMA calls numpy APIs (BSD licensed) and has no direct dependency on MKL. Users who install MKL-backed numpy wheels do so as a separate, optional runtime choice. Users requiring a pure GPL/FOSS stack can use standard numpy with OpenBLAS (the default), which works for datasets up to ~46k samples.
If MKL ILP64 is not available:
- GPU eigendecomposition: cuSOLVER on NVIDIA GPUs uses different integer interfaces
- Approximate methods: Randomized SVD or truncated eigendecomposition
- Sample subsetting: Use ~40k representative samples for kinship computation
JAMMA's current performance optimizations target Intel x86_64 Linux — the typical Databricks / HPC environment for large-scale GWAS:
- BLAS/LAPACK: Tuned for Intel MKL (shipped via
numpy-mklwheels). OpenBLAS works but is slower and segfaults above ~50k samples. - C extension: The OpenMP-parallelized
_lmm_accelextension provides a 5-7x speedup over pure NumPy for the LMM compute phase on mouse_hs1940. The margin grows with covariate count, because the Pab table recursion gets more expensive. - ARM / Apple Silicon: Runs correctly via Accelerate BLAS. Thread control
(
blas_threads()) is not available on Accelerate — Apple provides no public API andVECLIB_MAXIMUM_THREADSis only read at library load time. JAMMA detects this automatically and halves OpenMP threads in the C extension to avoid oversubscription with Accelerate's uncontrollable thread pool.JAMMA_BLAS_THREADShas no effect on Accelerate.
- Use the C extension for best performance — it is auto-compiled on first use and provides OpenMP-parallelized SNP processing
- Streaming mode (
numpy-streaming) works on all platforms and supports arbitrarily large datasets with full pipeline support (LOCO, multi-phenotype, HWE filtering) - Batch processing: JAMMA automatically batches kinship computation
- Memory: For very large datasets, the streaming backend (
numpy-streaming) is auto-selected when batch mode won't fit in memory
JAMMA collects local-only benchmark telemetry to help you track performance across runs. No data is ever transmitted to any server.
Each -lmm run appends one JSON line to ~/.jamma/benchmarks.jsonl with these fields:
| Field | Description |
|---|---|
timestamp |
ISO 8601 UTC timestamp |
jamma_version |
JAMMA version string |
n_samples |
Number of samples analyzed |
n_snps |
Number of SNPs tested |
backend |
Runner name (e.g. numpy, numpy-streaming) |
n_cvt |
Number of covariates (optional) |
lmm_mode |
LMM test mode: 1=Wald, 2=LRT, 3=Score, 4=All (optional) |
loco |
Whether LOCO mode was used (optional) |
kinship_s |
Kinship computation time in seconds (optional) |
lmm_s |
LMM computation time in seconds (optional) |
total_s |
Total pipeline time in seconds (optional) |
rotation_s |
Rotation time in seconds (optional) |
Additional optional fields may be present depending on configuration: n_chunks,
eigendecomp_s, peak_memory_gb, cpu_model, blas_backend, blas_threads,
total_ram_gb, numpy_version, platform. See jamma.core.telemetry.BenchmarkRecord
for the full schema.
No genotype data, phenotype data, or file paths are recorded.
Telemetry is written to ~/.jamma/benchmarks.jsonl as newline-delimited JSON.
Each line is independently parseable. The file is append-only; JAMMA never reads
it back. You can safely delete it at any time.
Disable telemetry with either:
-
CLI flag (per-run):
jamma --no-telemetry -lmm 1 -bfile data/study -k kinship.cXX.txt
-
Environment variable (persistent):
# JAMMA-specific export JAMMA_NO_TELEMETRY=1 # Or use the universal convention (also disables telemetry in other CLI tools) export DO_NOT_TRACK=1
JAMMA_NO_TELEMETRY disables telemetry for any non-empty value.
DO_NOT_TRACK follows the Do Not Track convention:
only DO_NOT_TRACK=1 opts out; DO_NOT_TRACK=0 explicitly opts in.
Kinship-only mode (-gk) does not emit telemetry regardless of these settings.
| Variable | Default | Description |
|---|---|---|
JAMMA_BACKEND |
auto-detect | Force backend: auto, numpy, or numpy-streaming. Auto-detect prefers C+NumPy, then NumPy fallback. |
JAMMA_BLAS_THREADS |
physical_cores |
Thread count for NumPy BLAS operations (eigendecomp, matmul). Controls MKL/OpenBLAS via threadpoolctl, not OpenMP. Linux only — has no effect on macOS Accelerate. |
JAMMA_LOCO_WORKERS |
1 |
Parallel chromosome workers in LOCO mode. Each worker holds a full K_loco matrix (n_samples^2 x 8 bytes), so increase with caution. |
JAMMA_NO_TELEMETRY |
(unset) | Set to any non-empty value to disable benchmark telemetry. See Telemetry. |
DO_NOT_TRACK |
(unset) | Universal telemetry opt-out convention. Set to 1 to disable JAMMA telemetry. See Telemetry. |
# Example: 4 BLAS threads, 2 LOCO workers
export JAMMA_BLAS_THREADS=4
export JAMMA_LOCO_WORKERS=2
jamma -lmm 1 -bfile data/my_study -loco -o outputNote: JAMMA_BLAS_THREADS scopes thread control to BLAS libraries (MKL, OpenBLAS)
and does not affect OpenMP (libgomp/libomp). It has no
effect on macOS Accelerate (which provides no thread-count API). If you have C
extensions compiled with -fopenmp, use OMP_NUM_THREADS separately.
JAMMA results match GEMMA within validated tolerances:
- Kinship matrices: < 1e-8 relative difference
- P-values (Wald/Score): < 1e-4 relative difference
- P-values (LRT): < 5e-3 relative difference (MLE subtraction amplification)
- Beta coefficients: < 1e-2 relative difference (lambda propagation)
- Log-likelihood (REML): < 1e-6 relative difference
- Log-likelihood (MLE/logl_H1): < 5e-3 relative difference on real data
- Significance calls: 100% agreement at all thresholds
- Effect directions and SNP rankings: identical
GEMMA uses Brent's method for lambda optimization; JAMMA uses grid search followed by golden section refinement. Both converge to within 1e-5 of the true optimum for strong-signal SNPs. However, weak-signal SNPs — where the optimization landscape is flat and lambda converges near the lower bound (1e-5) — can produce slightly different optima between the two methods. This propagates to per-SNP MLE log-likelihood (logl_H1) with up to ~0.14% relative difference on real datasets (observed on mouse_hs1940 at SNP index 596 of 10768). The quantities that drive scientific conclusions (p-values, effect directions, significance rankings) are unaffected.
See GEMMA_EQUIVALENCE.md for empirical validation and formal error propagation analysis.
JAMMA includes pre-flight memory checks that fail fast before OOM instead of crashing silently.
flowchart TD
subgraph PREFLIGHT["PRE-FLIGHT CHECK"]
direction TB
EST["Estimate peak memory<br/>(eigendecomp dominates)"]
BLAS["Detect BLAS backend<br/>+ ILP64 status"]
EST --> GATE{"Peak + 10%<br/>≤ available?"}
BLAS --> GATE
end
subgraph EIGEN_MEM["EIGENDECOMP MEMORY"]
direction TB
KM["K matrix<br/>8N² bytes"]
UM["U matrix<br/>8N² bytes"]
WK["DSYEVD workspace<br/>~8N² bytes"]
end
subgraph RUNTIME["RUNTIME CONTROLS"]
direction TB
INPLACE["In-place DSYEVD<br/>(saves 1× N²)"]
FALLBACK["DSYEVR fallback<br/>(O(N) workspace)"]
CHUNK["Streaming chunks<br/>O(N × chunk_size)"]
FLUSH["Per-chunk disk flush<br/>(no list accumulation)"]
end
GATE -->|sufficient| EIGEN_MEM
GATE -->|insufficient| FAIL["MemoryError<br/>(fail fast with suggestions)"]
EIGEN_MEM --> RUNTIME
style PREFLIGHT fill:#1a1a2e,stroke:#f5b461,color:#eee,stroke-width:2px
style EIGEN_MEM fill:#0f3460,stroke:#e94560,color:#eee,stroke-width:2px
style RUNTIME fill:#16213e,stroke:#2ecc71,color:#eee,stroke-width:2px
style EST fill:#f5b461,stroke:#d4943f,color:#1a1a2e
style BLAS fill:#f5b461,stroke:#d4943f,color:#1a1a2e
style GATE fill:#e94560,stroke:#c73550,color:#fff
style FAIL fill:#e74c3c,stroke:#c0392b,color:#fff
style KM fill:#e94560,stroke:#c73550,color:#fff
style UM fill:#e94560,stroke:#c73550,color:#fff
style WK fill:#e94560,stroke:#c73550,color:#fff
style INPLACE fill:#2ecc71,stroke:#27ae60,color:#1a1a2e
style FALLBACK fill:#2ecc71,stroke:#27ae60,color:#1a1a2e
style CHUNK fill:#2ecc71,stroke:#27ae60,color:#1a1a2e
style FLUSH fill:#2ecc71,stroke:#27ae60,color:#1a1a2e
By default, JAMMA estimates memory requirements before large allocations:
# Check memory estimate without running
jamma -lmm 1 -bfile data/large_study -k kinship.cXX.txt --mem-budget 64If the estimate exceeds available memory, you'll get a clear error:
MemoryError: LMM requires ~128.5GB but only 64.0GB available
Breakdown: kinship=74.5GB, eigendecomp=37.0GB, association=17.0GB
# Set explicit memory budget (GB)
jamma -lmm 1 ... --mem-budget 128
# Disable checks (use at your own risk)
jamma -lmm 1 ... --no-check-memoryfrom jamma.core.memory import estimate_workflow_memory, estimate_lmm_memory
# Full pipeline estimate (before starting anything)
full = estimate_workflow_memory(n_samples=200_000, n_snps=95_000)
print(f"Full pipeline peak: {full.total_gb:.1f}GB")
print(f"Eigendecomp workspace: {full.eigendecomp_workspace_gb:.1f}GB")
print(f"Available: {full.available_gb:.1f}GB")
print(f"Sufficient: {full.sufficient}")
# LMM-only estimate (after eigendecomp is done, kinship freed)
lmm = estimate_lmm_memory(n_samples=200_000, n_snps=95_000)
print(f"LMM phase: {lmm.total_gb:.1f}GB")JAMMA runs a pre-flight memory check before kinship and eigendecomposition. The check estimates peak memory (dominated by eigendecomposition: K + U + workspace) and applies a 10% safety margin based on empirical benchmarks. When vendor DSYEVR is available (via jlinalg BLAS dispatch), JAMMA automatically falls back from DSYEVD (faster, O(N^2) workspace) to DSYEVR (slower, O(N) workspace) when DSYEVD won't fit — this can increase the maximum sample count by ~40% for a given machine size.
Approximate sample limits by machine size:
| Machine RAM | ~Available | Max samples |
|---|---|---|
| 512GB | 490GB | ~142k |
| 256GB | 240GB | ~100k |
| 128GB | 120GB | ~70k |
| 64GB | 58GB | ~49k |
| 32GB | 28GB | ~34k |
| 16GB | 14GB | ~24k |
These limits assume the streaming pipeline (CLI default). Actual limits depend on available memory at runtime — other processes reduce headroom.
In-place imputation: JAMMA imputes missing genotypes in-place (no extra copy), so imputation does not add a full genotype matrix to peak memory.
If the pre-flight check rejects your run:
-
Free memory from other processes or previous runs
-
Use
--no-check-memoryto bypass the check (at your own risk):jamma -gk 1 --no-check-memory -bfile data/study jamma -lmm 1 --no-check-memory -bfile data/study -k kinship.txt
Small numerical differences (< 1e-5) are expected due to different optimization algorithms. Scientific conclusions (significance, rankings) should be identical. If you see larger differences, please open an issue.