Skip to content

Latest commit

 

History

89 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

Llama Short Manual

(c) 2026 – Roberto A. Foglietta <roberto.foglietta@gmail.com>, CC BY-NC-ND 4.0

  •  Click on the button to know how to  Sponsor me  this project and get in touch with me.

llama.cpp image

Revision 85

A short manual to run AI locally on your PC/laptop with decent performance despite minimal hardware requisites, focusing on optimizing memory management and presenting how to choose the best model to fit specific hardware limits. Backed by real-world benchmarks and configuration tests, this guide quickly evolves into a bottleneck root-cause investigation paper that exposes the critical roles of CPU thermal design and constraints over misleading burst benchmarks.


Minimal HW requirements

For a target model between 2B and 8B parameters:

  • CPU: Intel i5-8265 or Ryzen 2500U
  • RAM: not less than 16GB w/ Ubuntu
  • DDR4: clock 2400Mhz in dual-channel
  • GGUF: file size below 6GB (mem 8GB)

Target quantisation between Q4_0 (fast) and Q5_K_XL (fine).


Memory management plan

The main idea is to split the 16GB RAM in two halves, one for the OS/application usage and the other half for running the AI model.

In this scenario, few point to keep in consideration:

  • Using a HTTP server doesn't eat a lot of RAM but the browser.
  • The bigger the context size, the smaller the model should be.
  • On an Ubuntu essential configuration 8GB can be given to running the AI.

The Q4 and the context are those points on which we can save RAM when the AI local model is supposed to always run in concurrence with a decent desktop activity.

sudo prlimit --pid=$$ --memlock=$((6<<30)):$((8<<30))

For an optimised llama functioning the line above raises the current shell memory allocation limits respectively at 6GB (soft) and 8GB (hard). To make this change permanent for the user without the need to do sudo each time, the new limits should be set into /etc/security/limits.conf.


Vulkan SDK (optional)

Unless you have a serious videocard but an integrated one, the GPU will not effectively support your AI workload but slow down it.

However, if you do not try, you do not know. So, here below how to install the Vulkan SDK last version available for Ubuntu 22.04 (jammy) and 24.04 (noble):

distname=$(lsb_release -cs)
wget -qO- https://packages.lunarg.com/lunarg-signing-key-pub.asc |
  sudo tee /etc/apt/trusted.gpg.d/lunarg.asc
echo "deb https://packages.lunarg.com/vulkandistname main" |
  sudo tee /etc/apt/sources.list.d/lunarg-vulkan.list
sudo apt update
sudo apt install vulkan-sdk

This SDK already contains its own glslc package which conflicts with Ubuntu's one. In case of failure, you can choose for -DGGML_VULKAN=OFF in building or sudo apt purge glslc to cleanly install the Vulkan SDK's version.

Note

During Ubuntu installation, the hardware is probed and the libvulkan1 could be installed because it is functional to the graphical engine whether it is Xorg or Wayland. In such a case, you will compile Llama against LunarG's Vulkan loader libvulkan.so.1 but the code will dynamically link against the one from your system libraries. Unless you decide to proceed for a statically linked compilation (code duplication, cache underperformance, huge footprint) or re-install libvulkan1 and related packages from the LunagG's repository which is possible only for supported Ubuntu versions. However, managing these deployment intricacies is beyond the scope of this document.


Native llama quick build

Compile your own llama binary set isn't mandatory for the most common OS configurations:

  • llama.cpp releases by ggml.org for Ubuntu, MacOS and Windows,
  • supporting: CPU-only, Vulkan, ROCm, OpenVINO, CUDA (*) and HIP
  • (*) CUDA support available only for Windows x64 otherwise llamafile

To compile your llama native build w/ Vulkan support on Ubuntu (22.04 LTS, in my case):

sudo apt update
sudo apt install git build-essential cmake ccache \
  mesa-vulkan-drivers vulkan-tools libvulkan-dev libvulkan1 \
  glslang-tools glslang-dev glslang-tools libopenblas-dev \
  glslc pciutils libcurl4-openssl-dev
# optional and useless with -DLLAMA_SERVER_LLAMAUI=OFF
# sudo apt install nodejs npm

# copy 1:1 of the original project, currently at #e3471b3e7
git clone https://github.com/robang74/llama.cpp
cd llama.cpp

# if your GPU does not perform better than '-ngl 0' then
# set Vulkan OFF, add -DGGML_CPU_K_QUANTS=ON and before
# rebuild do rm -rf build to clean the previous build
# which is not strictly necessary but it 100% works.
cmake -B build -DGGML_VULKAN=ON -DGGML_BLAS=OFF \
  -DGGML_NATIVE=ON -DGGML_CPU_K_QUANTS=ON
cmake --build build --config Release -j --clean-first

Note

The OpenBLAS library is installed because it is supported but disabled because it may cause speed regression compared to the Llama ggml-CPU native backend. A similar regression is likely to occur if your Ubuntu installation isn't using libvulkan1 as per default installation or your GPU isn't powerful enough to compensate the PCI-express RAM⇆VRAM ping-pong overhead.


Testing after the build

The -ngl 0 excludes using the GPU completely, while the --mlock keep the model always in RAM:

cd build/bin
# Use this to call these binaries from anywhere
export PATH=$PWD:$PATH

model="$HOME/Downloads/Qwen3.5-4B-Q4_K_M.gguf"
opts="-ngl 0 --mlock --mmap --cpu-mask 0x0F --no-mmproj"
opts="$opts -ctk q8_0 -ctv q8_0 --swa-full --offline"
opts="$opts --temperature 0.7 --cpu-strict 1 -t 4"

./llama-cliopts -c 4096 -rea off -fa on -mmodel

The model is configured to reply without thinking and use in full the flash attention -fa on which reduces the consumption of the RAM compared to the same amount of tokens for the context -c 4096.

The context quantisation at 8-bit -ctk q8_0 -ctv q8_0 keeps a good precision but halves the consumption of the RAM compared with the 16-bit natural representation.

Moreover, for telegraphic-style answering mode, append:

  • -sys "You are a helpful assistant. Be concise in answering."

However this system prompt strongly influences the tests in a way that are much less comparable among different users / seeds, therefore it has not been used.


Running the llama server

Starting with the same for running llama-cli environment:

./llama-serveropts -c 4096 -rea off -fa on -mmodel \
    -np 1 --cache-ram 0

The last two options save RAM because the webserver is limited in running a single AI instance instead of the common four (parallelism), which each of them requires a context window cache allocated.

The server can be started manually or as system service or at the user login time, and the AI chatbot can be accessed at http://127.0.0.1:8080 by a web browser.

sudo apt install surf
aisurf() { GDK_BACKEND=x11 surf http://127.0.0.1:8080; }
aisurf

Including the minimalistic surf that has a very low RAM footprint (128MB for the whole browser, but a Chrome tab would be similar if already opened for other indispensable activities) and it is safe to use for browsing a fully trusted local address like the one provided by the llama web server.


Expected performance on i5-8365

Testing prompt:

What is the name of the capital of France?

Note that off-loading to the GPU is slower than CPU-only because the i5's GPU cannot handle all the layers:

# Type Model Name Size Read Write Mem File Fit
eq. tk/s tk/s GB GB
0¹ CRT Gemma-4 E2B-it-qat-UD Q2_K_XL gguf  ($${\color{lightgray}\textbf{wNa8o8}}$$) (4B) 70.7 22.4 ${\color{lightgreen}\textbf{》2.96《}}$$ ${\color{lightgreen}\textbf{》2.04《}}$$ 🟢
0² CRT Gemma-4 E2B-it-qat Q4_0 gguf  ($${\color{lightgreen}\textbf{full 32K @Q4{\_}0}}$$) (4B) ${\color{lightgreen}\textbf{》58.6《}}$$ ${\color{lightgreen}\textbf{》16.6《}}$$ 4.86 3.12 ✅
1 CDR Qwen-2.5 Coder 3B-it Q6_K gguf 3B 30.8 10.4 2.88 2.36 🟢
2 RPL Hermes-3 Llama-3.2 3B Q6_K_L gguf 3B 29.8 8.7 3.23 2.55 🟢
3 GNR Qwen-3.5 4B Q4_K_M gguf 4B ${\color{lightgreen}\textbf{》26.8《}}$$ 8.1 5.03 2.64 🟢
4¹ CRT Gemma-4 E4B-it-QAT Q4_0 gguf  ($${\color{lightgreen}\textbf{full 32K @Q4{\_}0}}$$) (8B) ${\color{lightgreen}\textbf{》26.4《}}$$ ${\color{lightgreen}\textbf{》9.2《}}$$ ${\color{lightblue}\textbf{》7.99《}}$$ 4.80 ✅
4² CRT Gemma-4 E4B-it-QAT Q4_0 gguf (8B) 22.5 9.0 ${\color{lightgray}\textbf{》7.53《}}$$ ${\color{lightgray}\textbf{》4.80《}}$$ ✔️
4³ CRT Gemma-4 E4B-it-obliterated Q4_K_M (8B) 23.2 7.5 7.40 4.97 —
4⁴ CRT Gemma-4 E4B-it-QAT UD-Q4_K_XL gguf (8B) 20.5 ${\color{lightgreen}\textbf{》9.2《}}$$ 6.99 3.93 🟢
5 SCI Phi-4-mini 3.8B-instruct Q5_K_M gguf 4B 21.3 9.3 ${\color{lightgreen}\textbf{》3.33《}}$$ 2.65 🟢
6 GNR Qwen-3.5 4B UD-Q5_K_XL gguf 4B ${\color{lightgray}\textbf{》18.3《}}$$ 7.3 4.15 3.08 ✔️
7 CDR NextCoder 7B i1-Q4_K_M gguf 7B 17.8 ${\color{lightgray}\textbf{》6.0《}}$$ ${\color{lightgray}\textbf{》7.86《}}$$ 4.36 ☑️
8 RSN DeepSeek-R1-dstl-Qwen 7B-uncensored i1-Q4_0 7B 15.9 6.6 8.01 4.14 —
9 GNR Apertus 8B-instruct-2509 UD-Q4_K_XL gguf 8B 13.7 ${\color{lightblue}\textbf{》5.3《}}$$ 7.61 4.78 ☑️
By Comparison:
A CRT Gemma-2 2B-it Q4_K_M 2B 47.4 13.8 3.13 2.15 1.59
B GNR Qwen-3.5 4B Q5_K_S gguf ➡ llamafile  ($${\color{lightgray}\textbf{3.75 GB}}$$) 4B 18.5 5.0 8.80 3.02 ✔️
C GNR Qwen-3.5 4B-MTP Q5_K_S gguf 4B 16.0 7.1 3.91 2.91 ✔️
Above Limits:
D SCI Hypnos i1-8B i1-IQ4_NL gguf  ($${\color{orange}\textbf{mem. 9 GB}}$$) 8B 16.9 ${\color{lightgreen}\textbf{》6.6《}}$$ ${\color{orange}\textbf{》8.64《}}$$ 4.36 ☑️
E CRT Gemma-4 12B-it UD-Q4_K_XL gguf  ($${\color{orange}\textbf{mem. 12 GB}}$$) 12B 8.9 ${\color{orange}\textbf{》3.6《}}$$ 12.2 6.86 🔶
F CRT Gemma-4 12B-it-QAT Q4_0 gguf  ($${\color{orange}\textbf{mem. 14 GB}}$$) 12B 9.1 ${\color{orange}\textbf{》4.0《}}$$ 13.2 6.50 🔶

Table's Notes

  • The human reading speed in English varies between 5 and 11 tk/s, on average 7.5 tk/s.
  • Some models are more verbose and their Wt/k drop, hence verbosity is a fair penalty.
  • Energy saving mode (max 15W TDP) otherwise i5-8365 gets hot and drops the frequency.
  • Tests were completed before adding --mmap, which by default is enabled, and --swa-full.

Data Evaluation

The prompt reading is usually faster (Rtk/s) than generation (Wtk/s) while the RAM consumption, analyzed via free, reveals the full impact of the model file and the context overhead (around 500-600MB extra). This wasn't obvious but free output remains consistent across various runs.

Threads parallelisation -t 4 should be related to the number of cores, ignoring the CPU threads. The CPU will throttle a bit above 50%, the performance will be the same, and the CPU will remain relatively colder and not fully busy.

Using -t 8 there is a regression in performances, while using --cpu-mask 0x0F the test results are much more stable and aligned with the maximum values recorded in the table.

Fundamentally, it is a matter of CPU temperature that rises from 45°C to 65°C in the first 10s of computing, then CPU starts to throttle down: more threads, more heat. A mask like 0x0F spreads the heat uniformly on the four physical cores, thus a more repeatable outcome.

While the Q4_0 might seem obsolete, it is way faster when the model is relatively big (7B) and the CPU is relatively old (i5-8th). In some models, distillation (or pruning) and uncensoring (or ablation) can spare a lot of RAM and improve speed.

By Comparison

I did as equivalent as possible tests on Qwen3.5-4B-Q5_K_S.llamafile and the most significative differences are: 1) it seems faster in loading the model in --chat mode; 2) much more pressure on the system RAM, not because the model rather than binary code redundancy; 3) apparently slower in answering. BTW, statistics are required to support these three claims.

Gemma 4's memory values collected are aligned with Google specifications. Hence, the E4B is equivalent to a 8B w/o the computational burden of a larger model. The most relevant aspect is about Gemma 4 Quantization Aware Training which suggests using Q4_0 for the KV caches is natively fine.

Therefore -ctk q4_0 -ctv q4_0 allows a relatively huge 32K context window -c((32<<10)) --swa-full while keeping the RAM usage within the 8GB limit. To grant having memory for a longer context window: 128K peaks at 9.66 GB, 64K at 8.55 GB. Instead, with the Gemma-4 12B-it-QAT Q4_0 and -ctk q4_0 -ctv q4_0 the most daring config is -c 4096 --swa-full within the 14GB limit.


Thermalisation

Considering Ubuntu base rootfs 22.04.5 and 24.04.4 are 28MB and a minimal Linux system w/ kernel 5.15 can happily run within 24MB of RAM, there is a good chance to run also the 12B Gemma 4 with a large context windows on a dedicated machine with only 16GB of RAM (2x4GB in quad-channel).

The real limit is set by the CPU's TDP and its thermal dissipation system, but desktop/mini PCs can easily deal with a 65W heat source, much more than 1.1-1.4 Kg laptops. While an old Xeon workstation can easily handle twice a consumer desktop.

PassMark and similar burst benchmarks are misleading for llama.cpp workloads. A laptop CPU (i5-8365U, 15W TDP) rated at 6000 points throttles to 800 MHz under sustained workload by 4 threads, while a workstation CPU (E5-1620 v4, 140W TDP) rated at 7000 points sustains 3.5 GHz on all 8 threads.

The real throughput factor is not 20% but 9× (140:15 = 875%), proportional to TDP budget and thermal design which is directly related with the mass of the computer. Surprisingly, AI workload is currently more similar to a steam locomotive in its relationship with energy rather than information technology, until someone shifts this paradigm.

Bare-metal minimum

A Thinkpad x280 or x390 are light, slim and compact laptops encased in an alloy chassis which is not heat-conductive like aluminium. Despite this limitation, a cheap cooling-pad can keep the max CPU temperature around 65°C and the SSD below 40°C. Without this aid, in compiling the Linux kernel with make -j8, it easily goes above 80°C and CPU's frequency drops from 1.8GHz to 800Mhz.

Considering a machine equipped with a Ryzen Pro 5 series 5000 and 16GB DDR4 3200Mhz in dual-channel, we can easily reach the conclusion that in the range [ €180, €260 ] such machine can work as a dedicated uAI-server providing a 12B model access by network, cabled or wifi indifferently. However, a dual-channel 32GB configuration is definitely more apt.

For running a basic Linux server 2GB of RAM is an "abundant luxury", therefore the Llama memory limits can be raised to 12GB (soft) and 14GB (hard). Considering the overall ratio in computational capacity, twice a Thinkpad X390, the expected throughput is 7.2 tk/s, in CPU-only mode and without specific low-level or ML optimisations.

Finally, looking at the llama pre-built releases, we can note that llamafile exists primarily for supporting CUDA on whatever OS apart from Windows x64 is a real pain. While llamafile with its 0.75GB extra on top of each AI model, detects and compiles on the fly the essential stuff for providing llama the support to deal with a local CUDA installation, if any.


Model loading time

A quick way to test the start time which includes the model loading is to pass as the first prompt the exit command. In this way it is possible to compare the starting time among various models and by a fair comparison with llamafile using the same model:

drpc() { sudo sh -c "sync; swapoff -a; echo 3 >/proc/sys/vm/drop_caches"; }

topt="$opts -c 4096 -rea off -fa on"

drpc; time -p ./llama-clitopt -p "/exit" \
  -m Qwen3.5-4B-Q4_K_M.gguf

drpc; echo "/exit" | time -p sh \
  ./Qwen3.5-4B-Q5_K_S.llamafile --chattopt

As anticipated the llamafile is faster at start-up time:

Model Size Quant. File Real User Sys Range
Qwen-3.5 4B Q4_K_M 2.64 5.64 3.94 1.51 min
5.78 4.08 1.75 max
Qwen-3.5 4B-UD Q5_K_XL 3.08 4.64 3.01 1.40 min
4.85 3.12 1.62 max
Qwen-3.5 4B-MTP Q5_K_S 2.91 4.09 2.91 1.18 min
4.20 3.00 1.27 max
.llamafile: 3.75
Qwen-3.5 4B Q5_K_S 3.02 3.88 0.76 1.59 min
4.00 1.08 1.81 max
  • The user timings aren't comparable with the .llama one due to the sh usage.
  • Two seconds (5.78 - 3.88) can be perceived but 4B-MTP load is just 5% slower.
  • The SSD hdparm throughput is 17GB/s cached reads, the model matters more than its size.

Benchmark screenshot example

The correct full approach includes checking also the resident size in memory of the llama instance running the model:

pmem() { grep -e "^Vm" /proc/$(pgrep1)/status; }
mpeak() { echo; free; pmem llama-cli | grep VmPeak; }

Dropping the cache before the run, and checking the free difference is the most straightforward way to check the pmem output:

$ drpc; sleep 15 && mpeak & free && ./llama-cliopts -c[32<<10] \
  -rea off -fa on -mmodel -p "What is the name of the capital of France?"
               total        used        free      shared  buff/cache   available
Mem:        16148688     4743292      927712      721412    10477684    10344416
Swap:              0           0           0
▄▄ ▄▄
██ ██
██ ██  ▀▀█▄ ███▄███▄  ▀▀█▄    ▄████ ████▄ ████▄
██ ██ ▄█▀██ ██ ██ ██ ▄█▀██    ██    ██ ██ ██ ██
██ ██ ▀█▄██ ██ ██ ██ ▀█▄██ ██ ▀████ ████▀ ████▀
                                    ██    ██
                                    ▀▀    ▀▀

build      : b9571-e3471b3e7
model      : Gemma-4-E2B-it-qat-UD-Q2_K_XL.gguf
modalities : text

available commands:
  /exit or Ctrl+C     stop or exit
  /regen              regenerate the last response
  /clear              clear the chat history
  /read <file>        add a text file
  /glob <pattern>     add text files using globbing pattern


> What is the name of the capital of France?

The capital of France is **Paris**.

[ Prompt: 65.8 t/s | Generation: 23.2 t/s ]
               total        used        free      shared  buff/cache   available
Mem:        16148688     5425408      228380      738572    10494900     7643900
Swap:              0           0           0

Choosing properly the AI model, it size and quantisation and aligning with it the KV cache size, despite a relatively huge 32K of context, and using a cheap cooling-pad, a strict allocation of the AI's threads on CPU's core, half of RAM and CPU threads maximum, a 2019 laptop can burst out some numbers about reading and writing tokens speeds that are quite impressing.


Conclusions

Running a local AI for a general purpose and/or sporadic use, the simplicity of LlamaFile approach wins but for everyone else it creates a certain rigidity in model choice which is not suitable or not even acceptable because it can strongly limit the choice and/or impact the performance.

Llamafile best choices

  • When simplicity is a necessity, LlamaFile is the way.
  • When flexibility is a must to have, Llama is the way.
AI Model Llamafile Intel AMD L3=16M Linux/MacM4 Windows
Qwen3.5 4B Q5_K_S 4.1 GB i5U 8-11th gen. Ryzen 5 2xxx 16 GB 16 GB
Apertus 8B-i 2509 5.9 GB !U: Xeon or H/P Ryzen 5 4xxx 16 GB 24 GB
GPT-oss 20b Q5_K_S 12 GB !U: Xeon W or i9 Ryzen 5 5xxx ' ' ' ' ↘ 32 GB
LFM2 24B-A2B Q5_K_M 16 GB "  "  "  " "  "  "  " 32 GB ↖ , , , ,

General rules of thumbs

  • The available free RAM size determines the model.
  • The CPU efficiency determines the size of the model.
  • The heat dissipation system determines the usability.
  • Llama !file allows flexibility on GGUF, like E4B QAT.
  • The OS determines the effective available free RAM.

About

A short manual to run AI locally on your PC/laptop

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors