Skip to content

Eval bug: Preset (models.ini) is slower then single model #298

Description

@dejan-vasic

Name and Version

./llama-cli --version
version: 10465 (fca3093)
built with GNU 11.4.0 for Linux x86_64

Operating systems

Linux

GGML backends

CUDA

Hardware

Quadro P4000 8GB

Models

Qwen3.6-35B-A3B-UD-Q4_K_M.gguf

Problem description & steps to reproduce

When I run llama turboquant with perset (models.ini) for model above I get 17 t/s but when I run it only for that specific model I get 24 t/s

First Bad Commit

No response

Relevant log output

Logs

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions