Each OS and architecture has exactly one pinned whisper.cpp build (whisper-lib-b4938), chosen without looking at the host's GPU:
- Linux x86-64 has only the CPU build. A host with an NVIDIA GPU still transcribes on the CPU: the final
small.en pass over a 9.4 s answer takes 2.14 s on a 32-thread CPU, against 0.13 s for the same sources built with CUDA. This is on the path to first audio for wg21.org's Talktron gateway host.
- Windows x86-64 has only the CUDA build. Its CUDA backend links the NVIDIA driver library directly, so a host without an NVIDIA driver cannot load it and has no speech.
llama-server already handles this with [local] llama_backend: auto runs nvidia-smi to pick a build, and an explicit value forces one.
Proposal
- Add a Linux x86-64 CUDA build and a Windows x86-64 CPU build to
whisper-lib.yml, published into whisper-lib-b4938 without replacing its five archives, which shipped gateways pin.
- Add
[stt] whisper_backend = auto | cpu | cuda, shaped like llama_backend. On Windows and Linux x86-64, auto picks the CUDA build when nvidia-smi reports an NVIDIA GPU and the CPU build otherwise. Other platforms keep their single build.
Each OS and architecture has exactly one pinned whisper.cpp build (
whisper-lib-b4938), chosen without looking at the host's GPU:small.enpass over a 9.4 s answer takes 2.14 s on a 32-thread CPU, against 0.13 s for the same sources built with CUDA. This is on the path to first audio for wg21.org's Talktron gateway host.llama-server already handles this with
[local] llama_backend:autorunsnvidia-smito pick a build, and an explicit value forces one.Proposal
whisper-lib.yml, published intowhisper-lib-b4938without replacing its five archives, which shipped gateways pin.[stt] whisper_backend = auto | cpu | cuda, shaped likellama_backend. On Windows and Linux x86-64,autopicks the CUDA build whennvidia-smireports an NVIDIA GPU and the CPU build otherwise. Other platforms keep their single build.