GPUs and presets¶
Which image to run for your GPU, where the accurate best preset runs fast, and the settings for a server with thousands of files to transcribe. Start with Server mode for the image, the models and the keys.
A GPU¶
The large speech model, Qwen3-ASR, runs on llama-server, and a GPU makes it many times faster than the CPU. Every image carries a llama-server build, uses the GPU it can open, and falls back to the CPU. Pick the image for your GPU:
| GPU | Image | Add to docker run |
|---|---|---|
| None | drumsergio/akou:0.5.5 |
Nothing |
| Intel (integrated or Arc) or AMD | drumsergio/akou:0.5.5-vulkan |
--device /dev/dri --group-add $(stat -c %g /dev/dri/renderD128) |
| NVIDIA | drumsergio/akou:0.5.5-cuda |
--gpus all, with the NVIDIA Container Toolkit on the host. The image carries the CUDA runtime; the host needs only the driver (570 or newer on x64) |
| Apple silicon | None: Docker on macOS has no GPU | Run akou on the Mac itself (A Mac as the server); it uses Metal |
--group-add gives the container's user the group that owns the render node on the host (render on most distributions). Without it the GPU is there but akou cannot open it, and it says so. In compose, the Vulkan image takes:
devices: ["/dev/dri:/dev/dri"]
group_add: ["993"] # the number `stat -c %g /dev/dri/renderD128` prints on the host
and the CUDA image:
To see what it chose:
gpu is the GPU backend (metal, vulkan, cuda, sycl or rocm) or null for the CPU. accelerator.device is the GPU's name as llama-server lists it, verified is true once llama-server itself confirmed the device, and reason says why when it runs on the CPU. The setting asr.accelerator overrides the choice: auto (the default), cpu, metal, vulkan, cuda, sycl or rocm, also as the environment variable AKOU_ACCELERATOR. auto never picks SYCL or ROCm, which need Intel's oneAPI or AMD's ROCm runtime on the host; Vulkan runs the same cards. OpenVINO is not offered: its llama.cpp backend does not run speech models yet.
The GPU runs the best preset's Qwen3-ASR (The best preset); Parakeet (fast) stays on the CPU, where it already runs far faster than real time.
The best preset¶
best runs Qwen3-ASR-1.7B, the most accurate open model akou knows for English and Spanish, as a child process of akou (llama-server, pinned to llama.cpp release b11200 and downloaded like a model). A job asks for it with preset=best; a server makes it the default for every job that names nothing with server.default_model:
asr.languages lists the languages the people you transcribe speak. Qwen picks the language of each stretch of audio itself, and sometimes names one nobody spoke (a filler heard as Chinese); with the list set, such a stretch is decoded again in the listed language the model scores higher. A job that sends language gets that language instead.
Where it runs is asr.accelerator:
| Machine | Setting | What runs |
|---|---|---|
A Mac with Apple silicon, akou run natively (akou serve) |
auto (the default) |
Metal. On a Mac mini M4 a 10-minute meeting with speaker labels took 94 s, a real-time factor of 0.16 |
| The Docker image, any Linux box | auto |
The GPU the image can open, else the CPU: the -vulkan image on an Intel or AMD GPU, the -cuda image on NVIDIA (A GPU). The plain image runs the CPU, several times slower |
| Linux or Windows, akou run natively, with an NVIDIA card | auto or cuda |
llama.cpp's CUDA build and NVIDIA's CUDA runtime, both downloaded with Qwen, so the host needs only the driver |
| Linux or Windows, akou run natively, with an Intel or AMD GPU | auto or vulkan |
llama.cpp's Vulkan build, through the GPU's Vulkan driver (Mesa on Linux) |
| Linux or Windows x64, akou run natively, with Intel's oneAPI or AMD's ROCm installed | sycl or rocm |
llama.cpp's SYCL or ROCm build, downloaded with Qwen. auto never picks these |
| Any machine, akou run natively | cpu |
llama.cpp's CPU build (on a Mac, the Metal build with no GPU device) |
Docker on a Mac has no Metal, so on a Mac run akou natively rather than in a container. A server elsewhere on the network (a Telegram-Archive box, for example) then reaches it by URL and key like any client.
GET /v1/server shows where Qwen runs, in the provider of its entry in engines (metal, vulkan, cuda, sycl, rocm or cpu), and gpu and accelerator say which GPU was found and why. A setting with no build here (metal on Linux, or sycl in an image) runs on what auto finds, the CPU when there is no GPU, and accelerator.reason says so. Natively, akou asks the build which devices it can open once best has downloaded it; a GPU it cannot open runs on the CPU build. For a build akou does not pin, compile llama-server on the machine and name it in asr.llamaServer in config.json (for example ["/opt/llama.cpp/build/bin/llama-server"]); akou adds the model and port arguments.
A large backlog¶
A client with thousands of files to send, such as Telegram-Archive transcribing a whole archive, leans on three settings:
| Setting | Default | What it does |
|---|---|---|
server.concurrency |
1 | Jobs run at once. Each running job loads its own copy of the model and uses asr.threads threads (default 2), so keep server.concurrency times asr.threads under the machine's cores: on a 20-thread box, 4 jobs of 4 threads leaves room for the rest |
server.queue_max |
1000 | Jobs queued or running at most, across every key. 0 means no limit |
server.queue_max_per_key |
500 | The same for one key, so one client cannot fill the queue. 0 means no limit |
Set them on the web page's settings, with PATCH /v1/config and an admin key, or in config.json as above. A new server.concurrency applies from the next submit or job end.
A submit past a limit is refused with 429 queue_full and a Retry-After header in seconds, before akou reads the upload; wait that long and send it again. A job may carry priority, from -10 to 10 (default 0): a higher one runs first, then the oldest. The queue is kept in jobs.db in the data volume, so a restart resumes it in the same order.
GET /v1/server and GET /healthz answer how the queue is doing, with no key:
{ "concurrency": 4, "max": 1000, "max_per_key": 500, "depth": 212, "queued": 208, "running": 4,
"jobs_last_hour": 610, "audio_seconds_last_hour": 21480, "mean_job_seconds": 23.5, "eta_seconds": 1246 }
audio_seconds_last_hour over 3600 is how many hours of audio the box transcribes per hour. eta_seconds is the time left at the pace of the last 50 jobs, and null until one has ended since the server started.