Skip to content

Configuration

akou keeps its settings in one file, ~/.config/akou/config.json: a JSON object with the dotted keys below, for example {"asr.segmentPause": 0.8, "api.port": 8476}. You rarely need to open it. The Settings page in the window changes the common ones and saves on its own, and the command line reads and writes any of them:

akou config show                        # every setting, its value and what it does
akou config set user.name "Jordan Lee"  # change one; it is checked before it is saved
akou config unset user.name             # back to the default

A key that is unknown, of the wrong type or out of range is refused with a message, and its default is used; the rest of the file still applies. A file that is not JSON is refused whole. Some settings take effect at the next start of akou, and akou config set says so when it saves one of them.

Keys marked file only name a program akou runs or an address your audio or transcripts go to. They are read from the file and never written over the local API, so the API token cannot become a way to run a chosen command or send your calls elsewhere. Keys marked window or file can also be set on the Settings page. Secrets, such as provider.apiKey, are never shown back; on macOS the API key is kept in the Keychain.

Environment variables

Variable What it does
AKOU_HOME Moves akou's home folder, and with it the config folder and the recordings. Meant for tests
AKOU_MODELS_DIR Overrides asr.modelsDir, the folder the speech models live in
AKOU_HEADLESS Sets app.headless: run with no window
AKOU_SERVER Sets server.enabled: server mode, with per-key access
AKOU_BEHIND_PROXY Sets server.behind_proxy: a reverse proxy with TLS is in front of akou. See Server mode
AKOU_ACCELERATOR Sets asr.accelerator: the GPU the large speech model runs on. See GPUs and presets
AKOU_URL Points the command line and akou mcp at a remote akou instead of the app on this machine
AKOU_API_KEY The key for AKOU_URL
AKOU_API_KEY_FILE A file holding the key for AKOU_URL, read instead of AKOU_API_KEY

Every setting

The tables below are generated from the settings registry in src/main/config/schema.ts, which is the one source of truth; bun scripts/settings-doc.ts refreshes them, and bun run check fails when they drift. Defaults that depend on the machine are shown for macOS.

recordings

Key Default What it does
recordings.root "~/Recordings/akou" Folder that holds a subfolder per workspace, each holding one folder per call.

user

Key Default What it does
user.name "" Your name, written into every call and used for the mic channel's label.

api

Key Default What it does
api.port 8476 Port of the local API on 127.0.0.1. 0 picks a free port; runtime.json has the one in use.
api.bind (file only) "" Address the API listens on in server mode: 127.0.0.1, 0.0.0.0 or ::, the binds akou's CLI on the same box reaches. Empty: 0.0.0.0. 0.0.0.0 and :: need server.behind_proxy. The app always listens on 127.0.0.1.

app

Key Default What it does
app.headless false Run with no window. Selected by the environment, never by command-line arguments.
app.hotkey "" Global shortcut that starts and stops a call, in accelerator form (Control+Shift+F9). Empty: Option+Command+R on macOS, Control+Shift+F9 elsewhere (never Control+Alt, which is AltGr on many layouts).
app.floatingIndicator true While a call records and the akou window is not in front, a small always-on-top bar with the time, the levels, Mute, Ask and Stop. It shows no transcript text, so it can stay up during a screen share.
app.openAtLogin false Start akou, with no window, when you log in, so the hotkey and the tray are always there.

share

Key Default What it does
share.bind "tailnet" Where a share link listens: tailnet, lan (plain HTTP, visible to that network) or an IPv4 address. Never every interface unless you type 0.0.0.0.
share.port 8477 Port of the read-only share link. 0 picks a free port.

server

Key Default What it does
server.enabled (file only) false Server mode: per-key access instead of the one token, bound to api.bind, for other programs to call over the network.
server.behind_proxy (file only) false A reverse proxy in front of akou terminates TLS. Required for any bind that is not loopback; with no server.public_host, any Host header is accepted.
server.public_host (file only) "" The host name clients use (akou.example, or with a port). Set: server mode accepts that Host header and loopback only.
server.trusted_proxies (file only) [] Addresses or CIDR blocks of the proxies whose X-Forwarded-For is believed for rate limits and audit. From any other peer the TCP address is the source.
server.max_upload_mb 512 Largest upload an upload route takes, in MiB. Every other route keeps the 64 KB JSON cap.
server.max_audio_minutes 240 Longest audio a file job transcribes, in minutes. A longer file fails as too_long before it is held in memory.
server.retain_days 7 Days a file job and its result are kept before they are deleted, as a client's delete would. The upload itself is deleted as soon as the job ends.
server.concurrency 1 File jobs run at once. Each running job has its own Worker with its models loaded and asr.threads threads, so keep this times asr.threads under the cores, and the memory for that many copies of the model.
server.queue_max 1000 File jobs queued or running at most, across keys. A submit past it is refused with 429 queue_full and Retry-After. 0: no limit.
server.queue_max_per_key 500 File jobs one key may have queued or running, so one client cannot fill the queue. A submit past it is refused with 429 queue_full and Retry-After. 0: no limit.
server.default_language "auto" The language a file job is transcribed in when its request sends language: auto or none, as Telegram-Archive does. A BCP-47 tag such as es, or auto to detect it.
server.default_diarize false Label speakers in a file job whose request has no diarize field. A request that sends diarize: false gets no labels.
server.default_model "auto" The model a file job runs when its request names none (preset: auto and no model): a preset name or an engine id from the model catalog. auto: best (Qwen3-ASR) wherever it is downloaded, else fast (Parakeet) when that is, else best on a GPU with 16 GB of memory and fast elsewhere; GET /v1/server auto says which and why.
server.auto_download true What a job naming a model that is not on disk gets. On (the default): it waits while akou downloads the model, each file checked against its pinned SHA-256, then runs. Off: it is refused with 409 preset_unavailable and the akou models pull line. Download on the Models page and akou models pull fetch either way.
server.models_max_gb 40 Largest the models folder may grow through downloads of one model (a job's, or Download on the Models page), in GB (10^9 bytes). 0: no cap. A download that would pass it is refused with 409 preset_unavailable, reason: models_max_gb.
server.models_unused_days 30 Days a model may go unused before akou deletes it, in the desktop app and in server mode. The default model, any model in use (a job, a worker, the recognizer) and one downloading are never deleted. 0: never delete.
server.remotes (file only) [] Other akou servers this one sends jobs to, one entry each: <url> <key file> [names], the key file holding a jobs key of that server. A job goes to a remote when this server cannot run it, or first when names (presets or model ids, comma-separated, * for all) lists it; the client still sees only this server.
server.admin_password_hash (file only) "" The web UI's admin password, hashed. Set it with akou admin set-password.
server.dictation_slots 1 Workers kept for dictating clients (interactive=true): they never take queued jobs, and a dictation is never refused by the queue limits. 0: a dictation queues like any job. Applies at the next start.
server.dictation_engine "auto" The preset or model a dictating client's request runs when it names none. auto: the server's default.

capture

Key Default What it does
capture.helper (file only) [] Command that starts the capture helper, before its own arguments. Empty: the akou-capture bundled with the app, else the one on PATH.
capture.mic "default" Microphone: default, none, or a device id (akou cannot list the ids yet).
capture.call "system" Call audio: system, none, or app:<id>[,<id>].
capture.warmStartSeconds 3 Wait for the helper to report capturing, after a helper has captured once this run.
capture.coldStartSeconds 10 Wait for the helper to report capturing on the first start of a run.
capture.stopSeconds 5 How long a helper may take to stop before it is killed.
capture.stallSeconds 10 No packet at all for this long means the helper is wedged and is restarted.
capture.deadRestartSeconds 60 A dead call side that lasts this long restarts the helper.
capture.queueSeconds 600 Audio kept per channel for the recognizer when it falls behind.

asr

Key Default What it does
asr.modelsDir "~/Library/Application Support/akou/models" Folder the speech models are downloaded into and loaded from.
asr.threads 2 Threads per recognizer.
asr.diarizer "nemotron" Who speaks when on the call channel: nemotron (NVIDIA Nemotron 3 Diarization, live at 2 s latency and in the final pass, through the akou-diarize helper) or embeddings (voice-embedding clusters live, pyannote in the final pass). akou models pull fetches what the choice needs; takes effect at the next start.
asr.accelerator "auto" The GPU the large speech model (Qwen3-ASR, on llama-server) runs on: auto, cpu, metal, vulkan (Intel and AMD, and NVIDIA without CUDA), cuda, sycl (Intel oneAPI) or rocm (AMD). auto picks Metal on Apple silicon, CUDA for an NVIDIA card, Vulkan for an Intel or AMD GPU, else the CPU, and never SYCL or ROCm. A build that cannot open the GPU falls back to the CPU; GET /v1/server says which runs and why. Natively it picks which pinned llama-server build to download; an image runs the build it carries. Natively sycl and rocm download llama.cpp's SYCL or ROCm build, which needs Intel's oneAPI or AMD's ROCm runtime on the host. asr.llamaServer runs an own build instead. Applies to the next job that starts llama-server.
asr.parakeet.decoding "greedy" How Parakeet decodes, live and in the final pass: greedy (the default) or beam. Beam search also steers decoding toward the call's vocabulary at boost 1.5, but on some meeting audio it returns whole spans empty. Vocabulary correction when reading and after the call applies with either. Takes effect at the next start.
asr.languages [] The languages Qwen3-ASR may choose among when a job's language is auto, as ISO 639 codes (for example en and es). An answer in another language is replaced by the decode, forced into one of these, that the model scores higher; a no-speech answer stays empty. Empty: whatever language the model names.
asr.llamaServer (file only) [] Command that starts an own llama-server for Qwen3-ASR, before the arguments akou adds (for example a build compiled on this machine). Empty: the pinned llama-server release for this platform and asr.accelerator, downloaded like a model.
asr.diarizeHelper (file only) [] Command that starts the diarization helper, before its own arguments. Empty: the akou-diarize bundled with the app, else the one on PATH.
asr.live "auto" The model that writes the live transcript of a call: auto, or a model's id (nemotron-3.5-560, nemotron-3.5-1120, nemotron-en-560, parakeet-tdt-0.6b-v3-fp32, or another chunk size of a streaming Nemotron: nemotron-en-80, nemotron-en-160, nemotron-en-1120, nemotron-3.5-80, nemotron-3.5-160, nemotron-3.5-320), as the Record row's Live panel saves it. nemotron: streaming Nemotron (asr.live.engine picks which), a word shown is never taken back. parakeet: Parakeet re-decodes each stretch between pauses, and words on screen can change. auto picks nemotron when its model is downloaded, else parakeet. A model that is not downloaded never runs. upgrade, the old value, is read as nemotron with asr.review.model qwen, and saved that way. akou start --live sets it for one call. A change applies from the next call; a running call keeps its model.
asr.review.model "none" A second pass during a call, none or a model's id (qwen3-asr-1.7b, parakeet-tdt-0.6b-v3-fp32; qwen and parakeet name the same): every asr.review.everySeconds, the sentences Nemotron finished since the last review are decoded again, whole, and the new words replace the live lines once. qwen: Qwen3-ASR, the most accurate, about 10 to 13 GB of memory during a call; it needs its llama-server, and it goes off for the rest of a call it cannot keep up with. The window offers it only on a machine with a GPU for it and 16 GB of memory; set here, it runs anyway. parakeet: Parakeet, on the processor, with at most about 100 MB more memory. none: the live lines stay as Nemotron wrote them. It reviews Nemotron's lines only, so a call whose live model is Parakeet runs none. A line someone edited, or fixed a word on, keeps their text. akou start --review sets it for one call. A change applies from the next call.
asr.final.model "auto" The model that writes the final transcript after a call: auto, or a model's id (qwen3-asr-1.7b, parakeet-tdt-0.6b-v3-fp32; qwen and parakeet name the same). qwen: Qwen3-ASR on its llama-server, the most accurate; it gives no word times, so each line keeps the times of the stretch it was cut from. While the pass runs it holds about 3 GB of memory and the GPU when there is one. Without a GPU it decodes on the processor, much slower, and a pass gets half the call's length plus 300 s before it is stopped as stuck, so on such a machine a long call can fail: set parakeet there. One Qwen pass runs at a time; another waits for it. parakeet: Parakeet, on the processor. auto picks qwen whenever its model and its llama-server are downloaded, else parakeet. A model that is not downloaded never runs: the setting falls back to the other model and says why in the log, and with neither downloaded no pass runs. On Qwen the pass does not need Parakeet on disk. Speaker labels are the same with either. A Qwen that cannot start, or fails twice in a row, fails the pass, and akou finalize --force runs it again. akou finalize --model sets it for one run, and is refused when that model is not downloaded; a pass stopped by a quit runs again at the next start on this setting's model. A change applies from the next pass.
asr.review.everySeconds 60 How often the second pass (asr.review.model) reviews, seconds: the reviewed text lands about this long after the words. akou start --review-every sets it for one call.
asr.live.engine "auto" The streaming model that writes the live transcript when asr.live resolves to nemotron: auto picks by asr.languages (English only: nemotron-en-560; Spanish only: nemotron-3.5-1120; anything else: nemotron-3.5-560, which follows a switch of language), or name one. The other chunk sizes (nemotron-en-80, nemotron-en-160, nemotron-en-1120, nemotron-3.5-80, nemotron-3.5-160, nemotron-3.5-320) run only when named: a shorter chunk writes a word sooner, and auto never picks one. A word it shows is never taken back. Its model is fetched with akou models pull <name>; while none is downloaded, live lines come from Parakeet re-decoding pauses. A change applies from the next call; a running call keeps its model.
asr.segmentPause 0.7 Silence that closes a live segment, seconds. Must be below asr.segmentWindow.
asr.segmentWindow 12 Longest live segment, seconds.
asr.qwenIdleMinutes 0 Stop the Qwen3-ASR server kept warm for dictation after this many idle minutes, to get its memory back; the next dictation starts it again. 0: never.

provider

Key Default What it does
provider.kind "harness" What answers questions and writes enhanced notes: your own Claude Code or Codex (harness), an OpenAI-compatible server, the Anthropic API with your key, or none (excerpts only).
provider.harness "auto" Which harness harness runs. auto: Claude Code if found, else Codex.
provider.harnessPath (file only) "" Pin the harness program by absolute path. Empty: look it up on PATH and through the login shell.
provider.baseUrl (window or file) "" Server address for openai-compatible (Ollama: http://127.0.0.1:11434/v1), or another Anthropic API address. Set it in the akou window or the config file, never over the API: it decides where your key and transcripts are sent.
provider.model "" Model id for openai-compatible (required) and anthropic (empty: the default model).
provider.apiKey "" API key for openai-compatible (optional) or anthropic (required). On macOS it is saved in the Keychain, never in the config file; elsewhere in the config file. Never shown back or logged.
provider.timeoutSeconds 60 How long an answer may take before akou shows the excerpts instead and says why.
provider.harnessResume false Reuse one Claude Code session for follow-up questions on a call (--resume), sending only what is new since the last question. Off until measured to cut the tokens per follow-up by at least 40 % (docs/providers.md). With it on, Claude Code keeps those sessions in its own history.

memo

Key Default What it does
memo.provider "auto" Whether the configured provider keeps the rolling memo during a call, every few minutes of new speech. auto: on for openai-compatible and anthropic, off for harness, because it would run your subscription unattended; on: the harness too; off: never. An agent can always write it with akou_memo_put.

export

Key Default What it does
export.dir "" Folder finished calls are exported into, one subfolder per workspace: Markdown with frontmatter, the event log and the audio. Empty: no export until you set it.
export.audio "link" How the export carries the audio: a link to the call's file, a copy, or nothing.

hooks

Key Default What it does
hooks (file only) [] Commands run after a call, each given the call as JSON on stdin: [{"stage": "call.ended" \| "final.done" \| "enhanced", "command": "…", "timeoutSec": 600, "workspace": "work"}]. File only: they are programs akou runs.

webhook

Key Default What it does
webhook.url (file only) "" Address the call is POSTed to at every hand-off stage, signed with webhook.secret. Empty: off. File only: it decides where your transcripts are sent.
webhook.secret "" Secret for the webhook's HMAC-SHA256 signature (X-Akou-Signature). The webhook stays off until it is set. Never shown back or logged.

vocab

Key Default What it does
vocab.extraFiles [] Extra vocabulary files layered over the global and workspace files.
vocab.languages [] Languages whose word lists tell a real word from a mishearing, so a vocabulary file never "corrects" a real word. Empty: every list akou ships (en, es). A language the recognizer detects in a call is added.

dictation

Key Default What it does
dictation.enabled false Dictation: hold the dictation key, speak, and the text is inserted where the cursor is (docs/ux/DICTATION.md). Off: the helper's dictate process is not started and no key is taken.
dictation.activation "hold-or-toggle" How the dictation key works: hold-or-toggle (a press of 300 ms or more is push-to-talk, a shorter tap latches listening on until the next tap), hold (push-to-talk only) or toggle (a tap starts, the next tap stops).
dictation.hotkey "" The dictation key: a modifier alone with its side (RightCommand, RightControl, RightOption, RightShift, the Left ones, Fn) or a chord (Control+Shift+Space). Empty: RightCommand on macOS, RightControl on Windows, Control+Shift+Space on Linux, where the desktop's shortcut portal binds chords only. A change applies at once; a key the helper cannot bind is refused and the old one stays.
dictation.hotkeyFixLast "" Opens the last dictation in the draft box to correct it and teach akou the word. Empty: Shift held before a modifier-only dictation key goes down (Shift+RightCommand), else Control+Shift+Period.
dictation.hotkeyDraft "" A second key that dictates into the draft box instead of the app, so you read the text before it goes anywhere. Empty: none.
dictation.hotkeyPasteLast "" Inserts the last dictation's text again. Empty: none.
dictation.silenceStopSeconds 30 A latched dictation (tapped on, not held) stops after this many seconds without speech, and its audio is still transcribed. 0: never.
dictation.maxMinutes 20 Any dictation stops at this length, with a warning a minute before; its audio is still transcribed.
dictation.mic "" The microphone dictation listens on, a device id from GET /devices. Empty: the system default.
dictation.preferBuiltInOverBluetooth true When the default microphone is a Bluetooth headset and the built-in one is there, dictate on the built-in one, so the headset keeps its good sound profile.
dictation.warmMic "auto" How long the microphone stays open. auto: from the first press until 30 s after each dictation, so a quick follow-up keeps its first syllable; always: while dictation is on; off: only while the key is down. Never kept open on a Bluetooth microphone.
dictation.engine "auto" What decodes a dictation. fast: Parakeet, already loaded, about 0.1 s for 5 s of speech; best: Qwen3-ASR, kept warm while dictation is on, falling back to fast when it fails or is too slow, and downloaded when missing (fast until it lands); auto: best where Qwen runs on a GPU and is downloaded, fast elsewhere; remote: another akou (dictation.remote.url), with no local model needed.
dictation.final "live" The text a dictation inserts, while dictation.engine is auto (fast, best and remote there win; the Dictation page sets both). live, the default and the fastest: the words the streaming model showed as you spoke, inserted the moment you let go, with no second decode, and less accurate than Parakeet; while no streaming model is downloaded, Parakeet inserts. parakeet: Parakeet decodes the whole recording at the release, with your word list. qwen: Qwen3-ASR, the most accurate, kept warm while dictation is on, as dictation.engine best. Whatever this says, the words while you speak come from the streaming model when one is downloaded (akou models pull nemotron-3.5-560), else from Parakeet twice a second.
dictation.localTimeoutSeconds 10 How long a local best may take, plus 0.2 s per second of audio, before the dictation is decoded with fast instead and says so.
dictation.remote.url (window or file) "" The other akou a remote dictation is sent to: an https address, or http to a loopback, private or Tailscale address only. Set it in the akou window or the config file, never over the API: it decides where your dictation audio goes.
dictation.remote.key "" A jobs key of the remote akou, sent as the bearer of every remote dictation. Never shown back or logged; applies to the next dictation.
dictation.remote.fallback "local" When the remote gives no transcript: local decodes on this machine and says so; error shows the error with Retry, Copy and Open draft. With no local model installed it is error.
dictation.remote.timeoutSeconds 6 How long the remote may take before any audio, plus 0.25 s per second of audio, before the fallback runs.
dictation.language "auto" The language of a dictation: auto lets the engine choose among dictation.languages; a code (en) forces it on best and on a remote. fast picks the language itself.
dictation.languages [] The languages an auto dictation may choose among. Empty: asr.languages, so editing this never changes how calls are transcribed.
dictation.glossary "off" Send your learned dictation words to the recognizer as context. Off until a measured evaluation shows it helps without inventing names; learned words are always applied as replacements either way.
dictation.glossaryMax 24 The most learned words sent as context, by recency and use; 24 is what the decoder and the remote route take.
dictation.insert "paste" How the text goes in: paste through the clipboard, which comes back afterwards; type as key presses, for remote desktops and fields that refuse a paste (a text with a line break is pasted, so no Return is pressed); clipboard only, and you paste.
dictation.sendKey "Enter" The key pressed to send after the text is in, once the app has read it: on Enter during a dictation, Ctrl+Enter in the draft box, or after every dictation with dictation.sendAlways. none: never.
dictation.sendAlways false Press the send key after every direct dictation.
dictation.restoreClipboard true Put the old clipboard back once the app has read the dictation. Off: the dictation stays in the clipboard.
dictation.smartSpacing true Add the spaces around the text, and lower-case its first word mid-sentence, from the text around the cursor.
dictation.trailingSpace false Where the text around the cursor cannot be read, end every dictation with a space.
dictation.spokenPunctuation false Replace spoken punctuation (comma, new line; Spanish coma, nueva línea) when it stands alone between pauses. Off until measured, since both engines punctuate already and period is often just a word.
dictation.fillers true Leave out filler words (um, uh; Spanish eh, este alone) from the inserted text; history keeps what was said.
dictation.spokenSend false A dictation ending in send it (Spanish envíalo) leaves those words out and presses the send key.
dictation.format "off" provider: pass the text through your configured provider first, with the prompt dictation.formatPrompt, to fix punctuation and casing. History keeps the raw text; a provider past the timeout is skipped. It runs on Retry and on every clip sent to POST /v1/dictations (akou dictate FILE) too, so a script posting clips asks the provider once per clip.
dictation.formatPrompt "default" The formatting prompt: default, or the name of a file in dictation-prompts/ in the config folder, without .md.
dictation.formatTimeoutSeconds 0 How long the formatting pass may take before the raw text is inserted. 0: 15 s for Claude Code, 4 s for an API or a local model.
dictation.muteMedia false Pause playing media while you dictate, through the system's media controls, and resume only what akou paused.
dictation.learn "ask" When you fix a word akou heard wrong: ask offers to learn it once, and ignoring the offer changes nothing; auto learns it with an Undo; off never looks.
dictation.readField true Read the field you dictated into, for smart spacing and to learn from your fixes there. Never a password field or a terminal; on macOS only once the Accessibility grant the paste needs is there.
dictation.learn.audioCheck true Before offering a word, decode the dictation again with it on Qwen3-ASR and offer it only if the audio agrees. Where Qwen3-ASR does not run here, the offer says it was not checked.
dictation.apps [] Per-app dictation rules, matched on the app that had the keyboard: [{"app": "com.example.chat", "mode": "draft-send", "insert": "paste", "sendKey": "Enter", "engine": "auto", "language": "en", "format": "off"}]. app is a bundle id (macOS), an executable name (Windows) or a window class (Linux); name, optional, is the name of the app the rule is shown by (Slack) and is never matched; a field left out follows the global setting. mode: draft opens the draft box instead of inserting, draft-send too with Enter there pressing the send key.
dictation.pill "top" Where the dictation pill shows listening and transcribing: top is the island at the top centre of the display. Off by default on Linux, where a compositor may give the pill the keyboard and the text would land in it, so turn it on there knowingly; the tray and the sounds carry the state instead.
dictation.pillPreview true Show the words as you speak on the pill's island. akou cannot hide its windows from screen capture yet (DK-P3), so a screen share shows them too: turn this off before sharing your screen if that matters.
dictation.sounds "auto" Cues at start, stop, cancel and done. auto: soft while the pill is off, silent while it shows, so a dictation is never both silent and invisible.
dictation.retainDays 30 Days a dictation's text and audio are kept; older ones leave only a tombstone. 0: only the last one, for paste last and fix last.
dictation.keepAudio true Keep each dictation's audio for Retry and for checking a learned word. Off: deleted once the offer to learn is closed.