Install the transcription service on the box
bash install.sh voice builds a python venv, downloads a 482 MB speech model,
starts v-code-voice.service on loopback, and points VCode at it. After
that the mic appears in the composer.
Why you would use it
Dictation is optional per machine. Run this on the box you dictate to. It is a
separate command from install/update because of the download size: most
machines never want it.
How to use it
Install VCode first (install-as-systemd-units,
bash install.sh install); this command needs the env file to exist.Make sure
uv,ffmpeg,curl,tarandsha256sumare onPATH. The script names any that are missing and stops.Run:
bash install.sh voice # 4 decode threads bash install.sh voice --threads 8 # and remember 8 for every later renderIt ends with
systemctl --user statusfor the units, so you can see the sidecar running.Follow it with
journalctl --user -u v-code-voice -fif a clip fails.
Run it from the main checkout: install, update, voice and voice-browser
refuse to run from a linked git worktree.
What you see
The script prints what it did and what it skipped: venv already there, left alone: …, model already there, left alone: …, added VCODE_STT_URL to …
or VCODE_STT_URL already set, left alone: …, and set VOICE_THREADS=8 in …. An env file with an empty VCODE_STT_URL gets the warning
warning: VCODE_STT_URL is empty in …, so voice stays off.
The sidecar's first journal line is voice sidecar: sherpa-onnx-nemo-parakeet-tdt-0.6b-v2-int8 on http://127.0.0.1:3446 with 4 threads. GET http://127.0.0.1:3446/health answers {"ok": true, "model": "…"}.
In the app, the mic (#mic) and the cog's Voice group (#settings-voice)
appear after the VCode restart the script performs.
Options and settings
| Option | Default | What it changes |
|---|---|---|
--threads N |
4 |
CPU threads sherpa-onnx decodes on. Written to the env file as VOICE_THREADS and read back by every later render |
VCODE_STT_URL |
set to http://127.0.0.1:3446 when absent |
Where VCode forwards the phone's clip. Never overwritten once present |
--state-dir DIR |
~/.local/state/v-code |
Where the venv (voice/venv) and model (voice/models/…) live |
--dry-run |
off | Print what would happen and change nothing |
| Sidecar port | 3446 |
Rendered into ExecStart; the voice unit reads no env file |
install --voice |
off | Does the same provisioning at the end of a fresh install |
Limits and known gaps
- The model is
sherpa-onnx-nemo-parakeet-tdt-0.6b-v2-int8, 482 MB, downloaded over https only (redirects included) and checked against a sha256 pinned ininstall.sh. A mismatch deletes the archive and stops the run. - English only. The model is fixed; the
modelfield VCode sends isparakeetbecause the OpenAI shape requires one, not because it selects anything. - The sidecar refuses a clip over 125 seconds of audio with 413, and a body over 26 MiB with 413 before reading it. Audio under 0.3 seconds comes back as an empty transcript rather than a guess: transducer models invent words on near-empty audio.
- ffmpeg is given 20 seconds per clip. Past that the clip fails with
ffmpeg gave up after 20sin the journal andthe audio could not be decodedon the phone. - One decode at a time: one loaded model behind one lock. Concurrent clips queue.
- The unit caps memory at 4G (a 119 second clip measured 2.1 GB peak RSS) and
tasks at 256, runs at
Nice=5, and restarts on failure every 5 seconds. A broken install exits 1 and logs one line per retry while VCode answers - The unit sets no
ProtectSystem,ProtectHome,PrivateTmp,PrivateDevicesorRestrictNamespaces: those need a mount namespace, which a--usermanager cannot always create. install,updateandstatusskipv-code-voice.serviceunless$STATE_DIR/voice/venv/bin/pythonand the model'stokens.txtboth exist, so a box without voice is left alone.- A re-run downloads nothing it already has.
uv pip installruns every time with--require-hashesagainst the generated lockvoice/requirements.txt. - Without
uvat update time the script warns and keeps the packages the venv has:warning: uv not found, skipped …/voice/requirements.txt.
Related
- point-dictation-at-another-service — another box, a GPU engine, or off
- dictate-live — live text, whose words this service fixes when you stop
- dictate-a-message — what the operator just enabled
- install-as-systemd-units — the install this one extends
- update-the-install — what a later
updatedoes with the voice unit - logs-and-troubleshooting — reading the journal when a clip fails
- environment-variables —
VCODE_STT_URLandVOICE_THREADS