VCode home

See the words appear while you speak

Agents: claude, codex, opencode · On: phone, desktop · optional

With Live text, the words land in the composer as you say them, not after you stop. A small speech model runs on the phone for that. When you stop, the audio goes to the VCode machine once, and its punctuated text replaces the live words.

Why you would use it

You want to watch the text build and stop as soon as it reads right, instead of watching a waveform and waiting for a transcript.

How to use it

Put the live engine's files on the box once, from a terminal on it:

  1. Install VCode and the transcription service first — see install-as-systemd-units and install-the-voice-sidecar. Live text needs the service: without it there is no mic.

  2. Make sure curl, tar, bzip2 and sha256sum are on PATH. The script names any that are missing and stops.

  3. Run:

    bash install.sh voice-browser
    

    It renders the units, reloads systemd and restarts VCode. The restart is only for a unit rendered before VCODE_VOICE_BROWSER_DIR existed — the files are served as soon as engine.json lands, and the check runs per request. /health then reports liveVoice: true.

Then, on the phone:

  1. Tap the cog (#settings-open). In the Voice group (#settings-voice), tap the Live text segment of Show the recording as (#voice-display) — see choose-the-dictation-engine.

  2. Put the caret where the words should go, or leave the box unfocused to add them at the end.

  3. Tap the mic (#mic) and speak. Each word shows up in the box (#draft) about half a second after you say it. The box stays in view the whole time; the waveform a recorded clip shows in its place does not appear here, nor during the machine's correction.

  4. Stop in any of these ways:

    • tap the mic again, and the last word is finished and kept;
    • stop talking, and it stops after the cog's silence window;
    • start typing, and it stops and keeps what landed;
    • tap the × beside the mic (#mic-cancel) to throw the dictation away and put the box back as it was when you tapped.
  5. The machine then fixes the words: the live text is replaced by the machine's, with punctuation and capitals.

Edit and send the message as you would anything you typed. Nothing is sent by voice.

What you see

Errors show as a toast, and the mic goes back to idle:

Options and settings

Option Default What it changes
#voice-display (voice.display) none saved, which is waveform live picks live text. Offered only where /health.liveVoice is true and the browser can run it
#voice-silence (voice.silenceSeconds) 3 Seconds with no new words before it stops, counted once the first word has landed. 0 stops only on the button or the cap
The cap 120 s It stops at two minutes whatever is said
VCODE_VOICE_BROWSER_DIR the unit sets $STATE_DIR/voice/browser Where the live engine's files live. Empty, or no finished engine.json, is no live engine
install.sh voice-browser --dry-run off Prints what it would fetch and change, and changes nothing

What install.sh voice-browser fetches, over https only (redirects included):

File Version Checked against
sherpa-onnx wasm release → the glue, the wasm and sherpa-onnx-asr.js v1.13.7 the tarball's sha256, then a sha256 each
sherpa-onnx-streaming-zipformer-en-2023-06-26 → encoder, decoder, joiner, tokens.txt asr-models release the tarball's sha256, then a sha256 each

The two tarballs are 175 MB and 310 MB; only the seven named files come out (86 MB), and each tarball is deleted once unpacked. A file already in place whose sha256 matches is not fetched again, so a re-run downloads nothing. engine.json is written last, so a half-finished run is no engine rather than a broken one. A box where an older run also fetched the in-browser Whisper engine has those files removed.

Limits and known gaps