See the words appear while you speak
With Live text, the words land in the composer as you say them, not after you stop. A small speech model runs on the phone for that. When you stop, the audio goes to the VCode machine once, and its punctuated text replaces the live words.
Why you would use it
You want to watch the text build and stop as soon as it reads right, instead of watching a waveform and waiting for a transcript.
How to use it
Put the live engine's files on the box once, from a terminal on it:
Install VCode and the transcription service first — see install-as-systemd-units and install-the-voice-sidecar. Live text needs the service: without it there is no mic.
Make sure
curl,tar,bzip2andsha256sumare onPATH. The script names any that are missing and stops.Run:
bash install.sh voice-browserIt renders the units, reloads systemd and restarts VCode. The restart is only for a unit rendered before
VCODE_VOICE_BROWSER_DIRexisted — the files are served as soon asengine.jsonlands, and the check runs per request./healththen reportsliveVoice: true.
Then, on the phone:
Tap the cog (
#settings-open). In the Voice group (#settings-voice), tap the Live text segment of Show the recording as (#voice-display) — see choose-the-dictation-engine.Put the caret where the words should go, or leave the box unfocused to add them at the end.
Tap the mic (
#mic) and speak. Each word shows up in the box (#draft) about half a second after you say it. The box stays in view the whole time; the waveform a recorded clip shows in its place does not appear here, nor during the machine's correction.Stop in any of these ways:
- tap the mic again, and the last word is finished and kept;
- stop talking, and it stops after the cog's silence window;
- start typing, and it stops and keeps what landed;
- tap the × beside the mic (
#mic-cancel) to throw the dictation away and put the box back as it was when you tapped.
The machine then fixes the words: the live text is replaced by the machine's, with punctuation and capitals.
Edit and send the message as you would anything you typed. Nothing is sent by voice.
What you see
- The first tap after the page loads downloads the model once: the mic shows a
download arrow with a filling ring,
Downloading the speech model, 42%, and the number beside it. ThenLoading the speech modelfor about a second while the model starts. Tap the mic during either to call the dictation off; the download carries on for next time. - While it listens the mic is a stop icon,
Stop recording, ringed by how much of the two minutes has gone, and the × is beside it. - The text is lowercase with a capital at the start of the box, a line, or after
a
.,?or!before the caret, and onI,I'm,I'll,I'veandI'd. There is no punctuation: the model does not produce any. - Words join the text before and after the caret with one space either side. Earlier words can still change while you speak, as the model revises its guess; only the dictated span changes, never the text around it.
- When the machine fixes the words, the mic shows the same faces a waveform clip does:
an upload arrow,
Sending the clip to be transcribed, with the share when it is known, thenTranscribing on this machine. Tap the mic then to cancel the correction and keep the live words; a new dictation needs one more tap. The correction gives up after 20 seconds. When the answer arrives, the dictated words are replaced by the machine's text. If you moved the caret meanwhile, it stays where you put it. If you typed, sent, or switched tabs in the meantime, the box is left alone and nothing is said.
Errors show as a toast, and the mic goes back to idle:
microphone permission denied,no microphone on this device, orthe microphone did not open: …model loaded — tap the mic again: the first tap waited for the download and the browser then refused the microphone, which iOS does once a tap is too old. The model is ready, so the next tap starts listeningthe speech model did not load: …, with the file and the HTTP statusthe live engine is not on this server (HTTP 404)the speech model loads from the server, which cannot be reachedthe live engine stopped: …when the worker dies mid-dictationthe speech model did not load: the engine stopped answeringwhen the load goes 30 seconds without progress, andthe live engine stopped answeringwhen a dictation hears nothing from the engine for 5 seconds. The next tap starts a fresh enginedictation stopped: the audio was interruptedwhen a call or Siri takes the audio and it does not come back within a second and a half. What landed stays… — the live text is keptwhen the machine's correction fails, with the service's own message in front. The live words stay, and no clip is kept for a retry. After 20 seconds the message says the request timed out after 20000ms and ends— the live text is kept, and the app shows itself offline until the next request gets through
Options and settings
| Option | Default | What it changes |
|---|---|---|
#voice-display (voice.display) |
none saved, which is waveform |
live picks live text. Offered only where /health.liveVoice is true and the browser can run it |
#voice-silence (voice.silenceSeconds) |
3 | Seconds with no new words before it stops, counted once the first word has landed. 0 stops only on the button or the cap |
| The cap | 120 s | It stops at two minutes whatever is said |
VCODE_VOICE_BROWSER_DIR |
the unit sets $STATE_DIR/voice/browser |
Where the live engine's files live. Empty, or no finished engine.json, is no live engine |
install.sh voice-browser --dry-run |
off | Prints what it would fetch and change, and changes nothing |
What install.sh voice-browser fetches, over https only (redirects included):
| File | Version | Checked against |
|---|---|---|
sherpa-onnx wasm release → the glue, the wasm and sherpa-onnx-asr.js |
v1.13.7 | the tarball's sha256, then a sha256 each |
sherpa-onnx-streaming-zipformer-en-2023-06-26 → encoder, decoder, joiner, tokens.txt |
asr-models release |
the tarball's sha256, then a sha256 each |
The two tarballs are 175 MB and 310 MB; only the seven named files come out
(86 MB), and each tarball is deleted once unpacked. A file already in place
whose sha256 matches is not fetched again, so a re-run downloads nothing.
engine.json is written last, so a half-finished run is no engine rather than a
broken one. A box where an older run also fetched the in-browser Whisper engine
has those files removed.
Limits and known gaps
- English only, with the accuracy of a 70 MB streaming model: expect wrong words on names, jargon and fast speech. Everything is lowercased, so acronyms and proper nouns lose their capitals. The machine's correction fixes most of this, when it runs.
- The correction runs only when the dictation ended with its text kept: a tap, the silence stop, the cap, or leaving the page. It does not run after cancel, typing, a tab switch or a send, when no words landed, or offline. There is no switch to keep the audio on the phone. An answer with no letters or digits in it leaves the live words alone.
- The correction replaces the whole dictated span with the machine's text. In
the middle of a sentence its first letter is lowercased, unless the word is
I, anI'm-style contraction, or an all-caps acronym, so a name at that spot loses its capital. Before a word that starts lowercase, its final full stop is dropped. - The worker keeps up to 120 seconds of 16 kHz audio on the phone until the session ends, for the correction; as a WAV that is at most about 3.8 MB.
- The download is 86 MB (the 13 MB wasm and the 72 MB model), once per browser.
The files are kept in Cache Storage (
portal-live-engine, a name kept from before the rename); a browser that clears site data downloads them again. - The model loads once per page. The first tap of each page load spends about a second loading it from Cache Storage before the mic opens.
- The model stays in the page's memory until the page closes.
- The browser needs
Worker,AudioWorkletNodeand WebAssembly. Without them Live text is not offered. engine.jsonis the allow-list for the live engine's files. A path it does not name is not served, and every named file must realpath to a regular file inside the directory: a symlink pointing out makes the box report no live engine. The files are served behind the login.- It stops when the page is hidden (you leave the app or lock the phone) and when you switch tabs; what landed stays in the box, and after a tab switch it stays with the tab you dictated into.
- Sending, or running a command from the box, ends the dictation first: the message is what had landed. If anything else changes the box during a dictation — a paste, or the tab closing — the dictation ends and writes nothing more.
- A dictation that hears nothing leaves the box as it was, including any text you had selected.
- A clip kept for a retry from a waveform recording goes to the machine again, not to live: the next tap retries it.
- The first tap of each page load fetches
engine.jsonfrom the server. Once the model is loaded, taps work with the server unreachable until the page reloads. - Checked in headless Chromium with a recorded voice. Not yet checked on a real phone microphone.
Related
- choose-the-dictation-engine — where Live text is picked
- dictate-a-message — the waveform recording, transcribed after you stop
- type-to-start-a-message — typing with the focus nowhere stops it too
- environment-variables —
VCODE_VOICE_BROWSER_DIR - stop-recording-on-silence — the silence window this engine also uses
- where-your-voice-goes — where the correction sends the audio
- install-the-voice-sidecar — the service that fixes the words