back to the survey
ai & agents · station 07

sotto

Push-to-talk dictation for Windows — hold a key, speak, release, and clean text lands at the cursor without the audio ever leaving the machine.

survey position
N 52.22° ·E 04.53°
elevation
980 m
field status
shipped
role
Builder
stack
Python · Whisper.cpp
year
2026
links
context

Dictation on a desktop usually means a cloud service: the microphone opens, the audio travels, and somewhere else a transcript comes back. For most text that trade is invisible; for some of it, it isn’t acceptable.

sotto takes the other route. Speech becomes text on the machine, and never leaves it — no cloud, no telemetry, no account watching over the microphone.

approach

Transcription runs on whisper.cpp on the user’s own hardware, and the optional cleanup pass runs on a local model, equally offline. Even the licence key is verified offline; there is no activation server to phone.

The interaction is deliberately small: hold a key, speak, release. English and Dutch each get their own key, spoken commands like “new line” handle formatting without touching the keyboard, and the text lands wherever the cursor is, in any application.

0 bytes out

audio is transcribed on the machine — no cloud, no telemetry, no activation server

result

Built for Windows 10/11: clean text in roughly two seconds, with a personal dictionary and snippets for the words the model wouldn’t otherwise know.

The privacy claim is structural rather than contractual — there is no server involved, so there is nothing to trust.

field readings
runson-device · windows 10/11
enginewhisper.cpp · local cleanup model
gesturehold · speak · release
languagesenglish · dutch