Skip to content
VoxSpica

Free · Open source · Offline

Your voice, turned into text — on your own machine

VoxSpica listens to your microphone and transcribes audio files locally with VOSK. No audio ever leaves your computer: no cloud, no account, no subscription, and it keeps working with the internet unplugged.

Windows 10/11 x64 · one installer per language · the language model is included

  • No audio leaves the computer
  • No account, no payment, no ads
  • Works with the network unplugged
  • Apache-2.0, source on GitHub
Live text while you speak, with pause and stop, a language selector and a searchable history of everything you have dictated.
Live text while you speak, with pause and stop, a language selector and a searchable history of everything you have dictated.

Why VoxSpica

Dictation without the strings attached

Cloud dictation services need a subscription, a network connection and your voice. VoxSpica does the recognition on your own CPU, with an open-source engine and models you download once.

  • Offline by design

    Recognition runs in a local process powered by VOSK (Kaldi ASR). After the model is downloaded there is no network call in the product at all — unplug the cable and it keeps transcribing.

  • Private by construction

    There is no endpoint to send audio to. Recordings, transcripts and history stay in your user profile, and the history database is a SQLite file on your disk.

  • Live transcription

    Text appears while you speak. Pause and resume, stop and keep the result — the model is warmed up in the background, so the record button responds immediately.

  • Files as well as microphone

    Transcribe recorded audio (WAV, and mp3/m4a/ogg with ffmpeg) or a live microphone. Automatic language detection picks the best installed model for a file.

  • A real history

    Every recognition is stored in SQLite with its date, language, model size, device and duration. Search it in the app, or from the command line.

  • A GUI and a CLI

    The window is for dictating; the command line is for scripting, with the same settings file, the same history and the same models. Voice, meeting, file — whatever the shell is doing.

  • One file to run

    The portable build is a single executable that unpacks its own libraries at launch. Python is not required; nothing else is installed.

  • Honest about accuracy

    A small model is good, a large one is noticeably better and takes longer to load. VoxSpica says which is which instead of pretending one number fits every language.

How it works

From sound to text in three steps

  1. Install in your language

    Each installer is localized and already contains a small recognition model for its language, so the first launch works offline with nothing to download.

  2. Press record and speak

    Choose the recognition language and the model size. VoxSpica loads the model in the background while the window is already usable.

  3. Get the text

    Watch it appear live, then copy it, save it next to the recording, or find it later in the searchable history.

Languages

33 recognition languages, 9 interface languages

Recognition and interface are separate: dictate in one language while the window speaks another. Every recognition language has a small model; large models exist for the languages VOSK publishes them for.

  • العربيةar
  • Português (BR)br
  • Catalàca
  • 中文(普通话)cn
  • Češtinacs
  • Deutschde
  • Ελληνικάel-gr
  • English (India)en-in
  • English (US)en-us
  • Esperantoeo
  • Españoles
  • فارسیfa
  • Françaisfr
  • ગુજરાતીgu
  • हिन्दीhi
  • Italianoit
  • 日本語ja
  • ქართულიka
  • 한국어ko
  • Кыргызчаky
  • Қазақшаkz
  • Nederlandsnl
  • Polskipl
  • Portuguêspt
  • Русскийru
  • Svenskasv
  • తెలుగుte
  • Тоҷикӣtg
  • Tagalogtl-ph
  • Türkçetr
  • Українськаuk
  • Oʻzbekchauz
  • Tiếng Việtvn

small + large · small model · large model

Interface languages

  • English
  • Русский
  • Українська
  • Беларуская
  • Deutsch
  • Français
  • Español
  • Italiano
  • 简体中文

Nine: English, Russian, Ukrainian, Belarusian, German, French, Spanish, Italian and Simplified Chinese. Seven of them have a localized installer — Ukrainian and Belarusian are supported by the application, but InstallShield ships no wizard strings for them.

What each installer brings

A localized setup wizard, the application in that language, and a ready small model for that language — which is why every installer is about 100 MB instead of 55 MB.

small model
29
large model
26

Download

Pick your language — the model comes with it

Seven installers, each in its own language, each with a small recognition model already inside. Install, launch, dictate: no first-run download.

Installers

Or the portable build

A single executable in a zip, no installer, no bundled model: the app asks for its interface language on the first launch and downloads a model for it. The smallest download, and the one to try first.

VoxSpica-0.1.3-win64.zip61.0 MB0.1.3

Download

Windows 10 / 11, x64 only.

Checksums — v0.1.3

The installers are not code signed yet, so Windows SmartScreen may warn about an unknown publisher. Choose “More info → Run anyway”.

FAQ

Questions worth asking before you install

Does my audio really stay on my computer?

Yes. Recognition is done by a local process — VOSK, an open-source Kaldi build — on your CPU. The product contains no upload endpoint, no telemetry and no account. Unplug the network and it keeps working.

Which Windows versions are supported?

Windows 10 and 11, 64-bit. There is no 32-bit build and no macOS or Linux build yet — the roadmap has them, the code does not.

Why is an installer about 100 MB instead of 55 MB?

The portable executable is 55 MB. Each installer adds a small recognition model for its language, so that the first launch already works offline instead of downloading 30–50 MB first.

How accurate is it, really?

It depends on the language and the model, and we would rather not give you one number. A small model is good for notes and short sentences; a large model is noticeably better and takes one to two minutes to load. Pick the language, then switch the size in Settings.

Do I need a graphics card?

No. Recognition runs on the CPU. A small model needs about 0.3 GB of RAM and starts in seconds; a large one wants a few gigabytes and a minute of loading.

Is it really free?

Yes — Apache-2.0, no subscription, no payment, no ads. The optional models come from the VOSK project under the Apache 2.0 licence as well.

What happens to my recordings?

Nothing is recorded unless you ask for it. Transcripts are written to a text file, and the history database is a SQLite file in your user profile. Delete the app, delete the folder, and it is gone.

Windows warns about an unknown publisher.

The binaries are not code signed yet. SmartScreen shows a warning on first run: choose “More info” → “Run anyway”. The sources are on GitHub, and the release page carries SHA-256 checksums.

Can I use it from the command line?

Yes. `VoxSpica.exe mic`, `VoxSpica.exe file meeting.mp3`, `VoxSpica.exe history --search "meeting"` — the CLI and the window share one settings file and one history database.

Dictate without giving your voice away

Free, offline, open source. Pick a language, install it, press record.