Skip to content
All work

Local AI voice assistant

Jarvis

A voice assistant that runs entirely on my own PC. It listens for its name, understands what you say and answers out loud, with no cloud and no API keys.

Role
Solo
Runs on
My own PC, no internet needed
Language
Python
AI models
Local, through Ollama

Why local

Most voice assistants send everything you say to a server. I wanted one that runs entirely on my own computer: it hears me, thinks and answers without anything leaving the PC, and without paying for an API.

How it works

Saying “Jarvis” wakes it up. Whisper turns speech into text on the graphics card, a local language model decides what to do, and the answer is spoken back. After each reply there's a 30-second window to keep talking without the wake word, and you can interrupt it mid-sentence.

  • Music: play, pause and queue songs on Spotify
  • Google: calendar events, Drive and Gmail
  • The PC: volume, apps, CPU and memory
  • The screen: reads on-screen text with local OCR, and can look through the webcam

Version two: measure before choosing

The second version, built on Nous Hermes models, is an agent that can run commands on the PC. Before choosing a model, I benchmarked them on my own graphics card, an RTX 5060 with 8 GB of memory.

The surprise: only about 6.5 GB of that memory is actually free, because the same card also drives the screen. A model that fits completely ran about 14 times faster than a bigger one that spills over to the CPU, and on three reasoning tests the smaller model gave the same answers.

Measured on an RTX 5060 8 GB
ModelOn the GPUSpeed
Hermes 3 8B (Q4)100%72 tokens/s
Hermes 4 14B (IQ3_M)80%about 8 tokens/s
Hermes 4 14B (Q4_K_M)58%5.4 tokens/s
The fast 8B model handles everything by default. The 14B only runs when you ask it to think hard.

Keeping it safe

Giving an AI a command line is risky, so every command is checked against a list of known read-only commands. Anything else, like deleting a file, has to be confirmed first, and if no one can confirm, the answer is no. A test creates a real file, tries to delete it and checks that it survives.

What it does

  • Speech recognition, the language model and the voice all run locally
  • Controls Spotify, volume, apps, Google Calendar and Gmail by voice
  • Reads the screen, and can look through the webcam when asked
  • A second version picks between a fast and a deep AI model, based on real benchmarks
  • Commands on the PC run only from an approved list; anything else asks first

Built with

  • Python
  • Ollama
  • Whisper
  • Nous Hermes
  • pywebview