Skip to case study

Voice + Vision Edge AI / 2026

Jarvis Bench Assistant

Built a voice-and-vision AI lab partner on a Raspberry Pi 5: wake word, speech-to-text, and speech output all run locally, with Claude as the reasoning brain, an overhead 16MP camera for vision turns, and AI-drawn wiring diagrams on a touchscreen HUD.

The spark

I wanted an Iron Man bench: say 'hey Jarvis, what's on the bench?' and have something actually look, through a camera, and answer out loud. It became a study in edge-first AI architecture: an event bus where audio, wake word, vision, reasoning, and the HUD are independent services, so losing a peripheral degrades capability instead of crashing the assistant.

The problem

Electronics bench work stalls constantly: wiring lookups, component identification, a third hand you don't have. Cloud voice assistants can't see the bench, and they die with the WiFi.

Who it affects
Anyone doing hardware work solo: every wiring question means putting down the iron, picking up a phone, and losing the thread.
Previous workflow
A phone in one hand, a datasheet PDF in the other, and a browser full of pinout tabs.
What it costs
Constant context switching, mis-wired components discovered after power-on, and assistants that either can't see your work or need the cloud for every word.

The solution

Built a voice-and-vision AI lab partner on a Raspberry Pi 5: wake word, speech-to-text, and speech output all run locally, with Claude as the reasoning brain, an overhead 16MP camera for vision turns, and AI-drawn wiring diagrams on a touchscreen HUD.

Outcome

Deployed over the bench via systemd with git push-to-deploy: 0.94+ wake-word confidence, ~2s on-device speech recognition, and streaming replies that start speaking while the model is still writing.

What I learned

Edge AI is a latency-and-trust budget. Splitting the stack (local ears and voice, cloud reasoning, streaming output) makes a Raspberry Pi feel instant, and a resilience layer that degrades instead of crashing matters more than any single feature.

Technical decisions

Raspberry Pi 5 + Python asyncio
Event-bus architecture where every capability is an independent service that bolts on without rewrites
Anthropic API (Claude)
Streaming tool-use loop: take photo, remember, draw diagram; persistent memory injected into the system prompt
openWakeWord + faster-whisper + Piper
The entire voice loop runs locally, no cloud audio, capped at 2 CPU threads to protect the realtime loop
OpenCV + AprilTag
Bench coordinate frame: fiducial calibration translating camera pixels to real-world millimetres
aiohttp + Server-Sent Events
Live command-center HUD showing transcript, camera frame, and rendered diagrams
systemd + git push-to-deploy
A bare repo with a post-receive hook auto-restarts the service on every push

The hard part

Making a $80 computer feel instant while talking to a cloud model. The answer was pipeline overlap: the wake word and transcription never leave the device, replies stream sentence-by-sentence into local TTS so speech starts mid-generation, and a half-duplex mic gate stops the assistant from hearing itself. The hardware fought back too: the budget autofocus lens has mechanical backlash that lands focus soft, so the camera sweeps past the calibrated focus value and keeps the sharpest frame.

Next steps

  • Projection mapping: drawing wiring guidance directly onto bench parts, camera calibration run in reverse
  • AprilTag-based object measurement on the bench surface
  • Presence and air-quality sensors feeding the same event bus