Voice + Vision Edge AI / 2026
Jarvis Bench Assistant
Built a voice-and-vision AI lab partner on a Raspberry Pi 5: wake word, speech-to-text, and speech output all run locally, with Claude as the reasoning brain, an overhead 16MP camera for vision turns, and AI-drawn wiring diagrams on a touchscreen HUD.
The spark
I wanted an Iron Man bench: say 'hey Jarvis, what's on the bench?' and have something actually look, through a camera, and answer out loud. It became a study in edge-first AI architecture: an event bus where audio, wake word, vision, reasoning, and the HUD are independent services, so losing a peripheral degrades capability instead of crashing the assistant.
The problem
Electronics bench work stalls constantly: wiring lookups, component identification, a third hand you don't have. Cloud voice assistants can't see the bench, and they die with the WiFi.
- Who it affects
- Anyone doing hardware work solo: every wiring question means putting down the iron, picking up a phone, and losing the thread.
- Previous workflow
- A phone in one hand, a datasheet PDF in the other, and a browser full of pinout tabs.
- What it costs
- Constant context switching, mis-wired components discovered after power-on, and assistants that either can't see your work or need the cloud for every word.
The solution
Built a voice-and-vision AI lab partner on a Raspberry Pi 5: wake word, speech-to-text, and speech output all run locally, with Claude as the reasoning brain, an overhead 16MP camera for vision turns, and AI-drawn wiring diagrams on a touchscreen HUD.
Outcome
Deployed over the bench via systemd with git push-to-deploy: 0.94+ wake-word confidence, ~2s on-device speech recognition, and streaming replies that start speaking while the model is still writing.
What I learned
Edge AI is a latency-and-trust budget. Splitting the stack (local ears and voice, cloud reasoning, streaming output) makes a Raspberry Pi feel instant, and a resilience layer that degrades instead of crashing matters more than any single feature.
Technical decisions
- Raspberry Pi 5 + Python asyncio
- Event-bus architecture where every capability is an independent service that bolts on without rewrites
- Anthropic API (Claude)
- Streaming tool-use loop: take photo, remember, draw diagram; persistent memory injected into the system prompt
- openWakeWord + faster-whisper + Piper
- The entire voice loop runs locally, no cloud audio, capped at 2 CPU threads to protect the realtime loop
- OpenCV + AprilTag
- Bench coordinate frame: fiducial calibration translating camera pixels to real-world millimetres
- aiohttp + Server-Sent Events
- Live command-center HUD showing transcript, camera frame, and rendered diagrams
- systemd + git push-to-deploy
- A bare repo with a post-receive hook auto-restarts the service on every push
The hard part
Making a $80 computer feel instant while talking to a cloud model. The answer was pipeline overlap: the wake word and transcription never leave the device, replies stream sentence-by-sentence into local TTS so speech starts mid-generation, and a half-duplex mic gate stops the assistant from hearing itself. The hardware fought back too: the budget autofocus lens has mechanical backlash that lands focus soft, so the camera sweeps past the calibrated focus value and keeps the sharpest frame.
Next steps
- Projection mapping: drawing wiring guidance directly onto bench parts, camera calibration run in reverse
- AprilTag-based object measurement on the bench surface
- Presence and air-quality sensors feeding the same event bus