SIDE PROJECT · RETIRED JULY 2026
Doorbell Intelligence
A local evidence layer that turned camera recordings into a searchable record of what happened and when.
99 hours from one camera since June 16, 2026.
Landed on the right recording, the moment at 01:08, for $0.000315 of hosted review.
Detection, tracking, speech, sound classification, and description search all ran on my own hardware.
The route, the database, and the stored evidence are gone. The architecture and the measurements stay here.
THE REASONING
Offloading raw video capture to camera onboard storage
The initial version sampled low-bitrate RTSP streams and reconstructed clips around bounding box triggers. Because the camera's local SD card recorded full-resolution 2560x1920 video at 60 fps with high-fidelity audio, streaming over WiFi consumed server resources while yielding degraded evidence.
The revised architecture offloads raw recording and motion alerts directly to camera firmware. Server-side jobs ingest completed video blocks for offline verification, tracking, multimodal embeddings, and search.
Evidence thresholds for detections, tracks, and appearance clusters
A package detection required two consecutive open-vocabulary scans ten seconds apart, positioned outside the delivery person's bounding box. Single-frame detections were discarded as noise.
Person crops required at least 128 pixels of vertical resolution and spatial agreement across three independent masked frames before enrolling in appearance clustering. Short or unstable tracks remained unprofiled events rather than noisy identities.
Multimodal review requests were capped at four candidate contact sheets, and bounding-box detections without visual confirmation were rejected.
On-device vision and audio model pipeline
Ingest verified each completed 2560 by 1920 recording at 60 frames per second, checking codec, resolution, duration, and audio. The job persisted in SQLite so interrupted work resumed where it stopped.
The local evidence layer was YOLO26n segmentation at 1280 pixels, ByteTrack tuned for sampled video, anonymous appearance grouping, one Whisper pass over the whole recording, ACLNet sound classification, and CLIP embeddings for image search. All of it was built around the source file rather than replacing it.
Same-footage benchmarking sent Whisper back to the CPU. On the GPU it was about 2.8 times slower, 41.3 seconds against 15.0, and it took 1.87 GB on a 2 GB card sitting beside the detector before the encoder ran out of memory. Detection stayed the only persistent GPU job.
Two-way audio pipeline and local speech latency
- Take one instruction. The job for this visit, from the resident. Background archive and transcription jobs pause during live visitor interactions to allocate device bandwidth and compute.
- Read the current turn. Encrypted camera stream, current frame, visitor transcribed locally. Each turn was grounded in what was visible and what had just been said.
- Produce one bounded reply. Instruction, current frame, local transcript, and short conversation state in. The next line out, not an open-ended plan.
- Speak through the camera. Kokoro generated the audio locally and go2rtc converted it to the doorbell's PCMA 8 kHz speaker format. The link was half-duplex, so it released the speaker after every sentence to hear the answer.
- Stop on purpose. Six replies, three minutes, or thirty seconds of quiet ended a session. Raw turn audio was deleted; instruction, text transcript, image, model, and cost were kept for review.
Multimodal event indexing and visual review retrieval
The query was a person walking a dog. It returned the correct high-resolution recording with the moment at 68 seconds, exercising the whole chain rather than the detector alone: ingest, tracking, moment grouping, local search, bounded review, and a link back to the source. Hosted review cost $0.000315, because local retrieval had already cut the archive to one recording and a few frames before anything left the house. The interface displayed the real number per answer and said cost unavailable when the provider omitted it, never an estimate.
Withheld here: recording id, date, frame, person, address, and camera account.
Six sampled frames of one reference clip produced 44 detections. Nobody checking their front door wants to read that, so the interface collapsed them into one reviewable moment carrying source, subjects, transcript, sounds, and evidence. No camera frame, neighbor, address, recording, transcript, or production account appears on this page.