← All projects

Image-Text-to-Text edge computing camera

2026 · Edge computing, custom build

A fully self-contained camera that reads every photo the moment you press the shutter and answers your custom prompt with a single sentence — 100% on-device, zero internet, zero cloud. Prompt and compare up-to-date tiny VLMs on the fly: Qwen3-VL (2B), MiniCPM-V 4.6 (1B), InternVL3.5 (2B), SmolVLM2 (2.2B), Moondream 2 (2B) and Ministral 3 (3B) composing the final sentence. Every result ships with its own telemetry — total time, input→output tokens, tok/s and a vision / generate / load breakdown — so encoder efficiency and real latency are laid bare shot after shot.

Image-to-text edge computing camera, front view
Front view
Image-to-text edge computing camera, side view
Side view