A real-time American Sign Language recognition and air-drawing web app that runs entirely in the browser.
SignSense is a self-directed project exploring on-device machine learning for accessibility. It reads American Sign Language from a live webcam feed and turns it into text in real time, and it also lets you draw in mid-air with hand gestures — all processed locally in the browser, with nothing sent to a server.
Most sign-language recognition tools depend on cloud services, which introduces latency, running costs, and a privacy trade-off: your camera feed has to leave your machine. I wanted to see how far I could push a fully client-side approach — fast enough to feel real-time, private by design, and with no backend to pay for or maintain.
I built SignSense solo, end to end: the landmark-extraction pipeline, the machine-learning models and their training, the real-time inference loop, and the React front end with the air-drawing canvas.
React and the Canvas API provide the interface, MediaPipe extracts hand and pose landmarks, and TensorFlow / Keras models handle static fingerspelling and GRU-based motion sequences. The inference pipeline runs on the user's device.
Real-time recognition needed to avoid the latency and privacy cost of uploading camera frames. I used MediaPipe landmarks as compact model inputs and kept inference in the browser, allowing both the recognition and air-drawing interactions to work without sending the webcam feed to a backend.
The live application is available for hands-on review. The current portfolio material does not document a formal accuracy study or performance benchmark, so no unsupported metric is presented.
This project developed my understanding of browser-based machine-learning pipelines, time-sequence modeling, camera interaction, and the trade-offs involved in keeping computer-vision features private and responsive on the client.