Fully Local Conversational AI Embedded in a Plush Toy
Developer msalsas created "The Philosopher Plush," an open-source DIY project that embeds a 100% local, privacy-centric conversational AI into a plush toy without relying on cloud services or external API keys. The physical toy acts as a lightweight client using a Raspberry Pi Zero WH that streams audio and camera feeds over WebSockets to a local machine running Whisper, Hermes 3 LLM, Kokoro TTS, and vision models. This project demonstrates how modern lightweight open-weight models (STT, LLM, TTS, and vision) can be combined to build fully private, real-time edge hardware applications. It provides a practical blueprint for developers looking to build smart voice assistants that eliminate cloud latency, API subscription costs, and privacy concerns. The hardware inside the plush includes a Raspberry Pi Zero WH connected to a microphone, speaker, camera, and a head servo motor. On the software backend, face recognition via dlib and emotion detection via FER+ allow the plush to adapt its behavior, while LangGraph manages per-person long-term conversational memory.
## BACKGROUND
Conversational AI systems typically require a pipeline combining Speech-to-Text (STT), Large Language Models (LLMs), and Text-to-Speech (TTS) components. Kokoro is an open-weight TTS model featuring 82 million parameters, while Nous Research's Hermes 3 is an open LLM fine-tuned on Meta's Llama 3.1 architecture designed for advanced reasoning and creative responses.