Back to Edge

ESP32 Special: Cloud Voices and an E-Ink Reader, No New On-Device LLM Yet

Prateek SinghOctober 7, 20263 min read9 views
ESP32 Special: Cloud Voices and an E-Ink Reader, No New On-Device LLM Yet

Two ESP32 voice projects lean on cloud AI while a maker's e-ink reader builds the hardware the next on-device model will need.

A Cloud-Backed Voice Assistant Rides a $10 ESP32 Board

On October 6, 2026, ElatoAI, an open-source voice-assistant firmware for Arduino ESP32 boards from developer akdeb, climbed GitHub's trending chart with 102 points for the day, pushing its lifetime total past 2,037 stars since the project's first commit in April 2025.

The firmware streams microphone audio from an ESP32-S3 over a secured WebSocket to a backend that can route to more than 100 different language models, including realtime APIs from OpenAI and Google, then plays the spoken reply back through a small speaker. Wake-word detection and audio capture happen on the chip; the actual language understanding does not. This is a cloud-with-a-microphone product, not on-device inference, and the project's own architecture notes are upfront about that split.

That honesty matters for this beat. ElatoAI is a usable reference design for anyone wiring a toy or appliance for a natural voice loop before a $5 chip can run a full LLM locally. It is a bridge, not the destination, but it is a well-documented one that a lot of hobbyists are clearly building on.

A $6 ESP32 Picks Up Phone Calls for ChatGPT, No Twilio Required

A Hacker News "Show HN" posted on October 6, 2026 demonstrates an ESP32 wired directly to a standard SIM card that answers incoming phone calls and lets ChatGPT carry the conversation, with no Twilio or other telephony-API middleman in between. The builder puts the microcontroller-side bill of materials at $6 in the demo video, which shows the board picking up a live call and routing the audio to a cloud model for a reply.

As with ElatoAI, the intelligence lives in the cloud. The ESP32's job is capturing and forwarding audio over the cellular network, not running a model on silicon. The HN thread picked at exactly that, with commenters asking what GSM module and audio codec path made direct SIM access possible without a carrier-grade voice API behind it.

It is a neat demonstration of how little hardware now sits between a landline-style phone number and a chatbot, and a reminder that "AI phone" builds on $6 microcontrollers still lean entirely on someone else's datacenter for the thinking part.

An ESP32-S3 E-Reader Builds the Hardware the Next On-Device Model Will Need

Posted to r/esp32 on October 7, 2026, a maker's write-up walks through "OctoFox Book," an e-ink reader built around a LILYGO Screen-4.7-S3 V2.4 board — an ESP32-S3 module paired with a 4.7-inch e-paper panel. The build adds physical page buttons, a custom 3D-printed case, and a client for the OPDS protocol so the device can browse and fetch titles straight from the builder's own self-hosted library server.

There is no language model on board here; this is firmware and enclosure design, not AI. It earns a spot on this beat because it is exactly the kind of ESP32-S3 platform — capable MCU, e-paper driver, enough flash and RAM for a real UI — that tiny on-device TTS or reading-companion models keep landing on once someone bolts inference onto a chassis like this one.

For now it is a clean, open build for anyone who wants a private, ad-free reading device, and a useful blueprint for the next ESP32-S3 project that tries to add a brain to it.

No new language model actually ran on an ESP32 chip in this sweep — just cloud voices routed through one, and a reader built for the model that hasn't landed yet. Some days the beat is the plumbing, not the brain.

References & Citations

Subscribe to new posts from theaivibe.org

No spam — just new posts. One-click unsubscribe.
Share this article

Related Posts