Hello guys,
Voice dictation on my Pebble was failing a lot: wrong words, or nothing at all. I'd had good results with Groq's Whisper API before (I use it with Hermes), so I wanted to try it as the backend for the watch. So my first idea was to change the speech-to-text provider in the Android code and build my own APK. Then I realized I'd lose the official build: update checks, sign-in, everything that comes with it. So instead I left the official app untouched and put a small "translation gateway" between the app and the internet.
The official app sends audio to Wispr Flow's REST API. On my phone I make that hostname resolve to my own server instead:
– Private DNS on Android pointed at a NextDNS profile, with a rewrite rule that maps the Wispr hostname to my VPS. Works on Wi-Fi and mobile data.
– My own CA installed on the phone. The app trusts user-installed CAs (it's in its network security config), so my server can present a certificate for that hostname and the app accepts it. No root, no repackaging.
And the gateway, A ~250-line Python script (standard library only) behind Caddy on my VPS with this steps:
-
Caddy terminates TLS with the private-CA certificate and proxies to the script on localhost.
-
The script receives the request in Wispr's JSON format (base64 WAV, 16 kHz mono), decodes it, and forwards the audio to Groq's `/audio/transcriptions` endpoint as multipart.
-
Groq returns the text, the script wraps it back into the JSON shape the app expects (`id` + `text`).
-
A small silence gate returns empty text for near-silent clips. Turns out pressing Back on the watch after a dictation starts a new recording, and Whisper loves to hallucinate "Thank you" on half a second of silence.
Once everything worked, transcriptions were still hit-or-miss. Capturing the WAVs showed the audio coming out of the watch was *very* quiet. I don't know if that's my hardware or just how it is, but a big part of it was me: the mic is on the lower-right edge of the case and I wasn't speaking into it. Speaking close to that spot fixed most of it.
On top of that I added an ffmpeg preprocessing step (trim the start/end clicks, high-pass, RNNoise, normalize speech level) before sending to Groq. With that it's working well for me now.
Happy to share more details if anyone wants to try it. Thanks for reading.
Source: r/pebble · by /u/geckotronic