Skip to content
DnsLister Forum

Where domain hunters compare notes

I routed my Pebble’s voice dictation to Groq (Whisper) without modifying the official app

Hello guys,

Voice dictation on my Pebble was failing a lot: wrong words, or nothing at all. I'd had good results with Groq's Whisper API before (I use it with Hermes), so I wanted to try it as the backend for the watch. So my first idea was to change the speech-to-text provider in the Android code and build my own APK. Then I realized I'd lose the official build: update checks, sign-in, everything that comes with it. So instead I left the official app untouched and put a small "translation gateway" between the app and the internet.

The official app sends audio to Wispr Flow's REST API. On my phone I make that hostname resolve to my own server instead:

– Private DNS on Android pointed at a NextDNS profile, with a rewrite rule that maps the Wispr hostname to my VPS. Works on Wi-Fi and mobile data.

– My own CA installed on the phone. The app trusts user-installed CAs (it's in its network security config), so my server can present a certificate for that hostname and the app accepts it. No root, no repackaging.

And the gateway, A ~250-line Python script (standard library only) behind Caddy on my VPS with this steps:

  1. Caddy terminates TLS with the private-CA certificate and proxies to the script on localhost.

  2. The script receives the request in Wispr's JSON format (base64 WAV, 16 kHz mono), decodes it, and forwards the audio to Groq's `/audio/transcriptions` endpoint as multipart.

  3. Groq returns the text, the script wraps it back into the JSON shape the app expects (`id` + `text`).

  4. A small silence gate returns empty text for near-silent clips. Turns out pressing Back on the watch after a dictation starts a new recording, and Whisper loves to hallucinate "Thank you" on half a second of silence.

Once everything worked, transcriptions were still hit-or-miss. Capturing the WAVs showed the audio coming out of the watch was *very* quiet. I don't know if that's my hardware or just how it is, but a big part of it was me: the mic is on the lower-right edge of the case and I wasn't speaking into it. Speaking close to that spot fixed most of it.

On top of that I added an ffmpeg preprocessing step (trim the start/end clicks, high-pass, RNNoise, normalize speech level) before sending to Groq. With that it's working well for me now.

Happy to share more details if anyone wants to try it. Thanks for reading.

Source: r/pebble · by /u/geckotronic

Leave a Reply

Your email address will not be published. Required fields are marked *