As much as I like seeing this, the space is really crowded, and I think the real value is not in single dictation for input, but in two other things: diarization (essential for call recordings, and a staple in video calling services, but hardly touched for in-room meetings and brainstorming) and the next step, which is cleaning it all up for actually useful notes. Not just the transcript, but better variations on meeting notes that you can customize, link to previous existing information, etc.
I think we're about to see a dictation revolution.
I've been coding with Claude, and I pretty much do it exclusively by voice now, simply because it's so much faster. And I kind of feel like Scotty from Star Trek IV, when he holds up a Mac mouse to try to talk into it.
I've been using VoiceInk [1] which looks like it's basically the same as this, but has been around for longer.
What has really made it work for me is using a Bluetooth media control like [2], where I use Karabiner Elements to remap its play/pause button to the dictation keyboard shortcut. I map rewind to option-backspace to delete the last word, fast-forward to shift-enter to insert a newline, a lower button to enter to submit my prompt, and the volume up/down buttons to scroll up/down. I've also just ordered to Xiaomi Bluetooth remote [3] that includes a microphone itself to see if I can get it to work by speaking directly into it and using it as the microphone -- there are a couple of open source projects to turn it into a Mac microphone directly. Since I'd like to be able to talk more quietly instead of into my Mac.
But what I'm REALLY waiting for is the ability to use one button for dictation, and a second button for issuing commands for a local LLM to interpret. So I hold down the voice button which transcribes "I think we need to catch that" and then the second button and go "change catch to cache, like c-a-c-h-e". Or hold down the second button and go "switch to VS code". I don't want something as finicky as macOS Voice Control, I want a local LLM I can speak naturally to.
I can very much see a future where I spend the majority of my "work" time using a Apple TV-type remote.
i forked handy and built 2 modes. One for dictation and one for quick agent actions that uses the pi harness with local models and jev style classifiers so it can read screens and interact with elements (via accessiblity tree, screenshots, os scripts, bash, mcp, etc). Its nice because in agent mode you can just give it instructions on how to respond (like dictation that can read the context of the current page or input). I've been pretty happy with it so far and debating if i should release it. Just not sure if the world needs yet another vibe coded agent system.
If folks like local-AI dictation, but want it optimized for meetings recordings (like Granola/Otter/Notion) check out Biscotti. Free, local, separates voices, voice identification, AI summaries, etc.
This is awesome. I really want to get something like this integrated directly into my agent IDE (https://getness.dev).
The main thing I need is some way of auto-sending a "end of message" button (like a newline would do) at the end of the message so that it auto-sends the chat.
Not sure the exact right way to implement that, but it seems like something that could be general purpose. Maybe a setting per-active app?
Many people are using Nemotron (NVIDIA provides Linux binaries for ARM and Intel that you can shove into a server and do audio over websockets for them).
I use handy.computer in push-to-talk mode. It inserts text in the active text box when you let go. It also supports auto-submitting at the end, so maybe that would help? I think it can send new line.
But trying to implement it directly into ness hasn't worked that well. It seems like betterwispr/wisprflow/etc all seem to belong better as their own app, and we just need good integration points between the push-to-talk app and the text consuming apps (like the IDE)
Nevertheless I dictate almost exclusively with Claude Code now, because it's such better ergonomics. And it's still significantly faster.
But a big benefit is that I'm never hunching, I'm never straining my wrists, none of that stuff. Dictation is just so much better for your neck and back.
To be honest, I used Wspr so much I only recently started realizing just HOW bad it gets trying to 'improve' your diction, especially coding. I have seen totally opperate directives. I had no idea it was potentially occuring, but it gets past a point where it is WAY to confident in guessing what you meant, sometimes to the extent of changing DON'T to DO, for example.
I’m building BetterWispr, a free and open-source voice dictation app for macOS.
I wanted to make voice typing more accessible without requiring a monthly subscription or sending every recording to a cloud server.
You can hold Option + Space, speak naturally, and have the transcription inserted directly into whatever app you’re using.
A few things I’ve been working on:
- Local speech recognition using Whisper, Parakeet, or Apple Speech
- Automatic cleanup of filler words and repeated phrases
- Custom vocabulary for names and technical terms
- Different writing styles depending on the app
- Optional meeting transcription and summaries
The project is open source under Apache 2.0.
I’m still improving the experience, especially around transcription accuracy, speed, and reliability across different Macs.
I’d love feedback from people who regularly use voice dictation.
What would make you switch from your current dictation tool to an open-source alternative?
https://developers.google.com/edge/foresight
I've been coding with Claude, and I pretty much do it exclusively by voice now, simply because it's so much faster. And I kind of feel like Scotty from Star Trek IV, when he holds up a Mac mouse to try to talk into it.
I've been using VoiceInk [1] which looks like it's basically the same as this, but has been around for longer.
What has really made it work for me is using a Bluetooth media control like [2], where I use Karabiner Elements to remap its play/pause button to the dictation keyboard shortcut. I map rewind to option-backspace to delete the last word, fast-forward to shift-enter to insert a newline, a lower button to enter to submit my prompt, and the volume up/down buttons to scroll up/down. I've also just ordered to Xiaomi Bluetooth remote [3] that includes a microphone itself to see if I can get it to work by speaking directly into it and using it as the microphone -- there are a couple of open source projects to turn it into a Mac microphone directly. Since I'd like to be able to talk more quietly instead of into my Mac.
But what I'm REALLY waiting for is the ability to use one button for dictation, and a second button for issuing commands for a local LLM to interpret. So I hold down the voice button which transcribes "I think we need to catch that" and then the second button and go "change catch to cache, like c-a-c-h-e". Or hold down the second button and go "switch to VS code". I don't want something as finicky as macOS Voice Control, I want a local LLM I can speak naturally to.
I can very much see a future where I spend the majority of my "work" time using a Apple TV-type remote.
[1] https://github.com/Beingpax/VoiceInk
[2] https://www.amazon.com/Satechi-Bluetooth-Multimedia-Remote-C...
[3] https://www.notebookcheck.net/Xiaomi-Bluetooth-Remote-2-Pro-...
Bonus points if you work in an office. This kind of workflow would be a nightmare.
…unless it means we get our private offices back.
https://github.com/scosman/Biscotti
Congrats on shipping!
And if MacOS 26 does it right - all of this is no longer necessary for us to vibe fork/ code on weekends :-)
P.S. Handy looks slick. May be able to ditch mine
The main thing I need is some way of auto-sending a "end of message" button (like a newline would do) at the end of the message so that it auto-sends the chat.
Not sure the exact right way to implement that, but it seems like something that could be general purpose. Maybe a setting per-active app?
But trying to implement it directly into ness hasn't worked that well. It seems like betterwispr/wisprflow/etc all seem to belong better as their own app, and we just need good integration points between the push-to-talk app and the text consuming apps (like the IDE)
Nevertheless I dictate almost exclusively with Claude Code now, because it's such better ergonomics. And it's still significantly faster.
But a big benefit is that I'm never hunching, I'm never straining my wrists, none of that stuff. Dictation is just so much better for your neck and back.
Are folks finding it lacking / needing alternatives?
I’m building BetterWispr, a free and open-source voice dictation app for macOS.
I wanted to make voice typing more accessible without requiring a monthly subscription or sending every recording to a cloud server.
You can hold Option + Space, speak naturally, and have the transcription inserted directly into whatever app you’re using.
A few things I’ve been working on:
- Local speech recognition using Whisper, Parakeet, or Apple Speech - Automatic cleanup of filler words and repeated phrases - Custom vocabulary for names and technical terms - Different writing styles depending on the app - Optional meeting transcription and summaries
The project is open source under Apache 2.0.
I’m still improving the experience, especially around transcription accuracy, speed, and reliability across different Macs.
I’d love feedback from people who regularly use voice dictation.
What would make you switch from your current dictation tool to an open-source alternative?
Website: https://betterwispr.com/