The tiny model is rough in pretty much everything that uses Whisper, it’s just too small. Small or medium is where it starts getting usable, and large-v3 is the good one but it wants a graphics card to be quick.
FYI, I make one of these myself, called TalkType. What fixed the accuracy problem for me was NVIDIA’s Parakeet model. It’s about as accurate as large-v3 and still fast on a regular processor with no graphics card, for English plus 24 European languages. TalkType has it as a choice, and Handy has it too. Hold-to-talk works on Wayland as well.
The tiny model is rough in pretty much everything that uses Whisper, it’s just too small. Small or medium is where it starts getting usable, and large-v3 is the good one but it wants a graphics card to be quick.
FYI, I make one of these myself, called TalkType. What fixed the accuracy problem for me was NVIDIA’s Parakeet model. It’s about as accurate as large-v3 and still fast on a regular processor with no graphics card, for English plus 24 European languages. TalkType has it as a choice, and Handy has it too. Hold-to-talk works on Wayland as well.