← All writing

Voice to Text… But What's Next?

As interfaces move from typing, to speaking, to thinking, the quality of our inner world will become a communication and safety problem.

Voice to Text… But What's Next?

WisprFlow, GPT-Live, and similar voice-to-text interfaces have changed the game for day-to-day communication. Most impactful is the experience of working with LLMs themselves: dictating instructions to your AI agents at a speed that was not possible before. For most people, speech can achieve higher WPM (words per minute) than typing.

But voice is still a physical interface. It depends on your environment, your body, and your ability to speak freely. A sleeping baby in your arms, a public sidewalk, bags in your hands, a grocery cart in front of you: there are plenty of situations where you cannot freely speak or type.

Constraints are healthy. Not every quiet moment needs to become another opportunity for productivity. Unfortunately, innovation culture rarely leaves constraints alone. Once typing is too slow and speaking is sometimes impossible, the next interface is the mind itself.

Elon Musk already planted the seeds for this next interface shift with Neuralink, years before the current text and voice AI revolution. On one hand, it is exciting to think what the upper limit on our communication looks like, where instructions can be dictated at the speed of your thoughts. At the same time, it is deeply concerning how this technology will be exploited.

To understand why that is so risky, it helps to remember that for all of human history, our communication discipline was built around external media (writing and speech), because these are the only media we have been able to observe and refine. Even then, speech is littered with stuttering and repetition. Today, this seems to be getting worse: 1) people have largely stopped writing by hand and from memory, 2) sophisticated language is hardly celebrated, and 3) there are infinite band-aids for poor communication ability. For example, WisprFlow attempts to polish the mess of stream-of-consciousness speech.

If speech is already messy, thought is even less edited. How many of you have any grip on your thought stream? If a brain-interface could suddenly transcribe Thought-to-Text, the output would be chaotic. Two decades ago, JK Rowling teased this as magic through Occlumency and Legilimency. If the technology to transcribe our thoughts is inevitable, it is imperative for us to train our attention as seriously as we train our language and speech: one, for better communication; but two, for our own safety.

The future will not only reward people who can speak well; it may reward people who can observe and regulate their own minds. In other words, if you do not have a meditation practice, you better start now.

A meditating figure held within a zen ensō

Disclaimer: this post was first drafted with WisprFlow 🙂