Voice-to-text is moving beyond simple transcription. Relay, a London startup founded by former Nothing employees Cookie Xu and Raymond Zhu, is betting that better dictation at work needs both smarter software and dedicated hardware.
Why Relay Is Starting With the Workday
Relay is entering a crowded voice-to-text field where tools already summarize meetings, clean up speech, remove filler words, and help with multilingual conversations. Its difference is hardware: the company is building a portable microphone called the Relay Q, designed for high-fidelity dictation wherever someone is working.
The software is debuting before the device. Relay is exclusive to macOS at launch, though the company plans to bring it to Windows, iOS, and Android. The Relay Q itself is not expected until early 2027.
Zhu says the company chose desktop first because work conversations are harder than casual phone chats. In front of a computer, people switch between brainstorming, messaging, formal notes, and quick instructions. That creates a harder test for voice input than a short message to a friend or family member.
The founding problem is practical: speaking into a computer in an office can be awkward. Zhu said that while using Wispr Flow's transcription software at Nothing, he and Xu found it useful for shaping ideas, but hard to use at the right volume around coworkers.
“In the office, we never really could get the right volume for input,” Zhu tells WIRED. “If I whisper too quietly, the Mac wouldn't be able to pick it up; too loud and my coworker next to me can hear. Sometimes I would actually pull out my earbud and go like this,” he says over a video call, holding an earbud in front of his mouth to demonstrate. “It calls for dedicated hardware.”
How The macOS App Works
Relay's current app can work with or without the dedicated microphone. During setup, users choose a hotkey that activates the tool inside any text field. The default is Fn. Pressing and holding that key turns on Relay and starts dictation.
The voice-to-text engine is powered by Google's Gemini models. The app is meant to understand messy speech, not merely convert audio into raw text. In testing described in the source article, it could recognize when the speaker corrected themselves and remove irrelevant fragments.
It is not flawless. The beta did not catch every word perfectly, and punctuation could go wrong. Emoji support was also inconsistent: asking for a heart emoji worked in one case, but adding it at the end of a sentence produced the written phrase “heart emoji.”
Relay can also handle contextual actions through features it calls Skills. A user can ask it to message a colleague in Slack, and Relay is intended to open the app, find the recipient, and draft the message. The Slack skill caused some trouble in testing, while creating a Google Calendar event worked better.
Preset Skills are available for apps including WhatsApp and Gmail. Users can also make their own and connect other apps. That makes Relay less like a dictation box and more like a voice layer for the desktop.
Customization Is Central To The Pitch
Relay is also trying to adapt to how different messages should sound. Users can add notes in the Relay app to guide its output. In the source article's example, the user told Relay not to add periods at the end of Slack messages, and the app followed that instruction.
The customization can also vary by person. Relay can detect people the user regularly speaks with, then adjust tone for different relationships. A message to a boss can be kept professional, while a message to friends can be casual.
Translation is part of the same idea. Relay can be set so that messages to a coworker in another country are translated into that person's language. The source does not describe every supported language or integration, but the direction is clear: Relay wants dictation to include intent, tone, and context.
That ambition also explains why the product is not just a microphone. The microphone is meant to solve the input problem. The app is meant to decide what the words should become.
The Privacy Trade-Off
Relay's automation depends on broad system access. The app needs microphone permission, and users can also allow it to record a snapshot of the screen whenever it is activated. That screen context is how Relay can carry out Skills.
The company says data is stored locally on the user's device, where users can manually edit or delete it. When data is sent for processing, it travels through Relay's servers on the way to Google's servers, but Relay says it does not retain the data.
Even with those limits, the source article notes that giving AI models visibility into an operating system brings privacy and security trade-offs. The comparison raised is Microsoft’s Recall, another example of AI features that depend on seeing what is happening on a user's computer.
For potential users, the value question is therefore tied to trust. Relay may save time in writing messages, scheduling, or translating workplace communication. But the most powerful features require permissions that some users and companies will want to examine closely.
What The Relay Q Adds
The Relay Q is a cylindrical device with a microphone at the end. Users can hold down the top while speaking to trigger Relay on a computer or phone, much like the keyboard hotkey. A double-tap can send the message.
The top button can also be removed from the cylinder and clipped to a shirt, so the user does not have to hold the device continuously. Viewed from above, the microphone resembles the letter Q, which explains the name.
Zhu says the Relay Q should pick up speech even when the user holds it near the mouth and whispers quietly. That is the core hardware promise: make voice input usable in shared spaces without forcing people to speak loudly at their desks.
The device will come with a charging dock. Relay is also making a separate MagSafe grip for the back of an iPhone or Qi2-enabled device, letting the Relay Q attach to a phone setup and trigger mobile voice-to-text. Mobile support is planned to launch alongside the Relay Q in the first quarter of 2027.
The Q will cost $150 and includes a year of the software. After that, the software will cost $15 per month. Buying the hardware is not required to use the software.
Relay is not the only company exploring voice hardware. Sandbar unveiled the Stream Ring last year, and Pebble later showed the Index 01. Both are rings meant for whispering thoughts into them and are starting to ship right about now to initial buyers. Relay considered a ring but chose a different form, with Zhu pointing to the number of smart rings already on the market.