AI assistants

Voice

Talk to Dex instead of typing. Voice is woven into Dex, so you can speak a question and hear the answer, speak a task and have an agent run it, and every dictation is kept in one searchable catalog.

The voice center

Voice in Deska is not a separate dictation box bolted on the side. It is the front door to Dex. A compact center sits at the bottom of the canvas where you can either type a question or hold to talk, and it carries you through the whole exchange: you speak, Dex listens, thinks, and speaks its answer back, all in one place.

It moves through a few clear states so you always know what is happening:

  • Listening: a live waveform shows it is hearing you. Press Esc to cancel or Enter to send.
  • Thinking: your words have been transcribed and handed to the model.
  • Speaking: the reply streams in, with small chips noting any actions Dex took on the canvas and a link to open the full thread.

You can also dictate while composing a typed message. Tap the mic, speak, and the transcription drops into the text field for you to edit before sending.

Push to talk

The fastest way to reach Dex by voice is to hold a key. Press and hold fn on macOS, speak, and release to send. On Windows, hold CtrlWin the same way, and on Linux the same chord is labeled CtrlSuper; the talk key is still in beta on both. On Linux it needs an X11 or XWayland session, because a native Wayland session hides global keys and Settings will say the listener is unavailable. The voice center takes over while the key is down and hands your words straight to the assistant.

The talk key is context aware. Hold it while a text field has focus and it dictates into that field, typing what you say at the cursor; hold it anywhere else and it talks to Dex, who does something and gives you an answer back. A quick tap cancels, and Esc cancels a dictation in progress, so nothing is sent.

Some Macs have no Apple fn key, and third-party external keyboards never send it. For those, open Settings → Keybindings → Push to talk and pick ControlOption as the hold key instead.

NoteOn macOS the talk key needs the Accessibility permission; Settings shows a Grant access button until it is given. And if tapping fn opens the emoji picker, tell macOS to let it go: System Settings → Keyboard → set “Press globe key to” to Do Nothing (the globe is printed on the fn key).

Hands-free conversation

Hold to talk is one exchange. For a longer back and forth, double-tap the talk key instead and Dex opens a continuous session with nothing held down. It listens, takes its turn, speaks the answer, then listens again, until you double-tap the key once more to end it.

It is the same key as push to talk, so whichever hold key you picked under Settings → Keybindings → Push to talk is also the one you double-tap.

Talk to an agent

Because voice flows into Dex, speaking is not limited to dictating text. You can ask for work to be done. Say “run the tests” or “what is failing in the build terminal?” and Dex uses what is open on your canvas to answer, or hands the task to a coding agent and reports back. It is the same assistant you reach with Cmd/CtrlL, just driven by your voice.

Which agent takes that work is yours to choose. The audio popover’s Coding agent row lists the agents installed on this computer (Claude Code, Codex, opencode), and each one opens onto the models its CLI advertises: the same list, in the same order, that the model chip on a task offers. Pick a model and every task Dex hands that agent starts on it; leave it on Default model and the agent decides, as before. The model is remembered per agent, so switching agents never carries the other one’s model over.

NoteA spoken turn runs on a capable default model, not a budget one. It drives the same multi-tool loop as a typed message, and lighter models were unreliable at following it, so Dex picks a model chosen for tool reliability. That means a spoken request costs about what the same request typed would.

Starting a dictation

To capture text rather than ask a question, use the dictation shortcut: CmdShiftV on macOS, or CtrlAltV on Windows and Linux. Because it is registered globally, it works even when Deska is in the background, so you can capture a thought without switching windows first. Inside Deska there is an even shorter path: hold the talk key while a text field has focus and it types what you say straight at the cursor.

  • Pick your microphone from the input picker if you have more than one.
  • Choose a reply language: Auto matches the language you speak, or pin it to English, Spanish, Portuguese, French, German, Italian, Japanese or Chinese.
  • Watch your balance: the voice minutes you have left this month are shown in Settings under Account, and on your account dashboard on the web.

Spoken replies

Dex can read its answers back to you. In the audio popover you pick an agent voice: Cove (calm and neutral), Ember (warm and expressive) or Sage (crisp and direct). Each has a preview, and your choice sticks across sessions.

There is nothing to configure. With no key set, synthesis runs on Deska’s managed voice service and draws on the same voice-minute allowance as dictation. If synthesis is unavailable, because you are signed out, out of minutes, or an error got in the way, Deska falls back to your operating system voice, which is functional but flat. Adding your own Fish Audio or OpenAI key under Settings → Provider keys is optional: it moves synthesis onto your key and budget instead of the allowance.

The edge glow

While you are talking to Dex, a soft, Apple-Intelligence-style glow breathes around the edges of the window and pulses in time with your voice: a gentle, ambient signal that it is listening, without a modal stealing the screen. It reacts to your microphone level in real time, reads best against Deska’s dark themes, and fades out the moment the exchange ends.

The voice catalog

Every dictation is kept. The Voice panel is a timeline of everything you have said, grouped by day, with a word count on each entry. From any entry you can:

  • Copy the text to the clipboard.
  • Insert it into whatever field currently has focus.
  • Delete it when you no longer need it.

For quick reuse without opening the full panel, the mic button also reveals a small flyout of your most recent dictations with the same copy, insert and delete actions.

A managed feature

Transcription runs on Deska’s servers, so there is nothing to set up. It is the one metered feature, and it is measured in minutes: 30 a month on Free, 300 on a paid plan. Spoken replies draw on the same allowance.

NoteIf you use up the month’s minutes, dictation pauses until they reset or you upgrade. Nothing else stops, including your coding agents. See Plans & billing for details.

Voice on your own key

Every plan, Lifetime included, comes with a voice allowance. On a subscription, voice transcription deliberately stays on that managed allowance even when you have stored provider keys, so your own keys are never spent on it. Only a Lifetime license runs voice on your own key.

Heads upOn Lifetime the OpenAI key is required, not optional. Until you add one, voice prompts you to add your OpenAI API key in Settings and will not transcribe.

Voice on the web & phone

The same voice experience travels to the companion app on the web and your phone. Tap the mic, speak, and Dex answers, connected back to the machine where your project lives. It is the easiest way to check on a long task when you are away from your desk.

💡 Ideas+🐛 BugsSuggest a feature or report a bug