When the context decides the answer, a useful prompt is long. It carries the constraint, the approach you already tried, the exception that broke the last answer, and the reason the obvious fix does not apply here. You know all of it before you touch the keyboard. Typing it takes two minutes, so you type a third of it in twenty seconds and send that.
The model answers the third you sent.
Speaking looks like the obvious repair. Say the whole two minutes out loud, and the parts you would have trimmed for effort survive.
Then you read the transcript. It begins twice. There is an "um" before the important clause, a sentence you abandoned halfway, and a correction you made out loud. Pasted into a prompt box, that is not clearly better than the short version you would have typed.
So the question is not whether to speak the prompt. It is how much of what you said should reach the text field.
The part you delete is usually the part the model needed
Watch which words disappear when you shorten a prompt by hand. The constraint survives, because it feels like the request. The context goes first: the framework version, the thing you tried on Tuesday, the reason you cannot change the schema. Effort decides, and effort has no idea which sentence carries the answer.
TalkTalkType counts one held recording as one dictation, whatever it contains. Forty words and four hundred words cost the same. That does not make you a better writer of prompts. It removes the specific reason you were cutting.
Speech breaks in different places than typing does
Typed prompts fail by omission. Spoken prompts fail by residue.
A held take collects the false start, the filler before a hard word, the half-sentence you dropped when a better phrasing arrived. None of that is a transcription error. The words were spoken, so a faithful transcript keeps them.
That is why a dictation app for prompts is not finished when the recognition is accurate. The remaining decision is editorial, and it belongs to whoever knows where the text is going.
The choice is Clean or Raw, not a row of presets
TalkTalkType exposes that decision as a single choice with two positions: Clean and Raw. There is no picker of destinations and no rewrite preset.
Clean performs faithful light cleanup. It keeps every spoken word, in the order you said it, at the register and brevity you intended, and removes only unambiguous hesitation fillers where removing them cannot change meaning. It does not paraphrase, expand, summarize, or make anything more polite. A short phrase stays short. The result comes back as a single line.
Raw returns the transcript with nothing done to it. No cleanup model runs. The first thing you see is what you said.
The two cover different takes. Clean suits the take where your exact wording is the point, including the awkward phrasing you chose deliberately and the part you meant to leave short. Raw suits a take you intend to edit yourself, or one holding a name the cleanup step might smooth away.
Neither one rewrites your prompt for you, and that is deliberate. When you press the hotkey, the app captures the focused application and classifies the field under the cursor with a fixed lookup, not a guess. The cleaned result arrives as a draft, and the first paste is faithful. If you want to reshape it — turn a spoken ramble into a tighter instruction — the optional Voice Patch step lets you speak a refinement before the text is inserted, without ever silently rewriting that first paste. Context-aware dictation on Mac follows how the destination is read and where adaptation does and does not happen.
Every plan includes Clean and Raw. Plans differ in weekly dictations, voice time, the clean-input cap, and recording length, not in which modes you can run.
The two modes side by side:
| Mode | What it preserves | What it changes | When to use it |
|---|---|---|---|
| Clean | Every spoken word, order, register, and intended brevity | Removes only unambiguous hesitation fillers; does not paraphrase, expand, summarize, or make anything more polite | A take where your exact wording is the point, including awkward phrasing you chose deliberately |
| Raw | The whole transcript, with nothing done to it | Nothing; no cleanup model runs | A take you intend to edit yourself, or one holding a name the cleanup step might smooth away |
What the meter counts decides what you dare to say
Free includes 30 dictations and 15 minutes of voice per week, with each recording capped at 60 seconds. Plus is listed at $1 per month for 80 dictations and 30 minutes. Pro is $7 for 500 dictations, 250 minutes, and a 120-second cap. Max is $18 for 1,200 dictations and 600 minutes. Current figures are on the TalkTalkType pricing page.
Cleanup has its own weekly pool. Free allows 20 clean inputs a week and Plus 60; Pro and Max carry no separate clean cap. A Raw take spends none of it, because nothing runs on the text.
Two habits follow from that arithmetic, and they are not the ones a word count would produce.
Restating a sentence inside a take is free. Say the constraint, hear that it came out wrong, say it again properly, and let the cleanup step drop the abandoned half. Starting over with a second take is not free, so it is worth finishing a bad take rather than releasing the key.
The 60-second cap on Free and Plus is the real boundary, and it is one you can feel while speaking. A minute is a lot of prompt. It is not enough for a stream of consciousness, which is the correct amount of pressure.
The prompt box already has your cursor in it
Hold Option-Space, speak, release. The text arrives in the field that had focus, which for this work is the prompt box in Cursor, ChatGPT, or a terminal you were already typing into. Direct paste uses macOS Accessibility permission, and clipboard delivery stays available when a protected app refuses automated input. The app runs in the menu bar on macOS 26 or later.
That destination matters more here than it does for a short message. A prompt is written next to the thing it refers to. Moving to a separate window to dictate, then carrying the result back, reintroduces the interruption you were trying to remove. System-wide dictation on Mac covers how the focused-cursor path behaves across apps.
Nothing about this stores the prompt on a server. Audio, raw transcripts, and cleaned text are not kept by the service, and the Mac app's short session history lives in memory and disappears when the app quits. Private dictation on Mac follows one recording through every stage of that path.
Start with the three prompts you shortened this week
Find them in your history. They are the ones where the answer was almost right, and where your follow-up message supplied the context you had left out.
Speak each one again as a single take, with the constraint, the prior attempt, and the exception included. Send the result. Compare it to the answer you got the first time, and notice whether your follow-up message was still necessary.
You can download TalkTalkType for Mac and run that comparison inside the free weekly allowance. What it will not settle is the take you are still composing in your head, the one where you do not yet know what you are asking. Speaking that one is a different experiment, and it usually needs the answer to a smaller question first.