Rehearse pitches, presentations, sales calls, and interviews in your browser. Get live feedback on filler words, words per minute, speaking ratio, longest pause, and volume consistency — with drill modes and a built-in teleprompter.
Your recording and audio measurements stay in this tab. If you turn on live transcription, your browser may send speech to its recognition service. VoxBoost does not receive or store that transcript.
Most people rehearse by repeating their pitch in front of a mirror and hoping for the best. VoxBoost AI gives you the same five metrics a speech coach would track — automatically, every take.
Free Practice, Call Opener (20s), Elevator (30s), Pitch (60s), Demo (2 min), or Presentation (5 min). Each mode auto-stops at the target length and adjusts WPM targets to match.
Paste talking points or a full script into the teleprompter. Set the scroll speed and it gently scrolls while you record — keep your eyes on the prompt, not the keyboard.
The big green button starts the take. Five metrics update live: fillers, words/min, speaking ratio, longest pause, volume consistency. The transcript scrolls underneath.
Stop saves the take to the list with all metrics, the full transcript, and audio playback. Compare up to 5 takes side-by-side, download the best one, delete the others.
Each metric tells you something different about how you sound. Together they paint a complete picture of your delivery without needing a coach in the room.
| Metric | What it measures | Good range | What it tells you |
|---|---|---|---|
| Filler words | Count of "um", "uh", "like", "you know", etc. detected in the transcript | 0–2 per minute | Drop below 2/min and you instantly sound more confident and prepared. The transcript highlights every instance so you know exactly when they crept in. |
| Words per minute | Total transcribed words ÷ minutes spoken | 140–170 WPM | Slower than 110 sounds tentative. Faster than 190 sounds rushed and is hard to follow over the phone. The "Speaking %" metric adjusts for pauses. |
| Speaking ratio | % of recording where your voice was above background level | 80–95% for monologue | For a pitch or presentation, you should be talking most of the time with short emphasis pauses. Long stretches under 80% usually mean you're stalling. |
| Longest pause | The single longest stretch of silence inside the take | under 2 seconds | Brief pauses (0.3–0.8s) punctuate ideas. Anything over 2–3 seconds feels like you've lost your place, even when you haven't — usually rehearse-able. |
| Volume consistency | 1–10 score based on how steady your loudness was | 7 or higher | Phone calls and podcasts need steady level. A low score means parts of your speech dropped quiet — listeners on the other end miss words. |
Each mode forces the constraint that real life will throw at you. Practice under the constraint and the real version feels easy.
No time limit. Warm up, work on accent and pronunciation, rehearse long material section by section. Best mode for first-time users.
The opening twenty seconds of a cold call decide whether you keep the conversation. Drill the same script ten times — by take five it'll feel natural.
Classic networking length. Who you are, what you do, why it matters, what you want. Stop at 30 seconds — discipline forces clarity.
The format used at most networking events, VC speed-dating, and "tell me about your company" moments. One full minute, no more.
The standard opening monologue length most VCs and enterprise buyers will give you before they start asking questions. Land it in two minutes.
Lightning talks, classroom sections, audition tapes, podcast monologues. Long enough to need pacing, short enough to memorize.
From bedroom rehearsals to sales-floor drills. The tool is the same — only the script changes.
Drill your opener until the first 20 seconds sound effortless. Pair with VoxBoost Rebuttals to also drill objection responses.
Rehearse your 30s, 60s, and 2-min pitches in the same session. Compare takes, ship the best one as a video script or audition tape.
Drill "Tell me about yourself" and "Why this company" until you stop saying "um" and start sounding like a hire. Listen back on the bus to the interview.
Track filler count and WPM across rehearsals. Watch your numbers improve week over week — the data makes the progress visible.
Group projects, dissertation defenses, speech class assignments. Practice into the camera roll, share the WebM with a study partner for feedback.
Set the language to the target, drill pronunciation. The transcript shows exactly which words the recognizer misheard — those are usually the ones to work on.
Rehearse your intro, ad reads, and segment transitions. The volume-consistency score helps you sound the same from minute one to minute thirty.
Test your hook lines and CTAs in 20-second drills. Find the version that flows best, then record the real take in your DAW.
Verifiers, intake calls, scripted disclosures. Steady pace and zero fillers translate into compliant calls that pass QA. Works great with Medicare Verifier.
Audition takes, dialect drills, demo reel prep. WAV export drops straight into ProTools, Reaper, or Adobe Audition for final mastering.
Share the page URL with new hires. They drill on their own, screenshot their best take's metrics, send to you. Faster than ride-alongs.
Best man speeches, eulogies, wedding toasts, conference keynotes. The metrics tell you when you've stopped reading and started speaking.
Most people stop practicing too early. A simple repeatable loop gets you to a great take faster than any amount of "let me just try it one more time".
Don't warm up. Don't look at the script. Hit record and just say it. This take will be ugly — that's the point. It sets your baseline.
Look at take 1's metrics. If fillers were high, slow down. If WPM was low, energize. If longest pause was 4 seconds, rehearse that part.
Now turn on the teleprompter at a comfortable scroll speed. Read the script — but keep eye contact with the camera position above the prompter.
Turn the teleprompter off. Hide your script. Say it from memory and intent, not from words. This is the take that sounds human.
Stand up. Smile while you speak. Imagine the actual person on the other end. This take usually beats take 1 by 50%+ on every metric.
Download the best take as WebM (to share) or WAV (to edit). Delete the others. You're done — total time around 10–15 minutes.
The opening twenty seconds of a cold call decide everything. They also happen to be the most rehearsable part of the call — same words every time. Use the Call Opener mode and follow this loop.
"Hi {Name}, this is {You} with {Company}. I know I'm calling out of the blue — quick reason for the call: we just helped {similar buyer} {specific outcome}. Worth a two-minute conversation to see if it could fit you?"
Run the Call Opener (20s) mode ten times in a row. By take five, the words drop and the meaning lands. By take ten, you can do it standing up, smiling, with your hands free of the script.
Once your opener is solid, drill the three most common objections in Free Practice mode. Use the VoxBoost Rebuttal Library for proven responses.
Every interview re-uses the same opening questions. Practice them once, win every interview for the rest of your career.
Use the 60-Second Pitch mode. Structure: where you are now → key strengths → why you're talking to them today. Drill until you can deliver it conversationally without sounding rehearsed.
Elevator Pitch (30s). Reference one specific thing from their site, their recent news, or their public roadmap. Generic answers tank you here.
Free Practice. Write a real weakness, the steps you're taking to fix it, and a concrete recent example. Drill until it sounds honest, not memorized.
Elevator Pitch (30s). Tie it to the role you're interviewing for without sounding entitled. Aim for filler count zero and pace 140-160.
The speech recognizer is, accidentally, an objective pronunciation coach. If it consistently mishears a word, that word probably needs work.
English (US), (UK), (AU), or (IN) — match the dialect you want to be understood in. The recognizer is trained on that variety and will be honest with you.
Use a passage you know — a news article, your LinkedIn bio, a poem. Compare the live transcript to the source. Mismatches reveal pronunciation issues.
If the recognizer consistently hears "sheep" as "ship" or "live" as "leave", drill those vowel pairs in Free Practice mode for 60 seconds at a time.
Each Stop saves the take with transcript. Over weeks, see your accuracy climb — and your filler count drop as you stop hesitating mid-sentence.
An honest comparison with the most popular speech-practice approaches.
| Feature | VoxBoost Practice | Voice Memos app | Speech-coach SaaS | Live coach |
|---|---|---|---|---|
| Free | Yes | Yes | $$ / mo | $$$ / hr |
| Filler-word count | Live | No | Yes | Yes |
| Words / minute | Live | No | Yes | Yes |
| Speaking ratio | Yes | No | Some | Yes |
| Longest-pause flag | Yes | No | Some | Yes |
| Volume consistency | Yes | No | Rare | Yes |
| Built-in teleprompter | Yes | No | Separate tool | N/A |
| Drill modes (30s / 60s / 2m) | Yes | No | Some | Yes |
| Multi-take comparison | 5 in-memory | Manual | Yes | In-session |
| Live transcript | Yes | No | Yes | No |
| Custom filler list | Yes | No | Rare | Verbal |
| WebM + WAV download | Yes | M4A only | Varies | N/A |
| Recording kept in the tab | Yes | Yes | Cloud | In person |
| No sign-up | Yes | Yes | Required | Required |
The questions we get asked most about how the tool works, where it works, and what the metrics mean.
It lets you rehearse out loud, listen back, and review delivery measurements such as speaking time, pauses, and volume consistency. If you choose to enable live transcription, it can also estimate words per minute and flag filler words. The recording stays in your browser tab.
Live transcription and filler-word counting use the Web Speech API, which is best supported in Chrome, Edge, Brave, and Opera on desktop and Android. Safari supports it on iOS and macOS in recent versions. Firefox has limited support, so on Firefox you'll still get audio-based metrics (pace estimate, speaking ratio, longest pause, volume consistency) but live transcript may not work.
The recording and audio measurements stay in your browser tab. If you turn on live transcription, your browser may send speech to its recognition provider; VoxBoost AI does not receive that stream. Leave transcription off if you want browser-only recording. You will still get speaking ratio, pause, and volume-consistency results, but filler words and words per minute will not be measured.
By default the tool flags: um, uh, ah, er, hmm, like, you know, sort of, kind of, basically, literally, actually, right, okay, so (when used as a sentence-opener filler). You can edit the list to add filler words specific to you — for example, some people overuse "obviously", "honestly", or "I mean".
For most professional speaking — sales calls, presentations, podcasts, conference talks — 140–170 words per minute is the sweet spot. Under 110 WPM sounds slow and tentative; over 190 WPM sounds rushed and is hard to follow, especially over the phone. Auctioneer-style speed is around 250+ WPM but is exhausting to listen to.
In a monologue practice (elevator pitch, presentation, audition tape), 85–95% speaking ratio is normal — most of the time should be talking, with short pauses for emphasis. In a discovery call or interview, 30–50% is ideal — you want the other person talking more than you. The tool measures only your audio so it's most useful for the monologue-style practice modes.
Brief pauses (0.3–0.8s) are good — they punctuate ideas and let the listener catch up. Pauses over 2–3 seconds in a pitch or presentation feel like you've lost your place, even if you haven't. The longest-pause metric helps you spot the one moment that breaks the flow so you can rehearse it specifically.
It compares the standard deviation of your loudness against your average loudness over the take. A score near 10 means you stay at roughly the same volume throughout — good for phone calls and podcasts. A lower score means parts of your speech were a lot quieter or louder than others, which can make you hard to hear on the parts that drop. Some variation is good (it shows energy); huge variation is not.
Paste any script into the teleprompter box. When you start recording, the text auto-scrolls at the pace you choose (slow / normal / fast). You can pause the scroll, change pace mid-take, or restart from the top. The teleprompter is purely visual — it doesn't read the script aloud.
Free Practice has no time limit. 30-Second Elevator Pitch forces you to land your message in 30 seconds. 60-Second Pitch is the standard pitch length used at most networking events. 2-Minute Demo is the length most VC firms allot for an initial pitch. Call Opener is 20 seconds — the window where most cold calls live or die. Each mode auto-stops at its time limit and tailors target metrics accordingly.
Yes. Each Stop creates a take in the takes list, with all of its metrics, the full transcript, and a playback bar. Keep up to 5 takes in memory so you can A/B different approaches before downloading the best one as WebM or WAV.
Yes. Set the speech-recognition language to the one you're learning. The transcript shows exactly which words it heard, which is great for accent feedback — if the recognizer mis-hears a word repeatedly, that word is likely pronounced unclearly. Pace and pause metrics help with rhythm.
That's exactly what the Call Opener and 60-Second Pitch modes are for. Paste your script into the teleprompter, set the scroll pace, and rehearse until your opener feels natural without the script. Pair with the VoxBoost AI Rebuttal Library to drill objection responses too.
Yes. Every take is downloadable as WebM (small, shareable) or WAV (uncompressed for editing). The transcript and metrics can also be copied as plain text from the results panel.
Accuracy varies with the browser, microphone, accent, room noise, language setting, and recognition provider. Filler counts and words-per-minute results are estimates, so verify important wording by listening to the recording.
Three possible reasons: your browser doesn't support the Web Speech API (try Chrome or Edge), you turned the live transcript off, or you genuinely didn't use any fillers in that take. The transcript box shows whatever the recognizer heard — if it's empty, recognition isn't running. If it has text but no count, you didn't say any flagged words.
Yes. Click "Edit fillers" under the live transcript, type your custom list as comma-separated words, and save. Common additions are "honestly", "obviously", "I mean", "right?", and any verbal tic specific to you (some people say "you know what I mean" constantly without realizing it).
The recorder, waveform, pace, ratio, pause, and volume metrics all work offline once the page is loaded. The live transcript needs an internet connection on most browsers because speech recognition is a cloud service. Turn the transcript off to record fully offline.
Pair the Practice Recorder with the rest of the VoxBoost AI suite for full pre-call workflow.