Do you actually need to remove them?
Some filler is natural and conversational — cutting every single one produces a clipped, slightly robotic rhythm that draws more attention than it removes. The real problem is density: a cluster of three or four in one sentence, or a nervous verbal tic that repeats throughout an episode. Judge by pattern, not by presence.
Editing them out manually
- Find the filler in the waveform and confirm the exact start and end of the sound.
- Cut on a zero-crossing where possible, to avoid an audible click at the edit point.
- Leave a natural-length pause behind rather than a hard splice — silence should sound like a breath, not an edit.
- Listen back at speed, not just at the cut, since rhythm problems only show up in context.
Transcript-based and AI-assisted removal
A newer generation of tools transcribes the episode and lets you delete filler words the same way you'd edit a text document, which is considerably faster than hunting through a waveform by ear. It's a genuinely useful step for long-form or interview-heavy shows — just treat it as a first pass, since automated transcripts occasionally misjudge a natural pause for a filler word worth cutting.
The order that actually works
Do content and filler-word editing first, on the raw recording — before noise reduction and mastering, not after. Cutting audio once it's already been mastered to a target loudness can leave small, audible jumps at each edit point. Editing first, then processing the finished cut as a whole, keeps the loudness and noise profile consistent across every edit.
Where Optivox fits
Optivox's job starts after your edit — noise cleanup, speaker balancing and mastering, run on the finished cut. Do your content and filler-word pass first, then let processing handle the rest in one consistent step.