This is a breakdown of an actual edit. A 94-minute raw recording became a 51-minute episode. The subject was a business interview with a guest who spoke well but at length. The host was practiced. The recording setup was decent -- two local tracks from a remote session, both reasonably clean.
Here is what happened between the raw file and the published episode.
The raw file (00:00)
94 minutes. Pre-interview small talk at the start, about 7 minutes. Technical check ("can you hear me, how does my audio sound"). A false start where the host asked the first question too early, stopped, and reset. Then the actual interview.
The guest's track had moderate HVAC noise in the background -- a consistent hum, identifiable as heating. The host's track was clean. The guest was on a decent USB condenser mic but clearly not in a treated space.
Initial level check showed the guest's average level about 4 dB lower than the host's. Both were within normal range but inconsistent enough to be noticeable in a side-by-side comparison.
First pass: structural decisions (00:00 - 94:00)
Before any audio processing, the full recording gets a first listen at 1.5x speed. The goal is to identify structural issues: topics to cut, sections to reorder, moments where the conversation went somewhere that should not be in the final episode.
In this case: the pre-interview talk came out, obviously. The false start came out. There was a section around the 52-minute mark where the guest went on a tangent about a previous employer that was not relevant to the episode topic and that the host visibly tried to redirect twice. That came out -- about 8 minutes. There was also a question around the 70-minute mark that the guest answered mostly by restating something they had already covered in more detail at 35 minutes. The 70-minute version came out; the 35-minute version was better.
After structural cuts: down to about 68 minutes of content to work with.
Technical processing pass (various)
Noise reduction on the guest track first. The HVAC hum was a consistent frequency profile, which is the kind of noise that responds well to spectral noise reduction. Sampled the noise profile from a quiet section early in the recording, applied the reduction at a conservative setting. Checked the result on headphones: the hum was substantially reduced without the voice sounding processed. Pushed the reduction slightly higher, listened again. At the higher setting, the "s" sounds on the guest's voice started to sound metallic -- the noise reduction was eating into the sibilance range. Backed off to the first setting.
Level normalization on both tracks. Target was -16 LUFS integrated for each track independently before combining. The guest track needed about 4 dB of gain after normalization; the host track was already close to target. Applied compression on the guest track to bring up the dynamic range before the final normalization -- the guest had a wide dynamic range, dropping quieter during longer answers.
Track alignment. Both tracks started within about 0.4 seconds of each other, which is typical for a remote session where both participants start recording simultaneously. Applied sync using the early audio handclap that was recorded as a sync marker. The alignment was clean at the start. At the 60-minute mark, drift had accumulated to about 1.8 seconds. The algorithmic correction handled this smoothly.
Filler word pass (throughout)
The guest was a fluid speaker but used "so" as a sentence starter repeatedly -- not a traditional filler word, but functionally similar when it appears 40 times in an hour. Left most of them in. Removed the ones that appeared at the start of answers immediately after a question, where they were clearly just a verbal runway before the actual content started.
The host had a habit of a slightly extended "uh" before follow-up questions. Removed the most prominent ones -- five total. Left several others because removing all of them would have created unnatural transitions into questions.
Both tracks had long silences between some segments. Silences over about 1.5 seconds in the combined mix were tightened to 0.8 seconds.
Final structure assembly (various)
After the filler pass, the cleaned material was about 58 minutes. Another structural review at 1.5x speed to check that the cuts were not creating context gaps -- that something said at the 30-minute mark still made sense after the section before it had been tightened.
Found one issue: the host asked a question that referenced something the guest had said "earlier." The "earlier" section had been cut. The host's reference now pointed to nothing. Pulled back some of the earlier content (about 45 seconds) to restore the referent, then trimmed elsewhere to compensate.
Final runtime after all cuts: 51 minutes. That is about a 46% reduction from the raw 94 minutes.
Export
MP3 at 192 kbps. Chapter markers at the major topic transitions identified during structural review. True peak check before export: verified under -1 dBTP. Final LUFS check on the exported file: -16.1 LUFS integrated, -14.3 LUFS short-term peak. Within spec.
What this took
Total time from raw file to export-ready: about 3.5 hours for a 94-minute source. Structural review and assembly was the largest component, roughly half the total time. Technical processing was about an hour. Filler pass and final review made up the rest.
The AI-assisted components (noise reduction, level normalization, track alignment, filler detection) covered a significant portion of what would otherwise have been manual work. The structural decisions -- what to cut, what to keep, where the tangent crossed the line from interesting into off-topic -- remained judgment calls throughout.