Sep 3, 2026, 5:26:24 AM • 8 min

A Producer’s Checklist for Recording Clean AI Voice Training Data

•••

An AI voice model can only be as good as the audio it learns from. Give it a clean, controlled recording and you get a model that works in a cinema mix. Give it a session with an air conditioner humming underneath, and you may end up with a voice model that sounds like it has strong opinions about indoor climate control.

Here’s a pragmatic session guide for engineers preparing AI voice cloning: the physical setup, file specs, and recording approaches that help you build a clean voice clone. Providers can tweak algorithms all day, but recording training data for AI voice is what truly calls the shots.

Key Takeaways

  • The model learns the whole recording. That includes the performer, the room, the air conditioner, and anything else that made it onto the track.

  • Keep master files raw and uncompressed. Deliver WAV or FLAC files with healthy headroom (peaking -12 to -6 dBFS). Skip the EQ, compression, and noise gates.

  • Range and consistency make the difference. Record varied material, match the setup across sessions, and give the model enough of the performer’s real voice to learn from.

Why Recording Quality Defines AI Voice Quality

Your brain is remarkably good at gaslighting you into thinking a room sounds fine. The AI algorithm, unfortunately, has zero capacity for self-deception.

The training software analyzes exact frequency patterns and harmonic structure. Your ear does something your model can't: it filters out echo and low background noise automatically, so you stop noticing them. The model doesn't filter anything. It treats every detail in the signal as information about the voice, including the artifacts you'd rather it ignored. 

Respeecher has built convincing models from difficult archival material, including Tommy Muñiz, Wilt Chamberlain, and Endurance. We have worked with old tape, mono, tight bandwidth, and a fragile wax-cylinder material more than a century old. But better voice training data quality means a shorter model build and more useful material for the engineers. 

Checklist Part 1: Room and Acoustic Environment

Before the session starts, give the room a quick reality check. You do not need a booth that looks like a spacecraft. You do need one that does not sound like a bathroom.

Microsoft's guidance for custom voice recording is similarly simple: use a small space with no noticeable echo or room tone, and soften reflective surfaces where needed.

Before you record, run through these voice cloning recording requirements:

  • Choose a small room. Avoid large, empty spaces with tile, glass, bare walls, or hard floors. If you can hear a clear echo, the model will hear it too.
  • Soften the reflections. Acoustic panels are great, but heavy curtains or a rack full of coats will do the job in a pinch. It might look absurd on set, but a ruined training run looks way worse.
  • Turn off the noise. Air conditioning, fans, fridges, buzzing lights, computers, and open windows all count. The model will not decide the HVAC was "just ambience."
  • Run the 30-second silence test. Record 30 seconds with nobody speaking.

  • Listen back on closed-back headphones. Any hum, hiss, traffic, or vibration you hear will sit underneath the voice too.
  • Set the microphone distance. Park the performer roughly a foot from the capsule. It is a small adjustment, but it can make a noticeable difference to the takes you keep.
  • Lock the setup before the first take. Mark the mic position and performer placement. If the session continues on another day, return to the same room, distance, and setup.

Those last two points are the practical core of most voice cloning recording requirements. They take a few minutes to set up and can save a lot of technical repair later.

Checklist Part 2: Microphone and Equipment

The room is under control. Now make sure the microphone hears the performer and set up the recording chain:

  • Use a cardioid studio microphone. A cardioid pickup pattern prioritises the voice in front of it and reduces what arrives from the sides. A studio condenser is a dependable default; Sennheiser, AKG, Shure SM7B, and quality USB condensers can all do the job. This is the microphone for AI voice training. A laptop mic is for meetings. A Bluetooth headset is for making meetings worse.
  • Put a pop filter in front of it. Plosives (the bursts of air in sounds such as “p,” “b,” and “t”) can overload the signal. A pop filter handles them before they become a waveform problem.
  • Match the connection to the mic. An XLR microphone runs through an audio interface with a clean preamp. A USB condenser connects directly to the recording computer. In a nutshell, keep the chain short.
  • Use a stable stand or boom. Handheld recording changes the distance from take to take, which changes the sound. Use a stand, set the height, and leave it alone unless the session needs a reset.
  • Monitor every take on closed-back headphones. The 30-second silence test confirmed that the room was usable. Listen for clicks, hum, traffic, clipping, and any new noise that appears once the session starts.

Checklist Part 3: Recording Format and Levels

The room is quiet. The microphone is behaving. Now do not lose the take in the export settings.

Set the recording specs:

  • Record uncompressed. For an AI voice cloning audio format, track your masters in WAV or FLAC and save MP3s for casual previews. At 24-bit / 48 kHz the training system gets enough detail to map the real vocal character. Compress it first and you feed the model the artifacts instead.
  • Leave headroom. Keep peaks between -12 and -6 dBFS. That gives you a healthy signal without pushing against the digital ceiling. If the meter hits red, stop and reset. A clipped take is not improved by everyone agreeing it was the best one.
  • Keep the files raw and dry. No EQ, compression, reverb, or noise reduction before the material reaches the model. Those tools alter the frequency structure of the voice. Save the polished version for later; the technical team needs the unvarnished one.
  • Plan for pitch calibration. For speech-to-speech work, calibration helps the system identify the source performer’s average pitch and adjust conversion more accurately. Record a separate calibration take if the technical team requests one. Recording clean audio for AI is the job; calibration helps the model make better use of it.

Not sure whether your material will work?
Respeecher’s team of 15+ sound professionals can review what you already have and tell you whether it is usable or whether another session would save time later. Talk to us about your project.


Checklist Part 4: What to Record and How

A model cannot give you a performance range it never heard. More audio helps, but the right audio helps more.

Plan the session material:

  • Set the right volume target. If you're figuring out how to record voice for AI cloning for a full text-to-speech voice, plan for at least one hour of clean, varied audio. Speech-to-speech conversion can work from less, often 20–30 minutes of source performance. De-aging, multilingual work, singing, or unusual character requirements may need more.
  • Give it different kinds of language. Include dialogue, monologue, technical copy, narration, short responses, and longer lines. An hour of one neutral script can make for a capable voice model with exactly one gear.
  • Record the emotional range. Capture neutral, upbeat, serious, intimate, irritated, and higher-energy delivery where the final production needs it. This is especially important in speech-to-speech projects, where the source actor’s emotion needs to survive the conversion.
  • Don't over-act, just talk naturally. Keep a comfortable pace and take extra care with names, figures, and odd jargon. The algorithm swallows every single habit whole, so those lazy sentence endings you hoped to gloss over will end up in the final clone.
  • Take alternate passes on hard material. Record several versions of difficult words, uncommon phonemes, accents, invented names, and emotionally demanding lines. It gives the technical team options, instead of one heroic take with a mystery noise halfway through.
  • Match conditions across sessions. If the recording happens over several days, use the same room, microphone, distance, gain settings, and direction. This is the practical core of any AI voice training data recording plan.

Quick Reference: Recording Checklist

Parameter

Recommended

Avoid

Room

Small, quiet room with no obvious echo; soften hard surfaces

Large empty rooms, tile, glass, bare walls

Microphone

Cardioid studio mic, pop filter, stable stand

Laptop mic, Bluetooth headset, handheld recording

Format

WAV or FLAC masters; 24-bit / 44.1–48 kHz where possible

MP3 or other lossy master files

Peak level

Peaks around -12 to -6 dBFS

Clipping at 0 dBFS or recording so quietly that noise takes over

Processing

Raw, dry files with no effects

EQ, compression, reverb, or pre-applied noise reduction

Recording volume

About one hour of varied material for TTS; often 20–30 minutes for STS

Short, repetitive clips with limited vocal range

Content

Dialogue, narration, technical language, varied emotion and pacing

One neutral read in a single style

What Happens After the Recording Session

Once the session wraps, the AI voice training data recording moves from the booth to the technical workflow. Here is what the team checks before a voice model is ready for use:

Engineering Audit

The raw files go to the engineering team. They run an audit to check for weak spots and evaluate whether your reference audio for voice cloning is truly solid, or if one quick retake now will save days of headaches later.

Pitch Calibration

Speech-to-speech projects need pitch calibration to align the source actor's range during conversion, though TTS workflows bypass this step. If required, you'll get brief instructions for a quick calibration take (the Respeecher API documentation breaks down the technical specs if you want a closer look).

Model Training

Once the audio passes inspection, the team builds the actual model. Timeline depends on the dataset size and project complexity, so just ask for a realistic ETA up front.

Talent Review & Final Sign-Off

The performer or their team gets to review the voice samples before anything is approved for use. That step is deliberate. In line with how Respeecher works with talent, consent stays active throughout the entire project rather than ending when the contract is signed.

Final thoughts

An AI voice session is not ordinary ADR, and it is not a standard VO date with a few extra file exports. You are recording the material a model will learn from. That means the room, the mic, the levels, and the script all affect the result.

When those basics are right, the model build starts clean. But when they aren't, someone eventually has to explain why the final voice comes with a faint but persistent interest in ventilation.

The checklist gives the session a solid foundation: a stable setup, clean source material, and enough performance range for the work ahead. Good recording training data for AI voice saves time, reduces avoidable re-records, and gives the technical team a better starting point.

Respeecher’s sound professionals support projects from session planning through final delivery. If you have a recording day ahead, see how we work in Film & TV production.image1

FAQ

A cardioid studio condenser is a sensible place to start. It focuses on the performer in front of the mic and picks up less of the room from the sides.

Sennheiser, AKG, a Shure SM7B, or a good USB condenser can all work. The exact model matters less than a quiet room, a pop filter, and a setup that stays put. A laptop mic can technically record a voice. So can a phone in a glass elevator.



Yes. A quiet room with soft furnishings, a cardioid mic, and a pop filter can get you more than enough of usable material.

But before the performer starts, record 30 seconds of your recording space and listen back on headphones. You may hear a fan, a fridge, traffic, or a mystery buzz that was left unnoticed. Better to meet it then.



 

Usually, no. Send the raw, dry files unless the technical team asks for something else. EQ, compression, reverb, and noise reduction can all change details the model needs to hear.

If the fan is audible, turn off the fan. That is rarely the most exciting part of a recording session, but it does work.



 

Keep the master recordings in WAV or FLAC. Aim for at least 16-bit / 44.1 kHz, and use 24-bit at 44.1 or 48 kHz if the session setup allows it.

MP3 is fine for a quick listen or a file that needs to travel fast, but it is not the format you want the model learning from. Even platforms like Microsoft specify uncompressed WAV files across 16-, 24-, or 32-bit depths for voice cloning workflows.



 

It depends on what you have and what the project needs. Respeecher has worked with difficult archival material when a new studio session was simply not an option, including Tommy Muñiz’s recordings for Los García and Wilt Chamberlain’s voice for Goliath.

 

So don’t write off older, noisy, or imperfect recordings before someone has heard them. Send over the reference audio for voice cloning and the team can assess it with you: what is usable, what may need a few extra takes, and whether a fresh session would make a meaningful difference.

One carefully prepared session can often cover it. For a full text-to-speech voice, plan for about an hour of clean, varied audio. Speech-to-speech projects may need less, depending on what the final production asks the voice to do.

If your session spreads across several days, remember that a massive part of how to record voice for AI cloning is keeping the environment stable. Use the same mic, lock in the performer's distance, and snap a quick photo of your gain knob so nothing drifts.

Glossary

Reference audio / training data

The recorded voice material an AI model analyzes to learn a speaker's unique tone, cadence, and acoustic habits.

dBFS (decibels full scale)

The digital audio meter scale where 0 dBFS represents the absolute maximum volume limit before severe digital distortion occurs.

Cardioid microphone

A directional microphone pattern that captures audio directly in front of the capsule while rejecting unwanted room reflections from the sides and back.

Dry recording

Unprocessed audio recorded without EQ, compression, or reverb, giving the model an untouched baseline of the performer’s real voice.

Room tone

The baseline ambient noise of an empty recording space that an AI model will inadvertently clone if not controlled.

Pitch calibration

A quick technical alignment pass used by Respeecher to fine-tune pitch accuracy, particularly for speech-to-speech voice conversion.

Clipping

Irreversible digital distortion that occurs when an audio signal exceeds 0 dBFS and permanently chops off the waveform peaks.
Previous Article
The Engineering Decisions Behind a Good Voice Agent
Clients:
Lucasfilm
Blumhouse productions
AloeBlacc
Calm
Deezer
Sony Interactive Entertainment
Edward Jones
Ylen
Iliad
Warner music France
Religion of sports
Digital domain
CMG Worldwide
Doyle Dane Bernbach
droga5
Sim Graphics
Veritone

Recommended Articles