A Producer’s Checklist for Recording Clean AI Voice Training Data

An AI voice model can only be as good as the audio it learns from. Give it a clean, controlled recording and you get a model that works in a cinema mix. Give it a session with an air conditioner humming underneath, and you may end up with a voice model that sounds like it has strong opinions about indoor climate control.
Here’s a pragmatic session guide for engineers preparing AI voice cloning: the physical setup, file specs, and recording approaches that help you build a clean voice clone. Providers can tweak algorithms all day, but recording training data for AI voice is what truly calls the shots.
Key Takeaways
-
The model learns the whole recording. That includes the performer, the room, the air conditioner, and anything else that made it onto the track.
-
Keep master files raw and uncompressed. Deliver WAV or FLAC files with healthy headroom (peaking -12 to -6 dBFS). Skip the EQ, compression, and noise gates.
-
Range and consistency make the difference. Record varied material, match the setup across sessions, and give the model enough of the performer’s real voice to learn from.
Why Recording Quality Defines AI Voice Quality
Your brain is remarkably good at gaslighting you into thinking a room sounds fine. The AI algorithm, unfortunately, has zero capacity for self-deception.
The training software analyzes exact frequency patterns and harmonic structure. Your ear does something your model can't: it filters out echo and low background noise automatically, so you stop noticing them. The model doesn't filter anything. It treats every detail in the signal as information about the voice, including the artifacts you'd rather it ignored.
Respeecher has built convincing models from difficult archival material, including Tommy Muñiz, Wilt Chamberlain, and Endurance. We have worked with old tape, mono, tight bandwidth, and a fragile wax-cylinder material more than a century old. But better voice training data quality means a shorter model build and more useful material for the engineers.
Checklist Part 1: Room and Acoustic Environment
Before the session starts, give the room a quick reality check. You do not need a booth that looks like a spacecraft. You do need one that does not sound like a bathroom.
Microsoft's guidance for custom voice recording is similarly simple: use a small space with no noticeable echo or room tone, and soften reflective surfaces where needed.
Before you record, run through these voice cloning recording requirements:
- Choose a small room. Avoid large, empty spaces with tile, glass, bare walls, or hard floors. If you can hear a clear echo, the model will hear it too.
- Soften the reflections. Acoustic panels are great, but heavy curtains or a rack full of coats will do the job in a pinch. It might look absurd on set, but a ruined training run looks way worse.
- Turn off the noise. Air conditioning, fans, fridges, buzzing lights, computers, and open windows all count. The model will not decide the HVAC was "just ambience."
-
Run the 30-second silence test. Record 30 seconds with nobody speaking.
- Listen back on closed-back headphones. Any hum, hiss, traffic, or vibration you hear will sit underneath the voice too.
- Set the microphone distance. Park the performer roughly a foot from the capsule. It is a small adjustment, but it can make a noticeable difference to the takes you keep.
- Lock the setup before the first take. Mark the mic position and performer placement. If the session continues on another day, return to the same room, distance, and setup.
Those last two points are the practical core of most voice cloning recording requirements. They take a few minutes to set up and can save a lot of technical repair later.
Checklist Part 2: Microphone and Equipment
The room is under control. Now make sure the microphone hears the performer and set up the recording chain:
- Use a cardioid studio microphone. A cardioid pickup pattern prioritises the voice in front of it and reduces what arrives from the sides. A studio condenser is a dependable default; Sennheiser, AKG, Shure SM7B, and quality USB condensers can all do the job. This is the microphone for AI voice training. A laptop mic is for meetings. A Bluetooth headset is for making meetings worse.
- Put a pop filter in front of it. Plosives (the bursts of air in sounds such as “p,” “b,” and “t”) can overload the signal. A pop filter handles them before they become a waveform problem.
- Match the connection to the mic. An XLR microphone runs through an audio interface with a clean preamp. A USB condenser connects directly to the recording computer. In a nutshell, keep the chain short.
- Use a stable stand or boom. Handheld recording changes the distance from take to take, which changes the sound. Use a stand, set the height, and leave it alone unless the session needs a reset.
-
Monitor every take on closed-back headphones. The 30-second silence test confirmed that the room was usable. Listen for clicks, hum, traffic, clipping, and any new noise that appears once the session starts.
Checklist Part 3: Recording Format and Levels
The room is quiet. The microphone is behaving. Now do not lose the take in the export settings.Set the recording specs:
- Record uncompressed. For an AI voice cloning audio format, track your masters in WAV or FLAC and save MP3s for casual previews. At 24-bit / 48 kHz the training system gets enough detail to map the real vocal character. Compress it first and you feed the model the artifacts instead.
- Leave headroom. Keep peaks between -12 and -6 dBFS. That gives you a healthy signal without pushing against the digital ceiling. If the meter hits red, stop and reset. A clipped take is not improved by everyone agreeing it was the best one.
- Keep the files raw and dry. No EQ, compression, reverb, or noise reduction before the material reaches the model. Those tools alter the frequency structure of the voice. Save the polished version for later; the technical team needs the unvarnished one.
- Plan for pitch calibration. For speech-to-speech work, calibration helps the system identify the source performer’s average pitch and adjust conversion more accurately. Record a separate calibration take if the technical team requests one. Recording clean audio for AI is the job; calibration helps the model make better use of it.
Not sure whether your material will work?
Respeecher’s team of 15+ sound professionals can review what you already have and tell you whether it is usable or whether another session would save time later. Talk to us about your project.
Checklist Part 4: What to Record and How
A model cannot give you a performance range it never heard. More audio helps, but the right audio helps more.
Plan the session material:
- Set the right volume target. If you're figuring out how to record voice for AI cloning for a full text-to-speech voice, plan for at least one hour of clean, varied audio. Speech-to-speech conversion can work from less, often 20–30 minutes of source performance. De-aging, multilingual work, singing, or unusual character requirements may need more.
- Give it different kinds of language. Include dialogue, monologue, technical copy, narration, short responses, and longer lines. An hour of one neutral script can make for a capable voice model with exactly one gear.
- Record the emotional range. Capture neutral, upbeat, serious, intimate, irritated, and higher-energy delivery where the final production needs it. This is especially important in speech-to-speech projects, where the source actor’s emotion needs to survive the conversion.
- Don't over-act, just talk naturally. Keep a comfortable pace and take extra care with names, figures, and odd jargon. The algorithm swallows every single habit whole, so those lazy sentence endings you hoped to gloss over will end up in the final clone.
- Take alternate passes on hard material. Record several versions of difficult words, uncommon phonemes, accents, invented names, and emotionally demanding lines. It gives the technical team options, instead of one heroic take with a mystery noise halfway through.
- Match conditions across sessions. If the recording happens over several days, use the same room, microphone, distance, gain settings, and direction. This is the practical core of any AI voice training data recording plan.
Quick Reference: Recording Checklist
|
Parameter |
Recommended |
Avoid |
|
Room |
Small, quiet room with no obvious echo; soften hard surfaces |
Large empty rooms, tile, glass, bare walls |
|
Microphone |
Cardioid studio mic, pop filter, stable stand |
Laptop mic, Bluetooth headset, handheld recording |
|
Format |
WAV or FLAC masters; 24-bit / 44.1–48 kHz where possible |
MP3 or other lossy master files |
|
Peak level |
Peaks around -12 to -6 dBFS |
Clipping at 0 dBFS or recording so quietly that noise takes over |
|
Processing |
Raw, dry files with no effects |
EQ, compression, reverb, or pre-applied noise reduction |
|
Recording volume |
About one hour of varied material for TTS; often 20–30 minutes for STS |
Short, repetitive clips with limited vocal range |
|
Content |
Dialogue, narration, technical language, varied emotion and pacing |
One neutral read in a single style |
What Happens After the Recording Session
Once the session wraps, the AI voice training data recording moves from the booth to the technical workflow. Here is what the team checks before a voice model is ready for use:
Engineering Audit
The raw files go to the engineering team. They run an audit to check for weak spots and evaluate whether your reference audio for voice cloning is truly solid, or if one quick retake now will save days of headaches later.
Pitch Calibration
Speech-to-speech projects need pitch calibration to align the source actor's range during conversion, though TTS workflows bypass this step. If required, you'll get brief instructions for a quick calibration take (the Respeecher API documentation breaks down the technical specs if you want a closer look).
Model Training
Once the audio passes inspection, the team builds the actual model. Timeline depends on the dataset size and project complexity, so just ask for a realistic ETA up front.
Talent Review & Final Sign-Off
The performer or their team gets to review the voice samples before anything is approved for use. That step is deliberate. In line with how Respeecher works with talent, consent stays active throughout the entire project rather than ending when the contract is signed.
Final thoughts
An AI voice session is not ordinary ADR, and it is not a standard VO date with a few extra file exports. You are recording the material a model will learn from. That means the room, the mic, the levels, and the script all affect the result.
When those basics are right, the model build starts clean. But when they aren't, someone eventually has to explain why the final voice comes with a faint but persistent interest in ventilation.
The checklist gives the session a solid foundation: a stable setup, clean source material, and enough performance range for the work ahead. Good recording training data for AI voice saves time, reduces avoidable re-records, and gives the technical team a better starting point.
Respeecher’s sound professionals support projects from session planning through final delivery. If you have a recording day ahead, see how we work in Film & TV production.
FAQ
A cardioid studio condenser is a sensible place to start. It focuses on the performer in front of the mic and picks up less of the room from the sides.
Sennheiser, AKG, a Shure SM7B, or a good USB condenser can all work. The exact model matters less than a quiet room, a pop filter, and a setup that stays put. A laptop mic can technically record a voice. So can a phone in a glass elevator.
Yes. A quiet room with soft furnishings, a cardioid mic, and a pop filter can get you more than enough of usable material.
But before the performer starts, record 30 seconds of your recording space and listen back on headphones. You may hear a fan, a fridge, traffic, or a mystery buzz that was left unnoticed. Better to meet it then.
Usually, no. Send the raw, dry files unless the technical team asks for something else. EQ, compression, reverb, and noise reduction can all change details the model needs to hear.
If the fan is audible, turn off the fan. That is rarely the most exciting part of a recording session, but it does work.
Keep the master recordings in WAV or FLAC. Aim for at least 16-bit / 44.1 kHz, and use 24-bit at 44.1 or 48 kHz if the session setup allows it.
MP3 is fine for a quick listen or a file that needs to travel fast, but it is not the format you want the model learning from. Even platforms like Microsoft specify uncompressed WAV files across 16-, 24-, or 32-bit depths for voice cloning workflows.
So don’t write off older, noisy, or imperfect recordings before someone has heard them. Send over the reference audio for voice cloning and the team can assess it with you: what is usable, what may need a few extra takes, and whether a fresh session would make a meaningful difference.
One carefully prepared session can often cover it. For a full text-to-speech voice, plan for about an hour of clean, varied audio. Speech-to-speech projects may need less, depending on what the final production asks the voice to do.
If your session spreads across several days, remember that a massive part of how to record voice for AI cloning is keeping the environment stable. Use the same mic, lock in the performer's distance, and snap a quick photo of your gain knob so nothing drifts.
Glossary
Reference audio / training data
The recorded voice material an AI model analyzes to learn a speaker's unique tone, cadence, and acoustic habits.
dBFS (decibels full scale)
Cardioid microphone
A directional microphone pattern that captures audio directly in front of the capsule while rejecting unwanted room reflections from the sides and back.
Dry recording
Unprocessed audio recorded without EQ, compression, or reverb, giving the model an untouched baseline of the performer’s real voice.
Room tone
The baseline ambient noise of an empty recording space that an AI model will inadvertently clone if not controlled.



