Imagera AI - AI content creation platform for generating images, cloning voices, creating avatars, and enhancing videos. Privacy Policy | Terms

Product Guide

How to Add an AI Voiceover That Fits Your Video

Illustration: a night-time editing desk with a condenser microphone, a script page with pencil timing marks, and a monitor showing kettle footage above a tall voice waveform and a quieter music waveform

Generate natural AI voices and narrations in seconds.

TL;DR

Write to the clip: about 65 words for 30 seconds if you want room to breathe, 75 to fill it. Generate the read with a stock voice in the 5-Engine Voice Studio or in your own cloned voice, and respell names the voice gets wrong. In our test, 67 words ran 26.0 and 27.2 seconds in two stock voices, 15 credits each. Mix in the Video Music Mixer: one pass for voice over the clip's own sound, or two passes (a 25% music bed, then the narration at 200%) for music under the voice. Each 30-second mixer pass is 30 credits.

Key takeaways

  1. A 67-word, 388-character script ran 26.0 s and 27.2 s in two stock voices (generated 2026-09-22).
  2. The two stock reads cost 15 credits each and came out 4.8 LU apart in loudness.
  3. A stock voice costs 15 credits for up to 1,388 characters; a cloned voice costs 10 credits per started 500 characters, 15 minimum.
  4. Each Video Music Mixer pass on a 30-second clip costs 30 credits.
  5. The mixer halves overlapping tracks, so narration at 200% over a 25% bed keeps the voice at its own level with an 18 dB gain gap to the bed.

To add an AI voiceover to a video, write the script to the clip's length first: about 65 to 75 words for 30 seconds. Generate the read with a stock voice in the 5-Engine Voice Studio, or in your own voice with the AI voice generator. Respell any name the voice gets wrong, then lay the audio under the footage in the Video Music Mixer. In our test, a 67-word script came back at 26.0 and 27.2 seconds in two stock voices, for 15 credits each.

A voiceover usually goes wrong in one of three ways. The read runs 38 seconds on a 30-second cut, the product name comes out wrong in the line that matters, or the music buries the narration. Trying another voice fixes none of them; the script and the mix do.

How many words fit your clip?

Voiceover timing guides put narration at 130–150 words per minute, commercial reads at 150–170 and explainers at 140–160 (GoTeleprompter, May 27, 2026). Plan on about 2.5 words per second to fill a clip, and about 2.2 if you want a beat of silence before the end card.

Clip lengthWords with room to breathe (130 wpm)Words to fill it (150 wpm, our tested pace)Approx. charactersCredits, stock voiceCredits, your own voice
30 seconds~65~75375–4351515
60 seconds~130~150750–8701520

Character counts assume 5.8 characters per word, the ratio of the sample script below. A stock voice costs 15 credits for up to 1,388 characters, about a minute and a half of speech. A cloned voice costs 10 credits per started 500 characters, with a 15-credit minimum.

Neither studio has a speed or pitch control, so only the words set the length. If the read runs long, cut words: at 2.5 words per second, five words is about two seconds.

Illustration: overhead view of a printed script split into five pencil-bracketed blocks beside a stopwatch and a tablet showing a kettle pouring into a coffee dripper

Illustration: marking each block of the script against the shot it has to cover.

A 30-second script, marked against the picture

Oskelo is an invented pour-over kettle. The script below is written for a 30-second product clip, one line per shot:

TimeShotLineWords
0:00–0:04Kettle on the counter, steam risingPour-over coffee goes wrong in the first ten seconds.9
0:04–0:11Close-up of grounds under a fast pourWater too hot scorches the grounds. Too cool, and the cup tastes flat. Most kettles leave you guessing.18
0:11–0:21The slow hero pourOskelo holds your water at the temperature you choose, and its narrow spout pours a slow, steady stream you can actually control.22
0:21–0:26Morning routineSet it once in the morning. Every cup after that tastes the same.13
0:26–0:30End cardOskelo. Better coffee, on purpose.5

That is 67 words and 388 characters. At 130 to 150 words per minute it should read in 27 to 31 seconds. The real reads took 26.0 and 27.2 seconds, at or below the fast end of that range, which is why the generated file's length is the only number to cut to.

How we tested two stock voices

  • Date and tool: 22 September 2026, the 5-Engine Voice Studio's default engine.
  • Input: the 67-word, 388-character Oskelo script above, read by two of the ten preset voices with default settings and no cues in the text.
  • Measured: each file's length, the pace in words per minute over that length, and its integrated loudness in LUFS.
  • Credits: 15 per read.

What the two reads delivered:

VoiceLengthPaceLoudnessCredits
Sarah, "Calm, reassuring female (US)"26.0 s154 wpm−13.8 LUFS15
Eric, "Smooth, trustworthy male (US)"27.2 s148 wpm−18.6 LUFS15
Real output: the 67-word Oskelo script in the stock voice Sarah, 26.0 seconds
Real output: the same script in the stock voice Eric, 27.2 seconds

The same words finished 1.2 seconds apart, so time the file before you cut the picture to it. The files also came out almost 5 dB apart in loudness (LUFS measures perceived loudness; higher is louder), so a mix setting that suits one voice may not suit the other.

Stock voice or your own: choose the route

Stock voiceYour own (cloned) voice
Where5-Engine Voice Studio, default engineAI voice generator
You provideThe scriptThe script plus a dry 3–10 second recording
Voices10 preset voices, each with a one-line descriptionThe person in your recording
Delivery controlsLanguage (Auto plus 11); cues in the textEmotion: the speaker's own, random, or custom sliders for joy, calm and six more
OutputMP3WAV
Credits15 up to 1,388 characters, 65 at 5,00010 per started 500 characters, 15 minimum, up to 5,000
Use it whenYou want a neutral narrator you can reuseThe brand already has a voice, or the channel is yours

Stock voice. Open the 5-Engine Voice Studio, keep the default engine, paste the script and pick a voice. Full stops and commas set the pauses, and the studio documents square-bracket cues such as [whispers] and [sighs]. If none of the ten voices fits, the studio's four other engines each have their own sound and price.

Your own voice. Record 3 to 10 seconds of one person speaking in a dry room, with no music, no echo and no second voice. A cupboard full of coats absorbs more echo than a bare kitchen. Emotion is set separately from the recording, so the same voice can read a calm intro and an energetic end card. There is no language setting, so treat non-English scripts as untested. Clone only your own voice or one you have written permission to use.

Illustration: a woman recording a short voice sample into a phone inside a closet lined with coats and a quilted blanket

Illustration: a dry, single-speaker reference recording is what cloning needs.

A voice you build from a written description in Voice Design does not appear in either studio; saved voices work in the Voice Changer.

Fix names before you pay for the full read

Neither studio has a pronunciation dictionary, so you fix pronunciation in the text. A 30-second script costs the same 15 credits as a single test line on either route, so generate the whole read and listen. For longer scripts, run the risky lines on their own first.

WrittenWhat to type for the voiceWhy
OskeloOss-keh-loHyphens split the syllables. For stress, try OSS-keh-lo, but listen: some voices spell capitals out
v2.1version two point oneNumbers and symbols read unpredictably
70–100°Cseventy to a hundred degreesRanges and units need words

Keep the respelled script for the voice and the correctly spelled one for captions, and put every respelling in a short glossary so the next video says the name the same way.

Mix the narration under the footage

The Video Music Mixer takes one video and one audio track and returns one MP4. It has two controls:

  • Track volume: 25, 50, 75, 100, 150 or 200 percent, applied only to the track you add.
  • Replace the original sound: on drops the clip's own audio; off mixes the new track over it.

The mixer never makes the video longer. Audio that runs past the last frame is cut, and the added track always starts on the first frame, with no offset control. When it replaces the sound, or when the clip has no sound of its own, the export ends where the shorter file ends. Sarah's 26-second read on a silent 30-second clip would come back as a 26-second video.

Illustration: a hand on a small mixer fader in front of a laptop showing kettle frames above a tall voice waveform and a much quieter music waveform

Illustration, not the mixer's actual screen: the voice should sit well above the music bed.

Choose the recipe that matches your clip:

Your clipPassesSettingsCredits (30 s clip)
Has sound worth keeping, such as room tone or product sounds1Video + narration, Replace off, Track volume 200%30
Silent, and you want voice only1Video + narration that runs to the last frame, Track volume 100%30
Needs music under the voice2Pass 1: video + music bed, Replace on, Track volume 25%. Pass 2: pass 1 result + narration, Replace off, Track volume 200%60

The two-pass order matters. Pass 1 replaces the sound, so the music bed must be at least as long as the clip; Music Factory can make an instrumental one. Pass 2 mixes instead of replacing, so a narration that finishes early leaves the bed playing rather than cutting the video short. You can pick the pass 1 export and your narration from the Generations tab in the upload window, with no download and re-upload.

What the level settings really do

The mixer changes gain; it does not measure or match loudness. It also halves both tracks while they overlap, then brings the remaining one back up once the shorter track ends:

  • Overlapping tracks are halved. In pass 2, narration at 200% plays at its original level while the bed plays at one-eighth of the music file's level (25%, then halved), a gain gap of about 18 dB. At 150%, the voice drops to three-quarters of its level and the gap is about 16 dB. In the one-pass recipe, the clip's own sound drops to half while the narration plays.
  • The gap is a gain difference, not a guarantee. A common rule of thumb keeps music at 10–20% of the speaking volume (Recorded, July 17, 2026). These settings land there only if the music and narration files start at similar loudness, and our two stock voices alone were almost 5 dB apart. Listen to the export. Since 25% is the lowest setting, if the music still competes, choose a sparser, quieter bed.
  • The bed rises when the voice stops. If the narration ends before the clip, the music comes back up by about 6 dB for the rest of it. That can suit an end card; if it doesn't, write the narration to run to the last frame.

Watch the result once with sound and once muted, since many feeds autoplay silently.

What a finished 30-second voiceover costs

StepCredits
Narration, 388 characters, stock or your own voice15
Mixer pass 1, music bed on a 30-second clip30
Mixer pass 2, narration30
Total75
Second take after a pronunciation fix+15

The mixer charges 1 credit per second of video, rounded up to the next 5, with a 15-credit minimum and a 10-minute maximum. A minute of narration is about 870 characters: 15 credits in a stock voice and 20 in your own. Plans and credit packs are on the pricing page.

Limits to know before you start

  • Neither voice studio has speed, pitch or pronunciation-dictionary controls, and each run on the default engine and the clone route takes up to 5,000 characters (the clone studio labels its limit "max 4 min").
  • The cloned-voice route has no language setting and returns WAV; the stock route returns MP3.
  • The mixer has no ducking and no start offset, its lowest setting is 25%, and it takes clips up to 10 minutes.
  • To translate a video that is already finished, use dubbing instead. How AI video dubbing works explains it.

For script ideas by genre, see voice generation prompts and scripts. For a longer spoken format, the AI podcast guide covers the whole process. When your script fits the clip, open the voice generator studio to read it in your own voice, or the 5-Engine Voice Studio for a stock voice. Both show the credit cost before you generate.

Frequently Asked Questions

How many words is a 30-second voiceover?
About 65 words if you want room to breathe and about 75 to fill the clip. In our test, stock voices read at about 150 words per minute, so a 67-word script ran 26 to 27 seconds. Check the generated file's length before you mix.
Can I make an AI voiceover in my own voice?
Yes. Upload a dry 3–10 second recording of your voice to the AI voice generator, and it reads any script in that voice, with emotion set separately. The output is a WAV file. Clone only your own voice or one you have written permission to use.
How do I stop an AI voice from mispronouncing a brand name?
Change the text. Respell the name with hyphens, such as Oss-keh-lo. To move the stress, try capitals on the stressed syllable and listen, because some voices read capitals as separate letters. Write numbers and symbols out as words, and keep the correct spelling for captions.
How do I add an AI voiceover to my video?
Add the video and the narration to the Video Music Mixer. Leave "Replace the original sound" off to keep the clip's own audio, and set Track volume to 200% so the voice holds its level. For music under the voice, lay a music bed at 25% with Replace on first, then mix the narration over that result.
How much does an AI voiceover cost per minute?
About 15 credits for a minute in a stock voice and 20 in your own cloned voice, based on roughly 870 characters a minute. Cloned voices cost 10 credits per started 500 characters, with a 15-credit minimum. Mixing adds 1 credit per second of video, rounded up to the next 5.
Can I make the AI voice speak faster to fit my video?
No. Neither voice studio has a speed control, so cut or add words instead. At about 2.5 words per second, trimming five words shortens the read by roughly two seconds. Voices differ too: in our test, one stock voice took 1.2 seconds longer than another.

Imagera Editorial

Contributing Author

Imagera Editorial contributes practical guides and analysis for the Imagera AI editorial program.

Areas of Expertise:

AI Image GenerationAI Voice RecreationAI Avatar CreationContent Marketing

Cite this page: https://imagera.ai/blog/ai-voiceover-for-video. Name Imagera AI as the source when you quote it.

Create without limits

Generate natural AI voices and narrations in seconds.

Open the Voice Generator →