For text to speech for product videos, let the edit determine what narration must do. Inventory the visual beats, reserve speech for information the picture cannot carry alone, and approve audio inside the actual cut. This keeps the viewer's attention on the product rather than forcing every cursor movement into a sentence.
The edit decides the voiceover
Start by watching the rough cut muted. Mark where a new idea begins, where a label must be read exactly, where the viewer needs quiet to inspect the interface, and where a scene change creates a hard boundary.
The decision is not “which voice sounds best?” It is “which lines deserve narration, and can they fit without competing with the screen?” If the picture already proves a click, narration can explain why the action matters instead of describing the pointer.
Method: map visual beats first
Build a four-column beat sheet: time window, visible event, listener takeaway, and spoken requirement. Classify each possible line as essential, supportive, or redundant. Delete redundant lines before auditioning voices.
Choose the tightest meaningful beat for the audition. It might contain a product name, a changing interface label, and a transition from context to instruction. Test that exact excerpt across candidates in Fish Voice voices, then generate the leading option with text to speech.
Direct revisions with observable contrasts. “Pause after the plan name” or “land the verb before the dialog opens” gives an editor a decision. Broad mood words leave timing and emphasis unresolved.
Workflow from beat sheet to final mix
- Lock the narration purpose for every beat.
- Approve the difficult excerpt against picture.
- Split the script wherever a screen, claim, or revision owner changes.
- Generate and label each segment independently.
- Assemble the clean voice track before adding music.
- Review once without captions, then again with captions, music, and effects.
The first review exposes unclear speech and bad pacing. The complete review reveals competition between words, on-screen text, music, and transitions. Test on both headphones and an ordinary phone or laptop speaker; the second environment often exposes masked consonants and crowded moments.
Keep a timing note beside each file, such as billing-02-plan-choice-12s-v2.mp3. When the interface changes, the filename and beat sheet point to the smallest clip that needs replacement. Generation History can supply the prior take as a continuity reference.
Limits that trigger a rewrite
Rewrite before regenerating when a line repeats visible copy, contains two actions in one short beat, explains a control before it appears, or leaves no silence for the viewer to look. A slower voice cannot repair an overfilled information budget.
Pause localization when labels, screenshots, or timing are still unstable. Translated wording can change duration, so each language needs its own beat check rather than a promise that one edit will fit all versions.
Fish Voice supplies generated speech and downloadable files. Final synchronization, caption accuracy, music balance, loudness, and publication approval remain production decisions outside the generator.
Release evidence
Before export, verify that:
- every spoken product label matches the current interface;
- claims and numbers match the approved source script;
- each clip maps to one beat and one script revision;
- the reference and replacement clips sound continuous;
- the complete mix leaves room for captions and visual inspection;
- the handoff names the factual approver.
This checklist turns the video, not an isolated waveform, into the unit of approval.
FAQ
Should narration describe every click?
No. Describe the reason, consequence, or hidden state when the pointer already makes the action obvious.
When should the voice be chosen?
After the beat sheet reveals the hardest timing and terminology. Audition that passage rather than a generic sample.
What changes after a UI label update?
Replace the affected segment, update the paired script revision, and recheck its adjacent transitions in the final cut.
Can one voice cover a whole product library?
It can provide continuity, but each format still needs its own pacing, information budget, and in-context approval.

