RapidVideoMaker
Sign up free

How to Add Word-by-Word Animated Captions to a Video (Free)

How to Add Word-by-Word Animated Captions to a Video (Free)

Short-form video keeps people watching for one reason more than any other: captions that move. Not static subtitles sitting at the bottom of the frame — words that pop in one at a time, in sync with the voice, the way TikTok's native caption tool and apps like CapCut popularized. This is called word-by-word (or "karaoke-style") text reveal, and it's built directly into RapidVideoMaker's image-to-video mode.

What word-by-word reveal actually does

Instead of showing a full sentence at once, word-by-word reveal displays your text overlay one word at a time, timed to a fixed pace you control. Each word appears, holds briefly, and gets replaced by the next — the same rhythm viewers are used to from short-form video captions, minus the manual keyframing a traditional editor requires.

Technically, this generates a sequence of overlay layers, each one showing the words revealed so far — the more words in your text, the more layers stack up, which is why there's a practical cap (50 words) on how long a single reveal can run.

Reveal speed and animation

Two settings control how the words land:

  • Reveal speed — from 0.3 to 5 words per second, default 1.5. Slower speeds suit a reflective voiceover or a quote meant to be read carefully; faster speeds match energetic, punchy narration.
  • Fade animation — each word can pop in instantly, or fade in and out smoothly as it's replaced. Fade reads as more polished on a finished video; instant reveal feels punchier and matches fast-cut content.

One thing worth knowing: word-by-word reveal is mutually exclusive with the other text entrance animations (slide, bounce) and with the general on-screen text effects — enabling it automatically disables those, since the word-level timing already handles the entrance for you.

Syncing captions with a voiceover

Word-by-word reveal is most effective when the pace of the words roughly matches the pace of speech underneath. If you don't already have a voiceover, RapidVideoMaker's text-to-speech mode can generate one from a script in 48 voices across 20 languages — generate the voiceover first, listen to its pacing, then set the reveal speed to match roughly one word per spoken word.

How to add word-by-word captions with RapidVideoMaker

  1. Upload one image (JPG/PNG) and one MP3 — the voiceover or music track the video will be built around.
  2. Add a text overlay: type your caption, choose its position (top, center, bottom), and pick a background box if the image is busy.
  3. Set the reveal mode to word-by-word, choose a reveal speed, and optionally enable the fade animation.
  4. Choose a motion effect for the image (or leave it static) and pick your fit mode.
  5. Generate — the output duration matches your MP3 exactly, words revealing in sync as it plays.

The same options are available through the API documentation as parameters of the image_to_video mode — useful for generating captioned video at scale from a batch of scripts.

Practical tips

  • Keep captions short. With a 50-word cap and a readable pace, word-by-word works best for a punchy hook or a key quote — not a full paragraph.
  • Match the reveal speed to your voiceover's actual pace rather than guessing; a mismatch between spoken words and revealed words is more distracting than no captions at all.
  • Use a background box behind the text on photos with busy or light backgrounds — without it, word-by-word text over a bright image becomes unreadable, especially on smaller mobile screens.

Frequently asked questions

Can I use word-by-word reveal with a long paragraph?
Not really — there's a 50-word cap on this mode, since each additional word adds another rendering layer. For longer narration, keep the on-screen text to a short highlight or key line, and let the voiceover carry the rest.

Does word-by-word reveal work with any font or animation style?
Yes for position, background box, and color — the reveal mode and the fade animation are the two settings specific to word-by-word. Other entrance animations (slide, bounce) are disabled automatically since word-by-word already handles the entrance.

Do I need a voiceover for this to look good?
No — it works with music alone too, though captions synced to nothing in particular tend to look best as a short quote or hook rather than a long stream of text. Voiceover sync is where the effect is most convincing.

Is this the same as automatic speech-to-text captions?
No. Word-by-word reveal displays text you write, timed at a speed you set — it doesn't transcribe audio automatically. Pair it with RapidVideoMaker's text-to-speech mode if you want the voice and the captions to originate from the exact same script.

Ready to create your own videos?

Try RapidVideoMaker