Short-form video keeps people watching for one reason more than any other: captions that move. Not static subtitles sitting at the bottom of the frame — words that pop in one at a time, in sync with the voice, the way TikTok's native caption tool and apps like CapCut popularized. This is called word-by-word (or "karaoke-style") text reveal, and it's built directly into RapidVideoMaker's image-to-video mode.
What word-by-word reveal actually does
Instead of showing a full sentence at once, word-by-word reveal displays your text overlay one word at a time, timed to a fixed pace you control. Each word appears, holds briefly, and gets replaced by the next — the same rhythm viewers are used to from short-form video captions, minus the manual keyframing a traditional editor requires.
Technically, this generates a sequence of overlay layers, each one showing the words revealed so far — the more words in your text, the more layers stack up, which is why there's a practical cap (50 words) on how long a single reveal can run.
Reveal speed and animation
Two settings control how the words land:
- Reveal speed — from 0.3 to 5 words per second, default 1.5. Slower speeds suit a reflective voiceover or a quote meant to be read carefully; faster speeds match energetic, punchy narration.
- Fade animation — each word can pop in instantly, or fade in and out smoothly as it's replaced. Fade reads as more polished on a finished video; instant reveal feels punchier and matches fast-cut content.
One thing worth knowing: word-by-word reveal is mutually exclusive with the other text entrance animations (slide, bounce) and with the general on-screen text effects — enabling it automatically disables those, since the word-level timing already handles the entrance for you.
Syncing captions with a voiceover
Word-by-word reveal is most effective when the pace of the words roughly matches the pace of speech underneath. If you don't already have a voiceover, RapidVideoMaker's text-to-speech mode can generate one from a script in 48 voices across 20 languages — generate the voiceover first, listen to its pacing, then set the reveal speed to match roughly one word per spoken word.
How to add word-by-word captions with RapidVideoMaker
- Upload one image (JPG/PNG) and one MP3 — the voiceover or music track the video will be built around.
- Add a text overlay: type your caption, choose its position (top, center, bottom), and pick a background box if the image is busy.
- Set the reveal mode to word-by-word, choose a reveal speed, and optionally enable the fade animation.
- Choose a motion effect for the image (or leave it static) and pick your fit mode.
- Generate — the output duration matches your MP3 exactly, words revealing in sync as it plays.
The same options are available through the API documentation as parameters of the image_to_video mode — useful for generating captioned video at scale from a batch of scripts.
Practical tips
- Keep captions short. With a 50-word cap and a readable pace, word-by-word works best for a punchy hook or a key quote — not a full paragraph.
- Match the reveal speed to your voiceover's actual pace rather than guessing; a mismatch between spoken words and revealed words is more distracting than no captions at all.
- Use a background box behind the text on photos with busy or light backgrounds — without it, word-by-word text over a bright image becomes unreadable, especially on smaller mobile screens.
Frequently asked questions
Can I use word-by-word reveal with a long paragraph?
Not really — there's a 50-word cap on this mode, since each additional word adds another rendering layer. For longer narration, keep the on-screen text to a short highlight or key line, and let the voiceover carry the rest.
Does word-by-word reveal work with any font or animation style?
Yes for position, background box, and color — the reveal mode and the fade animation are the two settings specific to word-by-word. Other entrance animations (slide, bounce) are disabled automatically since word-by-word already handles the entrance.
Do I need a voiceover for this to look good?
No — it works with music alone too, though captions synced to nothing in particular tend to look best as a short quote or hook rather than a long stream of text. Voiceover sync is where the effect is most convincing.
Is this the same as automatic speech-to-text captions?
No. Word-by-word reveal displays text you write, timed at a speed you set — it doesn't transcribe audio automatically. Pair it with RapidVideoMaker's text-to-speech mode if you want the voice and the captions to originate from the exact same script.