A music video is the one format where the edit is not a creative choice - the song already told you where the cuts go. VideoCue reads the vocal timing off your uploaded track and storyboards one shot per lyric line against it.
This is the working path from an mp3 on your desktop to a rendered music video starring your own characters.
Start from the song
On /app/videos/new, pick From a song. Drop in an MP3, WAV, M4A, AAC, or FLAC up to 100 MB. There is no voice step and no length step: the song is the entire audio bus, and its duration is the video's duration.
Paste the lyrics
The single biggest quality lever on this whole path. Transcription models are trained on speech, and sung vocals - held notes, harmonies, heavy production - are exactly what they handle worst.
Paste the lyrics and we stop transcribing and start force-aligning: the text is known, so every word gets matched to a real timestamp in the audio. Cuts land on the syllable instead of near it. Leave the box blank and we transcribe the track ourselves, which works, but the timing is looser.
One lyric line per line of text. Don't bother with section markers - the instrumental gaps are detected from the audio itself.
Cast, then Look
Cast the characters who appear in the video and mark one as primary. Then pick a Look - it decides the grade, the framing, and the art direction that every shot inherits. Both are editable on the video's Edit page afterwards, which is exactly why the upload lands as a draft instead of firing the pipeline immediately.
No lip sync in v1. Your characters perform in the video; they don't mouth the words.
What happens on generate
- Lyrics are transcribed or force-aligned to word-level timestamps.
- One beat per lyric line, starting on the first word of the line.
- Instrumental stretches - the intro, solos, the outro - get their own beats, subdivided at the track's energy peaks so an instrumental break cuts on the loud moments instead of holding one shot for forty seconds.
- An LLM storyboards visuals only. The timings are ours and get re-stamped after the response, so a hallucinated duration can never reach the render.
- Stills plus motion clips, following whatever motion-clip percentage the channel already uses. There's no new knob to learn.
The rights part, plainly
The upload form requires a rights acknowledgement, and it isn't decorative. Rendering someone else's recording for yourself is fine. Publishing that render to YouTube will draw a Content ID claim - the rights holder can monetise, mute, or block your video, and the claim lands against your channel.
Use your own music, a track you've licensed for commercial use, or something genuinely in the public domain. See picking music that doesn't fight your voiceover for the licensing options that hold up.
Bring it to VideoCue
Build the performer first with build your first character, pick an aesthetic with author a custom Look, then upload the song and generate.