Making a music video used to mean a camera, a location, a crew and a day you could not afford. Now it means a photo and a song.
The approach here is different from most AI video tools. Instead of generating abstract visuals over your audio, you upload a picture of a person and the video is that person singing the track, with their mouth driven from the audio itself.
Start with the photo
The photo matters more than anything else you will choose, so spend a moment on it.
What works:
- Front facing. The face roughly towards the camera.
- Well lit. Daylight is ideal. Even indoor lighting is fine.
- Face reasonably large in the frame. A head and shoulders shot beats a full body one.
- Unobstructed. No sunglasses, no hand over the mouth, nothing across the face.
What causes a weak result:
- Heavy shadow across one side of the face
- Extreme angles, especially profile shots
- Very low resolution or a heavily compressed screenshot
- Sunglasses, which remove the eyes the model needs
An ordinary phone selfie taken near a window is usually better than a carefully staged photo in bad light.
Then pick the song
There are three ways to get one, and they suit different situations:
- The library. Over a thousand tracks, ready to go. Fastest route, and best if you want something recognisable.
- Upload your own. If you make music, this is the one. Your recording, your rights, and the video becomes promotion for your actual track.
- Generate one. Describe a vibe and get an original with vocals and instruments. This is a Pro feature.
Then trim it. This is the step people skip and it matters twice over: short clips cost less to render, and short-form only ever wanted the hook anyway.
Find the four bars that work with no context. If the hook lands at 1:20, your video starts at 1:20. Nobody waits through a build.
Choose how it performs
There are two modes and they suit different jobs.
Lip sync gives you one close-up shot of the person singing. It is cheaper, it is faster, and for a hook clip it is usually the right choice. The performance is the content, so a single shot is enough.
Full music video builds a sequence of cinematic scenes with your performer across them, cut against the lyrics. Use this when you want the video to have somewhere to go across a longer track.
If you are testing which hook works, use lip sync. If you are making the piece for a release, use the full video.
Draft before you render
Draft quality costs less than full quality. Use it to check whether the photo and the song work together before you spend on the real thing.
Almost nobody gets the combination right first time. The photo that looked perfect can turn out to be lit badly for this, or the song section can turn out to be the wrong four bars. Finding that out on a draft is much cheaper than finding it out on a full render.
Export for where you are posting
Vertical for TikTok, Reels and Shorts. Landscape only if YouTube is the main destination.
This is not a detail. Letterboxed wide video on a vertical platform reads as reposted content and gets suppressed accordingly.
Keep the face centred and slightly high. The bottom fifth of the screen is captions and the sound name, and the right edge is the button stack.
Then make more
One video gets one chance with the algorithm. Five videos of the same hook, with different photos, styles or sections, get five chances and tell you which one works.
That is only realistic because each one takes minutes rather than an evening. It is the entire reason generation matters for music promotion, and it is the difference between artists who post daily and artists who post once per release.
Common mistakes
- Starting at the start of the song. Start at the hook.
- A badly lit photo. This is the single most common cause of a poor result.
- Rendering full quality first. Draft, check, then spend.
- Wide export on a vertical platform. Black bars get you downranked.
- Making one video. Make five.
- Using someone else's photo without asking. Do not do this. For friends and family, ask first.
