One close-up shot of your singer performing
Lip sync mode produces one close-up shot of your singer performing the track, with the mouth driven directly from the audio.
It is the cheaper and more direct of the two modes, and it suits short clips where the performance is the whole point.
Last updated 2026-07-29
Three engines
Choose between sharpest mouth sync, a faster cheaper draft with camera motion, or the lowest-cost option.
Cheaper than a full video
A single shot costs less to render than a sequence of scenes, which makes it the sensible choice for hook clips and for testing.
Driven by your track
The mouth movement comes straight from the audio rather than being approximated.
Ideal for hook clips
A close-up performance of the strongest four bars is exactly the shape short-form rewards.
Step 1
The audio drives the performance, so both are needed up front.
Step 2
Best quality for a release, fast for a draft, or the cheapest for volume testing.
Step 3
Cost scales with length, so trimming to the hook keeps it cheap.
Step 4
Portrait for short-form, landscape for YouTube.
The best-quality engine produces the sharpest mouth sync and costs the most. It is what you want for anything you are actually releasing.
The faster engine renders a draft with camera motion at a lower cost, which is useful for checking whether a photo suits a track. The cheapest option is best for testing at volume before committing.
Lip sync is the right mode when the performance is the content: a hook clip, a snippet of a verse, a single memorable line.
A full music video is the right mode when you want the video to have somewhere to go across a longer track. Many people use both, generating lip sync clips for volume and a full video for the release itself.
Cost scales with the length of the audio and depends on which engine you choose. Trimming to the hook makes a large difference, and the app shows the exact cost before you generate.
Best quality for anything you are releasing, faster options for checking an idea or testing many variations. The app shows the token cost for each before you commit.
Yes. The mouth movement is driven from the audio, so it works with your own recordings as well as tracks generated in Muzie.

© 2026 APHVN Limited. All rights reserved.
APHVN Limited is registered in the United Kingdom. Company number 15968200.
https://aphvn.com
Support
Follow Us
Discord