Create AI avatar videos from a single image
Turn one image into a realistic AI avatar video. Combine the image and audio to create natural facial expressions, head movement, and lifelike animation.

OmniHuman 1.5 is ByteDance's AI avatar and digital human model. It turns a single image and an audio clip into a realistic talking video with natural lip sync, expressive faces, and lifelike movement. One image and audio are enough to create avatar, lip-synced, and digital human videos.
Create now
OMNIHUMAN 1.5 工作台
上傳一張圖片和一段音訊,設定輸出方式,確認點數後生成數位人影片。
多人圖片可上傳目標人物的遮罩圖。
0/300
-1 為隨機;填寫相同非負整數可提高結果一致性。
在這裡查看 OmniHuman 1.5 的生成進度和已完成影片。
正在載入作品…
Create realistic AI avatar, lip-synced, and digital human videos from one image and an audio clip.
Turn one image into a realistic AI avatar video. Combine the image and audio to create natural facial expressions, head movement, and lifelike animation.

Match speech to mouth movement so the subject speaks naturally with realistic lips and expressive facial details.

Use the meaning and tone of the audio to guide matching emotions, gestures, and body language for a more engaging video.

Go beyond a traditional talking-photo tool with natural upper-body and full-body movement.

Use the model for digital human generators, virtual idols, avatar platforms, AI video tools, and creative products.

Prepare a clear image and an audio clip before creating a digital human video.
| Item | Media or format | Guidance |
|---|---|---|
| Image | JPG/JPEG, PNG, or WebP; up to 10 MB | A well-lit, front-facing image with a clear face is recommended. Pets and anime subjects are also supported. |
| Audio | MP3, WAV, AAC, OGG, or MP4 audio; up to 10 MB | Audio must be under 60 seconds. For more consistent results, 15 seconds or less is recommended. |
| Video | MP4 at 720P or 1080P | The generated subject's movement follows the supplied audio. |
Prepare an image, add audio, and review the synchronized result.
Upload a clear image. A well-lit portrait with visible facial details works best.
Add spoken, sung, or recorded audio to guide the mouth, expressions, and body movement.
Generate and review the synchronized MP4 digital human video.
Use an image and audio for virtual creators, talking photos, music, and brand content.
Use a single image and audio file to create realistic virtual influencer content and engaging short videos for creators, agencies, and brands.
Bring natural speech synchronization to talking-photo experiences, avatar generators, virtual presenters, and personalized videos.
Turn songs, recordings, and audio tracks into expressive music videos with precise lip sync and emotion-rich performances.
Create brand spokespeople, product explainers, and promotional videos while keeping the character consistent across campaigns.
OmniHuman 1.5 is ByteDance's AI avatar and digital human model. It turns one image and audio into a video with synchronized lips, facial expressions, and natural movement.
Provide one image and an audio clip. The model analyzes the voice, emotion, and subject to generate a video synchronized with the audio.
Yes. OmniHuman 1.5 matches mouth movement to speech while creating natural facial expressions and movement.
JPG/JPEG, PNG, and WebP images up to 10 MB are supported. A clear, front-facing image is recommended.
Audio of 15 seconds or less is recommended. It must be shorter than 60 seconds; longer clips may reduce visual quality.
Yes. The model supports people, pets, anime, cartoon characters, and stylized subjects.
Yes. In addition to facial and head movement, OmniHuman 1.5 supports natural upper-body and full-body animation.
The generated video is an MP4 file suitable for social, web, and content-creation workflows.
OmniHuman 1.5 is ByteDance's AI avatar and digital human model. It turns a single image and an audio clip into a realistic talking video with natural lip sync, expressive faces, and lifelike movement. One image and audio are enough to create avatar, lip-synced, and digital human videos.
Create now