OmniHuman 1.5

OmniHuman 1.5 is ByteDance's AI avatar and digital human model. It turns a single image and an audio clip into a realistic talking video with natural lip sync, expressive faces, and lifelike movement. One image and audio are enough to create avatar, lip-synced, and digital human videos.

Create now

OMNIHUMAN 1.5 STUDIO

Bring a portrait to life with audio

Upload one image and an audio clip, choose the output, review the credit quote, then create your video.

Portrait image
Driving audio

For multiple subjects, upload a mask for the one you want to animate.

0/300

Resolution

-1 is random; reuse a non-negative number for greater consistency.

Your creations

Follow your OmniHuman 1.5 tasks and revisit finished videos.

Loading creations…

OmniHuman 1.5 core capabilities

Create realistic AI avatar, lip-synced, and digital human videos from one image and an audio clip.

Create AI avatar videos from a single image

Turn one image into a realistic AI avatar video. Combine the image and audio to create natural facial expressions, head movement, and lifelike animation.

Create precise lip-synced avatar videos with OmniHuman 1.5

Match speech to mouth movement so the subject speaks naturally with realistic lips and expressive facial details.

Create emotion-aware OmniHuman 1.5 digital human videos

Use the meaning and tone of the audio to guide matching emotions, gestures, and body language for a more engaging video.

Create full-body AI avatar animation with OmniHuman 1.5

Go beyond a traditional talking-photo tool with natural upper-body and full-body movement.

Use OmniHuman 1.5 in AI video applications

Use the model for digital human generators, virtual idols, avatar platforms, AI video tools, and creative products.

Image, audio, and video output

Prepare a clear image and an audio clip before creating a digital human video.

ItemMedia or formatGuidance
ImageJPG/JPEG, PNG, or WebP; up to 10 MBA well-lit, front-facing image with a clear face is recommended. Pets and anime subjects are also supported.
AudioMP3, WAV, AAC, OGG, or MP4 audio; up to 10 MBAudio must be under 60 seconds. For more consistent results, 15 seconds or less is recommended.
VideoMP4 at 720P or 1080PThe generated subject's movement follows the supplied audio.

From one image to an avatar video

Prepare an image, add audio, and review the synchronized result.

  1. 01

    Upload a clear image. A well-lit portrait with visible facial details works best.

  2. 02

    Add spoken, sung, or recorded audio to guide the mouth, expressions, and body movement.

  3. 03

    Generate and review the synchronized MP4 digital human video.

Where OmniHuman 1.5 fits

Use an image and audio for virtual creators, talking photos, music, and brand content.

1

Virtual creators and influencers

Use a single image and audio file to create realistic virtual influencer content and engaging short videos for creators, agencies, and brands.

2

Lip-synced avatar videos

Bring natural speech synchronization to talking-photo experiences, avatar generators, virtual presenters, and personalized videos.

3

AI music videos

Turn songs, recordings, and audio tracks into expressive music videos with precise lip sync and emotion-rich performances.

4

Digital spokesperson videos

Create brand spokespeople, product explainers, and promotional videos while keeping the character consistent across campaigns.

OmniHuman 1.5 FAQ

What is OmniHuman 1.5?+

OmniHuman 1.5 is ByteDance's AI avatar and digital human model. It turns one image and audio into a video with synchronized lips, facial expressions, and natural movement.

How does it create a digital human video?+

Provide one image and an audio clip. The model analyzes the voice, emotion, and subject to generate a video synchronized with the audio.

Can it create lip-synced videos?+

Yes. OmniHuman 1.5 matches mouth movement to speech while creating natural facial expressions and movement.

Which image formats are supported?+

JPG/JPEG, PNG, and WebP images up to 10 MB are supported. A clear, front-facing image is recommended.

How long should the audio be?+

Audio of 15 seconds or less is recommended. It must be shorter than 60 seconds; longer clips may reduce visual quality.

Can it animate anime or cartoon characters?+

Yes. The model supports people, pets, anime, cartoon characters, and stylized subjects.

Can it show upper-body or full-body movement?+

Yes. In addition to facial and head movement, OmniHuman 1.5 supports natural upper-body and full-body animation.

What format is the video?+

The generated video is an MP4 file suitable for social, web, and content-creation workflows.

See how digital human videos are made

OmniHuman 1.5 is ByteDance's AI avatar and digital human model. It turns a single image and an audio clip into a realistic talking video with natural lip sync, expressive faces, and lifelike movement. One image and audio are enough to create avatar, lip-synced, and digital human videos.

Create now
See how it works