Post #2938718
2026-05-28 17:57 UTC
Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation
The method adopts a causal dual-stream DiT model to generate synchronized avatar video and speech in a streaming manner. Hallo-Live reaches 20.38 FPS with 0.94 s latency on two NVIDIA H200 GPUs, while preserving strong lip-sync accuracy, visual fidelity, and speech quality.
Built on #LMDB
https://github.com/fudan-generative-vision/Hallo-Live
Replies (0)
No replies.