Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction
Apple researchers have introduced a new framework designed to improve the synchronization between audio and video when generating content from text prompts. By addressing limitations in how models process text conditions, the approach aims to resolve alignment issues that have historically hindered the realism of AI-generated audiovisual media.