← Search

Yixin Li

2 accepted papers

2025

Multiview Equivariance Improves 3D Correspondence Understanding with Minimal Feature Finetuning

ICLR 2025poster

Vision foundation models, particularly the ViT family, have revolutionized image understanding by providing rich semantic features. However, despite their success in 2D comprehension, their abilities on grasping 3D spatial relationships are still unclear. In this work, we evaluate and enhance the 3D…

2025

SwapTalk: Audio-Driven Talking Face Generation with One-Shot Customization in Latent Space

ICASSP 2025accepted

Combining face-swapping with lip synchronization offers a cost-effective solution for generating customized talking faces. However, directly cascading existing models can introduce significant interference and reduce video clarity due to limited interaction space in the low-level RGB domain. To solv…

Cited by 0SourceScholar