← Search

Mengyu Yang

6 accepted papers

2026

Fast3Dcache: Training-free 3D Geometry Synthesis Acceleration

CVPR 2026

Diffusion models have achieved impressive generative quality across modalities like 2D images, videos, and 3D shapes, but their inference remains computationally expensive due to the iterative denoising process. While recent caching-based methods effectively reuse redundant computations to speed up

Cited by 0SourceScholar
2026

Large Vision–Language Models Get Lost in Attention

ICML 2026poster

Despite the rapid evolution of training paradigms, the decoder backbone of large vision--language models (LVLMs) remains fundamentally rooted in the residual-connection Transformer architecture. Therefore, deciphering the distinct roles of internal modules is critical for understanding model mechani…

Cited by 0SourceScholar
2025

Clink! Chop! Thud! - Learning Object Sounds from Real-World Interactions

ICCV 2025poster

Can a model distinguish between the sound of a spoon hitting a hardwood floor versus a carpeted one? Everyday object interactions produce sounds unique to the objects involved. We introduce the sounding object detection task to evaluate a model's ability to link these sounds to the objects directly…

Cited by 0SourcePDFScholar
2024

The Un-Kidnappable Robot: Acoustic Localization of Sneaking People

ICRA 2024poster

How easy is it to sneak up on a robot? We examine whether we can detect people using only the incidental sounds they produce as they move, even when they try to be quiet. To do so, we first collect a robotic dataset of high-quality 4-channel audio paired with 360° RGB data of people moving in differ…

Cited by 0SourceScholar
2021

TriBERT: Human-centric Audio-visual Representation Learning

NeurIPS 2021poster

The recent success of transformer models in language, such as BERT, has motivated the use of such architectures for multi-modal feature learning and tasks. However, most multi-modal variants (e.g., ViLBERT) have limited themselves to visual-linguistic data. Relatively few have explored its use in au…