← Search

Yiwen Wu

4 accepted papers

2025

AdaV: Adaptive Text-visual Redirection for Vision-Language Models

ACL 2025finding

The success of Vision-Language Models (VLMs) often relies on high-resolution schemes that preserve image details, while these approaches also generate an excess of visual tokens, leading to a substantial decrease in model efficiency. A typical VLM includes a visual encoder, a text encoder, and an LL…

2025

Capturing the Unseen: Vision-Free Facial Motion Capture Using Inertial Measurement Units

AAAI 2025technical

We present Capturing the Unseen (CAPUS), a novel facial motion capture (MoCap) technique that operates without visual signals. CAPUS leverages miniaturized Inertial Measurement Units (IMUs) as a new sensing modality for facial motion capture. While IMUs have become essential in full-body MoCap for t…

Cited by 0SourcePDFScholar
2025

SLIM: Let LLM Learn More and Forget Less with Soft LoRA and Identity Mixture

NAACL 2025long

Despite the recent efforts from the NLP community, balancing the training budget, downstream performance, and general capabilities of large language models (LLM) remains a challenge in many applications. Training the entire model for downstream tasks is expensive, and could easily result in catastro…

Cited by 3SourcePDFScholar
2025

TokMan:Tokenize Manhattan Mask Optimization for Inverse Lithography

NeurIPS 2025poster

Manhattan representations, defined by axis-aligned, orthogonal structures, are widely used in vision, robotics, and semiconductor design for their geometric regularity and algorithmic simplicity. In integrated circuit (IC) design, Manhattan geometry is key for routing, design rule checking, and lith…

Cited by 0SourceScholar