← Search

Fu-En Yang

14 accepted papers

2026

Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning

CVPR 2026

Vision-Language-Action (VLA) tasks require reasoning over complex visual scenes and executing adaptive actions in dynamic environments. While recent studies on reasoning VLAs show that explicit chain-of-thought (CoT) can improve generalization, they suffer from high inference latency due to lengthy

Cited by 0SourceScholar
2026

Frequency Switching Mechanism for Parameter-Efficient Multi-Task Learning

CVPR 2026

Multi-task learning (MTL) aims to enable a single model to solve multiple tasks efficiently; however, current parameter-efficient fine-tuning (PEFT) methods remain largely limited to single-task adaptation. We introduce Free Sinewich, a parameter-efficient multi-task learning framework that enables

Cited by 0SourceScholar
2026

Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them

ICML 2026poster

Video diffusion models can generate visually stunning content, yet frequently produce motion that violates physical laws, objects accelerate implausibly or vanish mid-trajectory. We reveal a surprising finding: a 2-step generation often exhibits better physical consistency than a 50-step output from…

Cited by 0SourceScholar
2025

LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long Videos

ICCV 2025poster

LongSplat addresses critical challenges in novel view synthesis (NVS) from casually captured long videos characterized by irregular camera motion, unknown camera poses, and expansive scenes. Current methods often suffer from pose drift, inaccurate geometry initialization, and severe memory limitatio…

2025

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning

NeurIPS 2025poster

Vision-language-action (VLA) reasoning tasks require agents to interpret multimodal instructions, perform long-horizon planning, and act adaptively in dynamic environments. Existing approaches typically train VLA models in an end-to-end fashion, directly mapping inputs to actions without explicit re…

Cited by 0SourceScholar
2025

VideoMage: Multi-Subject and Motion Customization of Text-to-Video Diffusion Models

CVPR 2025poster

Customized text-to-video generation aims to produce high-quality videos that incorporate user-specified subject identities or motion patterns. However, existing methods mainly focus on personalizing a single concept, either subject identity or motion pattern, limiting their effectiveness for multipl…

Cited by 0SourcePDFScholar
2024

Language-Guided Transformer for Federated Multi-Label Classification

AAAI 2024technical

Federated Learning (FL) is an emerging paradigm that enables multiple users to collaboratively train a robust model in a privacy-preserving manner without sharing their private data. Most existing approaches of FL only consider traditional single-label image classification, ignoring the impact when…

2024

RAPPER: Reinforced Rationale-Prompted Paradigm for Natural Language Explanation in Visual Question Answering

ICLR 2024poster

Natural Language Explanation (NLE) in vision and language tasks aims to provide human-understandable explanations for the associated decision-making process. In practice, one might encounter explanations which lack informativeness or contradict visual-grounded facts, known as implausibility and hall…

Cited by 2SourcePDFScholar
2024

Receler: Reliable Concept Erasing of Text-to-Image Diffusion Models via Lightweight Erasers

ECCV 2024poster

"Concept erasure in text-to-image diffusion models aims to disable pre-trained diffusion models from generating images related to a target concept. To perform reliable concept erasure, the properties of robustness and locality are desirable. The former refrains the model from producing images associ…

2024

Select and Distill: Selective Dual-Teacher Knowledge Transfer for Continual Learning on Vision-Language Models

ECCV 2024poster

"Large-scale vision-language models (VLMs) have shown a strong zero-shot generalization capability on unseen-domain data. However, adapting pre-trained VLMs to a sequence of downstream tasks often leads to the forgetting of previously learned knowledge and a reduction in zero-shot classification per…

Cited by 11SourcePDFScholar
2023

Efficient Model Personalization in Federated Learning via Client-Specific Prompt Generation

ICCV 2023poster

Federated learning (FL) emerges as a decentralized learning framework which trains models from multiple distributed clients without sharing their data to preserve privacy. Recently, large-scale pre-trained models (e.g., Vision Transformer) have shown a strong capability of deriving robust representa…

Cited by 46PDFScholar
2021

Adversarial Teacher-Student Representation Learning for Domain Generalization

NeurIPS 2021spotlight

Domain generalization (DG) aims to transfer the learning task from a single or multiple source domains to unseen target domains. To extract and leverage the information which exhibits sufficient generalization ability, we propose a simple yet effective approach of Adversarial Teacher-Student Represe…

Cited by 71SourcePDFScholar
2021

LayoutTransformer: Scene Layout Generation With Conceptual and Spatial Diversity

CVPR 2021poster

When translating text inputs into layouts or images, existing works typically require explicit descriptions of each object in a scene, including their spatial information or the associated relationships. To better exploit the text input, so that implicit objects or relationships can be properly infe…

Cited by 42PDFcodeScholar
2020

Learning Identity-Invariant Motion Representations for Cross-ID Face Reenactment

CVPR 2020poster

Human face reenactment aims at transferring motion patterns from one face (from a source-domain video) to an-other (in the target domain with the identity of interest).While recent works report impressive results, they are notable to handle multiple identities in a unified model. In this paper, we p…

Cited by 43PDFcodeScholar