← Search

Min-Hung Chen

25 accepted papers

2026

4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation

CVPR 2026

Despite advances in Multimodal LLMs (MLLMs), their ability to reason over 3D structures and temporal dynamics remains limited, constrained by weak 4D perception and temporal understanding. Existing 3D and 4D Video Question Answering (VQA) benchmarks also emphasize static scenes and lack region-level

Cited by 0SourcecodeScholar
2026

Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning

CVPR 2026

Vision-Language-Action (VLA) tasks require reasoning over complex visual scenes and executing adaptive actions in dynamic environments. While recent studies on reasoning VLAs show that explicit chain-of-thought (CoT) can improve generalization, they suffer from high inference latency due to lengthy

Cited by 0SourceScholar
2026

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

ICML 2026poster

As language models become increasingly capable, users expect them to provide not only accurate responses but also behaviors aligned with diverse human preferences across a variety of scenarios. To achieve this, Reinforcement learning (RL) pipelines have begun incorporating multiple rewards, each cap…

Cited by 0SourceScholar
2026

Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them

ICML 2026poster

Video diffusion models can generate visually stunning content, yet frequently produce motion that violates physical laws, objects accelerate implausibly or vanish mid-trajectory. We reveal a surprising finding: a 2-step generation often exhibits better physical consistency than a 50-step output from…

Cited by 0SourceScholar
2026

V2V-GoT: Vehicle-To-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Models and Graph-Of-Thoughts

ICRA 2026poster

Current state-of-the-art autonomous vehicles could face safety critical situations when their local sensors are occluded by large objects on the road nearby. Vehicle-to-vehicle (V2V) cooperative autonomous driving is proposed to address this problem. More recent work further adopts a new approach th…

2026

V2V-LLM: Vehicle-To-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Models

ICRA 2026poster

Current autonomous driving vehicles rely mainly on their individual sensors to understand surrounding scenes and plan for future trajectories, which can be unreliable when the sensors are malfunctioning or occluded. To address this problem, cooperative perception methods via vehicle-to-vehicle (V2V)…

2025

AuraFusion360: Augmented Unseen Region Alignment for Reference-based 360deg Unbounded Scene Inpainting

CVPR 2025poster

Three-dimensional scene inpainting is crucial for applications from virtual reality to architectural visualization, yet existing methods struggle with view consistency and geometric accuracy in 360deg unbounded scenes. We present AuraFusion360, a novel reference-based method that enables high-qualit…

Cited by 0SourcePDFScholar
2025

BlurDM: A Blur Diffusion Model for Image Deblurring

NeurIPS 2025poster

Diffusion models show promise for dynamic scene deblurring; however, existing studies often fail to leverage the intrinsic nature of the blurring process within diffusion models, limiting their full potential. To address it, we present a Blur Diffusion Model (BlurDM), which seamlessly integrates the…

Cited by 0SourcecodeScholar
2025

HERMES: temporal-coHERent long-forM understanding with Episodes and Semantics

ICCV 2025poster

Long-form video understanding presents unique challenges that extend beyond traditional short-video analysis approaches, particularly in capturing long-range dependencies, processing redundant information efficiently, and extracting high-level semantic concepts. To address these challenges, we propo…

2025

Hymba: A Hybrid-head Architecture for Small Language Models

ICLR 2025spotlight

We propose Hymba, a family of small language models featuring a hybrid-head parallel architecture that integrates attention mechanisms and state space models (SSMs) within the same layer, offering parallel and complementary processing of the same inputs. In this hybrid-head module, attention heads p…

Cited by 12SourcePDFScholar
2025

LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long Videos

ICCV 2025poster

LongSplat addresses critical challenges in novel view synthesis (NVS) from casually captured long videos characterized by irregular camera motion, unknown camera poses, and expansive scenes. Current methods often suffer from pose drift, inaccurate geometry initialization, and severe memory limitatio…

2025

MovieCORE: COgnitive REasoning in Movies

EMNLP 2025

This paper introduces MovieCORE, a novel video question answering (VQA) dataset designed to probe deeper cognitive understanding of movie content. Unlike existing datasets that focus on surface-level comprehension, MovieCORE emphasizes questions that engage System-2 thinking while remaining specific

2025

Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks

CVPR 2025poster

We present Omni-RGPT, a multimodal large language model designed to facilitate region-level comprehension for both images and videos. To achieve consistent region representation across spatio-temporal dimensions, we introduce Token Mark, a set of tokens highlighting the target regions within the vis…

Cited by 2SourcePDFScholar
2025

SANER: Annotation-free Societal Attribute Neutralizer for Debiasing CLIP

ICLR 2025poster

Large-scale vision-language models, such as CLIP, are known to contain societal bias regarding protected attributes (e.g., gender, age). This paper aims to address the problems of societal bias in CLIP. Although previous studies have proposed to debias societal bias through adversarial learning or t…

Cited by 2SourcePDFScholar
2025

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning

NeurIPS 2025poster

Vision-language-action (VLA) reasoning tasks require agents to interpret multimodal instructions, perform long-horizon planning, and act adaptively in dynamic environments. Existing approaches typically train VLA models in an end-to-end fashion, directly mapping inputs to actions without explicit re…

Cited by 0SourceScholar
2024

Diffusion-Reward Adversarial Imitation Learning

NeurIPS 2024poster

Imitation learning aims to learn a policy from observing expert demonstrations without access to reward signals from environments. Generative adversarial imitation learning (GAIL) formulates imitation learning as adversarial learning, employing a generator policy learning to imitate expert behaviors…

2024

DoRA: Weight-Decomposed Low-Rank Adaptation

ICML 2024oral

Among the widely used parameter-efficient fine-tuning (PEFT) methods, LoRA and its variants have gained considerable popularity because of avoiding additional inference costs. However, there still often exists an accuracy gap between these methods and full fine-tuning (FT). In this work, we first in…

2024

Image-Text Co-Decomposition for Text-Supervised Semantic Segmentation

CVPR 2024poster

This paper addresses text-supervised semantic segmentation aiming to learn a model capable of segmenting arbitrary visual concepts within images by using only image-text pairs without dense annotations. Existing methods have demonstrated that contrastive learning on image-text pairs effectively alig…

2024

PartDistill: 3D Shape Part Segmentation by Vision-Language Model Distillation

CVPR 2024poster

This paper proposes a cross-modal distillation framework PartDistill which transfers 2D knowledge from vision-language models (VLMs) to facilitate 3D shape part segmentation. PartDistill addresses three major challenges in this task: the lack of 3D segmentation in invisible or undetected regions in…

2024

Probabilistic 3D Multi-Object Cooperative Tracking for Autonomous Driving via Differentiable Multi-Sensor Kalman Filter

ICRA 2024poster

Current state-of-the-art autonomous driving vehicles mainly rely on each individual sensor system to perform perception tasks. Such a framework’s reliability could be limited by occlusion or sensor failure. To address this issue, more recent research proposes using vehicle-to-vehicle (V2V) communica…

Cited by 8SourcecodeScholar
2023

2D-3D Interlaced Transformer for Point Cloud Segmentation with Scene-Level Supervision

ICCV 2023poster

We present a Multimodal Interlaced Transformer (MIT) that jointly considers 2D and 3D data for weakly supervised point cloud segmentation. Research studies have shown that 2D and 3D features are complementary for point cloud segmentation. However, existing methods require extra 2D annotations to ach…

Cited by 16PDFScholar
2023

Learning Continuous Exposure Value Representations for Single-Image HDR Reconstruction

ICCV 2023poster

Deep learning is commonly used to produce impressive results in reconstructing HDR images from LDR images. LDR stack-based methods are used for single-image HDR reconstruction, generating an HDR image from a deep learning generated LDR stack. However, current methods generate the LDR stack with pred…

Cited by 10PDFScholar
2020

Action Segmentation With Joint Self-Supervised Temporal Domain Adaptation

CVPR 2020poster

Despite the recent progress of fully-supervised action segmentation techniques, the performance is still not fully satisfactory. One main challenge is the problem of spatiotemporal variations (e.g. different people may perform the same activity in various ways). Therefore, we exploit unlabeled video…

Cited by 149PDFcodeScholar
2020

Interpretable Self-Attention Temporal Reasoning for Driving Behavior Understanding

ICASSP 2020accepted

Performing driving behaviors based on causal reasoning is essential to ensure driving safety. In this work, we investigated how state-of-the-art 3D Convolutional Neural Networks (CNNs) perform on classifying driving behaviors based on causal reasoning. We proposed a perturbation-based visual explana…

Cited by 0SourceScholar
2019

Temporal Attentive Alignment for Large-Scale Video Domain Adaptation

ICCV 2019oral

Although various image-based domain adaptation (DA) techniques have been proposed in recent years, domain shift in videos is still not well-explored. Most previous works only evaluate performance on small-scale datasets which are saturated. Therefore, we first propose two large-scale video DA datase…

Cited by 243PDFcodeScholar