← Search

Kun-Yu Lin

20 accepted papers

2026

EgoTraj-Bench: Towards Robust Trajectory Prediction under Ego-View Noisy Observations

ICRA 2026poster

Reliable trajectory prediction from an ego-centric perspective is crucial for robotic navigation in human-centric environments. However, existing methods typically assume noiseless observation histories, failing to account for the perceptual artifacts inherent in first-person vision, such as occlusi…

2026

Reason with Thumbnails, Answer with Focus: An Efficient and Effective Paradigm for Multimodal Grounded Visual Reasoning

ICML 2026poster

To enhance the interpretability of multimodal large language models' outputs, recent efforts explored Grounded Visual Reasoning (GVR), in which the model is trained to select relevant image regions before answering the question. However, the multi-round ``ground-then-answer'' and reasoning nature of…

Cited by 0SourceScholar
2025

Decoupled Distillation to Erase: A General Unlearning Method for Any Class-centric Tasks

CVPR 2025highlight

In this work, we present DEcoupLEd Distillation To Erase (DELETE), a general and strong unlearning method for any class-centric tasks. To derive this, we first propose a theoretical framework to analyze the general form of unlearning loss and decompose it into forgetting and retention terms. Through…

Cited by 2SourcePDFScholar
2025

Exploring the Limits of Vision-Language-Action Manipulation in Cross-task Generalization

NeurIPS 2025poster

The generalization capabilities of vision-language-action (VLA) models to unseen tasks are crucial to achieving general-purpose robotic manipulation in open-world settings. However, the cross-task generalization capabilities of existing VLA models remain significantly underexplored. To address this…

Cited by 0SourceScholar
2025

Less Static, More Private: Towards Transferable Privacy-Preserving Action Recognition by Generative Decoupled Learning

ICCV 2025poster

This work focuses on the task of privacy-preserving action recognition (PPAR), which aims to protect individual privacy in action videos without compromising recognition performance. Despite recent advancements, existing PPAR models still struggle with video domain shifts. To address this challenge,…

Cited by 0SourcePDFScholar
2025

Mitigating the Human-Robot Domain Discrepancy in Visual Pre-training for Robotic Manipulation

CVPR 2025poster

Learning generalizable visual representations across different embodied environments is essential for effective robotic manipulation in real-world scenarios. However, the limited scale and diversity of robot demonstration data pose a significant challenge. Recent research has explored leveraging lar…

Cited by 8SourcePDFScholar
2025

Modeling Multiple Normal Action Representations for Error Detection in Procedural Tasks

CVPR 2025poster

Error detection in procedural activities is essential for consistent and correct outcomes in AR-assisted and robotic systems. Existing methods often focus on temporal ordering errors or rely on static prototypes to represent normal actions. However, these approaches typically overlook the common sce…

2025

ParGo: Bridging Vision-Language with Partial and Global Views

AAAI 2025technical

This work presents ParGo, a novel Partial-Global projector designed to connect the vision and language modalities for Multimodal Large Language Models (MLLMs). Unlike previous works that rely on global attention-based projectors, our ParGo bridges the representation gap between the separately pre-tr…

2025

Person De-reidentification: A Variation-guided Identity Shift Modeling

CVPR 2025poster

Person re-identification (ReID) is to associate images of individuals from different camera views against cross-view variations. Like other surveillance technologies, Re-ID faces serious privacy challenges, particularly the potential for unauthorized tracking. Although various tasks (e.g., face reco…

Cited by 0SourcePDFScholar
2025

ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations

ICCV 2025poster

Referring video object segmentation (RVOS) aims to segment target objects throughout a video based on a text description. This is challenging as it involves deep vision-language understanding, pixel-level dense prediction and spatiotemporal reasoning. Despite notable progress in recent years, existi…

Cited by 0SourcePDFScholar
2025

ViSpeak: Visual Instruction Feedback in Streaming Videos

ICCV 2025poster

Recent advances in Large Multi-modal Models (LMMs) are primarily focused on offline video understanding. Instead, streaming video understanding poses great challenges to recent models due to its time-sensitive, omni-modal and interactive characteristics. In this work, we aim to extend the streaming…

2023

AsyFOD: An Asymmetric Adaptation Paradigm for Few-Shot Domain Adaptive Object Detection

CVPR 2023poster

In this work, we study few-shot domain adaptive object detection (FSDAOD), where only a few target labeled images are available for training in addition to sufficient source labeled images. Critically, in FSDAOD, the data-scarcity in the target domain leads to an extreme data imbalance between the s…

2023

Diversifying Spatial-Temporal Perception for Video Domain Generalization

NeurIPS 2023poster

Video domain generalization aims to learn generalizable video classification models for unseen target domains by training in a source domain. A critical challenge of video domain generalization is to defend against the heavy reliance on domain-specific cues extracted from the source domain when reco…

2023

Event-Guided Procedure Planning from Instructional Videos with Text Supervision

ICCV 2023poster

In this work, we focus on the task of procedure planning from instructional videos with text supervision, where a model aims to predict an action sequence to transform the initial visual state into the goal visual state. A critical challenge of this task is the large semantic gap between observed vi…

Cited by 20PDFScholar
2023

Generating Anomalies for Video Anomaly Detection With Prompt-Based Feature Mapping

CVPR 2023poster

Anomaly detection in surveillance videos is a challenging computer vision task where only normal videos are available during training. Recent work released the first virtual anomaly detection dataset to assist real-world detection. However, an anomaly gap exists because the anomalies are bounded in…

Cited by 42SourcePDFScholar
2021

Graph-Based High-Order Relation Modeling for Long-Term Action Recognition

CVPR 2021poster

Long-term actions involve many important visual concepts, e.g., objects, motions, and sub-actions, and there are various relations among these concepts, which we call basic relations. These basic relations will jointly affect each other during the temporal evolution of long-term actions, which forms…

Cited by 67PDFScholar