← Search

Da Chen

19 accepted papers

2026

Beyond Reassembly: Fractured Object Recovery with Missing Parts

CVPR 2026

We propose a novel learning-based task named fractured object recovery. Unlike the previous fractured object reassembly task that only aligns existing parts with overlaps, our task aims to recover the complete shape by not only reassembling irrelevant parts but also predicting missing parts. Our tas

Cited by 0SourceScholar
2026

EchoBat: Echo-Vision Enhancement and Echo-Layered Sampling for Video LLMs Hallucination Mitigation

AAAI 2026technical

Recent advancements in multimodal large language models (MLLMs) have shown remarkable progress in video understanding. However, video MLLMs (VideoMLLMs) still suffer from hallucinations, generating nonsensical or irrelevant content. This issue partly stems from over-reliance on pre-trained knowledge

Cited by 0SourcePDFScholar
2026

ExPO-HM: Learning to Explain-then-Detect for Hateful Meme Detection

ICLR 2026poster

Hateful memes have emerged as a particularly challenging form of online abuse, motivating the development of automated detection systems. Most prior approaches rely on direct detection, producing only binary predictions. Such models fail to provide the context and explanations that real-world modera…

Cited by 0SourcecodeScholar
2026

Interpretability Transfer from Language to Vision via Sparse Autoencoders

ICML 2026poster

Recent advances in language model interpretability using sparse autoencoders (SAEs) have yet to effectively translate to the visual domain, mainly due to the difficulty and ambiguity of labeling visual concepts. In this paper, we introduce Visual Interpretability via SAE Transfer Alignment (VISTA), …

Cited by 0SourceScholar
2026

POLIA: Policy Optimization with Visual-Object-Level Intrinsic Advantage for Multimodal Reasoning

ICML 2026poster

Recent advances in group-based reinforcement learning (RL) greatly improve LLMs' ability in text reasoning. Yet, these methods lack sufficient modeling of multimodal information, leading to significant reasoning hallucination. In this work, we propose POLIA, a novel group-based RL method with visual…

Cited by 0SourceScholar
2025

Rethinking Few Shot CLIP Benchmarks: A Critical Analysis in the Inductive Setting

ICCV 2025poster

CLIP is a foundational model with transferable classification performance in the few-shot setting. Several methods have shown improved performance of CLIP using few-shot examples. However, so far all these techniques have been benchmarked using standard few-shot datasets. We argue that this mode of…

2025

ShieldHead: Decoding-time Safeguard for Large Language Models

ACL 2025finding

In light of the widespread deployment of Large Language Models (LLMs), the responsibility for safeguarding and regulating LLM-generated content has taken on heightened significance. Recent advancements in LLM-based moderation methods, e.g., LlamaGuard, have demonstrated remarkable promise in identif…

2023

COCO-O: A Benchmark for Object Detectors under Natural Distribution Shifts

ICCV 2023poster

Practical object detection application can lose its effectiveness on image inputs with natural distribution shifts. This problem leads the research community to pay more attention on the robustness of detectors under Out-Of-Distribution (OOD) inputs. Existing works construct datasets to benchmark th…

Cited by 23PDFcodeScholar
2023

Prompt Switch: Efficient CLIP Adaptation for Text-Video Retrieval

ICCV 2023poster

In text-video retrieval, recent works have benefited from the powerful learning capabilities of pre-trained text-image foundation models (e.g., CLIP) by adapting them to the video domain. A critical problem for them is how to effectively capture the rich semantics inside the video using the image en…

Cited by 40PDFcodeScholar
2021

A New Tubular Structure Tracking Algorithm Based On Curvature-Penalized Perceptual Grouping

ICASSP 2021accepted

In this paper, we propose a new minimal path-based framework for minimally interactive tubular structure tracking in conjunction with a perceptual grouping scheme. The minimal path models have shown great advantages in tubular structures tracing. However, they suffer from shortcuts or short branches…

Cited by 0SourceScholar
2021

Multiple Pairwise Ranking Networks for Personalized Video Summarization

ICCV 2021poster

In this paper, we investigate video summarization in the supervised setting. Since video summarization is subjective to the preference of the end-user, the design of a unique model is limited. In this work, we propose a model that provides personalized video summaries by conditioning the summarizati…

Cited by 29PDFScholar
2021

Self-Supervised Learning for Few-Shot Image Classification

ICASSP 2021accepted

Few-shot image classification aims to classify unseen classes with limited labelled samples. Recent works benefit from the meta-learning process with episodic tasks and can fast adapt to class from training to testing. Due to the limited number of samples for each task, the initial embedding network…

Cited by 0SourceScholar
2020

Hierarchical Sequence Representation with Graph Network

ICASSP 2020accepted

Video classification problem is a challenging task in computer vision. The performance of this task is highly relied on the scale of training data and the effectiveness of video embedding via a robust embedding network. Unsupervised solutions such as feature average pooling technique, as a simple la…

Cited by 0SourceScholar
2016

A New Finsler Minimal Path Model With Curvature Penalization for Image Segmentation and Closed Contour Detection

CVPR 2016poster

In this paper, we propose a new curvature penalized minimal path model for image segmentation via closed contour detection based on the weighted Euler elastica curves, firstly introduced to the field of computer vision in [22]. Our image segmentation method extracts a collection of curvature penaliz…

Cited by 20PDFScholar