← Search

Chenyou Fan

13 accepted papers

2026

A Novel Fine-Tuned CLIP-OOD Detection Method with Double Loss Constraint Through Optimal Transport Semantic Alignment

AAAI 2026technical

Detecting Out-Of-Distribution (OOD) samples in image classification is crucial for model reliability. With the rise of Vision-Language Models (VLMs), CLIP-OOD has become a research hotspot. However, we observe the Low Focus Attention phenomenon from the image encoders of CLIP, which means the attent

Cited by 0SourcePDFScholar
2025

Forward KL Regularized Preference Optimization for Aligning Diffusion Policies

AAAI 2025technical

Diffusion models have achieved remarkable success in sequential decision-making by leveraging the highly expressive model capabilities in policy learning. A central problem for learning diffusion policies is to align the policy output with human intents in various tasks. To achieve this, previous me…

Cited by 3SourcePDFScholar
2025

Heterogeneous Federated Learning with Scalable Server Mixture-of-Experts

IJCAI 2025

Classical Federated Learning (FL) encounters significant challenges when deploying large models on power-constrained clients. To tackle this, we propose an asymmetric FL mechanism that enables the aggregation of compact client models into a comprehensive server model. We design the server model as a

Cited by 0SourcePDFScholar
2025

Hybrid Feature Fusion for Enhancing Medical Document Embedding

ICASSP 2025accepted

Despite the strong capabilities of large language models in generative tasks, issues related to information unreliability and hallucinations pose significant challenges in high-precision fields, such as drug analysis and recommendations in the medical domain. In this work, we introduce the HFFN mode…

Cited by 0SourceScholar
2025

Learning Robust Stereo Matching in the Wild with Selective Mixture-of-Experts

ICCV 2025poster

Recently, learning-based stereo matching networks have advanced significantly.However, they often lack robustness and struggle to achieve impressive cross-domain performance due to domain shifts and imbalanced disparity distributions among diverse datasets.Leveraging Vision Foundation Models (VFMs)…

2025

Task-Agnostic Pre-training and Task-Guided Fine-tuning for Versatile Diffusion Planner

ICML 2025poster

Diffusion models have demonstrated their capabilities in modeling trajectories of multi-tasks. However, existing multi-task planners or policies typically rely on task-specific demonstrations via multi-task imitation, or require task-specific reward labels to facilitate policy optimization via Reinf…

Cited by 0SourcePDFScholar
2022

Where to Attack: A Dynamic Locator Model for Backdoor Attack in Text Classifications

COLING 2022main

Nowadays, deep-learning based NLP models are usually trained with large-scale third-party data which can be easily injected with malicious backdoors. Thus, BackDoor Attack (BDA) study has become a trending research to help promote the robustness of an NLP system. Text-based BDA aims to train a poiso…

2020

Projection Robust Wasserstein Distance and Riemannian Optimization

NeurIPS 2020spotlight

Projection robust Wasserstein (PRW) distance, or Wasserstein projection pursuit (WPP), is a robust variant of the Wasserstein distance. Recent work suggests that this quantity is more robust than the standard Wasserstein distance, in particular when comparing probability measures in high-dimensions.…

2019

Heterogeneous Memory Enhanced Multimodal Attention Model for Video Question Answering

CVPR 2019poster

In this paper, we propose a novel end-to-end trainable Video Question Answering (VideoQA) framework with three major components: 1) a new heterogeneous memory which can effectively learn global context information from appearance and motion features; 2) a redesigned question memory which helps under…

Cited by 342PDFcodeScholar
2018

Joint Person Segmentation and Identification in Synchronized First- and Third-person Videos

ECCV 2018poster

In a world of pervasive cameras, public spaces are often captured from multiple perspectives by cameras of different types, both fixed and mobile. An important problem is to organize these heterogeneous collections of videos by finding connections between them, such as identifying correspondences be…

Cited by 46SourcePDFScholar
2017

Identifying First-Person Camera Wearers in Third-Person Videos

CVPR 2017poster

We consider scenarios in which we wish to perform joint scene understanding, object tracking, activity recognition, and other tasks in scenarios in which multiple people are wearing body-worn cameras while a third-person static camera also captures the scene. To do this, we need to establ…

Cited by 77PDFScholar