← Search

Yingxian Chen

5 accepted papers

2026

Learning to See through Illumination Extremes with Event Streaming in Multimodal Large Language Models

CVPR 2026

Multimodal Large Language Models (MLLMs) perform strong vision-language reasoning under standard conditions but fail in extreme illumination, where RGB inputs lose irrevocable structure and semantics. We propose Event-MLLM, an event-enhanced model that performs all-light visual reasoning by dynamica

Cited by 0SourceScholar
2025

Aligning Effective Tokens with Video Anomaly in Large Language Models

ICCV 2025poster

Understanding abnormal events in videos is a vital and challenging task that has garnered significant attention in a wide range of applications. Although current video understanding Multi-modal Large Language Models (MLLMs) are capable of analyzing general videos, they often struggle to handle anoma…

Cited by 0SourcePDFScholar
2024

Can OOD Object Detectors Learn from Foundation Models?

ECCV 2024poster

"Out-of-distribution (OOD) object detection is a challenging task due to the absence of open-set OOD data. Inspired by recent advancements in text-to-image generative models, such as Stable Diffusion, we study the potential of generative models trained on large-scale open-set data to synthesize OOD…

2023

Communication Resources Constrained Hierarchical Federated Learning for End-to-End Autonomous Driving

IROS 2023poster

While federated learning (FL) improves the generalization of end-to-end autonomous driving by model aggregation, the conventional single-hop FL (SFL) suffers from slow convergence rate due to long-range communications among vehicles and cloud server. Hierarchical federated learning (HFL) overcomes s…

Cited by 20SourcecodeScholar
2023

MGFN: Magnitude-Contrastive Glance-and-Focus Network for Weakly-Supervised Video Anomaly Detection

AAAI 2023technical

Weakly supervised detection of anomalies in surveillance videos is a challenging task. Going beyond existing works that have deficient capabilities to localize anomalies in long videos, we propose a novel glance and focus network to effectively integrate spatial-temporal information for accurate ano…