← Search

Ruoyu Chen

8 accepted papers

2026

Benchmarking Trustworthiness in Multimodal LLMs for Video Understanding

AAAI 2026technical

Recent advancements in multimodal large language models for video understanding (videoLLMs) have enhanced their capacity to process complex spatiotemporal data. However, challenges such as factual inaccuracies, harmful content, biases, hallucinations, and privacy risks compromise their reliability.

Cited by 0SourcePDFScholar
2026

Deconstructing Positional Information: From Attention Logits to Training Biases

ICLR 2026poster

Positional encodings, a mechanism for incorporating sequential information into the Transformer model, are central to contemporary research on neural architectures. Previous work has largely focused on understanding their function through the principle of distance attenuation, where proximity dictat…

Cited by 0SourceScholar
2026

PhaseWin Search Framework Enable Efficient Object-Level Interpretation

CVPR 2026

Attribution is essential for interpreting object-level foundation models. Recent methods based on submodular subset selection have achieved high faithfulness, but their efficiency limitations hinder practical deployment in real-world scenarios. To address this, we propose PhaseWin, a novel phase-win

Cited by 0SourcecodeScholar
2026

Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation

CVPR 2026

Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in aligning visual inputs with natural language outputs. Yet, the extent to which generated tokens depend on visual modalities remains poorly understood, limiting interpretability and reliability. In this work, we pre

Cited by 0SourcecodeScholar
2025

Interpreting Object-level Foundation Models via Visual Precision Search

CVPR 2025highlight

Advances in multimodal pre-training have propelled object-level foundation models, such as Grounding DINO and Florence-2, in tasks like visual grounding and object detection. However, interpreting these models' decisions has grown increasingly challenging. Existing interpretable attribution methods…

2024

Less is More: Fewer Interpretable Region via Submodular Subset Selection

ICLR 2024oral

Image attribution algorithms aim to identify important regions that are highly relevant to model decisions. Although existing attribution solutions can effectively assign importance to target elements, they still face the following challenges: 1) existing attribution methods generate inaccurate smal…

2024

MCL-NER: Cross-Lingual Named Entity Recognition via Multi-View Contrastive Learning

AAAI 2024technical

Cross-lingual named entity recognition (CrossNER) faces challenges stemming from uneven performance due to the scarcity of multilingual corpora, especially for non-English data. While prior efforts mainly focus on data-driven transfer methods, a significant aspect that has not been fully explored is…

Cited by 18SourcePDFScholar
2024

Tackling the Data Heterogeneity in Asynchronous Federated Learning with Cached Update Calibration

ICLR 2024poster

Asynchronous federated learning, which enables local clients to send their model update asynchronously to the server without waiting for others, has recently emerged for its improved efficiency and scalability over traditional synchronized federated learning. In this paper, we study how the asynchro…

Cited by 18SourcePDFScholar