← Search

Yue Ding

20 accepted papers

2026

AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration

ICLR 2026poster

Audiovisual video captioning aims to generate semantically rich descriptions with temporal alignment between visual and auditory events, thereby benefiting both video understanding and generation. In this paper, we present **AVoCaDO**, a powerful audiovisual video captioner driven by the temporal or…

Cited by 0SourceScholar
2026

Dejavu: Towards Experience Feedback Learning for Embodied Intelligence

CVPR 2026

Embodied agents face a fundamental limitation: once deployed in real-world environments, they cannot easily acquire new knowledge to improve task performance. In this paper, we propose Dejavu, a general post-deployment learning framework that augments a frozen Vision-Language-Action (VLA) policy wit

Cited by 0SourcecodeScholar
2026

MVP-Nav: Multi-layer Value Map Planner Navigator

RSS 2026poster

Zero-Shot Object Goal Navigation (ZSON) is an important task for robots. While Multimodal Large Language Models (MLLMs) have empowered robots with significant semantic reasoning capabilities, current RGB-only navigation methods still struggle to align high-level discrete logic with low-level continu…

Cited by 0SourceScholar
2026

OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models

ICML 2026poster

Omni-modal Large Language Models (Omni-LLMs) have demonstrated strong capabilities in audio-video understanding tasks. However, their reliance on long multimodal token sequences leads to substantial computational overhead. Despite this challenge, token compression methods designed for Omni-LLMs rema…

Cited by 0SourceScholar
2026

RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark

CVPR 2026

The integration of visual understanding and generation into unified multimodal models represents a significant stride toward general-purpose AI. However, a fundamental question remains unanswered by existing benchmarks: does this architectural unification actually enable synergetic interaction betwe

Cited by 0SourcecodeScholar
2026

Recovering Hidden Reward in Diffusion-Based Policies

ICML 2026poster

This paper introduces EnergyFlow, a framework that unifies generative action modeling with inverse reinforcement learning by parameterizing a scalar energy function whose gradient is the denoising field. We establish that under maximum-entropy optimality, the score function learned via denoising sco…

Cited by 0SourceScholar
2026

Towards Safe Reasoning in Large Reasoning Models via Corrective Intervention

ICLR 2026poster

Although Large Reasoning Models (LRMs) have progressed in solving complex problems, their chain-of-thought (CoT) reasoning often contains harmful content that can persist even when the final responses appear safe. We show that this issue still remains in existing methods which overlook the unique si…

Cited by 0SourceScholar
2025

Attention-guided Self-reflection for Zero-shot Hallucination Detection in Large Language Models

EMNLP 2025

Hallucination has emerged as a significant barrier to the effective application of Large Language Models (LLMs). In this work, we introduce a novel Attention-Guided SElf-Reflection (AGSER) approach for zero-shot hallucination detection in LLMs. The AGSER method utilizes attention contributions to ca

Cited by 0SourcePDFScholar
2025

BECAME: Bayesian Continual Learning with Adaptive Model Merging

ICML 2025poster

Continual Learning (CL) strives to learn incrementally across tasks while mitigating catastrophic forgetting. A key challenge in CL is balancing stability (retaining prior knowledge) and plasticity (learning new tasks). While representative gradient projection methods ensure stability, they often li…

2025

Discretized Gaussian Representation for Tomographic Reconstruction

ICCV 2025poster

Computed Tomography (CT) enables detailed cross-sectional imaging but continues to face challenges in balancing reconstruction quality and computational efficiency. While deep learning-based methods have significantly improved image quality and noise reduction, they typically require large-scale tra…

2025

How Does Topology Bias Distort Message Passing in Graph Recommender? A Dirichlet Energy Perspective

NeurIPS 2025poster

Graph-based recommender systems have achieved remarkable effectiveness by modeling high-order interactions between users and items. However, such approaches are significantly undermined by popularity bias, which distorts the interaction graph’s structure—referred to as topology bias. This leads to o…

Cited by 0SourcecodeScholar
2025

SHARP: Steering Hallucination in LVLMs via Representation Engineering

EMNLP 2025

Despite their impressive capabilities, Large Vision-Language Models (LVLMs) frequently generate responses that are plausible but incorrect or unsupported—commonly referred to as hallucinations. In this study, we investigate whether different types of hallucinations are reflected in the model’s inter

Cited by 0SourcePDFScholar
2025

Sparse Meets Dense: Unified Generative Recommendations with Cascaded Sparse-Dense Representations

NeurIPS 2025poster

Generative models have recently gained attention in recommendation systems by directly predicting item identifiers from user interaction sequences. However, existing methods suffer from significant information loss due to the separation of stages such as quantization and sequence modeling, hindering…

Cited by 0SourceScholar
2024

FedHCA2: Towards Hetero-Client Federated Multi-Task Learning

CVPR 2024poster

Federated Learning (FL) enables joint training across distributed clients using their local data privately. Federated Multi-Task Learning (FMTL) builds on FL to handle multiple tasks assuming model congruity that identical model architecture is deployed in each client. To relax this assumption and t…

2024

InterpGNN: Understand and Improve Generalization Ability of Transdutive GNNs through the Lens of Interplay between Train and Test Nodes

ICLR 2024poster

Transductive node prediction has been a popular learning setting in Graph Neural Networks (GNNs). It has been widely observed that the shortage of information flow between the distant nodes and intra-batch nodes (for large-scale graphs) often hurt the generalization of GNNs which overwhelmingly adop…

Cited by 1SourcePDFScholar
2024

Task Indicating Transformer for Task-Conditional Dense Predictions

ICASSP 2024accepted

The task-conditional model is a distinctive stream for efficient multi-task learning. Existing works encounter a critical limitation in learning task-agnostic and task-specific representations, primarily due to shortcomings in global context modeling arising from CNN-based architectures, as well as…

Cited by 0SourceScholar
2024

UNIDEAL: Curriculum Knowledge Distillation Federated Learning

ICASSP 2024accepted

Federated Learning (FL) has emerged as a promising approach to enable collaborative learning among multiple clients while preserving data privacy. However, cross-domain FL tasks, where clients possess data from different domains or distributions, remain a challenging problem due to the inherent hete…

Cited by 0SourceScholar
2024

YOLO-Med : Multi-Task Interaction Network for Biomedical Images

ICASSP 2024accepted

Object detection and semantic segmentation are pivotal components in biomedical image analysis. Current single-task networks exhibit promising outcomes in both detection and segmentation tasks. Multi-task networks have gained prominence due to their capability to simultaneously tackle segmentation a…

Cited by 0SourceScholar
2023

Spammer Detection on Short Video Applications: A new Challenge and Baselines

ICASSP 2023accepted

Users can interact with the advertisements and share their impressions through the review system on short video applications. However, spammers may post false or malicious comments to mislead normal users due to profit-driven reasons, damaging the community’s positive atmosphere. In this paper, we i…

Cited by 0SourceScholar