← Search

Ye Zhang

28 accepted papers

2026

Beyond Missing Modalities: Hypergraph Conditioned Diffusion for Uncertainty-Aware Multimodal Emotion Recognition

CVPR 2026

Multimodal Emotion Recognition in Conversations (MERC) aims to understand emotions expressed in each utterance by effectively integrating audio, text, and visual modalities. However, in real-world scenarios, unavoidable missing modalities often degrade multimodal interpretation performance. To addre

Cited by 0SourceScholar
2026

D2MDM2: A Brain-Inspired Deep Network Based on DDM Decision-Making Mechanism for Remote Sensing Change Detection

IJCAI 2026

Remote sensing change detection (RSCD) aims to identify changed regions in bitemporal images. However, conventional one-step modeling suffers from performance degradation caused by imaging temporal differences (e.g., illumination disturbances, seasonal variations). To address this issue, we formulat

Cited by 0Scholar
2026

MangoBench: A Benchmark for Multi-Agent Goal-Conditioned Offline Reinforcement Learning

CVPR 2026

Offline Multi-Agent Reinforcement Learning (MARL) is critical for coordinating multiple agents in costly and unsafe environments, yet existing methods struggle with high sensitivity to reward functions and weak generalization to new goals, limiting its practical impact. Inspired by single-agent Offl

Cited by 0SourceScholar
2026

PaQ-DETR: Learning Pattern and Quality-Aware Dynamic Queries for Object Detection

CVPR 2026

Detection Transformer (DETR) has redefined object detection by casting it as a set prediction task within an end-to-end framework. Despite its elegance, DETR and its variants still rely on fixed learnable queries and suffer from severe query utilization imbalance, which limits adaptability and leave

Cited by 0SourceScholar
2026

Seeing Motion, Generating Action: Explicit Motion-Aware Policy for Robotic Action Generation

ICRA 2026poster

Imitation learning (IL) offers a scalable framework for teaching robots complex manipulation skills from human demonstrations. However, conventional end-to-end visuomotor IL models often suffer from poor performance and robustness due to the significant modality mismatch between high-dimensional vis…

Cited by 0Scholar
2025

$U2$ Frame: A Unified and Unsupervised Learning Framework for LiDAR-Based Loop Closing

ICRA 2025

Loop closing is critically important in Simultaneous Localization and Mapping (SLAM) due to its ability to correct accumulated localization errors. However, existing methods are hindered by the difficulty of acquiring pose labels and the unreliability of ground truth data. In this paper, we propose

Cited by 0SourcecodeScholar
2025

3D Whole-Body Pose Estimation Using Graph High-Resolution Network for Humanoid Robot Teleoperation

ICRA 2025

In the realm of robotics, teleoperation plays a pivotal role in performing high-risk or intricate tasks, and obtaining precise 3D whole-body pose is crucial for this purpose. Traditional two-stage methods have limitations in estimating different body parts, leading to complex systems and higher esti

Cited by 0SourcecodeScholar
2025

AIQViT: Architecture-Informed Post-Training Quantization for Vision Transformers

AAAI 2025technical

Post-training quantization (PTQ) has emerged as a promising solution for reducing the storage and computational cost of vision transformers (ViTs). Recent advances primarily target at crafting quantizers to deal with peculiar activations characterized by ViTs. However, most existing methods underest…

Cited by 0SourcePDFScholar
2025

CPath-Omni: A Unified Multimodal Foundation Model for Patch and Whole Slide Image Analysis in Computational Pathology

CVPR 2025poster

The emergence of large multimodal models (LMMs) has brought significant advancements to pathology. Previous research has primarily focused on separately training patch-level and whole-slide image (WSI)-level models, limiting the integration of learned knowledge across patches and WSIs and resulting…

2025

Category Prompt Mamba Network for Nuclei Segmentation and Classification

AAAI 2025technical

Nuclei segmentation and classification provide an essential basis for tumor immune microenvironment analysis. The previous nuclei segmentation and classification models require splitting large images into smaller patches for training, leading to two significant issues. First, nuclei at the borders o…

Cited by 0SourcePDFScholar
2025

Correlated Multiple IHC Virtual Staining for Breast Histopathological Images

ICASSP 2025accepted

Immunohistochemistry (IHC) examination is essential for determining breast cancer subtypes and provides critical prognostic factors to guide treatment decisions. However, the complex and expensive preparation of IHC staining limits its widespread use in clinical practice. Recent advancements in gene…

Cited by 0SourceScholar
2025

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation

EMNLP 2025

Intent detection, a core component of natural language understanding, has considerably evolved as a crucial mechanism in safeguarding large language models (LLMs). While prior work has applied intent detection to enhance LLMs’ moderation guardrails, showing a significant success against content-leve

Cited by 0SourcePDFScholar
2025

Federated Dialogue-Semantic Diffusion for Emotion Recognition under Incomplete Modalities

NeurIPS 2025poster

Multimodal Emotion Recognition in Conversations (MERC) enhances emotional understanding through the fusion of multimodal signals. However, unpredictable modality absence in real-world scenarios significantly degrades the performance of existing methods. Conventional missing-modality recovery approac…

Cited by 0SourceScholar
2025

Multi-Modality Test-Time Adaptation for Semantic Segmentation in Robotic Perception

ICRA 2025

Test-Time Adaptation (TTA) adjusts pre-trained models in unlabeled unseen environments during the test phase, making it more practical for robotic applications. However, the constant changes of the physical world create significant domain gaps between the received data during robot deployment and th

Cited by 0SourceScholar
2025

Multi-scale Context Intertwining for Panoramic Renal Pathology Segmentation

ICASSP 2025accepted

Panoramic segmentation of renal pathological tissues plays a crucial role in diagnosing renal carcinoma and other kidney-related diseases. The multi-scale nature of kidney tissues, which requires different magnification levels for accurate analysis, presents a significant challenge for segmentation…

Cited by 0SourceScholar
2025

Multi-type MOOCs Recommendation: Leveraging Deep Multi-Relational Representation and Hierarchical Reasoning

AAAI 2025technical

Massive open online courses (MOOCs) recommendation provides online courses tailored to learners' individual preferences. Existing literature is limited by: 1) Ignoring the interrelations among courses, knowledge concepts, and videos, which leads to suboptimal recommendation performance; 2) Neglectin…

Cited by 0SourcePDFScholar
2025

OT-StainNet: Optimal Transport Driven Semantic Matching for Weakly Paired H&E-to-IHC Stain Transfer

AAAI 2025technical

Immunohistochemistry (IHC) examination is essential for characterizing tumor subtypes, providing prognostic information, and developing personalized treatment plans. However, IHC staining preparation is more complex and expensive compared to Hematoxylin and Eosin (H&E) staining, limiting its widespr…

Cited by 0SourcePDFScholar
2025

Progressive Correspondence Regenerator for Robust 3D Registration

CVPR 2025poster

Obtaining enough high-quality correspondences is crucial for robust registration. Existing correspondence refinement methods mostly follow the paradigm of outlier removal, which either fails to correctly identify the accurate correspondences under extreme outlier ratios, or select too few correct co…

2025

Revolutionizing Encrypted Traffic Classification with MH-Net: A Multi-View Heterogeneous Graph Model

AAAI 2025technical

With the growing significance of network security, the classification of encrypted traffic has emerged as an urgent challenge. Traditional byte-based traffic analysis methods are constrained by the rigid granularity of information and fail to fully exploit the diverse correlations between bytes. To…

2025

SaMam: Style-aware State Space Model for Arbitrary Image Style Transfer

CVPR 2025highlight

Global effective receptive field plays a crucial role for image style transfer (ST) to obtain high-quality stylized results. However, existing ST backbones (e.g., CNNs and Transformers) suffer huge computational complexity to achieve global receptive fields. Recently, the State Space Model (SSM), es…

2025

Self-Distilled Stereo Matching: Real-Time Domain Generalization for Robotic Depth Perception

IROS 2025

While human vision inherently achieves robust cross-domain depth estimation through binocular coordination, robotic systems employing stereo matching still confront significant challenges in maintaining robustness across domains when performing real-time environmental depth perception. Furthermore,

Cited by 0SourceScholar
2025

The Four Color Theorem for Cell Instance Segmentation

ICML 2025poster

Cell instance segmentation is critical to analyzing biomedical images, yet accurately distinguishing tightly touching cells remains a persistent challenge. Existing instance segmentation frameworks, including detection-based, contour-based, and distance mapping-based approaches, have made significan…

2024

Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

EMNLP 2024finding

Large language models (LLMs) are increasingly being adopted in a wide range of real-world applications. Despite their impressive performance, recent studies have shown that LLMs are vulnerable to deliberately crafted adversarial prompts even when aligned via Reinforcement Learning from Human Feedbac…

2024

LoS: Local Structure-Guided Stereo Matching

CVPR 2024poster

Estimating disparities in challenging areas is difficult and limits the performance of stereo matching models. In this paper we exploit local structure information (LSI) to enhance stereo matching. Specifically our LSI comprises a series of key elements including the slant plane (parameterised by di…

Cited by 14SourcePDFScholar
2024

MRMLREC: A Two-Stage Approach for Addressing Data Sparsity in MOOC Video Recommendation (Student Abstract)

AAAI 2024technical

With the abundance of learning resources available on massive open online courses (MOOCs) platforms, the issue of interactive data sparsity has emerged as a significant challenge.This paper introduces MRMLREC, an efficient MOOC video recommendation which consists of two main stages: multi-relational…

Cited by 4SourcePDFScholar
2021

Effective Sequence-to-Sequence Dialogue State Tracking

EMNLP 2021main

Sequence-to-sequence models have been applied to a wide variety of NLP tasks, but how to properly use them for dialogue state tracking has not been systematically investigated. In this paper, we study this problem from the perspectives of pre-training objectives as well as the formats of context rep…

2021

Graph-Based Asynchronous Event Processing for Rapid Object Recognition

ICCV 2021poster

Different from traditional video cameras, event cameras capture asynchronous events stream in which each event encodes pixel location, trigger time, and the polarity of the brightness changes. In this paper, we introduce a novel graph-based framework for event cameras, namely SlideGCN. Unlike some r…

Cited by 110PDFScholar
2016

Energy-efficient pilot and data power allocation in massive MIMO communication systems based on MMSE channel estimation

ICASSP 2016accepted

This paper addresses the pilot and data power allocation issue in time division duplexing (TDD) massive multi-user multiple-input multiple-output (MU-MIMO) systems. By using minimum mean square error (MMSE) channel estimation along with a maximum-ratio combining (MRC) detector for the uplink transmi…

Cited by 0SourceScholar