← Search

Shu Zhang

20 accepted papers

2026

Decoding 3D Perception via BrainSSD: Synergistic Fusion of EEG Representations from Static and Dynamic Visual Streams

CVPR 2026

Understanding how the brain constructs coherent 3D visual percepts from multifaceted experiences remains a pivotal yet underexplored challenge. To investigate this, we introduce BrainSSD, a novel framework for decoding 3D representations from electroencephalography (EEG) signals. The core of BrainSS

Cited by 0SourceScholar
2026

EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation

CVPR 2026

Sound effects build an essential layer of multimodal storytelling, shaping the emotional atmosphere and the narrative semantics of videos. Despite recent advancement in video-text-to-audio (VT2A), the current formulation faces three key limitations: (1) an imbalance between visual and textual condit

Cited by 0SourceScholar
2026

Few-Shot Hybrid Incremental Learning:Continually Learning under Data Scarcity and Task Uncertainty

CVPR 2026

The increasing complexity of real-world deployment requires intelligent agents to effectively adapt to non-stationary data streams with stochastic increments under data scarcity. We formally define this challenge as the Few-Shot Hybrid Incremental Learning (FSHIL) paradigm, which reveals a critical

Cited by 0SourceScholar
2026

Structured Diversity Control: A Dual-Level Framework for Group-Aware Multi-Agent Coordination

ICRA 2026poster

Controlling the behavioral diversity is a pivotal challenge in multi-agent reinforcement learning (MARL), particularly in complex collaborative scenarios. While existing methods attempt to regulate behavioral diversity by directly differentiating across all agents, they lack deep characterization an…

2025

HARP: Human-Assisted Regrouping With Permutation Invariant Critic for Multi-Agent Reinforcement Learning

ICRA 2025

Human-in-the-loop reinforcement learning integrates human expertise to accelerate agent learning and provide critical guidance and feedback in complex fields. However, many existing approaches focus on single-agent tasks and require continuous human involvement during the training process, significa

Cited by 1SourcecodeScholar
2025

MDD-5k: A New Diagnostic Conversation Dataset for Mental Disorders Synthesized via Neuro-Symbolic LLM Agents

AAAI 2025technical

The clinical diagnosis of most mental disorders primarily relies on the conversations between psychiatrist and patient. The creation of such diagnostic conversation datasets is promising to boost the AI mental healthcare community. However, directly collecting the conversations in real diagnosis sce…

2024

A Method for X-Ray Image Landmarks Localization using Cyclic Coordinate-Guided Strategy

ICASSP 2024accepted

In this study, we present a novel method for pinpointing landmarks in X-ray images, which simultaneously offers computational efficiency and localization precision. Our method leverages a cyclic coordinate-guided strategy that requires fewer model parameters and lower computational costs than tradit…

Cited by 0SourceScholar
2024

HIVE: Harnessing Human Feedback for Instructional Visual Editing

CVPR 2024poster

Incorporating human feedback has been shown to be crucial to align text generated by large language models to human preferences. We hypothesize that state-of-the-art instructional image editing models where outputs are generated based on an input image and an editing instruction could similarly bene…

2024

ULIP-2: Towards Scalable Multimodal Pre-training for 3D Understanding

CVPR 2024poster

Recent advancements in multimodal pre-training have shown promising efficacy in 3D representation learning by aligning multimodal features across 3D shapes their 2D counterparts and language descriptions. However the methods used by existing frameworks to curate such multimodal data in particular la…

2023

Fairness-guided Few-shot Prompting for Large Language Models

NeurIPS 2023poster

Large language models have demonstrated surprising ability to perform in-context learning, i.e., these models can be directly applied to solve numerous downstream tasks by conditioning on a prompt constructed by a few input-output examples. However, prior research has shown that in-context learning…

Cited by 82SourcePDFScholar
2023

GlueGen: Plug and Play Multi-modal Encoders for X-to-image Generation

ICCV 2023poster

Text-to-image (T2I) models based on diffusion processes have achieved remarkable success in controllable image generation using user-provided captions. However, the tight coupling between the current text encoder and image decoder in T2I models makes it challenging to replace or upgrade. Such change…

Cited by 26PDFcodeScholar
2023

UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild

NeurIPS 2023poster

Achieving machine autonomy and human control often represent divergent objectives in the design of interactive AI systems. Visual generative foundation models such as Stable Diffusion show promise in navigating these goals, especially when prompted with arbitrary languages. However, they often fall…

2022

Use All the Labels: A Hierarchical Multi-Label Contrastive Learning Framework

CVPR 2022poster

Current contrastive learning frameworks focus on leveraging a single supervisory signal to learn representations, which limits the efficacy on unseen data and downstream tasks. In this paper, we present a hierarchical multi-label representation learning framework that can leverage all available labe…

Cited by 102PDFcodeScholar
2019

Heterogeneous Memory Enhanced Multimodal Attention Model for Video Question Answering

CVPR 2019poster

In this paper, we propose a novel end-to-end trainable Video Question Answering (VideoQA) framework with three major components: 1) a new heterogeneous memory which can effectively learn global context information from appearance and motion features; 2) a redesigned question memory which helps under…

Cited by 342PDFcodeScholar
2019

Towards Rich Feature Discovery With Class Activation Maps Augmentation for Person Re-Identification

CVPR 2019poster

The fundamental challenge of small inter-person variation requires Person Re-Identification (Re-ID) models to capture sufficient fine-grained information. This paper proposes to discover diverse discriminative visual cues without extra assistance, e.g., pose estimation, human parsing. Specifically,…

Cited by 311PDFScholar
2017

Beyond Face Rotation: Global and Local Perception GAN for Photorealistic and Identity Preserving Frontal View Synthesis

ICCV 2017poster

Photorealistic frontal view synthesis from a single face image has a wide range of applications in the field of face recognition. Although data-driven deep learning methods have been proposed to address this problem by seeking solutions from ample face data, this problem is still challenging because…

Cited by 849PDFScholar