← Search

Dan Guo

41 accepted papers

2026

A Closer Look at Knowledge Distillation in Spiking Neural Network Training

AAAI 2026technical

Spiking Neural Networks (SNNs) become popular due to excellent energy efficiency, yet facing challenges for effective model training. Recent works improve this by introducing knowledge distillation (KD) techniques, with the pre-trained artificial neural networks (ANNs) used as teachers and the targe

Cited by 2SourcePDFScholar
2026

AgentMental: An Interactive Multi-Agent Framework for Explainable and Adaptive Mental Health Assessment

AAAI 2026technical

Mental health assessment is crucial for early intervention and effective treatment, yet traditional clinician-based approaches are limited by the shortage of qualified professionals. Recent advances in artificial intelligence have sparked growing interest in automated psychological assessment, yet m

Cited by 0SourcePDFScholar
2026

Bidirectional Counterfactual Distillation for Review-Based Recommendation

AAAI 2026technical

Review-based recommendation methods typically integrate multiple behaviors, including interactions, reviews, and ratings, to model user preferences. To effectively extract preference signals from diverse behaviors, some studies train multiple student models to capture distinct behavioral patterns, a

Cited by 0SourcePDFScholar
2026

CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization

AAAI 2026technical

The Dense Audio-Visual Event Localization (DAVEL) task aims to temporally localize events in untrimmed videos that occur simultaneously in both the audio and visual modalities. This paper explores DAVEL under a new and more challenging weakly-supervised setting (W-DAVEL task), where only video-level

Cited by 0SourcePDFScholar
2026

FLOW: Optimal Transport-Driven Feature Warping for Generalized Remote Physiological Measurement

CVPR 2026

Remote photoplethysmography (rPPG) enables non-contact physiological measurement from facial videos but often suffers from severe performance degradation under domain shifts. Traditional STMap-based methods [??] rely on predefined spatio-temporal representations that offer engineered robustness but

Cited by 0SourceScholar
2026

Face-Guided Sentiment Boundary Enhancement for Weakly-Supervised Temporal Sentiment Localization

CVPR 2026

Point-level weakly-supervised temporal sentiment localization (P-WTSL) aims to detect sentiment-relevant segments in untrimmed multimodal videos using timestamp sentiment annotations, which greatly reduces the costly frame-level labeling. To tackle the intrinsic challenges of imprecise sentiment bou

Cited by 0SourcecodeScholar
2026

Generalizable and Actionable Parts Pose Estimation with Symmetry Annotation-Free Learning Strategy

ICML 2026poster

Urgently needed generalizable robot object interaction and manipulation requires high-quality Cross-Category object perception. As a pioneer of this area, Generalizable and Actionable Parts (GAParts) understanding has attracted increasing attention from relevant researchers. However, most recent wor…

Cited by 0SourceScholar
2026

LinProVSR: Linguistics-Knowledge Guided Progressive Disambiguation Network for Visual Speech Recognition

AAAI 2026technical

Visual Speech Recognition (VSR), commonly known as lipreading, enables the recognition of spoken text by analyzing lip visual features. Due to the subtlety of lip movements, its recognition is much harder than other motion recognition tasks. Existing VSR models face the challenge of viseme ambiguit

Cited by 0SourcePDFScholar
2026

PHASE-Net: Physics-Grounded Harmonic Attention System for Efficient Remote Photoplethysmography Measurement

CVPR 2026

Remote photoplethysmography (rPPG) measurement enables non-contact physiological monitoring but suffers from accuracy degradation under head motion and illumination changes. Existing deep learning methods are mostly heuristic and lack theoretical grounding, limiting robustness and interpretability.

Cited by 0SourcecodeScholar
2026

SIAM: Towards Generalizable Articulated Object Modeling via Single Robot-Object Interaction

AAAI 2026technical

Articulated object modeling, which represents interconnected rigid bodies with their geometry, part segmentation, articulation tree, and physical properties, is crucial for robotic perception and manipulation. Recently existing methods like SAGCI leverage Interactive Perception (IP) to refine models

Cited by 0SourcePDFScholar
2026

SIMTOKEN: A SIMPLE BASELINE FOR REFERRING AUDIO-VISUAL SEGMENTATION

ICASSP 2026poster

Referring Audio-Visual Segmentation (Ref-AVS) aims to segment specific objects in videos based on natural language expressions involving audio, vision, and text information. This task poses significant challenges in cross-modal reasoning and fine-grained object localization. In this paper, we propos…

Cited by 0SourcePDFScholar
2026

See the Emotion: A Facial Emoji Proxy Modeling for EEG Emotion Recognition

ICML 2026poster

Despite the high accuracy of EEG-based emotion recognition, existing models remain opaque "black boxes", lacking semantic grounding between abstract neural features and human-interpretable states. In this paper, we reframe EEG explainability as a cross-modal generation task, shifting the paradigm fr…

Cited by 0SourceScholar
2025

AugRefer: Advancing 3D Visual Grounding via Cross-Modal Augmentation and Spatial Relation-based Referring

AAAI 2025technical

3D visual grounding (3DVG), which aims to correlate a natural language description with the target object within a 3D scene, is a significant yet challenging task. Despite recent advancements in this domain, existing approaches commonly encounter a shortage: a limited amount and diversity of text-3D…

Cited by 1SourcePDFScholar
2025

Dense Audio-Visual Event Localization Under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration

AAAI 2025technical

In the field of audio-visual learning, most research tasks focus exclusively on short videos. This paper focuses on the more practical Dense Audio-Visual Event Localization (DAVEL) task, advancing audio-visual scene understanding for longer, untrimmed videos. This task seeks to identify and temporal…

2025

Discrete to Continuous: Generating Smooth Transition Poses from Sign Language Observations

CVPR 2025poster

Generating continuous sign language videos from discrete segments is challenging due to the need for smooth transitions that preserve natural flow and meaning. Traditional approaches that simply concatenate isolated signs often result in abrupt transitions, disrupting video coherence. To address thi…

Cited by 0SourcePDFScholar
2025

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering

CVPR 2025poster

We introduce EgoTextVQA, a novel and rigorously constructed benchmark for egocentric QA assistance involving scene text. EgoTextVQA contains 1.5K ego-view videos and 7K scene-text aware questions that reflect real user needs in outdoor driving and indoor house-keeping activities. The questions are d…

2025

Linguistics-Vision Monotonic Consistent Network for Sign Language Production

ICASSP 2025accepted

Sign Language Production (SLP) aims to generate sign videos corresponding to spoken language sentences, where the conversion of sign Glosses to Poses (G2P) is the key step. Due to the cross-modal semantic gap and the lack of word-action correspondence labels for strong supervision alignment, the SLP…

Cited by 0SourceScholar
2025

MMAD: Multi-label Micro-Action Detection in Videos

ICCV 2025poster

Human body actions are an important form of non-verbal communication in social interactions. This paper specifically focuses on a subset of body actions known as micro-actions, which are subtle, low-intensity body movements with promising applications in human emotion analysis. In real-world scenari…

2025

MOL-Mamba: Enhancing Molecular Representation with Structural & Electronic Insights

AAAI 2025technical

Molecular representation learning plays a crucial role in various downstream tasks, such as molecular property prediction and drug design. To accurately represent molecules, Graph Neural Networks (GNNs) and Graph Transformers (GTs) have shown potential in the realm of self-supervised pretraining. Ho…

2025

Moderating the Generalization of Score-based Generative Model

ICCV 2025poster

Score-based Generative Models (SGMs) have demonstrated remarkable generalization capabilities, e.g. generating unseen, but natural data. However, the greater the generalization power, the more likely the unintended generalization, and the more dangerous the abuse. Despite these concerns, research on…

2025

Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing

AAAI 2025technical

The Audio-Visual Video Parsing task aims to recognize and temporally localize all events occurring in either the audio or visual stream, or both. Capturing accurate event semantics for each audio/visual segment is vital. Prior works directly utilize the extracted holistic audio and visual features f…

2025

Patch-level Sounding Object Tracking for Audio-Visual Question Answering

AAAI 2025technical

Answering questions related to audio-visual scenes, i.e., the AVQA task, is becoming increasingly popular. A critical challenge is accurately identifying and tracking sounding objects related to the question along the timeline. In this paper, we present a new Patch-level Sounding Object Tracking (PS…

Cited by 6SourcePDFScholar
2025

PhysDiff: Physiology-based Dynamicity Disentangled Diffusion Model for Remote Physiological Measurement

AAAI 2025technical

Recent works on remote PhotoPlethysmoGraphy (rPPG) estimation typically use techniques like CNNs and Transformers to encode implicit features from facial videos for prediction. These methods learn to directly map facial videos to the static values of rPPG signals, overlooking the inherent dynamic ch…

2025

Prototypical Calibrating Ambiguous Samples for Micro-Action Recognition

AAAI 2025technical

Micro-Action Recognition (MAR) has gained increasing attention due to its crucial role as a form of non-verbal communication in social interactions, with promising potential for applications in human communication and emotion analysis. However, current approaches often overlook the inherent ambiguit…

2025

Sign-IDD: Iconicity Disentangled Diffusion for Sign Language Production

AAAI 2025technical

Sign Language Production (SLP) aims to generate semantically consistent sign videos from textual statements, where the conversion from textual glosses to sign poses (G2P) is a crucial step. Existing G2P methods typically treat sign poses as discrete three-dimensional coordinates and directly fit the…

2025

Text-Infused Audio-Visual Video Parsing with Semantic-Aware Multimodal Contrastive Learning

ICASSP 2025accepted

The Audio-Visual Video Parsing task aims to recognize events occurring in video segments for each modality. Presently, the excellent performance in handling video parsing is shown by generating pseudo labels at the segment level. However, these approaches still suffer from adequate semantic learning…

Cited by 0SourceScholar
2025

Towards Open-Vocabulary Audio-Visual Event Localization

CVPR 2025poster

The Audio-Visual Event Localization (AVEL) task aims to temporally locate and classify video events that are both audible and visible.Most research in this field assumes a closed-set setting, which restricts these models' ability to handle test data containing event categories absent (unseen) during…

2024

EulerMormer: Robust Eulerian Motion Magnification via Dynamic Filtering within Transformer

AAAI 2024technical

Video Motion Magnification (VMM) aims to break the resolution limit of human visual perception capability and reveal the imperceptible minor motion that contains valuable information in the macroscopic domain. However, challenges arise in this task due to photon noise inevitably introduced by photog…

2024

Frequency Decoupling for Motion Magnification via Multi-Level Isomorphic Architecture

CVPR 2024poster

Video Motion Magnification (VMM) aims to reveal subtle and imperceptible motion information of objects in the macroscopic world. Prior methods directly model the motion field from the Eulerian perspective by Representation Learning that separates shape and texture or Multi-domain Learning from phase…

2024

KPA-Tracker: Towards Robust and Real-Time Category-Level Articulated Object 6D Pose Tracking

AAAI 2024technical

Our life is populated with articulated objects. Current category-level articulation estimation works largely focus on predicting part-level 6D poses on static point cloud observations. In this paper, we tackle the problem of category-level online robust and real-time 6D pose tracking of articulated…

2024

Object-Aware Adaptive-Positivity Learning for Audio-Visual Question Answering

AAAI 2024technical

This paper focuses on the Audio-Visual Question Answering (AVQA) task that aims to answer questions derived from untrimmed audible videos. To generate accurate answers, an AVQA model is expected to find the most informative audio-visual clues relevant to the given questions. In this paper, we propos…

2024

Text-Based Occluded Person Re-identification via Multi-Granularity Contrastive Consistency Learning

AAAI 2024technical

Text-based Person Re-identification (T-ReID), which aims at retrieving a specific pedestrian image from a collection of images via text-based information, has received significant attention. However, previous research has overlooked a challenging yet practical form of T-ReID: dealing with image gall…

2024

Towards Understanding Future: Consistency Guided Probabilistic Modeling for Action Anticipation

AAAI 2024technical

Action anticipation aims to infer the action in the unobserved segment (future segment) with the observed segment (past segment). Existing methods focus on learning key past semantics to predict the future, but they do not model the temporal continuity between the past and the future. However, past…

Cited by 3SourcePDFScholar
2022

A Label-Aware Autoregressive Framework for Cross-Domain NER

NAACL 2022findings

Cross-domain named entity recognition (NER) aims to borrow the entity information from the source domain to help the entity recognition in the target domain with limited labeled data. Despite the promising performance of existing approaches, most of them focus on reducing the discrepancy of token re…

2022

Audio—Visual Segmentation

ECCV 2022poster

"We propose to explore a new problem called audio-visual segmentation (AVS), in which the goal is to output a pixel-level map of the object(s) that produce sound at the time of the image frame. To facilitate this research, we construct the first audio-visual segmentation benchmark (AVSBench), provid…