← Search

Xinhua Zeng

6 accepted papers

2025

Guiding Inter-domain Class Balancing With Salient Features For Domain Adaptive Object Detection

ICASSP 2025accepted

Although multi-scale alignment has improved domain adaptive object detection by addressing data distribution differences and annotation challenges, little attention has been given to class distribution differences between domains. Additionally, the utilization of feature information across different…

Cited by 0SourceScholar
2025

Stephanie: Step-by-Step Dialogues for Mimicking Human Interactions in Social Conversations

NAACL 2025findings

In the rapidly evolving field of natural language processing, dialogue systems primarily employ a single-step dialogue paradigm. Although this paradigm is commonly adopted, it lacks the depth and fluidity of human interactions and does not appear natural. We introduce a novel **Step**-by-Step Dialog…

Cited by 2SourcePDFScholar
2024

Denoising Diffusion-Augmented Hybrid Video Anomaly Detection via Reconstructing Noised Frames

IJCAI 2024poster

Video Anomaly Detection (VAD) is crucial for enhancing security and surveillance systems through automatic identification of irregular events, thereby enabling timely responses and augmenting overall situational awareness. Although existing methods have achieved decent detection performances on benc…

Cited by 3SourcePDFScholar
2024

DyHGDAT: Dynamic Hypergraph Dual Attention Network for multi-agent trajectory prediction

ICRA 2024poster

Modeling the interactions among agents based on their historical trajectories is key to precise multi-agent trajectory prediction. Hypergraph Convolutional Networks (HGCN) have become a proper choice for capturing high-order interactions among agents in this field. However, most existing works only…

Cited by 0SourceScholar
2023

Spatial-Temporal Graph Convolutional Network Boosted Flow-Frame Prediction For Video Anomaly Detection

ICASSP 2023accepted

Video Anomaly Detection (VAD) is a critical technology for intelligent surveillance systems and remains a challenging task in the signal processing community. An intuitive idea for VAD is to use a two-stream network to learn appearance and motion normality, respectively. However, existing approaches…

Cited by 0SourceScholar
2022

Look, Listen and Pay More Attention: Fusing Multi-Modal Information for Video Violence Detection

ICASSP 2022accepted

Violence detection is an essential and challenging problem in the computer vision community. Most existing works focus on single modal data analysis, which is not effective when multi-modality is available. Therefore, we propose a two-stage multi-modal information fusion method for violence detectio…

Cited by 0SourceScholar