← Search

Dan Zeng

23 accepted papers

2026

Beyond Predictive Resampling: Learning Input-Agnostic Downsampling for Efficient Aligned Vision Recognition

AAAI 2026technical

Images are typically sampled on a uniform grid,despite their non-uniform information distribution—some regions are rich in content while others are not. The mismatch leads to inefficient computation allocation in deep learning models. To address this, recent studies have proposed predictive downsamp

Cited by 0SourcePDFScholar
2026

CapeNext: Rethinking and Refining Dynamic Support Information for Category-Agnostic Pose Estimation

AAAI 2026technical

Recent research in Category-Agnostic Pose Estimation (CAPE) has adopted fixed textual keypoint description as semantic prior for two-stage pose matching frameworks. While this paradigm enhances robustness and flexibility by disentangling the dependency of support images, our critical analysis reveal

Cited by 0SourcePDFScholar
2025

3CAD: A Large-Scale Real-World 3C Product Dataset for Unsupervised Anomaly Detection

AAAI 2025technical

Industrial anomaly detection achieves progress thanks to datasets such as MVTec-AD and VisA. However, they suffer from limitations in terms of the number of defect samples, types of defects, and availability of real-world scenes. These constraints inhibit researchers from further exploring the perfo…

2025

Integrating Low-Level Visual Cues for Enhanced Unsupervised Semantic Segmentation

AAAI 2025technical

Unsupervised semantic segmentation algorithms aim to identify meaningful semantic groups without annotations. Recent approaches leveraging self-supervised transformers as pre-training backbones have successfully obtained high-level dense features that effectively express semantic coherence. However,…

Cited by 0SourcePDFScholar
2025

Learning Occlusion-Robust Vision Transformers for Real-Time UAV Tracking

CVPR 2025poster

Single-stream architectures using Vision Transformer (ViT) backbones show great potential for real-time UAV tracking recently. However, frequent occlusions from obstacles like buildings and trees expose a major drawback: these models often lack strategies to handle occlusions effectively. New method…

2025

MambaNUT: Nighttime UAV Tracking via Mamba-based Adaptive Curriculum Learning

IROS 2025

Harnessing low-light enhancement and domain adaptation, nighttime UAV tracking has made substantial strides. However, over-reliance on image enhancement, limited high-quality nighttime data, and a lack of integration between daytime and nighttime trackers hinder the development of an end-to-end trai

Cited by 4SourcecodeScholar
2025

MixPrompt: Efficient Mixed Prompting for Multimodal Semantic Segmentation

NeurIPS 2025poster

Recent advances in multimodal semantic segmentation show that incorporating auxiliary inputs—such as depth or thermal images—can significantly improve performance over single-modality (RGB-only) approaches. However, most existing solutions rely on parallel backbone networks and complex fusion module…

Cited by 0SourceScholar
2024

3DBench: A Scalable 3D Benchmark and Instruction-Tuning Dataset

IJCAI 2024poster

Evaluating the performance of Multi-modal Large Language Models (MLLMs), integrating both point cloud and language, presents significant challenges. The lack of a comprehensive assessment hampers determining whether these models truly represent advancements, thereby impeding further progress in the…

2024

Coupled Confusion Correction: Learning from Crowds with Sparse Annotations

AAAI 2024technical

As the size of the datasets getting larger, accurately annotating such datasets is becoming more impractical due to the expensiveness on both time and economy. Therefore, crowd-sourcing has been widely adopted to alleviate the cost of collecting labels, which also inevitably introduces label noise a…

2024

DiffMOT: A Real-time Diffusion-based Multiple Object Tracker with Non-linear Prediction

CVPR 2024poster

In Multiple Object Tracking objects often exhibit non-linear motion of acceleration and deceleration with irregular direction changes. Tacking-by-detection (TBD) trackers with Kalman Filter motion prediction work well in pedestrian-dominant scenarios but fall short in complex situations when multipl…

Cited by 21SourcePDFScholar
2024

M3D: Dataset Condensation by Minimizing Maximum Mean Discrepancy

AAAI 2024technical

Training state-of-the-art (SOTA) deep models often requires extensive data, resulting in substantial training and storage costs. To address these challenges, dataset condensation has been developed to learn a small synthetic set that preserves essential information from the original large-scale data…

2024

Masked Face Recognition with Generative-to-Discriminative Representations

ICML 2024spotlight

Masked face recognition is important for social good but challenged by diverse occlusions that cause insufficient or inaccurate representations. In this work, we propose a unified deep network to learn generative-to-discriminative representations for facilitating masked face recognition. To this end…

Cited by 4SourcePDFScholar
2023

Adaptive and Background-Aware Vision Transformer for Real-Time UAV Tracking

ICCV 2023poster

While discriminative correlation filters (DCF)-based trackers prevail in UAV tracking for their favorable efficiency, lightweight convolutional neural network (CNN)-based trackers using filter pruning have also demonstrated remarkable efficiency and precision. However, the use of pure vision transfo…

Cited by 38PDFcodeScholar
2023

Analyzing and Combating Attribute Bias for Face Restoration

IJCAI 2023poster

Face restoration (FR) recovers high resolution (HR) faces from low resolution (LR) faces and is challenging due to its ill-posed nature. With years of development, existing methods can produce quality HR faces with realistic details. However, we observe that key facial attributes (e.g., age and gend…

2023

Bootstrapping Multi-View Representations for Fake News Detection

AAAI 2023technical

Previous researches on multimedia fake news detection include a series of complex feature extraction and fusion networks to gather useful information from the news. However, how cross-modal consistency relates to the fidelity of news and how features from different modalities affect the decision-mak…

2023

Model Conversion via Differentially Private Data-Free Distillation

IJCAI 2023poster

While massive valuable deep models trained on large-scale data have been released to facilitate the artificial intelligence community, they may encounter attacks in deployment which leads to privacy leakage of training data. In this work, we propose a learning approach termed differentially private…

2022

Face2Exp: Combating Data Biases for Facial Expression Recognition

CVPR 2022poster

Facial expression recognition (FER) is challenging due to the class imbalance caused by data collection. Existing studies tackle the data bias problem using only labeled facial expression dataset. Orthogonal to existing FER methods, we propose to utilize large unlabeled face recognition (FR) dataset…

Cited by 129PDFcodeScholar
2022

Genre-Conditioned Long-Term 3D Dance Generation Driven by Music

ICASSP 2022accepted

Dancing to music is an artistic behavior of humans, however, letting machines generate dances from music is still challenging. Most existing works have been made progress in tackling the problem of motion prediction conditioned by music, yet they rarely consider the importance of the musical genre.…

Cited by 0SourceScholar
2021

Detecting Deepfake Videos with Temporal Dropout 3DCNN

IJCAI 2021poster

While the abuse of deepfake technology has brought about a serious impact on human society, the detection of deepfake videos is still very challenging due to their highly photorealistic synthesis on each frame. To address that, this paper aims to leverage the possible inconsistent cues among video f…

2021

Neural Architecture Search for Joint Human Parsing and Pose Estimation

ICCV 2021poster

Human parsing and pose estimation are crucial for the understanding of human behaviors. Since these tasks are closely related, employing one unified model to perform two tasks simultaneously allows them to benefit from each other. However, since human parsing is a pixel-wise classification process w…

Cited by 26PDFcodeScholar
2020

Semi-Siamese Training for Shallow Face Learning

ECCV 2020poster

Most existing public face datasets, such as MS-Celeb-1M and VGGFace2, provide abundant information in both breadth (large number of IDs) and depth (sufficient number of samples) for training. However, in many real-world scenarios of face recognition, the training dataset is limited in depth, $ extit…

2018

Disentangling Features in 3D Face Shapes for Joint Face Reconstruction and Recognition

CVPR 2018poster

This paper proposes an encoder-decoder network to disentangle shape features during 3D face shape reconstruction from single 2D images, such that the tasks of learning discriminative shape features for face recognition and reconstructing accurate 3D face shapes can be done simultaneously. Unlike exi…

Cited by 129SourcePDFScholar