← Search

Dongming Zhang

6 accepted papers

2026

Cat-PO: Cross-modal Adaptive Token-rewards for Preference Optimization in Truthful Multimodal LLMs

ICLR 2026poster

Multi-modal Large Language Models (MLLMs) have shown remarkable generative capabilities across multi-modal tasks, yet remain plagued by hallucinations where generated textual contents are semantically inconsistent with the input images. This work reveals that existing multi-modal preference optimiza…

Cited by 0SourceScholar
2023

Unilaterally Aggregated Contrastive Learning with Hierarchical Augmentation for Anomaly Detection

ICCV 2023poster

Anomaly detection (AD), aiming to find samples that deviate from the training distribution, is essential in safety-critical applications. Though recent self-supervised learning based attempts achieve promising results by creating virtual outliers, their training objectives are less faithful to AD wh…

Cited by 10PDFScholar
2022

Representation Learning for Compressed Video Action Recognition via Attentive Cross-modal Interaction with Motion Enhancement

IJCAI 2022poster

Compressed video action recognition has recently drawn growing attention, since it remarkably reduces the storage and computational cost via replacing raw videos by sparsely sampled RGB frames and compressed motion cues (e.g., motion vectors and residuals). However, this task severely suffers from t…

Cited by 19SourcePDFScholar
2022

Video Anomaly Detection by Solving Decoupled Spatio-Temporal Jigsaw Puzzles

ECCV 2022poster

"Video Anomaly Detection (VAD) is an important topic in computer vision. Motivated by the recent advances in self-supervised learning, this paper addresses VAD by solving an intuitive yet challenging pretext task, i.e., spatio-temporal jigsaw puzzles, which is cast as a multi-label fine-grained clas…

2019

APE-GAN: Adversarial Perturbation Elimination with GAN

ICASSP 2019accepted

Although Deep Neural Networks could achieve state-of-the-art performance while recongnizing images, they often suffer a tremendous defeat from adversarial examples-inputs generated by utilizing imperceptible but intentional perturbations to samples from the datasets. So far, very few methods have pr…

Cited by 0SourceScholar
2017

Task-Driven Dynamic Fusion: Reducing Ambiguity in Video Description

CVPR 2017poster

Integrating complementary features from multiple channels is expected to solve the description ambiguity problem in video captioning, whereas inappropriate fusion strategies often harm rather than help the performance. Existing static fusion methods in video captioning such as concatenation and summ…

Cited by 84PDFScholar