← Search

Chenshuang Zhang

7 accepted papers

2026

Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow

CVPR 2026

Vision-Language Models (VLMs) have demonstrated strong capability in a wide range of tasks such as visual recognition, document parsing, and visual grounding. Nevertheless, recent work shows that while VLMs often manage to capture the correct image region corresponding to the question, they do not n

Cited by 0SourcecodeScholar
2026

SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models

CVPR 2026

Despite the success of audio-visual large-language models (LLMs), they can produce plausible but ungrounded outputs, termed hallucination. Existing benchmarks focus on environmental sounds (e.g., dog barking) to indicate event occurrence. In contrast, human speech carries fundamentally different, ri

Cited by 0SourcecodeScholar
2025

Video Diffusion Models Excel at Tracking Similar-Looking Objects Without Supervision

NeurIPS 2025poster

Distinguishing visually similar objects by their motion remains a critical challenge in computer vision. Although supervised trackers show promise, contemporary self-supervised trackers struggle when visual cues become ambiguous, limiting their scalability and generalization without extensive labele…

Cited by 0SourceScholar
2024

ImageNet-D: Benchmarking Neural Network Robustness on Diffusion Synthetic Object

CVPR 2024highlight

We establish rigorous benchmarks for visual perception robustness. Synthetic images such as ImageNet-C ImageNet-9 and Stylized ImageNet provide specific type of evaluation over synthetic corruptions backgrounds and textures yet those robustness benchmarks are restricted in specified variations and h…

2023

A Survey on Masked Autoencoder for Visual Self-supervised Learning

IJCAI 2023poster

With the increasing popularity of masked autoencoders, self-supervised learning (SSL) in vision undertakes a similar trajectory as in NLP. Specifically, generative pretext tasks with the masked prediction have become a de facto standard SSL practice in NLP (e.g., BERT). By contrast, early attempts a…

Cited by 12SourcePDFScholar
2022

Decoupled Adversarial Contrastive Learning for Self-Supervised Adversarial Robustness

ECCV 2022poster

"\textit{Adversarial training} (AT) for robust representation learning and \textit{self-supervised learning} (SSL) for unsupervised representation learning are two active research fields. Integrating AT into SSL, multiple prior works have accomplished a highly significant yet challenging task: learn…

2022

How Does SimSiam Avoid Collapse Without Negative Samples? A Unified Understanding with Self-supervised Contrastive Learning

ICLR 2022poster

To avoid collapse in self-supervised learning (SSL), a contrastive loss is widely used but often requires a large number of negative samples. Without negative samples yet achieving competitive performance, a recent work~\citep{chen2021exploring} has attracted significant attention for providing a mi…

Cited by 98SourcePDFScholar