← Search

Cheng Zou

6 accepted papers

2026

ACTIVE-o3 : Empowering MLLMs with Active Perception via Pure Reinforcement Learning

ICML 2026poster

Active vision, also known as active perception, refers to actively selecting where and how to look in order to gather task-relevant information. It is a critical component of efficient perception and decision-making in humans and advanced embodied agents. With the rise of Multimodal Large Language M…

Cited by 0SourceScholar
2025

DS-BTIAN: A Novel Deep-Shallow Bidirectional Transformer Interactive Attention Network for Multimodal Emotion Recognition

ICASSP 2025accepted

In this work, we propose a novel Deep-Shallow Bidirectional Transformer Interactive Attention Network (DS-BTIAN) designed for robust multimodal emotion recognition. DS-BTIAN leverages pre-trained Wav2Vec2.0 and BERT models for efficient feature extraction without the need for finetuning, enhancing r…

Cited by 0SourceScholar
2024

StyleTokenizer: Defining Image Style by a Single Instance for Controlling Diffusion Models

ECCV 2024poster

"Despite the burst of innovative methods for controlling the diffusion process, effectively controlling image styles in text-to-image generation remains a challenging task. Many adapter-based methods impose image representation conditions on the denoising process to accomplish image control. However…

2023

DC-Former: Diverse and Compact Transformer for Person Re-identification

AAAI 2023technical

In person re-identification (ReID) task, it is still challenging to learn discriminative representation by deep learning, due to limited data. Generally speaking, the model will get better performance when increasing the amount of data. The addition of similar classes strengthens the ability of the…

2022

Improving Human-Object Interaction Detection via Phrase Learning and Label Composition

AAAI 2022technical

Human-Object Interaction (HOI) detection is a fundamental task in high-level human-centric scene understanding. We propose PhraseHOI, containing a HOI branch and a novel phrase branch, to leverage language prior and improve relation expression. Specifically, the phrase branch is supervised by semant…

Cited by 44SourcePDFScholar
2021

End-to-End Human Object Interaction Detection With HOI Transformer

CVPR 2021poster

We propose HOI Transformer to tackle human object interaction (HOI) detection in an end-to-end manner. Current approaches either decouple HOI task into separated stages of object detection and interaction classification or introduce surrogate interaction problem. In contrast, our method, named HOI T…

Cited by 266PDFcodeScholar