← Search

Ning Zhou

7 accepted papers

2026

Composite-Attribute Person Re-Identification via Pose-Guided Disentanglement

CVPR 2026

Recent advancements in vision-language models have enabled multi-modal person re-identification (Re-ID), where the system takes both an image and a text query to identify matching individuals. While previous state-of-the-art methods perform well with detailed, sentence-level descriptions, we found t

Cited by 0SourceScholar
2026

Dynamics: Language-Based Representation for Inferring Rigid-Body Dynamics From Videos

CVPR 2026

Inferring rigid-body physical states and properties from monocular videos is a fundamental step toward physics-based perception and simulation. Existing approaches assume specific underlying physical systems, object types, and camera poses, which are unable to generalize to complex real-world settin

Cited by 0SourceScholar
2026

Reinforcing Structured Chain-of-Thought for Video Understanding

CVPR 2026

Multi-modal Large Language Models (MLLMs) show promise in video understanding. However, their reasoning often suffers from thinking drift and weak temporal comprehension, even when enhanced by Reinforcement Learning (RL) techniques like Group Relative Policy Optimization (GRPO). Moreover, existing R

Cited by 0SourceScholar
2025

Pose as a Modality: A Psychology-Inspired Network for Personality Recognition with a New Multimodal Dataset

AAAI 2025technical

In recent years, predicting Big Five personality traits from multimodal data has received significant attention in artificial intelligence (AI). However, existing computational models often fail to achieve satisfactory performance. Psychological research has shown a strong correlation between pose a…

Cited by 0SourcePDFScholar
2023

PADCLIP: Pseudo-labeling with Adaptive Debiasing in CLIP for Unsupervised Domain Adaptation

ICCV 2023poster

Traditional Unsupervised Domain Adaptation (UDA) leverages the labeled source domain to tackle the learning tasks on the unlabeled target domain. It can be more challenging when a large domain gap exists between the source and the target domain. A more practical setting is to utilize a large-scale p…

Cited by 71PDFScholar
2020

Adaptive Fractional Dilated Convolution Network for Image Aesthetics Assessment

CVPR 2020poster

To leverage deep learning for image aesthetics assessment, one critical but unsolved issue is how to seamlessly incorporate the information of image aspect ratios to learn more robust models. In this paper, an adaptive fractional dilated convolution (AFDC), which is aspect-ratio-embedded, compositio…

Cited by 113PDFScholar