← Search

Zhe Cao

21 accepted papers

2026

OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs

ICLR 2026poster

Recent advances in multimodal large language models (MLLMs) have demonstrated substantial potential in video understanding. However, existing benchmarks fail to comprehensively evaluate synergistic reasoning capabilities across audio and visual modalities, often neglecting either one of the modaliti…

Cited by 0SourcecodeScholar
2026

RoadSceneBench: A Lightweight Benchmark for Mid-Level Road Scene Understanding

CVPR 2026

Understanding mid-level road semantics, which capture the structural and contextual cues that link low-level perception to high-level planning, is essential for reliable autonomous driving and digital map construction. However, existing benchmarks primarily target perception tasks such as detection

Cited by 0SourcecodeScholar
2026

SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak Attacks

ICLR 2026poster

With the rapid advancement of Large Language Models (LLMs), the safety of LLMs has been a critical concern requiring precise assessment. Current benchmarks primarily concentrate on single-turn dialogues or a single jailbreak attack method to assess the safety. Additionally, these benchmarks have not…

Cited by 0SourcecodeScholar
2026

T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation

ICML 2026poster

Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically synchronized audio from natural language, yet its evaluation remains fragmented, often relying on unimodal metrics or narrowly scoped benchmarks that fail to capture cross-modal alignment, instruction…

Cited by 0SourceScholar
2025

Consensus Graph-Based Spectral Ensemble Clustering via Low-Rank Tensor Learning

ICASSP 2025accepted

Ensemble clustering using co-association matrices integrates multiple base clusterings but often overlooks interactions between crucial samples and base clusterings. This neglect can introduce noise and lead to information loss and instability. To address these issues, we propose the Consensus Graph…

Cited by 0SourceScholar
2025

Dimensionality-Reduced Spatial Bipartite Graph Clustering for Hyperspectral and LiDAR Data

ICASSP 2025accepted

The growing volume of remote sensing (RS) data highlights the need for enhanced data integration and processing. While combining hyperspectral and LiDAR data improves analysis by addressing spectral variability, challenges persist due to the high dimensionality, noise, and outliers in hyperspectral…

Cited by 0SourceScholar
2025

Ego4o: Egocentric Human Motion Capture and Understanding from Multi-Modal Input

CVPR 2025poster

This work focuses on tracking and understanding human motion using consumer wearable devices, such as VR/AR headsets, smart glasses, cellphones, and smartwatches. These devices provide diverse, multi-modal sensor inputs, including egocentric images, and 1-3 sparse IMU sensors in varied combinations.…

Cited by 0SourcePDFScholar
2025

IRGPT: Understanding Real-world Infrared Image with Bi-cross-modal Curriculum on Large-scale Benchmark

ICCV 2025poster

Real-world infrared imagery presents unique challenges for vision-language models due to the scarcity of aligned text data and domain-specific characteristics. Although existing methods have advanced the field, their reliance on synthetic infrared images generated through style transfer from visible…

Cited by 0SourcePDFScholar
2025

MMCSBench: A Fine-Grained Benchmark for Large Vision-Language Models in Camouflage Scenes

NeurIPS 2025poster

Current camouflaged object detection methods predominantly follow discriminative segmentation paradigms and heavily rely on predefined categories present in the training data, limiting their generalization to unseen or emerging camouflage objects. This limitation is further compounded by the labor-i…

Cited by 0SourceScholar
2025

RADCI: A Synchronized Radar-RGBT Object Detecting-Tracking Dataset And A Benchmark

ICASSP 2025accepted

High-quality perception is crucial in autonomous driving and monitoring systems, where millimeter-wave radar and infrared cameras play important roles due to their robustness and reliability under harsh conditions. Both technologies can serve as low-cost supplements to optical image detection, impro…

Cited by 0SourceScholar
2025

Self-Supervised Localized Topology Consistency for Noise-Robust Hyperspectral Image Classification

ICASSP 2025accepted

Label noise in hyperspectral image classification (HIC) can severely degrade model performance by leading to incorrect predictions and overfitting, especially as erroneous labels propagate and compound throughout the training process. To address this, we propose a robust learning framework called Se…

Cited by 0SourceScholar
2024

Egocentric Whole-Body Motion Capture with FisheyeViT and Diffusion-Based Motion Refinement

CVPR 2024poster

In this work we explore egocentric whole-body motion capture using a single fisheye camera which simultaneously estimates human body and hand motion. This task presents significant challenges due to three factors: the lack of high-quality datasets fisheye camera distortion and human body self-occlus…

Cited by 21SourcePDFScholar
2024

Exploring Intrinsic Language-specific Subspaces in Fine-tuning Multilingual Neural Machine Translation

EMNLP 2024main

Multilingual neural machine translation models support fine-tuning hundreds of languages simultaneously. However, fine-tuning on full parameters solely is inefficient potentially leading to negative interactions among languages. In this work, we demonstrate that the fine-tuning for a language occurs…

2024

Learning Camouflaged Object Detection from Noisy Pseudo Label

ECCV 2024poster

"Existing Camouflaged Object Detection (COD) methods rely heavily on large-scale pixel-annotated training sets, which are both time-consuming and labor-intensive. Although weakly supervised methods offer higher annotation efficiency, their performance is far behind due to the unclear visual demarcat…

2024

Mind the Boundary: Coreset Selection via Reconstructing the Decision Boundary

ICML 2024poster

Existing paradigms of pushing the state of the art require exponentially more training data in many fields. Coreset selection seeks to mitigate this growing demand by identifying the most efficient subset of training data. In this paper, we delve into geometry-based coreset methods and preliminarily…

Cited by 11SourcePDFScholar
2020

Long-term Human Motion Prediction with Scene Context

ECCV 2020poster

Human movement is goal-directed and influenced by the spatial layout of the objects in the scene. To plan future human motion, it is crucial to perceive the environment -- imagine how hard it is to navigate a new room with lights off. Existing works on predicting human motion do not pay attention to…

2019

Learning Independent Object Motion From Unlabelled Stereoscopic Videos

CVPR 2019poster

We present a system for learning motion maps of independently moving objects from stereo videos. The only annotations used in our system are 2D object bounding boxes which introduce the notion of objects in our system. Unlike prior learning based approaches which have focused on predicting dense opt…

Cited by 37PDFScholar
2017

Realtime Multi-Person 2D Pose Estimation Using Part Affinity Fields

CVPR 2017oral

We present an approach to efficiently detect the 2D pose of multiple people in an image. The approach uses a nonparametric representation, which we refer to as Part Affinity Fields (PAFs), to learn to associate body parts with individuals in the image. The architecture encodes global context, allowi…

Cited by 9379PDFcodeScholar