← Search

Jihua Zhu

19 accepted papers

2026

Memory Matters: Boosting Training-Free Zero-Shot Temporal Action Localization with a Learnable Lookup Table

CVPR 2026

Zero-Shot Temporal Action Localization (ZS-TAL) aims to classify and localize actions in untrimmed videos that are unseen during training. Existing training-based ZS-TAL methods typically rely on fine-tuning models on large-scale annotated training data. This can be impractical in real-world applica

Cited by 0SourceScholar
2026

Point-SRA: Self-Representation Alignment for 3D Representation Learning

AAAI 2026technical

Masked autoencoders (MAE) have become a dominant paradigm in 3D representation learning, setting new performance benchmarks across various downstream tasks. Existing methods with fixed mask ratios neglect multi-level representational correlations and intrinsic geometric structures, while relying on

Cited by 0SourcePDFScholar
2025

Boundary-Aware Temporal Dynamic Pseudo-Supervision Pairs Generation for Zero-Shot Natural Language Video Localization

AAAI 2025technical

Zero-shot Natural Language Video Localization (NLVL) aims to automatically generate moments and corresponding pseudo queries from raw videos for the training of the localization model without any manual annotations. Existing approaches typically produce pseudo queries as simple words, which overlook…

Cited by 0SourcePDFScholar
2025

Multiple Rotation Averaging with Constrained Reweighting Deep Matrix Factorization

ICRA 2025

Multiple rotation averaging plays a crucial role in computer vision and robotics domains. The conventional optimization-based methods optimize a nonlinear cost function based on certain noise assumptions, while most previous learning-based methods require ground truth labels in the supervised traini

Cited by 0SourceScholar
2025

SYKI-SVC: Advancing Singing Voice Conversion with Post-Processing Innovations and an Open-Source Professional Testset

ICASSP 2025accepted

Singing voice conversion aims to transform a source singing voice into that of a target singer while preserving the original lyrics, melody, and various vocal techniques. In this paper, we propose a high-fidelity singing voice conversion system. Our system builds upon the SVCC T02 framework and cons…

Cited by 0SourceScholar
2024

Cross-Modal Information-Guided Network Using Contrastive Learning for Point Cloud Registration

RA-L 2024

The majority of point cloud registration methods currently rely on extracting features from points. However, these methods are limited by their dependence on information obtained from a single modality of points, which can result in deficiencies such as inadequate perception of global features and a

Cited by 14SourcecodeScholar
2024

Hyperbolic Image-and-Pointcloud Contrastive Learning for 3D Classification

IROS 2024poster

3D contrastive representation learning has exhibited remarkable efficacy across various downstream tasks. However, existing contrastive learning paradigms based on cosine similarity fail to deeply explore the potential intra-modal hierarchical and cross-modal semantic correlations about multi-modal…

Cited by 0SourceScholar
2024

Matching Distance and Geometric Distribution Aided Learning Multiview Point Cloud Registration

RA-L 2024

Multiview point cloud registration plays a crucial role in robotics, automation, and computer vision fields. This letter concentrates on pose graph construction and motion synchronization within multiview registration. Previous methods for pose graph construction often pruned fully connected graphs

Cited by 8SourcecodeScholar
2024

Watch Your Head: Assembling Projection Heads to Save the Reliability of Federated Models

AAAI 2024technical

Federated learning encounters substantial challenges with heterogeneous data, leading to performance degradation and convergence issues. While considerable progress has been achieved in mitigating such an impact, the reliability aspect of federated models has been largely disregarded. In this study,…

2023

DualGenerator: Information Interaction-Based Generative Network for Point Cloud Completion

RA-L 2023

Point cloud completion estimates complete shapes from incomplete point clouds to obtain higher-quality point cloud data. Most existing methods only consider global object features, ignoring spatial and semantic information of adjacent points. They cannot distinguish structural information well betwe

Cited by 8SourceScholar
2022

C-CAM: Causal CAM for Weakly Supervised Semantic Segmentation on Medical Image

CVPR 2022poster

Recently, many excellent weakly supervised semantic segmentation (WSSS) works are proposed based on class activation mapping (CAM). However, there are few works that consider the characteristics of medical images. In this paper, we find that there are mainly two challenges of medical images in WSSS:…

Cited by 118PDFcodeScholar
2022

Coarse-to-Fine Generative Modeling for Graphic Layouts

AAAI 2022technical

Even though graphic layout generation has attracted growing attention recently, it is still challenging to synthesis realistic and diverse layouts, due to the complicated element relationships and varied element arrangements. In this work, we seek to improve the performance of layout generation by i…

Cited by 43SourcePDFScholar
2021

Robust Motion Averaging under Maximum Correntropy Criterion

ICRA 2021poster

Recently, the motion averaging method has been introduced as an effective means to solve the multi-view registration problem. This method aims to recover global motions from a set of relative motions, where the original method is sensitive to outliers due to using the Frobenius norm error in the opt…

Cited by 9SourceScholar
2021

Show and Speak: Directly Synthesize Spoken Description of Images

ICASSP 2021accepted

This paper proposes a new model, referred to as the show and speak (SAS) model that, for the first time, is able to directly synthesize spoken descriptions of images, bypassing the need for any text or phonemes. The basic structure of SAS is an encoder-decoder architecture that takes an image as inp…

Cited by 0SourceScholar