← Search

Jun He

23 accepted papers

2026

ActAvatar: Temporally-Aware Precise Action Control for Talking Avatars

CVPR 2026

Despite significant advances in talking avatar generation, existing methods face critical challenges: insufficient text-following capability for diverse actions, lack of temporal alignment between actions and audio content, and dependency on additional control signals such as pose skeletons. We pres

Cited by 0SourceScholar
2026

MajutsuCity: Language-driven Aesthetic-adaptive City Generation with Controllable 3D Assets and Layouts

CVPR 2026

Generating realistic 3D cities is fundamental to world models, virtual reality, and game development, where an ideal urban scene must satisfy both stylistic diversity, fine-grained, and controllability. However, existing methods struggle to balance the creative flexibility offered by text-based gene

Cited by 0SourcecodeScholar
2026

UrbanFeel:A Comprehensive Benchmark for Temporal and Perceptual Understanding of City Scenes through Human Perspective

ICLR 2026poster

Urban development impacts over half of the global population, making human-centered understanding of its structural and perceptual changes essential for smart city planning. While Multimodal Large Language Models (MLLMs) have shown remarkable capabilities across various domains, existing benchmarks…

Cited by 0SourcecodeScholar
2025

BLINK-Twice: You see, but do you observe? A Reasoning Benchmark on Visual Perception

NeurIPS 2025poster

Recently, Multimodal Large Language Models (MLLMs) have made rapid progress, particularly in enhancing their reasoning capabilities. However, existing reasoning benchmarks still primarily assess language-based reasoning, often treating visual input as replaceable context. To address this gap, we int…

Cited by 0SourcecodeScholar
2025

DualTalk: Dual-Speaker Interaction for 3D Talking Head Conversations

CVPR 2025poster

In face-to-face conversations, individuals need to switch between speaking and listening roles seamlessly. Existing 3D talking head generation models focus solely on speaking or listening, neglecting the natural dynamics of interactive conversation, which leads to unnatural interactions and awkward…

2025

LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models

ICLR 2025spotlight

With the rapid development of AI-generated content, the future internet may be inundated with synthetic data, making the discrimination of authentic and credible multimodal data increasingly challenging. Synthetic data detection has thus garnered widespread attention, and the performance of large mu…

2025

Leveraging BEV Paradigm for Ground-to-Aerial Image Synthesis

ICCV 2025poster

Ground-to-aerial image synthesis focuses on generating realistic aerial images from corresponding ground street view images while maintaining consistent content layout, simulating a top-down view. The significant viewpoint difference leads to domain gaps between views, and dense urban scenes limit t…

2025

MEGADance: Mixture-of-Experts Architecture for Genre-Aware 3D Dance Generation

NeurIPS 2025poster

Music-driven 3D dance generation has attracted increasing attention in recent years, with promising applications in choreography, virtual reality, and creative content creation. Previous research has generated promising realistic dance movement from audio signals. However, traditional methods underu…

Cited by 0SourceScholar
2025

MalImgDA: Diffusion-based Data Augmentation for Long-tailed Malware Family Classification

ICASSP 2025accepted

With the rapid improvement of machine learning technology, leveraging machine learning methods for malware classification has emerged as a viable approach. However, under real-world circumstance, the imbalanced or long-tailed distribution among various malware families, poses a critical challenge to…

Cited by 0SourceScholar
2025

OmniSync: Towards Universal Lip Synchronization via Diffusion Transformers

NeurIPS 2025spotlight

Lip synchronization is the task of aligning a speaker’s lip movements in video with corresponding speech audio, and it is essential for creating realistic, expressive video content. However, existing methods often rely on reference frames and masked-frame inpainting, which limit their robustness to…

Cited by 0SourceScholar
2025

Scene4U: Hierarchical Layered 3D Scene Reconstruction from Single Panoramic Image for Your Immerse Exploration

CVPR 2025poster

The reconstruction of immersive and realistic 3D scenes holds significant practical importance in various fields of computer vision and computer graphics. Typically, immersive and realistic scenes should be free from obstructions by dynamic objects, maintain global texture consistency, and allow for…

Cited by 0SourcePDFScholar
2025

SwiftPrune: Hessian-Free Weight Pruning for Large Language Models

EMNLP 2025

Post-training pruning, as one of the key techniques for compressing large language models (LLMs), plays a vital role in lightweight model deployment and model sparsity. However, current mainstream pruning methods dependent on the Hessian matrix face significant limitations in both pruning speed and

Cited by 0SourcePDFScholar
2024

SyncTalk: The Devil is in the Synchronization for Talking Head Synthesis

CVPR 2024poster

Achieving high synchronization in the synthesis of realistic speech-driven talking head videos presents a significant challenge. Traditional Generative Adversarial Networks (GAN) struggle to maintain consistent facial identity while Neural Radiance Fields (NeRF) methods although they can address thi…

2023

EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face Animation

ICCV 2023poster

Speech-driven 3D face animation aims to generate realistic facial expressions that match the speech content and emotion. However, existing methods often neglect emotional facial expressions or fail to disentangle them from speech content. To address this issue, this paper proposes an end-to-end neur…

Cited by 117PDFcodeScholar
2023

GIDP: Learning a Good Initialization and Inducing Descriptor Post-enhancing for Large-scale Place Recognition

ICRA 2023poster

Large-scale place recognition is a fundamental but challenging task, which plays an increasingly important role in autonomous driving and robotics. Existing methods have achieved acceptable good performance, however, most of them are concentrating on designing elaborate global descriptor learning ne…

Cited by 0SourceScholar
2023

MM-PCQA: Multi-Modal Learning for No-reference Point Cloud Quality Assessment

IJCAI 2023poster

The visual quality of point clouds has been greatly emphasized since the ever-increasing 3D vision applications are expected to provide cost-effective and high-quality experiences for users. Looking back on the development of point cloud quality assessment (PCQA), the visual quality is usually eval…

2023

Reconstruction-Aware Prior Distillation for Semi-supervised Point Cloud Completion

IJCAI 2023poster

Real-world sensors often produce incomplete, irregular, and noisy point clouds, making point cloud completion increasingly important. However, most existing completion methods rely on large paired datasets for training, which is labor-intensive. This paper proposes RaPD, a novel semi-supervised poin…

Cited by 14SourcePDFScholar
2022

An MRC Framework for Semantic Role Labeling

COLING 2022main

Semantic Role Labeling (SRL) aims at recognizing the predicate-argument structure of a sentence and can be decomposed into two subtasks: predicate disambiguation and argument labeling. Prior work deals with these two tasks independently, which ignores the semantic connection between the two tasks. I…

2022

Object Level Depth Reconstruction for Category Level 6D Object Pose Estimation from Monocular RGB Image

ECCV 2022poster

"Recently, RGBD-based category-level 6D object pose estimation has achieved promising improvement in performance, however, the requirement of depth information prohibits broader applications. In order to relieve this problem, this paper proposes a novel approach named Object Level Depth reconstructi…

Cited by 34SourcePDFScholar
2022

SVT-Net: Super Light-Weight Sparse Voxel Transformer for Large Scale Place Recognition

AAAI 2022technical

Simultaneous Localization and Mapping (SLAM) and Autonomous Driving are becoming increasingly more important in recent years. Point cloud-based large scale place recognition is the spine of them. While many models have been proposed and have achieved acceptable performance by learning short-range lo…

Cited by 78SourcePDFScholar
2021

MPDNet: A 3D Missing Part Detection Network Based on Point Cloud Segmentation

ICASSP 2021accepted

Utilizing computer vision technologies for machinery missing part detection has been a hot research topic recently. Most of existing methods take images as input and utilize 2D object detection pipelines for detecting fault regions. However, 2D models can’t handle the situation when occlusion exists…

Cited by 0SourceScholar
2021

Simultaneous Control of Terrain Adaptation and Wheel Speed Allocation for a Planetary Rover With an Active Suspension System

RA-L 2021

Active suspensions are important features of many recent planetary rovers. For such a rover, the control strategy is crucial to its performance. This letter presents the control method for a planetary rover equipped with an active suspension system. The control algorithm is based on the estimation o

Cited by 9SourceScholar