← Search

Jiaqi Yang

50 accepted papers

2026

DepthMesh: A Dual-End Complementary Online Depth Estimation and Mesh Reconstruction

ICRA 2026poster

We present a novel dual-end complementary method for online depth estimation and mesh reconstruction, termed DepthMesh. Unlike most existing state-of-the-art methods that produce either only depth online or surface mesh offline, our method tightly couples online multiview depth estimation and Trunca…

Cited by 0Scholar
2026

DynOPETs: A Versatile Benchmark for Dynamic Object Pose Estimation and Tracking in Moving Camera Scenarios

ICRA 2026poster

In the realm of object pose estimation, scenarios involving both dynamic objects and moving cameras are prevalent. However, the scarcity of corresponding real-world datasets significantly hinders the development and evaluation of robust pose estimation models. This is largely attributed to the inher…

2026

Efficient Feature-Free Initialization for Monocular Visual-Inertial Systems Using A Feed-Forward 3D Model

RSS 2026poster

Fast and reliable initialization is critical for monocular visual–inertial navigation systems (VINS), as it establishes the starting conditions for subsequent state estimation. Despite steady progress, most existing methods heavily rely on visual feature correspondences and require 3-4 seconds of se…

Cited by 0SourceScholar
2026

H2-Surv: Hierarchical Hyperbolic Multimodal Representation Learning for Survival Prediction

CVPR 2026

Cancer survival prediction through multimodal learning that combines histopathology images with genomic data represents a promising research direction. However, current approaches still suffer from two key limitations. First, most methods operate in a Euclidean feature space, which makes it difficul

Cited by 0SourceScholar
2026

Hg-I2P: Bridging Modalities for Generalizable Image-to-Point-Cloud Registration via Heterogeneous Graphs

CVPR 2026

Image-to-point-cloud (I2P) registration aims to align 2D images with 3D point clouds by establishing reliable 2D-3D correspondences. The drastic modality gap between images and point clouds makes it challenging to learn features that are both discriminative and generalizable, leading to severe perfo

Cited by 0SourcecodeScholar
2026

Instilling an Active Mind in Avatars via Cognitive Simulation

ICLR 2026oral

Current video avatar models can generate fluid animations but struggle to capture a character's authentic essence, primarily synchronizing motion with low-level audio cues instead of understanding higher-level semantics like emotion or intent. To bridge this gap, we propose a novel framework for gen…

Cited by 0SourcecodeScholar
2026

InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions

ICLR 2026poster

End-to-end human animation with rich multi-modal conditions, e.g., text, image and audio has achieved remarkable advancements in recent years. However, most existing methods could only animate a single subject and inject conditions in a global manner, ignoring scenarios that multiple concepts could…

Cited by 0SourceScholar
2026

Metric, Inertially Aligned Monocular State Estimation Via Kinetodynamic Priors

ICRA 2026poster

Accurate state estimation for flexible robotic systems poses significant challenges, particularly for platforms with dynamically deforming structures that invalidate rigid-body assumptions. This paper addresses this problem and enables the extension of existing rigid-body pose estimation methods to …

2026

Topology-aware Feature Propagation for Unsupervised Non-rigid Point Cloud Correspondence

CVPR 2026

Unsupervised non-rigid point cloud correspondence aims to predict point-to-point correspondences without annotations. Existing methods leverage the spatial-relation-based feature propagation strategy that includes non-physical connections, which are sensitive to non-rigid deformation. To address thi

Cited by 0SourceScholar
2025

ArgMatch: Adaptive Refinement Gathering for Efficient Dense Matching

ICCV 2025poster

Establishing dense correspondences is crucial yet computationally demanding in multi-view tasks. Although coarse-to-fine schemes mitigate computational costs, their efficiency remains limited by the substantial demands of heavy feature extractors and global matchers. In this paper, we propose Adapti…

2025

BoxDreamer: Dreaming Box Corners for Generalizable Object Pose Estimation

ICCV 2025poster

This paper presents a generalizable RGB-based approach for object pose estimation, specifically designed to address challenges in sparse-view settings. While existing methods can estimate the poses of unseen objects, their generalization ability remains limited in scenarios involving occlusions and…

Cited by 0SourcePDFScholar
2025

CyberHost: A One-stage Diffusion Framework for Audio-driven Talking Body Generation

ICLR 2025oral

Diffusion-based video generation technology has advanced significantly, catalyzing a proliferation of research in human animation. While breakthroughs have been made in driving human animation through various modalities for portraits, most of current solutions for human body animation still focus on…

Cited by 0SourcePDFScholar
2025

DFF: Decision-Focused Fine-Tuning for Smarter Predict-Then-Optimize with Limited Data

AAAI 2025technical

Decision-focused learning (DFL) offers an end-to-end approach to the predict-then-optimize (PO) framework by training predictive models directly on decision loss (DL), enhancing decision-making performance within PO contexts. However, the implementation of DFL poses distinct challenges. Primarily, D…

Cited by 0SourcePDFScholar
2025

DynOPETs: A Versatile Benchmark for Dynamic Object Pose Estimation and Tracking in Moving Camera Scenarios

RA-L 2025

In the realm of object pose estimation, scenarios involving both dynamic objects and moving cameras are prevalent. However, the scarcity of corresponding real-world datasets significantly hinders the development and evaluation of robust pose estimation models. This is largely attributed to the inher

Cited by 0SourceScholar
2025

FADA: Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG Distillation

CVPR 2025poster

Diffusion-based audio-driven talking avatar methods have recently gained attention for their high-fidelity, vivid, and expressive results. However, their slow inference speed limits practical applications. Despite the development of various distillation techniques for diffusion models, we found that…

2025

GS-EVT: Cross-Modal Event Camera Tracking Based on Gaussian Splatting

ICRA 2025

Reliable self-localization is a foundational skill for many intelligent mobile platforms. This paper explores the use of event cameras for motion tracking thereby providing a solution with inherent robustness under difficult dynamics and illumination. In order to circumvent the challenge of event ca

Cited by 3SourceScholar
2025

HyperGCT: A Dynamic Hyper-GNN-Learned Geometric Constraint for 3D Registration

ICCV 2025poster

Geometric constraints between feature matches are critical in 3D point cloud registration problems. Existing approaches typically model unordered matches as a consistency graph and sample consistent matches to generate hypotheses. However, explicit graph construction introduces noise, posing great c…

2025

Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency

ICLR 2025oral

With the introduction of video diffusion model, audio-conditioned human video generation has recently achieved significant breakthroughs in both the naturalness of motion and the synthesis of portrait details. Due to the limited control of audio signals in driving human motion, existing methods ofte…

2025

MinCD-PnP: Learning 2D-3D Correspondences with Approximate Blind PnP

ICCV 2025poster

Image-to-point-cloud (I2P) registration is a fundamental problem in computer vision, focusing on establishing 2D-3D correspondences between an image and a point cloud. Recently, the differentiable perspective-n-point (PnP) has been widely used to supervise I2P registration networks by enforcing proj…

2025

MobilePortrait: Real-Time One-Shot Neural Head Avatars on Mobile Devices

CVPR 2025poster

Existing neural head avatars methods have achieved significant progress in the image quality and motion range of portrait animation. However, these methods prioritize effectiveness over computational overhead. This paper presents MobilePortrait, a lightweight one-shot neural head avatars method that…

Cited by 8SourcePDFScholar
2025

OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

ICCV 2025poster

End-to-end human animation, such as audio-driven talking human generation, has undergone notable advancements in the recent few years. However, existing methods still struggle to scale up as large general video generation models, limiting their potential in real applications. In this paper, we propo…

Cited by 0SourcePDFScholar
2025

SPU-IMR: Self-supervised Arbitrary-scale Point Cloud Upsampling via Iterative Mask-recovery Network

AAAI 2025technical

Point cloud upsampling aims to generate dense and uniformly distributed point sets from sparse point clouds. Existing point cloud upsampling methods typically approach the task as an interpolation problem. They achieve upsampling by performing local interpolation between point clouds or in the featu…

2025

Top-I2P: Explore Open-Domain Image-to-Point Cloud Registration Using Topology Relationship

IJCAI 2025

Image-to-point cloud (I2P) registration is a fundamental task in computer vision, which aims to align pixels in 2D images with corresponding points in 3D point clouds. While deep learning based methods dominate this field, they often fail to generalize to the open domain. In this paper, we address o

Cited by 0SourcePDFScholar
2025

Unlocking Generalization Power in LiDAR Point Cloud Registration

CVPR 2025highlight

In real-world environments, a LiDAR point cloud registration method with robust generalization capabilities (across varying distances and datasets) is crucial for ensuring safety in autonomous driving and other LiDAR-based applications. However, current methods fall short in achieving this level of…

2024

APSeg: Auto-Prompt Network for Cross-Domain Few-Shot Semantic Segmentation

CVPR 2024poster

Few-shot semantic segmentation (FSS) endeavors to segment unseen classes with only a few labeled samples. Current FSS methods are commonly built on the assumption that their training and application scenarios share similar domains and their performances degrade significantly while applied to a disti…

Cited by 16SourcePDFScholar
2024

H-LegalKI: A Hierarchical Legal Knowledge Integration Framework for Legal Community Question Answering

EMNLP 2024finding

Legal question answering (LQA) aims to bridge the gap between the limited availability of legal professionals and the high demand for legal assistance. Traditional LQA approaches typically either select the optimal answers from an answer set or extract answers from law texts. However, they often str…

2024

MV-ROPE: Multi-view Constraints for Robust Category-level Object Pose and Size Estimation

IROS 2024poster

Recently there has been a growing interest in category-level object pose and size estimation, and prevailing methods commonly rely on single view RGB-D images. However, one disadvantage of such methods is that they require accurate depth maps which cannot be produced by consumer-grade sensors. Furth…

Cited by 2SourceScholar
2024

MonoSample: Synthetic 3D Data Augmentation Method in Monocular 3D Object Detection

RA-L 2024

In the context of autonomous driving, it is both critical and challenging to locate 3D objects by using a calibrated RGB image. Current methods typically utilize heteroscedastic aleatoric uncertainty loss to regress the depth of objects, thereby reducing the impact of noisy input while also ensuring

Cited by 4SourceScholar
2024

RE-SORT: Removing Spurious Correlation in Multilevel Interaction for CTR Prediction

UAI 2024poster

Click-through rate (CTR) prediction is a critical task in recommendation systems, serving as the ultimate filtering step to sort items for a user. Most recent cutting-edge methods primarily focus on investigating complex implicit and explicit feature interactions; however, these methods neglect the…

2024

Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis

ICLR 2024spotlight

One-shot 3D talking portrait generation aims to reconstruct a 3D avatar from an unseen image, and then animate it with a reference video or audio to generate a talking portrait video. The existing methods fail to simultaneously achieve the goals of accurate 3D avatar reconstruction and stable talkin…

2023

Learning Zero-Shot Cooperation with Humans, Assuming Humans Are Biased

ICLR 2023poster

There is a recent trend of applying multi-agent reinforcement learning (MARL) to train an agent that can cooperate with humans in a zero-shot fashion without using any human data. The typical workflow is to first repeatedly run self-play (SP) to build a policy pool and then train the final adaptive…

2023

MixCycle: Mixup Assisted Semi-Supervised 3D Single Object Tracking with Cycle Consistency

ICCV 2023poster

3D single object tracking (SOT) is an indispensable part of automated driving. Existing approaches rely heavily on large, densely labeled datasets. However, annotating point clouds is both costly and time-consuming. Inspired by the great success of cycle tracking in unsupervised 2D SOT, we introduce…

Cited by 6PDFcodeScholar
2023

Multi-View Inverse Rendering for Large-Scale Real-World Indoor Scenes

CVPR 2023poster

We present a efficient multi-view inverse rendering method for large-scale real-world indoor scenes that reconstructs global illumination and physically-reasonable SVBRDFs. Unlike previous representations, where the global illumination of large scenes is simplified as multiple environment maps, we p…

2023

Revisiting Event-Based Video Frame Interpolation

IROS 2023poster

Dynamic vision sensors or event cameras provide rich complementary information for video frame interpolation. Existing state-of-the-art methods follow the paradigm of combining both synthesis-based and warping networks. However, few of those methods fully respect the intrinsic characteristics of eve…

Cited by 4SourceScholar
2023

iQuery: Instruments As Queries for Audio-Visual Sound Separation

CVPR 2023poster

Current audio-visual separation methods share a standard architecture design where an audio encoder-decoder network is fused with visual encoding features at the encoder bottleneck. This design confounds the learning of multi-modal feature encoding with robust sound decoding for audio separation. To…

2022

DEVO: Depth-Event Camera Visual Odometry in Challenging Conditions

ICRA 2022poster

We present a novel real-time visual odometry framework for a stereo setup of a depth and high-resolution event camera. Our framework balances accuracy and robustness against computational efficiency towards strong performance in challenging scenarios. We extend conventional edge-based semi-dense vis…

Cited by 65SourceScholar
2022

Phasic Self-Imitative Reduction for Sparse-Reward Goal-Conditioned Reinforcement Learning

ICML 2022spotlight

It has been a recent trend to leverage the power of supervised learning (SL) towards more effective reinforcement learning (RL) methods. We propose a novel phasic solution by alternating online RL and offline SL for tackling sparse-reward goal-conditioned problems. In the online phase, we perform RL…

Cited by 22SourcePDFScholar
2022

PhyIR: Physics-Based Inverse Rendering for Panoramic Indoor Images

CVPR 2022poster

Inverse rendering of complex material such as glossy, metal and mirror material is a long-standing ill-posed problem in this area, which has not been well solved. Previous approaches cannot tackle them well due to simplified BRDF and unsuitable illumination representations. In this paper, we present…

Cited by 26PDFcodeScholar
2022

Revisiting Some Common Practices in Cooperative Multi-Agent Reinforcement Learning

ICML 2022spotlight

Many advances in cooperative multi-agent reinforcement learning (MARL) are based on two common design principles: value decomposition and parameter sharing. A typical MARL algorithm of this fashion decomposes a centralized Q-function into local Q-networks with parameters shared across agents. Such a…

Cited by 49SourcePDFScholar
2022

Unsupervised Learning of 3D Semantic Keypoints with Mutual Reconstruction

ECCV 2022poster

"Semantic 3D keypoints are category-level semantic consistent points on 3D objects. Detecting 3D semantic keypoints is a foundation for a number of 3D vision tasks but remains challenging, due to the ambiguity of semantic information, especially when the objects are represented by unordered 3D point…

2022

VECtor: A Versatile Event-Centric Benchmark for Multi-Sensor SLAM

RA-L 2022

Event cameras have recently gained in popularity as they hold strong potential to complement regular cameras in situations of high dynamics or challenging illumination. An important problem that may benefit from the addition of an event camera is given by Simultaneous Localization And Mapping (SLAM)

Cited by 103SourceScholar
2021

Going Beyond Linear RL: Sample Efficient Neural Function Approximation

NeurIPS 2021poster

Deep Reinforcement Learning (RL) powered by neural net approximation of the Q function has had enormous empirical success. While the theory of RL has traditionally focused on linear function approximation (or eluder dimension) approaches, little is known about nonlinear RL with neural net approximat…

Cited by 10SourcePDFScholar
2021

Improved Variance-Aware Confidence Sets for Linear Bandits and Linear Mixture MDP

NeurIPS 2021poster

This paper presents new \emph{variance-aware} confidence sets for linear bandits and linear mixture Markov Decision Processes (MDPs). With the new confidence sets, we obtain the follow regret bounds: For linear bandits, we obtain an $\widetilde{O}(\mathrm{poly}(d)\sqrt{1 + \sum_{k=1}^{K}\sigma_k^2}…

Cited by 45SourcePDFScholar
2021

Optimal Gradient-based Algorithms for Non-concave Bandit Optimization

NeurIPS 2021poster

Bandit problems with linear or concave reward have been extensively studied, but relatively few works have studied bandits with non-concave reward. This work considers a large family of bandit problems where the unknown underlying reward function is non-concave, including the low-rank generalized li…

Cited by 18SourcePDFScholar
2021

Provable Model-based Nonlinear Bandit and Reinforcement Learning: Shelve Optimism, Embrace Virtual Curvature

NeurIPS 2021poster

This paper studies model-based bandit and reinforcement learning (RL) with nonlinear function approximations. We propose to study convergence to approximate local maxima because we show that global convergence is statistically intractable even for one-layer neural net bandit with a deterministic rew…

Cited by 47SourcePDFScholar