← Search

Qi Ye

40 accepted papers

2026

DiffPBR: Point-Based Rendering via Spatial-Aware Residual Diffusion

ICLR 2026poster

Neural radiance fields and 3D Gaussian splatting (3DGS) have significantly advanced 3D reconstruction and novel view synthesis (NVS). Yet, achieving high-fidelity and view-consistent renderings directly from point clouds---without costly per-scene optimization---remains a core challenge. In this wor…

Cited by 0SourceScholar
2026

MIMIC: Mask-Injected Manipulation Video Generation with Interaction Control

ICLR 2026poster

Embodied intelligence faces a fundamental bottleneck from limited large-scale interaction data. Video generation offers a scalable alternative, but manipulation videos remain particularly challenging, as they require capturing subtle, contact-rich dynamics. Despite recent advances, video diffusion m…

Cited by 0SourceScholar
2026

The Developments and Challenges towards Dexterous and Embodied Robotic Manipulation: A Survey

ICRA 2026poster

Achieving human-like dexterous robotic manipulation remains a central goal and a pivotal challenge in robotics. The development of Artificial Intelligence (AI) has allowed rapid progress in robotic manipulation. This survey summarizes the evolution of robotic manipulation from mechanical programming…

2025

A3GS: Arbitrary Artistic Style into Arbitrary 3D Gaussian Splatting

ICCV 2025poster

Recently, the field of 3D scene stylization has attracted considerable attention, particularly for applications in the metaverse. A key challenge is rapidly transferring the style of an arbitrary reference image to a 3D scene while faithfully preserving its content structure and spatial layout. Work…

Cited by 0SourcePDFScholar
2025

CMQCIC-Bench: A Chinese Benchmark for Evaluating Large Language Models in Medical Quality Control Indicator Calculation

ACL 2025finding

Medical quality control indicators are essential to assess the qualifications of healthcare institutions for medical services. With the impressive performance of large language models (LLMs) like GPT-4 in the medical field, leveraging these technologies for the Medical Quality Control Indicator Calc…

2025

Hand-held Object Reconstruction from RGB Video with Dynamic Interaction

CVPR 2025poster

This work aims to reconstruct the 3D geometry of a rigid object manipulated by one or both hands using monocular RGB video. Previous methods rely on Structure-from-Motion or hand priors to estimate relative motion between the object and camera, which typically assume textured objects or single-hand…

2025

IMQC: A Large Language Model Platform for Medical Quality Control

AAAI 2025technical

Medical quality control (MQC) indicators are essential for evaluating the performance of healthcare institutions to ensure high-quality patient care. In this paper, we report the design, implementation, and deployment of the Intelligent EMR-LLM platform for Medical Quality Control (IMQC), a large la…

2025

IntrinsicControlNet: Cross-distribution Image Generation with Real and Unreal

ICCV 2025poster

Realistic images are usually produced by simulating light transportation results of 3D scenes using rendering engines. This framework can precisely control the output but is usually weak at producing photo-like images. Alternatively, diffusion models have seen great success in photorealistic image g…

Cited by 0SourcePDFScholar
2025

MinosEval: Distinguishing Factoid and Non-Factoid for Tailored Open-Ended QA Evaluation with LLMs

ACL 2025finding

Open-ended question answering (QA) is a key task for evaluating the capabilities of large language models (LLMs). Compared to closed-ended QA, it demands longer answer statements, more nuanced reasoning processes, and diverse expressions, making refined and interpretable automatic evaluation both cr…

Cited by 0SourcePDFScholar
2025

Multi-Robot Autonomous 3D Reconstruction Using Gaussian Splatting With Semantic Guidance

RA-L 2025

Implicit neural representations and 3D Gaussian splatting (3DGS) have shown great potential for scene reconstruction. Recent studies have expanded their applications in autonomous reconstruction through task assignment methods. However, these methods are mainly limited to a single robot, and rapid r

Cited by 4SourceScholar
2025

OpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language Model

NeurIPS 2025oral

Understanding and synthesizing realistic 3D hand-object interactions (HOI) is critical for applications ranging from immersive AR/VR to dexterous robotics. Existing methods struggle with generalization, performing well on closed-set objects and predefined tasks but failing to handle unseen objects o…

Cited by 0SourceScholar
2025

RFMPose: Generative Category-level Object Pose Estimation via Riemannian Flow Matching

NeurIPS 2025poster

We introduce RFMPose, a novel generative framework for category-level 6D object pose estimation that learns deterministic pose trajectories through Riemannian Flow Matching (RFM). Existing discriminative approaches struggle with multi-hypothesis predictions (e.g., symmetry ambiguities) and often req…

Cited by 0SourceScholar
2025

Temporal-Spatial Representation Fusion for Dexterous Manipulation Learning with Unpaired Visual-Action Data

IROS 2025

Supervised behavioral cloning using robot visual-action data has been widely investigated in robot manipulation. However, these methods typically require simultaneous acquisition of visual and action data, which makes them difficult to utilize unpaired visual-action datasets: e.g. videos on Internet

Cited by 0SourceScholar
2025

UpViTaL: Unpaired Visual-Tactile Self-Supervised Representation Learning for Dexterous Robotic Manipulation

ICRA 2025

Visual and tactile pretraining have been extensively studied in dexterous robot manipulation tasks. However, existing methods typically require the simultaneous acquisition of visual and tactile data, making it difficult to utilize low-cost, unpaired visual-tactile datasets. Moreover, these methods

Cited by 1SourceScholar
2025

VTAO-BiManip: Masked Visual-Tactile-Action Pre-training with Object Understanding for Bimanual Dexterous Manipulation

IROS 2025

Bimanual dexterous manipulation remains a significant challenge in robotics due to the high DoFs of each hand and their coordination. Existing single-hand manipulation techniques often leverage human demonstrations to guide RL methods but fail to generalize to complex bimanual tasks involving multip

Cited by 4SourceScholar
2025

VTDexManip: A Dataset and Benchmark for Visual-tactile Pretraining and Dexterous Manipulation with Reinforcement Learning

ICLR 2025poster

Vision and touch are the most commonly used senses in human manipulation. While leveraging human manipulation videos for robotic task pretraining has shown promise in prior works, it is limited to image and language modalities and deployment to simple parallel grippers. In this paper, aiming to addr…

2024

Autonomous Implicit Indoor Scene Reconstruction with Frontier Exploration

ICRA 2024poster

Implicit neural representations have demonstrated significant promise for 3D scene reconstruction. Recent works have extended their applications to autonomous implicit reconstruction through the Next Best View (NBV) based method. However, the NBV method cannot guarantee complete scene coverage and o…

Cited by 3SourceScholar
2024

CAMInterHand: Cooperative Attention for Multi-View Interactive Hand Pose and Mesh Reconstruction

ICRA 2024poster

Interactive hand mesh reconstruction from singleview images poses a significant challenge with the severe occlusion and depth ambiguity inherent in interactive hand gestures. Recent approaches that employ probabilistic models and tokenpruned techniques have shown decent results in multi-view human b…

Cited by 0SourceScholar
2024

Error-aware Sampling in Adaptive Shells for Neural Surface Reconstruction

IJCAI 2024poster

Neural implicit surfaces with signed distance functions (SDFs) achieve superior quality in 3D geometry reconstruction. However, training SDFs is time-consuming because it requires a great number of samples to calculate accurate weight distributions and a considerable amount of samples sampled from t…

2024

In-Hand 3D Object Reconstruction from a Monocular RGB Video

AAAI 2024technical

Our work aims to reconstruct a 3D object that is held and rotated by a hand in front of a static RGB camera. Previous methods that use implicit neural representations to recover the geometry of a generic hand-held object from multi-view images achieved compelling results in the visible part of the o…

2024

InterRep: A Visual Interaction Representation for Robotic Grasping

ICRA 2024poster

Recently, pre-trained vision models have gained significant attention in motor control, showcasing impressive performance across diverse robotic learning tasks. While previous works predominantly concentrate on the significance of the pre-training phase, the equally important task of extracting more…

Cited by 1SourceScholar
2024

Masked Visual-Tactile Pre-training for Robot Manipulation

ICRA 2024poster

Recent works on the pretraining for robot manipulation have demonstrated that representations learning from large human manipulation data can generalize well to new manipulation tasks and environments. However, these approaches mainly focus on human vision or natural language, neglecting tactile fee…

Cited by 6SourceScholar
2024

TPGP: Temporal-Parametric Optimization with Deep Grasp Prior for Dexterous Motion Planning

ICRA 2024poster

Grasping motion planning aims to find a feasible grasping trajectory in the configuration space given an input target grasp. While optimizing grasp motion with two or three-fingered grippers has been well studied, the study on natural grasp motion planning with a dexterous hand remains a very challe…

Cited by 2SourceScholar
2023

Contact2Grasp: 3D Grasp Synthesis via Hand-Object Contact Constraint

IJCAI 2023poster

3D grasp synthesis generates grasping poses given an input object. Existing works tackle the problem by learning a direct mapping from objects to the distributions of grasping poses. However, because the physical contact is sensitive to small changes in pose, the high-nonlinear mapping between 3D ob…

Cited by 15SourcePDFScholar
2023

DexRepNet: Learning Dexterous Robotic Grasping Network with Geometric and Spatial Hand-Object Representations

IROS 2023poster

Robotic dexterous grasping is a challenging problem due to the high degree of freedom (DoF) and complex contacts of multi-fingered robotic hands. Existing deep re-inforcement learning (DRL) based methods leverage human demonstrations to reduce sample complexity due to the high dimensional action spa…

Cited by 18SourceScholar
2023

Diff-LfD: Contact-aware Model-based Learning from Visual Demonstration for Robotic Manipulation via Differentiable Physics-based Simulation and Rendering

CoRL 2023oral

Learning from Demonstration (LfD) is an efficient technique for robots to acquire new skills through expert observation, significantly mitigating the need for laborious manual reward function design. This paper introduces a novel framework for model-based LfD in the context of robotic manipulation.…

Cited by 19SourceScholar
2023

Efficient View Path Planning for Autonomous Implicit Reconstruction

ICRA 2023poster

Implicit neural representations have shown promising potential for 3D scene reconstruction. Recent work applies it to autonomous 3D reconstruction by learning information gain for view path planning. Effective as it is, the computation of the information gain is expensive, and compared with that usi…

Cited by 20SourceScholar
2023

F&F Attack: Adversarial Attack against Multiple Object Trackers by Inducing False Negatives and False Positives

ICCV 2023poster

Multi-object tracking (MOT) aims to build moving trajectories for number-agnostic objects. Modern multi-object trackers commonly follow the tracking-by-detection strategy. Therefore, fooling detectors can be an effective solution but it usually requires attacks in multiple successive frames, resulti…

Cited by 10PDFScholar
2023

I2-SDF: Intrinsic Indoor Scene Reconstruction and Editing via Raytracing in Neural SDFs

CVPR 2023poster

In this work, we present I^2-SDF, a new method for intrinsic indoor scene reconstruction and editing using differentiable Monte Carlo raytracing on neural signed distance fields (SDFs). Our holistic neural SDF-based framework jointly recovers the underlying shapes, incident radiance and materials fr…

2023

ImmFusion: Robust mmWave-RGB Fusion for 3D Human Body Reconstruction in All Weather Conditions

ICRA 2023poster

3D human reconstruction from RGB images achieves decent results in good weather conditions but degrades dramatically in rough weather. Complementary, mmWave radars have been employed to reconstruct 3D human joints and meshes in rough weather. However, combining RGB and mmWave signals for robust all-…

Cited by 24SourceScholar
2023

Implementation and Optimization of Grasping Learning with Dual-modal Soft Gripper

ICRA 2023poster

Robust and efficient grasping of different objects is still an open problem due to the difficulty of integrating multidisciplinary knowledge such as gripper ontology design, perception, control, and learning. In recent years, learning-based methods have achieved excellent results in grasping various…

Cited by 4SourceScholar
2023

InterTracker: Discovering and Tracking General Objects Interacting with Hands in the Wild

IROS 2023poster

Understanding human interaction with objects is an important research topic for embodied Artificial Intelligence and identifying the objects that humans are interacting with is a primary problem for interaction understanding. Existing methods rely on frame-based detectors to locate interacting objec…

Cited by 1SourceScholar
2023

NeurAR: Neural Uncertainty for Autonomous 3D Reconstruction With Implicit Neural Representations

RA-L 2023

Implicit neural representations have shown compelling results in offline 3D reconstruction and also recently demonstrated the potential for online SLAM systems. However, applying them to autonomous 3D reconstruction, where a robot is required to explore a scene and plan a view path for the reconstru

Cited by 91SourceScholar
2023

Seal-3D: Interactive Pixel-Level Editing for Neural Radiance Fields

ICCV 2023poster

With the popularity of implicit neural representations, or neural radiance fields (NeRF), there is a pressing need for editing methods to interact with the implicit 3D models for tasks like post-processing reconstructed scenes and 3D content creation. While previous works have explored NeRF editing…

Cited by 18PDFcodeScholar
2022

Few-Shot Learning with Improved Local Representations via Bias Rectify Module

ICASSP 2022accepted

Recent approaches based on metric learning have achieved great progress in few-shot learning. However, most of them are limited to image-level representation manners, which fail to properly deal with the intra-class variations and spatial knowledge and thus produce undesirable performance. In this p…

Cited by 0SourceScholar
2021

Geometry-Based Distance Decomposition for Monocular 3D Object Detection

ICCV 2021poster

Monocular 3D object detection is of great significance for autonomous driving but remains challenging. The core challenge is to predict the distance of objects in the absence of explicit depth information. Unlike regressing the distance as a single variable in most existing methods, we propose a nov…

Cited by 168PDFcodeScholar
2020

The Phong Surface: Efficient 3D Model Fitting using Lifted Optimization

ECCV 2020poster

Realtime perceptual and interaction capabilities in mixed reality require a range of 3D tracking problems to be solved at low latency on resource-constrained hardware such as head-mounted devices. Indeed, for devices such as HoloLens 2 where the CPU and GPU are left available for applications, multi…

Cited by 17SourcePDFScholar
2017

BigHand2.2M Benchmark: Hand Pose Dataset and State of the Art Analysis

CVPR 2017poster

In this paper we introduce a large-scale hand pose dataset, collected using a novel capture method. Existing datasets are either generated synthetically or captured using depth sensors: synthetic datasets exhibit a certain level of appearance difference from real depth images, and real datasets are…

Cited by 329PDFScholar