← Search

Tao Lu

33 accepted papers

2026

CoRe: Combined Rewards with Vision-Language Model Feedback for Preference-Aligned Reinforcement Learning

ICML 2026poster

Reward design remains a central challenge in reinforcement learning (RL). Hand-crafted rewards are often difficult to specify and may lead to suboptimal policies, while learned rewards from preferences can suffer from inefficiency and unstable training. Inspired by the dual nature of human learning …

Cited by 0SourceScholar
2026

Otil: Accelerating Diffusion Model Inference via Communication-Efficient Multi-GPU Parallelism

CVPR 2026

Diffusion models (DMs) have recently achieved remarkable success across diverse modalities, including high-fidelity image and video synthesis.However, their inherent step sequential denoising process introduces substantial cumulative latency, which significantly degrades user experience. While exist

Cited by 0SourceScholar
2026

PackUV: Packed Gaussian UV Maps for 4D Volumetric Video

CVPR 2026

Volumetric videos offer immersive 4D experiences, but remain difficult to reconstruct, store, and stream at scale. Existing Gaussian Splatting based methods achieve high-quality reconstruction but break down on long sequences, temporal inconsistency, and fail under large motions and disocclusions. M

Cited by 0SourceScholar
2026

Turbo-GS: Accelerating 3D Gaussian Fitting for High-Resolution Radiance Fields

CVPR 2026

Novel-view synthesis plays a crucial role in computer vision with applications in 3D reconstruction, mixed reality, and robotics. Recent approaches, such as 3D Gaussian Splatting (3DGS), have emerged as state-of-the-art solutions, offering high-quality novel view synthesis in real time. However, tra

Cited by 0SourcecodeScholar
2025

End-to-End Entity-Predicate Association Reasoning for Dynamic Scene Graph Generation

ICCV 2025poster

Dynamic Scene Graph Generation (DSGG) aims to comprehensively understand videos by abstracting them into visual triplets <subject, predicate, object>. Most existing methods focus on capturing temporal dependencies, but overlook crucial visual relationship dependencies between entities and predicates…

2025

GaRe: Relightable 3D Gaussian Splatting for Outdoor Scenes from Unconstrained Photo Collections

ICCV 2025poster

We propose a 3D Gaussian splatting-based framework for outdoor relighting that leverages intrinsic image decomposition to precisely integrate sunlight, sky radiance, and indirect lighting from unconstrained photo collections. Unlike prior methods that compress the per-image global illumination into…

Cited by 0SourcePDFScholar
2025

Horizon-GS: Unified 3D Gaussian Splatting for Large-Scale Aerial-to-Ground Scenes

CVPR 2025poster

Seamless integration of both aerial and street view images remains a significant challenge in neural scene reconstruction and rendering. Existing methods predominantly focus on single domain, limiting their applications in immersive environments, which demand extensive free view exploration with lar…

Cited by 1SourcePDFScholar
2025

MISCGrasp: Leveraging Multiple Integrated Scales and Contrastive Learning for Enhanced Volumetric Grasping

IROS 2025

Robotic grasping faces challenges in adapting to objects with varying shapes and sizes. In this paper, we introduce MISCGrasp, a volumetric grasping method that integrates multi-scale feature extraction with contrastive feature enhancement for self-adaptive grasping. We propose a query-based interac

Cited by 2SourcecodeScholar
2025

Matching While Perceiving: Enhance Image Feature Matching with Applicable Semantic Amalgamation

AAAI 2025technical

Image feature matching is a cardinal problem in computer vision, aiming to establish accurate correspondences between two-view images. Existing methods are constrained by the performance of feature extractors and struggle to capture local information affected by sparse texture or occlusions. Recogni…

2025

NeuGrasp: Generalizable Neural Surface Reconstruction with Background Priors for Material-Agnostic Object Grasp Detection

ICRA 2025

Robotic grasping in scenes with transparent and specular objects presents great challenges for methods relying on accurate depth information. In this paper, we introduce NeuGrasp, a neural surface reconstruction method that leverages background priors for material-agnostic grasp detection. NeuGrasp

Cited by 3SourcecodeScholar
2025

SENIOR: Efficient Query Selection and Preference-Guided Exploration in Preference-based Reinforcement Learning

IROS 2025

Preference-based Reinforcement Learning (PbRL) methods provide a solution to avoid reward engineering by learning reward models based on human preferences. However, poor feedback- and sample- efficiency still remain the problems that hinder the application of PbRL. In this paper, we present a novel

Cited by 0SourcecodeScholar
2024

A Robust Mutual-Reinforcing Framework for 3D Multi-Modal Medical Image Fusion Based on Visual-Semantic Consistency

AAAI 2024technical

This work proposes a robust 3D medical image fusion framework to establish a mutual-reinforcing mechanism between visual fusion and lesion segmentation, achieving their double improvement. Specifically, we explore the consistency between vision and semantics by sharing feature fusion modules. Throug…

2024

GSDF: 3DGS Meets SDF for Improved Neural Rendering and Reconstruction

NeurIPS 2024poster

Representing 3D scenes from multiview images remains a core challenge in computer vision and graphics, requiring both reliable rendering and reconstruction, which often conflicts due to the mismatched prioritization of image quality over precise underlying scene geometry. Although both neural implic…

Cited by 3SourcePDFScholar
2023

GPDAN: Grasp Pose Domain Adaptation Network for Sim-to-Real 6-DoF Object Grasping

RA-L 2023

In this letter, we propose a novel Grasp Pose Domain Adaptation Network (GPDAN) to achieve sim-to-real domain adaptation for 6-DoF grasp pose detection. The main task of GPDAN is to detect feasible 6-DoF grasp poses in cluttered scenes. A point-wise self-supervised domain classification module with

Cited by 16SourceScholar
2023

LinK: Linear Kernel for LiDAR-Based 3D Perception

CVPR 2023poster

Extending the success of 2D Large Kernel to 3D perception is challenging due to: 1. the cubically-increasing overhead in processing 3D data; 2. the optimization difficulties from data scarcity and sparsity. Previous work has taken the first step to scale up the kernel size from 3x3x3 to 7x7x7 by int…

2023

SparseBEV: High-Performance Sparse 3D Object Detection from Multi-Camera Videos

ICCV 2023poster

Camera-based 3D object detection in BEV (Bird's Eye View) space has drawn great attention over the past few years. Dense detectors typically follow a two-stage pipeline by first constructing a dense BEV feature and then performing object detection in BEV space, which suffers from complex view transf…

Cited by 127PDFcodeScholar
2023

Structure-Aware Multi-Feature Co-Learning for Dual Branch Face Super Resolution

ICASSP 2023accepted

Recently, face super-resolution has achieved pleasing performance. Numerous works have shown that texture features and structural information play a crucial role for super-resolution reconstruction. However, effective co-learning of both has been limiting the performance improvement of existing stat…

Cited by 0SourceScholar
2022

CamLiFlow: Bidirectional Camera-LiDAR Fusion for Joint Optical Flow and Scene Flow Estimation

CVPR 2022oral

In this paper, we study the problem of jointly estimating the optical flow and scene flow from synchronized 2D and 3D data. Previous methods either employ a complex pipeline that splits the joint task into independent stages, or fuse 2D and 3D information in an "early-fusion" or "late-fusion" manner…

Cited by 80PDFcodeScholar
2022

Degrade Is Upgrade: Learning Degradation for Low-Light Image Enhancement

AAAI 2022technical

Low-light image enhancement aims to improve an image's visibility while keeping its visual naturalness. Different from existing methods, which tend to accomplish the relighting task directly, we investigate the intrinsic degradation and relight the low-light image while refining the details and colo…

2022

Meta-Residual Policy Learning: Zero-Trial Robot Skill Adaptation via Knowledge Fusion

RA-L 2022

Adapting the mastered manipulation skill to novel objects is still challenging for robots. Recent works have attempted to endow the robot with the ability to adapt to unseen tasks by leveraging meta-learning. However, these methods are data-hungry in the training phase, which limits their applicatio

Cited by 23SourcecodeScholar
2021

DIMSAN: Fast Exploration with the Synergy between Density-based Intrinsic Motivation and Self-adaptive Action Noise

ICRA 2021poster

Exploration in environments with sparse rewards remains a challenging problem in Deep Reinforcement Learning (DRL). For the off-policy method, it usually needs a large number of training samples. With the growing dimensions of state and action space, this method becomes more and more sample-ineffici…

Cited by 0SourceScholar
2019

Localizing Discriminative Visual Landmarks for Place Recognition

ICRA 2019poster

We address the problem of visual place recognition with perceptual changes. The fundamental problem of visual place recognition is generating robust image representations which are not only insensitive to environmental changes but also distinguishable to different places. Taking advantage of the fea…

Cited by 66SourceScholar
2019

Self-modeling Tracking Control of Crawler Fire Fighting Robot Based on Causal Network

IROS 2019poster

In this paper, a self-modeling method based on a causal network is proposed for the tracking control of the Crawler Fire Fighting Robot (CFFR). The method mainly consists of two parts, one is a motion model, based on data driving, learning to establish the correspondence between control signal seque…

Cited by 0SourceScholar
2016

L1-L1 norms for face super-resolution with mixed Gaussian-impulse noise

ICASSP 2016accepted

In real world surveillance application, the captured faces are often low resolution (LR) and corrupted by mixed Gaussian-impulse noise during the acquisition and transmission processes. In this paper, we propose an effective patch-based face super-resolution method to reconstruct a high resolution (…

Cited by 0SourceScholar