← Search

Liu Liu

89 accepted papers

2026

3D-Fixer: Coarse-to-Fine In-place Completion for 3D Scenes from a Single Image

CVPR 2026

Compositional 3D scene generation from a single view requires the simultaneous recovery of scene layout and 3D assets. Existing approaches mainly fall into two categories: feed-forward generation methods and per-instance generation methods. The former directly predict 3D assets with explicit 6DoF po

Cited by 0SourcecodeScholar
2026

ContextPRM: Leveraging Contextual Coherence for multi-domain Test-Time Scaling

ICLR 2026poster

Process reward models (PRMs) have demonstrated significant efficacy in enhancing the mathematical reasoning capabilities of large language models (LLMs) by leveraging test-time scaling (TTS). However, while most PRMs exhibit substantial gains in mathematical domains, the scarcity of domain-specific…

Cited by 0SourceScholar
2026

CrystalDiT: Simple Diffusion Transformers for Crystal Generation

AAAI 2026technical

We present CrystalDiT, a diffusion transformer for crystal structure generation that achieves state-of-the-art performance by challenging the trend of architectural complexity. Instead of intricate, multi-stream designs, CrystalDiT employs a unified transformer that imposes a powerful inductive bias

Cited by 0SourcePDFScholar
2026

DICArt: Advancing Category-level Articulated Object Pose Estimation in Discrete State-Spaces

CVPR 2026

Articulated object pose estimation is a core task in embodied AI and computer vision. Existing methods typically regress poses in a continuous space, but often struggle with 1) navigating a large, complex search space and 2) failing to incorporate intrinsic kinematic constraints. In this paper, we i

Cited by 0SourceScholar
2026

DiffVL: Diffusion-Based Visual Localization on 2D Maps Via BEV-Conditioned GPS Denoising

ICRA 2026poster

Accurate visual localization is crucial for autonomous driving, yet existing methods face a fundamental dilemma: While high-definition (HD) maps provide high-precision localization references, their costly construction and maintenance hinder scalability, which drives research toward standard-definit…

2026

Efficient Quantization of Mixture-of-Experts with Theoretical Generalization Guarantees

ICLR 2026poster

Sparse Mixture-of-Experts (MoE) allows scaling of language and vision models efficiently by activating only a small subset of experts per input. While this reduces computation, the large number of parameters still incurs substantial memory overhead during inference. Post-training quantization has be…

Cited by 0SourcecodeScholar
2026

Exploring Category-level Articulated Object Pose Tracking on SE(3) Manifolds

AAAI 2026technical

Articulated objects are prevalent in daily life and robotic manipulation tasks. However, compared to rigid objects, pose tracking for articulated objects remains an underexplored problem due to their inherent kinematic constraints. To address these challenges, this work proposes a novel point-pair-b

Cited by 0SourcePDFScholar
2026

FRBAT: Conditionally-Visible Physical Backdoor Attack via Fluorescence

AAAI 2026technical

Deep neural networks are increasingly vulnerable to physically deployable backdoor attacks, which manipulate real-world objects to induce targeted model failures. However, current physical backdoor attacks predominantly rely on perpetually visible triggers appended to target objects. These methods i

Cited by 0SourcePDFScholar
2026

Generalizable and Actionable Parts Pose Estimation with Symmetry Annotation-Free Learning Strategy

ICML 2026poster

Urgently needed generalizable robot object interaction and manipulation requires high-quality Cross-Category object perception. As a pioneer of this area, Generalizable and Actionable Parts (GAParts) understanding has attracted increasing attention from relevant researchers. However, most recent wor…

Cited by 0SourceScholar
2026

IGFuse: Interactive 3D Gaussian Scene Reconstruction via Multi-Scans Fusion

AAAI 2026technical

Reconstructing complete and interactive 3D scenes remains a fundamental challenge in computer vision and robotics, particularly due to persistent object occlusions and limited sensor coverage. Even multi-view observations from a single scene scan often fail to capture the full structural details. Ex

Cited by 0SourcePDFScholar
2026

MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes

AAAI 2026technical

Understanding group-level social interactions in public spaces is crucial for urban planning, informing the design of socially vibrant and inclusive environments. Detecting such interactions from images involves interpreting subtle visual cues such as relations, proximity and co-movement – semantica

Cited by 0SourcePDFScholar
2026

ManiSplat: Manipulation Trajectory Synthesis from Monocular Video via Decoupled 3D Gaussian Splatting

IJCAI 2026

Reconstructing dynamic and interactive 3D scenes from real-world observations remains a fundamental challenge in computer vision and robotics. While recent advances in 3D Gaussian Splatting have enabled high-fidelity static reconstruction, extending it to interactive environments with articulated ro

Cited by 0Scholar
2026

REArtGS++: Generalizable Articulation Reconstruction with Temporal Geometry Constraint via Planar Gaussian Splatting

CVPR 2026

Articulated objects are pervasive in daily environments, such as drawers and refrigerators. Towards their part-level surface reconstruction and joint parameter estimation, REArtGS [??] introduces a category-agnostic approach using multi-view RGB images at two different states. However, we observe th

Cited by 0SourceScholar
2026

Remodeling Semantic Relationships in Vision-Language Fine-Tuning

AAAI 2026technical

Vision-language fine-tuning has emerged as an efficient paradigm for constructing multimodal foundation models. While textual context often highlights semantic relationships within an image, existing fine-tuning methods typically overlook this information when aligning vision and language, thus lead

Cited by 0SourcePDFScholar
2026

SIAM: Towards Generalizable Articulated Object Modeling via Single Robot-Object Interaction

AAAI 2026technical

Articulated object modeling, which represents interconnected rigid bodies with their geometry, part segmentation, articulation tree, and physical properties, is crucial for robotic perception and manipulation. Recently existing methods like SAGCI leverage Interactive Perception (IP) to refine models

Cited by 0SourcePDFScholar
2026

Uni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View Images

CVPR 2026

Reconstructing and semantically interpreting 3D scenes from sparse 2D views remains a fundamental challenge in computer vision. Conventional methods often decouple semantic understanding from reconstruction or necessitate costly per-scene optimization, thereby restricting their scalability and gener

Cited by 0SourcecodeScholar
2025

ArtGS: 3D Gaussian Splatting for Interactive Visual-Physical Modeling and Manipulation of Articulated Objects

IROS 2025

Articulated object manipulation remains a critical challenge in robotics due to the complex kinematic constraints and the limited physical reasoning of existing methods. In this work, we introduce ArtGS, a novel framework that extends 3D Gaussian Splatting (3DGS) by integrating visual-physical model

Cited by 8SourceScholar
2025

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding

CVPR 2025poster

The challenge in LLM-based video understanding lies in preserving visual and semantic information in long videos while maintaining a memory-affordable token count. However, redundancy and correspondence in videos have hindered the performance potential of existing methods. Through statistical learni…

Cited by 2SourcePDFScholar
2025

FruitMMBench: A Multi-modal Benchmark for Fruit Quality Assessment

ICASSP 2025accepted

The rapid advancement of Large Vision-Language Models (LVLMs) has brought notable improvements in tasks like visual recognition and multi-modal understanding, demonstrating significant potential in real-world applications. However, their performances on issues related to daily life such as fruit qua…

Cited by 0SourceScholar
2025

GaPT-DAR: Category-level Garments Pose Tracking via Integrated 2D Deformation and 3D Reconstruction

CVPR 2025poster

Garments are common in daily life and are important for embodied intelligence community. Current category-level garments pose tracking works focus on predicting point-wise canonical correspondence and learning a shape deformation in point cloud sequences. In this paper, motivated by the 2D warping s…

Cited by 0SourcePDFScholar
2025

GaussTR: Foundation Model-Aligned Gaussian Transformer for Self-Supervised 3D Spatial Understanding

CVPR 2025poster

3D Semantic Occupancy Prediction is fundamental for spatial understanding, yet existing approaches face challenges in scalability and generalization due to their reliance on extensive labeled data and computationally intensive voxel-wise representations. In this paper, we introduce GaussTR, a novel…

2025

Generalizable Articulated Object Perception with Superpoints

ICASSP 2025accepted

Manipulating articulated objects with robotic arms is challenging due to the complex kinematic structure, which requires precise part segmentation for efficient manipulation. In this work, we introduce a novel superpoint-based perception method designed to improve part segmentation in 3D point cloud…

Cited by 0SourceScholar
2025

Generalizable and Actionable Part Detection and Manipulation with SAM-rectified Segmentation and Iterative Pose Refinement

IROS 2025

The ability to perform cross-category object perception and manipulation is highly desirable in building intelligent robots. One promising approach is to define the concept of Generalizable and Actionable Parts (GAParts), such as buttons and handles, on both seen and unseen object categories. Howeve

Cited by 0SourceScholar
2025

GeoFlow-SLAM: A Robust Tightly-Coupled RGBD-Inertial and Legged Odometry Fusion SLAM for Dynamic Legged Robotics

IROS 2025

This paper presents GeoFlow-SLAM, a robust and effective Tightly-Coupled RGBD-Inertial and Legged Odometry Fusion SLAM for legged robotics undergoing aggressive and high-frequency motions. By integrating geometric consistency, legged odometry constraints, and dual-stream optical flow (GeoFlow), our

Cited by 1SourcecodeScholar
2025

HDT: Hierarchical Discrete Transformer for Multivariate Time Series Forecasting

AAAI 2025technical

Generative models have gained significant attention in multivariate time series forecasting (MTS), particularly due to their ability to generate high-fidelity samples. Forecasting the probability distribution of multivariate time series is a challenging yet practical task. Although some recent attem…

2025

MONTROSE: LLM-driven Monte Carlo Tree Search Self-Refinement for Cross-Domain Rumor Detection

ACL 2025finding

With the emergence of new topics on social media as sources of rumor dissemination, addressing the distribution shifts between source and target domains remains a crucial task in cross-domain rumor detection. Existing feature alignment methods, which aim to reduce the discrepancies between domains,…

2025

Multi-Task Vehicle Routing Solver via Mixture of Specialized Experts under State-Decomposable MDP

NeurIPS 2025poster

Existing neural methods for multi-task vehicle routing problems (VRPs) typically learn unified solvers to handle multiple constraints simultaneously. However, they often underutilize the compositional structure of VRP variants, each derivable from a common set of basis VRP variants. This critical ov…

Cited by 0SourceScholar
2025

Pre-defined Keypoints Promote Category-level Articulation Pose Estimation via Multi-Modal Alignment

IJCAI 2025

Articulations are essential in everyday interactions, yet traditional RGB-based pose estimation methods often struggle with issues such as lighting variations and shadows. To overcome these challenges, we propose a novel Pre-defined keypoint based framework for category-level articulation pose estim

Cited by 0SourcePDFScholar
2025

Principled Data Selection for Alignment: The Hidden Risks of Difficult Examples

ICML 2025poster

The alignment of large language models (LLMs) often assumes that using more clean data yields better outcomes, overlooking the match between model capacity and example difficulty. Challenging this, we propose a new principle: *Preference data vary in difficulty, and overly difficult examples hinder…

2025

ProBench: Judging Multimodal Foundation Models on Open-ended Multi-domain Expert Tasks

ACL 2025finding

Solving expert-level multimodal tasks is a key milestone in general intelligence. As the capabilities of multimodal large language models (MLLMs) continue to evolve, evaluation of frontier multimodal intelligence becomes necessary yet challenging. In this work, we introduce ProBench, a benchmark of…

Cited by 0SourcePDFScholar
2025

REArtGS: Reconstructing and Generating Articulated Objects via 3D Gaussian Splatting with Geometric and Motion Constraints

NeurIPS 2025poster

Articulated objects, as prevalent entities in human life, their 3D representations play crucial roles across various applications. However, achieving both high-fidelity textured surface reconstruction and dynamic generation for articulated objects remains challenging for existing methods. In this pa…

Cited by 0SourceScholar
2025

R^2-Art: Category-Level Articulation Pose Estimation from Single RGB Image via Cascade Render Strategy

AAAI 2025technical

Human life is filled with articulated objects. Previous works for estimating the pose of category-level articulated objects rely on costly 3D point clouds or RGB-D images. In this paper, our goal is to estimate category-level articulation poses from a single RGB image, where we propose R2-Art, a nov…

Cited by 0SourcePDFScholar
2025

SIGMA: Refining Large Language Model Reasoning via Sibling-Guided Monte Carlo Augmentation

NeurIPS 2025poster

Enhancing large language models by simply scaling up datasets has begun to yield diminishing returns, shifting the spotlight to data quality. Monte Carlo Tree Search (MCTS) has emerged as a powerful technique for generating high-quality chain-of-thought data, yet conventional approaches typically re…

Cited by 0SourceScholar
2025

Storyboard-guided Alignment for Fine-grained Video Action Recognition

NeurIPS 2025poster

Fine-grained video action recognition can be formulated as a video–text matching problem. Previous approaches primarily rely on global video semantics to consolidate video embeddings, often leading to misaligned video–text pairs due to inaccurate atomic-level action understanding. This inaccuracy ar…

Cited by 0SourceScholar
2025

Test-time Adapted Reinforcement Learning with Action Entropy Regularization

ICML 2025poster

Offline reinforcement learning is widely applied in multiple fields due to its advantages in efficiency and risk control. However, a major problem it faces is the distribution shift between offline datasets and online environments. This mismatch leads to out-of-distribution (OOD) state-action pairs…

Cited by 0SourcePDFScholar
2025

Towards Robust Category-level Articulation Pose Estimation via Integrated Differentiable Rendering

ICASSP 2025accepted

Accurate object pose estimation is crucial for embodied intelligence tasks such as manipulation, grasping, and human-robot interaction. However, due to the inherent characteristics of articulated objects, such as kinematic constraints and self-occlusion, pose estimation for articulated objects has r…

Cited by 0SourceScholar
2025

UniAff: A Unified Representation of Affordances for Tool Usage and Articulation with Vision-Language Models

ICRA 2025

Previous studies on robotic manipulation are based on a limited understanding of the underlying 3D motion constraints and affordances. To address these challenges, we propose a comprehensive paradigm, termed UniAff, that integrates 3D object-centric manipulation and task understanding in a unified f

Cited by 10SourceScholar
2024

EfficientCAPER: An End-to-End Framework for Fast and Robust Category-Level Articulated Object Pose Estimation

NeurIPS 2024poster

Human life is populated with articulated objects. Pose estimation for category-level articulated objects is a significant challenge due to their inherent complexity and diverse kinematic structures. Current methods for this task usually meet the problems of insufficient consideration of kinematic co…

Cited by 0SourcePDFScholar
2024

Event-based Few-shot Fine-grained Human Action Recognition

IROS 2024poster

Few-shot fine-grained human (FGH) action recognition is crucial in the context of human-robot interaction within open-set real-world environments. Existing works mainly focus on features extracted from RGB frames. However, their performances are drastically impacted in challenging scenarios, such as…

Cited by 1SourceScholar
2024

GAMMA: Generalizable Articulation Modeling and Manipulation for Articulated Objects

ICRA 2024poster

Articulated objects like cabinets and doors are widespread in daily life. However, directly manipulating 3D articulated objects is challenging because they have diverse geometrical shapes, semantic categories, and kinetic constraints. Prior works mostly focused on recognizing and manipulating articu…

Cited by 15SourcecodeScholar
2024

KPA-Tracker: Towards Robust and Real-Time Category-Level Articulated Object 6D Pose Tracking

AAAI 2024technical

Our life is populated with articulated objects. Current category-level articulation estimation works largely focus on predicting part-level 6D poses on static point cloud observations. In this paper, we tackle the problem of category-level online robust and real-time 6D pose tracking of articulated…

2024

RPMArt: Towards Robust Perception and Manipulation for Articulated Objects

IROS 2024poster

Articulated objects are commonly found in daily life. It is essential that robots can exhibit robust perception and manipulation skills for articulated objects in real-world robotic applications. However, existing methods for articulated objects insufficiently address noise in point clouds and strug…

Cited by 4SourcecodeScholar
2024

Rethinking 3D Convolution in $\ell_p$-norm Space

NeurIPS 2024spotlight

Convolution is a fundamental operation in the 3D backbone. However, under certain conditions, the feature extraction ability of traditional convolution methods may be weakened. In this paper, we introduce a new convolution method based on $\ell_p$-norm. For theoretical support, we prove the univer…

Cited by 9SourcePDFScholar
2024

Thermal-NeRF: Neural Radiance Fields from an Infrared Camera

IROS 2024poster

In recent years, Neural Radiance Fields (NeRFs) have demonstrated significant potential in encoding highly-detailed 3D geometry and environmental appearance, positioning themselves as a promising alternative to traditional explicit representation for 3D scene reconstruction. However, the predominant…

Cited by 13SourcecodeScholar
2024

U-COPE: Taking a Further Step to Universal 9D Category-level Object Pose Estimation

ECCV 2024poster

"Rigid and articulated objects are common in our daily lives. Pose estimation tasks for both types of objects have been extensively studied within their respective domains. However, a universal framework capable of estimating the pose of both rigid and articulated objects has yet to be reported. In…

Cited by 3SourcePDFScholar
2023

BEEF: Bi-Compatible Class-Incremental Learning via Energy-Based Expansion and Fusion

ICLR 2023poster

Neural networks suffer from catastrophic forgetting when sequentially learning tasks phase-by-phase, making them inapplicable in dynamically updated systems. Class-incremental learning (CIL) aims to enable neural networks to learn different categories at multi-stages. Recently, dynamic-structure-bas…

2023

Class-Aware Patch Embedding Adaptation for Few-Shot Image Classification

ICCV 2023poster

"A picture is worth a thousand words", significantly beyond mere a categorization. Accompanied by that, many patches of the image could have completely irrelevant meanings with the categorization if they were independently observed. This could significantly reduce the efficiency of a large family of…

Cited by 32PDFcodeScholar
2023

K3DN: Disparity-Aware Kernel Estimation for Dual-Pixel Defocus Deblurring

CVPR 2023poster

The dual-pixel (DP) sensor captures a two-view image pair in a single snapshot by splitting each pixel in half. The disparity occurs in defocus blurred regions between the two views of the DP pair, while the in-focus sharp regions have zero disparity. This motivates us to propose a K3DN framework fo…

Cited by 12SourcePDFScholar
2023

One-Shot Neural Band Selection for Spectral Recovery

ICASSP 2023accepted

Band selection has a great impact on the spectral recovery quality. To solve this ill-posed inverse problem, most band selection methods adopt hand-crafted priors or exploit clustering or sparse regularization constraints to find most prominent bands. These methods are either very slow due to the co…

Cited by 0SourceScholar
2023

Reject Decoding via Language-Vision Models for Text-to-Image Synthesis

AAAI 2023technical

Transformer-based text-to-image synthesis generates images from abstractive textual conditions and achieves prompt results. Since transformer-based models predict visual tokens step by step in testing, where the early error is hard to be corrected and would be propagated. To alleviate this issue, th…

2023

Semi-Supervised Contrastive Learning with Soft Mask Attention for Facial Action Unit Detection

ICASSP 2023accepted

This paper presents a novel facial action unit (AU) detection method by simultaneously improving AU feature’s discriminative ability and alleviating the AU data scarcity problem. We design a supervised AU soft mask attention scheme to learn local AU features by integrating prior expert knowledge. To…

Cited by 0SourceScholar
2022

AKB-48: A Real-World Articulated Object Knowledge Base

CVPR 2022poster

Human life is populated with articulated objects. A comprehensive understanding of articulated objects, namely appearance, structure, physics property, and semantics, will benefit many research communities. As current articulated object understanding solutions are usually based on synthetic object d…

Cited by 91PDFcodeScholar
2022

Balancing Stability and Plasticity through Advanced Null Space in Continual Learning

ECCV 2022poster

"Continual learning is a learning paradigm that learns tasks sequentially with resources constraints, in which the key challenge is stability-plasticity dilemma, i.e., it is uneasy to simultaneously have the stability to prevent catastrophic forgetting of old tasks and the plasticity to learn new ta…

Cited by 48SourcePDFScholar
2022

Channelized Axial Attention – considering Channel Relation within Spatial Attention for Semantic Segmentation

AAAI 2022technical

Spatial and channel attentions, modelling the semantic interdependencies in spatial and channel dimensions respectively, have recently been widely used for semantic segmentation. However, computing spatial and channel attentions separately sometimes causes errors, especially for those difficult case…

Cited by 45SourcePDFScholar
2022

Escaping from the Barren Plateau via Gaussian Initializations in Deep Variational Quantum Circuits

NeurIPS 2022accept

Variational quantum circuits have been widely employed in quantum simulation and quantum machine learning in recent years. However, quantum circuits with random structures have poor trainability due to the exponentially vanishing gradient with respect to the circuit depth and the qubit number. This…

Cited by 79SourcePDFScholar
2022

OakInk: A Large-Scale Knowledge Repository for Understanding Hand-Object Interaction

CVPR 2022poster

Learning how humans manipulate objects requires machines to acquire knowledge from two perspectives: one for understanding object affordances and the other for learning human's interactions based on the affordances. Even though these two knowledge bases are crucial, we find that current databases la…

Cited by 97PDFcodeScholar
2022

Online Continual Learning with Contrastive Vision Transformer

ECCV 2022poster

"Online continual learning (online CL) studies the problem of learning sequential tasks from an online data stream without task boundaries, aiming to adapt to new data while alleviating catastrophic forgetting on the past tasks. This paper proposes a framework Contrastive Vision Transformer (CVT), w…

Cited by 43SourcePDFScholar
2022

Resistance Training Using Prior Bias: Toward Unbiased Scene Graph Generation

AAAI 2022technical

Scene Graph Generation (SGG) aims to build a structured representation of a scene using objects and pairwise relationships, which benefits downstream tasks. However, current SGG methods usually suffer from sub-optimal scene graph generation because of the long-tailed distribution of training data. T…

2022

Text-to-Image Synthesis Based on Object-Guided Joint-Decoding Transformer

CVPR 2022poster

Object-guided text-to-image synthesis aims to generate images from natural language descriptions built by two-step frameworks, i.e., the model generates the layout and then synthesizes images from the layout and captions. However, such frameworks have two issues: 1) complex structure, since generati…

Cited by 17PDFScholar
2022

UMIX: Improving Importance Weighting for Subpopulation Shift via Uncertainty-Aware Mixup

NeurIPS 2022accept

Subpopulation shift widely exists in many real-world machine learning applications, referring to the training and test distributions containing the same subpopulation groups but varying in subpopulation frequencies. Importance reweighting is a normal way to handle the subpopulation shift issue by im…

2021

Activity Image-to-Video Retrieval by Disentangling Appearance and Motion

AAAI 2021technical

With the rapid emergence of video data, image-to-video retrieval has attracted much attention. There are two types of image-to-video retrieval: instance-based and activity-based. The former task aims to retrieve videos containing the same main objects as the query image, while the latter focuses on…

Cited by 26SourcePDFScholar
2021

Contrastive Graph Poisson Networks: Semi-Supervised Learning with Extremely Limited Labels

NeurIPS 2021poster

Graph Neural Networks (GNNs) have achieved remarkable performance in the task of semi-supervised node classification. However, most existing GNN models require sufficient labeled data for effective network training. Their performance can be seriously degraded when labels are extremely limited. To ad…

Cited by 65SourcePDFScholar
2021

Weak-shot Fine-grained Classification via Similarity Transfer

NeurIPS 2021poster

Recognizing fine-grained categories remains a challenging task, due to the subtle distinctions among different subordinate categories, which results in the need of abundant annotated samples. To alleviate the data-hungry problem, we consider the problem of learning novel categories from web data wit…

2020

Boosting Deep Neural Network Efficiency with Dual-Module Inference

ICML 2020poster

Using deep neural networks (DNNs) in machine learning tasks is promising in delivering high-quality results but challenging to meet stringent latency requirements and energy constraints because of the memory-bound and the compute-bound execution pattern of DNNs. We propose a big-little dual-module i…

2020

Channel Attention Based Iterative Residual Learning for Depth Map Super-Resolution

CVPR 2020poster

Despite the remarkable progresses made in deep learning based depth map super-resolution (DSR), how to tackle real-world degradation in low-resolution (LR) depth maps remains a major challenge. Existing DSR model is generally trained and tested on synthetic dataset, which is very different from what…

Cited by 103PDFScholar
2020

DoveNet: Deep Image Harmonization via Domain Verification

CVPR 2020poster

Image composition is an important operation in image processing, but the inconsistency between foreground and background significantly degrades the quality of composite image. Image harmonization, aiming to make the foreground compatible with the background, is a promising yet challenging task. Howe…

Cited by 260PDFcodeScholar
2020

Joint 3D Instance Segmentation and Object Detection for Autonomous Driving

CVPR 2020poster

Currently, in Autonomous Driving (AD), most of the 3D object detection frameworks (either anchor- or anchor-free-based) consider the detection as a Bounding Box (BBox) regression problem. However, this compact representation is not sufficient to explore all the information of the objects. To tackle…

Cited by 132PDFScholar
2020

Robust compressed sensing using generative models

NeurIPS 2020poster

We consider estimating a high dimensional signal in $\R^n$ using a sublinear number of linear measurements. In analogy to classical compressed sensing, here we assume a generative model as a prior, that is, we assume the signal is represented by a deep generative model $G: \R^k \rightarrow \R^n$. Cl…

2020

Solving the Blind Perspective-n-Point Problem End-To-End With Robust Differentiable Geometric Optimization

ECCV 2020poster

Blind Perspective-n-Point (PnP) is the problem of estimating the position and orientation of a camera relative to a scene, given 2D image points and 3D scene points, without prior knowledge of the 2D-3D correspondences. Solving for pose and correspondences simultaneously is extremely challenging sin…

2019

Dynamic Sparse Graph for Efficient Deep Learning

ICLR 2019poster

We propose to execute deep neural networks (DNNs) with dynamic and sparse graph (DSG) structure for compressive memory and accelerative execution during both training and inference. The great success of DNNs motivates the pursuing of lightweight models for the deployment onto embedded devices. Howev…

Cited by 69SourcePDFScholar
2019

Furcax: End-to-end Monaural Speech Separation Based on Deep Gated (De)convolutional Neural Networks with Adversarial Example Training

ICASSP 2019accepted

Deep gated convolutional networks have been proved to be very effective in single channel speech separation. However current state-of-the-art framework often considers training the gated convolutional networks in time-frequency (TF) domain. Such an approach will result in limited perceptual score, s…

Cited by 0SourceScholar
2019

Spatial-Aware Feature Aggregation for Image based Cross-View Geo-Localization

NeurIPS 2019poster

In this paper, we develop a new deep network to explicitly address these inherent differences between ground and aerial views. We observe there exist some approximate domain correspondences between ground and aerial images. Specifically, pixels lying on the same azimuth direction in an aerial image…