← Search

Yuan Liu

98 accepted papers

2026

AT-Field: Rethinking the Games in Adversarial Training

AAAI 2026technical

Adversarial training is often modeled as a two-player zero-sum game, relying on strong assumptions that limit its practical guidance. In this paper, we instead analyze the interactions between training samples and show that even the fundamental objective—minimizing training loss—may not converge. To

Cited by 0SourcePDFScholar
2026

BuildAnyPoint: 3D Building Structured Abstraction from Diverse Point Clouds

CVPR 2026

We introduce BuildAnyPoint, a novel generative framework for structured 3D building reconstruction from point clouds with diverse distributions, such as those captured by airborne LiDAR and Structure-from-Motion.To recover artist-created building abstraction in this highly underconstrained setting,

Cited by 0SourceScholar
2026

Closed-Loop Multimodal Sensory Training Enhances the Proprioceptive-Motor Pathway: Low-Load Automaticity and Fine Motor Control

ICRA 2026poster

Designing reliable upper-limb human-machine interfaces (HMIs) with low attentional demand requires strengthening the proprioceptive-motor pathway (PMP). We propose a closed-loop multimodal sensory training that maps three robot joint angles to six bidirectional electrotactile channels and combines v…

Cited by 0Scholar
2026

Dynamic Gaussian Scene Reconstruction from Unsynchronized Videos

AAAI 2026technical

Multi-view video reconstruction plays a vital role in computer vision, enabling applications in film production, virtual reality, and motion analysis. While recent advances such as 3D Gaussian Splatting have demonstrated impressive capabilities in dynamic scene reconstruction, they typically rely on

Cited by 0SourcePDFScholar
2026

FEDCOMPASS: FEDERATED CLUSTERED AND PERIODIC AGGREGATION FRAMEWORK FOR HYBRID CLASSICAL-QUANTUM MODELS

ICASSP 2026poster

Federated learning enables collaborative model training across decentralized clients under privacy constraints. Quantum computing offers potential for alleviating computational and communication burdens in federated learning, yet hybrid classical-quantum federated learning remains susceptible to per…

Cited by 0SourcePDFScholar
2026

FinMMDocR: Benchmarking Financial Multimodal Reasoning with Scenario Awareness, Document Understanding, and Multi-Step Computation

AAAI 2026technical

We introduce FinMMDocR, a novel bilingual multimodal benchmark for evaluating multimodal large language models (MLLMs) on real-world financial numerical reasoning. Compared to existing benchmarks, our work delivers three major advancements. (1) Scenario Awareness: 57.9% of 1,200 expert-annotated pro

Cited by 0SourcePDFScholar
2026

GAGS: Granularity-Aware Feature Distillation for Language Gaussian Splatting

AAAI 2026technical

3D open-vocabulary scene understanding, which accurately perceives complex semantic properties of objects in space, has gained significant attention in recent years. In this paper, we propose GAGS, a framework that distills 2D CLIP features into 3D Gaussian splatting, enabling open-vocabulary querie

Cited by 0SourcePDFScholar
2026

Kronecker Generative Networks: A General Neural Architecture for Parameter-Efficient Learning Across Classification Tasks

ICML 2026poster

Modern neural networks derive much of their effectiveness from rich connectivity patterns. Yet, existing architectures often fix the topology at either the sparse or dense extremes, thereby limiting structural flexibility and analysis. We propose Kronecker Generative Networks (KGNs), an algebraic fr…

Cited by 0SourceScholar
2026

MEGS^{2}: Memory-Efficient Gaussian Splatting via Spherical Gaussians and Unified Pruning

ICLR 2026poster

3D Gaussian Splatting (3DGS) has emerged as a dominant novel-view synthesis technique, but its high memory consumption severely limits its applicability on edge devices. A growing number of 3DGS compression methods have been proposed to make 3DGS more efficient, yet most only focus on storage compre…

Cited by 0SourcecodeScholar
2026

MV2UV: Generating High-quality UV Texture Maps with Multiview Prompts

CVPR 2026

Generating high-quality textures for 3D assets is a challenging task. Existing multiview texture generation methods suffer from the multiview inconsistency and missing textures on unseen parts, while UV inpainting texture methods do not generalize well due to insufficient UV data and cannot well uti

Cited by 0SourceScholar
2026

MeshMosaic: Scaling Artist Mesh Generation via Local-to-Global Assembly

CVPR 2026

Scaling artist-designed meshes to high triangle numbers remains challenging for autoregressive generative models. Existing transformer-based methods suffer from long-sequence bottlenecks and limited quantization resolution, primarily due to the large number of tokens required and constrained quantiz

Cited by 0SourcecodeScholar
2026

Neural Honeytrace: Plug&Play Watermarking Framework against Model Extraction Attacks

ICML 2026poster

Triggerable watermarking enables model owners to assert ownership against model extraction attacks. However, most existing approaches require additional training, which limits post-deployment flexibility, and the lack of clear theoretical foundations makes them vulnerable to adaptive attacks. In thi…

Cited by 0SourceScholar
2026

NeuralGS: Bridging Neural Fields and 3D Gaussian Splatting for Compact 3D Representations

AAAI 2026technical

3D Gaussian Splatting (3DGS) achieves impressive quality and rendering speed, but with millions of 3D Gaussians and significant storage and transmission costs. In this paper, we aim to develop a simple yet effective method called NeuralGS that compresses the original 3DGS into a compact representati

Cited by 0SourcePDFScholar
2026

PAT3D: Physics-Augmented Text-to-3D Scene Generation

ICLR 2026poster

We introduce PAT3D, the first physics-augmented text-to-3D scene generation framework that integrates vision–language models with physics-based simulation to produce physically plausible, simulation-ready, and intersection-free 3D scenes. Given a text prompt, PAT3D generates 3D objects, infers their…

Cited by 0SourcecodeScholar
2026

PFAvatar: Pose-Fusion 3D Personalized Avatar Reconstruction from Real-World Outfit-of-the-Day Photos

AAAI 2026technical

We propose PFAvatar (Pose-Fusion Avatar), a new method that reconstructs high-quality 3D avatars from Outfit of the Day (OOTD) photos, which exhibit diverse poses, occlusions, and complex backgrounds. Our method consists of two stages: (1) fine-tuning a pose-aware diffusion model from few-shot OOTD

Cited by 0SourcePDFScholar
2026

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs

CVPR 2026

Multimodal Large Language Models (MLLMs) have recently demonstrated remarkable capabilities in cross-modal understanding and generation. However, the rapid growth of visual token sequences--especially in long-video and streaming scenarios--poses a major challenge to their scalability and real-world

Cited by 0SourcecodeScholar
2026

PartSAM: A Scalable Promptable Part Segmentation Model Trained on Native 3D Data

ICLR 2026poster

Segmenting 3D objects into parts is a long-standing challenge in computer vision. To overcome taxonomy constraints and generalize to unseen 3D objects, recent works turn to open-world part segmentation. These approaches typically transfer supervision from 2D foundation models, such as SAM, by liftin…

Cited by 0SourcecodeScholar
2026

Partial Weakly-Supervised Oriented Object Detection

CVPR 2026

The growing demand for oriented object detection (OOD) across various domains has driven significant research in this area. However, the high cost of dataset annotation remains a major concern. Current mainstream OOD algorithms can be mainly categorized into three types: (1) fully supervised methods

Cited by 0SourcecodeScholar
2026

Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs

ICLR 2026poster

Multimodal large language models (MLLMs) have advanced rapidly in recent years. However, existing approaches for vision tasks often rely on indirect representations, such as generating coordinates as text for detection, which limits performance and prevents dense prediction tasks like segmentation.…

Cited by 0SourcecodeScholar
2026

RefAny3D: 3D Asset-Referenced Diffusion Models for Image Generation

ICLR 2026poster

In this paper, we propose a 3D asset-referenced diffusion model for image generation, exploring how to integrate 3D assets into image diffusion models. Existing reference-based image generation methods leverage large-scale pretrained diffusion models and demonstrate strong capability in generating d…

Cited by 0SourcecodeScholar
2026

SCoT: Teaching 3D-LLMs to Think Spatially with Million-scale CoT Annotations

ICLR 2026poster

Recent advances in 3D Large Language Models (3D-LLMs) show strong potential in understanding and interacting with 3D environments, yet their training data typically lack explicit reasoning processes, limiting complex spatial reasoning and task planning. To address this, we annotate SCoT, a million-s…

Cited by 0SourcecodeScholar
2026

Towards Long-Lived Robots: Continual Learning VLA Models via Reinforcement Fine-Tuning

RSS 2026poster

Pretrained on large-scale and diverse datasets, VLA models demonstrate strong generalization and adaptability as general-purpose robotic policies. However, Supervised Fine-Tuning (SFT), which serves as the primary mechanism for adapting VLAs to downstream domains, requires substantial amounts of tas…

Cited by 0SourceScholar
2026

UniSH: Unifying Scene and Human Reconstruction in a Feed-Forward Pass

CVPR 2026

We present UniSH, a unified, feed-forward framework for joint metric-scale 3D scene and human reconstruction. A key challenge in this domain is the scarcity of large-scale, annotated real-world data, forcing a reliance on synthetic datasets. This reliance introduces a significant sim-to-real domain

Cited by 0SourcecodeScholar
2026

WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference

ICML 2026oral

Autoregressive (AR) generation is the standard decoding paradigm for Large Language Models (LLMs), but its token-by-token nature limits parallelism at inference time. Diffusion Language Models (DLLMs) offer parallel decoding by recovering multiple masked tokens per step; however, in practice they of…

Cited by 0SourceScholar
2025

A Tightly Coupled Inertial-Sonar Fusion for Localization of Underwater Robots

IROS 2025

This paper proposes a tightly coupled fusion method for inertial and forward-looking sonar (FLS) data, integrating underwater image observations from the FLS into the inertial odometry for underwater robot localization. Since the FLS images provide only horizontal plane information, this work focuse

Cited by 0SourceScholar
2025

Align3R: Aligned Monocular Depth Estimation for Dynamic Videos

CVPR 2025highlight

Recent developments in monocular depth estimation methods enable high-quality depth estimation of single-view images but fail to estimate consistent video depth across different frames. Recent works address this problem by applying a video diffusion model to generate video depth conditioned on the i…

Cited by 14SourcePDFScholar
2025

CADDreamer: CAD Object Generation from Single-view Images

CVPR 2025highlight

The field of diffusion-based 3D generation has experienced tremendous progress in recent times. However, existing 3D generative models often produce overly dense and unstructured meshes, which are in stark contrast to the compact, structured and clear-edged CAD models created by human modelers. We i…

Cited by 0SourcePDFScholar
2025

CityAnchor: City-scale 3D Visual Grounding with Multi-modality LLMs

ICLR 2025poster

In this paper, we present a 3D visual grounding method called CityAnchor for localizing an urban object in a city-scale point cloud. Recent developments in multiview reconstruction enable us to reconstruct city-scale point clouds but how to conduct visual grounding on such a large-scale urban point…

Cited by 0SourcePDFScholar
2025

Contactless Vital Sign Monitoring for Multiple People Using a Millimeter-wave MIMO Radar

ICASSP 2025accepted

Radar technology offers much appeal for contactless vital sign monitoring. While most radar-based approaches achieve reasonable performance for single-person scenarios, they suffer from inaccurate vital sign estimates for multiple individuals, especially when the subjects occupy the same range bin.…

Cited by 0SourceScholar
2025

DICE: End-to-end Deformation Capture of Hand-Face Interactions from a Single Image

ICLR 2025poster

Reconstructing 3D hand-face interactions with deformations from a single image is a challenging yet crucial task with broad applications in AR, VR, and gaming. The challenges stem from self-occlusions during single-view hand-face interactions, diverse spatial relationships between hands and face, co…

2025

EasyHOI: Unleashing the Power of Large Models for Reconstructing Hand-Object Interactions in the Wild

CVPR 2025poster

Our work aims to reconstruct hand-object interactions from a single-view image, which is a fundamental but ill-posed task.Unlike methods that reconstruct from videos, multi-view images, or predefined 3D templates, single-view reconstruction faces significant challenges due to inherent ambiguities an…

2025

Epona: Autoregressive Diffusion World Model for Autonomous Driving

ICCV 2025poster

Diffusion models have demonstrated exceptional visual quality in video generation, making them promising for autonomous driving world modeling. However, existing video diffusion-based world models struggle with flexible-length, long-horizon predictions and integrating trajectory planning. This is be…

2025

EquiBench: Benchmarking Large Language Models’ Reasoning about Program Semantics via Equivalence Checking

EMNLP 2025

As large language models (LLMs) become integral to code-related tasks, a central question emerges: Do LLMs truly understand program semantics? We introduce EquiBench, a new benchmark for evaluating LLMs through equivalence checking, i.e., determining whether two programs produce identical outputs fo

2025

Exploiting Motion Prior for Accurate Pose Estimation of Dashboard Cameras

RA-L 2025

Dashboard cameras (dashcams) record millions of driving videos daily, offering a valuable potential data source for various applications, including driving map production and updates. A necessary step for utilizing these dashcam data involves the estimation of camera poses. However, the low-quality

Cited by 1SourceScholar
2025

Fine-Grained Global Modeling Learning for Personalized Federated Sequential Recommender

ICASSP 2025accepted

Personalized sequential recommender has become a key task in the consumer electronics domain. Existing methods for personalized sequential recommenders primarily focus on modeling user behavior and have achieved satisfactory recommender results. However, the inherent quadratic computational complexi…

Cited by 0SourceScholar
2025

GauUpdate: New Object Insertion in 3D Gaussian Fields with Consistent Global Illumination

ICCV 2025poster

3D Gaussian Splatting (3DGS) is a prevailing technique to reconstruct large-scale 3D scenes from multiview images for novel view synthesis, like a room, a block, and even a city. Such large-scale scenes are not static with changes constantly happening in these scenes, like a new building being built…

Cited by 0SourcePDFScholar
2025

Implementing Finite Impulse Response Filters on Quantum Computers

ICASSP 2025accepted

While signal processing is a mature area, its connections with quantum computing have received less attention. In this work, we propose approaches that perform classical discrete-time signal processing using quantum systems. Our approaches encode the classical discrete-time input signal into quantum…

Cited by 0SourceScholar
2025

Judge and Improve: Towards a Better Reasoning of Knowledge Graphs with Large Language Models

EMNLP 2025

Graph Neural Networks (GNNs) have shown immense potential in improving the performance of large-scale models by effectively incorporating structured relational information. However, current approaches face two key challenges: (1) achieving robust semantic alignment between graph representations and

Cited by 0SourcePDFScholar
2025

MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks

ICLR 2025poster

We present MEGA-Bench, an evaluation suite that scales multimodal evaluation to over 500 real-world tasks, to address the highly heterogeneous daily use cases of end users. Our objective is to optimize for a set of high-quality data samples that cover a highly diverse and rich set of multimodal task…

2025

ML$^2$-GCL: Manifold Learning Inspired Lightweight Graph Contrastive Learning

ICML 2025poster

Graph contrastive learning has attracted great interest as a dominant and promising self-supervised representation learning approach in recent years. While existing works follow the basic principle of pulling positive pairs closer and pushing negative pairs far away, they still suffer from several c…

2025

MVTokenFlow: High-quality 4D Content Generation using Multiview Token Flow

ICLR 2025poster

In this paper, we present MVTokenFlow for high-quality 4D content creation from monocular videos. Recent advancements in generative models such as video diffusion models and multiview diffusion models enable us to create videos or 3D models. However, extending these generative models for dynamic 4D…

2025

Micro-UAV with Ant-Inspired Bistable Gripper for Adaptive Perching and Wildlife Detection

IROS 2025

With the global ecological environment facing continuous deterioration, effective monitoring of arboreal birds in complex canopy environments remains challenging due to limitations of conventional drones in endurance, size, and habitat disturbance. To address these challenges, this paper presents an

Cited by 0SourceScholar
2025

MoDGS: Dynamic Gaussian Splatting from Casually-captured Monocular Videos with Depth Priors

ICLR 2025poster

In this paper, we propose MoDGS, a new pipeline to render novel-view images in dynamic scenes using only casually captured monocular videos. Previous monocular dynamic NeRF or Gaussian Splatting methods strongly rely on the rapid movement of input cameras to construct multiview consistency but fail…

2025

M²N: A Progressive Macro-to-Micro 3D Modeling Scheme for Unveiling Drug-Target Affinity

AAAI 2025technical

Accurate drug-target affinity (DTA) prediction holds significant potential in the field of artificial intelligence (AI)-based drug discovery. However, existing methods primarily operate at a single scale, specifically at the macro (residue) scale for target proteins and the micro (atom) scale for dr…

Cited by 0SourcePDFScholar
2025

OmniDiff: A Comprehensive Benchmark for Fine-grained Image Difference Captioning

ICCV 2025poster

Image Difference Captioning (IDC) aims to generate natural language descriptions of subtle differences between image pairs, requiring both precise visual change localization and coherent semantic expression. Despite recent advancements, existing datasets often lack breadth and depth, limiting their…

Cited by 0SourcePDFScholar
2025

OpenAnimals: Revisiting Person Re-Identification for Animals Towards Better Generalization

ICCV 2025poster

This paper addresses the challenge of animal re-identification, an emerging field that shares similarities with person re-identification but presents unique complexities due to the diverse species, environments and poses. To facilitate research in this domain, we introduce OpenAnimals, a flexible an…

2025

POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document Conversion

EMNLP 2025

High-quality labeled data is essential for training accurate document conversion models, particularly in domains with complex formats such as tables, formulas, and multi-column text. However, manual annotation is both costly and time-consuming, while automatic labeling using existing models often la

2025

PSHuman: Photorealistic Single-image 3D Human Reconstruction using Cross-Scale Multiview Diffusion and Explicit Remeshing

CVPR 2025poster

Photorealistic 3D human modeling is essential for various applications and has seen tremendous progress. However, existing methods for monocular full-body reconstruction, typically relying on front and/or predicted back view, still struggle with satisfactory performance due to the ill-posed nature o…

2025

PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video Generation

NeurIPS 2025poster

Existing video generation models excel at producing photo-realistic videos from text or images, but often lack physical plausibility and 3D controllability. To overcome these limitations, we introduce PhysCtrl, a novel framework for physics-grounded image-to-video generation with physical parameters…

Cited by 0SourceScholar
2025

RLBCD: Residual-guided Latent Brownian-bridge Co-Diffusion for Anatomical-to-Metabolic Image Synthesis

IJCAI 2025

While metabolic imaging can facilitate early diagnosis by revealing physiological changes of lesions, it is limited by high cost, high radiation risk, and potential renal impairment. Thus, developing an effective approach for Anatomical-to-Metabolic Image Synthesis (A2MIS) is highly required. Howeve

Cited by 0SourcePDFScholar
2025

RRM: Robust Reward Model Training Mitigates Reward Hacking

ICLR 2025poster

Reward models (RMs) play a pivotal role in aligning large language models (LLMs) with human preferences. However, traditional RM training, which relies on response pairs tied to specific prompts, struggles to disentangle prompt-driven preferences from prompt-independent artifacts, such as response l…

Cited by 0SourcePDFScholar
2025

Reward-Guided Prompt Evolving in Reinforcement Learning for LLMs

ICML 2025poster

Existing reinforcement learning (RL) methods for large language models (LLMs) rely on static prompt sets, where prompts are curated a priori, and sampled in a fixed schedule for training, regardless of their usefulness to the RL process. We design `eva`, the first method that allows LLMs to prioriti…

Cited by 0SourcePDFScholar
2025

Sparse Bayesian Network for Fast Micro-Doppler Analysis

ICASSP 2025accepted

Micro-Doppler Analysis (MDA) of rigid-body targets is crucial for various practical downstream tasks such as target imaging and recognition. Radar echoes from micro-moving targets typically represent non-stationary signals and are often described using the parameterized Time-Varying Auto Regressive…

Cited by 0SourceScholar
2025

SyncHuman: Synchronizing 2D and 3D Generative Models for Single-view Human Reconstruction

NeurIPS 2025poster

Photorealistic 3D full-body human reconstruction from a single image is a critical yet challenging task for applications in films and video games due to inherent ambiguities and severe self-occlusions. While recent approaches leverage SMPL estimation and SMPL-conditioned image generative models to h…

Cited by 0SourceScholar
2025

TrackingWorld: World-centric Monocular 3D Tracking of Almost All Pixels

NeurIPS 2025poster

Monocular 3D tracking aims to capture the long-term motion of pixels in 3D space from a single monocular video and has witnessed rapid progress in recent years. However, we argue that the existing monocular 3D tracking methods still fall short in separating the camera motion from foreground dynamic…

Cited by 0SourceScholar
2025

VistaDream: Sampling multiview consistent images for single-view scene reconstruction

ICCV 2025poster

In this paper, we propose VistaDream, a novel framework to reconstruct a 3D scene from a single-view image. Recent diffusion models enable generating high-quality novel-view images from a single-view input image. Most existing methods only concentrate on building the consistency between the input im…

Cited by 0SourcePDFScholar
2025

🎧MOSPA: Human Motion Generation Driven by Spatial Audio

NeurIPS 2025spotlight

Enabling virtual humans to dynamically and realistically respond to diverse auditory stimuli remains a key challenge in character animation, demanding the integration of perceptual modeling and motion synthesis. Despite its significance, this task remains largely unexplored. Most previous works have…

Cited by 0SourcecodeScholar
2024

"EMDM: Efficient Motion Diffusion Model for Fast, High-Quality Human Motion Generation"

ECCV 2024poster

"We introduce Efficient Motion Diffusion Model (EMDM) for fast and high-quality human motion generation. Current state-of-the-art generative diffusion models have produced impressive results but struggle to achieve fast generation without sacrificing quality. On the one hand, previous works, like mo…

2024

Debris Sensing Based on Leo Constellation: An Intersatellite Channel Parameter Estimation Approach

ICASSP 2024accepted

Space debris detection and tracking, a key enabler for Space Situational Awareness (SSA), poses two inherent challenges: (1) small-sized targets (e.g., 1 − 10 cm) posing detection difficulties for conventional ground-based radars (GBRs) and optical measurements; (2) large number resulting in a costl…

Cited by 0SourceScholar
2024

Disentangled Clothed Avatar Generation from Text Descriptions

ECCV 2024poster

"In this paper, we introduce a novel text-to-avatar generation method that separately generates the human body and the clothes and allows high-quality animation on the generated avatar. While recent advancements in text-to-avatar generation have yielded diverse human avatars from text prompts, these…

Cited by 24SourcePDFScholar
2024

Era3D: High-Resolution Multiview Diffusion using Efficient Row-wise Attention

NeurIPS 2024poster

In this paper, we introduce **Era3D**, a novel multiview diffusion method that generates high-resolution multiview images from a single-view image. Despite significant advancements in multiview generation, existing methods still suffer from camera prior mismatch, inefficacy, and low resolution, resu…

Cited by 7SourcePDFScholar
2024

Explicitly Guided Information Interaction Network for Cross-modal Point Cloud Completion

ECCV 2024poster

"∗ Equal contribution Corresponding authorIn this paper, we explore a novel framework, EGIInet (Explicitly Guided Information Interaction Network), a model for View-guided Point cloud Completion (ViPC) task, which aims to restore a complete point cloud from a partial one with a single view image. In…

2024

FreeReg: Image-to-Point Cloud Registration Leveraging Pretrained Diffusion Models and Monocular Depth Estimators

ICLR 2024poster

Matching cross-modality features between images and point clouds is a fundamental problem for image-to-point cloud registration. However, due to the modality difference between images and points, it is difficult to learn robust and discriminative cross-modality features by existing metric learning m…

2024

IntrinsicAnything: Learning Diffusion Priors for Inverse Rendering Under Unknown Illumination

ECCV 2024poster

"† Corresponding author. This paper aims to recover object materials from posed images captured under an unknown static lighting condition. Recent methods solve this task by optimizing material parameters through differentiable physically based rendering. However, due to the coupling between object…

2024

Surf-D: Generating High-Quality Surfaces of Arbitrary Topologies Using Diffusion Models

ECCV 2024poster

"We present Surf-D, a novel method for generating high-quality 3D shapes as Surfaces with arbitrary topologies using Diffusion models. Previous methods explored shape generation with different representations and they suffer from limited topologies and poor geometry details. To generate high-quality…

Cited by 1SourcePDFScholar
2024

SyncDreamer: Generating Multiview-consistent Images from a Single-view Image

ICLR 2024spotlight

In this paper, we present a novel diffusion model called SyncDreamer that generates multiview-consistent images from a single-view image. Using pretrained large-scale 2D diffusion models, recent work Zero123 demonstrates the ability to generate plausible novel views from a single-view image of an ob…

2024

Wonder3D: Single Image to 3D using Cross-Domain Diffusion

CVPR 2024highlight

In this work we introduce Wonder3D a novel method for generating high-fidelity textured meshes from single-view images with remarkable efficiency. Recent methods based on the Score Distillation Sampling (SDS) loss methods have shown the potential to recover 3D geometry from 2D diffusion priors but t…

Cited by 414SourcePDFScholar
2023

F2-NeRF: Fast Neural Radiance Field Training With Free Camera Trajectories

CVPR 2023highlight

This paper presents a novel grid-based NeRF called F^2-NeRF (Fast-Free-NeRF) for novel view synthesis, which enables arbitrary input camera trajectories and only costs a few minutes for training. Existing fast grid-based NeRF training frameworks, like Instant-NGP, Plenoxels, DVGO, or TensoRF, are ma…

2023

Improving Pixel-based MIM by Reducing Wasted Modeling Capability

ICCV 2023poster

There has been significant progress in Masked Image Modeling (MIM). Existing MIM methods can be broadly categorized into two groups based on the reconstruction target: pixel-based and tokenizer-based approaches. The former offers a simpler pipeline and lower computational cost, but it is known to be…

Cited by 48PDFcodeScholar
2023

Mitigating and Evaluating Static Bias of Action Representations in the Background and the Foreground

ICCV 2023oral

In video action recognition, shortcut static features can interfere with the learning of motion features, resulting in poor out-of-distribution (OOD) generalization. The video background is clearly a source of static bias, but the video foreground, such as the clothing of the actor, can also provide…

Cited by 18PDFcodeScholar
2023

NeuralUDF: Learning Unsigned Distance Fields for Multi-View Reconstruction of Surfaces With Arbitrary Topologies

CVPR 2023poster

We present a novel method, called NeuralUDF, for reconstructing surfaces with arbitrary topologies from 2D images via volume rendering. Recent advances in neural rendering based reconstruction have achieved compelling results. However, these methods are limited to objects with closed surfaces since…

Cited by 67SourcePDFScholar
2023

Robust Multiview Point Cloud Registration With Reliable Pose Graph Initialization and History Reweighting

CVPR 2023poster

In this paper, we present a new method for the multiview registration of point cloud. Previous multiview registration methods rely on exhaustive pairwise registration to construct a densely-connected pose graph and apply Iteratively Reweighted Least Square (IRLS) on the pose graph to compute the sca…

2022

Bi-Directional Modality Fusion Network For Audio-Visual Event Localization

ICASSP 2022accepted

Audio and visual signals stimulate many audio-visual sensory neurons of persons to generate audio-visual contents, helping humans perceive the world. Most of the existing audio-visual event localization approaches focus on generating audio-visual features by fusing the audio and visual modalities fo…

Cited by 0SourceScholar
2022

Diagnosing failures of fairness transfer across distribution shift in real-world medical settings

NeurIPS 2022accept

Diagnosing and mitigating changes in model fairness under distribution shift is an important component of the safe deployment of machine learning in healthcare settings. Importantly, the success of any mitigation strategy strongly depends on the \textit{structure} of the shift. Despite this, there h…

Cited by 74SourcePDFScholar
2022

Gen6D: Generalizable Model-Free 6-DoF Object Pose Estimation from RGB Images

ECCV 2022poster

"In this paper, we present a generalizable model-free 6-DoF object pose estimator called Gen6D. Existing generalizable pose estimators either need the high-quality object models or require additional depth maps or object masks in test time, which significantly limits their application scope. In cont…

2022

Joint Source Localization and Association Through Overcomplete Representation Under Multipath Propagation Environment

ICASSP 2022accepted

This work addresses the source localization and association problem in a multipath propagation environment. By focusing on the limitation of the prior information in practical applications, we propose a target localization and association method based on iterative optimization with semi-unitary cons…

Cited by 0SourceScholar
2022

Multichannel Noise Reduction Using Dilated Multichannel U-Net and Pre-Trained Single-Channel Network

ICASSP 2022accepted

Pre-trained single-channel neural networks have become more prevalent for noise reduction in recent years. However, unlike their multichannel counterparts, these monoaural approaches do not exploit spatial information during the optimization process. Furthermore, while multichannel neural networks e…

Cited by 0SourceScholar
2022

Neural Rays for Occlusion-Aware Image-Based Rendering

CVPR 2022poster

We present a new neural representation, called Neural Ray (NeuRay), for the novel view synthesis task. Recent works construct radiance fields from image features of input views to render novel view images, which enables the generalization to new scenes. However, due to occlusions, a 3D point may be…

Cited by 234PDFcodeScholar
2021

AdaFit: Rethinking Learning-Based Normal Estimation on Point Clouds

ICCV 2021poster

This paper presents a neural network for robust normal estimation on point clouds, named AdaFit, that can deal with point clouds with noise and density variations. Existing works use a network to learn point-wise weights for weighted least squares surface fitting to estimate the normals, which has d…

Cited by 56PDFcodeScholar
2021

Contrast and Order Representations for Video Self-Supervised Learning

ICCV 2021poster

This paper studies the problem of learning self-supervised representations on videos. In contrast to image modality that only requires appearance information on objects or scenes, video needs to further explore the relations between multiple frames/clips along the temporal dimension. However, the re…

Cited by 77PDFcodeScholar
2021

Improving 3D Object Detection With Channel-Wise Transformer

ICCV 2021poster

Though 3D object detection from point clouds has achieved rapid progress in recent years, the lack of flexible and high-performance proposal refinement remains a great hurdle for existing state-of-the-art two-stage detectors. Previous works on refining 3D proposals have relied on human-designed comp…

Cited by 298PDFcodeScholar
2021

NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view Reconstruction

NeurIPS 2021spotlight

We present a novel neural surface reconstruction method, called NeuS, for reconstructing objects and scenes with high fidelity from 2D image inputs. Existing neural surface reconstruction approaches, such as DVR [Niemeyer et al., 2020] and IDR [Yariv et al., 2020], require foreground mask as supervi…

2021

Point2Skeleton: Learning Skeletal Representations from Point Clouds

CVPR 2021poster

We introduce Point2Skeleton, an unsupervised method to learn skeletal representations from point clouds. Existing skeletonization methods are limited to tubular shapes and the stringent requirement of watertight input, while our method aims to produce more generalized skeletal representations for co…

Cited by 70PDFScholar
2021

The Blessings of Unlabeled Background in Untrimmed Videos

CVPR 2021poster

Weakly-supervised Temporal Action Localization (WTAL) aims to detect the action segments with only video-level action labels in training. The key challenge is how to distinguish the action of interest segments from the background, which is unlabelled even on the video-level. While previous works tre…

Cited by 50PDFcodeScholar
2020

Improving Motion Planning for Surgical Robot with Active Constraints

IROS 2020poster

In this paper, an improved motion planning scheme is proposed for surgical robot control with multiple active constraints, including joint constraints, joint velocity constraints and remote center of motion constraints. It introduces an improved recurrent neural network (RNN) to optimize the online…

Cited by 6SourceScholar
2020

Object Tracking using Spatio-Temporal Networks for Future Prediction Location

ECCV 2020poster

We introduce an object tracking algorithm that predicts the future locations of the target object and assists the tracker to handle object occlusion. Given a few frames of an object that are extracted from a complete input sequence, we aim to predict the object’s location in the future frames. To fa…

Cited by 31SourcePDFScholar
2019

GIFT: Learning Transformation-Invariant Dense Visual Descriptors via Group CNNs

NeurIPS 2019poster

Finding local correspondences between images with different viewpoints requires local descriptors that are robust against geometric transformations. An approach for transformation invariance is to integrate out the transformations by pooling the features extracted from transformed versions of an ima…

Cited by 109SourcePDFScholar
2019

Investigation of Modeling Units for Mandarin Speech Recognition Using Dfsmn-ctc-smbr

ICASSP 2019accepted

The choice of acoustic modeling units is critical to acoustic modeling in large vocabulary continuous speech recognition (LVCSR) tasks. The recent connectionist temporal classification (CTC) based acoustic models have more options for the choice of modeling units. In this work, we propose a DFSMN-CT…

Cited by 0SourceScholar
2019

PVNet: Pixel-Wise Voting Network for 6DoF Pose Estimation

CVPR 2019oral

This paper addresses the challenge of 6DoF pose estimation from a single RGB image under severe occlusion or truncation. Many recent works have shown that a two-stage approach, which first detects keypoints and then solves a Perspective-n-Point (PnP) problem for pose estimation, achieves remarkable…

Cited by 1355PDFcodeScholar
2018

Altitude Measurement of Low-Angle Target Under Complex Terrain Environment for Meter-Wave Radar

ICASSP 2018accepted

For modern meter-wave radar, the performance of low-angle target altitude measurement is limited by multipath phenomenon, especially in the complex terrain environment where the multipath signal is perturbed by irregular surface. To address this problem, a practical signal model for meter-wave radar…

Cited by 0SourceScholar
2017

A novel actuation configuration of robotic hand and the mechanical implementation via postural synergies

ICRA 2017poster

How to design a robotic hand for reproducing the move characteristics of human hand joints is a big challenge in robotics. In this paper, we present an approach to determine the actuation configuration based on the statistical results of hand joint angle in different grasps. A relationship between t…

Cited by 8SourceScholar