← Search

Lei Yang

114 accepted papers

2026

AdaThinkDrive: Adaptive Thinking Via Reinforcement Learning for Autonomous Driving

ICRA 2026poster

While reasoning technology like Chain-of-Thought (CoT) has been widely adopted in Vision-Language-Action (VLA) models, it demonstrates promising capabilities in end-to-end autonomous driving. However, recent efforts to integrate CoT reasoning often fall short in simple scenarios, introducing unneces…

2026

Approximated Collision Detection for Contact-Rich Dexterous Manipulation with Nonnegative Least Squares

ICRA 2026poster

Collision detection between robotic hands and manipulated objects is crucial to model predictive control (MPC) for contact-rich dexterous manipulation. Based on the Gilbert-Johnson-Keerthi (GJK) algorithm and the expanding polytope algorithm (EPA), the GJK-EPA method has achieved success while requi…

Cited by 0Scholar
2026

Beyond the Clouds: Reliable and Cloud-Aware Spatiotemporal Fusion via Adversarial Regression Wavelets

IJCAI 2026

Spatiotemporal fusion (STF) bridges the gap between temporal and spatial resolutions in satellite imagery, enabling effective monitoring of Earth's surface dynamics. However, existing methods rely on cloud-free reference images, a constraint that fails in realistic, cloud-prone scenarios. To overcom

Cited by 0Scholar
2026

ConsistCompose: Unified Multimodal Layout Control for Image Composition

CVPR 2026

Unified multimodal models that couple visual understanding with image generation have advanced rapidly, yet most systems still focus on visual grounding--aligning language with image regions--while their generative counterpart, linguistic-embedded layout-grounded generation(LELG) for layout-controll

Cited by 0SourcecodeScholar
2026

Cross-Modal Unlearning via Influential Neuron Path Editing in Multimodal Large Language Models

AAAI 2026technical

Multimodal Large Language Models (MLLMs) extend foundation models to real-world applications by integrating inputs such as text and vision. However, their broad knowledge capacity raises growing concerns about privacy leakage, toxicity mitigation, and intellectual property violations. Machine Unlear

Cited by 0SourcePDFScholar
2026

DSSG: Dual-Stream Semantic Guidance for Source-Fully-Free Adaptation of Vision-Language Models

IJCAI 2026

Source-Fully-Free Domain Adaptation (SFF-DA) has emerged as a strategic paradigm to adapt Vision-Language Models (VLMs) without any access to source data or task-specific source models. However, we identify a critical "Dual Semantic Drift" that hinders this process: static drift caused by the stagna

Cited by 0Scholar
2026

DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving

ICML 2026poster

End-to-end autonomous driving systems are increasingly integrating Vision-Language Model (VLM) architectures, incorporating text reasoning or visual reasoning to enhance the robustness and accuracy of driving decisions. However, the reasoning mechanisms employed in most methods are direct adaptation…

Cited by 0SourceScholar
2026

EmbodMocap: In-the-Wild 4D Human-Scene Reconstruction for Embodied Agents

CVPR 2026

Human behaviors in the real world naturally encode rich, long-term contextual information that can be leveraged to train embodied agents for perception, understanding, and acting.However, existing capture systems typically rely on costly studio setups and wearable devices, limiting the large-scale c

Cited by 0SourcecodeScholar
2026

Enhancing the Security of Visual Speaker Authentication Based on Dynamic Lip-Print Analysis

CVPR 2026

In recent years, face-based authentication methods are gradually replacing traditional methods across various applications, offering enhanced security and user convenience. However, these methods are threatened by the continuously evolving DeepFake techniques. In this paper, a novel Visual Speaker A

Cited by 0SourceScholar
2026

FedLAGC: Towards High Performance System-Heterogeneous Federated Learning via Layer-Adaptive Submodel Extraction and Gradient Correction

AAAI 2026technical

Federated learning has emerged as a promising paradigm for collaborative model training while preserving data privacy. However, many existing FL methods implicitly assume that clients have sufficient computational and storage resources, making them less applicable in real-world scenarios with severe

Cited by 0SourcePDFScholar
2026

GUI-ReWalk: Massive Data Generation for GUI Agent via Stochastic Exploration and Intent-Aware Reasoning

IJCAI 2026

Graphical User Interface (GUI) Agents, powered by large language and vision-language models, hold promise for enabling end-to-end automation in digital environments. However, their progress is fundamentally constrained by the scarcity of scalable, high-quality trajectory data. Existing data collecti

Cited by 0Scholar
2026

HiddenEcho: Mitigating Noise Amplification in Differentially Private LLMs with Hidden-State Correction

ICLR 2026poster

The rise of large language models (LLMs) has driven the adoption of Model-as-a-Service (MaaS). However, transmitting raw text to servers raises critical privacy concerns. Existing approaches employ deep neural networks (DNNs) or differential privacy (DP) to perturb inputs. Yet, these approaches suff…

Cited by 0SourceScholar
2026

MEANSE: EFFICIENT GENERATIVE SPEECH ENHANCEMENT WITH MEAN FLOWS

ICASSP 2026poster

Speech enhancement (SE) improves degraded speech's quality, with generative models like flow matching gaining attention for their outstanding perceptual quality. However, the flow-based model requires multiple numbers of function evaluations (NFEs) to achieve stable and satisfactory performance, lea…

Cited by 0SourcePDFScholar
2026

Optimized Decoupling Control and Cross-Domain Validation of a V-Shaped Quadrant-Habitat Fan-Wing Vehicle - Optimized for Underwater Operations

RA-L 2026

Multi-habitat robots can operate across diverse environments, yet mainstream platforms still depend on fixed-wing/propeller actuation and “multi-controller–multi-hardware” architectures that hinder lightweight, integrated designs. This paper proposes a fan-wing–centric aquatic cross-habitat scheme:

Cited by 0SourceScholar
2026

PHOTONS: Pose-Free Human-Centric Photo-Realistic Real-Time Novel View Synthesis from Sparse Views

AAAI 2026technical

We present PHOTONS (Pose-Free Human-Centric Photo-Realistic Real-Time Novel View Synthesis from Sparse Views), a real-time framework for novel view synthesis without requiring camera calibration. Our method reconstructs consistent 3D Gaussian point clouds and synthesizes 2K photo-realistic novel vie

Cited by 0SourcePDFScholar
2026

Phased DMD: Few-step Distribution Matching Distillation via Score Matching within Subintervals

CVPR 2026

Distribution Matching Distillation (DMD) distills score-based generative models into efficient one-step generators, without requiring a one-to-one correspondence with the sampling trajectories of their teachers. Yet, the limited capacity of one-step distilled models compromises generative diversity

Cited by 0SourcecodeScholar
2026

SGMD: Score Gradient Matching Distillation for Few-Step Video Diffusion Distillation

ICML 2026poster

Distribution Matching Distillation (DMD) is a widely used paradigm for accelerating inference in few-step video diffusion models. However, DMD-style training faces a structural bottleneck: the student-side auxiliary score network (the fake score) must closely track a continuously evolving generator.…

Cited by 0SourceScholar
2026

Scaling Spatial Intelligence with Multimodal Foundation Models

CVPR 2026

Despite remarkable progress, multimodal foundation models still exhibit surprising deficiencies in spatial intelligence. In this work, we explore scaling up multimodal foundation models to cultivate spatial intelligence within the SenseNova-SI family, built upon established multimodal foundations in

Cited by 0SourcecodeScholar
2026

The Quest for Generalizable Motion Generation: Data, Model, and Evaluation

ICLR 2026poster

Despite recent advances in 3D human motion generation (MoGen) on standard benchmarks, existing models still face a fundamental bottleneck in their generalization capability. In contrast, adjacent generative fields, most notably video generation (ViGen), have demonstrated remarkable generalization in…

Cited by 0SourcecodeScholar
2025

A Novel Diffusion Model for Pairwise Geoscience Data Generation with Unbalanced Training Dataset

AAAI 2025technical

Recently, the advent of generative AI technologies has made transformational impacts on our daily lives, yet its application in scientific applications remains in its early stages. Data scarcity is a major, well-known barrier in data-driven scientific computing, so physics-guided generative AI hold…

Cited by 0SourcePDFScholar
2025

ADHMR: Aligning Diffusion-based Human Mesh Recovery via Direct Preference Optimization

ICML 2025poster

Human mesh recovery (HMR) from a single image is inherently ill-posed due to depth ambiguity and occlusions. Probabilistic methods have tried to solve this by generating numerous plausible 3D human mesh predictions, but they often exhibit misalignment with 2D image observations and weak robustness t…

2025

Clients Collaborate: Flexible Differentially Private Federated Learning with Guaranteed Improvement of Utility-Privacy Trade-off

ICML 2025poster

To defend against privacy leakage of user data, differential privacy is widely used in federated learning, but it is not free. The addition of noise randomly disrupts the semantic integrity of the model and this disturbance accumulates with increased communication rounds. In this paper, we introduce…

2025

DCIS: Efficient Length Extrapolation of LLMs via Divide-and-Conquer Scaling Factor Search

EMNLP 2025

Large language models (LLMs) based on the Transformer architecture usually have their context length limited due to the high training cost. Recent advancements extend the context window by adjusting the scaling factors of RoPE and fine-tuning. However, suboptimal initialization of these factors resu

2025

DPoser-X: Diffusion Model as Robust 3D Whole-body Human Pose Prior

ICCV 2025poster

We present DPoser-X, a diffusion-based prior model for 3D whole-body human poses. Building a versatile and robust full-body human pose prior remains challenging due to the inherent complexity of articulated human poses and the scarcity of high-quality whole-body pose datasets. To address these limit…

Cited by 0SourcePDFScholar
2025

Disco4D: Disentangled 4D Human Generation and Animation from a Single Image

CVPR 2025poster

We present Disco4D, a novel Gaussian Splatting framework for 4D human generation and animation from a single image. Different from existing methods, Disco4D distinctively disentangles clothings (with Gaussian models) from the human body (with SMPL-X model), significantly enhancing the generation det…

2025

Discrete Distribution Networks

ICLR 2025poster

We introduce a novel generative model, the Discrete Distribution Networks (DDN), that approximates data distribution using hierarchical discrete distributions. We posit that since the features within a network inherently capture distributional information, enabling the network to generate multiple s…

2025

Divide-and-Conquer Variational Bayesian Inference for Multi-task Learning of High-resolution SAR Imagery

ICASSP 2025accepted

Conventional statistical-driven synthetic aperture radar (SAR) imaging algorithms can only encode a single and/or static prior, leading to that limited features can be accessed quantitatively. To this end, a novel multi-task learning framework is proposed by devising a divide-and-conquer variational…

Cited by 0SourceScholar
2025

Don't Shake the Wheel: Momentum-Aware Planning in End-to-End Autonomous Driving

CVPR 2025poster

End-to-end autonomous driving frameworks enable seamless integration of perception and planning but often rely on one-shot trajectory prediction, which may lead to unstable control and vulnerability to occlusions in single-frame perception. To address this, we propose the Momentum-Aware Driving (Mom…

2025

EgoLife: Towards Egocentric Life Assistant

CVPR 2025poster

We introduce EgoLife, a project to develop an egocentric life assistant that accompanies and enhances personal efficiency through AI-powered wearable glasses. To lay the foundation for this assistant, we conducted a comprehensive data collection study where six participants lived together for one we…

2025

Enhancing Visual Forced Alignment with Local Context-Aware Feature Extraction and Multi-Task Learning

ICASSP 2025accepted

This paper introduces a novel approach to Visual Forced Alignment (VFA), aiming to accurately synchronize utterances with corresponding lip movements, without relying on audio cues. We propose a novel VFA approach that integrates a local context-aware feature extractor and employs multitask learning…

Cited by 0SourceScholar
2025

GauSS-MI: Gaussian Splatting Shannon Mutual Information for Active 3D Reconstruction

RSS 2025poster

This research tackles the challenge of real-time active view selection and uncertainty quantification on visual quality for active 3D reconstruction. Visual quality is a critical aspect of 3D reconstruction. Recent advancements such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) h…

Cited by 0PDFcodeScholar
2025

MeshAnything: Artist-Created Mesh Generation with Autoregressive Transformers

ICLR 2025poster

Recently, 3D assets created via reconstruction and generation have matched the quality of manually crafted assets, highlighting their potential for replacement. However, this potential is largely unrealized because these assets always need to be converted to meshes for 3D industry applications, and…

2025

Praetor: A Fine-Grained Generative LLM Evaluator with Instance-Level Customizable Evaluation Criteria

ACL 2025long

With the increasing capability of large language models (LLMs), LLM-as-a-judge has emerged as a new evaluation paradigm. Compared with traditional automatic and manual evaluation, LLM evaluators exhibit better interpretability and efficiency. Despite this, existing LLM evaluators suffer from limited…

2025

SIMS: Simulating Stylized Human-Scene Interactions with Retrieval-Augmented Script Generation

ICCV 2025poster

Simulating stylized human-scene interactions (HSI) in physical environments is a challenging yet fascinating task. Prior works emphasize long-term execution but fall short in achieving both diverse style and physical plausibility. To tackle this challenge, we introduce a novel hierarchical framework…

Cited by 0SourcePDFScholar
2025

SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters

CVPR 2025poster

Human beings are social animals. How to equip 3D autonomous characters with similar social intelligence that can perceive, understand and interact with humans remains an open yet foundamental problem. In this paper, we introduce SOLAMI, the first end-to-end Social vision-Language-Action (VLA) Modeli…

Cited by 2SourcePDFScholar
2025

StarGen: A Spatiotemporal Autoregression Framework with Video Diffusion Model for Scalable and Controllable Scene Generation

CVPR 2025poster

Recent advances in large reconstruction and generative models have significantly improved scene reconstruction and novel view generation. However, due to compute limitations, each inference with these large models is confined to a small area, making long-range consistent scene generation challenging…

Cited by 1SourcePDFScholar
2025

TokensGen: Harnessing Condensed Tokens for Long Video Generation

ICCV 2025poster

Generating consistent long videos is a complex challenge: while diffusion-based generative models generate visually impressive short clips, extending them to longer durations often leads to memory bottlenecks and long-term inconsistency. In this paper, we propose TokensGen, a novel two-stage framewo…

Cited by 0SourcePDFScholar
2025

Towards a Pairwise Ranking Model with Orderliness and Monotonicity for Label Enhancement

NeurIPS 2025spotlight

Label distribution in recent years has been applied in a diverse array of complex decision-making tasks. To address the availability of label distributions, label enhancement has been established as an effective learning paradigm that aims to automatically infer label distributions from readily avai…

Cited by 0SourceScholar
2025

TraF-Align: Trajectory-aware Feature Alignment for Asynchronous Multi-agent Perception

CVPR 2025poster

Cooperative perception presents significant potential for enhancing the sensing capabilities of individual vehicles, however, inter-agent latency remains a critical challenge. Latencies cause misalignments in both spatial and semantic features, complicating the fusion of real-time observations from…

2025

Trajectory attention for fine-grained video motion control

ICLR 2025poster

Recent advancements in video generation have been greatly driven by video diffusion models, with camera motion control emerging as a crucial challenge in creating view-customized visual content. This paper introduces trajectory attention, a novel approach that performs attention along available pixe…

Cited by 0SourcePDFScholar
2025

ULRVT II: A Novel Upper Limb Rehabilitation Robot with Joint Synergy Control and Evaluation for Virtual Training*

IROS 2025

Global population aging has led to a sharp increase in patients of upper limb motor dysfunction. Robot assisted virtual training, as a novel solution, can offer safe and precise assistance for upper limb rehabilitation. However, it remains a critical challenge to compensate virtual interaction force

Cited by 0SourceScholar
2025

V2X-Radar: A Multi-modal Dataset with 4D Radar for Cooperative Perception

NeurIPS 2025spotlight

Modern autonomous vehicle perception systems often struggle with occlusions and limited perception range. Previous studies have demonstrated the effectiveness of cooperative perception in extending the perception range and overcoming occlusions, thereby enhancing the safety of autonomous driving. In…

Cited by 0SourceScholar
2025

sEMG-Based Continues Motion Prediction of Shoulder exoskeleton Control Using the VGANet Model

IROS 2025

Wearable exoskeleton robots play a crucial role in promoting upper limb function recovery. To enhance human-robot interaction and achieve precise control, continuous prediction of limb joint angles is required. This paper proposes a decoupled network model (VGANet) based on Variable Graph Convolutio

Cited by 0SourceScholar
2024

AiOS: All-in-One-Stage Expressive Human Pose and Shape Estimation

CVPR 2024poster

Expressive human pose and shape estimation (a.k.a. 3D whole-body mesh recovery) involves the human body hand and expression estimation. Most existing methods have tackled this task in a two-stage manner first detecting the human body part with an off-the-shelf detection model and then inferring the…

2024

AttriHuman-3D: Editable 3D Human Avatar Generation with Attribute Decomposition and Indexing

CVPR 2024poster

Editable 3D-aware generation which supports user-interacted editing has witnessed rapid development recently. However existing editable 3D GANs either fail to achieve high-accuracy local editing or suffer from huge computational costs. We propose AttriHuman-3D an editable 3D human generation model w…

Cited by 9SourcePDFScholar
2024

Contrastive Deep Nonnegative Matrix Factorization For Community Detection

ICASSP 2024accepted

Recently, nonnegative matrix factorization (NMF) has been widely adopted for community detection, because of its better interpretability. However, the existing NMF-based methods have the following three problems: 1) they directly transform the original network into community membership space, so it…

Cited by 0SourceScholar
2024

Differentiable Convex Polyhedra Optimization from Multi-view Images

ECCV 2024poster

"This paper presents a novel approach for the differentiable rendering of convex polyhedra, addressing the limitations of recent methods that rely on implicit field supervision. Our technique introduces a strategy that combines non-differentiable computation of hyperplane intersection through dualit…

2024

Digital Life Project: Autonomous 3D Characters with Social Intelligence

CVPR 2024poster

In this work we present Digital Life Project a framework utilizing language as the universal medium to build autonomous 3D characters who are capable of engaging in social interactions and expressing with articulated body motions thereby simulating life in a digital environment. Our framework compri…

Cited by 30SourcePDFScholar
2024

Duolando: Follower GPT with Off-Policy Reinforcement Learning for Dance Accompaniment

ICLR 2024poster

We introduce a novel task within the field of human motion generation, termed dance accompaniment, which necessitates the generation of responsive movements from a dance partner, the "follower", synchronized with the lead dancer’s movements and the underlying musical rhythm. Unlike existing solo or…

Cited by 19SourcePDFScholar
2024

Efficient Planar Fabric Repositioning: Deformation-Aware RRT* for Non-Prehensile Fabric Manipulation

RA-L 2024

Fabrics present significant challenges to robotic manipulation due to their complex dynamics and infinite degrees of freedom. This letter proposes a non-prehensile approach to aligning a fabric cut piece to a specified target pose, which is a common step for many garment manufacturing tasks. Compare

Cited by 4SourceScholar
2024

Fspen: an Ultra-Lightweight Network for Real Time Speech Enahncment

ICASSP 2024accepted

Deep learning-based speech enhancement methods have shown promising result in recent years. However, in practical applications, the model size and computational complexity are important factors that limit their use in end-products. Therefore, in products that require real-time speech enhancement wit…

Cited by 25SourceScholar
2024

FuxiTranyu: A Multilingual Large Language Model Trained with Balanced Data

EMNLP 2024industry

Large language models (LLMs) have demonstrated prowess in a wide range of tasks. However, many LLMs exhibit significant performance discrepancies between high- and low-resource languages. To mitigate this challenge, we present FuxiTranyu, an open-source multilingual LLM, which is designed to satisfy…

2024

GaussianEditor: Swift and Controllable 3D Editing with Gaussian Splatting

CVPR 2024poster

3D editing plays a crucial role in many areas such as gaming and virtual reality. Traditional 3D editing methods which rely on representations like meshes and point clouds often fall short in realistically depicting complex scenes. On the other hand methods based on implicit 3D representations like…

2024

GraphBEV: Towards Robust BEV Feature Alignment for Multi-Modal 3D Object Detection

ECCV 2024poster

"Integrating LiDAR and camera information into Bird’s-Eye-View (BEV) representation has emerged as a crucial aspect of 3D object detection in autonomous driving. However, existing methods are susceptible to the inaccurate calibration relationship between LiDAR and the camera sensor. Such inaccuracie…

2024

IT3D: Improved Text-to-3D Generation with Explicit View Synthesis

AAAI 2024technical

Recent strides in Text-to-3D techniques have been propelled by distilling knowledge from powerful large text-to-image diffusion models (LDMs). Nonetheless, existing Text-to-3D approaches often grapple with challenges such as over-saturation, inadequate detailing, and unrealistic outputs. This study…

2024

Language-Augmented Symbolic Planner for Open-World Task Planning

RSS 2024poster

Enabling robotic agents to perform complex long-horizon tasks has been a long-standing goal in robotics and artificial intelligence (AI). Despite the potential shown by large language models (LLMs), their planning capabilities remain limited to short-horizon tasks and they are unable to replace the…

2024

Large Motion Model for Unified Multi-Modal Motion Generation

ECCV 2024poster

"Human motion generation, a cornerstone technique in animation and video production, has widespread applications in various tasks like text-to-motion and music-to-dance. Previous works focus on developing specialist models tailored for each task without scalability. In this work, we present Large Mo…

Cited by 27SourcePDFScholar
2024

Learning Dense Correspondence for NeRF-Based Face Reenactment

AAAI 2024technical

Face reenactment is challenging due to the need to establish dense correspondence between various face representations for motion transfer. Recent studies have utilized Neural Radiance Field (NeRF) as fundamental representation, which further enhanced the performance of multi-view face reenactment i…

Cited by 11SourcePDFScholar
2024

Let All Be Whitened: Multi-Teacher Distillation for Efficient Visual Retrieval

AAAI 2024technical

Visual retrieval aims to search for the most relevant visual items, e.g., images and videos, from a candidate gallery with a given query item. Accuracy and efficiency are two competing objectives in retrieval tasks. Instead of crafting a new method pursuing further improvement on accuracy, in this p…

2024

RoboFusion: Towards Robust Multi-Modal 3D Object Detection via SAM

IJCAI 2024poster

Multi-modal 3D object detectors are dedicated to exploring secure and reliable perception systems for autonomous driving (AD). Although achieving state-of-the-art (SOTA) performance on clean benchmark datasets, they tend to overlook the complexity and harsh conditions of real-world environments. Wit…

2024

Speaker-Adaptive Lipreading Via Spatio-Temporal Information Learning

ICASSP 2024accepted

Lipreading has been rapidly developed recently with the help of large-scale datasets and large models. Despite the significant progress made, the performance of lipreading models still falls short when dealing with unseen speakers. Therefore, it is necessary to utilize the speaker’s videos for fine-…

Cited by 0SourceScholar
2024

ZeRO++: Extremely Efficient Collective Communication for Large Model Training

ICLR 2024poster

Zero Redundancy Optimizer (ZeRO) has been used to train a wide range of large language models on massive GPU clusters due to its ease of use, efficiency, and good scalability. However, when training on low-bandwidth clusters, and/or when small batch size per GPU is used, ZeRO’s effective throughput…

Cited by 9SourcePDFScholar
2023

BEVHeight: A Robust Framework for Vision-Based Roadside 3D Object Detection

CVPR 2023poster

While most recent autonomous driving system focuses on developing perception methods on ego-vehicle sensors, people tend to overlook an alternative approach to leverage intelligent roadside cameras to extend the perception ability beyond the visual range. We discover that the state-of-the-art vision…

2023

Boundary-Aware Backward-Compatible Representation via Adversarial Learning in Image Retrieval

CVPR 2023poster

Image retrieval plays an important role in the Internet world. Usually, the core parts of mainstream visual retrieval systems include an online service of the embedding model and a large-scale vector database. For traditional model upgrades, the old model will not be replaced by the new one until th…

2023

ConKI: Contrastive Knowledge Injection for Multimodal Sentiment Analysis

ACL 2023findings

Multimodal Sentiment Analysis leverages multimodal signals to detect the sentiment of a speaker. Previous approaches concentrate on performing multimodal fusion and representation learning based on general knowledge obtained from pretrained models, which neglects the effect of domain-specific knowle…

2023

DNA-Rendering: A Diverse Neural Actor Repository for High-Fidelity Human-Centric Rendering

ICCV 2023poster

Realistic human-centric rendering plays a key role in both computer vision and computer graphics. Rapid progress has been made in the algorithm aspect over the years, yet existing human-centric rendering datasets and benchmarks are rather impoverished in terms of diversity (e.g., outfit's fabric/mat…

Cited by 61PDFcodeScholar
2023

FineMoGen: Fine-Grained Spatio-Temporal Motion Generation and Editing

NeurIPS 2023poster

Text-driven motion generation has achieved substantial progress with the emergence of diffusion models. However, existing methods still struggle to generate complex motion sequences that correspond to fine-grained descriptions, depicting detailed and accurate spatio-temporal actions.This lack of fin…

2023

GraphAlign: Enhancing Accurate Feature Alignment by Graph matching for Multi-Modal 3D Object Detection

ICCV 2023poster

LiDAR and cameras are complementary sensors for 3D object detection in autonomous driving. However, it is challenging to explore the unnatural interaction between point clouds and images, and the critical factor is how to conduct feature alignment of heterogeneous modalities. Currently, many methods…

Cited by 45PDFScholar
2023

Hierarchical Temporal Transformer for 3D Hand Pose Estimation and Action Recognition From Egocentric RGB Videos

CVPR 2023poster

Understanding dynamic hand motions and actions from egocentric RGB videos is a fundamental yet challenging task due to self-occlusion and ambiguity. To address occlusion and ambiguity, we develop a transformer-based framework to exploit temporal information for robust estimation. Noticing the differ…

2023

OmniObject3D: Large-Vocabulary 3D Object Dataset for Realistic Perception, Reconstruction and Generation

CVPR 2023poster

Recent advances in modeling 3D objects mostly rely on synthetic datasets due to the lack of large-scale real-scanned 3D databases. To facilitate the development of 3D perception, reconstruction, and generation in the real world, we propose OmniObject3D, a large vocabulary 3D object dataset with mass…

Cited by 214SourcePDFScholar
2023

PrimDiffusion: Volumetric Primitives Diffusion for 3D Human Generation

NeurIPS 2023poster

We present PrimDiffusion, the first diffusion-based framework for 3D human generation. Devising diffusion models for 3D human generation is difficult due to the intensive computational cost of 3D representations and the articulated topology of 3D humans. To tackle these challenges, our key insight i…

2023

ReMoDiffuse: Retrieval-Augmented Motion Diffusion Model

ICCV 2023poster

3D human motion generation is crucial for creative industry. Recent advances rely on generative models with domain knowledge for text-driven motion generation, leading to substantial progress in capturing common motions. However, the performance on more diverse motions remains unsatisfactory. In thi…

Cited by 167PDFcodeScholar
2023

RenderMe-360: A Large Digital Asset Library and Benchmarks Towards High-fidelity Head Avatars

NeurIPS 2023poster

Synthesizing high-fidelity head avatars is a central problem for computer vision and graphics. While head avatar synthesis algorithms have advanced rapidly, the best ones still face great obstacles in real-world scenarios. One of the vital causes is the inadequate datasets -- 1) current public data…

2023

SHERF: Generalizable Human NeRF from a Single Image

ICCV 2023poster

Existing Human NeRF methods for reconstructing 3D humans typically rely on multiple 2D images from multi-view cameras or monocular videos captured from fixed camera views. However, in real-world scenarios, human images are often captured from random camera angles, presenting challenges for high-qual…

Cited by 82PDFcodeScholar
2023

SMPLer-X: Scaling Up Expressive Human Pose and Shape Estimation

NeurIPS 2023poster

Expressive human pose and shape estimation (EHPS) unifies body, hands, and face motion capture with numerous applications. Despite encouraging progress, current state-of-the-art methods still depend largely on a confined set of training datasets. In this work, we investigate scaling up EHPS towards…

2023

Search-Map-Search: A Frame Selection Paradigm for Action Recognition

CVPR 2023poster

Despite the success of deep learning in video understanding tasks, processing every frame in a video is computationally expensive and often unnecessary in real-time applications. Frame selection aims to extract the most informative and representative frames to help a model better understand video co…

2023

Surface Extraction from Neural Unsigned Distance Fields

ICCV 2023poster

We propose a method, named DualMesh-UDF, to extract a surface from unsigned distance functions (UDFs), encoded by neural networks, or neural UDFs. Neural UDFs are becoming increasingly popular for surface representation because of their versatility in presenting surfaces with arbitrary topologies, a…

Cited by 8PDFcodeScholar
2023

SynBody: Synthetic Dataset with Layered Human Models for 3D Human Perception and Modeling

ICCV 2023poster

Synthetic data has emerged as a promising source for 3D human research as it offers low-cost access to large-scale human datasets. To advance the diversity and annotation quality of human models, we introduce a new synthetic dataset, SynBody, with three appealing features: 1) a clothed parametric hu…

Cited by 48PDFcodeScholar
2023

Target Speaker Extraction with Ultra-Short Reference Speech by VE-VE Framework

ICASSP 2023accepted

The goal of target speaker extraction (TSE) is to extract the target speaker's voice from the mixture speech of multiple speakers. It needs to enroll the speech of the target speaker in advance as a reference. However, in practical applications, too long reference speech during enrollment will decre…

Cited by 0SourceScholar
2023

Towards Robust and Expressive Whole-body Human Pose and Shape Estimation

NeurIPS 2023poster

Whole-body pose and shape estimation aims to jointly predict different behaviors (e.g., pose, hand gesture, facial expression) of the entire human body from a monocular image. Existing methods often exhibit suboptimal performance due to the complexity of in-the-wild scenarios. We argue that the pred…

2023

TransVCL: Attention-Enhanced Video Copy Localization Network with Flexible Supervision

AAAI 2023technical

Video copy localization aims to precisely localize all the copied segments within a pair of untrimmed videos in video retrieval applications. Previous methods typically start from frame-to-frame similarity matrix generated by cosine similarity between frame-level features of the input video pair, an…

2023

Zolly: Zoom Focal Length Correctly for Perspective-Distorted Human Mesh Reconstruction

ICCV 2023oral

As it is hard to calibrate single-view RGB images in the wild, existing 3D human mesh reconstruction (3DHMR) methods either use a constant large focal length or estimate one based on the background environment context, which can not tackle the problem of the torso, limb, hand or face distortion caus…

Cited by 39PDFcodeScholar
2022

A Large-Scale Comprehensive Dataset and Copy-Overlap Aware Evaluation Protocol for Segment-Level Video Copy Detection

CVPR 2022poster

In this paper, we introduce VCSL (Video Copy Segment Localization), a new comprehensive segment-level annotated video copy dataset. Compared with existing copy detection datasets restricted by either video-level annotation or small-scale, VCSL not only has two orders of magnitude more segment-level…

Cited by 18PDFcodeScholar
2022

Benchmarking and Analyzing 3D Human Pose and Shape Estimation Beyond Algorithms

NeurIPS 2022accept

3D human pose and shape estimation (a.k.a. ``human mesh recovery'') has achieved substantial progress. Researchers mainly focus on the development of novel algorithms, while less attention has been paid to other critical factors involved. This could lead to less optimal baselines, hindering the fair…

2022

DISP6D: Disentangled Implicit Shape and Pose Learning for Scalable 6D Pose Estimation

ECCV 2022poster

"Scalable 6D pose estimation for rigid objects from RGB images aims at handling multiple objects and generalizing to novel objects. Building on a well-known auto-encoding framework to cope with object symmetry and the lack of labeled training data, we achieve scalability by disentangling the latent…

2022

DeciWatch: A Simple Baseline for 10× Efficient 2D and 3D Pose Estimation

ECCV 2022poster

"This paper proposes a simple baseline framework for video-based 2D/3D human pose estimation that can achieve 10 times efficiency improvement over existing works without any performance degradation, named DeciWatch. Unlike current solutions that estimate each frame in a video, DeciWatch introduces a…

2022

HuMMan: Multi-modal 4D Human Dataset for Versatile Sensing and Modeling

ECCV 2022poster

"4D human sensing and modeling are fundamental tasks in vision and graphics with numerous applications. With the advances of new sensors and algorithms, there is an increasing demand for more versatile datasets. In this work, we contribute HuMMan, a large-scale multi-modal 4D human dataset with 1000…

Cited by 125SourcePDFScholar
2022

Learn to Predict How Humans Manipulate Large-Sized Objects From Interactive Motions

RA-L 2022

Understanding human intentions during interactions has been a long-lasting theme, that has applications in human-robot interaction, virtual reality and surveillance. In this study, we focus on full-body human interactions with large-sized daily objects and aim to predict the future states of objects

Cited by 35SourceScholar
2022

Not All Models Are Equal: Predicting Model Transferability in a Self-Challenging Fisher Space

ECCV 2022poster

"This paper addresses an important problem of ranking the pre-trained deep neural networks and screening the most transferable ones for downstream tasks. It is challenging because the ground-truth model ranking for each task can only be generated by fine-tuning the pre-trained models on the target d…

2022

Practical Stereo Matching via Cascaded Recurrent Network With Adaptive Correlation

CVPR 2022oral

With the advent of convolutional neural networks, stereo matching algorithms have recently gained tremendous progress. However, it remains a great challenge to accurately extract disparities from real-world image pairs taken by consumer-level devices like smartphones, due to practical complicating f…

Cited by 318PDFcodeScholar
2022

Skim: Skipping Memory Lstm for Low-Latency Real-Time Continuous Speech Separation

ICASSP 2022accepted

Continuous speech separation for meeting pre-processing has recently become a focused research topic. Compared to the data in utterance-level speech separation, the meeting-style audio stream lasts longer, has an uncertain number of speakers. We adopt the time-domain speech separation method and the…

Cited by 0SourceScholar
2022

SmoothNet: A Plug-and-Play Network for Refining Human Poses in Videos

ECCV 2022poster

"When analyzing human motion videos, the output jitters from existing pose estimators are highly-unbalanced with varied estimation errors across frames. Most frames in a video are relatively easy to estimate and only suffer from slight jitters. In contrast, for rarely seen or occluded actions, the e…

2022

Visual-tactile Sensing for Real-time Liquid Volume Estimation in Grasping

IROS 2022poster

We propose a deep visuo-tactile model for real-time estimation of the liquid inside a deformable container in a proprioceptive way. We fuse two sensory modalities, i.e., the raw visual inputs from the RGB camera and the tactile cues from our specific tactile sensor without any extra sensor calibrati…

Cited by 16SourceScholar
2021

Geometry Uncertainty Projection Network for Monocular 3D Object Detection

ICCV 2021poster

Monocular 3D object detection has received increasing attention due to the wide application in autonomous driving. Existing works mainly focus on introducing geometry projection to predict depth priors for each object. Despite their impressive progress, these methods neglect the geometry leverage ef…

Cited by 263PDFcodeScholar
2021

Learning Skeletal Graph Neural Networks for Hard 3D Pose Estimation

ICCV 2021poster

Various deep learning techniques have been proposed to solve the single-view 2D-to-3D pose estimation problem. While the average prediction accuracy has been improved significantly over the years, the performance on hard poses with depth ambiguity, self-occlusion, and complex or rare poses is still…

Cited by 163PDFScholar
2021

Model-Based Robust Tracking Control Without Observers for Soft Bending Actuators

RA-L 2021

It is of great importance to improve the performance of soft bending actuators with simpler control laws. Considering that the velocity information of soft actuators cannot be measured by conventional sensors due to a contradiction between the soft body and the rigid state, this letter addresses a r

Cited by 20SourceScholar
2021

Refer-It-in-RGBD: A Bottom-Up Approach for 3D Visual Grounding in RGBD Images

CVPR 2021poster

Grounding referring expressions in RGBD image has been an emerging field. We present a novel task of 3D visual grounding in single-view RGBD image where the referred objects are often only partially scanned due to occlusion. In contrast to previous works that directly generate object proposals for g…

Cited by 43PDFScholar
2021

SimiGrad: Fine-Grained Adaptive Batching for Large Scale Training using Gradient Similarity Measurement

NeurIPS 2021poster

Large scale training requires massive parallelism to finish the training within a reasonable amount of time. To support massive parallelism, large batch training is the key enabler but often at the cost of generalization performance. Existing works explore adaptive batching or hand-tuned static larg…

2020

Caption-Supervised Face Recognition: Training a State-of-the-Art Face Model without Manual Annotation

ECCV 2020poster

The advances over the past several years have pushed the performance of face recognition to an amazing level. This great success, to a large extent, is built on top of millions of annotated samples. However, as we endeavor to take the performance to the next level, the reliance on annotated data bec…

Cited by 28SourcePDFScholar
2020

Edge Enhanced Implicit Orientation Learning With Geometric Prior for 6D Pose Estimation

RA-L 2020

Estimating 6D poses of rigid objects from RGB images is an important but challenging task. This is especially true for textureless objects with strong symmetry, since they have only sparse visual features to be leveraged for the task and their symmetry leads to pose ambiguity. The implicit encoding

Cited by 34SourcecodeScholar
2020

Learn to Propagate Reliably on Noisy Affinity Graphs

ECCV 2020poster

Recent works have shown that exploiting unlabeled data through label propagation can substantially reduce the labeling cost, which has been a critical issue in developing visual recognition models. Yet, how to propagate labels reliably, especially on a dataset with unknown outliers, remains an open…

Cited by 16SourcePDFScholar
2020

Learning to Cluster Faces via Confidence and Connectivity Estimation

CVPR 2020poster

Face clustering is an essential tool for exploiting the unlabeled face data, and has a wide range of applications including face annotation and retrieval. Recent works show that supervised clustering can result in noticeable performance gain. However, they usually involve heuristic steps and require…

Cited by 116PDFcodeScholar
2020

Mapping in a Cycle: Sinkhorn Regularized Unsupervised Learning for Point Cloud Shapes

ECCV 2020poster

We propose an unsupervised learning framework with the pretext task of finding dense correspondences between point cloud shapes from the same category based on the cycle-consistency formulation. In order to learn discriminative pointwise features from point cloud data, we incorporate in the formulat…

2020

SemanticAdv: Generating Adversarial Examples via Attribute-conditioned Image Editing

ECCV 2020poster

Deep neural networks (DNNs) have achieved great successes in various vision applications due to their strong expressive power. However, recent studies have shown that DNNs are vulnerable to adversarial examples which are manipulated instances targeting to mislead DNNs to make incorrect predictions.…

Cited by 203SourcePDFScholar
2019

Learning to Cluster Faces on an Affinity Graph

CVPR 2019oral

Face recognition sees remarkable progress in recent years, and its performance has reached a very high level. Taking it to a next level requires substantially larger data, which would involve prohibitive annotation cost. Hence, exploiting unlabeled data becomes an appealing alternative. Recent works…

Cited by 161PDFcodeScholar
2015

Ground moving target imaging by synthetic aperture radar based on an unified framework of keystone transformation

ICASSP 2015accepted

This paper presents a new SAR ground moving target imaging (GMTIm) algorithm based on an unified framework of Keystone transformation (KT). To combat the inherent range-azimuth coupling, an tandem two-step strategy is designed, where the range decoupling is implemented by polar format algorithm (PFA…

Cited by 0SourceScholar
2015

Grounding English Commands to Reward Functions

RSS 2015poster

As intelligent robots become more prevalent, methods to make interaction with the robots more accessible are increasingly important. Communicating the tasks that a person wants the robot to carry out via natural language, and training the robot to ground the natural language through demonstration, a…