← Search

ye yuan

101 accepted papers

2026

CARI4D: Category Agnostic 4D Reconstruction of Human-Object Interaction

CVPR 2026

Accurate capture of human-object interaction from ubiquitous sensors like RGB cameras is important for applications in human understanding, gaming, and robot learning. However, inferring 4D interactions from a single RGB view is highly challenging due to the unknown object and human information, dep

Cited by 0SourcecodeScholar
2026

FlowMAP: Flow Matching for Generalizable Agent Planning

ICML 2026poster

Agent planning faces dynamic heterogeneity—nonstationary observations, dynamics, and objectives with sparse, delayed rewards—which dominant methods largely ignore, leading to poor generalization under environment shifts. We propose Flow-Matching for Agent Planning (FlowMAP), which formulates plannin…

Cited by 0SourceScholar
2026

LoC-Decomp: LLM Autoformalization via Logical Concept Decomposition and Iterative Feedback Correction

ICLR 2026poster

Autoformalization—the process of converting natural language mathematical statements into machine-verifiable formal code—plays a critical role in ensuring the reliability of mathematical reasoning generated by large language models (LLMs). Recent studies show that LLMs exhibit strong potential in au…

Cited by 0SourcecodeScholar
2026

Safe Multi-Agent Reinforcement Learning via Distributional Safety Critic and Maximum Entropy Optimization

AAAI 2026technical

Deploying multi-agent reinforcement learning (MARL) in safety-critical systems faces significant challenges due to insufficient agent exploration and inadequate safety constraint guarantees. Current approaches are constrained by two fundamental limitations: inefficient exploration leading to subopti

Cited by 0SourcePDFScholar
2026

Support-Proximity Augmented Diffusion Estimation for Offline Black-Box Optimization

ICML 2026poster

Offline black-box optimization aims to discover novel designs with high property scores using only a static dataset, a task fundamentally challenged by the out-of-distribution (OOD) extrapolation problem. Existing approaches typically bifurcate into inverse methods, which struggle with the ill-posed…

Cited by 0SourceScholar
2026

Training Diffusion Language Models for Black-Box Optimization

ICML 2026spotlight

We study offline black-box optimization (BBO), aiming to discover improved designs from an offline dataset of designs and labels, a problem common in robotics, DNA, and materials science with limited labeled samples. While recent work applies autoregressive LLMs to BBO by formatting tasks as natural…

Cited by 0SourceScholar
2026

Uncertainty-Constrained Trustworthiness for Graph Learning

ICML 2026poster

Graph learning has been increasingly deployed in critical and sensitive domains, raising pressing demands for trustworthiness-robustness, fairness, and beyond. However, these properties are often undermined by various perturbations, which induce distributional uncertainty and compromise the trustwor…

Cited by 0SourceScholar
2026

VIRAL: Visual Sim-to-Real at Scale for Humanoid Loco-Manipulation

CVPR 2026

A key barrier to the real-world deployment of humanoid robots is the lack of autonomous loco-manipulation skills. We introduce VIRAL, a visual sim-to-real framework that learns humanoid loco-manipulation entirely in simulation and deploys it zero-shot to real hardware. VIRAL follows a teacher-studen

Cited by 0SourcecodeScholar
2025

AdaHuman: Animatable Detailed 3D Human Generation with Compositional Multiview Diffusion

ICCV 2025poster

Existing methods for image-to-3D avatar generation struggle to produce highly detailed, animation-ready avatars suitable for real-world applications. We introduce AdaHuman, a novel framework that generates high-fidelity animatable 3D avatars from a single in-the-wild image. AdaHuman incorporates two…

Cited by 0SourcePDFScholar
2025

BLADE: Single-view Body Mesh Estimation through Accurate Depth Estimation

CVPR 2025poster

Single-image human mesh recovery is a challenging task due to the ill-posed nature of simultaneous body shape, pose, and camera estimation. Existing estimators work well on images taken from afar, but they break down as the person moves close to the camera. Moreover, current methods fail to achieve…

Cited by 0SourcePDFScholar
2025

EmoRLTalk: Speech-Driven Emotional Facial Animation With Offline Reinforcement Learning

IROS 2025

In recent years, significant breakthroughs have been made in audio-guided 3D facial animation. However, existing methods mainly focus on lip shape and audio consistency and still face key challenges to achieve alignment between facial emotions and speech emotions. To overcome this limitation, we int

Cited by 0SourceScholar
2025

Fira: Can We Achieve Full-rank Training of LLMs Under Low-rank Constraint?

NeurIPS 2025poster

Low-rank training has emerged as a promising approach for reducing memory usage in training Large Language Models (LLMs). Previous methods either rely on decomposing weight matrices (e.g., LoRA), or seek to decompose gradient matrices (e.g., GaLore) to ensure reduced memory consumption. However, bot…

Cited by 0SourcecodeScholar
2025

GENMO: A GENeralist Model for Human MOtion

ICCV 2025poster

Human motion modeling traditionally separates motion generation and estimation into distinct tasks with specialized models. Motion generation models focus on creating diverse, realistic motions from inputs like text, audio, or keyframes, while motion estimation models aim to reconstruct accurate mot…

Cited by 0SourcePDFScholar
2025

Generalization Bounds for Kolmogorov-Arnold Networks (KANs) and Enhanced KANs with Lower Lipschitz Complexity

NeurIPS 2025poster

Kolmogorov-Arnold Networks (KANs) have demonstrated remarkable expressive capacity and predictive power in symbolic learning. However, existing generalization errors of KANs primarily focus on approximation errors while neglecting estimation errors, leading to a suboptimal bias-variance trade-off an…

Cited by 0SourceScholar
2025

GeoMan: Temporally Consistent Human Geometry Estimation using Image-to-Video Diffusion

ICCV 2025poster

Estimating accurate and temporally consistent 3D human geometry from videos is a challenging problem in computer vision. Existing methods, primarily optimized for single images, often suffer from temporal inconsistencies and fail to capture fine-grained dynamic details. To address these limitations,…

Cited by 0SourcePDFScholar
2025

Handling Label Noise via Instance-Level Difficulty Modeling and Dynamic Optimization

NeurIPS 2025poster

Recent studies indicate that deep neural networks degrade in generalization performance under noisy supervision. Existing methods focus on isolating clean subsets or correcting noisy labels, facing limitations such as high computational costs, heavy hyperparameter tuning process, and coarse-grained…

Cited by 0SourcecodeScholar
2025

Harnessing Diversity for Important Data Selection in Pretraining Large Language Models

ICLR 2025spotlight

Data selection is of great significance in pretraining large language models, given the variation in quality within the large-scale available training corpora. To achieve this, researchers are currently investigating the use of data influence to measure the importance of data instances, $i.e.,$ a…

Cited by 8SourcePDFScholar
2025

MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation

NAACL 2025long

Large Multimodal Models (LMMs) exhibit impressive cross-modal understanding and reasoning abilities, often assessed through multiple-choice questions (MCQs) that include an image, a question, and several options. However, many benchmarks used for such evaluations suffer from systematic biases. Remar…

2025

Real-Time Neural Denoising with Render-Aware Knowledge Distillation

AAAI 2025technical

Real-time Monte Carlo (MC) ray tracing with low sampling rates demands a denoising algorithm that adeptly balances the trade-off between quality and efficiency. Previous works have paid much attention on designing delicate denoising architecture while ignoring model compression. In this work, we pre…

Cited by 0SourcePDFScholar
2025

Robust Dwell Time Allocation for Multiple Ballistic Reentry Target Tracking in Phased Array Radar

ICASSP 2025accepted

Phased array radar (PAR) is shown to provide an enhanced performance for target tracking due to its beam agility and ability for time resource allocation. Existing algorithms for PAR resource allocation often consider standard and simplified dynamic models for target motion, which are not suitable f…

Cited by 0SourceScholar
2025

Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification

ACL 2025long

Chain-of-Thought (CoT) prompting has become the de facto method to elicit reasoning capabilities from large language models (LLMs). However, to mitigate hallucinations in CoT that are notoriously difficult to detect, current methods such as process reward models (PRMs) or self-consistency operate as…

2025

SimAvatar: Simulation-Ready Avatars with Layered Hair and Clothing

CVPR 2025poster

We introduce SimAvatar, a framework designed to generate simulation-ready clothed 3D human avatars from a text prompt. Current text-driven human avatar generation methods either model hair, clothing and human body using a unified geometry or produce hair and garments that are not easily adaptable fo…

Cited by 1SourcePDFScholar
2025

Towards Multi-Table Learning: A Novel Paradigm for Complementarity Quantification and Integration

NeurIPS 2025spotlight

Multi-table data integrate various entities and attributes, with potential interconnections between them. However, existing tabular learning methods often struggle to describe and leverage the underlying complementarity across distinct tables. To address this limitation, we propose the first unified…

Cited by 0SourceScholar
2025

VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning

NeurIPS 2025poster

Reinforcement fine-tuning (RFT) has shown great promise in achieving humanlevel reasoning capabilities of Large Language Models (LLMs), and has recently been extended to MLLMs. Nevertheless, reasoning about videos, which is a fundamental aspect of human intelligence, remains a persistent challenge d…

Cited by 0SourcecodeScholar
2024

A robust inlier identification algorithm for point cloud registration via $\mathbf{\ell_0}$-minimization

NeurIPS 2024poster

Correspondences in point cloud registration are prone to outliers, significantly reducing registration accuracy and highlighting the need for precise inlier identification. In this paper, we propose a robust inlier identification algorithm for point cloud registration by reformulating the convention…

Cited by 1SourcePDFScholar
2024

An Iterative Min-Min Optimization Method for Sparse Bayesian Learning

ICML 2024poster

As a well-known machine learning algorithm, sparse Bayesian learning (SBL) can find sparse representations in linearly probabilistic models by imposing a sparsity-promoting prior on model coefficients. However, classical SBL algorithms lack the essential theoretical guarantees of global convergence.…

Cited by 1SourcePDFScholar
2024

Anti-Deception Jamming Power Optimization Strategy for Multi-Target Tracking Tasks in Multi-Radar Systems

ICASSP 2024accepted

In this paper, a power optimization (PO) strategy is proposed to combat deception jamming in multi-radar systems (MRSs) performing multi-target tracking (MTT). As a crucial parameter for distinguishing between physical and false targets in MRS under deception jamming, we propose integrating the dece…

Cited by 0SourceScholar
2024

COIN: Control-Inpainting Diffusion Prior for Human and Camera Motion Estimation

ECCV 2024poster

"Estimating global human motion from moving cameras is challenging due to the entanglement of human and camera motions. To mitigate the ambiguity, existing methods leverage learned human motion priors, which however often result in oversmoothed motions with misaligned 2D projections. To tackle this…

2024

EquiPocket: an E(3)-Equivariant Geometric Graph Neural Network for Ligand Binding Site Prediction

ICML 2024oral

Predicting the binding sites of target proteins plays a fundamental role in drug discovery. Most existing deep-learning methods consider a protein as a 3D image by spatially clustering its atoms into voxels and then feed the voxelized protein into a 3D CNN for prediction. However, the CNN-based meth…

2024

FDNet: Feature Decoupling Framework for Trajectory Prediction

IROS 2024poster

Trajectory prediction plays a significant role in autonomous driving, with current challenges primarily focused on capturing complex interactions in traffic scenes. Previous methods usually directly encode non-interactive and interactive information together, and then decode them for trajectory pred…

Cited by 1SourceScholar
2024

GAvatar: Animatable 3D Gaussian Avatars with Implicit Mesh Learning

CVPR 2024highlight

Gaussian splatting has emerged as a powerful 3D representation that harnesses the advantages of both explicit (mesh) and implicit (NeRF) 3D representations. In this paper we seek to leverage Gaussian splatting to generate realistic animatable avatars from textual descriptions addressing the limitati…

Cited by 41SourcePDFScholar
2024

Harmon: Whole-Body Motion Generation of Humanoid Robots from Language Descriptions

CoRL 2024poster

Humanoid robots, with their human-like embodiment, have the potential to integrate seamlessly into human environments. Critical to their coexistence and cooperation with humans is the ability to understand natural language communications and exhibit human-like behaviors. This work focuses on generat…

Cited by 9SourcecodeScholar
2024

Keypoint-based Progressive Chain-of-Thought Distillation for LLMs

ICML 2024poster

Chain-of-thought distillation is a powerful technique for transferring reasoning abilities from large language models (LLMs) to smaller student models. Previous methods typically require the student to mimic the step-by-step rationale produced by LLMs, often facing the following challenges: (i) Toke…

Cited by 2SourcePDFScholar
2024

LaKD: Length-agnostic Knowledge Distillation for Trajectory Prediction with Any Length Observations

NeurIPS 2024poster

Trajectory prediction is a crucial technology to help systems avoid traffic accidents, ensuring safe autonomous driving. Previous methods typically use a fixed-length and sufficiently long trajectory of an agent as observations to predict its future trajectory. However, in real-world scenarios, we o…

Cited by 1SourcePDFScholar
2024

Learning to Extract Structured Entities Using Language Models

EMNLP 2024main

Recent advances in machine learning have significantly impacted the field of information extraction, with Language Models (LMs) playing a pivotal role in extracting structured information from unstructured text. Prior works typically represent information extraction as triplet-centric and use classi…

2024

Measuring Vision-Language STEM Skills of Neural Models

ICLR 2024poster

We introduce a new challenge to test the STEM skills of neural models. The problems in the real world often require solutions, combining knowledge from STEM (science, technology, engineering, and math). Unlike existing datasets, our dataset requires the understanding of multimodal vision-language in…

2024

On the Road to Portability: Compressing End-to-End Motion Planner for Autonomous Driving

CVPR 2024poster

End-to-end motion planning models equipped with deep neural networks have shown great potential for enabling full autonomous driving. However the oversized neural networks render them impractical for deployment on resource-constrained systems which unavoidably requires more computational time and re…

2024

Online Learning Based Shape Control for a Soft Manipulator Based on Spatial Features Feedback

RA-L 2024

Although soft manipulators are endowed with compliance and flexibility, most control strategies focus on end-effector control and lack shape control ability. This letter aims to design a shape controller for the soft manipulator. Firstly, we establish a modified forward kinematics model (FKM) based

Cited by 5SourceScholar
2024

PACER+: On-Demand Pedestrian Animation Controller in Driving Scenarios

CVPR 2024poster

We address the challenge of content diversity and controllability in pedestrian simulation for driving scenarios. Recent pedestrian animation frameworks have a significant limitation wherein they primarily focus on either following trajectory or the content of the reference video consequently overlo…

Cited by 15SourcePDFScholar
2024

Preparing Lessons for Progressive Training on Language Models

AAAI 2024technical

The rapid progress of Transformers in artificial intelligence has come at the cost of increased resource consumption and greenhouse gas emissions due to growing model sizes. Prior work suggests using pretrained small models to improve training efficiency, but this approach may not be suitable for ne…

2023

BCDiff: Bidirectional Consistent Diffusion for Instantaneous Trajectory Prediction

NeurIPS 2023poster

The objective of pedestrian trajectory prediction is to estimate the future paths of pedestrians by leveraging historical observations, which plays a vital role in ensuring the safety of self-driving vehicles and navigation robots. Previous works usually rely on a sufficient amount of observation ti…

Cited by 28SourcePDFScholar
2023

DT-Solver: Automated Theorem Proving with Dynamic-Tree Sampling Guided by Proof-level Value Function

ACL 2023long

Recent advances in neural theorem-proving resort to large language models and tree searches. When proving a theorem, a language model advises single-step actions based on the current proving state and the tree search finds a sequence of correct steps using actions given by the language model. Howeve…

Cited by 35SourcePDFScholar
2023

Importance-aware Co-teaching for Offline Model-based Optimization

NeurIPS 2023poster

Offline model-based optimization aims to find a design that maximizes a property of interest using only an offline dataset, with applications in robot, protein, and molecule design, among others. A prevalent approach is gradient ascent, where a proxy model is trained on the offline dataset and then…

2023

Learning Human Dynamics in Autonomous Driving Scenarios

ICCV 2023poster

Simulation has emerged as an indispensable tool for scaling and accelerating the development of self-driving systems. A critical aspect of this is simulating realistic and diverse human behavior and intent. In this work, we propose a holistic framework for learning physically plausible human dynamic…

Cited by 23PDFScholar
2023

NeRFool: Uncovering the Vulnerability of Generalizable Neural Radiance Fields against Adversarial Perturbations

ICML 2023poster

Generalizable Neural Radiance Fields (GNeRF) are one of the most promising real-world solutions for novel view synthesis, thanks to their cross-scene generalization capability and thus the possibility of instant rendering on new scenes. While adversarial robustness is essential for real-world applic…

2023

RGB-Only Reconstruction of Tabletop Scenes for Collision-Free Manipulator Control

ICRA 2023poster

We present a system for collision-free control of a robot manipulator that uses only RGB views of the world. Perceptual input of a tabletop scene is provided by multiple images of an RGB camera (without depth) that is either handheld or mounted on the robot end effector. A NeRF-like process is used…

Cited by 14SourcecodeScholar
2023

Reusing Pretrained Models by Multi-linear Operators for Efficient Training

NeurIPS 2023poster

Training large models from scratch usually costs a substantial amount of resources. Towards this problem, recent studies such as bert2BERT and LiGO have reused small pretrained models to initialize a large model (termed the ``target model''), leading to a considerable acceleration in training. Despi…

Cited by 16SourcePDFScholar
2023

TREA: Tree-Structure Reasoning Schema for Conversational Recommendation

ACL 2023long

Conversational recommender systems (CRS) aim to timely trace the dynamic interests of users through dialogues and generate relevant responses for item recommendations. Recently, various external knowledge bases (especially knowledge graphs) are incorporated into CRS to enhance the understanding of c…

2023

TRIGO: Benchmarking Formal Mathematical Proof Reduction for Generative Language Models

EMNLP 2023long main

Automated theorem proving (ATP) has become an appealing domain for exploring the reasoning ability of the recent successful generative language models. However, current ATP benchmarks are mainly focus on symbolic inference, but rarely involve the understanding of complex number combination reasoni…

Cited by 0SourcecodeScholar
2023

Trace and Pace: Controllable Pedestrian Animation via Guided Trajectory Diffusion

CVPR 2023poster

We introduce a method for generating realistic pedestrian trajectories and full-body animations that can be controlled to meet user-defined goals. We draw on recent advances in guided diffusion modeling to achieve test-time controllability of trajectories, which is normally only associated with rule…

Cited by 118SourcePDFScholar
2023

Wheel Vision: Wheel-Terrain Interaction Measurement and Analysis Using a Sensorized Transparent Wheel on Deformable Terrains

RA-L 2023

The off-road locomotion of wheeled mobile robots (WMRs) over soft terrains can be quite challenging due to the complicated wheel-terrain interaction (WTI). To avoid unforeseen non-geometric hazards such as excessive sinkage or slippage, it is crucial to oversee these terrain-related uncertainties. H

Cited by 13SourceScholar
2023

Wheel-Terrain Contact Geometry Estimation and Interaction Analysis Using Aside-Wheel Camera Over Deformable Terrains

RA-L 2023

Wheeled mobile robots (WMRs) have been proven to be quite competitive and useful in outdoor missions. However, they may face serious sinkage or slippage on deformable terrains, and even get stuck or damaged, thereby causing mission failure. To mitigate these risks, it is essential to closely monitor

Cited by 10SourceScholar
2022

Design of a Soft Gripper With Improved Microfluidic Tactile Sensors for Classification of Deformable Objects

RA-L 2022

Tactile object recognition is vital for robotic handling systems; however, existing technologies that concentrate on tactile sensors with high modulus are not suitable for soft grippers to classify deformable objects. In this letter, we integrated an indenter layer into the traditional microfluidic

Cited by 19SourceScholar
2022

GLAMR: Global Occlusion-Aware Human Mesh Recovery With Dynamic Cameras

CVPR 2022oral

We present an approach for 3D global human mesh recovery from monocular videos recorded with dynamic cameras. Our approach is robust to severe and long-term occlusions and tracks human bodies even when they go outside the camera's field of view. To achieve this, we first propose a deep generative mo…

Cited by 136PDFcodeScholar
2022

PALT: Parameter-Lite Transfer of Language Models for Knowledge Graph Completion

EMNLP 2022finding

This paper presents a parameter-lite transfer learning approach of pretrained language models (LM) for knowledge graph (KG) completion. Instead of finetuning, which modifies all LM parameters, we only tune a few new parameters while keeping the original LM parameters fixed. We establish this via ref…

2022

Sen-Glove: A Lightweight Wearable Glove for Hand Assistance with Soft Joint Sensing

ICRA 2022poster

Perception and portability are critical issues for wearable gloves in hand assistive engineering. However, available wearable gloves either lack flexible sensing or are bulky. In this paper, we present a tendon-driven lightweight wearable glove with soft joint sensing, Sen-Glove. Sen-Glove is equipp…

Cited by 12SourceScholar
2022

Syntax-Aware Network for Handwritten Mathematical Expression Recognition

CVPR 2022poster

Handwritten mathematical expression recognition (HMER) is a challenging task that has many potential applications. Recent methods for HMER have achieved outstanding performance with an encoder-decoder architecture. However, these methods adhere to the paradigm that the prediction is made "from one c…

Cited by 98PDFScholar
2022

Transform2Act: Learning a Transform-and-Control Policy for Efficient Agent Design

ICLR 2022oral

An agent's functionality is largely determined by its design, i.e., skeletal structure and joint attributes (e.g., length, size, strength). However, finding the optimal agent design for a given function is extremely challenging since the problem is inherently combinatorial and the design space is pr…

2022

When Counting Meets HMER: Counting-Aware Network for Handwritten Mathematical Expression Recognition

ECCV 2022poster

"Recently, most handwritten mathematical expression recognition (HMER) methods adopt the encoder-decoder networks, which directly predict the markup sequences from formula images with the attention mechanism. However, such methods may fail to accurately read formulas with complicated structure or ge…

2021

AgentFormer: Agent-Aware Transformers for Socio-Temporal Multi-Agent Forecasting

ICCV 2021poster

Predicting accurate future trajectories of multiple agents is essential for autonomous systems but is challenging due to the complex interaction between agents and the uncertainty in each agent's future behavior. Forecasting multi-agent trajectories requires modeling two key dimensions: (1) time dim…

Cited by 594PDFcodeScholar
2021

AnchorFace: An Anchor-based Facial Landmark Detector Across Large Poses

AAAI 2021technical

Facial landmark localization aims to detect the predefined points of human faces, and the topic has been rapidly improved with the recent development of neural network based methods. However, it remains a challenging task when dealing with faces in unconstrained scenarios, especially with large pose…

2021

Dynamics-regulated kinematic policy for egocentric pose estimation

NeurIPS 2021poster

We propose a method for object-aware 3D egocentric pose estimation that tightly integrates kinematics modeling, dynamics modeling, and scene object information. Unlike prior kinematics or dynamics-based approaches where the two components are used disjointly, we synergize the two approaches via dyna…

2021

FSCE: Few-Shot Object Detection via Contrastive Proposal Encoding

CVPR 2021poster

Emerging interests have been brought to recognize previously unseen objects given very few training examples, known as few-shot object detection (FSOD). Recent researches demonstrate that good feature embedding is the key to reach favorable few-shot learning performance. We observe object proposals…

Cited by 531PDFcodeScholar
2021

SimPoE: Simulated Character Control for 3D Human Pose Estimation

CVPR 2021poster

Accurate estimation of 3D human motion from monocular video requires modeling both kinematics (body motion without physical forces) and dynamics (motion with physical forces). To demonstrate this, we present SimPoE, a Simulation-based approach for 3D human Pose Estimation, which integrates image-bas…

Cited by 163PDFScholar
2021

Temporal Knowledge Consistency for Unsupervised Visual Representation Learning

ICCV 2021poster

The instance discrimination paradigm has become dominant in unsupervised learning. It always adopts a teacher-student framework, in which the teacher provides embedded knowledge as a supervision signal for the student. The student learns meaningful representations by enforcing instance spatial consi…

Cited by 13PDFcodeScholar
2021

Unsupervised Active Learning via Subspace Learning

AAAI 2021technical

Unsupervised active learning has been an active research topic in machine learning community, with the purpose of choosing representative samples to be labelled in an unsupervised manner. Previous works usually take the minimization of data reconstruction loss as the criterion to select representati…

Cited by 18SourcePDFScholar
2020

Beta R-CNN: Looking into Pedestrian Detection from Another Perspective

NeurIPS 2020poster

Recently significant progress has been made in pedestrian detection, but it remains challenging to achieve high performance in occluded and crowded scenes. It could be mostly attributed to the widely used representation of pedestrians, i.e., 2Daxis-aligned bounding box, which just describes the appr…

2020

Efficient Non-Line-of-Sight Imaging from Transient Sinograms

ECCV 2020poster

Non-line-of-sight (NLOS) imaging techniques use light that diffusely reflects off of visible surfaces (e.g., walls) to see around corners. One approach involves using pulsed lasers and ultrafast sensors to measure the travel time of multiply scattered light. Unlike existing NLOS techniques that gene…

Cited by 48SourcePDFScholar
2020

Generative Hybrid Representations for Activity Forecasting With No-Regret Learning

CVPR 2020oral

Automatically reasoning about future human behaviors is a difficult problem but has significant practical applications to assistive systems. Part of this difficulty stems from learning systems' inability to represent all kinds of behaviors. Some behaviors, such as motion, are best described with con…

Cited by 36PDFScholar
2020

Optical Non-Line-of-Sight Physics-Based 3D Human Pose Estimation

CVPR 2020poster

We describe a method for 3D human pose estimation from transient images (i.e., a 3D spatio-temporal histogram of photons) acquired by an optical non-line-of-sight (NLOS) imaging system. Our method can perceive 3D human pose by 'looking around corners' through the use of light indirectly reflected by…

Cited by 90PDFcodeScholar
2020

Residual Force Control for Agile Human Behavior Imitation and Extended Motion Synthesis

NeurIPS 2020poster

Reinforcement learning has shown great promise for synthesizing realistic human behaviors by learning humanoid control policies from motion capture data. However, it is still very challenging to reproduce sophisticated human skills like ballet dance, or to stably imitate long-term human behaviors wi…

2020

Scalable Graph Neural Networks via Bidirectional Propagation

NeurIPS 2020poster

Graph Neural Networks (GNN) are an emerging field for learning on non-Euclidean data. Recently, there has been increased interest in designing GNN that scales to large graphs. Most existing methods use "graph sampling" or "layer-wise sampling" techniques to reduce training time; However, these metho…

2020

Self-PU: Self Boosted and Calibrated Positive-Unlabeled Training

ICML 2020poster

Many real-world applications have to tackle the Positive-Unlabeled (PU) learning problem, i.e., learning binary classifiers from a large amount of unlabeled data and a few labeled positive examples. While current state-of-the-art methods employ importance reweighting to design various biased or unbi…

2020

Simultaneous Arrival Matching for New Spatial Crowdsourcing Platforms

IJCAI 2020poster

In recent years, 3D spatial crowdsourcing platforms become popular, in which users and workers travel together to their assigned workplaces for services, such as InterestingSport and Nanguache. A typical problem over 3D spatial crowdsourcing platforms is to match users with suitable workers and work…

Cited by 0SourcePDFScholar
2020

State-Aware Tracker for Real-Time Video Object Segmentation

CVPR 2020poster

In this work, we address the task of semi-supervised video object segmentation (VOS) and explore how to make efficient use of video property to tackle the challenge of semi-supervision. We propose a novel pipeline called State-Aware Tracker (SAT), which can produce accurate segmentation results with…

Cited by 149PDFcodeScholar
2020

Uncertainty Quantification for Deep Context-Aware Mobile Activity Recognition and Unknown Context Discovery

AISTATS 2020poster

Activity recognition in wearable computing faces two key challenges: i) activity characteristics may be context-dependent and change under different contexts or situations; ii) unknown contexts and activities may occur from time to time, requiring flexibility and adaptability of the algorithm. We de…

Cited by 19SourcePDFScholar
2020

pbSGD: Powered Stochastic Gradient Descent Methods for Accelerated Non-Convex Optimization

IJCAI 2020poster

We propose a novel technique for improving the stochastic gradient descent (SGD) method to train deep networks, which we term pbSGD. The proposed pbSGD method simply raises the stochastic gradient to a certain power elementwise during iterations and introduces only one additional parameter, namely,…

2019

ABD-Net: Attentive but Diverse Person Re-Identification

ICCV 2019poster

Attention mechanisms have been found effective for person re-identification (Re-ID). However, the learned "attentive" features are often not naturally uncorrelated or "diverse", which compromises the retrieval performance based on the Euclidean distance. We advocate the complementary powers of atten…

Cited by 672PDFcodeScholar
2017

Words or Characters? Fine-grained Gating for Reading Comprehension

ICLR 2017poster

Previous work combines word-level and character-level representations using concatenation or scalar weighting, which is suboptimal for high-level tasks like reading comprehension. We present a fine-grained gating mechanism to dynamically combine word-level and character-level representations based o…

Cited by 100SourcecodeScholar
2016

Review Networks for Caption Generation

NeurIPS 2016poster

We propose a novel extension of the encoder-decoder framework, called a review network. The review network is generic and can enhance any existing encoder- decoder model: in this paper, we consider RNN decoders with both CNN and RNN encoders. The review network performs a number of review steps with…