← Search

Can Wang

45 accepted papers

2026

Ani3DHuman: Photorealistic 3D Human Animation with Self-guided Stochastic Sampling

CVPR 2026

Current 3D human animation methods fail at photorealism: kinematics-based approaches lack non-rigid dynamics like clothing, while methods reconstructing from generated videos suffer from low-quality artifacts and identity loss. To overcome these limitations, we present Ani3DHuman, a framework that m

Cited by 0SourcecodeScholar
2026

DICE: Distilling Classifier-Free Guidance into Text Embeddings

AAAI 2026technical

Text-to-image diffusion models are capable of generating high-quality images, but suboptimal pre-trained text representations often result in these images failing to align closely with the given text prompts. Classifier-free guidance (CFG) is a popular and effective technique for improving text-imag

Cited by 0SourcePDFScholar
2026

DetAny4D: Detect Anything 4D Temporally in a Streaming RGB Video

CVPR 2026

Reliable 4D object detection, which refers to 3D object detection in streaming video, is crucial for perceiving and understanding the real world. Existing open-set 4D object detection methods typically make predictions on a frame-by-frame basis without modeling temporal consistency, or rely on compl

Cited by 0SourcecodeScholar
2026

Falsdo: Benchmarking Artifact-Controlled Multimodal Fake News Verification via Failure-Aligned Auditing

IJCAI 2026

Recent generative AI renders multimodal misinformation structurally harder to detect, making reliable detection dependent on semantic verification grounded in verifiable evidence. However, current benchmarks often fail to isolate true semantic checking from superficial shortcut exploitation. We intr

Cited by 0Scholar
2026

From Perception to Planning: Evolving Ego-Centric Task-Oriented Spatiotemporal Reasoning via Curriculum Learning

ICML 2026poster

Modern vision-language models achieve strong performance in static perception, but remain limited in the complex spatiotemporal reasoning required for embodied, egocentric tasks. A major source of failure is their reliance on temporal priors learned from passive video data, which often leads to spat…

Cited by 0SourceScholar
2026

MLLMSplat: A 2D MLLM-Powered Framework for 3D Gaussian Splatting Understanding, Generation, and Editing

CVPR 2026

3D Gaussian Splatting (3DGS) has emerged as a mainstream representation for 3D scenes, drawing increasing research attention to its understanding, generation, and editing. However, existing studies remain limited to low-level perception, low-quality generation, and low-efficiency editing, lagging fa

Cited by 0SourceScholar
2026

On The Surprising Effectiveness of a Single Global Merging in Decentralized Learning

ICLR 2026oral

Decentralized learning provides a scalable alternative to parameter-server-based training, yet its performance is often hindered by limited peer-to-peer communication. In this paper, we study how communication should be scheduled over time to improve global generalization, including determining whe…

Cited by 0SourceScholar
2026

SpatialLogic-Bench: A Diagnostic Benchmark for Task-Oriented Spatiotemporal Reasoning

AAAI 2026technical

Vision-Language Models (VLMs) have made significant progress in static perception, but their ability to understand dynamic task-oriented reasoning remains unclear. Existing benchmarks mainly focus on static spatial relationships and lack systematic assessment of dynamic reasoning capabilities. To th

Cited by 0SourcePDFScholar
2026

Temporal Motif-aware Graph Test-time Adaptation for OOD Blockchain Anomaly Detection

IJCAI 2026

The ever-evolving transaction patterns have significantly hindered anomaly detection on emerging cryptocurrency blockchains due to the vast number of addresses and diverse anomalous behaviors. Recently, advanced Graph Anomaly Detection (GAD) approaches applied to blockchains have faced two critical

Cited by 0Scholar
2025

A Framework for Effective Invocation Methods of Various LLM Services

COLING 2025main

Large Language Models (LLMs) have shown impressive abilities in solving various natural language processing tasks and are now widely offered as services. LLM services enable users to accomplish tasks without requiring specialized knowledge, simply by paying service providers. However, numerous provi…

2025

Advancing Loss Functions in Recommender Systems: A Comparative Study with a Rényi Divergence-Based Solution

AAAI 2025technical

Loss functions play a pivotal role in optimizing recommendation models. Among various loss functions, Softmax Loss (SL) and Cosine Contrastive Loss (CCL) are particularly effective. Their theoretical connections and differences warrant in-depth exploration. This work conducts comprehensive analyses…

2025

Automated Dual-Micropipette Coordination Microinjection for Batch Zebrafish Larvae Based on Pose Estimation

IROS 2025

Zebrafish are widely used in the biomedical field, as an ideal model for microinjection. In automated zebrafish microinjection, posture adjustment is the first and key step, which takes a lot of skill, and injection success assessment is a challenging task. Constrained by these two aspects, it is di

Cited by 0SourceScholar
2025

Knowledge Distillation with Refined Logits

ICCV 2025poster

Recent research on knowledge distillation has increasingly focused on logit distillation because of its simplicity, effectiveness, and versatility in model compression. In this paper, we introduce Refined Logit Distillation (RLD) to address the limitations of current logit distillation methods. Our…

2025

M^2LLM: Multi-view Molecular Representation Learning with Large Language Models

IJCAI 2025

Accurate molecular property prediction is a critical challenge with wide-ranging applications in chemistry, materials science, and drug discovery. Molecular representation methods, including fingerprints and graph neural networks (GNNs), achieve state-of-the-art results by effectively deriving featu

Cited by 0SourcePDFScholar
2025

Recognizing Actions from Robotic View for Natural Human-Robot Interaction

ICCV 2025poster

Natural Human-Robot Interaction (N-HRI) requires robots to recognize human actions at varying distances and states, regardless of whether the robot itself is in motion or stationary. This setup is more flexible and practical than conventional human action recognition tasks. However, existing benchma…

2025

SLOT-MPC: A Hierarchical Whole-Body Model Predictive Controller to Enhance Simultaneous Localization and Object Tracking for UAVs

RA-L 2025

This paper proposes <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">SLOT-MPC</i>, a hierarchical model predictive control framework for a system of multirotor Unmmaned Aerial Vehicle (UAV), which aims to minimize uncertainty in estimating both ego-mo

Cited by 1SourceScholar
2025

sEMG-Based Joint Angle Estimation via Hierarchical Spiking Attentional Feature Decomposition Network

RA-L 2025

Surface electromyography (sEMG) has demonstrated significant potential in simultaneous and proportional control (SPC). However, existing algorithms for predicting joint angles based on sEMG often suffer from high inference costs or are limited to specific subjects rather than multi-subject scenarios

Cited by 2SourcecodeScholar
2024

Fast ODE-based Sampling for Diffusion Models in Around 5 Steps

CVPR 2024highlight

Sampling from diffusion models can be treated as solving the corresponding ordinary differential equations (ODEs) with the aim of obtaining an accurate solution with as few number of function evaluations (NFE) as possible. Recently various fast samplers utilizing higher-order ODE solvers have emerge…

2024

Generating 6-D Trajectories for Omnidirectional Multirotor Aerial Vehicles in Cluttered Environments

RA-L 2024

As fully-actuated systems, omnidirectional multirotor aerial vehicles (OMAVs) have more flexible maneuverability and advantages in aggressive flight in cluttered environments than traditional underactuated MAVs. This letter aims to achieve safe flight of OMAVs in cluttered environments. We present a

Cited by 1SourceScholar
2024

On the Trajectory Regularity of ODE-based Diffusion Sampling

ICML 2024poster

Diffusion-based generative models use stochastic differential equations (SDEs) and their equivalent ordinary differential equations (ODEs) to establish a smooth connection between a complex data distribution and a tractable prior distribution. In this paper, we identify several intriguing trajectory…

2024

PSL: Rethinking and Improving Softmax Loss from Pairwise Perspective for Recommendation

NeurIPS 2024poster

Softmax Loss (SL) is widely applied in recommender systems (RS) and has demonstrated effectiveness. This work analyzes SL from a pairwise perspective, revealing two significant limitations: 1) the relationship between SL and conventional ranking metrics like DCG is not sufficiently tight; 2) SL is h…

2024

Simple and Fast Distillation of Diffusion Models

NeurIPS 2024poster

Diffusion-based generative models have demonstrated their powerful performance across various tasks, but this comes at a cost of the slow sampling speed. To achieve both efficient and high-quality synthesis, various distillation-based accelerated sampling methods have been developed recently. Howeve…

2023

AvatarCraft: Transforming Text into Neural Human Avatars with Parameterized Shape and Pose Control

ICCV 2023poster

Neural implicit fields are powerful for representing 3D scenes and generating high-quality novel views, but it remains challenging to use such implicit representations for creating a 3D human avatar with a specific identity and artistic style that can be easily animated. Our proposed method, AvatarC…

Cited by 81PDFcodeScholar
2023

Continuous Estimation of Lower Limb Joint Angles From Multi-Stream Signals Based on Knowledge Tracing

RA-L 2023

Multi-stream signals are increasingly being used in robot-assisted rehabilitation training, where the timely and accurate prediction of a patient's motor intentions is frequently required to provide simultaneous and proportional control strategies. However, existing methods for motion intent predict

Cited by 21SourceScholar
2023

OpenGSL: A Comprehensive Benchmark for Graph Structure Learning

NeurIPS 2023poster

Graph Neural Networks (GNNs) have emerged as the *de facto* standard for representation learning on graphs, owing to their ability to effectively integrate graph topology and node attributes. However, the inherent suboptimal nature of node connections, resulting from the complex and contingent forma…

2023

Robust Sequence Networked Submodular Maximization

AAAI 2023technical

In this paper, we study the Robust optimization for sequence Networked submodular maximization (RoseNets) problem. We interweave the robust optimization with the sequence networked submodular maximization. The elements are connected by a directed acyclic graph and the objective function is not subm…

Cited by 0SourcePDFScholar
2022

A Novel Method for Detecting Misclassifications of the Locomotion Mode in Lower-Limb Exoskeleton Robot Control

RA-L 2022

Lower-limb exoskeleton robots can support hemiplegic patients’ affected limbs and assist in their rehabilitation. In order to set effective control strategies, it is necessary to obtain the user’s motion intention accurately and timeously. These requirements pose many challenges. The surface electro

Cited by 23SourceScholar
2022

CLIP-NeRF: Text-and-Image Driven Manipulation of Neural Radiance Fields

CVPR 2022poster

We present CLIP-NeRF, a multi-modal 3D object manipulation method for neural radiance fields (NeRF). By leveraging the joint language-image embedding space of the recent Contrastive Language-Image Pre-Training (CLIP) model, we propose a unified framework that allows manipulating NeRF in a user-frien…

Cited by 458PDFcodeScholar
2022

Design and Characteristics of 3D Magnetically Steerable Guidewire System for Minimally Invasive Surgery

RA-L 2022

Endovascular techniques have been increasingly adapted to medical applications as a minimally-invasive treatment approach to diagnose and treat various vascular diseases. Guidewires are the basic devices in endovascular surgery. The conventional practice of endovascular procedures has a number of dr

Cited by 39SourceScholar
2022

Knowledge Distillation With the Reused Teacher Classifier

CVPR 2022poster

Knowledge distillation aims to compress a powerful yet cumbersome teacher model into a lightweight student model without much sacrifice of performance. For this purpose, various approaches have been proposed over the past few years, generally with elaborately designed knowledge representations, whic…

Cited by 246PDFcodeScholar
2022

Pseudo-Labeled Auto-Curriculum Learning for Semi-Supervised Keypoint Localization

ICLR 2022poster

Localizing keypoints of an object is a basic visual problem. However, supervised learning of a keypoint localization network often requires a large amount of data, which is expensive and time-consuming to obtain. To remedy this, there is an ever-growing interest in semi-supervised learning (SSL), wh…

Cited by 20SourcePDFScholar
2021

Cross-Layer Distillation with Semantic Calibration

AAAI 2021technical

Recently proposed knowledge distillation approaches based on feature-map transfer validate that intermediate layers of a teacher model can serve as effective targets for training a student model to obtain better generalization ability. Existing studies mainly focus on particular representation forms…

2021

Distilling Holistic Knowledge With Graph Neural Networks

ICCV 2021poster

Knowledge Distillation (KD) aims at transferring knowledge from a larger well-optimized teacher network to a smaller learnable student network. Existing KD methods have mainly considered two types of knowledge, namely the individual knowledge and the relational knowledge. However, these two types of…

Cited by 80PDFcodeScholar
2021

Human Pose Regression With Residual Log-Likelihood Estimation

ICCV 2021poster

Heatmap-based methods dominate in the field of human pose estimation by modelling the output distribution through likelihood heatmaps. In contrast, regression-based methods are more efficient but suffer from inferior performance. In this work, we explore maximum likelihood estimation (MLE) to develo…

Cited by 277PDFcodeScholar
2020

Adaptive Graph Convolutional Recurrent Network for Traffic Forecasting

NeurIPS 2020poster

Modeling complex spatial and temporal correlations in the correlated time series data is indispensable for understanding the traffic dynamics and predicting the future status of an evolving traffic system. Recent works focus on designing complicated graph neural network architectures to capture shar…

2020

HMOR: Hierarchical Multi-Person Ordinal Relations for Monocular Multi-Person 3D Pose Estimation

ECCV 2020poster

Remarkable progress has been made in 3D human pose estimation from a monocular RGB camera. However, only a few studies explored 3D multi-person cases. In this paper, we attempt to address the lack of a global perspective of the top-down approaches by introducing a novel form of supervision - Hierarc…

Cited by 75SourcePDFScholar
2020

Whole-Body Human Pose Estimation in the Wild

ECCV 2020poster

This paper investigates the task of 2D human whole-body pose estimation, which aims to localize dense landmarks on the entire human body including face, hands, body, and feet. As existing datasets do not have whole-body annotations, previous methods have to assemble different deep models trained ind…

2019

CrowdPose: Efficient Crowded Scenes Pose Estimation and a New Benchmark

CVPR 2019oral

Multi-person pose estimation is fundamental to many computer vision tasks and has made significant progress in recent years. However, few previous methods explored the problem of pose estimation in crowded scenes while it remains challenging and inevitable in many scenarios. Moreover, current benchm…

Cited by 699PDFcodeScholar
2015

A predictive model for narrow passage path planner by using Support Vector Machine in changing environments

ICRA 2015poster

Narrow passages in changing environments create huge difficulties, since locations and shapes of narrow passages in Configuration Space(C-space) change frequently. It is very important for a planner to identify narrow passages in real time and boost valid points within them effectively. A novel narr…

Cited by 5SourceScholar