← Search

Jingya Wang

58 accepted papers

2026

Bridging the Gap in Autonomous Science: The Corpus and Benchmark for Biological Protocol Reasoning

ICML 2026poster

The realization of autonomous scientific experimentation is currently limited by LLMs' struggle to grasp the strict procedural logic and accuracy required by biological protocols. To address this fundamental challenge, we present **BioProBench**, a comprehensive resource for procedural reasoning in …

Cited by 0SourceScholar
2026

Conformal Reliability: A New Evaluation Metric for Conditional Generation

ICML 2026poster

Conditional generative models have recently achieved remarkable success in various applications. However, a suitable metric for evaluating the reliability of these models, which takes into account their inherent uncertainty, is still lacking. Existing metrics, which typically assess a single output,…

Cited by 0SourceScholar
2026

Diffusion Bridge or Flow Matching? A Unifying Framework and Comparative Analysis

ICML 2026poster

Diffusion Bridge and Flow Matching have both demonstrated compelling empirical performance in transformation between arbitrary distributions. However, there remains confusion about which approach is generally preferable, and the substantial discrepancies in their modeling assumptions and practical i…

Cited by 0SourceScholar
2026

DiscoForcing: A Unified Framework for Real-Time Audio-Driven Character Control with Diffusion Forcing

ICML 2026poster

We study real-time audio-responsive character control as a deployment-faithful problem: strictly causal, bounded-latency streaming that must generate coherent full-body motion at interactive frame rates while the audio condition can change abruptly (tempo shifts, drops, or user edits). Prior music-t…

Cited by 0SourceScholar
2026

Human-Object Interaction via Automatically Designed VLM-Guided Motion Policy

ICLR 2026poster

Human-object interaction (HOI) synthesis is crucial for applications in animation, simulation, and robotics. However, existing approaches either rely on expensive motion capture data or require manual reward engineering, limiting their scalability and generalizability. In this work, we introduce the…

Cited by 0SourcecodeScholar
2026

InterAgent: Physics-based Multi-agent Command Execution via Diffusion on Interaction Graphs

CVPR 2026

Humanoid agents are expected to emulate the complex coordination inherent in human social behaviors. However, existing methods are largely confined to single-agent scenarios, overlooking the physically plausible interplay essential for multi-agent interactions. To bridge this gap, we propose InterAg

Cited by 0SourcecodeScholar
2026

MIMIC: Mask-Injected Manipulation Video Generation with Interaction Control

ICLR 2026poster

Embodied intelligence faces a fundamental bottleneck from limited large-scale interaction data. Video generation offers a scalable alternative, but manipulation videos remain particularly challenging, as they require capturing subtle, contact-rich dynamics. Despite recent advances, video diffusion m…

Cited by 0SourceScholar
2026

Neural Dynamics Self-Attention for Spiking Transformers

ICLR 2026poster

Integrating Spiking Neural Networks (SNNs) with Transformer architectures offers a promising pathway to balance energy efficiency and performance, particularly for edge vision applications. However, existing Spiking Transformers face two critical challenges: i) a substantial performance gap relative…

Cited by 0SourceScholar
2026

Sample from What You See: Visuomotor Policy Learning via Diffusion Bridge with Observation-Embedded Stochastic Differential Equation

ICML 2026poster

Imitation learning with diffusion models has advanced robotic control by capturing the multi-modal action distributions. However, existing methods typically treat observations only as high-level conditions to the denoising network, rather than integrating them into the stochastic dynamics of the dif…

Cited by 0SourceScholar
2026

Sample-Efficient Diffusion-based Reinforcement Learning with Critic Guidance

ICML 2026poster

Recent advances in reinforcement learning (RL) have achieved great successes by leveraging the multimodality and exploration capability of diffusion policies. Among these approaches, one representative branch focuses on the weighted-based policy optimization. This design enables better exploration c…

Cited by 0SourceScholar
2026

Training-Free ANN-to-SNN Conversion for High-Performance Spiking Transformers

AAAI 2026technical

Leveraging the event-driven paradigm, Spiking Neural Networks (SNNs) offer a promising approach for constructing energy-efficient Transformer architectures. Compared to directly trained Spiking Transformers, ANN-to-SNN conversion methods bypass the high training costs. However, existing methods stil

Cited by 0SourcePDFScholar
2026

UniHM: Unified Dexterous Hand Manipulation with Vision Language Model

ICLR 2026poster

Planning physically feasible dexterous hand manipulation is a central challenge in robotic manipulation and Embodied AI. Prior work typically relies on object-centric cues or precise hand-object interaction sequences, foregoing the rich, compositional guidance of open-vocabulary instruction. We intr…

Cited by 0SourcecodeScholar
2025

AffordDP: Generalizable Diffusion Policy with Transferable Affordance

CVPR 2025poster

Diffusion-based policies have shown impressive performance in robotic manipulation tasks while struggling with out-of-domain distributions. Recent efforts attempted to enhance generalization by improving the visual feature encoding for diffusion policy. However, their generalization is typically lim…

Cited by 5SourcePDFScholar
2025

Bipolar Self-attention for Spiking Transformers

NeurIPS 2025spotlight

Harnessing the event-driven characteristic, Spiking Neural Networks (SNNs) present a promising avenue toward energy-efficient Transformer architectures. However, existing Spiking Transformers still suffer significant performance gaps compared to their Artificial Neural Network counterparts. Through…

Cited by 0SourceScholar
2025

Capturing the Unseen: Vision-Free Facial Motion Capture Using Inertial Measurement Units

AAAI 2025technical

We present Capturing the Unseen (CAPUS), a novel facial motion capture (MoCap) technique that operates without visual signals. CAPUS leverages miniaturized Inertial Measurement Units (IMUs) as a new sensing modality for facial motion capture. While IMUs have become essential in full-body MoCap for t…

Cited by 0SourcePDFScholar
2025

Dendritic Resonate-and-Fire Neuron for Effective and Efficient Long Sequence Modeling

NeurIPS 2025poster

The explosive growth in sequence length has intensified the demand for effective and efficient long sequence modeling. Benefiting from intrinsic oscillatory membrane dynamics, Resonate-and-Fire (RF) neurons can efficiently extract frequency components from input signals and encode them into spatiote…

Cited by 0SourceScholar
2025

GenPO: Generative Diffusion Models Meet On-Policy Reinforcement Learning

NeurIPS 2025poster

Recent advances in reinforcement learning (RL) have demonstrated the powerful exploration capabilities and multimodality of generative diffusion-based policies. While substantial progress has been made in offline RL and off-policy RL settings, integrating diffusion policies into on-policy frameworks…

Cited by 0SourceScholar
2025

IDE: A Multi-Agent-Driven Iterative Framework for Dynamic Evaluation of LLMs

ICASSP 2025accepted

With the widespread use of large language models (LLMs) in natural language processing, traditional evaluation methods based on static datasets have become inadequate to fully capture their performance and generalization capabilities. To address this challenge, we propose an Iterative Dynamic Evalua…

Cited by 0SourceScholar
2025

LithoSim: A Large, Holistic Lithography Simulation Benchmark for AI-Driven Semiconductor Manufacturing

NeurIPS 2025poster

Lithography orchestrates a symphony of light, mask and photochemicals to transfer the integrated circuit patterns onto the wafer. Lithography simulation serves as the critical nexus between circuit design and manufacturing, where its speed and accuracy fundamentally govern the optimization quality o…

Cited by 0SourcecodeScholar
2025

Multi-modal Multi-platform Person Re-Identification: Benchmark and Method

ICCV 2025poster

Conventional person re-identification (ReID) research is often limited to single-modality sensor data from static cameras, which fails to address the complexities of real-world scenarios where multi-modal signals are increasingly prevalent. For instance, consider an urban ReID system integrating sta…

Cited by 0SourcePDFScholar
2025

NLPrompt: Noise-Label Prompt Learning for Vision-Language Models

CVPR 2025highlight

The emergence of vision-language foundation models, such as CLIP, has revolutionized image-text representation, enabling a broad range of applications via prompt learning. Despite its promise, real-world datasets often contain noisy labels that can degrade prompt learning performance. In this paper,…

2025

OpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language Model

NeurIPS 2025oral

Understanding and synthesizing realistic 3D hand-object interactions (HOI) is critical for applications ranging from immersive AR/VR to dexterous robotics. Existing methods struggle with generalization, performing well on closed-set objects and predefined tasks but failing to handle unseen objects o…

Cited by 0SourceScholar
2025

SMGDiff: Soccer Motion Generation using Diffusion Probabilistic Models

ICCV 2025poster

Soccer is a globally renowned sport with significant applications in video games and VR/AR. However, generating realistic soccer motions remains challenging due to the intricate interactions between the player and the ball. In this paper, we introduce SMGDiff, a novel two-stage framework for generat…

Cited by 0SourcePDFScholar
2025

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model

CVPR 2025poster

3D affordance segmentation aims to link human instructions to touchable regions of 3D objects for embodied manipulations. Existing efforts typically adhere to single-object, single-affordance paradigms, where each affordance type or explicit instruction strictly corresponds to a specific affordance…

Cited by 3SourcePDFScholar
2025

TokMan:Tokenize Manhattan Mask Optimization for Inverse Lithography

NeurIPS 2025poster

Manhattan representations, defined by axis-aligned, orthogonal structures, are widely used in vision, robotics, and semiconductor design for their geometric regularity and algorithmic simplicity. In integrated circuit (IC) design, Manhattan geometry is key for routing, design rule checking, and lith…

Cited by 0SourceScholar
2025

Towards Immersive Human-X Interaction: A Real-Time Framework for Physically Plausible Motion Synthesis

ICCV 2025poster

Real-time synthesis of physically plausible human interactions remains a critical challenge for immersive VR/AR systems and humanoid robotics. While existing methods demonstrate progress in kinematic motion generation, they often fail to address the fundamental tension between real-time responsivene…

Cited by 0SourcePDFScholar
2025

UniDB: A Unified Diffusion Bridge Framework via Stochastic Optimal Control

ICML 2025spotlight

Recent advances in diffusion bridge models leverage Doob’s $h$-transform to establish fixed endpoints between distributions, demonstrating promising results in image translation and restoration tasks. However, these approaches frequently produce blurred or excessively smoothed image details and lack…

2024

A Unified Diffusion Framework for Scene-aware Human Motion Estimation from Sparse Signals

CVPR 2024poster

Estimating full-body human motion via sparse tracking signals from head-mounted displays and hand controllers in 3D scenes is crucial to applications in AR/VR. One of the biggest challenges to this task is the one-to-many mapping from sparse observations to dense full-body motions which endowed inhe…

2024

BOTH2Hands: Inferring 3D Hands from Both Text Prompts and Body Dynamics

CVPR 2024poster

The recently emerging text-to-motion advances have spired numerous attempts for convenient and interactive human motion generation. Yet existing methods are largely limited to generating body motions only without considering the rich two-hand motions let alone handling various conditions like body d…

2024

Diffusion-based Reinforcement Learning via Q-weighted Variational Policy Optimization

NeurIPS 2024poster

Diffusion models have garnered widespread attention in Reinforcement Learning (RL) for their powerful expressiveness and multimodality. It has been verified that utilizing diffusion policies can significantly improve the performance of RL algorithms in continuous control tasks by overcoming the limi…

2024

Global and Local Prompts Cooperation via Optimal Transport for Federated Learning

CVPR 2024poster

Prompt learning in pretrained visual-language models has shown remarkable flexibility across various downstream tasks. Leveraging its inherent lightweight nature recent research attempted to integrate the powerful pretrained models into federated learning frameworks to simultaneously reduce communic…

2024

Guidance with Spherical Gaussian Constraint for Conditional Diffusion

ICML 2024poster

Recent advances in diffusion models attempt to handle conditional generative tasks by utilizing a differentiable loss function for guidance without the need for additional training. While these methods achieved certain success, they often compromise on sample quality and require small guidance step…

2024

HOI-M^3: Capture Multiple Humans and Objects Interaction within Contextual Environment

CVPR 2024highlight

Humans naturally interact with both others and the surrounding multiple objects engaging in various social activities. However recent advances in modeling human-object interactions mostly focus on perceiving isolated individuals and objects due to fundamental data scarcity. In this paper we introduc…

Cited by 12SourcePDFScholar
2024

Harmonizing Generalization and Personalization in Federated Prompt Learning

ICML 2024poster

Federated Prompt Learning (FPL) incorporates large pre-trained Vision-Language models (VLM) into federated learning through prompt tuning. The transferable representations and remarkable generalization capacity of VLM make them highly compatible with the integration of federated learning. Addressing…

2024

HybridGait: A Benchmark for Spatial-Temporal Cloth-Changing Gait Recognition with Hybrid Explorations

AAAI 2024technical

Existing gait recognition benchmarks mostly include minor clothing variations in the laboratory environments, but lack persistent changes in appearance over time and space. In this paper, we propose the first in-the-wild benchmark CCGait for cloth-changing gait recognition, which incorporates divers…

2024

I'M HOI: Inertia-aware Monocular Capture of 3D Human-Object Interactions

CVPR 2024poster

We are living in a world surrounded by diverse and "smart" devices with rich modalities of sensing ability. Conveniently capturing the interactions between us humans and these objects remains far-reaching. In this paper we present I'm-HOI a monocular scheme to faithfully capture the 3D motions of bo…

Cited by 7SourcePDFScholar
2024

LiveHPS: LiDAR-based Scene-level Human Pose and Shape Estimation in Free Environment

CVPR 2024highlight

For human-centric large-scale scenes fine-grained modeling for 3D human global pose and shape is significant for scene understanding and can benefit many real-world applications. In this paper we present LiveHPS a novel single-LiDAR-based approach for scene-level human pose and shape estimation with…

Cited by 15SourcePDFScholar
2024

Unsupervised Cross-Domain Image Retrieval via Prototypical Optimal Transport

AAAI 2024technical

Unsupervised cross-domain image retrieval (UCIR) aims to retrieve images sharing the same category across diverse domains without relying on labeled data. Prior approaches have typically decomposed the UCIR problem into two distinct tasks: intra-domain representation learning and cross-domain featur…

2023

Alternating Differentiation for Optimization Layers

ICLR 2023poster

The idea of embedding optimization problems into deep neural networks as optimization layers to encode constraints and inductive priors has taken hold in recent years. Most existing methods focus on implicitly differentiating Karush–Kuhn–Tucker (KKT) conditions in a way that requires expensive compu…

2023

CSOT: Curriculum and Structure-Aware Optimal Transport for Learning with Noisy Labels

NeurIPS 2023poster

Learning with noisy labels (LNL) poses a significant challenge in training a well-generalized model while avoiding overfitting to corrupted labels. Recent advances have achieved impressive performance by identifying clean labels and correcting corrupted labels for training. However, the current appr…

2023

Fed-CO$_{2}$: Cooperation of Online and Offline Models for Severe Data Heterogeneity in Federated Learning

NeurIPS 2023poster

Federated Learning (FL) has emerged as a promising distributed learning paradigm that enables multiple clients to learn a global model collaboratively without sharing their private data. However, the effectiveness of FL is highly dependent on the quality of the data that is being used for training.…

2023

HybridCap: Inertia-Aid Monocular Capture of Challenging Human Motions

AAAI 2023technical

Monocular 3D motion capture (mocap) is beneficial to many applications. The use of a single camera, however, often fails to handle occlusions of different body parts and hence it is limited to capture relatively simple movements. We present a light-weight, hybrid mocap technique called HybridCap tha…

2023

IKOL: Inverse Kinematics Optimization Layer for 3D Human Pose and Shape Estimation via Gauss-Newton Differentiation

AAAI 2023technical

This paper presents an inverse kinematic optimization layer (IKOL) for 3D human pose and shape estimation that leverages the strength of both optimization- and regression-based methods within an end-to-end framework. IKOL involves a nonconvex optimization that establishes an implicit mapping from an…

2023

Knowledge-Aware Federated Active Learning with Non-IID Data

ICCV 2023poster

Federated learning enables multiple decentralized clients to learn collaboratively without sharing local data. However, the expensive annotation cost on local clients remains an obstacle in utilizing local data. In this paper, we propose a federated active learning paradigm to efficiently learn a gl…

Cited by 25PDFcodeScholar
2023

Lifelong Person Re-identification via Knowledge Refreshing and Consolidation

AAAI 2023technical

Lifelong person re-identification (LReID) is in significant demand for real-world development as a large amount of ReID data is captured from diverse locations over time and cannot be accessed at once inherently. However, a key challenge for LReID is how to incrementally preserve old knowledge and g…

2023

NeuralDome: A Neural Modeling Pipeline on Multi-View Human-Object Interactions

CVPR 2023poster

Humans constantly interact with objects in daily life tasks. Capturing such processes and subsequently conducting visual inferences from a fixed viewpoint suffers from occlusions, shape and texture ambiguities, motions, etc. To mitigate the problem, it is essential to build a training dataset that c…

2023

Reduced Policy Optimization for Continuous Control with Hard Constraints

NeurIPS 2023poster

Recent advances in constrained reinforcement learning (RL) have endowed reinforcement learning with certain safety guarantees. However, deploying existing constrained RL algorithms in continuous control tasks with general hard constraints remains challenging, particularly in those situations with no…

2023

StackFLOW: Monocular Human-Object Reconstruction by Stacked Normalizing Flow with Offset

IJCAI 2023poster

Modeling and capturing the 3D spatial arrangement of the human and the object is the key to perceiving 3D human-object interaction from monocular images. In this work, we propose to use the Human-Object Offset between anchors which are densely sampled from the surface of human mesh and object mesh t…

2023

Two Sides of The Same Coin: Bridging Deep Equilibrium Models and Neural ODEs via Homotopy Continuation

NeurIPS 2023poster

Deep Equilibrium Models (DEQs) and Neural Ordinary Differential Equations (Neural ODEs) are two branches of implicit models that have achieved remarkable success owing to their superior performance and low memory consumption. While both are implicit models, DEQs and Neural ODEs are derived from diff…

2023

Weakly Supervised 3D Multi-Person Pose Estimation for Large-Scale Scenes Based on Monocular Camera and Single LiDAR

AAAI 2023technical

Depth estimation is usually ill-posed and ambiguous for monocular camera-based 3D multi-person pose estimation. Since LiDAR can capture accurate depth information in long-range scenes, it can benefit both the global localization of individuals and the 3D pose estimation by providing rich geometry fe…

2022

Unified Optimal Transport Framework for Universal Domain Adaptation

NeurIPS 2022accept

Universal Domain Adaptation (UniDA) aims to transfer knowledge from a source domain to a target domain without any constraints on label sets. Since both domains may hold private classes, identifying target common samples for domain alignment is an essential issue in UniDA. Most existing methods requ…

2019

Deep Reinforcement Active Learning for Human-in-the-Loop Person Re-Identification

ICCV 2019oral

Most existing person re-identification(Re-ID) approaches achieve superior results based on the assumption that a large amount of pre-labelled data is usually available and can be put into training phrase all at once. However, this assumption is not applicable to most real-world deployment of the Re-…

Cited by 118PDFScholar
2018

Transferable Joint Attribute-Identity Deep Learning for Unsupervised Person Re-Identification

CVPR 2018poster

Most existing person re-identification (re-id) methods require supervised model learning from a separate large set of pairwise labelled training data for every single camera pair. This significantly limits their scalability and usability in real-world large scale deployments with the need for perfor…

Cited by 728SourcePDFScholar
2017

Attribute Recognition by Joint Recurrent Learning of Context and Correlation

ICCV 2017poster

Recognising semantic pedestrian attributes in surveillance images is a challenging task for computer vision, particularly when the imaging quality is poor with complex background clutter and uncontrolled viewing conditions, and the number of labelled training data is small. In this work, we formulat…

Cited by 181PDFScholar