← Search

Zhiyuan Zhang

48 accepted papers

2026

3D Dynamics-Aware Manipulation: Endowing Manipulation Policies with 3D Foresight

ICRA 2026poster

The incorporation of world modeling into manipulation policy learning has pushed the boundary of manipulation performance. However, existing efforts simply model the 2D visual dynamics, which is insufficient for robust manipulation when target tasks involve prominent depth-wise movement. To address …

2026

CityLens: Evaluating Large Vision-Language Models for Urban Socioeconomic Sensing

ICLR 2026poster

Understanding urban socioeconomic conditions through visual data is a challenging yet essential task for sustainable urban development and policy planning. In this work, we introduce CityLens, a comprehensive benchmark designed to evaluate the capabilities of Large Vision-Language Models (LVLMs) in…

Cited by 0SourcecodeScholar
2026

EmbryoDiff: A Conditional Diffusion Framework with Multi-Focal Feature Fusion for Fine-Grained Embryo Developmental Stage Recognition

AAAI 2026technical

Identification of fine-grained embryo developmental stages during In Vitro Fertilization (IVF) is crucial for assessing embryo viability. Although recent deep learning methods have achieved promising accuracy, existing discriminative models fail to utilize the distributional prior of embryonic devel

Cited by 0SourcePDFScholar
2026

SegPVSG: Panoptic Video Scene Graph Generation via Temporal Focusing and Generative Augmentation

ICML 2026poster

Panoptic Video Scene Graph Generation (PVSG) aims to identify relations between pixel-level entities in a video, serving as a novel paradigm for structured video parsing. However, this task faces two key challenges. First, the interactions between entities are temporally fragmented and sparse, meani…

Cited by 0SourceScholar
2026

Spatial Retrieval Augmented Autonomous Driving

CVPR 2026

Existing autonomous driving systems rely on onboard sensors (cameras, LiDAR, IMU, etc) for environmental perception. However, this paradigm is limited by the drive-time perception horizon and often fails under limited view scope, occlusion or extreme conditions such as darkness and rain. In contrast

Cited by 0SourcecodeScholar
2026

TrajTok: What makes for a good trajectory tokenizer in behavior generation?

ICLR 2026poster

Behavior generation in autonomous driving aims to simulate dynamic driving scenarios from recorded driving logs. A popular approach is to apply next-token-prediction with discrete trajectory tokenization. In this work, we explore what makes a good trajectory tokenizer from the perspective of logged…

Cited by 0SourcecodeScholar
2025

Closed-Loop Open-Vocabulary Mobile Manipulation with GPT-4V

ICRA 2025

Autonomous robot navigation and manipulation in open environments require reasoning and replanning with closed-loop feedback. In this work, we present COME-robot, the first closed-loop robotic system utilizing the GPT-4V vision-language foundation model for open-ended reasoning and adaptive planning

Cited by 62SourceScholar
2025

ControlVLA: Few-shot Object-centric Adaptation for Pre-trained Vision-Language-Action Models

CoRL 2025poster

Learning real-world robotic manipulation is challenging, particularly when limited demonstrations are available. Existing methods for few-shot manipulation often rely on simulation-augmented data or pre-built modules like grasping and pose estimation, which struggle with sim-to-real gaps and lack ex…

Cited by 0SourceScholar
2025

DriveTransformer: Unified Transformer for Scalable End-to-End Autonomous Driving

ICLR 2025poster

End-to-end autonomous driving (E2E-AD) has emerged as a trend in the field of autonomous driving, promising a data-driven, scalable approach to system design. However, existing E2E-AD methods usually adopt the sequential paradigm of perception-prediction-planning, which leads to cumulative errors an…

2025

GenM3: Generative Pretrained Multi-path Motion Model for Text Conditional Human Motion Generation

ICCV 2025poster

Scaling up motion datasets is crucial to enhance motion generation capabilities. However, training on large-scale multi-source datasets introduces data heterogeneity challenges due to variations in motion content. To address this, we propose Generative Pretrained Multi-path Motion Model (GenM^3), a…

Cited by 0SourcePDFScholar
2025

Improving Multimodal Human Pose Estimation by Adversarial Modality Enhancement†

ICASSP 2025accepted

Human pose estimation in computer vision predominantly focuses on the visible modality, with limited research on the infrared modality. No existing methods demonstrate robust performance across both modalities, missing their complementary strengths. This gap arises from the lack of a multimodal benc…

Cited by 0SourceScholar
2025

Information-Bottleneck Driven Binary Neural Network for Change Detection

ICCV 2025poster

In this paper, we propose Binarized Change Detection (BiCD), the first binary neural network (BNN) designed specifically for change detection. Conventional network binarization approaches, which directly quantize both weights and activations in change detection models, severely limit the network's a…

2025

Residual Descent Differential Dynamic Game (RD3G) - A Fast Newton Solver for Constrained General Sum Games

ICRA 2025

We present Residual Descent Differential Dynamic Game (RD3G), a Newton-based solver for constrained multiagent game-control problems. The proposed solver seeks a local Nash equilibrium for games where agents are coupled through their rewards and state constraints. By maintaining a dynamic set of act

Cited by 1SourceScholar
2024

Bench2Drive: Towards Multi-Ability Benchmarking of Closed-Loop End-To-End Autonomous Driving

NeurIPS 2024poster

In an era marked by the rapid scaling of foundation models, autonomous driving technologies are approaching a transformative threshold where end-to-end autonomous driving (E2E-AD) emerges due to its potential of scaling up in the data-driven manner. However, existing E2E-AD methods are mostly evalua…

Cited by 39SourcePDFScholar
2024

Binary Amplitude-Only Hologram Generation for Acoustic End-Effector Design by Physics-based deep learning

IROS 2024poster

Acoustic holography has emerged as a cutting-edge technique for constructing a micro-robot acoustic end-effector for non-contact manipulation. As one of typical implementations of acoustic holography, Binary Amplitude-Only Hologram (BAOH) featured with a simple structure provides an efficient altern…

Cited by 0SourceScholar
2024

BuzzRacer: A Palm-sized Autonomous Vehicle Platform for Testing Multi-Agent Adversarial Decision-Making

IROS 2024poster

We present BuzzRacer, a palm-sized autonomous vehicle platform suitable for multi-agent autonomous racing. BuzzRacer consists of two parts. First, a software framework with multiple racetrack environments, dynamic simulation, visualization, and control pipelines. Second, a miniature autonomous vehic…

Cited by 2SourceScholar
2024

DSPy: Compiling Declarative Language Model Calls into State-of-the-Art Pipelines

ICLR 2024spotlight

The ML community is rapidly exploring techniques for prompting language models (LMs) and for stacking them into pipelines that solve complex tasks. Unfortunately, existing LM pipelines are typically implemented using hard-coded “prompt templates”, i.e. lengthy strings discovered via trial and error.…

2024

Enhancing Byzantine-Resistant Aggregations with Client Embedding

EMNLP 2024finding

Byzantine-resistant aggregations detect poisonous clients and discard them to ensure that the global model is not poisoned or attacked by malicious clients. However, these aggregations are mainly conducted on the parameter space, and the parameter distances cannot reflect the data distribution diver…

Cited by 0SourcePDFScholar
2024

GelRoller: A Rolling Vision-based Tactile Sensor for Large Surface Reconstruction Using Self-Supervised Photometric Stereo Method

ICRA 2024poster

Accurate perception of the surrounding environment stands as a primary objective for robots. Through tactile interaction, vision-based tactile sensors provide the capability to capture high-resolution and multi-modal surface information of objects, thereby facilitating robots in achieving more dexte…

Cited by 2SourcecodeScholar
2024

Scaling Up Dynamic Human-Scene Interaction Modeling

CVPR 2024highlight

Confronting the challenges of data scarcity and advanced motion synthesis in human-scene interaction modeling we introduce the TRUMANS dataset alongside a novel HSI motion synthesis method. TRUMANS stands as the most comprehensive motion-captured HSI dataset currently available encompassing over 15…

Cited by 54SourcePDFScholar
2023

Diffusion Theory as a Scalpel: Detecting and Purifying Poisonous Dimensions in Pre-trained Language Models Caused by Backdoor or Bias

ACL 2023findings

Pre-trained Language Models (PLMs) may be poisonous with backdoors or bias injected by the suspicious attacker during the fine-tuning process. A core challenge of purifying potentially poisonous PLMs is precisely finding poisonous dimensions. To settle this issue, we propose the Fine-purifying appro…

Cited by 7SourcePDFScholar
2023

Fed-FA: Theoretically Modeling Client Data Divergence for Federated Language Backdoor Defense

NeurIPS 2023poster

Federated learning algorithms enable neural network models to be trained across multiple decentralized edge devices without sharing private data. However, they are susceptible to backdoor attacks launched by malicious clients. Existing robust federated aggregation algorithms heuristically detect and…

Cited by 4SourcePDFScholar
2023

Full-Body Articulated Human-Object Interaction

ICCV 2023poster

Fine-grained capture of 3D Human-Object Interactions (HOIs) boosts human activity understanding and facilitates various downstream visual tasks. Prior models mostly assume that humans interact with rigid objects using only a few body parts, limiting their scope. In this paper, we address the challen…

Cited by 90PDFcodeScholar
2023

Risk-Aware Model Predictive Path Integral Control Using Conditional Value-at-Risk

ICRA 2023poster

In this paper, we present a novel Model Predictive Control method for autonomous robot planning and control subject to arbitrary forms of uncertainty. The proposed Risk-Aware Model Predictive Path Integral (RA-MPPI) control utilizes the Conditional Value-at-Risk (CVaR) measure to generate optimal co…

Cited by 39SourceScholar
2023

SonoRotor: An Acoustic Rotational Robotic Platform for Zebrafish Embryos and Larvae

RA-L 2023

Rotation manipulation is an essential component of biological microscopy and can become integral to multidisciplinary research and applications. On-chip rotation of microobjects with spherical shapes like biological cells and model organisms has been demonstrated based on advanced microfluidic techn

Cited by 10SourceScholar
2022

CVFNet: Real-time 3D Object Detection by Learning Cross View Features

IROS 2022poster

In recent years 3D object detection from LiDAR point clouds has made great progress thanks to the development of deep learning technologies. Although voxel or point based methods are popular in 3D object detection, they usually involve time-consuming operations such as 3D convolutions on voxels or b…

Cited by 20SourceScholar
2022

Context-Aware Video Reconstruction for Rolling Shutter Cameras

CVPR 2022poster

With the ubiquity of rolling shutter (RS) cameras, it is becoming increasingly attractive to recover the latent global shutter (GS) video from two consecutive RS frames, which also places a higher demand on realism. Existing solutions, using deep neural networks or optimization, achieve promising pe…

Cited by 30PDFcodeScholar
2022

Dim-Krum: Backdoor-Resistant Federated Learning for NLP with Dimension-wise Krum-Based Aggregation

EMNLP 2022finding

Despite the potential of federated learning, it is known to be vulnerable to backdoor attacks. Many robust federated aggregation methods are proposed to reduce the potential backdoor risk. However, they are mainly validated in the CV field. In this paper, we find that NLP backdoors are hard to defen…

Cited by 15SourcePDFScholar
2022

End-to-End Learning the Partial Permutation Matrix for Robust 3D Point Cloud Registration

AAAI 2022technical

Even though considerable progress has been made in deep learning-based 3D point cloud processing, how to obtain accurate correspondences for robust registration remains a major challenge because existing hard assignment methods cannot deal with outliers naturally. Alternatively, the soft matching-ba…

Cited by 33SourcePDFScholar
2022

Expose Backdoors on the Way: A Feature-Based Efficient Defense against Textual Backdoor Attacks

EMNLP 2022finding

Natural language processing (NLP) models are known to be vulnerable to backdoor attacks, which poses a newly arisen threat to NLP models. Prior online backdoor defense methods for NLP models only focus on the anomalies at either the input or output level, still suffering from fragility to adaptive a…

2022

Fine-mixing: Mitigating Backdoors in Fine-tuned Language Models

EMNLP 2022finding

Deep Neural Networks (DNNs) are known to be vulnerable to backdoor attacks. In Natural Language Processing (NLP), DNNs are often backdoored during the fine-tuning process of a large-scale Pre-trained Language Model (PLM) with poisoned samples. Although the clean weights of PLMs are readily available…

2022

GA-SAM: Gradient-Strength based Adaptive Sharpness-Aware Minimization for Improved Generalization

EMNLP 2022main

Recently, Sharpness-Aware Minimization (SAM) algorithm has shown state-of-the-art generalization abilities in vision tasks. It demonstrates that flat minima tend to imply better generalization abilities. However, it has some difficulty implying SAM to some natural language tasks, especially to model…

Cited by 0SourcePDFScholar
2022

How to Inject Backdoors with Better Consistency: Logit Anchoring on Clean Data

ICLR 2022poster

Since training a large-scale backdoored model from scratch requires a large training dataset, several recent attacks have considered to inject backdoors into a trained clean model without altering model behaviors on the clean data. Previous work finds that backdoors can be injected into a trained cl…

Cited by 41SourcePDFScholar
2022

Trajectory Distribution Control for Model Predictive Path Integral Control using Covariance Steering

ICRA 2022poster

This paper presents a novel control approach for autonomous systems operating under uncertainty. We combine Model Predictive Path Integral (MPPI) control with Covariance Steering (CS) theory to obtain a robust controller for general nonlinear systems. The proposed Covariance-Controlled Model Predict…

Cited by 69SourceScholar
2021

Be Careful about Poisoned Word Embeddings: Exploring the Vulnerability of the Embedding Layers in NLP Models

NAACL 2021long

Recent studies have revealed a security threat to natural language processing (NLP) models, called the Backdoor Attack. Victim models can maintain competitive performance on clean samples while behaving abnormally on samples with a specific trigger word inserted. Previous backdoor attacking methods…

2021

Exploring the Vulnerability of Deep Neural Networks: A Study of Parameter Corruption

AAAI 2021technical

We argue that the vulnerability of model parameters is of crucial value to the study of model robustness and generalization but little research has been devoted to understanding this matter. In this work, we propose an indicator to measure the robustness of neural network parameters by exploiting th…

Cited by 41SourcePDFScholar
2021

Neural Network Surgery: Injecting Data Patterns into Pre-trained Models with Minimal Instance-wise Side Effects

NAACL 2021long

Side effects during neural network tuning are typically measured by overall accuracy changes. However, we find that even with similar overall accuracy, existing tuning methods result in non-negligible instance-wise side effects. Motivated by neuroscientific evidence and theoretical results, we demon…

Cited by 13SourcePDFScholar
2021

Raise a Child in Large Language Model: Towards Effective and Generalizable Fine-tuning

EMNLP 2021main

Recent pretrained language models extend from millions to billions of parameters. Thus the need to fine-tune an extremely large pretrained model with a limited training corpus arises in various downstream tasks. In this paper, we propose a straightforward yet effective fine-tuning technique, Child-T…

Cited by 202SourcePDFScholar
2021

Soft-CCD Algorithm for Inverse Kinematics of Soft Continuum Manipulators

IROS 2021poster

To date, soft robots have been increasingly designed and analyzed, especially, Soft Continuum Manipulators (SCMs). Due to dexterous deformability, their Inverse Kinematics (IK) is still difficult to solve. Cyclic Coordinate Descent (CCD) algorithm is one of the classical optimization algorithms to s…

Cited by 6SourceScholar
2020

Is the Skip Connection Provable to Reform the Neural Network Loss Landscape?

IJCAI 2020poster

The residual network is now one of the most effective structures in deep learning, which utilizes the skip connections to “guarantee" the performance will not get worse. However, the non-convexity of the neural network makes it unclear whether the skip connections do provably improve the learning ab…

Cited by 0SourcePDFScholar
2020

Rethinking Skip Connection with Layer Normalization

COLING 2020main

Skip connection is a widely-used technique to improve the performance and the convergence of deep neural networks, which is believed to relieve the difficulty in optimization due to non-linearity by propagating a linear component through the neural network layers. However, from another point of view…

Cited by 0SourcePDFScholar
2019

ShellNet: Efficient Point Cloud Convolutional Neural Networks Using Concentric Shells Statistics

ICCV 2019oral

Deep learning with 3D data has progressed significantly since the introduction of convolutional neural networks that can handle point order ambiguity in point cloud data. While being able to achieve good accuracies in various scene understanding tasks, previous methods often have low training speed…

Cited by 481PDFcodeScholar
2019

Understanding and Improving Layer Normalization

NeurIPS 2019poster

Layer normalization (LayerNorm) is a technique to normalize the distributions of intermediate layers. It enables smoother gradients, faster training, and better generalization accuracy. However, it is still unclear where the effectiveness stems from. In this paper, our main contribution is to take a…