← Search

Chenyu Zhang

19 accepted papers

2026

Kinematics-Aware Diffusion Policy With Consistent 3D Observation and Action Space for Whole-Arm Robotic Manipulation

RA-L 2026

Full-configuration control of robotic manipulators with awareness of whole-arm kinematics is crucial for many manipulation scenarios involving body collision avoidance or body-object interactions, making it insufficient to consider only the end-effector poses in policy learning. The typical approach

Cited by 1SourceScholar
2026

Probing Newtonian Mechanics in Video Generative Models with Real Physical Systems

ICML 2026poster

Recent advances in image and video generation raise hopes that these models possess world modeling capabilities—the ability to generate realistic, physically plausible videos. This could revolutionize applications in robotics, autonomous driving, and scientific simulation. However, before treating t…

Cited by 0SourceScholar
2026

Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoning

AAAI 2026technical

Text-to-Image (T2I) models typically deploy safety mechanisms to prevent the generation of sensitive images. Unfortunately, recent jailbreaking attack methods manually design instructions for the LLM to generate adversarial prompts, which effectively exposing safety vulnerabilities of T2I models.

Cited by 0SourcePDFScholar
2026

RobotArena $\infty$: Unlimited Robot Benchmarking via Real-to-Sim Translation

ICLR 2026poster

The pursuit of robot generalists—instructable agents capable of performing diverse tasks across diverse environments—demands rigorous and scalable evaluation. Yet real-world testing of robot policies remains fundamentally constrained: it is labor-intensive, slow, unsafe at scale, and difficult to re…

Cited by 0SourcecodeScholar
2026

Scaling Equitable Reflection Assessment in Education via Large Language Models and Role-Based Feedback Agents

AAAI 2026technical

Formative feedback is widely recognized as one of the most effective drivers of student learning, yet it remains difficult to implement equitably at scale. In large or low-resource courses, instructors often lack the time, staffing, and bandwidth required to review and respond to every student refle

Cited by 0SourcePDFScholar
2026

T2I-RiskyPrompt: A Benchmark for Safety Evaluation, Attack, and Defense on Text-to-Image Model

AAAI 2026technical

Using risky text prompts, such as pornography and violent prompts, to test the safety of text-to-image (T2I) models is a critical task. However, existing risky prompt datasets are limited in three key areas: 1) limited risky categories, 2) coarse-grained annotation, and 3) low effectiveness. To addr

Cited by 0SourcePDFScholar
2025

Boosting the Transferability of Adversarial Examples via Local Mixup and Adaptive Step Size

ICASSP 2025accepted

Adversarial examples are one critical security threat to various visual applications, where injected human-imperceptible perturbations confuse the output. Generating transferable adversarial examples in the black-box setting is crucial but challenging in practice. Existing input-diversity-based meth…

Cited by 0SourceScholar
2025

SCAP: Transductive Test-Time Adaptation via Supportive Clique-based Attribute Prompting

CVPR 2025poster

Vision-language models (VLMs) encounter considerable challenges when adapting to domain shifts stemming from changes in data distribution. Test-time adaptation (TTA) has emerged as a promising approach to enhance VLM performance under such conditions. In practice, test data often arrives in batches,…

2025

Stochastic Semi-Gradient Descent for Learning Mean Field Games with Population-Aware Function Approximation

ICLR 2025poster

Mean field games (MFGs) model interactions in large-population multi-agent systems through population distributions. Traditional learning methods for MFGs are based on fixed-point iteration (FPI), where policy updates and induced population distributions are computed separately and sequentially. How…

Cited by 1SourcePDFScholar
2025

TRCE: Towards Reliable Malicious Concept Erasure in Text-to-Image Diffusion Models

ICCV 2025poster

Recent advances in text-to-image diffusion models enable photorealistic image generation, but they also risk producing malicious content, such as NSFW images. To mitigate risk, concept erasure methods are studied to facilitate the model to unlearn specific concepts. However, current studies struggle…

2024

Finite-Time Analysis of On-Policy Heterogeneous Federated Reinforcement Learning

ICLR 2024poster

Federated reinforcement learning (FRL) has emerged as a promising paradigm for reducing the sample complexity of reinforcement learning tasks by exploiting information from different agents. However, when each agent interacts with a potentially different environment, little to nothing is known theor…

Cited by 20SourcePDFScholar
2024

Graphon Mean Field Games with a Representative Player: Analysis and Learning Algorithm

ICML 2024poster

We propose a discrete time graphon game formulation on continuous state and action spaces using a representative player to study stochastic games with heterogeneous interaction among agents. This formulation admits both conceptual and mathematical advantages, compared to a widely adopted formulation…

Cited by 4SourcePDFScholar
2023

Geo-Seq2seq: Twitter User Geolocation on Noisy Data through Sequence to Sequence Learning

ACL 2023findings

Location information can support social media analyses by providing geographic context. Some of the most accurate and popular Twitter geolocation systems rely on rule-based methods that examine the user-provided profile location, which fail to handle informal or noisy location names. We propose Geo-…

Cited by 4SourcePDFScholar
2023

Skeleton Graph-Based Ultrasound-CT Non-Rigid Registration

RA-L 2023

Autonomous ultrasound (US) scanning has attracted increased attention, and it has been seen as a potential solution to overcome the limitations of conventional US examinations, such as inter-operator variations. However, it is still challenging to autonomously and accurately transfer a planned scan

Cited by 14SourceScholar
2023

Underwater and Surface Aquatic Locomotion of Soft Biomimetic Robot Based on Bending Rolled Dielectric Elastomer Actuators

IROS 2023poster

All-around, real-time navigation and sensing across the water environments by miniature soft robotics are promising, for their merits of small size, high agility and good compliance to the unstructured surroundings. In this paper, we propose and demonstrate a mantas-like soft aquatic robot which pro…

Cited by 5SourceScholar
2022

SwapMix: Diagnosing and Regularizing the Over-Reliance on Visual Context in Visual Question Answering

CVPR 2022poster

While Visual Question Answering (VQA) has progressed rapidly, previous works raise concerns about robustness of current VQA models. In this work, we study the robustness of VQA models from a novel perspective: visual context. We suggest that the models over-rely on the visual context, i.e., irreleva…

Cited by 65PDFcodeScholar
2022

Visual Commonsense in Pretrained Unimodal and Multimodal Models

NAACL 2022long

Our commonsense knowledge about objects includes their typical visual attributes; we know that bananas are typically yellow or green, and not purple. Text and image corpora, being subject to reporting bias, represent this world-knowledge to varying degrees of faithfulness. In this paper, we investig…