← Search

Ziyang Chen

48 accepted papers

2026

CRAFT: Long-Horizon Cable Routing Algorithm and Low-Friction Caging Gripper

ICRA 2026poster

Cable routing is a common manipulation task in assembly and manufacturing, yet it remains challenging due to the deformable nature of cables and the constraints of cluttered routing environments. In this paper, we present CRAFT: Cable Routing Around Fixtures using Two grippers, a novel hardware plus…

Cited by 0Scholar
2026

EntropyLong: Effective Long-Context Training via Predictive Uncertainty

ICLR 2026poster

Training long-context language models to capture long-range dependencies requires specialized data construction. Current approaches, such as generic text concatenation or heuristic-based variants, frequently fail to guarantee genuine long-range dependencies. We propose \textbf{EntropyLong}, a novel…

Cited by 0SourceScholar
2026

LiteLong: Resource-Efficient Long-Context Data Synthesis for LLMs

AAAI 2026technical

High-quality long-context data is essential for training large language models (LLMs) capable of processing extensive documents, yet existing synthesis approaches using relevance-based aggregation face challenges of computational efficiency. We present LiteLong, a resource-efficient method for synth

Cited by 0SourcePDFScholar
2026

Post-Training LLMs as Better Decision-Making Agents: A Regret-Minimization Approach

ICML 2026poster

Large language models (LLMs) are increasingly deployed as agents for decision-making (DM) in interactive and dynamic environments. However, since they are not originally designed for DM, recent studies show that LLMs struggle in basic online DM settings. We introduce ITERATIVE REGRET-MINIMIZATION FI…

Cited by 0SourceScholar
2026

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

ICLR 2026poster

The development of autonomous agents for complex, long-horizon tasks is a central goal in AI. However, dominant training paradigms face a critical limitation: reinforcement learning (RL) methods that optimize solely for final task success often reinforce flawed or inefficient reasoning paths, a prob…

Cited by 0SourceScholar
2026

RoboVista: Evaluating Vision Language Models for Diverse Robot Applications

RSS 2026poster

Diverse applications for robotics, such as industry and agriculture, require robots to operate across various embodiments, changing visual conditions, and complex planning. Vision–Language Models (VLMs) offer a promising foundation for general-purpose and interpretable robotic reasoning. Aligning VL…

Cited by 0SourceScholar
2026

STITCH 2.0: Extending Augmented Suturing with EKF Needle Estimation and Thread Management

ICRA 2026poster

Suturing is a high-precision task performed at the end of procedures when surgeon fatigue may increase errors, highlighting the need for robot assistance. Previous autonomous suturing works, such as STITCH 1.0 [1], struggle to fully close wounds due to inaccurate needle tracking, thread tangling, an…

2026

SplitFlux: Learning to Decouple Content and Style from a Single Image

CVPR 2026

Disentangling image content and style is essential for customized image generation. Existing SDXL-based methods struggle to achieve high-quality results, while the recently proposed Flux model fails to achieve effective content-style separation due to its underexplored characteristics. To address th

Cited by 0SourcecodeScholar
2025

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding

ACL 2025finding

Vision-language Models (VLMs) have shown remarkable capabilities in advancing general artificial intelligence, yet the irrational encoding of visual positions persists in inhibiting the models’ comprehensive perception performance across different levels of granularity. In this work, we propose Pyra…

2025

Boosting Open-Vocabulary Object Detection Performance via Class-Agnostic Pseudo-Labels and MultiModal Hybrid Knowledge

ICASSP 2025accepted

Open-vocabulary object detection (OVD) is a significant task identifying objects from categories not included in the training set. Our comprehensive analysis reveals two main issues with existing OVD models: poor generalization of localization network to novel categories and poor quality of class em…

Cited by 0SourceScholar
2025

GPS as a Control Signal for Image Generation

CVPR 2025poster

We show that the GPS tags contained in photo metadata provide a useful control signal for image generation. We train GPS-to-image models and use them for tasks that require a fine-grained understanding of how images vary within a city. In particular, we train a diffusion model to generate images con…

Cited by 0SourcePDFScholar
2025

Gradient Alignment Improves Test-Time Adaptation for Medical Image Segmentation

AAAI 2025technical

Although recent years have witnessed significant advancements in medical image segmentation, the pervasive issue of domain shift among medical images from diverse centres hinders the effective deployment of pre-trained models. Many Test-time Adaptation (TTA) methods have been proposed to address thi…

2025

Inference Based Multi-Object Reactive Search in a Partially Known Environment With Temporal Logic Specifications

ICRA 2025

Efficiently searching for multiple objects in a partially known environment, where only the names and locations of landmarks are available, presents significant challenges. Existing search algorithms in the literature fail to fully utilize prior knowledge to improve search efficiency, and exhibit si

Cited by 0SourceScholar
2025

Logic-in-Frames: Dynamic Keyframe Search via Visual Semantic-Logical Verification for Long Video Understanding

NeurIPS 2025poster

Understanding long video content is a complex endeavor that often relies on densely sampled frame captions or end-to-end feature selectors, yet these techniques commonly overlook the logical relationships between textual queries and visual elements. In practice, computational constraints necessitate…

Cited by 0SourceScholar
2025

STITCH 2.0: Extending Augmented Suturing With EKF Needle Estimation and Thread Management

RA-L 2025

Surgical suturing is a high-precision task that impacts patient healing and scarring. Suturing skill varies widely between surgeons, highlighting the need for robot assistance. Previous robot suturing works, such as STITCH 1.0 [1], struggle to fully close wounds due to inaccurate needle tracking and

Cited by 1SourceScholar
2025

Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding

CVPR 2025poster

Recent advancements in multimodal large language models (MLLMs) have significantly improved performance in visual question answering. However, they often suffer from hallucinations. In this work, hallucinations are categorized into two main types: initial hallucinations and snowball hallucinations.…

Cited by 0SourcePDFScholar
2025

Surgical D-Knot: Augmented Dexterity for Tying Double Knots by Monitoring Optical Flow in Monocular Attention Windows

IROS 2025

Knot tying is a fundamental dexterous surgical subtask that is a key step in suturing. One challenge to robot augmentation is limited depth perception due to the small baseline of surgical endoscopic cameras. In this work, we present Surgical D-Knot: an augmented dexterity pipeline combining learned

Cited by 0SourceScholar
2025

ToolExpNet: Optimizing Multi-Tool Selection in LLMs with Similarity and Dependency-Aware Experience Networks

ACL 2025finding

Tool learning enhances Large Language Models’ (LLMs) dynamic interaction with external tools, improving their ability to solve complex problems. However, current empirical methods, which primarily focus on isolated tools learning, still struggle with accurate multi-tool selection due to issues like…

Cited by 0SourcePDFScholar
2025

Video-Guided Foley Sound Generation with Multimodal Controls

CVPR 2025poster

Generating sound effects for videos often requires creating artistic sound effects that diverge significantly from real-life sources and flexible control in the sound design. To address this problem, we introduce *MultiFoley*, a model designed for video-guided sound generation that supports multimod…

Cited by 10SourcePDFScholar
2024

Binding Touch to Everything: Learning Unified Multimodal Tactile Representations

CVPR 2024poster

The ability to associate touch with other modalities has huge implications for humans and computational systems. However multimodal learning with touch remains challenging due to the expensive data collection process and non-standardized sensor outputs. We introduce UniTouch a unified tactile model…

Cited by 53SourcePDFScholar
2024

Continual Self-supervised Learning: Towards Universal Multi-modal Medical Data Representation Learning

CVPR 2024highlight

Self-supervised learning (SSL) is an efficient pre-training method for medical image analysis. However current research is mostly confined to certain modalities consuming considerable time and resources without achieving universality across different modalities. A straightforward solution is combini…

2024

Each Test Image Deserves A Specific Prompt: Continual Test-Time Adaptation for 2D Medical Image Segmentation

CVPR 2024poster

Distribution shift widely exists in medical images acquired from different medical centres and poses a significant obstacle to deploying the pre-trained semantic segmentation model in real-world applications. Test-time adaptation has proven its effectiveness in tackling the cross-domain distribution…

2024

Fast Temporal Logic Mission Planning of Multiple Robots: A Planning Decision Tree Approach

RA-L 2024

This work develops a fast mission planning framework named planning decision tree (PDT), that can handle large-scale multi-robot systems with temporal logic specifications in real time. Specifically, PDT builds a tree incrementally to represent the task progress. The system states are modeled by bot

Cited by 8SourceScholar
2024

MoCha-Stereo: Motif Channel Attention Network for Stereo Matching

CVPR 2024poster

Learning-based stereo matching techniques have made significant progress. However existing methods inevitably lose geometrical structure information during the feature channel generation process resulting in edge detail mismatches. In this paper the Motif Channel Attention Stereo Matching Network (M…

2024

Projection-Based Fast and Safe Policy Optimization for Reinforcement Learning

ICRA 2024poster

While reinforcement learning (RL) attracts increasing research attention, maximizing the return while keeping the agent safe at the same time remains an open problem. Motivated to address this challenge, this work proposes a new Fast and Safe Policy Optimization (FSPO) algorithm, which consists of t…

Cited by 1SourceScholar
2024

Real Acoustic Fields: An Audio-Visual Room Acoustics Dataset and Benchmark

CVPR 2024highlight

We present a new dataset called Real Acoustic Fields (RAF) that captures real acoustic room data from multiple modalities. The dataset includes high-quality and densely captured room impulse response data paired with multi-view images and precise 6DoF pose tracking data for sound emitters and listen…

Cited by 13SourcePDFScholar
2024

Temporal Knowledge Question Answering via Abstract Reasoning Induction

ACL 2024long

In this study, we address the challenge of enhancing temporal knowledge reasoning in Large Language Models (LLMs). LLMs often struggle with this task, leading to the generation of inaccurate or misleading responses. This issue mainly arises from their limited ability to handle evolving factual knowl…

2024

Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?

NeurIPS 2024poster

How can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified metrics, unfair comparisons, and short-term outcome pressure. As a consequence, good performance on standard benchmarks…

2024

Toward a framework integrating augmented reality and virtual fixtures for safer robot-assisted lymphadenectomy

ICRA 2024poster

Lymphadenectomy generally accompanies various oncology surgeries to remove infected cancer cells. However, there are two limitations in robot-assisted lymphadenectomy: 1) lymph nodes are not visible during operation since they are hidden by the superficial fat layer; 2) intra-operative bleeding may…

Cited by 0SourceScholar
2023

A Hierarchical Decoupling Approach for Fast Temporal Logic Motion Planning

ICRA 2023poster

Fast motion planning is of great significance, espe-cially when a timely mission is desired. However, the complexity of motion planning can grow drastically with the increase of environment details and mission complexity. This challenge can be further exacerbated if the tasks are coupled with the de…

Cited by 3SourceScholar
2023

Conditional Generation of Audio From Video via Foley Analogies

CVPR 2023poster

The sound effects that designers add to videos are designed to convey a particular artistic effect and, thus, may be quite different from a scene's true sound. Inspired by the challenges of creating a soundtrack for a video that differs from its true sound, but that nonetheless matches the actions o…

2023

Fast Task Allocation of Heterogeneous Robots With Temporal Logic and Inter-Task Constraints

RA-L 2023

This work develops a fast task allocation framework for heterogeneous multi-robot systems subject to both temporal logic and inter-task constraints. The considered inter-task constraints include unrelated tasks, compatible tasks, and exclusive tasks. To specify such inter-task relationships, we exte

Cited by 24SourceScholar
2023

Large Language Models Meet Harry Potter: A Dataset for Aligning Dialogue Agents with Characters

EMNLP 2023long findings

In recent years, Dialogue-style Large Language Models (LLMs) such as ChatGPT and GPT4 have demonstrated immense potential in constructing open-domain dialogue agents. However, aligning these agents with specific characters or individuals remains a considerable challenge due to the complexities of ch…

Cited by 0SourceScholar
2023

Self-Supervised Video Forensics by Audio-Visual Anomaly Detection

CVPR 2023highlight

Manipulated videos often contain subtle inconsistencies between their visual and audio signals. We propose a video forensics method, based on anomaly detection, that can identify these inconsistencies, and that can be trained solely using real, unlabeled data. We train an autoregressive model to gen…

2023

Sound Localization from Motion: Jointly Learning Sound Direction and Camera Rotation

ICCV 2023poster

The images and sounds that we perceive undergo subtle but geometrically consistent changes as we rotate our heads. In this paper, we use these cues to solve a problem we call Sound Localization from Motion (SLfM): jointly estimating camera rotation and localizing sound sources. We learn to solve the…

Cited by 11PDFcodeScholar
2022

Sound Localization by Self-Supervised Time Delay Estimation

ECCV 2022poster

"Sounds reach one microphone in a stereo pair sooner than the other, resulting in an interaural time delay that conveys their directions. Estimating a sound’s time delay requires finding correspondences between the signals recorded by each microphone. We propose to learn these correspondences throug…