← Search

ziyang zhang

31 accepted papers

2026

Design, Optimization and Experiment of a Detachable Bi-Segment Hybrid Aerial-Underwater Vehicle With Morphable Arms

RA-L 2026

Effective ocean observation requires hybrid aerial-underwater vehicles (HAUVs) to operate efficiently across both air and water. However, such cross-domain missions often lead to significant hydro-aerodynamic conflicts and endurance limitations caused by structural redundancy. This letter proposes N

Cited by 0SourceScholar
2026

Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch

ICML 2026poster

Deterministic inference is increasingly critical for large language model (LLM) applications such as LLM-as-a-judge evaluation, multi-agent systems, and Reinforcement Learning (RL). However, existing LLM serving frameworks exhibit non-deterministic behavior: identical inputs can yield different outp…

Cited by 0SourceScholar
2026

SciTS: Scientific Time Series Understanding and Generation with LLMs

ICLR 2026poster

The scientific reasoning ability of large language models (LLMs) has recently attracted significant attention. Time series, as a fundamental modality in scientific data, presents unique challenges that are often overlooked in current multimodal LLMs, which either encode numerical sequences as text o…

Cited by 0SourceScholar
2025

Bayesian WeakS-to-Strong from Text Classification to Generation

ICLR 2025poster

Advances in large language models raise the question of how alignment techniques will adapt as models become increasingly complex and humans will only be able to supervise them weakly. Weak-to-Strong mimics such a scenario where weak model supervision attempts to harness the full capabilities of a m…

Cited by 0SourcePDFScholar
2025

Debiasing Federated Learning with Correlated Client Participation

ICLR 2025poster

In cross-device federated learning (FL) with millions of mobile clients, only a small subset of clients participate in training in every communication round, and Federated Averaging (FedAvg) is the most popular algorithm in practice. Existing analyses of FedAvg usually assume the participating clie…

Cited by 0SourcePDFScholar
2025

E2B: A Single Modality Point-Based Tracker with Event Cameras

ICRA 2025

High-speed object tracking holds significant relevance across robotic domains, such as drones and autonomous driving. Compared to conventional cameras, event cameras are equipped with the ability to capture object motion information at exceptionally high temporal resolution with relatively low power

Cited by 1SourceScholar
2025

E4: Energy-Efficient DNN Inference for Edge Video Analytics via Early Exiting and DVFS

AAAI 2025technical

Deep neural network (DNN) models are increasingly popular in edge video analytic applications. However, the computeintensive nature of DNN models pose challenges for energyefficient inference on resource-constrained edge devices. Most existing solutions focus on optimizing DNN inference latency and…

Cited by 0SourcePDFScholar
2025

GA-TEB: Goal-Adaptive Framework for Efficient Navigation Based on Goal Lines

ICRA 2025

In crowd navigation, the local goal plays a crucial role in trajectory initialization, optimization, and evaluation. Recognizing that when the global goal is distant, the robot's primary objective is avoiding collisions, making it less critical to pass through the exact local goal point, this work i

Cited by 3SourceScholar
2025

LMR-BENCH: Evaluating LLM Agent’s Ability on Reproducing Language Modeling Research

EMNLP 2025

Large language model (LLM) agents have demonstrated remarkable potential in advancing scientific discovery. However, their capability in the fundamental yet crucial task of reproducing code from research papers, especially in the NLP domain, remains underexplored. This task includes unique complex r

2025

MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual Representations

CVPR 2025poster

Despite significant progress in Vision-Language Pre-training (VLP), current approaches predominantly emphasize feature extraction and cross-modal comprehension, with limited attention to generating or transforming visual content. This gap hinders the model's ability to synthesize coherent and novel…

2025

Real-Time Consistent Monocular Depth Recovery System for Dynamic Environments

IROS 2025

Monocular depth estimation is essential for applications such as autonomous navigation and 3D reconstruction. However, achieving accurate and temporally consistent depth estimation in dynamic environments remains challenging due to scale ambiguity, sensitivity to dynamic objects, and inconsistent de

Cited by 1SourceScholar
2025

SemiETS: Integrating Spatial and Content Consistencies for Semi-Supervised End-to-end Text Spotting

CVPR 2025poster

Most previous scene text spotting methods rely on high-quality manual annotations to achieve promising performance. To reduce their expensive costs, we study semi-supervised text spotting (SSTS) to exploit useful information from unlabeled images. However, directly applying existing semi-supervised…

2025

TCM-Ladder: A Benchmark for Multimodal Question Answering on Traditional Chinese Medicine

NeurIPS 2025poster

Traditional Chinese Medicine (TCM), as an effective alternative medicine, has been receiving increasing attention. In recent years, the rapid development of large language models (LLMs) tailored for TCM has highlighted the urgent need for an objective and comprehensive evaluation framework to assess…

Cited by 0SourcecodeScholar
2024

PARDEN, Can You Repeat That? Defending against Jailbreaks via Repetition

ICML 2024poster

Large language models (LLMs) have shown success in many natural language processing tasks. Despite rigorous safety alignment processes, supposedly safety-aligned LLMs like Llama 2 and Claude 2 are still susceptible to jailbreaks, leading to security risks and abuse of the models. One option to mitig…

2023

Audio-Driven High Definetion and Lip-Synchronized Talking Face Generation Based on Face Reenactment

ICASSP 2023accepted

Generating audio-driven photo-realistic talking face has received intensive attention due to its ability to bring more new human-computer interaction experiences. However, previous works struggled to balance high definition, lip synchronization, and low customization costs, which would degrade the u…

Cited by 0SourceScholar
2023

Deep Directly-Trained Spiking Neural Networks for Object Detection

ICCV 2023poster

Spiking neural networks (SNNs) are brain-inspired energy-efficient models that encode information in spatiotemporal dynamics. Recently, deep SNNs trained directly have shown great success in achieving high performance on classification tasks with very few time steps. However, how to design a directl…

Cited by 102PDFcodeScholar
2023

Dual Memory Aggregation Network for Event-Based Object Detection with Learnable Representation

AAAI 2023technical

Event-based cameras are bio-inspired sensors that capture brightness change of every pixel in an asynchronous manner. Compared with frame-based sensors, event cameras have microsecond-level latency and high dynamic range, hence showing great potential for object detection under high-speed motion and…

2023

Inherent Redundancy in Spiking Neural Networks

ICCV 2023poster

Spiking Neural Networks (SNNs) are well known as a promising energy-efficient alternative to conventional artificial neural networks. Subject to the preconceived impression that SNNs are sparse firing, the analysis and optimization of inherent redundancy in SNNs have been largely overlooked, thus th…

Cited by 26PDFcodeScholar
2023

Test-Time Training-Free Domain Adaptation

ICASSP 2023accepted

Deploying deep learning models to new environments is very challenging. Domain adaptation (DA) is a promising paradigm to solve the problem by collecting and adapting to unlabeled data in new environments. Though research efforts have led to steady performance improvement over the past decade, DA al…

Cited by 0SourceScholar
2022

Audio-Driven Stylized Gesture Generation with Flow-Based Model

ECCV 2022poster

"Generating stylized audio-driven gestures for robots and virtual avatars has attracted increasing considerations recently. Existing methods require style labels (e.g. speaker identities), or complex preprocessing of the data to obtain style control parameters. In this paper, we propose a new end-to…

Cited by 28SourcePDFScholar
2022

Brain-Inspired Multilayer Perceptron With Spiking Neurons

CVPR 2022poster

Recently, Multilayer Perceptron (MLP) becomes the hotspot in the field of computer vision tasks. Without inductive bias, MLPs perform well on feature extraction and achieve amazing results. However, due to the simplicity of their structures, the performance highly depends on the local features commu…

Cited by 38PDFScholar
2022

Deep Reinforcement Learning for Robot Collision Avoidance With Self-State-Attention and Sensor Fusion

RA-L 2022

3D LiDAR sensors can provide 3D point clouds of the environment, and are widely used in automobile navigation; while 2D LiDAR sensors can only provide point cloud in a 2D sweeping plane, and then are only used for navigating robots of small height, e.g., floor mopping robots. In this letter, we prop

Cited by 56SourceScholar
2022

Discrete Time Convolution for Fast Event-Based Stereo

CVPR 2022poster

Inspired by biological retina, dynamical vision sensor transmits events of instantaneous changes of pixel intensity, giving it a series of advantages over traditional frame-based camera, such as high dynamical range, high temporal resolution and low power consumption. However, extracting information…

Cited by 32PDFcodeScholar
2022

Hub-Pathway: Transfer Learning from A Hub of Pre-trained Models

NeurIPS 2022accept

Transfer learning aims to leverage knowledge from pre-trained models to benefit the target task. Prior transfer learning work mainly transfers from a single model. However, with the emergence of deep models pre-trained from different resources, model hubs consisting of diverse models with various ar…

Cited by 8SourcePDFScholar
2022

Meta Talk: Learning To Data-Efficiently Generate Audio-Driven Lip-Synchronized Talking Face With High Definition

ICASSP 2022accepted

Audio-driven talking face, driving talking face by audio, has received considerable attention in multi-modal learning due to its widespread use in virtual reality. However, long-time recording of target high-quality video is needed by most existing audio-driven talking face studies, which significan…

Cited by 0SourceScholar
2022

TimeReplayer: Unlocking the Potential of Event Cameras for Video Interpolation

CVPR 2022poster

Recording fast motion in a high FPS (frame-per-second) requires expensive high-speed cameras. As an alternative, interpolating low-FPS videos from commodity cameras has attracted significant attention. If only low-FPS videos are available, motion assumptions (linear or quadratic) are necessary to in…

Cited by 37PDFScholar
2022

Video Interpolation by Event-Driven Anisotropic Adjustment of Optical Flow

ECCV 2022poster

"Video frame interpolation is a challenging task due to the ever-changing real-world scene. Previous methods often calculate the bi-directional optical flows and then predict the intermediate optical flows under the linear motion assumptions, leading to isotropic intermediate flow generation. Follow…

Cited by 15SourcePDFScholar