← Search

Ziyi Zhang

25 accepted papers

2026

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs

ICML 2026poster

Recent advances in Omni-Multimodal Large Language Models (Omni-MLLMs) have enabled strong integration of vision, audio, and language. However, their audio-visual intelligence (AVI) remains insufficiently evaluated due to the lack of systematic and comprehensive benchmarks. We introduce AVI-Bench, a …

Cited by 0SourceScholar
2026

ConFlux: Multivariate Time Series in Flux, One Unified Forecast in Confluence

ICML 2026oral

Real-world multivariate time series are inherently in flux: different variables evolve asynchronously and interact in complex, time-varying ways, yet accurate forecasting requires these dispersed signals to converge into a single unified prediction. This structural mismatch between dynamic, heteroge…

Cited by 0SourceScholar
2026

FrontierCS: Evolving Challenges for Evolving Intelligence

ICML 2026poster

We introduce FrontierCS, a benchmark of 240 open-ended problems across diverse areas of computer science, designed and reviewed by experts, including CS PhDs and top-tier competitive programming participants and problem setters. Unlike existing benchmarks that focus on tasks with known optimal solut…

Cited by 0SourceScholar
2026

Goal-Oriented Control Strategies for Soft Growing Robots

ICRA 2026poster

Soft growing robots, as highly mobile pneumatic membrane robots, are limited in control performance due to their soft structure and nonlinear mechanical properties, especially under dynamic conditions. Therefore, developing reliable control strategies for the robot is essential. This study proposes …

Cited by 0SourceScholar
2026

Optimization Method for Surrogate Function in Spiking Neural Networks Based on Membrane Potential Distribution

AAAI 2026technical

Spiking Neural Networks (SNNs) offer promising energy efficiency and temporal sparsity for edge intelligence, but their training remains difficult due to gradient mismatch, membrane potential drift, and discretization errors. In this paper, we propose a membrane potential-guided surrogate optimizati

Cited by 0SourcePDFScholar
2026

Refining Few-Step Text-to-Multiview Diffusion via Reinforcement Learning

CVPR 2026

Text-to-multiview (T2MV) diffusion models have shown great promise in generating multiple views of a scene from a single text prompt. While few-step backbones enable real-time T2MV generation, they often compromise key aspects of generation quality, such as per-view fidelity and cross-view consisten

Cited by 0SourcecodeScholar
2026

Stealing Split Learning Bottom Models by Recovering Embedding Geometry

CVPR 2026

Vertical federated learning (VFL) trains models by splitting computation across clients and a server that only exchange intermediate embeddings. Recent work shows that a server even if honest-but-curious can steal a client's bottom model by querying the system and regressing on the returned embeddin

Cited by 0SourceScholar
2025

A Unified Supervised and Unsupervised Dialogue Topic Segmentation Framework Based on Utterance Pair Modeling

NAACL 2025long

The Dialogue Topic Segmentation task aims to divide a dialogue into different topic paragraphs in order to better understand the structure and content of the dialogue. Due to the short sentences, serious references and non-standard language in the dialogue, it is difficult to determine the boundarie…

Cited by 0SourcePDFScholar
2025

Are We in the AI-Generated Text World Already? Quantifying and Monitoring AIGT on Social Media

ACL 2025long

Social media platforms are experiencing a growing presence of AI-Generated Texts (AIGTs). However, the misuse of AIGTs could have profound implications for public opinion, such as spreading misinformation and manipulating narratives. Despite its importance, it remains unclear how prevalent AIGTs are…

2025

FC-Attack: Jailbreaking Multimodal Large Language Models via Auto-Generated Flowcharts

EMNLP 2025

Multimodal Large Language Models (MLLMs) have become powerful and widely adopted in some practical applications.However, recent research has revealed their vulnerability to multimodal jailbreak attacks, whereby the model can be induced to generate harmful content, leading to safety risks. Although m

2025

Goal-Oriented Control Strategies for Soft Growing Robots

RA-L 2025

Soft growing robots, as highly mobile pneumatic membrane robots, are limited in control performance due to their soft structure and nonlinear mechanical properties, especially under dynamic conditions. Therefore, developing reliable control strategies for the robot is essential. This study proposes

Cited by 1SourceScholar
2025

Learning to Stabilize Unknown LTI Systems on a Single Trajectory under Stochastic Noise

UAI 2025

We study the problem of learning to stabilize unknown noisy Linear Time-Invariant (LTI) systems on a single trajectory. The state-of-the-art guarantees that the system is stabilized before the system state reaches $2^{O(k \log n)}$ in $L^2$-norm, where $n$ is the state dimension, and $k$ is the dime

Cited by 0SourcePDFScholar
2025

MAITFuse: Multi-Dimension Adaptive Interaction Transform Network For Infrared-visible Image Fusion

ICASSP 2025accepted

In recent years, Transformers have achieved significant success in image fusion. These methods utilize self-attention mechanism across different spatial or channel dimensions and have demonstrated impressive performance. However, existing methods only optimize along a single dimension and struggle t…

Cited by 0SourceScholar
2025

Stabilizing LTI Systems under Partial Observability: Sample Complexity and Fundamental Limits

NeurIPS 2025poster

We study the problem of stabilizing an unknown partially observable linear time-invariant (LTI) system. For fully observable systems, leveraging an unstable/stable subspace decomposition approach, state-of-art sample complexity is independent from system dimension $n$ and only scales with respect to…

Cited by 0SourceScholar
2025

TMANet: Triple Multi-Scale Attention based Network with Boundary Association Loss for Superpixel Segmentation

ICASSP 2025accepted

Superpixel segmentation with deep learning has been proposed in recent years and is widely employed to reduce the input image primitives for subsequent computer vision tasks. In this paper, we propose a Triple Multi-Scale Attention based Network (TMANet) for superpixel segmentation. First, aiming to…

Cited by 0SourceScholar
2024

Confronting Reward Overoptimization for Diffusion Models: A Perspective of Inductive and Primacy Biases

ICML 2024poster

Bridging the gap between diffusion models and human preferences is crucial for their integration into practical generative workflows. While optimizing downstream reward models has emerged as a promising alignment strategy, concerns arise regarding the risk of excessive optimization with learned rewa…

2024

DEL: Discrete Element Learner for Learning 3D Particle Dynamics with Neural Rendering

NeurIPS 2024poster

Learning-based simulators show great potential for simulating particle dynamics when 3D groundtruth is available, but per-particle correspondences are not always accessible. The development of neural rendering presents a new solution to this field to learn 3D dynamics from 2D images by inverse rende…

Cited by 0SourcePDFScholar
2024

Diving into Underwater: Segment Anything Model Guided Underwater Salient Instance Segmentation and A Large-scale Dataset

ICML 2024poster

With the breakthrough of large models, Segment Anything Model (SAM) and its extensions have been attempted to apply in diverse tasks of computer vision. Underwater salient instance segmentation is a foundational and vital step for various underwater vision tasks, which often suffer from low segmenta…

2024

EvGGS: A Collaborative Learning Framework for Event-based Generalizable Gaussian Splatting

ICML 2024poster

Event cameras offer promising advantages such as high dynamic range and low latency, making them well-suited for challenging lighting conditions and fast-moving scenarios. However, reconstructing 3D scenes from raw event streams is difficult because event data is sparse and does not carry absolute c…

2024

Learning Robust Generalizable Radiance Field with Visibility and Feature Augmented Point Representation

ICLR 2024poster

This paper introduces a novel paradigm for the generalizable neural radiance field (NeRF). Previous generic NeRFs combine multiview stereo techniques with image-based neural rendering, yielding impressive results, while suffering from three issues. First, occlusions often result in inconsistent feat…

Cited by 6SourcePDFScholar
2023

RankMatch: Fostering Confidence and Consistency in Learning with Noisy Labels

ICCV 2023poster

Learning with noisy labels (LNL) is one of the most important and challenging problems in weakly-supervised learning. Recent advances adopt the sample selection strategy to mitigate the interference of noisy labels and use small-loss criteria to select clean samples. However, the one-dimensional los…

Cited by 14PDFScholar
2022

Dite-HRNet: Dynamic Lightweight High-Resolution Network for Human Pose Estimation

IJCAI 2022poster

A high-resolution network exhibits remarkable capability in extracting multi-scale features for human pose estimation, but fails to capture long-range interactions between joints and has high computational complexity. To address these problems, we present a Dynamic lightweight High-Resolution Networ…

2022

Divide and Contrast: Source-free Domain Adaptation via Adaptive Contrastive Learning

NeurIPS 2022accept

We investigate a practical domain adaptation task, called source-free domain adaptation (SFUDA), where the source pretrained model is adapted to the target domain without access to the source data. Existing techniques mainly leverage self-supervised pseudo-labeling to achieve class-wise global align…