← Search

Yibo Zhang

17 accepted papers

2026

Hidden in the Noise: Unveiling Backdoors in Audio LLMs Alignment Through Latent Acoustic Pattern Triggers

AAAI 2026technical

As Audio Large Language Models (ALLMs) emerge as powerful tools for speech processing, their safety implications demand urgent attention. While considerable research has explored textual and vision safety, audio’s distinct characteristics present significant challenges. This paper first investigates

Cited by 0SourcePDFScholar
2026

Optimal Design of Integrated Aerial Platforms with Passive Joints

ICRA 2026poster

The Integrated Aerial Platform (IAP) uses multiple quadrotor sub-vehicles, acting as independent thrust generators, connected to a central platform via passive joints. This setup allows the sub-vehicles to collectively apply forces and torques to the central platform, achieving full six-degree-of-fr…

Cited by 0SourceScholar
2025

DecoupledGaussian: Object-Scene Decoupling for Physics-Based Interaction

CVPR 2025poster

We present DecoupledGaussian, a novel system that decouples static objects from their contacted surfaces captured in-the-wild videos, a key prerequisite for realistic Newtonian-based physical simulations. Unlike prior methods focused on synthetic data or elastic jittering along the contact surface,…

2025

Diff3DS: Generating View-Consistent 3D Sketch via Differentiable Curve Rendering

ICLR 2025poster

3D sketches are widely used for visually representing the 3D shape and structure of objects or scenes. However, the creation of 3D sketch often requires users to possess professional artistic skills. Existing research efforts primarily focus on enhancing the ability of interactive sketch generation…

2025

FreeAlign: Superior Text-Image Alignment by Modulating Prompt Attention

ICASSP 2025accepted

In recent years, Text-to-Image (T2I) models have made remarkable advancements, yet accurate accurate association of attributes remains a key challenge. This paper presents FreeAlign, a novel training-free framework designed to enhance attribute alignment in T2I generation. By modulating attention an…

Cited by 0SourceScholar
2025

MIO: A Foundation Model on Multimodal Tokens

EMNLP 2025

In this paper, we introduce MIO, a novel foundation model built on multimodal tokens, capable of understanding and generating speech, text, images, and videos in an end-to-end, autoregressive manner. While the emergence of large language models (LLMs) and multimodal large language models (MM-LLMs) p

2025

Resilient Test-Time Adaptation by Mitigating Batch-Normalization Overfitting

ICASSP 2025accepted

Test-time domain adaptation adjusts a source domain model to accommodate previously unseen domain shifts in a target domain during inference. In real-world scenarios, domain shifts continually evolve, and test data are often non-independent and identically distributed (non-i.i.d.). Existing methods…

Cited by 0SourceScholar
2025

Stable Control Visual AutoRegressive Model: Precise and Efficient Image Generation via Scale Alignment

ICASSP 2025accepted

Although diffusion models advance condition-based visual generation, they suffer from speed and cost issues, unlike faster AutoRegressive methods that are limited in performance. To address these, we introduce the Stable Control Visual AutoRegressive Model (SCVAR). SCVAR ensures stable control by al…

Cited by 0SourceScholar
2025

Test-Time Adaptation on Noisy Data via Model-Pruning-Based Filtering and Flatness-Aware Entropy Minimization

AAAI 2025technical

Test-time adaptation (TTA) deals with domain shifts during inference by training models based on only unlabeled test samples. Test samples may include noisy samples, which degrade domain adaptation. Existing methods rely on the model's output prediction to detect and filter noisy samples, and furthe…

2024

3D-SceneDreamer: Text-Driven 3D-Consistent Scene Generation

CVPR 2024poster

Text-driven 3D scene generation techniques have made rapid progress in recent years. Their success is mainly attributed to using existing generative models to iteratively perform image warping and inpainting to generate 3D scenes. However these methods heavily rely on the outputs of existing models…

Cited by 8SourcePDFScholar
2024

Can Public Large Language Models Help Private Cross-device Federated Learning?

NAACL 2024findings

We study (differentially) private federated learning (FL) of language models. The language models in cross-device FL are relatively small, which can be trained with meaningful formal user-level differential privacy (DP) guarantees when massive parallelism in training is enabled by the participation…

Cited by 45SourcePDFScholar
2024

From Bottom to Top: Extending the Potential of Parameter Efficient Fine-Tuning

EMNLP 2024main

With the proliferation of large language models, Parameter Efficient Fine-Tuning (PEFT) method, which freeze pre-trained parameters and only fine-tune a few task-specific parameters, are playing an increasingly important role. However, previous work primarily applied uniform operations across all la…

Cited by 0SourcePDFScholar
2024

PositionID: LLMs can Control Lengths, Copy and Paste with Explicit Positional Awareness

EMNLP 2024finding

Large Language Models (LLMs) demonstrate impressive capabilities across various domains, including role-playing, creative writing, mathematical reasoning, and coding. Despite these advancements, LLMs still encounter challenges with length control, frequently failing to adhere to specific length cons…

2023

DialogMI: A Dialogue Model Based on Enhancing Dialogue Mutual Information

ICASSP 2023accepted

Most of the open-domain dialogue models tend to perform insufficiently in generating informative response. The possible reason is that they lack the capability of enhancing the mutual information between generated responses and dialogue history. To address this issue, we present a novel task of the…

Cited by 0SourceScholar
2023

TextObfuscator: Making Pre-trained Language Model a Privacy Protector via Obfuscating Word Representations

ACL 2023findings

In real-world applications, pre-trained language models are typically deployed on the cloud, allowing clients to upload data and perform compute-intensive inference remotely. To avoid sharing sensitive data directly with service providers, clients can upload numerical representations rather than pla…

2018

Design and Implementation of a Novel Aerial Manipulator with Tandem Ducted Fans

IROS 2018poster

This paper proposes a novel aerial manipulator with tandem ducted fans, which takes both trafficability and effective loading into account. The aerial manipulator is particularly suitable for grasping in complex and narrow environment, in which traditional multi-rotor and helicopter would be inacces…

Cited by 10SourceScholar