← Search

Yi Yu

69 accepted papers

2026

All Vehicles Can Lie: Efficient Adversarial Defense in Fully Untrusted-Vehicle Collaborative Perception via Pseudo-Random Bayesian Inference

CVPR 2026

Collaborative perception (CP) enables multiple vehicles to augment their individual perception capacities through the exchange of feature-level sensory data. However, this fusion mechanism is inherently vulnerable to adversarial attacks, especially in fully untrusted-vehicle environments. Existing d

Cited by 0SourceScholar
2026

Every Little Bit Helps: Exploring Better Utilization of Unlabeled Data for Semi-supervised Singing Melody Extraction Using Multi-bands Diffusion Model

AAAI 2026technical

Semi-supervised singing melody extraction (SSME) is one of the key tasks in the field of music information retrieval (MIR). Recently, several SSME methods have been proposed and achieved remarkable successes. However, existing methods are still facing two critical issues: firstly, there is a lack of

Cited by 0SourcePDFScholar
2026

From Pretrain to Pain: Adversarial Vulnerability of Video Foundation Models Without Task Knowledge

AAAI 2026technical

Large-scale Video Foundation Models (VFMs) have significantly advanced various video-related tasks, either through task-specific models or Multi-modal Large Language Models (MLLMs). However, the open accessibility of VFMs also introduces critical security risks, as adversaries can exploit full knowl

Cited by 0SourcePDFScholar
2026

Origami-Inspired 3-DoF Robotic Joint Design With Pneumatic Actuators

RA-L 2026

Soft pneumatic joints, as the core components of soft robotic manipulators, often face a trade-off between multiple degrees of freedom (DoF) mobility and load capacity. Inspired by the Yoshimura origami pattern, this paper presents a rigid–soft hybrid pneumatic joint. Through spatially ordered place

Cited by 0SourceScholar
2026

Partial Weakly-Supervised Oriented Object Detection

CVPR 2026

The growing demand for oriented object detection (OOD) across various domains has driven significant research in this area. However, the high cost of dataset annotation remains a major concern. Current mainstream OOD algorithms can be mainly categorized into three types: (1) fully supervised methods

Cited by 0SourcecodeScholar
2026

Point2RBox-v3: Self-Bootstrapping from Point Annotations via Integrated Pseudo-Label Refinement and Utilization

ICLR 2026poster

Driven by the growing need for Oriented Object Detection (OOD), learning from point annotations under a weakly-supervised framework has emerged as a promising alternative to costly and laborious manual labeling. In this paper, we discuss two deficiencies in existing point-supervised methods: ineffic…

Cited by 0SourcecodeScholar
2026

Reallocating Attention Across Layers to Reduce Multimodal Hallucination

CVPR 2026

Multimodal large reasoning models (MLRMs) often suffer from hallucinations that stem not only from insufficient visual grounding but also from imbalanced allocation between perception and reasoning processes. Building upon recent interpretability findings suggesting a staged division of attention ac

Cited by 0SourcecodeScholar
2026

SAVER: Mitigating Hallucinations in Large Vision-Language Models via Style-Aware Visual Early Revision

AAAI 2026technical

Large Vision-Language Models (LVLMs) recently achieve significant breakthroughs in understanding complex visual-textual contexts. However, hallucination issues still limit their real-world applicability. Although previous mitigation methods effectively reduce hallucinations in photographic images, t

Cited by 0SourcePDFScholar
2026

TEAR: Temporal-aware Automated Red-teaming for Text-to-Video Models

CVPR 2026

Text-to-Video (T2V) models are capable of synthesizing high-quality, temporally coherent dynamic video content, but the diverse generation also inherently introduces critical safety challenges. Existing safety evaluation methods, which focus on static image and text generation, are insufficient to c

Cited by 0SourceScholar
2026

Time Is All It Takes: Spike-Retiming Attacks on Event-Driven Spiking Neural Networks

ICLR 2026poster

Spiking neural networks (SNNs) compute with discrete spikes and exploit temporal structure, yet most adversarial attacks change intensities or event counts instead of timing. We study a timing-only adversary that retimes existing spikes while preserving spike counts and amplitudes in event-driven SN…

Cited by 0SourcecodeScholar
2026

When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action Models

CVPR 2026

Vision-Language-Action (VLA) models are vulnerable to adversarial attacks, yet universal and transferable attacks remain underexplored, as most existing patches overfit to a single model and fail in black-box settings. To address this gap, we present a systematic study of universal, transferable adv

Cited by 0SourcecodeScholar
2025

A Mamba-based Network for Semi-supervised Singing Melody Extraction Using Confidence Binary Regularization

ICASSP 2025accepted

Singing melody extraction (SME) is a key task in the field of music information retrieval. However, existing methods are facing several limitations: firstly, prior models use transformers to capture the contextual dependencies, which requires quadratic computation resulting in low efficiency in the…

Cited by 0SourceScholar
2025

Backdoor Attacks Against No-Reference Image Quality Assessment Models via a Scalable Trigger

AAAI 2025technical

No-Reference Image Quality Assessment (NR-IQA), responsible for assessing the quality of a single input image without using any reference, plays a critical role in evaluating and optimizing computer vision systems, e.g., low-light enhancement. Recent research indicates that NR-IQA models are suscep…

2025

Enduring, Efficient and Robust Trajectory Prediction Attack in Autonomous Driving via Optimization-Driven Multi-Frame Perturbation Framework

CVPR 2025highlight

Trajectory prediction plays a crucial role in autonomous driving systems, and exploring its vulnerability has garnered widespread attention. However, existing trajectory prediction attack methods often rely on single-point attacks to make efficient perturbations. This limits their applications in re…

2025

Enhancing Video-Text Matching via Sparse Stratified Sampling

ICASSP 2025accepted

Video-text matching is a critical task in multimedia retrieval, but traditional methods often fail to capture the diversity and depth of video content due to inefficient and inaccurate frame sampling. We propose a novel sparse stratified sampling technique that can substantially improve the video-te…

Cited by 0SourceScholar
2025

EvoBench: Towards Real-world LLM-Generated Text Detection Benchmarking for Evolving Large Language Models

ACL 2025finding

With the widespread of Large Language Models (LLMs), there has been an increasing need to detect LLM-generated texts, prompting extensive research in this area. However, existing detection methods mainly evaluate on static benchmarks, which neglect the evolving nature of LLMs. Relying on existing st…

Cited by 0SourcePDFScholar
2025

MTL-UE: Learning to Learn Nothing for Multi-Task Learning

ICML 2025poster

Most existing unlearnable strategies focus on preventing unauthorized users from training single-task learning (STL) models with personal data. Nevertheless, the paradigm has recently shifted towards multi-task data and multi-task learning (MTL), targeting generalist and foundation models that can h…

2025

One Example Shown, Many Concepts Known! Counterexample-Driven Conceptual Reasoning in Mathematical LLMs

ICML 2025poster

Leveraging mathematical Large Language Models (LLMs) for proof generation is a fundamental topic in LLMs research. We argue that the ability of current LLMs to prove statements largely depends on whether they have encountered the relevant proof process during training. This reliance limits their dee…

Cited by 3SourcePDFScholar
2025

Point2RBox-v2: Rethinking Point-supervised Oriented Object Detection with Spatial Layout Among Instances

CVPR 2025poster

With the rapidly increasing demand for oriented object detection (OOD), recent research involving weakly-supervised detectors for learning OOD from point annotations has gained great attention. In this paper, we rethink this challenging task setting with the layout among instances and present Point2…

2025

PointOBB-v2: Towards Simpler, Faster, and Stronger Single Point Supervised Oriented Object Detection

ICLR 2025poster

Single point supervised oriented object detection has gained attention and made initial progress within the community. Diverse from those approaches relying on one-shot samples or powerful pretrained models (e.g. SAM), PointOBB has shown promise due to its prior-free feature. In this paper, we propo…

Cited by 22SourcePDFScholar
2025

Temporal Unlearnable Examples: Preventing Personal Video Data from Unauthorized Exploitation by Object Tracking

ICCV 2025poster

With the rise of social media, vast amounts of user-uploaded videos (e.g., YouTube) are utilized as training data for Visual Object Tracking (VOT). However, the VOT community has largely overlooked video data-privacy issues, as many private videos have been collected and used for training commercial…

Cited by 0SourcePDFScholar
2025

Theoretical Insights in Model Inversion Robustness and Conditional Entropy Maximization for Collaborative Inference Systems

CVPR 2025highlight

By locally encoding raw data into intermediate features, collaborative inference enables end users to leverage powerful deep learning models without exposure of sensitive raw data to cloud servers. However, recent studies have revealed that these intermediate features may not sufficiently preserve p…

2025

Visual Entity-Centric Prompting for Knowledge Retrieval in Knowledge-based VQA

ICASSP 2025accepted

External knowledge provides critical clues for knowledge-based visual question answering (KB-VQA), while the implicit knowledge in images is difficult to capture in order to construct effective queries for knowledge bases. To this end, we propose a visual entity-centric prompting for knowledge retri…

Cited by 0SourceScholar
2025

X-Boundary: Establishing Exact Safety Boundary to Shield LLMs from Jailbreak Attacks without Compromising Usability

EMNLP 2025

With the widespread application of large language models (LLMs) across various domains, techniques for enhancing their security have progressed rapidly. In this paper, we reveal that although existing defense methods can improve the robustness of LLMs against jailbreaks, they compromise usability, i

2024

A Steered Response Power Approach with Bilinear Prediction-Based Trade-Off Prewhitening for Speaker Localization

ICASSP 2024accepted

This paper studies the problem of acoustic source localization in room environments. It presents an improved steered response power (SRP) approach with low-complexity and trade-off prewhitening. This method consists of two steps. In the first one, the linear predictor that is used to model the speec…

Cited by 0SourceScholar
2024

Benchmarking Adversarial Robustness of Image Shadow Removal with Shadow-Adaptive Attacks

ICASSP 2024accepted

Shadow removal is a task aimed at erasing regional shadows present in images and reinstating visually pleasing natural scenes with consistent illumination. While recent deep learning techniques have demonstrated impressive performance in image shadow removal, their robustness against adversarial att…

Cited by 0SourceScholar
2024

Mitigating the Curse of Dimensionality for Certified Robustness via Dual Randomized Smoothing

ICLR 2024poster

Randomized Smoothing (RS) has been proven a promising method for endowing an arbitrary image classifier with certified robustness. However, the substantial uncertainty inherent in the high-dimensional isotropic Gaussian noise imposes the curse of dimensionality on RS. Specifically, the upper bound o…

2024

Point2RBox: Combine Knowledge from Synthetic Visual Patterns for End-to-end Oriented Object Detection with Single Point Supervision

CVPR 2024poster

With the rapidly increasing demand for oriented object detection (OOD) recent research involving weakly-supervised detectors for learning rotated box (RBox) from the horizontal box (HBox) has attracted more and more attention. In this paper we explore a more challenging yet label-efficient setting n…

Cited by 14SourcePDFScholar
2024

PointOBB: Learning Oriented Object Detection via Single Point Supervision

CVPR 2024poster

Single point-supervised object detection is gaining attention due to its cost-effectiveness. However existing approaches focus on generating horizontal bounding boxes (HBBs) while ignoring oriented bounding boxes (OBBs) commonly used for objects in aerial images. This paper proposes PointOBB the fir…

2024

Progressive Divide-and-Conquer via Subsampling Decomposition for Accelerated MRI

CVPR 2024highlight

Deep unfolding networks (DUN) have emerged as a popular iterative framework for accelerated magnetic resonance imaging (MRI) reconstruction. However conventional DUN aims to reconstruct all the missing information within the entire space in each iteration. Thus it could be challenging when dealing w…

2024

Purify Unlearnable Examples via Rate-Constrained Variational Autoencoders

ICML 2024poster

Unlearnable examples (UEs) seek to maximize testing error by making subtle modifications to training examples that are correctly labeled. Defenses against these poisoning attacks can be categorized based on whether specific interventions are adopted during training. The first approach is training-ti…

2024

Scalable Motion Style Transfer with Constrained Diffusion Generation

AAAI 2024technical

Current training of motion style transfer systems relies on consistency losses across style domains to preserve contents, hindering its scalable application to a large number of domains and private data. Recent image transfer works show the potential of independent training on each domain by leverag…

2024

Semantic Enrichment for Video Question Answering with Gated Graph Neural Networks

ICASSP 2024accepted

Video Question Answering (VideoQA) is a complex task that requires a deep understanding of a video to accurately answer questions. Existing methods often struggle to effectively integrate the visual and language-based semantic information, subsequently leading to an incomplete understanding of video…

Cited by 0SourceScholar
2024

Towards Physical World Backdoor Attacks against Skeleton Action Recognition

ECCV 2024poster

"Skeleton Action Recognition (SAR) has attracted significant interest for its efficient representation of the human skeletal structure. Despite its advancements, recent studies have raised security concerns in SAR models, particularly their vulnerability to adversarial attacks. However, such strateg…

Cited by 3SourcePDFScholar
2024

Transferable Adversarial Attacks on SAM and Its Downstream Models

NeurIPS 2024poster

The utilization of large foundational models has a dilemma: while fine-tuning downstream tasks from them holds promise for making use of the well-generalized knowledge in practical applications, their open accessibility also poses threats of adverse usage. This paper, for the first time, explores th…

2023

A Frequency-Domain Recursive Least-Squares Adaptive Filtering Algorithm Based On A Kronecker Product Decomposition

ICASSP 2023accepted

This paper proposes a frequency-domain recursive least-squares (RLS) adaptive filtering algorithm for identifying time-varying acoustic systems in noisy environments. The Kronecker product (KP) is employed to decompose the model filter of the acoustic channel impulse response into two sets of short…

Cited by 0SourceScholar
2023

Backdoor Attacks Against Deep Image Compression via Adaptive Frequency Trigger

CVPR 2023poster

Recent deep-learning-based compression methods have achieved superior performance compared with traditional approaches. However, deep learning models have proven to be vulnerable to backdoor attacks, where some specific trigger patterns added to the input can lead to malicious behavior of the models…

Cited by 57SourcePDFScholar
2023

Change point detection and inference in multivariate non-parametric models under mixing conditions

NeurIPS 2023poster

This paper addresses the problem of localizing and inferring multiple change points, in non-parametric multivariate time series settings. Specifically, we consider a multivariate time series with potentially short-range dependence, whose underlying distributions have Hölder smooth densities and can…

Cited by 15SourcePDFScholar
2023

ExposureDiffusion: Learning to Expose for Low-light Image Enhancement

ICCV 2023poster

Previous raw image-based low-light image enhancement methods predominantly relied on feed-forward neural networks to learn deterministic mappings from low-light to normally-exposed images. However, they failed to capture critical distribution information, leading to visually undesirable results. Thi…

Cited by 65PDFcodeScholar
2023

Frame-Level Multi-Label Playing Technique Detection Using Multi-Scale Network and Self-Attention Mechanism

ICASSP 2023accepted

Instrument playing technique (IPT) is a key element of musical presentation. However, most of the existing works for IPT detection only concern monophonic music signals, yet little has been done to detect IPTs in polyphonic instrumental solo pieces with overlapping IPTs or mixed IPTs. In this paper,…

Cited by 0SourceScholar
2023

H2RBox-v2: Incorporating Symmetry for Boosting Horizontal Box Supervised Oriented Object Detection

NeurIPS 2023poster

With the rapidly increasing demand for oriented object detection, e.g. in autonomous driving and remote sensing, the recently proposed paradigm involving weakly-supervised detector H2RBox for learning rotated box (RBox) from the more readily-available horizontal box (HBox) has shown promise. This pa…

Cited by 42SourcePDFScholar
2023

Raw Image Reconstruction With Learned Compact Metadata

CVPR 2023poster

While raw images exhibit advantages over sRGB images (e.g. linearity and fine-grained quantization level), they are not widely used by common users due to the large storage requirements. Very recent works propose to compress raw images by designing the sampling masks in the raw image pixel space, le…

2022

Change-point Detection for Sparse and Dense Functional Data in General Dimensions

NeurIPS 2022accept

We study the problem of change-point detection and localisation for functional data sequentially observed on a general $d$-dimensional space, where we allow the functional curves to be either sparsely or densely sampled. Data of this form naturally arise in a wide range of applications such as biolo…

2022

Deepchorus: A Hybrid Model of Multi-Scale Convolution And Self-Attention for Chorus Detection

ICASSP 2022accepted

Chorus detection is a challenging problem in musical signal processing as the chorus often repeats more than once in popular songs, usually with rich instruments and complex rhythm forms. Most of the existing works focus on the receptiveness of chorus sections based on some explicit features such as…

Cited by 0SourceScholar
2022

Denoising and change point localisation in piecewise-constant high-dimensional regression coefficients

AISTATS 2022poster

We study the theoretical properties of the fused lasso procedure originally proposed by Tibshirani et al. (2005) in the context of a linear regression model in which the regression coefficient are totally ordered and assumed to be sparse and piecewise constant. Despite its popularity, to the best of…

Cited by 11SourcePDFScholar
2022

Feature Distillation Interaction Weighting Network for Lightweight Image Super-resolution

AAAI 2022technical

Convolutional neural networks based single-image superresolution (SISR) has made great progress in recent years. However, it is difficult to apply these methods to real-world scenarios due to the computational and memory cost. Meanwhile, how to take full advantage of the intermediate features under…

2022

Lightweight Bimodal Network for Single-Image Super-Resolution via Symmetric CNN and Recursive Transformer

IJCAI 2022poster

Single-image super-resolution (SISR) has achieved significant breakthroughs with the development of deep learning. However, these methods are difficult to be applied in real-world scenarios since they are inevitably accompanied by the problems of computational and memory costs caused by the complex…

2022

Towards Robust Rain Removal Against Adversarial Attacks: A Comprehensive Benchmark Analysis and Beyond

CVPR 2022poster

Rain removal aims to remove rain streaks from images/videos and reduce the disruptive effects caused by rain. It not only enhances image/video visibility but also allows many computer vision algorithms to function properly. This paper makes the first attempt to conduct a comprehensive study on the r…

Cited by 66PDFcodeScholar
2021

Localizing Changes in High-Dimensional Regression Models

AISTATS 2021poster

This paper addresses the problem of localizing change points in high-dimensional linear regression models with piecewise constant regression coefficients. We develop a dynamic programming approach to estimate the locations of the change points whose performance improves upon the current state-of-the…

Cited by 49SourcePDFScholar
2021

Multi-TimeLine Summarization (MTLS): Improving Timeline Summarization by Generating Multiple Summaries

ACL 2021long

In this paper, we address a novel task, Multiple TimeLine Summarization (MTLS), which extends the flexibility and versatility of Time-Line Summarization (TLS). Given any collection of time-stamped news articles, MTLS automatically discovers important yet different stories and generates a correspondi…

2021

Robust Recursive Least M-Estimate Adaptive Filter for the Identification of Low-Rank Acoustic Systems

ICASSP 2021accepted

To identify acoustic systems (which are low-rank in nature) in non-Gaussian and Gaussian noise, a robust recursive least M-estimate adaptive filtering algorithm is developed in this paper by applying the nearest Kronecker product to decompose the acoustic impulse response. Two M-estimators, i.e., th…

Cited by 0SourceScholar
2021

Singer Identification Using Deep Timbre Feature Learning with KNN-NET

ICASSP 2021accepted

In this paper, we study the issue of automatic singer identification (SID) in popular music recordings, which aims to recognize who sang a given piece of song. The main challenge for this investigation lies in the fact that a singer’s singing voice changes and intertwines with the signal of backgrou…

Cited by 0SourceScholar
2020

Robust Frequency-Domain Recursive Least M-Estimate Adaptive Filter For Acoustic System Identification

ICASSP 2020accepted

To identify acoustic systems in non-Gaussian and Gaussian noises, a robust frequency-domain recursive least M-estimate (FRLM) adaptive filtering algorithm is proposed. The cost function of the adaptive filter is defined by using a robust time-domain M-estimator, while its update equation is derived…

Cited by 0SourceScholar
2018

Robust Diffusion Recursive Least Squares Estimation with Side Information for Networked Agents

ICASSP 2018accepted

This work develops a robust diffusion recursive least squares algorithm to mitigate the performance degradation often experienced in networks of agents in the presence of impulsive noise. This algorithm minimizes an exponentially weighted least-squares cost function subject to a time-dependent const…

Cited by 0SourceScholar