← Search

Yue Zhou

45 accepted papers

2026

AdaptCLIP: Adapting CLIP for Universal Visual Anomaly Detection

AAAI 2026technical

Universal visual anomaly detection aims to identify anomalies from novel or unseen vision domains without additional fine-tuning, which is critical in open scenarios. Recent studies have demonstrated that pre-trained vision-language models like CLIP exhibit strong generalization with just zero or a

Cited by 0SourcePDFScholar
2026

Deep Residual Injection for Full-Spectrum Forensic Signal Perception in Multimodal Large Language Models

ICML 2026poster

Multimodal large language models (MLLMs) have been increasingly adopted in forensics for their robust semantic understanding. As AI-generated images become realistic, semantic-level inconsistencies alone are often insufficient for reliable detection. This motivates a critical question: *whether MLLM…

Cited by 0SourceScholar
2026

Foundation Visual Encoders Are Secretly Few-Shot Anomaly Detectors

ICLR 2026poster

Few-shot anomaly detection streamlines and simplifies industrial safety inspection. However, limited samples make accurate differentiation between normal and abnormal features challenging, and even more so under category-agnostic conditions. Large-scale pre-training of foundation visual encoders has…

Cited by 0SourcecodeScholar
2026

Hilbert Curve-Encoded Rotation-Equivariant Oriented Object Detector with Locality-Preserving Spatial Mapping

AAAI 2026technical

Arbitrary-Oriented Object Detection (AOOD) has found broad applications in embodied intelligence, autonomous driving, and satellite remote sensing. However, current AOOD frameworks face challenges in ineffective feature extraction and orientation regression inaccuracy. Inspired by Hilbert curve

Cited by 0SourcePDFScholar
2026

Partial Weakly-Supervised Oriented Object Detection

CVPR 2026

The growing demand for oriented object detection (OOD) across various domains has driven significant research in this area. However, the high cost of dataset annotation remains a major concern. Current mainstream OOD algorithms can be mainly categorized into three types: (1) fully supervised methods

Cited by 0SourcecodeScholar
2026

Point2RBox-v3: Self-Bootstrapping from Point Annotations via Integrated Pseudo-Label Refinement and Utilization

ICLR 2026poster

Driven by the growing need for Oriented Object Detection (OOD), learning from point annotations under a weakly-supervised framework has emerged as a promising alternative to costly and laborious manual labeling. In this paper, we discuss two deficiencies in existing point-supervised methods: ineffic…

Cited by 0SourcecodeScholar
2026

RECODE: A Benchmark for Research Code DEvelopment with Interactive Human Feedback

ICLR 2026poster

Large language models (LLMs) show the promise in supporting scientific research implementation, yet their ability to generate correct and executable code remains limited. Existing works largely adopt one-shot settings, ignoring the iterative and feedback-driven nature of realistic workflows of scien…

Cited by 0SourcecodeScholar
2026

Scalable Event Cloud Network for Event-based Classification

ICML 2026oral

Event cameras are biologically inspired sensors garnering significant attention from both industry and academia. Mainstream methods favor frame and voxel representations, which reach a satisfactory performance while introducing time-consuming transformations, bulky models, and sacrificing fine-grain…

Cited by 0SourceScholar
2026

Where It Moves, It Matters: Referring Surgical Instrument Segmentation via Motion

AAAI 2026technical

Enabling intuitive, language-driven interaction with surgical scenes is a critical step toward intelligent operating rooms and autonomous surgical robotic assistance. However, the task of referring segmentation, localizing surgical instruments based on natural language descriptions, remains underexp

Cited by 0SourcePDFScholar
2025

Breaking Latent Prior Bias in Detectors for Generalizable AIGC Image Detection

NeurIPS 2025poster

Current AIGC detectors often achieve near-perfect accuracy on images produced by the same generator used for training but struggle to generalize to outputs from unseen generators. We trace this failure in part to latent prior bias: detectors learn shortcuts tied to patterns stemming from the initial…

Cited by 0SourceScholar
2025

ClearSight: Human Vision-Inspired Solutions for Event-Based Motion Deblurring

ICCV 2025poster

Motion deblurring addresses the challenge of image blur caused by camera or scene movement. Event cameras provide motion information that is encoded in the asynchronous event streams. To efficiently leverage the temporal information of event streams, we employ Spiking Neural Networks (SNNs) for moti…

Cited by 0SourcePDFScholar
2025

Guard Me If You Know Me: Protecting Specific Face-Identity from Deepfakes

NeurIPS 2025poster

Securing personal identity against deepfake attacks is increasingly critical in the digital age, especially for celebrities and political figures whose faces are easily accessible and frequently targeted. Most existing deepfake detection methods focus on general-purpose scenarios and often ignore th…

Cited by 0SourcecodeScholar
2025

PersonaGym: Evaluating Persona Agents and LLMs

EMNLP 2025

Persona agents, which are LLM agents conditioned to act according to an assigned persona, enable contextually rich and user-aligned interactions across domains like education and healthcare.However, evaluating how faithfully these agents adhere to their personas remains a significant challenge, part

Cited by 0SourcePDFScholar
2025

Representation Purification for End-to-End Speech Translation

COLING 2025main

Speech-to-text translation (ST) is a cross-modal task that involves converting spoken language into text in a different language. Previous research primarily focused on enhancing speech translation by facilitating knowledge transfer from machine translation, exploring various methods to bridge the g…

2025

SACR: Self-training with Saliency-Augmented Consistency Regularization for Few-Shot Learners

ICASSP 2025accepted

Pre-trained language models have made significant strides in natural language processing tasks, enabling flexible fine-tuning for downstream applications. However, in few-shot learning scenarios, pre-trained models face challenges related to overfitting due to limited training samples, which hinders…

Cited by 0SourceScholar
2025

TestNUC: Enhancing Test-Time Computing Approaches and Scaling through Neighboring Unlabeled Data Consistency

ACL 2025long

Test-time computing approaches, which leverage additional computational resources during inference, have been proven effective in enhancing large language model performance. This work introduces a novel, linearly scaling approach, TestNUC, that improves test-time predictions by leveraging the local…

2025

Text4Seg: Reimagining Image Segmentation as Text Generation

ICLR 2025poster

Multimodal Large Language Models (MLLMs) have shown exceptional capabilities in vision-language tasks; however, effectively integrating image segmentation into these models remains a significant challenge. In this paper, we introduce Text4Seg, a novel text-as-mask paradigm that casts image segmentat…

2025

Towards Data Contamination Detection for Modern Large Language Models: Limitations, Inconsistencies, and Oracle Challenges

COLING 2025main

As large language models achieve increasingly impressive results, questions arise about whether such performance is from generalizability or mere data memorization. Thus, numerous data contamination detection methods have been proposed. However, these approaches are often validated with traditional…

2025

Unveiling Performance Challenges of Large Language Models in Low-Resource Healthcare: A Demographic Fairness Perspective

COLING 2025main

This paper studies the performance of large language models (LLMs), particularly regarding demographic fairness, in solving real-world healthcare tasks. We evaluate state-of-the-art LLMs with three prevalent learning frameworks across six diverse healthcare tasks and find significant challenges in a…

2025

VLForgery Face Triad: Detection, Localization and Attribution via Multimodal Large Language Models

NeurIPS 2025poster

Faces synthesized by diffusion models (DMs) with high-quality and controllable attributes pose a significant challenge for Deepfake detection. Most state-of-the-art detectors only yield a binary decision, incapable of forgery localization, attribution of forgery methods, and providing analysis on th…

Cited by 0SourceScholar
2025

Veracity Bias and Beyond: Uncovering LLMs’ Hidden Beliefs in Problem-Solving Reasoning

ACL 2025long

Despite LLMs’ explicit alignment against demographic stereotypes, they have been shown to exhibit biases under various social contexts. In this work, we find that LLMs exhibit concerning biases in how they associate solution veracity with demographics. Through experiments across five human value-ali…

Cited by 0SourcePDFScholar
2024

A Simple and Effective Point-based Network for Event Camera 6-DOFs Pose Relocalization

CVPR 2024poster

Event cameras exhibit remarkable attributes such as high dynamic range asynchronicity and low latency making them highly suitable for vision tasks that involve high-speed motion in challenging lighting conditions. These cameras implicitly capture movement and depth information in events making them…

Cited by 12SourcePDFScholar
2024

CLIF: Complementary Leaky Integrate-and-Fire Neuron for Spiking Neural Networks

ICML 2024spotlight

Spiking neural networks (SNNs) are promising brain-inspired energy-efficient models. Compared to conventional deep Artificial Neural Networks (ANNs), SNNs exhibit superior efficiency and capability to process temporal information. However, it remains a challenge to train SNNs due to their undifferen…

2024

ImplicitAVE: An Open-Source Dataset and Multimodal LLMs Benchmark for Implicit Attribute Value Extraction

ACL 2024findings

Existing datasets for attribute value extraction (AVE) predominantly focus on explicit attribute values while neglecting the implicit ones, lack product images, are often not publicly available, and lack an in-depth human inspection across diverse domains. To address these limitations, we present Im…

2024

Large Language Models Are Involuntary Truth-Tellers: Exploiting Fallacy Failure for Jailbreak Attacks

EMNLP 2024main

We find that language models have difficulties generating fallacious and deceptive reasoning. When asked to generate deceptive outputs, language models tend to leak honest counterparts but believe them to be false. Exploiting this deficiency, we propose a jailbreak attack method that elicits an alig…

2024

Memory-Augmented speech-to-text Translation with Multi-Scale Context Translation Strategy

ICASSP 2024accepted

End-to-end speech-to-text translation (ST) has demonstrated promising results on sentence-level translation. In real-world scenarios, audio is typically long and requires cross-sentence contextual connections for translation. Sentence-level ST models are facing challenges since they lack the ability…

Cited by 0SourceScholar
2024

Modeling Low-Resource Health Coaching Dialogues via Neuro-Symbolic Goal Summarization and Text-Units-Text Generation

COLING 2024main

Health coaching helps patients achieve personalized and lifestyle-related goals, effectively managing chronic conditions and alleviating mental health issues. It is particularly beneficial, however cost-prohibitive, for low-socioeconomic status populations due to its highly personalized and labor-in…

2024

Paraphrase and Solve: Exploring and Exploiting the Impact of Surface Form on Mathematical Reasoning in Large Language Models

NAACL 2024long

This paper studies the relationship between the surface form of a mathematical problem and its solvability by large language models. We find that subtle alterations in the surface form can significantly impact the answer distribution and the solve rate, exposing the language model’s lack of robustne…

2024

SpikePoint: An Efficient Point-based Spiking Neural Network for Event Cameras Action Recognition

ICLR 2024spotlight

Event cameras are bio-inspired sensors that respond to local changes in light intensity and feature low latency, high energy efficiency, and high dynamic range. Meanwhile, Spiking Neural Networks (SNNs) have gained significant attention due to their remarkable efficiency and fault tolerance. By syne…

Cited by 25SourcePDFScholar
2023

DeCrisisMB: Debiased Semi-Supervised Learning for Crisis Tweet Classification via Memory Bank

EMNLP 2023long findings

During crisis events, people often use social media platforms such as Twitter to disseminate information about the situation, warnings, advice, and support. Emergency relief organizations leverage such information to acquire timely crisis circumstances and expedite rescue operations. While existing…

Cited by 0SourcecodeScholar
2023

GradPU: Positive-Unlabeled Learning via Gradient Penalty and Positive Upweighting

AAAI 2023technical

Positive-unlabeled learning is an essential problem in many real-world applications with only labeled positive and unlabeled data, especially when the negative samples are difficult to identify. Most existing positive-unlabeled learning methods will inevitably overfit the positive class to some exte…

Cited by 8SourcePDFScholar
2023

H2RBox-v2: Incorporating Symmetry for Boosting Horizontal Box Supervised Oriented Object Detection

NeurIPS 2023poster

With the rapidly increasing demand for oriented object detection, e.g. in autonomous driving and remote sensing, the recently proposed paradigm involving weakly-supervised detector H2RBox for learning rotated box (RBox) from the more readily-available horizontal box (HBox) has shown promise. This pa…

Cited by 42SourcePDFScholar
2023

H2RBox: Horizontal Box Annotation is All You Need for Oriented Object Detection

ICLR 2023poster

Oriented object detection emerges in many applications from aerial images to autonomous driving, while many existing detection benchmarks are annotated with horizontal bounding box only which is also less costive than fine-grained rotated box, leading to a gap between the readily available training…

2023

Improving Chinese Pop Song and Hokkien Gezi Opera Singing Voice Synthesis by Enhancing Local Modeling

EMNLP 2023long main

Singing Voice Synthesis (SVS) strives to synthesize pleasing vocals based on music scores and lyrics. The current acoustic models based on Transformer usually process the entire sequence globally and use a simple L1 loss. However, this approach overlooks the significance of local modeling within the…

Cited by 0SourcecodeScholar
2023

Multimodal Tremor Suppression of the Wrist Using FES and Electric Motors-A Simulation Study

RA-L 2023

Wearable technologies have shown promising results in tremor management, making them a feasible alternative to current treatments. Devices based on active actuation, such as electric motors, show high tremor suppression rate, but are heavy and bulky. In contrast, devices based on functional electric

Cited by 8SourceScholar
2023

The KFIoU Loss for Rotated Object Detection

ICLR 2023poster

Differing from the well-developed horizontal object detection area whereby the computing-friendly IoU based loss is readily adopted and well fits with the detection metrics, rotation detectors often involve a more complicated loss based on SkewIoU which is unfriendly to gradient-based training. In t…

Cited by 239SourcePDFScholar
2022

Towards Enhancing Health Coaching Dialogue in Low-Resource Settings

COLING 2022main

Health coaching helps patients identify and accomplish lifestyle-related goals, effectively improving the control of chronic diseases and mitigating mental health conditions. However, health coaching is cost-prohibitive due to its highly personalized and labor-intensive nature. In this paper, we pro…

2021

Analysis of the Effect of Common Disturbances on the Safety of a Wearable Tremor Suppression Device

RA-L 2021

The advent of wearable technology has enabled a large number of externally worn mechatronic devices to be developed and tested on people with movement disorders. The complexity of these disorders and the variety of conditions across different patients have resulted in a pressing demand for the incor

Cited by 6SourceScholar
2021

Dense Label Encoding for Boundary Discontinuity Free Rotation Detection

CVPR 2021poster

Rotation detection serves as a fundamental building block in many visual applications involving aerial image, scene text, and face etc. Differing from the dominant regression-based approaches for orientation estimation, this paper explores a relatively less-studied methodology based on classificatio…

Cited by 350PDFScholar