← Search

Le Yang

29 accepted papers

2026

Beyond Pixels: Mining Compressed Domain Artifacts for Efficient AI-Generated Video Detection

ICML 2026poster

With the rapid advancement of high-fidelity video generation models, robust AI-generated video (AIGV) detection has become increasingly needed. While most AIGV detection methods operate in the decoded pixel domain, we observe that detection in the pixel domain inevitably entangles task-irrelevant se…

Cited by 0SourceScholar
2026

DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning

ICLR 2026poster

Large Vision-Language Models excel at multimodal understanding but struggle to deeply integrate visual information into their predominantly text-based reasoning processes, a key challenge in mirroring human cognition. To address this, we introduce DeepEyes, a model that learns to ``think with images…

Cited by 0SourcecodeScholar
2026

Privacy on the Fly: A Predictive Adversarial Transformation Network for Mobile Sensor Data

AAAI 2026technical

Mobile motion sensors such as accelerometers and gyroscopes are now ubiquitously accessible by third-party apps via standard APIs. While enabling rich functionalities like activity recognition and step counting, this openness has also enabled unregulated inference of sensitive user traits, such as g

Cited by 0SourcePDFScholar
2025

D3: Training-Free AI-Generated Video Detection Using Second-Order Features

ICCV 2025poster

The evolution of video generation techniques, such as Sora, has made it increasingly easy to produce high-fidelity AI-generated videos, raising public concern over the dissemination of synthetic content. However, existing detection methodologies remain limited by their insufficient exploration of te…

2025

Deep Learning Based Topography Aware Gas Source Localization with Mobile Robot

ICRA 2025

Gas source localization in complex environments is critical for applications such as environmental monitoring, industrial safety, and disaster response. Traditional methods often struggle with the challenges posed by a lack of environmental topography integration, especially when interactions betwee

Cited by 0SourceScholar
2025

Hierarchical Similarity Loss Enhanced Depth and Structural Fidelity in Monocular RGB-to-Depth Mapping with Adversarial Training

ICASSP 2025accepted

The conversion of monocular RGB images to depth maps is crucial in robotic applications. Current supervised learning approaches, dependent on high-quality RGB-Depth pairs, struggle with indoor environments characterized by multiple objects and fluctuating lighting, leading to inaccurate and unstable…

Cited by 0SourceScholar
2025

HybridGS: High-Efficiency Gaussian Splatting Data Compression using Dual-Channel Sparse Representation and Point Cloud Encoder

ICML 2025poster

Most existing 3D Gaussian Splatting (3DGS) compression schemes focus on producing compact 3DGS representation via implicit data embedding. They have long encoding and decoding times and highly customized data format, making it difficult for widespread deployment. This paper presents a new 3DGS compr…

2025

Jailbreak-AudioBench: In-Depth Evaluation and Analysis of Jailbreak Threats for Large Audio Language Models

NeurIPS 2025poster

Large Language Models (LLMs) demonstrate impressive zero-shot performance across a wide range of natural language processing tasks. Integrating various modality encoders further expands their capabilities, giving rise to Multimodal Large Language Models (MLLMs) that process not only text but also vi…

Cited by 0SourceScholar
2025

M3ADD: A Novel Benchmark for Physiology Signal-based Automatic Depression Detection with Multimodal Multitask Multievent Framework

ICASSP 2025accepted

The prevalence of depression is escalating, especially among youth, which has become a critical mental health concern. Current assessment methods, relying heavily on questionnaires, clinical observations, and AI-driven analyses, are limited by their focus on single-event data, failing to encapsulate…

Cited by 0SourceScholar
2025

Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace Projection

CVPR 2025poster

Recent studies have shown that large vision-language models (LVLMs) often suffer from the issue of object hallucinations (OH). To mitigate this issue, we introduce an efficient method that edits the model weights based on an unsafe subspace, which we call HalluSpace in this paper. With truthful and…

2024

Enhanced Face Recognition using Intra-class Incoherence Constraint

ICLR 2024spotlight

The current face recognition (FR) algorithms has achieved a high level of accuracy, making further improvements increasingly challenging. While existing FR algorithms primarily focus on optimizing margins and loss functions, limited attention has been given to exploring the feature representation sp…

Cited by 3SourcePDFScholar
2024

Fine-grained Dynamic Network for Generic Event Boundary Detection

ECCV 2024poster

"Generic event boundary detection (GEBD) aims at pinpointing event boundaries naturally perceived by humans, playing a crucial role in understanding long-form videos. Given the diverse nature of generic boundaries, spanning different video appearances, objects, and actions, this task remains challen…

2024

Unveiling Typographic Deceptions: Insights of the Typographic Vulnerability in Large Vision-Language Models

ECCV 2024poster

"Large Vision-Language Models (LVLMs) rely on vision encoders and Large Language Models (LLMs) to exhibit remarkable capabilities on various multi-modal tasks in the joint space of vision and language. However, typographic attacks, which disrupt Vision-Language Models (VLMs) such as Contrastive Lang…

2023

Optimizing Distributed Multi-Sensor Multi-Target Tracking Algorithm Based On Labeled Multi-Bernoulli Filter

ICASSP 2023accepted

In this paper, we propose an improved distributed fusion algorithm under the Labeled multi-Bernoulli (LMB) filter framework. Firstly, the LMB parameter set is augmented by a new group variable, which is able to record the matching information of the neighbour sensors. Then the matching LMB component…

Cited by 0SourceScholar
2021

CondenseNet V2: Sparse Feature Reactivation for Deep Networks

CVPR 2021poster

Reusing features in deep networks through dense connectivity is an effective way to achieve high computational efficiency. The recent proposed CondenseNet has shown that this mechanism can be further improved if redundant features are removed. In this paper, we propose an alternative approach named…

Cited by 90PDFcodeScholar
2021

Kld Minimization-Based Constrained Measurement Filtering For Two-Step TDOA Indoor Tracking

ICASSP 2021accepted

This paper presents an enhanced two-step method for tracking an indoor point target using the time difference of arrival (TDOA) measurements from an ultra wideband (UWB) positioning system. Again, the algorithm preprocesses the raw TDOAs and then feeds the results to a recursively bounded grid-based…

Cited by 0SourceScholar
2021

Revisiting Locally Supervised Learning: an Alternative to End-to-end Training

ICLR 2021poster

Due to the need to store the intermediate activations for back-propagation, end-to-end (E2E) training of deep networks usually suffers from high GPUs memory footprint. This paper aims to address this problem by revisiting the locally supervised learning, where a network is split into gradient-isolat…

2020

Glance and Focus: a Dynamic Approach to Reducing Spatial Redundancy in Image Classification

NeurIPS 2020poster

The accuracy of deep convolutional neural networks (CNNs) generally improves when fueled with high resolution images. However, this often comes at a high computational cost and high memory footprint. Inspired by the fact that not all regions in an image are task-relevant, we propose a novel framewor…

2020

Resolution Adaptive Networks for Efficient Inference

CVPR 2020poster

Adaptive inference is an effective mechanism to achieve a dynamic tradeoff between accuracy and computational cost in deep networks. Existing works mainly exploit architecture redundancy in network depth or width. In this paper, we focus on spatial redundancy of input samples and propose a novel Res…

Cited by 308PDFcodeScholar
2020

Robust Tdoa Indoor Tracking Using Constrained Measurement Filtering and Grid-Based Filtering

ICASSP 2020accepted

This paper considers exploiting the time difference of arrival (TDOA) measurements from a ultra wideband (UWB) indoor positioning system to locate a moving point target. In indoor environments, measured TDOAs are subject to large errors due to multipath and/or non-line-of-sight (NLOS) propagation. B…

Cited by 0SourceScholar
2018

Reinforcement Cutting-Agent Learning for Video Object Segmentation

CVPR 2018poster

Video object segmentation is a fundamental yet challenging task in computer vision community. In this paper, we formulate this problem as a Markov Decision Process, where agents are learned to segment object regions under a deep reinforcement learning framework. Essentially, learning agents for segm…

Cited by 105SourcePDFScholar
2017

Moving target localization in multistatic sonar using time delays, Doppler shifts and arrival angles

ICASSP 2017accepted

Identifying the location of a target is a fundamental application in multistatic sonar. Numerous attempts have been made to improve the accuracy, computational efficiency and robustness of target positioning. Previous studies mostly use time delay and angle measurements for localization, or time del…

Cited by 0SourceScholar
2017

SPFTN: A Self-Paced Fine-Tuning Network for Segmenting Objects in Weakly Labelled Videos

CVPR 2017poster

Object segmentation in weakly labelled videos is an interesting yet challenging task, which aims at learning to perform category-specific video object segmentation by only using video-level tags. Existing works in this research area might still have some limitations, e.g., lack of effective DNN-base…

Cited by 61PDFScholar