← Search

Junyu Gao

41 accepted papers

2026

Beyond Prompt Degradation: Prototype-guided Dual-pool Prompting for Incremental Object Detection

CVPR 2026

Incremental Object Detection (IOD) aims to continuously learn new object categories without forgetting previously learned ones. Recently, prompt-based methods have gained popularity for their replay-free design and parameter efficiency. However, due to prompt coupling and prompt drift, these methods

Cited by 0SourcecodeScholar
2026

Dual-level Adaptation for Multi-Object Tracking: Building Test-Time Calibration from Experience and Intuition

CVPR 2026

Multiple Object Tracking (MOT) has long been a fundamental task in computer vision, with broad applications in various real-world scenarios. However, due to distribution shifts in appearance, motion pattern, and catagory between the training and testing data, model performance degrades considerably

Cited by 0SourcecodeScholar
2026

Exploring Efficient Open-Vocabulary Segmentation in the Remote Sensing

AAAI 2026technical

Open-Vocabulary Remote Sensing Image Segmentation (OVRSIS), an emerging task that adapts Open-Vocabulary Segmentation (OVS) to the remote sensing (RS) domain, remains underexplored due to the absence of a unified evaluation benchmark and the domain gap between natural and RS images. To bridge these

Cited by 0SourcePDFScholar
2026

Exploring the Underwater World Segmentation without Extra Training

CVPR 2026

Accurate segmentation of marine organisms is vital for biodiversity monitoring and ecological assessment, yet existing datasets and models remain largely limited to terrestrial scenes. To bridge this gap, we introduce **AquaOV255**, the first large-scale and fine-grained underwater segmentation data

Cited by 0SourcecodeScholar
2026

HulluEdit: Single-Pass Evidence-Consistent Subspace Editing for Mitigating Hallucinations in Large Vision-Language Models

CVPR 2026

Object hallucination in Large Vision-Language Models (LVLMs) significantly hinders their reliable deployment. Existing methods struggle to balance efficiency and accuracy: they often require expensive reference models and multiple forward passes, or apply static edits that risk suppressing genuine v

Cited by 0SourcecodeScholar
2026

Inconsistency Biases in Dynamic Data Pruning

ICLR 2026poster

Dynamic data pruning accelerates training by focusing on informative samples. However, comparing importance scores across different model states introduces inconsistency (score context drift), and variable selection rates bias gradient dynamics over time (temporal gradient bias). We introduce RePB (…

Cited by 0SourcecodeScholar
2026

IntroSVG: Learning from Rendering Feedback for Text-to-SVG Generation via an Introspective Generator-Critic Framework

CVPR 2026

Scalable Vector Graphics (SVG) are central to digital design due to their inherent scalability and editability. Despite significant advancements in content generation enabled by Visual Language Models (VLMs), existing text-to-SVG generation methods are limited by a core challenge: the autoregressive

Cited by 0SourceScholar
2026

MedMamba: Multi-View State Space Models with Adaptive Graph Learning for Medical Time Series Classification

ICML 2026poster

Medical time series are central to healthcare, enabling continuous monitoring and supporting timely clinical decisions. Despite recent progress, existing methods struggle to jointly model local-global dynamics and handle nonstationarities like baseline drift, while often failing to capture latent ch…

Cited by 0SourceScholar
2026

Progressive Cross-Modal Causal Intervention for Long-Term Action Recognition

CVPR 2026

Intricate correlations among atomic actions and inherent visual confounders in long-term action recognition (LTAR) contribute to the persistent challenges in this domain. While methods based on vision-language models that employ label text for supervision offer potential for handling visual confound

Cited by 0SourcecodeScholar
2026

Reasoning via Implicit Self-supervised Emergence for Instruction Segmentation

AAAI 2026technical

We challenge the assumption that complex instruction-guided segmentation tasks necessitate equally complex and explicit supervision. This paper introduces RISE (Reasoning via Implicit Self-supervised Emergence), a framework that learns intricate compositional reasoning, spanning spatial relations to

Cited by 0SourcePDFScholar
2026

SC$^{2}$-WM: A Self-Correcting World Model with Closed-Loop Feedback for Vision-and-Language Navigation in Continuous Environments

ICML 2026poster

Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to make fine-grained navigation decisions under partial observability. However, most existing methods rely on open-loop execution, lacking mechanisms to detect and correct internal state drift during inference. We pro…

Cited by 0SourceScholar
2025

Enhancing Low-Rank Adaptation with Recoverability-Based Reinforcement Pruning for Object Counting

AAAI 2025technical

Object counting is crucial for understanding the distribution of objects in different scenarios. Recently, many object counting networks have been designed to be more complex to achieve marginal improvements, leading to excessive time spent on model design. With the development of large models (LMs)…

Cited by 0SourcePDFScholar
2025

LLMs Caught in the Crossfire: Malware Requests and Jailbreak Challenges

ACL 2025long

The widespread adoption of Large Language Models (LLMs) has heightened concerns about their security, particularly their vulnerability to jailbreak attacks that leverage crafted prompts to generate malicious outputs. While prior research has been conducted on general security capabilities of LLMs, t…

2025

NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments

ICCV 2025poster

Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to execute sequential navigation actions in complex environments guided by natural language instructions. Current approaches often struggle with generalizing to novel environments and adapting to ongoing changes durin…

2025

PUO-Bench: A Panel Understanding and Operation Benchmark with A Privacy-Preserving Framework

NeurIPS 2025poster

Recent advancements in Vision-Language Models (VLMs) have enabled GUI agents to leverage visual features for interface understanding and operation in the digital world. However, limited research has addressed the interpretation and interaction with control panels in real-world settings. To bridge th…

Cited by 0SourceScholar
2025

Unity in Diversity: Video Editing via Gradient-Latent Purification

CVPR 2025poster

Recently, text-driven video editing methods that optimize target latent representations have garnered significant attention and demonstrated promising results. However, these methods rely on self-supervised objectives to compute the gradients needed for updating latent representations, which inevita…

2025

WebUIBench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in WebUI-to-Code

ACL 2025finding

With the rapid advancement of Generative AI technology, Multimodal Large Language Models(MLLMs) have the potential to act as AI software engineers capable of executing complex web application development. Considering that the model requires a confluence of multidimensional sub-capabilities to addres…

2024

Combating Data Imbalances in Federated Semi-supervised Learning with Dual Regulators

AAAI 2024technical

Federated learning has become a popular method to learn from decentralized heterogeneous data. Federated semi-supervised learning (FSSL) emerges to train models from a small fraction of labeled data due to label scarcity on decentralized clients. Existing FSSL methods assume independent and identica…

Cited by 8SourcePDFScholar
2024

Conjugated Semantic Pool Improves OOD Detection with Pre-trained Vision-Language Models

NeurIPS 2024poster

A straightforward pipeline for zero-shot out-of-distribution (OOD) detection involves selecting potential OOD labels from an extensive semantic pool and then leveraging a pre-trained vision-language model to perform classification on both in-distribution (ID) and OOD labels. In this paper, we theori…

2024

Fast-Slow Test-Time Adaptation for Online Vision-and-Language Navigation

ICML 2024poster

The ability to accurately comprehend natural language instructions and navigate to the target location is essential for an embodied agent. Such agents are typically required to execute user instructions in an online manner, leading us to explore the use of unlabeled test samples for effective online…

2023

Cascade Evidential Learning for Open-World Weakly-Supervised Temporal Action Localization

CVPR 2023poster

Targeting at recognizing and localizing action instances with only video-level labels during training, Weakly-supervised Temporal Action Localization (WTAL) has achieved significant progress in recent years. However, living in the dynamically changing open world where unknown actions constantly spri…

Cited by 21SourcePDFScholar
2023

Collecting Cross-Modal Presence-Absence Evidence for Weakly-Supervised Audio-Visual Event Perception

CVPR 2023poster

With only video-level event labels, this paper targets at the task of weakly-supervised audio-visual event perception (WS-AVEP), which aims to temporally localize and categorize events belonging to each modality. Despite the recent progress, most existing approaches either ignore the unsynchronized…

2022

DR.VIC: Decomposition and Reasoning for Video Individual Counting

CVPR 2022poster

Pedestrian counting is a fundamental tool for understanding pedestrian patterns and crowd flow analysis. Existing works (e.g., image-level pedestrian counting, crossline crowd counting et al.) either only focus on the image-level counting or are constrained to the manual annotation of lines. In this…

Cited by 28PDFcodeScholar
2022

Dual-Evidential Learning for Weakly-Supervised Temporal Action Localization

ECCV 2022poster

"Weakly-supervised temporal action localization (WS-TAL) aims to localize the action instances and recognize their categories with only video-level labels. Despite great progress, existing methods suffer from severe action-background ambiguity, which mainly comes from background noise introduced by…

2022

Fine-Grained Temporal Contrastive Learning for Weakly-Supervised Temporal Action Localization

CVPR 2022poster

We target at the task of weakly-supervised action localization (WSAL), where only video-level action labels are available during model training. Despite the recent progress, existing methods mainly embrace a localization-by-classification paradigm and overlook the fruitful fine-grained temporal dist…

Cited by 104PDFcodeScholar
2021

Fast Video Moment Retrieval

ICCV 2021poster

This paper targets at fast video moment retrieval (fast VMR), aiming to localize the target moment efficiently and accurately as queried by a given natural language sentence. We argue that most existing VMR approaches can be divided into three modules namely video encoder, text encoder, and cross-mo…

Cited by 137PDFScholar
2020

Unsupervised Semantic Aggregation and Deformable Template Matching for Semi-Supervised Learning

NeurIPS 2020poster

Unlabeled data learning has attracted considerable attention recently. However, it is still elusive to extract the expected high-level semantic feature with mere unsupervised learning. In the meantime, semi-supervised learning (SSL) demonstrates a promising future in leveraging few samples. In this…

2017

Embedding structured contour and location prior in siamesed fully convolutional networks for road detection

ICRA 2017poster

Road detection from the perspective of moving vehicles is a challenging issue in autonomous driving. Recently, many deep learning methods spring up for this task because they can extract high-level local features to find road regions from raw RGB data, such as Convolutional Neural Networks (CNN) and…

Cited by 307SourceScholar