← Search

Jun Long

15 accepted papers

2026

Injection Without Distortion: Geometrically Constrained Knowledge Enhancement for Vision-Language Models

AAAI 2026technical

Vision-Language Models (VLMs) are widely used in tasks like Open-Vocabulary Object Detection and zero-shot Classification, owing to their powerful generalization. However, recent research reveals that VLMs exhibit significant performance instability when tasked with recognizing concepts at varying g

Cited by 0SourcePDFScholar
2026

Unlearning without Forgetting: Securely Removing Targeted Concepts from Large-Scale Vision-Language Open-Vocabulary Detectors

CVPR 2026

Open-vocabulary detectors (OvOD) inherit tightly coupled cross-modal knowledge from web-scale pretraining, creating privacy, copyright, and compliance risks. Existing machine unlearning methods face geometric entanglement interference in OvOD: forgetting updates inevitably distort preserved knowledg

Cited by 0SourceScholar
2025

A Reinforcement Learning Agent Controlled Multi-branch Small Object Detection Framework

ICASSP 2025accepted

The past few years have witnessed the immense development of small object detection, which is aimed at detecting size-limited targets in high-resolution images. The prevailing methods focus on extracting fine-grained information by expanding the receptive fields and then generating the potential sma…

Cited by 0SourceScholar
2025

Enhancing Extrapolation Reasoning on Temporal Knowledge Graphs with Logic Rules and Queries

ICASSP 2025accepted

Extrapolation reasoning on Temporal Knowledge Graphs (TKGs) plays a pivotal role in various systems, including retrieval, recommendation, and Q&A. Traditional TKG reasoning methods tend to emphasize modeling the local and global features of facts, often overlooking the alignment with query semantics…

Cited by 0SourceScholar
2025

Enhancing Session-Based Recommendation with Hypergraph Motifs and Contrastive Learning

ICASSP 2025accepted

Session-based recommendation (SBR) provides personalized recommendations by analyzing the interactions of anonymous session users. Recent approaches based on graph neural networks (GNNs) focus on pairwise relations to infer potential user preferences. However, real-world user interactions are often…

Cited by 0SourceScholar
2025

Harmonizing for defect visibility with Fine-Grained Hierarchical Interaction Learning

ICASSP 2025accepted

Defect detection is a fundamental task in industrial image analysis, crucial for identifying and delineating defect regions. However, existing models, often struggle to learn critical features effectively under conditions of noisy interference. In this study, we introduce the Fine-Grained Hierarchic…

Cited by 0SourceScholar
2025

HieClip: Hierarchical CLIP with Explicit Alignment for Zero-Shot Anomaly Detection

ICASSP 2025accepted

Large image-language models(LLM) have made significant progress in zero-shot anomaly detection(ZSAD), however, the semantic gap between images and text limits their performance in hierarchical learning. In this paper, we propose the hierarchical alignment clip(HieClip) framework, to achieve hierarch…

Cited by 0SourceScholar
2025

Noisy Correspondence Rectification via Asymmetric Similarity Learning

AAAI 2025technical

Cross-modal matching shows enormous potential to recognize objects across different sensory modalities, which is fundamental to numerous visual-language tasks like image-text retrieval and visual captioning. Existing works generally rely on massive and well-aligned data pairs for model training. Unf…

Cited by 0SourcePDFScholar
2025

Statistical Model-driven Similarity Hashing: Bridging Modalities for Efficient Unsupervised Retrieval

AAAI 2025technical

Unsupervised deep cross-modal hash retrieval aims to map multi-modal features into binary hash codes without labels, which is of interest due to its storage efficiency, query speed and convenient applications. However, existing approaches suffer from two main limitations: (1) Slightly insufficient c…

Cited by 0SourcePDFScholar
2024

Beyond the Limit of Weight-Sharing: Pioneering Space-Evolving NAS with Large Language Models

ICASSP 2024accepted

Large language models (LLMs) offer impressive performance across diverse fields, but their increasing complexity raises both design costs and the need for specialized expertise. These challenges are intensified for Neural Architecture Search (NAS) methods reliant on weight-sharing techniques. This p…

Cited by 0SourceScholar
2024

Detecting Any instruction-to-answer interaction relationship:Universal Instruction-to-Answer Navigator for Med-VQA

ICML 2024poster

Medical Visual Question Answering (Med-VQA) interprets complex medical imagery using user instructions for precise diagnostics, yet faces challenges due to diverse, inadequately annotated images. In this paper, we introduce the Universal Instruction-Vision Navigator (Uni-Med) framework for extractin…

2024

HTCCN: Temporal Causal Convolutional Networks with Hawkes Process for Extrapolation Reasoning in Temporal Knowledge Graphs

NAACL 2024long

Temporal knowledge graphs (TKGs) serve as powerful tools for storing and modeling dynamic facts, holding immense potential in anticipating future facts. Since future facts are inherently unknowable, effectively modeling the intricate temporal structure of historical facts becomes paramount for accur…

Cited by 4SourcePDFScholar
2023

Stacking-Based Attention Temporal Convolutional Network for Action Segmentation

ICASSP 2023accepted

Action segmentation plays an important role in video understanding, which is implemented by frame-wise action classification. Recent works on action segmentation capture long-term dependencies by increasing temporal convolution layers in Temporal Convolution Networks (TCNs). However, high layers in…

Cited by 0SourceScholar