← Search

Zijun Long

5 accepted papers

2026

Is Training Necessary for Anomaly Detection?

ICML 2026poster

Current state-of-the-art multi-class unsupervised anomaly detection (MUAD) methods rely on training encoder–decoder models to reconstruct anomaly-free features. We first show these approaches have an inherent fidelity–stability dilemma in how they detect anomalies via reconstruction residuals. We th…

Cited by 0SourceScholar
2026

RoboEye: Enhancing 2D Robotic Object Identification with Selective 3D Geometric Keypoint Matching

ICRA 2026poster

The rapid growing number of product categories in large-scale e-commerce makes accurate object identification for automated packing in warehouses substantially more difficult. As the catalog grows, intra-class variability and a long tail of rare or visually similar items increase, and—when combined …

2024

LaCViT: A Label-Aware Contrastive Fine-Tuning Framework for Vision Transformers

ICASSP 2024accepted

Vision Transformers (ViTs) have emerged as popular models in computer vision, demonstrating state-of-the-art performance across various tasks. This success typically follows a two-stage strategy involving pre-training on large-scale datasets using self-supervised signals, such as masked random patch…

Cited by 0SourceScholar
2024

Multiway-Adapter: Adapting Multimodal Large Language Models for Scalable Image-Text Retrieval

ICASSP 2024accepted

As Multimodal Large Language Models (MLLMs) grow in size, adapting them to specialized tasks becomes increasingly challenging due to high computational and memory demands. Indeed, traditional fine-tuning methods are costly, due to the need for extensive, task-specific training. While efficient adapt…

Cited by 0SourceScholar
2024

RoboLLM: Robotic Vision Tasks Grounded on Multimodal Large Language Models

ICRA 2024poster

Robotic vision applications often necessitate a wide range of visual perception tasks, such as object detection, segmentation, and identification. While there have been substantial advances in these individual tasks, integrating specialized models into a unified vision pipeline presents significant…

Cited by 21SourcecodeScholar