← Search

Long Qian

10 accepted papers

2026

MATHPHYS-GUIDED COARSE-TO-FINE ANOMALY SYNTHESIS WITH SQE-DRIVEN BI-LEVEL OPTIMIZATION FOR ANOMALY DETECTION

ICASSP 2026oral

Currently, industrial anomaly detection suffers from two bottlenecks: (i) the rarity of real-world defect images and (ii) the opacity of sample quality when synthetic data are used. Existing synthetic strategies (e.g., cut-and-paste) overlook the underlying physical causes of defects, leading to inc…

Cited by 0SourcePDFScholar
2026

Quality-Aware Language-Conditioned Local Auto-Regressive Anomaly Synthesis and Detection

AAAI 2026technical

Despite substantial progress in anomaly synthesis, existing diffusion-based and coarse inpainting pipelines commonly suffer from structural deficiencies such as micro-structural discontinuities, limited semantic controllability, and inefficient generation. To overcome these limitations, we introduce

Cited by 0SourcePDFScholar
2025

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training

CVPR 2025poster

Video Large Language Models (Video-LLMs) have recently shown strong performance in basic video understanding tasks, such as captioning and coarse-grained question answering, but struggle with compositional reasoning that requires multi-step spatio-temporal inference across object relations, interact…

Cited by 4SourcePDFScholar
2024

Grounded Answers for Multi-agent Decision-making Problem through Generative World Model

NeurIPS 2024poster

Recent progress in generative models has stimulated significant innovations in many fields, such as image generation and chatbots. Despite their success, these models often produce sketchy and misleading solutions for complex multi-agent decision-making problems because they miss the trial-and-error…

Cited by 10SourcePDFScholar
2024

Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning

ICML 2024poster

Large Language Models (LLMs) demonstrate remarkable proficiency in comprehending and handling text-based tasks. Many efforts are being made to transfer these attributes to video modality, which are termed Video-LLMs. However, existing Video-LLMs can only capture the coarse-grained semantics and are…

2022

Compositional Temporal Grounding With Structured Variational Cross-Graph Correspondence Learning

CVPR 2022poster

Temporal grounding in videos aims to localize one target video segment that semantically corresponds to a given query sentence. Thanks to the semantic diversity of natural language descriptions, temporal grounding allows activity grounding beyond pre-defined classes and has received increasing atten…

Cited by 80PDFcodeScholar
2022

Fine-Grained Semantically Aligned Vision-Language Pre-Training

NeurIPS 2022accept

Large-scale vision-language pre-training has shown impressive advances in a wide range of downstream tasks. Existing methods mainly model the cross-modal alignment by the similarity of the global representations of images and text, or advanced cross-modal attention upon image and text features. Howe…

2020

FlexiVision: Teleporting the Surgeon’s Eyes via Robotic Flexible Endoscope and Head-Mounted Display

IROS 2020poster

A flexible endoscope introduces more dexterity to the image capturing in endoscopic surgery. However, manual control or automatic control based on instrument tracking does not handle the misorientation between the endoscopic video and the surgeon. We propose an automatic flexible endoscope control m…

Cited by 21SourceScholar
2019

Augmented Reality Assisted Instrument Insertion and Tool Manipulation for the First Assistant in Robotic Surgery

ICRA 2019poster

In robotic-assisted laparoscopic surgery, the first assistant (FA) stands at the bedside assisting the intervention, while the surgeon sits at the console teleoperating the robot. Tasks for the FA include navigating new instruments into the surgeon's field-of-view and passing in or retracting materi…

Cited by 44SourceScholar