← Search

Qiankun Li

11 accepted papers

2026

Granulon: Awakening Pixel-Level Visual Encoders with Adaptive Multi-Granularity Semantics for MLLM

CVPR 2026

Recent advances in multimodal large language models largely rely on CLIP-based visual encoders, which emphasize global semantic alignment but struggle with fine-grained visual understanding. In contrast, DINOv3 provides strong pixel-level perception yet lacks coarse-grained semantic abstraction, lea

Cited by 0SourcecodeScholar
2026

HG-Lane: High-Fidelity Generation of Lane Scenes under Adverse Weather and Lighting Conditions without Re-annotation

CVPR 2026

Lane detection is a crucial task in autonomous driving, which is conducive to ensuring the safe operation of vehicles. However, current datasets like CULane and TuSimple have relatively limited data under extreme weather conditions, such as rain, snow and fog, which makes detection models unreliable

Cited by 0SourcecodeScholar
2026

Mem-T: Densifying Rewards for Long-Horizon Memory Agents

ICML 2026poster

Memory agents, which depart from predefined memory-processing pipelines by endogenously managing the processing, storage, and retrieval of memories, have garnered increasing attention for their autonomy and adaptability. However, existing training paradigms remain constrained: agents often traverse …

Cited by 0SourceScholar
2026

Memoria-Bench: A Comprehensive Benchmark for Evaluating Memory in Long-Horizon Autonomous Agents

ICML 2026poster

Memory is a core capability of autonomous agents, yet existing benchmarks evaluate it primarily in constrained settings such as short dialogues or synthetic tasks, failing to reflect realistic agent deployments. We present \textbf{Memoria-Bench}, a benchmark for evaluating agent memory grounded in c…

Cited by 0SourceScholar
2026

OralGPT-Omni: A Versatile Dental Multimodal Large Language Model

CVPR 2026

Multimodal Large Language Models (MLLMs) have exhibited immense potential across numerous medical specialties, yet dentistry remains underexplored, in part due to limited domain-specific data, scarce dental expert annotations, insufficient modality-specific modeling, and challenges in reliability. I

Cited by 0SourceScholar
2026

Reallocating Attention Across Layers to Reduce Multimodal Hallucination

CVPR 2026

Multimodal large reasoning models (MLRMs) often suffer from hallucinations that stem not only from insufficient visual grounding but also from imbalanced allocation between perception and reasoning processes. Building upon recent interpretability findings suggesting a staged division of attention ac

Cited by 0SourcecodeScholar
2025

From Pixels to Views: Learning Angular-Aware and Physics-Consistent Representations for Light Field Microscopy

NeurIPS 2025poster

Light field microscopy (LFM) has become an emerging tool in neuroscience for large-scale neural imaging in vivo, with XLFM (eXtended Light Field Microscopy) notable for its single-exposure volumetric imaging, broad field of view, and high temporal resolution. However, learning-based 3D reconstructi…

Cited by 0SourcecodeScholar
2025

Unleashing Foundation Vision Models: Adaptive Transfer for Diverse Data-Limited Scientific Domains

NeurIPS 2025poster

In the big data era, the computer vision field benefits from large-scale datasets such as LAION-2B, LAION-400M, and ImageNet-21K, Kinetics, on which popular models like the ViT and ConvNeXt series have been pre-trained, acquiring substantial knowledge. However, numerous downstream tasks in speciali…

Cited by 0SourcecodeScholar
2023

LABANet: Lead-Assisting Backbone Attention Network for Oral Multi-Pathology Segmentation

ICASSP 2023accepted

This paper presents a Lead-Assisting Backbone Attention Network (LABANet), which is able to perform multi-pathology instance segmentation of dental panoramic X-rays. A Lead-Assisting Attention Backbone (LAAB), containing two Swin-Transformers, is first developed for feature extraction. The following…

Cited by 0SourceScholar