← Search

Changlin Li

23 accepted papers

2026

DisCa: Accelerating Video Diffusion Transformers with Distillation-Compatible Learnable Feature Caching

CVPR 2026

While diffusion models have achieved great success in the field of video generation, this progress is accompanied by a rapidly escalating computational burden. Among the existing acceleration methods, Feature Caching is popular due to its training-free property and considerable speedup performance,

Cited by 0SourcecodeScholar
2026

Do VLMs Perceive or Recall? Probing Visual Perception vs. Memory with Classic Visual Illusions

CVPR 2026

Large Vision-Language Models (VLMs) often answer classic visual illusions "correctly" on original images, yet persist with the same responses when illusion factors are inverted, even though the visual change is obvious to humans. This raises a fundamental question: do VLMs perceive visual changes or

Cited by 0SourceScholar
2026

Efficient Training for Human Video Generation with Entropy-Guided Prioritized Progressive Learning

CVPR 2026

Human video generation has advanced rapidly with the development of diffusion models, but the high computational cost and substantial memory consumption associated with training these models on high-resolution, multi-frame data pose significant challenges. In this paper, we propose Entropy-Guided Pr

Cited by 0SourcecodeScholar
2026

GeoSense: Internalizing Geometric Necessity Perception for Multimodal Reasoning

ICML 2026poster

Advancing towards artificial superintelligence requires rich and intelligent perceptual capabilities. A critical frontier in this pursuit is overcoming the limited spatial understanding of Multimodal Large Language Models (MLLMs), where geometry information is essential. Existing methods often addre…

Cited by 0SourceScholar
2026

Progressive Online Video Understanding with Evidence-Aligned Timing and Transparent Decisions

ICLR 2026poster

Visual agents operating in the wild must respond to queries precisely when sufficient evidence first appears in a video stream, a critical capability that is overlooked by conventional video LLMs evaluated in offline settings. The shift to an online, streaming paradigm introduces significant challen…

Cited by 0SourceScholar
2024

DiFiNet: Boundary-Aware Semantic Differentiation and Filtration Network for Nested Named Entity Recognition

ACL 2024long

Nested Named Entity Recognition (Nested NER) entails identifying and classifying entity spans within the text, including the detection of named entities that are embedded within external entities. Prior approaches primarily employ span-based techniques, utilizing the power of exhaustive searches to…

Cited by 2SourcePDFScholar
2024

Perception-Oriented Video Frame Interpolation via Asymmetric Blending

CVPR 2024poster

Previous methods for Video Frame Interpolation (VFI) have encountered challenges notably the manifestation of blur and ghosting effects. These issues can be traced back to two pivotal factors: unavoidable motion errors and misalignment in supervision. In practice motion estimates often prove to be e…

2024

Predicting the Unpredictable: Uncertainty-Aware Reasoning over Temporal Knowledge Graphs via Diffusion Process

ACL 2024findings

Temporal Knowledge Graph (TKG) reasoning seeks to predict future incomplete facts leveraging historical data. While existing approaches have shown effectiveness in addressing the task through various perspectives, such as graph learning and logic rules, they are limited in capturing the indeterminac…

Cited by 0SourcePDFScholar
2023

GrowCLIP: Data-Aware Automatic Model Growing for Large-scale Contrastive Language-Image Pre-Training

ICCV 2023poster

Cross-modal pre-training has shown impressive performance on a wide range of downstream tasks, benefiting from massive image-text pairs collected from the Internet. In practice, online data are growing constantly, highlighting the importance of the ability of pre-trained model to learn from data tha…

Cited by 5PDFcodeScholar
2023

MixReorg: Cross-Modal Mixed Patch Reorganization is a Good Mask Learner for Open-World Semantic Segmentation

ICCV 2023poster

Recently, semantic segmentation models trained with image-level text supervision have shown promising results in challenging open-world scenarios. However, these models still face difficulties in learning fine-grained semantic alignment at the pixel level and predicting accurate object masks. To add…

Cited by 19PDFScholar
2023

TTC4MCP: Monocular Collision Prediction Based on Self-Supervised TTC Estimation

IROS 2023poster

Vision-based collision prediction for autonomous driving is a challenging task due to the dynamic movement of vehicles and diverse types of obstacles. Most existing methods rely on object detection algorithms, which only predict predefined collision targets, such as vehicles and pedestrians, and can…

Cited by 2SourceScholar
2023

ViewCo: Discovering Text-Supervised Segmentation Masks via Multi-View Semantic Consistency

ICLR 2023poster

Recently, great success has been made in learning visual representations from text supervision, facilitating the emergence of text-supervised semantic segmentation. However, existing works focus on pixel grouping and cross-modal semantic alignment, while ignoring the correspondence among multiple au…

2022

Arch-Graph: Acyclic Architecture Relation Predictor for Task-Transferable Neural Architecture Search

CVPR 2022poster

Neural Architecture Search (NAS) aims to find efficient models for multiple tasks. Beyond seeking solutions for a single task, there are surging interests in transferring network design knowledge across multiple tasks. In this line of research, effectively modeling task correlations is vital yet hig…

Cited by 25PDFcodeScholar
2022

Automated Progressive Learning for Efficient Training of Vision Transformers

CVPR 2022poster

Recent advances in vision Transformers (ViTs) have come with a voracious appetite for computing power, high-lighting the urgent need to develop efficient training methods for ViTs. Progressive learning, a training scheme where the model capacity grows progressively during training, has started showi…

Cited by 49PDFcodeScholar
2022

Beyond Fixation: Dynamic Window Visual Transformer

CVPR 2022poster

Recently, a surge of interest in visual transformers is to reduce the computational cost by limiting the calculation of self-attention to a local window. Most current work uses a fixed single-scale window for modeling by default, ignoring the impact of window size on model performance. However, this…

Cited by 41PDFcodeScholar
2022

Continual Object Detection via Prototypical Task Correlation Guided Gating Mechanism

CVPR 2022poster

Continual learning is a challenging real-world problem for constructing a mature AI system when data are provided in a streaming fashion. Despite recent progress in continual classification, the researches of continual object detection are impeded by the diverse sizes and numbers of objects in each…

Cited by 44PDFcodeScholar
2022

Look Back and Forth: Video Super-Resolution With Explicit Temporal Difference Modeling

CVPR 2022poster

Temporal modeling is crucial for video super-resolution. Most of the video super-resolution methods adopt the optical flow or deformable convolution for explicitly motion compensation. However, such temporal modeling techniques increase the model complexity and might fail in case of occlusion or com…

Cited by 60PDFcodeScholar
2021

BossNAS: Exploring Hybrid CNN-Transformers With Block-Wisely Self-Supervised Neural Architecture Search

ICCV 2021poster

A myriad of recent breakthroughs in hand-crafted neural architectures for visual recognition have highlighted the urgent need to explore hybrid architectures consisting of diversified building blocks. Meanwhile, neural architecture search methods are surging with an expectation to reduce human effor…

Cited by 142PDFcodeScholar
2021

Joint Depth and Normal Estimation from Real-world Time-of-flight Raw Data

IROS 2021poster

We present a novel approach to joint depth and normal estimation for time-of-flight (ToF) sensors. Our model learns to predict the high-quality depth and normal maps jointly from ToF raw sensor data. To achieve this, we meticulously constructed the first large-scale dataset (named ToF-100) with pair…

Cited by 4SourceScholar
2021

Pi-NAS: Improving Neural Architecture Search by Reducing Supernet Training Consistency Shift

ICCV 2021poster

Recently proposed neural architecture search (NAS) methods co-train billions of architectures in a supernet and estimate their potential accuracy using the network weights detached from the supernet. However, the ranking correlation between the architectures' predicted accuracy and their actual capa…

Cited by 22PDFcodeScholar
2020

Block-Wisely Supervised Neural Architecture Search With Knowledge Distillation

CVPR 2020poster

Neural Architecture Search (NAS), aiming at automatically designing network architectures by machines, is expected to bring about a new revolution in machine learning. Despite these high expectation, the effectiveness and efficiency of existing NAS solutions are unclear, with some recent works going…

Cited by 244PDFcodeScholar