← Search

Cuiling Lan

33 accepted papers

2026

Temperature as a Meta-Policy: Adaptive Temperature in LLM Reinforcement Learning

ICLR 2026poster

Temperature is a crucial hyperparameter in large language models (LLMs), controlling the trade-off between exploration and exploitation during text generation. High temperatures encourage diverse but noisy outputs, while low temperatures produce focused outputs but may cause premature convergence. Y…

Cited by 0SourceScholar
2025

TIV-Diffusion: Towards Object-Centric Movement for Text-driven Image to Video Generation

AAAI 2025technical

Text-driven Image to Video Generation (TI2V) aims to generate controllable video given the first frame and corresponding textual description. The primary challenges of this task lie in two parts: (i) how to identify the target objects and ensure the consistency between the movement trajectory and th…

Cited by 1SourcePDFScholar
2024

Diffusion Model with Cross Attention as an Inductive Bias for Disentanglement

NeurIPS 2024spotlight

Disentangled representation learning strives to extract the intrinsic factors within the observed data. Factoring these representations in an unsupervised manner is notably challenging and usually requires tailored loss functions or specific structural designs. In this paper, we introduce a new pers…

Cited by 6SourcePDFScholar
2024

Slot-VLM: Object-Event Slots for Video-Language Modeling

NeurIPS 2024poster

Video-Language Models (VLMs), powered by the advancements in Large Language Models (LLMs), are charting new frontiers in video understanding. A pivotal challenge is the development of an effective method to encapsulate video content into a set of representative tokens to align with LLMs. In this wor…

Cited by 0SourcePDFScholar
2024

Text Grouping Adapter: Adapting Pre-trained Text Detector for Layout Analysis

CVPR 2024poster

Significant progress has been made in scene text detection models since the rise of deep learning but scene text layout analysis which aims to group detected text instances as paragraphs has not kept pace. Previous works either treated text detection and grouping using separate models or train a mod…

Cited by 1SourcePDFScholar
2024

UCIP: A Universal Framework for Compressed Image Super-Resolution using Dynamic Prompt

ECCV 2024poster

"Compressed Image Super-resolution (CSR) aims to simultaneously super-resolve the compressed images and tackle the challenging hybrid distortions caused by compression. However, existing works on CSR usually focus on single compression codec, , JPEG, ignoring the diverse traditional or learning-base…

2023

Adaptive Frequency Filters As Efficient Global Token Mixers

ICCV 2023poster

Recent vision transformers, large-kernel CNNs and MLPs have attained remarkable successes in broad vision tasks thanks to their effective information fusion in the global scope. However, their efficient deployments, especially on mobile devices, still suffer from noteworthy challenges due to the hea…

Cited by 69PDFcodeScholar
2023

Deep Frequency Filtering for Domain Generalization

CVPR 2023poster

Improving the generalization ability of Deep Neural Networks (DNNs) is critical for their practical uses, which has been a longstanding challenge. Some theoretical studies have uncovered that DNNs have preferences for some frequency components in the learning process and indicated that this may affe…

Cited by 63SourcePDFScholar
2023

Learning Distortion Invariant Representation for Image Restoration From a Causality Perspective

CVPR 2023poster

In recent years, we have witnessed the great advancement of Deep neural networks (DNNs) in image restoration. However, a critical limitation is that they cannot generalize well to real-world degradations with different degrees or types. In this paper, we are the first to propose a novel training str…

2023

Shatter and Gather: Learning Referring Image Segmentation with Text Supervision

ICCV 2023poster

Referring image segmentation, the task of segmenting any arbitrary entities described in free-form texts, opens up a variety of vision applications. However, manual labeling of training data for this task is prohibitively costly, leading to lack of labeled data for training. We address this issue b…

Cited by 22PDFcodeScholar
2023

Template-guided Hierarchical Feature Restoration for Anomaly Detection

ICCV 2023poster

Targeting for detecting anomalies of various sizes for complicated normal patterns, we propose a Template-guided Hierarchical Feature Restoration method, which introduces two key techniques, bottleneck compression and template-guided compensation, for anomaly-free feature restoration. Specially, our…

Cited by 34PDFScholar
2023

Versatile Neural Processes for Learning Implicit Neural Representations

ICLR 2023poster

Representing a signal as a continuous function parameterized by neural network (a.k.a. Implicit Neural Representations, INRs) has attracted increasing attention in recent years. Neural Processes (NPs), which model the distributions over functions conditioned on partial observations (context set), pr…

2023

WEDGE: Web-Image Assisted Domain Generalization for Semantic Segmentation

ICRA 2023poster

Domain generalization for semantic segmentation is highly demanded in real applications, where a trained model is expected to work well in previously unseen domains. One challenge lies in the lack of data which could cover the diverse distributions of the possible unseen domains for training. In thi…

Cited by 27SourceScholar
2022

Lifelong Unsupervised Domain Adaptive Person Re-Identification With Coordinated Anti-Forgetting and Adaptation

CVPR 2022poster

Unsupervised domain adaptive person re-identification (ReID) has been extensively investigated to mitigate the adverse effects of domain gaps. Those works assume the target domain data can be accessible all at once. However, for the real-world streaming data, this hinders the timely adaptation to ch…

Cited by 42PDFScholar
2022

Mask-based Latent Reconstruction for Reinforcement Learning

NeurIPS 2022accept

For deep reinforcement learning (RL) from pixels, learning effective state representations is crucial for achieving high performance. However, in practice, limited experience and high-dimensional inputs prevent effective representation learning. To address this, motivated by the success of mask-base…

2022

ReSTR: Convolution-Free Referring Image Segmentation Using Transformers

CVPR 2022poster

Referring image segmentation is an advanced semantic segmentation task where target is not a predefined class but is described in natural language. Most of existing methods for this task rely heavily on convolutional neural networks, which however have trouble capturing long-range dependencies betwe…

Cited by 176PDFScholar
2021

Exploiting Sample Uncertainty for Domain Adaptive Person Re-Identification

AAAI 2021technical

Many unsupervised domain adaptive (UDA) person ReID approaches combine clustering-based pseudo-label prediction with feature fine-tuning. However, because of domain gap, the pseudo-labels are not always reliable and there are noisy/incorrect labels. This would mislead the feature representation lea…

Cited by 190SourcePDFScholar
2021

Generalizing to Unseen Domains: A Survey on Domain Generalization

IJCAI 2021poster

Domain generalization (DG), i.e., out-of-distribution generalization, has attracted increased interests in recent years. Domain generalization deals with a challenging setting where one or several different but related domain(s) are given, and the goal is to learn a model that can generalize to an u…

2021

MetaAlign: Coordinating Domain Alignment and Classification for Unsupervised Domain Adaptation

CVPR 2021poster

For unsupervised domain adaptation (UDA), to alleviate the effect of domain shift, many approaches align the source and target domains in the feature space by adversarial learning or by explicitly aligning their statistics. However, the optimization objective of such domain alignment is generally no…

Cited by 137PDFScholar
2021

PlayVirtual: Augmenting Cycle-Consistent Virtual Trajectories for Reinforcement Learning

NeurIPS 2021poster

Learning good feature representations is important for deep reinforcement learning (RL). However, with limited experience, RL often suffers from data inefficiency for training. For un-experienced or less-experienced trajectories (i.e., state-action sequences), the lack of data limits the use of them…

2021

Re-Energizing Domain Discriminator With Sample Relabeling for Adversarial Domain Adaptation

ICCV 2021poster

Many unsupervised domain adaptation (UDA) methods exploit domain adversarial training to align the features to reduce domain gap, where a feature extractor is trained to fool a domain discriminator in order to have aligned feature distributions. The discrimination capability of the domain classifier…

Cited by 18PDFScholar
2021

ToAlign: Task-Oriented Alignment for Unsupervised Domain Adaptation

NeurIPS 2021poster

Unsupervised domain adaptive classifcation intends to improve the classifcation performance on unlabeled target domain. To alleviate the adverse effect of domain shift, many approaches align the source and target domains in the feature space. However, a feature is usually taken as a whole for alignm…

2021

Uncertainty-Aware Few-Shot Image Classification

IJCAI 2021poster

Few-shot image classification learns to recognize new categories from limited labelled data. Metric learning based approaches have been widely investigated, where a query sample is classified by finding the nearest prototype from the support set based on their feature similarities. A neural network…

Cited by 30SourcePDFScholar
2020

Global Distance-distributions Separation for Unsupervised Person Re-identification

ECCV 2020poster

Supervised person re-identification (ReID) often has poor scalability and usability in real-world deployments due to domain gaps and the lack of annotations for the target domain data. Unsupervised person ReID through domain adaptation is attractive yet challenging. Existing unsupervised ReID approa…

Cited by 90SourcePDFScholar
2020

Multi-Granularity Reference-Aided Attentive Feature Aggregation for Video-Based Person Re-Identification

CVPR 2020poster

Video-based person re-identification (reID) aims at matching the same person across video clips. It is a challenging task due to the existence of redundancy among frames, newly revealed appearance, occlusion, and motion blurs. In this paper, we propose an attentive feature aggregation module, namely…

Cited by 138PDFScholar
2020

Relation-Aware Global Attention for Person Re-Identification

CVPR 2020poster

For person re-identification (re-id), attention mechanisms have become attractive as they aim at strengthening discriminative features and suppressing irrelevant ones, which matches well the key of re-id, i.e., discriminative feature learning. Previous approaches typically learn attention using loca…

Cited by 713PDFcodeScholar
2020

Semantics-Guided Neural Networks for Efficient Skeleton-Based Human Action Recognition

CVPR 2020poster

Skeleton-based human action recognition has attracted great interest thanks to the easy accessibility of the human skeleton data. Recently, there is a trend of using very deep feedforward neural networks to model the 3D coordinates of joints without considering the computational efficiency. In this…

Cited by 635PDFcodeScholar
2020

Style Normalization and Restitution for Generalizable Person Re-Identification

CVPR 2020poster

Existing fully-supervised person re-identification (ReID) methods usually suffer from poor generalization capability caused by domain gaps. The key to solving this problem lies in filtering out identity-irrelevant interference and learning domain-invariant person representations. In this paper, we a…

Cited by 448PDFcodeScholar
2018

Adding Attentiveness to the Neurons in Recurrent Neural Networks

ECCV 2018poster

Recurrent neural networks (RNNs) are capable of modeling the temporal dynamics of complex sequential information. However, the structures of existing RNN neurons mainly focus on controlling the contributions of current and historical information but do not explore the different importance levels of…

Cited by 105SourcePDFScholar
2017

Human Pose Estimation Using Global and Local Normalization

ICCV 2017poster

In this paper, we address the problem of estimating the positions of human joints, i.e., articulated pose estimation. Recent state-of-the-art solutions model two key issues, joint detection and spatial configuration refinement, together using convolutional neural networks. Our work mainly focuses on…

Cited by 82PDFScholar
2017

View Adaptive Recurrent Neural Networks for High Performance Human Action Recognition From Skeleton Data

ICCV 2017poster

Skeleton-based human action recognition has recently attracted increasing attention due to the popularity of 3D skeleton data. One main challenge lies in the large view variations in captured human actions. We propose a novel view adaptation scheme to automatically regulate observation viewpoints du…

Cited by 683PDFcodeScholar