← Search

Kaiqi Huang

54 accepted papers

2026

CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos

AAAI 2026technical

Recent advances in large language models (LLMs) have improved reasoning in text and image domains, yet achieving robust video reasoning remains a significant challenge. Existing video benchmarks mainly assess shallow understanding and reasoning and allow models to exploit global context, failing to

Cited by 0SourcePDFScholar
2026

Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos

ICML 2026poster

Without incurring significant computational overhead, train-free long video generation aims to enable foundation video generation models to produce longer videos. Frame-level autoregressive frameworks, e.g., FIFO-diffusion, offer the advantage of generating infinitely long videos with constant memor…

Cited by 0SourceScholar
2026

ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency Constraints

AAAI 2026technical

Video generation models have achieved remarkable progress, particularly excelling in realistic scenarios; however, their performance degrades notably in imaginative scenarios. These prompts often involve rarely co-occurring concepts with long-distance semantic relationships, falling outside training

Cited by 0SourcePDFScholar
2026

LATENT TEMPORAL DISCREPANCY AS MOTION PRIOR: A LOSS-WEIGHTING STRATEGY FOR DYNAMIC FIDELITY IN T2V

ICASSP 2026oral

Video generation models have achieved notable progress in static scenarios, yet their performance in motion video generation remains limited, with quality degrading under drastic dynamic changes. This is due to noise disrupting temporal coherence and increasing the difficulty of learning dynamic reg…

Cited by 0SourcePDFScholar
2026

NarrLV: Towards a Comprehensive Narrative-Centric Evaluation for Long Video Generation

ICLR 2026poster

With the rapid development of foundation video generation technologies, long video generation models have exhibited promising research potential thanks to expanded content creation space. Recent studies reveal that the goal of long video generation tasks is not only to extend video duration but also…

Cited by 0SourceScholar
2026

No-Regret Strategy Solving in Imperfect-Information Games via Pre-Trained Embedding

AAAI 2026technical

High-quality information set abstraction remains a core challenge in solving large-scale imperfect-information extensive-form games (IIEFGs)--such as no-limit Texas Hold’em--where the finite nature of spatial resources hinders solving strategies for the full game. State-of-the-art AI methods rely on

Cited by 0SourcePDFScholar
2026

RefRea: Reference-Guided Reasoning with Meta-Cognition for Accurate Language Model Agents

AAAI 2026technical

In recent years, with the rapid development of large language models (LLMs), LLM-based agents have achieved remarkable progress across a wide range of tasks. However, reasoning inconsistencies in LLMs still significantly limit the performance of agents in complex decision-making scenarios. Cognitive

Cited by 0SourcePDFScholar
2025

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking

ICCV 2025poster

Vision-language tracking aims to locate the target object in the video sequence using a template patch and a language description provided in the initial frame. To achieve robust tracking, especially in complex long-term scenarios that reflect real-world conditions as recently highlighted by MGIT, i…

2025

CSTrack: Enhancing RGB-X Tracking via Compact Spatiotemporal Features

ICML 2025poster

Effectively modeling and utilizing spatiotemporal features from RGB and other modalities (e.g., depth, thermal, and event data, denoted as X) is the core of RGB-X tracker design. Existing methods often employ two parallel branches to separately process the RGB and X input streams, requiring the mod…

2025

Constructive Conflict-Driven Multi-Agent Reinforcement Learning for Strategic Diversity

IJCAI 2025

In recent years, diversity has emerged as a useful mechanism to enhance the efficiency of multi-agent reinforcement learning (MARL). However, existing methods predominantly focus on designing policies based on individual agent characteristics, often neglecting the interplay and mutual influence amon

Cited by 0SourcePDFScholar
2025

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues

ICASSP 2025accepted

Vision-Language Tracking (VLT) aims to localize a target in video sequences using a visual template and language description. While textual cues enhance tracking potential, current datasets typically contain much more image data than text, limiting the ability of VLT methods to align the two modalit…

Cited by 0SourceScholar
2025

LLM Data Selection and Utilization via Dynamic Bi-level Optimization

ICML 2025poster

While large-scale training data is fundamental for developing capable large language models (LLMs), strategically selecting high-quality data has emerged as a critical approach to enhance training efficiency and reduce computational costs. Current data selection methodologies predominantly rely on s…

Cited by 0SourcePDFScholar
2025

Sequential Preference Optimization: Multi-Dimensional Preference Alignment with Implicit Reward Modeling

AAAI 2025technical

Human preference alignment is critical in building powerful and reliable large language models (LLMs). However, current methods either ignore the multi-dimensionality of human preferences (e.g. helpfulness and harmlessness) or struggle with the complexity of managing multiple reward models. To addre…

2025

Task-Parameterized Dynamic Movement Primitives With Reinforcement Learning for Improved Motion Planning

RA-L 2025

Online trajectory planning in unstructured environments poses significant challenges for mobile robots, particularly when navigating complex obstacles. Traditional learning-from-demonstration (LfD) methods depend on offline datasets, limiting their ability to adapt to varying obstacle shapes and dyn

Cited by 2SourceScholar
2025

Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards

NeurIPS 2025poster

Chain of thought reasoning has demonstrated remarkable success in large language models, yet its adaptation to vision-language reasoning remains an open challenge with unclear best practices. Existing attempts typically employ reasoning chains at a coarse-grained level, which struggles to perform fi…

Cited by 0SourcecodeScholar
2024

ADMN: Agent-Driven Modular Network for Dynamic Parameter Sharing in Cooperative Multi-Agent Reinforcement Learning

IJCAI 2024poster

Parameter sharing is a common strategy in multi-agent reinforcement learning (MARL) to make the training more efficient and scalable. However, applying parameter sharing among agents indiscriminately hinders the emergence of agents diversity and degrades the final cooperative performance. To better…

Cited by 1SourcePDFScholar
2024

Beyond Accuracy: Tracking more like Human via Visual Search

NeurIPS 2024poster

Human visual search ability enables efficient and accurate tracking of an arbitrary moving target, which is a significant research interest in cognitive neuroscience. The recently proposed Central-Peripheral Dichotomy (CPD) theory sheds light on how humans effectively process visual information and…

2024

DDAE: Towards Deep Dynamic Vision BERT Pretraining

AAAI 2024technical

Recently, masked image modeling (MIM) has demonstrated promising prospects in self-supervised representation learning. However, existing MIM frameworks recover all masked patches equivalently, ignoring that the reconstruction difficulty of different patches can vary sharply due to their diverse dist…

Cited by 1SourcePDFScholar
2024

MemVLT: Vision-Language Tracking with Adaptive Memory-based Prompts

NeurIPS 2024poster

Vision-language tracking (VLT) enhances traditional visual object tracking by integrating language descriptions, requiring the tracker to flexibly understand complex and diverse text in addition to visual information. However, most existing vision-language trackers still overly rely on initial fixed…

Cited by 5SourcePDFScholar
2024

PeLK: Parameter-efficient Large Kernel ConvNets with Peripheral Convolution

CVPR 2024poster

Recently some large kernel convnets strike back with appealing performance and efficiency. However given the square complexity of convolution scaling up kernels can bring about an enormous amount of parameters and the proliferated parameters can induce severe optimization problem. Due to these issue…

Cited by 46SourcePDFScholar
2024

Revealing the Dark Secrets of Extremely Large Kernel ConvNets on Robustness

ICML 2024poster

Robustness is a vital aspect to consider when deploying deep learning models into the wild. Numerous studies have been dedicated to the study of the robustness of vision transformers (ViTs), which have dominated as the mainstream backbone choice for vision tasks since the dawn of 2020s. Recently, so…

2024

TAPE: Leveraging Agent Topology for Cooperative Multi-Agent Policy Gradient

AAAI 2024technical

Multi-Agent Policy Gradient (MAPG) has made significant progress in recent years. However, centralized critics in state-of-the-art MAPG methods still face the centralized-decentralized mismatch (CDM) issue, which means sub-optimal actions by some agents will affect other agent's policy learning. Whi…

2024

Task-Wise Prompt Query Function for Rehearsal-Free Continual Learning

ICASSP 2024accepted

Continual learning (CL) aims to enable a model to retain knowledge of old tasks while learning new ones. One effective approach to CL is based on data rehearsal method. However, this approach increases the cost of storing data and cannot be used when data from old tasks is unavailable for some reaso…

Cited by 0SourceScholar
2023

A Multi-modal Global Instance Tracking Benchmark (MGIT): Better Locating Target in Complex Spatio-temporal and Causal Relationship

NeurIPS 2023poster

Tracking an arbitrary moving target in a video sequence is the foundation for high-level tasks like video understanding. Although existing visual-based trackers have demonstrated good tracking capabilities in short video sequences, they always perform poorly in complex environments, as represented b…

Cited by 13SourcePDFScholar
2023

Re-parameterizing Your Optimizers rather than Architectures

ICLR 2023poster

The well-designed structures in neural networks reflect the prior knowledge incorporated into the models. However, though different models have various priors, we are used to training them with model-agnostic optimizers such as SGD. In this paper, we propose to incorporate model-specific prior knowl…

2023

Subspace-Aware Exploration for Sparse-Reward Multi-Agent Tasks

AAAI 2023technical

Exploration under sparse rewards is a key challenge for multi-agent reinforcement learning problems. One possible solution to this issue is to exploit inherent task structures for an acceleration of exploration. In this paper, we present a novel exploration approach, which encodes a special structur…

Cited by 8SourcePDFScholar
2022

InsPro: Propagating Instance Query and Proposal for Online Video Instance Segmentation

NeurIPS 2022accept

Video instance segmentation (VIS) aims at segmenting and tracking objects in videos. Prior methods typically generate frame-level or clip-level object instances first and then associate them by either additional tracking heads or complex instance matching algorithms. This explicit instance associati…

Cited by 19SourcePDFScholar
2022

Learning Disentangled Attribute Representations for Robust Pedestrian Attribute Recognition

AAAI 2022technical

Although various methods have been proposed for pedestrian attribute recognition, most studies follow the same feature learning mechanism, ie, learning a shared pedestrian image feature to classify multiple attributes. However, this mechanism leads to low-confidence predictions and non-robustness of…

Cited by 39SourcePDFScholar
2022

PanopticDepth: A Unified Framework for Depth-Aware Panoptic Segmentation

CVPR 2022poster

This paper presents a unified framework for depth-aware panoptic segmentation (DPS), which aims to reconstruct 3D scene with instance-level semantics from one single image. Prior works address this problem by simply adding a dense depth regression head to panoptic segmentation (PS) networks, resulti…

Cited by 29PDFcodeScholar
2022

QueryProp: Object Query Propagation for High-Performance Video Object Detection

AAAI 2022technical

Video object detection has been an important yet challenging topic in computer vision. Traditional methods mainly focus on designing the image-level or box-level feature propagation strategies to exploit temporal information. This paper argues that with a more effective and efficient feature propaga…

Cited by 36SourcePDFScholar
2021

Learning to Reweight Imaginary Transitions for Model-Based Reinforcement Learning

AAAI 2021technical

Model-based reinforcement learning (RL) is more sample efficient than model-free RL by using imaginary trajectories generated by the learned dynamics model. When the model is inaccurate or biased, imaginary trajectories may be deleterious for training the action-value and policy functions. To allevi…

2021

Spatial and Semantic Consistency Regularizations for Pedestrian Attribute Recognition

ICCV 2021poster

While recent studies on pedestrian attribute recognition have shown remarkable progress in leveraging complicated networks and attention mechanisms, most of them neglect the inter-image relations and an important prior: spatial consistency and semantic consistency of attributes under surveillance sc…

Cited by 79PDFScholar
2019

SSAP: Single-Shot Instance Segmentation With Affinity Pyramid

ICCV 2019poster

Recently, proposal-free instance segmentation has received increasing attention due to its concise and efficient pipeline. Generally, proposal-free methods generate instance-agnostic semantic segmentation labels and instance-aware features to group pixels into different object instances. However, pr…

Cited by 316PDFScholar
2019

Towards Rich Feature Discovery With Class Activation Maps Augmentation for Person Re-Identification

CVPR 2019poster

The fundamental challenge of small inter-person variation requires Person Re-Identification (Re-ID) models to capture sufficient fine-grained information. This paper proposes to discover diverse discriminative visual cues without extra assistance, e.g., pose estimation, human parsing. Specifically,…

Cited by 311PDFScholar
2018

A2-RL: Aesthetics Aware Reinforcement Learning for Image Cropping

CVPR 2018poster

Image cropping aims at improving the aesthetic quality of images by adjusting their composition. Most weakly supervised cropping methods (without bounding box supervision) rely on the sliding window mechanism. The sliding window mechanism requires fixed aspect ratios and limits the cropping region w…

2018

Adversarially Occluded Samples for Person Re-Identification

CVPR 2018poster

Person re-identification (ReID) is the task of retrieving particular persons across different cameras. Despite its great progress in recent years, it is still confronted with challenges like pose variation, occlusion, and similar appearance among different persons. The large gap between training and…

Cited by 305SourcePDFScholar
2018

Discriminative Learning of Latent Features for Zero-Shot Recognition

CVPR 2018poster

Zero-shot learning (ZSL) aims to recognize unseen image categories by learning an embedding space between image and semantic representations. For years, among existing works, it has been the center task to learn the proper mapping matrices aligning the visual and semantic space, whilst the importanc…

Cited by 196SourcePDFScholar
2017

Beyond Triplet Loss: A Deep Quadruplet Network for Person Re-Identification

CVPR 2017spotlight

Person re-identification (ReID) is an important task in wide area video surveillance which focuses on identifying people across different cameras. Recently, deep learning networks with a triplet loss become a common framework for person ReID. However, the triplet loss pays main attentions on obtaini…

Cited by 1545PDFScholar
2017

Deep Crisp Boundaries

CVPR 2017poster

Edge detection had made significant progress with the help of deep Convolutional Networks (ConvNet). ConvNet based edge detectors approached human level performance on standard benchmarks. We provide a systematical study of these detector outputs, and show that they failed to accurately localize edg…

Cited by 142PDFScholar
2017

Learning Deep Context-Aware Features Over Body and Latent Parts for Person Re-Identification

CVPR 2017poster

Person Re-identification (ReID) is to identify the same person across different cameras. It is a challenging task due to the large variations in person pose, occlusion, background clutter, etc. How to extract powerful features is a fundamental problem in ReID and is still an open problem today. In t…

Cited by 829PDFScholar
2017

Locality-Sensitive Deconvolution Networks With Gated Fusion for RGB-D Indoor Semantic Segmentation

CVPR 2017poster

This paper focuses on indoor semantic segmentation using RGB-D data. Although the commonly used deconvolution networks (DeconvNet) have achieved impressive results on this task, we find there is still room for improvements in two aspects. One is about the boundary segmentation. DeconvNet aggregates…

Cited by 273PDFScholar
2016

ReD-SFA: Relation Discovery Based Slow Feature Analysis for Trajectory Clustering

CVPR 2016poster

For spectral embedding/clustering, it is still an open problem on how to construct an relation graph to reflect the intrinsic structures in data. In this paper, we proposed an approach, named Relation Discovery based Slow Feature Analysis (ReD-SFA), for feature learning and graph construction simult…

Cited by 15PDFScholar
2015

Beyond Tree Structure Models: A New Occlusion Aware Graphical Model for Human Pose Estimation

ICCV 2015poster

Occlusion is a main challenge for human pose estimation, which is largely ignored in popular tree structure models. The tree structure model is simple and convenient for exact inference, but short in modeling the occlusion coherence especially in the case of self-occlusion. We propose an occlusion a…

Cited by 28PDFScholar
2015

GRSA: Generalized Range Swap Algorithm for the Efficient Optimization of MRFs

CVPR 2015poster

Markov Random Field (MRF) is an important tool and has been widely used in many vision tasks. Thus, the optimization of MRFs is a problem of fundamental importance. Recently, Veskler and Kumar et. al propose the range move algorithms, which are one of the most successful solvers to this problem. How…

Cited by 8SourcePDFScholar
2015

Query Adaptive Similarity Measure for RGB-D Object Recognition

ICCV 2015poster

This paper studies the problem of improving the top-1 accuracy of RGB-D object recognition. Despite of the impressive top-5 accuracies achieved by existing methods, their top-1 accuracies are not very satisfactory. The reasons are in two-fold: (1) existing similarity measures are sensitive to object…

Cited by 18PDFScholar