← Search

Houqiang Li

150 accepted papers

2026

CoD: A Diffusion Foundation Model for Image Compression

CVPR 2026

Existing diffusion codecs typically build on text-to-image diffusion foundation models like Stable Diffusion.However, text conditioning is suboptimal from a compression perspective, hindering the potential of downstream diffusion codecs, particularly at ultra-low bitrates.To address it, we introduce

Cited by 0SourcecodeScholar
2026

D-Nav: End-to-End Dynamic UAV Navigation with Dual-Resolution Motion Awareness

RSS 2026poster

Autonomous navigation in dense, dynamic clutter remains a fundamental challenge for Unmanned Aerial Vehicles (UAVs) due to the heterogeneous obstacle scales and complex motion patterns. Existing methods often rely on fragile explicit tracking or noise-sensitive implicit flow estimation, both of whic…

Cited by 0SourceScholar
2026

Disentangling Length Bias in Preference Learning via Response-Conditioned Modeling

ICLR 2026poster

Reinforcement Learning from Human Feedback (RLHF) has achieved considerable success in aligning large language models (LLMs) by modeling human preferences with a learnable reward model and employing a reinforcement learning algorithm to maximize the reward model's scores. However, these reward model…

Cited by 0SourceScholar
2026

DocR1: Evidence Page-Guided GRPO for Multi-Page Document Understanding

AAAI 2026technical

Understanding multi-page documents poses a significant challenge for multimodal large language models (MLLMs), as it requires fine-grained visual comprehension and multi-hop reasoning across pages. While prior work has explored reinforcement learning (RL) for enhancing advanced reasoning in MLLMs, i

Cited by 0SourcePDFScholar
2026

EchoGen: Generating Visual Echoes in Any Scene via Feed-Forward Subject-Driven Auto-Regressive Model

ICLR 2026poster

Subject-driven generation is a critical task in creative AI; yet current state-of-the-art methods present a stark trade-off. They either rely on computationally expensive, per-subject fine-tuning, sacrificing efficiency and zero-shot capability, or employ feed-forward architectures built on diffusio…

Cited by 0SourcecodeScholar
2026

Exploring Mode Connectivity in Krylov Subspace for Domain Generalization

ICLR 2026poster

This paper explores the geometric characteristics of loss landscapes to enhance domain generalization (DG) in deep neural networks. Existing methods mainly leverage the local flatness around minima for improved generalization. However, recent theoretical studies indicate that flatness does not univ…

Cited by 0SourceScholar
2026

Flatness Guided Test-Time Adaptation for Vision-Language Models

ICLR 2026poster

Test-time adaptation (TTA) of Vision-Language Models (VLMs) has emerged as a technique for tackling distribution shifts during the test time. Recent research indicates that the test-time adaptation is intrinsically linked to the model's training history. However, existing TTA methods, such as Test-…

Cited by 0SourceScholar
2026

Gait-Adaptive Perceptive Humanoid Locomotion With Real-Time Under-Base Terrain Reconstruction

RA-L 2026

For full-size humanoid robots, reliable locomotion on complex terrains—such as long staircases—remains challenging, even with recent advances in reinforcement-learning-based control. In such settings, limited perception, ambiguous terrain cues, and insufficient adaptation of gait timing can cause ev

Cited by 9SourcecodeScholar
2026

Generative Video Compression with One-Dimensional Latent Representation

CVPR 2026

Recent advancements in generative video codec (GVC) typically encode video into a 2D latent grid and employ high-capacity generative decoders for reconstruction. However, this paradigm still leaves two key challenges in fully exploiting spatial-temporal redundancy: Spatially, the 2D latent grid inev

Cited by 0SourceScholar
2026

Geometry-Aware Dataset Condensation for Diffusion Model Training

ICML 2026poster

Dataset condensation aims to construct compact datasets from real data via synthesis or selection. However, existing approaches are ill-suited for diffusion model training: synthetic data generation often yields low-fidelity samples unsuitable for authentic modeling, while real subset selection typi…

Cited by 0SourceScholar
2026

Learning Surgical Robotic Manipulation with 3D Spatial Priors

CVPR 2026

Achieving 3D spatial awareness is crucial for surgical robotic manipulation, where precise and delicate operations are required. Existing methods either explicitly reconstruct the surgical scene prior to manipulation, or enhance multi-view features by adding wrist-mounted cameras to supplement the d

Cited by 0SourceScholar
2026

Primary-Fine Decoupling for Action Generation in Robotic Imitation

ICLR 2026poster

Multi-modal distribution in robotic manipulation action sequences poses critical challenges for imitation learning. To this end, existing approaches often model the action space as either a discrete set of tokens or a continuous, latent-variable distribution. However, both approaches present trade-…

Cited by 0SourceScholar
2026

Real-Time and Lightweight Diffusion Image Compression

ICML 2026poster

Recent advanced diffusion methods typically derive strong generative priors by scaling diffusion transformers. However, scaling fails to generalize when adapted for real-time compression scenarios that demand lightweight models. In this paper, we explore the design of real-time and lightweight diffu…

Cited by 0SourceScholar
2026

Rethinking Long-tailed Dataset Distillation: A Uni-Level Framework with Unbiased Recovery and Relabeling

AAAI 2026technical

Dataset distillation creates a small distilled set that enables efficient training by capturing key information from the full dataset. While existing dataset distillation methods perform well on balanced datasets, they struggle under long-tailed distributions, where imbalanced class frequencies indu

Cited by 0SourcePDFScholar
2026

Self-supervised Hierarchical Visual Reasoning with World Model

ICML 2026poster

3D open-world environments with adversarial opponents remain a core challenge for reinforcement learning due to their vast state spaces. Effective reasoning representations are essential in such settings. While existing self-supervised visual foresight reasoning approaches often suffer from multi-st…

Cited by 0SourceScholar
2026

Structural Action Transformer for 3D Dexterous Manipulation

CVPR 2026

Achieving human-level dexterity in robots via imitation learning from heterogeneous datasets is hindered by the challenge of cross-embodiment skill transfer, particularly for high-DoF robotic hands. Existing methods, often relying on 2D observations and temporal-centric action representation, strugg

Cited by 0SourcecodeScholar
2026

UAST: Unified Active Search and Tracking for Arbitrary Targets with UAVs

CVPR 2026

Active search and tracking of arbitrary targets by Unmanned Aerial Vehicles (UAVs) in cluttered environments remains a highly challenging problem. Existing methods either construct complex modular pipelines, leading to substantial computational costs, or adopt end-to-end controllers that often fail

Cited by 0SourcecodeScholar
2026

VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models

ICLR 2026poster

Visual reasoning is a core component of human intelligence and a critical capability for advanced multimodal models. Yet current reasoning evaluations of multimodal large language models (MLLMs) often rely on text descriptions and allow language-based reasoning shortcuts, failing to measure genuine…

Cited by 0SourcecodeScholar
2025

Active Perception Meets Rule-Guided RL: A Two-Phase Approach for Precise Object Navigation in Complex Environments

ICCV 2025poster

Object Goal Navigation (ObjectNav) in unknown environments presents significant challenges, particularly in Open-Vocabulary Mobile Manipulation (OVMM), where robots must efficiently explore large spaces, locate small objects, and accurately position themselves for subsequent manipulation. Existing a…

2025

Controllable Style Arithmetic with Language Models

ACL 2025long

Language models have shown remarkable capabilities in text generation, but precisely controlling their linguistic style remains challenging. Existing methods either lack fine-grained control, require extensive computation, or introduce significant latency. We propose Style Arithmetic (SA), a novel p…

2025

DP-Habitat: Bridging the Gap Between Simulation and Reality for Visual Navigation in Dynamic Pedestrian Environments

ICRA 2025

Visual navigation in dynamic environments poses a considerable challenge, particularly in scenarios with diverse pedestrian behaviors. Traditional simulators primarily focus on static scenes, while existing dynamic pedestrian simulators often suffer limitations such as monotonous pedestrian models,

Cited by 0SourcecodeScholar
2025

Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding

NeurIPS 2025poster

Long-form video understanding presents significant challenges due to extensive temporal-spatial complexity and the difficulty of question answering under such extended contexts. While Large Language Models (LLMs) have demonstrated considerable advancements in video analysis capabilities and long co…

Cited by 0SourceScholar
2025

DesignDiffusion: High-Quality Text-to-Design Image Generation with Diffusion Models

CVPR 2025poster

In this paper, we present DesignDiffusion, a simple yet effective framework for the novel task of synthesizing design images from textual descriptions. A primary challenge lies in generating accurate and style-consistent textual and visual content. Existing works in a related task of visual text gen…

2025

EG4D: Explicit Generation of 4D Object without Score Distillation

ICLR 2025poster

In recent years, the increasing demand for dynamic 3D assets in design and gaming applications has given rise to powerful generative pipelines capable of synthesizing high-quality 4D objects. Previous methods generally rely on score distillation sampling (SDS) algorithm to infer the unseen views a…

2025

Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling

NeurIPS 2025poster

Outcome‑reward reinforcement learning (RL) is a common—and increasingly significant—way to refine the step‑by‑step reasoning of multimodal large language models (MLLMs). In the multiple‑choice setting—a dominant format for multimodal reasoning benchmarks—the paradigm faces a significant yet often ov…

Cited by 0SourceScholar
2025

Incremental Transformer: Efficient Encoder for Incremented Text Over MRC and Conversation Tasks

COLING 2025main

Some encoder inputs such as conversation histories are frequently extended with short additional inputs like new responses. However, to obtain the real-time encoding of the extended input, existing Transformer-based encoders like BERT have to encode the whole extended input again without utilizing t…

Cited by 0SourcePDFScholar
2025

Interpret and Improve In-Context Learning via the Lens of Input-Label Mappings

ACL 2025long

Large language models (LLMs) excel at downstream NLP tasks through in-context learning (ICL) with a few demonstrations of input–label pairs. However, the internal mechanisms behind ICL remain under-explored, particularly the mappings between inputs and labels. In this work, we reverse-engineer ICL b…

Cited by 0SourcePDFScholar
2025

Leveraging Visual Captions for Enhanced Zero-Shot HOI Detection

ICASSP 2025accepted

Zero-shot Human-Object Interaction (HOI) detection aims to identify both seen and unseen HOI categories in an image. Most existing methods rely on semantic knowledge distilled from CLIP to find novel interactions but fail to fully exploit the powerful generalization ability of vision-language models…

Cited by 0SourceScholar
2025

Make-It-Animatable: An Efficient Framework for Authoring Animation-Ready 3D Characters

CVPR 2025highlight

3D characters are essential to modern creative industries, but making them animatable often demands extensive manual work in tasks like rigging and skinning. Existing automatic rigging tools face several limitations, including the necessity for manual annotations, rigid skeleton topologies, and limi…

2025

Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering

NeurIPS 2025poster

Multimodal large language models (MLLMs) have achieved remarkable progress in video understanding. However, hallucination, where the model generates plausible yet incorrect outputs, persists as a significant and under-addressed challenge in the video domain. Among existing solutions, activation engi…

Cited by 0SourceScholar
2025

Multi-Level Optimal Transport for Universal Cross-Tokenizer Knowledge Distillation on Language Models

AAAI 2025technical

Knowledge distillation (KD) has become a prevalent technique for compressing large language models (LLMs). Existing KD methods are constrained by the need for identical tokenizers (i.e., vocabularies) between teacher and student models, limiting their versatility in handling LLMs of different archit…

2025

OPTICAL: Leveraging Optimal Transport for Contribution Allocation in Dataset Distillation

CVPR 2025highlight

The demands for increasingly large-scale datasets pose substantial storage and computation challenges to building deep learning models. Dataset distillation methods, especially those via sample generation techniques, rise in response to condensing large original datasets into small synthetic ones wh…

Cited by 0SourcePDFScholar
2025

Optimizing Distributional Geometry Alignment with Optimal Transport for Generative Dataset Distillation

NeurIPS 2025poster

Dataset distillation seeks to synthesize a compact distilled dataset, enabling models trained on it to achieve performance comparable to models trained on the full dataset. Recent methods for large-scale datasets focus on matching global distributional statistics (e.g., mean and variance), but overl…

Cited by 0SourceScholar
2025

RaCFormer: Towards High-Quality 3D Object Detection via Query-based Radar-Camera Fusion

CVPR 2025poster

We propose Radar-Camera fusion transformer (RaCFormer) to boost the accuracy of 3D object detection by the following insight. The Radar-Camera fusion in outdoor 3D scene perception is capped by the image-to-BEV transformation-if the depth of pixels is not accurately estimated, the naive combination…

2025

Robust Multimodal Large Language Models Against Modality Conflict

ICML 2025poster

Despite the impressive capabilities of multimodal large language models (MLLMs) in vision-language tasks, they are prone to hallucinations in real-world scenarios. This paper investigates the hallucination phenomenon in MLLMs from the perspective of modality conflict. Unlike existing works focusing…

Cited by 0SourcePDFScholar
2025

S3R-GS: Streamlining the Pipeline for Large-Scale Street Scene Reconstruction

ICCV 2025poster

Recently, 3D Gaussian Splatting (3DGS) has reshaped the field of photorealistic 3D reconstruction, achieving impressive rendering quality and speed. However, when applied to large-scale street scenes, existing methods suffer from rapidly escalating per-viewpoint reconstruction costs as scene size in…

2025

Self-Classification Enhancement and Correction for Weakly Supervised Object Detection

IJCAI 2025

In recent years, weakly supervised object detection (WSOD) has attracted much attention due to its low labeling cost. The success of recent WSOD models is often ascribed to the two-stage multi-class classification (MCC) task, i.e., multiple instance learning and online classification refinement. Des

Cited by 0SourcePDFScholar
2025

SmartEraser: Remove Anything from Images using Masked-Region Guidance

CVPR 2025poster

Object removal has so far been dominated by the mask-and-inpaint paradigm, where the masked region is excluded from the input, leaving models relying on unmasked areas to inpaint the missing region. However, this approach lacks contextual information for the masked area, often resulting in unstable…

Cited by 2SourcePDFScholar
2025

TinySAM: Pushing the Envelope for Efficient Segment Anything Model

AAAI 2025technical

Recently segment anything model (SAM) has shown powerful segmentation capability and has drawn great attention in computer vision fields. Massive following works have developed various applications based on the pre-trained SAM and achieved impressive performance on downstream vision tasks. However,…

2025

Towards Practical Real-Time Neural Video Compression

CVPR 2025poster

We introduce a practical real-time neural video codec (NVC) designed to deliver high compression ratio, low latency and broad versatility. In practice, the coding speed of NVCs depends on 1) computational costs, and 2) non-computational operational costs, such as memory I/O and the number of functio…

2025

Uni-Sign: Toward Unified Sign Language Understanding at Scale

ICLR 2025poster

Sign language pre-training has gained increasing attention for its ability to enhance performance across various sign language understanding (SLU) tasks. However, existing methods often suffer from a gap between pre-training and fine-tuning, leading to suboptimal results. To address this, we propose…

2025

Visual Evidence Prompting Mitigates Hallucinations in Large Vision-Language Models

ACL 2025long

Large Vision-Language Models (LVLMs) have shown impressive progress by integrating visual perception with linguistic understanding to produce contextually grounded outputs. Despite these advancements achieved, LVLMs still suffer from the hallucination problem, e.g., they tend to produce content that…

Cited by 0SourcePDFScholar
2024

BoolQuestions: Does Dense Retrieval Understand Boolean Logic in Language?

EMNLP 2024finding

Dense retrieval, which aims to encode the semantic information of arbitrary text into dense vector representations or embeddings, has emerged as an effective and efficient paradigm for text retrieval, consequently becoming an essential component in various natural language processing systems. These…

2024

Forest2Seq: Revitalizing Order Prior for Sequential Indoor Scene Synthesis

ECCV 2024poster

"Synthesizing realistic 3D indoor scenes is a challenging task that traditionally relies on manual arrangement and annotation by expert designers. Recent advances in autoregressive models have automated this process, but they often lack semantic understanding of the relationships and hierarchies pre…

Cited by 7SourcePDFScholar
2024

From Yes-Men to Truth-Tellers: Addressing Sycophancy in Large Language Models with Pinpoint Tuning

ICML 2024poster

Large Language Models (LLMs) tend to prioritize adherence to user prompts over providing veracious responses, leading to the sycophancy issue. When challenged by users, LLMs tend to admit mistakes and provide inaccurate responses even if they initially provided the correct answer. Recent works propo…

Cited by 10SourcePDFScholar
2024

Generative Latent Coding for Ultra-Low Bitrate Image Compression

CVPR 2024poster

Most existing image compression approaches perform transform coding in the pixel space to reduce its spatial redundancy. However they encounter difficulties in achieving both high-realism and high-fidelity at low bitrate as the pixel-space distortion may not align with human perception. To address t…

Cited by 13SourcePDFScholar
2024

Image2Sentence based Asymmetrical Zero-shot Composed Image Retrieval

ICLR 2024spotlight

The task of composed image retrieval (CIR) aims to retrieve images based on the query image and the text describing the users' intent. Existing methods have made great progress with the advanced large vision-language (VL) model in CIR task, however, they generally suffer from two main issues: lack…

Cited by 12SourcePDFScholar
2024

Instance-aware Exploration-Verification-Exploitation for Instance ImageGoal Navigation

CVPR 2024poster

As a new embodied vision task Instance ImageGoal Navigation (IIN) aims to navigate to a specified object depicted by a goal image in an unexplored environment. The main challenge of this task lies in identifying the target object from different viewpoints while rejecting similar distractors. Existin…

2024

InstructDiffusion: A Generalist Modeling Interface for Vision Tasks

CVPR 2024poster

We present InstructDiffusion a unified and generic framework for aligning computer vision tasks with human instructions. Unlike existing approaches that integrate prior knowledge and pre-define the output space (e.g. categories and coordinates) for each vision task we cast diverse vision tasks into…

Cited by 109SourcePDFScholar
2024

Interpretable Composition Attribution Enhancement for Visio-linguistic Compositional Understanding

EMNLP 2024main

Contrastively trained vision-language models such as CLIP have achieved remarkable progress in vision and language representation learning. Despite the promising progress, their proficiency in compositional reasoning over attributes and relations (e.g., distinguishing between “the car is underneath…

Cited by 0SourcePDFScholar
2024

KGDM: A Diffusion Model to Capture Multiple Relation Semantics for Knowledge Graph Embedding

AAAI 2024technical

Knowledge graph embedding (KGE) is an efficient and scalable method for knowledge graph completion. However, most existing KGE methods suffer from the challenge of multiple relation semantics, which often degrades their performance. This is because most KGE methods learn fixed continuous vectors for…

2024

Learning Label Dependencies for Visual Information Extraction

IJCAI 2024poster

Visual Information Extraction (VIE), which aims to extract structured information from visually rich document images, has drawn much attention due to its wide applications in document understanding. However, previous methods often treat the VIE task as a sequence labeling problem and ignore the labe…

Cited by 0SourcePDFScholar
2024

Long-term Temporal Context Gathering for Neural Video Compression

ECCV 2024poster

"Most existing neural video codecs (NVCs) only extract short-term temporal context by optical flow-based motion compensation. However, such short-term temporal context suffers from error propagation and lacks awareness of long-term relevant information. This limits their performance, particularly in…

2024

Revisiting Open-Set Panoptic Segmentation

AAAI 2024technical

In this paper, we focus on the open-set panoptic segmentation (OPS) task to circumvent the data explosion problem. Different from the close-set setting, OPS targets to detect both known and unknown categories, where the latter is not annotated during training. Different from existing work that only…

Cited by 0SourcePDFScholar
2024

SUF: Stabilized Unconstrained Fine-Tuning for Offline-to-Online Reinforcement Learning

AAAI 2024technical

Offline-to-online reinforcement learning (RL) provides a promising solution to improving suboptimal offline pre-trained policies through online fine-tuning. However, one efficient method, unconstrained fine-tuning, often suffers from severe policy collapse due to excessive distribution shift. To ens…

Cited by 3SourcePDFScholar
2024

Sinkhorn Distance Minimization for Knowledge Distillation

COLING 2024main

Knowledge distillation (KD) has been widely adopted to compress large language models (LLMs). Existing KD methods investigate various divergence measures including the Kullback-Leibler (KL), reverse Kullback-Leibler (RKL), and Jensen-Shannon (JS) divergences. However, due to limitations inherent in…

2024

TabPedia: Towards Comprehensive Visual Table Understanding with Concept Synergy

NeurIPS 2024poster

Tables contain factual and quantitative data accompanied by various structures and contents that pose challenges for machine comprehension. Previous methods generally design task-specific architectures and objectives for individual tasks, resulting in modal isolation and intricate workflows. In this…

2024

Trustworthy Alignment of Retrieval-Augmented Large Language Models via Reinforcement Learning

ICML 2024poster

Trustworthiness is an essential prerequisite for the real-world application of large language models. In this paper, we focus on the trustworthiness of language models with respect to retrieval augmentation. Despite being supported with external evidence, retrieval-augmented generation still suffers…

2023

$\mathcal{O}$-GNN: incorporating ring priors into molecular modeling

ICLR 2023poster

Cyclic compounds that contain at least one ring play an important role in drug design. Despite the recent success of molecular modeling with graph neural networks (GNNs), few models explicitly take rings in compounds into consideration, consequently limiting the expressiveness of the models. In this…

Cited by 0SourcePDFScholar
2023

AltFreezing for More General Video Face Forgery Detection

CVPR 2023highlight

Existing face forgery detection models try to discriminate fake images by detecting only spatial artifacts (e.g., generative artifacts, blending) or mainly temporal artifacts (e.g., flickering, discontinuity). They may experience significant performance degradation when facing out-domain artifacts.…

2023

BEST: BERT Pre-training for Sign Language Recognition with Coupling Tokenization

AAAI 2023technical

In this work, we are dedicated to leveraging the BERT pre-training success and modeling the domain-specific statistics to fertilize the sign language recognition~(SLR) model. Considering the dominance of hand and body in sign language expression, we organize them as pose triplet units and feed them…

Cited by 37SourcePDFScholar
2023

CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Alignment

ICLR 2023poster

Pre-trained image-text models, like CLIP, have demonstrated the strong power of vision-language representation learned from a large scale of web-collected image-text data. In light of the well-learned visual features, there are works that transfer image representation to the video domain and achieve…

2023

CLIP4HOI: Towards Adapting CLIP for Practical Zero-Shot HOI Detection

NeurIPS 2023poster

Zero-shot Human-Object Interaction (HOI) detection aims to identify both seen and unseen HOI categories. A strong zero-shot HOI detector is supposed to be not only capable of discriminating novel interactions but also robust to positional distribution discrepancy between seen and unseen categories w…

Cited by 21SourcePDFScholar
2023

Cyclic-Bootstrap Labeling for Weakly Supervised Object Detection

ICCV 2023poster

Recent progress in weakly supervised object detection is featured by a combination of multiple instance detection networks (MIDN) and ordinal online refinement. However, with only image-level annotation, MIDN inevitably assigns high scores to some unexpected region proposals when generating pseudo l…

Cited by 12PDFcodeScholar
2023

DIFFER:Decomposing Individual Reward for Fair Experience Replay in Multi-Agent Reinforcement Learning

NeurIPS 2023poster

Cooperative multi-agent reinforcement learning (MARL) is a challenging task, as agents must learn complex and diverse individual strategies from a shared team reward. However, existing methods struggle to distinguish and exploit important individual experiences, as they lack an effective way to deco…

Cited by 2SourcePDFScholar
2023

DIRE for Diffusion-Generated Image Detection

ICCV 2023poster

Diffusion models have shown remarkable success in visual synthesis, but have also raised concerns about potential abuse for malicious purposes. In this paper, we seek to build a detector for telling apart real images from diffusion-generated images. We find that existing detectors struggle to detect…

Cited by 221PDFcodeScholar
2023

Focus on Your Target: A Dual Teacher-Student Framework for Domain-Adaptive Semantic Segmentation

ICCV 2023poster

We study unsupervised domain adaptation (UDA) for semantic segmentation. Currently, a popular UDA framework lies in self-training which endows the model with two-fold abilities: (i) learning reliable semantics from the labeled images in the source domain, and (ii) adapting to the target domain via g…

Cited by 12PDFcodeScholar
2023

HandNeRF: Neural Radiance Fields for Animatable Interacting Hands

CVPR 2023poster

We propose a novel framework to reconstruct accurate appearance and geometry with neural radiance fields (NeRF) for interacting hands, enabling the rendering of photo-realistic images and videos for gesture animation from arbitrary views. Given multi-view images of a single hand or interacting hands…

Cited by 29SourcePDFScholar
2023

Learning robust representation for reinforcement learning with distractions by reward sequence prediction

UAI 2023poster

Reinforcement learning algorithms have achieved remarkable success in acquiring behavioral skills directly from pixel inputs. However, their application in real-world scenarios presents challenges due to their sensitivity to visual distractions (e.g., changes in viewpoint and light). A key factor co…

2023

Low-Light Video Enhancement with Synthetic Event Guidance

AAAI 2023technical

Low-light video enhancement (LLVE) is an important yet challenging task with many applications such as photographing and autonomous driving. Unlike single image low-light enhancement, most LLVE methods utilize temporal information from adjacent frames to restore the color and remove the noise of the…

Cited by 29SourcePDFScholar
2023

MA2CL:Masked Attentive Contrastive Learning for Multi-Agent Reinforcement Learning

IJCAI 2023poster

Recent approaches have utilized self-supervised auxiliary tasks as representation learning to improve the performance and sample efficiency of vision-based reinforcement learning algorithms in single-agent settings. However, in multi-agent reinforcement learning (MARL), these techniques face challen…

2023

Making Better Decision by Directly Planning in Continuous Control

ICLR 2023poster

By properly utilizing the learned environment model, model-based reinforcement learning methods can improve the sample efficiency for decision-making problems. Beyond using the learned environment model to train a policy, the success of MCTS-based methods shows that directly incorporating the learne…

2023

Masked Motion Predictors are Strong 3D Action Representation Learners

ICCV 2023poster

In 3D human action recognition, limited supervised data makes it challenging to fully tap into the modeling potential of powerful networks such as transformers. As a result, researchers have been actively investigating effective self-supervised pre-training strategies. In this work, we show that ins…

Cited by 46PDFcodeScholar
2023

Multi-Agent First Order Constrained Optimization in Policy Space

NeurIPS 2023poster

In the realm of multi-agent reinforcement learning (MARL), achieving high performance is crucial for a successful multi-agent system. Meanwhile, the ability to avoid unsafe actions is becoming an urgent and imperative problem to solve for real-life applications. Whereas, it is still challenging to…

Cited by 3SourcePDFScholar
2023

NUWA-XL: Diffusion over Diffusion for eXtremely Long Video Generation

ACL 2023long

In this paper, we propose NUWA-XL, a novel Diffusion over Diffusion architecture for eXtremely Long video generation. Most current work generates long videos segment by segment sequentially, which normally leads to the gap between training on short videos and inferring long videos, and the sequentia…

Cited by 118SourcePDFScholar
2023

SimFIR: A Simple Framework for Fisheye Image Rectification with Self-supervised Representation Learning

ICCV 2023poster

In fisheye images, rich distinct distortion patterns are regularly distributed in the image plane. These distortion patterns are independent of the visual content and provide informative cues for rectification. To make the best of such rectification cues, we introduce SimFIR, a simple framework for…

Cited by 31PDFScholar
2023

Stare at What You See: Masked Image Modeling Without Reconstruction

CVPR 2023poster

Masked Autoencoders (MAE) have been prevailing paradigms for large-scale vision representation pre-training. By reconstructing masked image patches from a small portion of visible image regions, MAE forces the model to infer semantic correlation within an image. Recently, some approaches apply seman…

Cited by 33SourcePDFScholar
2023

State Sequences Prediction via Fourier Transform for Representation Learning

NeurIPS 2023spotlight

While deep reinforcement learning (RL) has been demonstrated effective in solving complex control tasks, sample efficiency remains a key challenge due to the large amounts of data required for remarkable performance. Existing research explores the application of representation learning for data-effi…

2022

${\mathsf{EZFusion}}$: A Close Look at the Integration of LiDAR, Millimeter-Wave Radar, and Camera for Accurate 3D Object Detection and Tracking

RA-L 2022

A recent trend is to combine multiple sensors ( <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">i.e.</i> , cameras, LiDARs and millimeter-wave Radars) to achieve robust multi-modal perception for autonomous systems such as self-driving vehicles. Alth

Cited by 11SourceScholar
2022

CMD: Self-Supervised 3D Action Representation Learning with Cross-Modal Mutual Distillation

ECCV 2022poster

"In 3D action recognition, there exists rich complementary information between skeleton modalities. Nevertheless, how to model and utilize this information remains a challenging problem for self-supervised 3D action representation learning. In this work, we formulate the cross-modal interaction as a…

2022

CMT: Context-Matching-Guided Transformer for 3D Tracking in Point Clouds

ECCV 2022poster

"How to effectively match the target template features with the search area is the core problem in point-cloud-based 3D single object tracking. However, in the literature, most of the methods focus on devising sophisticated matching modules at point-level, while overlooking the rich spatial context…

Cited by 27SourcePDFScholar
2022

Domain-Agnostic Prior for Transfer Semantic Segmentation

CVPR 2022poster

Unsupervised domain adaptation (UDA) is an important topic in the computer vision community. The key difficulty lies in defining a common property between the source and target domains so that the source-domain features can align with the target-domain semantics. In this paper, we present a simple a…

Cited by 45PDFScholar
2022

Equivalence Analysis between Counterfactual Regret Minimization and Online Mirror Descent

ICML 2022spotlight

Follow-the-Regularized-Leader (FTRL) and Online Mirror Descent (OMD) are regret minimization algorithms for Online Convex Optimization (OCO), they are mathematically elegant but less practical in solving Extensive-Form Games (EFGs). Counterfactual Regret Minimization (CFR) is a technique for approxi…

2022

Geometric Representation Learning for Document Image Rectification

ECCV 2022poster

"In document image rectification, there exist rich geometric constraints between the distorted image and the ground truth one. How- ever, such geometric constraints are largely ignored in existing advanced solutions, which limits the rectification performance. To this end, we present DocGeoNet for d…

2022

LDSA: Learning Dynamic Subtask Assignment in Cooperative Multi-Agent Reinforcement Learning

NeurIPS 2022accept

Cooperative multi-agent reinforcement learning (MARL) has made prominent progress in recent years. For training efficiency and scalability, most of the MARL algorithms make all agents share the same policy or value network. However, in many complex multi-agent tasks, different agents are expected to…

Cited by 45SourcePDFScholar
2022

Large-Scale Pre-Training for Person Re-Identification With Noisy Labels

CVPR 2022poster

This paper aims to address the problem of pre-training for person re-identification (Re-ID) with noisy labels. To setup the pre-training task, we apply a simple online multi-object tracking system on raw videos of an existing unlabeled Re-ID dataset "LUPerson" and build the Noisy Labeled variant cal…

Cited by 81PDFcodeScholar
2022

Learning Robust Policy against Disturbance in Transition Dynamics via State-Conservative Policy Optimization

AAAI 2022technical

Deep reinforcement learning algorithms can perform poorly in real-world tasks due to the discrepancy between source and target environments. This discrepancy is commonly viewed as the disturbance in transition dynamics. Many existing algorithms learn robust policies by modeling the disturbance and a…

Cited by 24SourcePDFScholar
2022

Learning Token-Based Representation for Image Retrieval

AAAI 2022technical

In image retrieval, deep local features learned in a data-driven manner have been demonstrated effective to improve retrieval performance. To realize efficient retrieval on large image database, some approaches quantize deep local features with a large codebook and match images with aggregated match…

2022

Neural-based Mixture Probabilistic Query Embedding for Answering FOL queries on Knowledge Graphs

EMNLP 2022main

Query embedding (QE)—which aims to embed entities and first-order logical (FOL) queries in a vector space, has shown great power in answering FOL queries on knowledge graphs (KGs). Existing QE methods divide a complex query into a sequence of mini-queries according to its computation graph and perfo…

Cited by 11SourcePDFScholar
2022

Sample-Efficient Reinforcement Learning via Conservative Model-Based Actor-Critic

AAAI 2022technical

Model-based reinforcement learning algorithms, which aim to learn a model of the environment to make decisions, are more sample efficient than their model-free counterparts. The sample efficiency of model-based approaches relies on whether the model can well approximate the environment. However, lea…

Cited by 45SourcePDFScholar
2022

TAPE: Task-Agnostic Prior Embedding for Image Restoration

ECCV 2022poster

"Learning a generalized prior for natural image restoration is an important yet challenging task. Early methods mostly involved handcrafted priors including normalized sparsity, â„“0 gradients, dark channel priors, etc.. Recently, deep neural networks have been used to learn various image priors but…

Cited by 65SourcePDFScholar
2022

Uformer: A General U-Shaped Transformer for Image Restoration

CVPR 2022poster

In this paper, we present Uformer, an effective and efficient Transformer-based architecture for image restoration, in which we build a hierarchical encoder-decoder network using the Transformer block. In Uformer, there are two core designs. First, we introduce a novel locally-enhanced window (LeWin…

Cited by 1992PDFcodeScholar
2021

3D Local Convolutional Neural Networks for Gait Recognition

ICCV 2021poster

The goal of gait recognition is to learn the unique spatio-temporal pattern about the human body shape from its temporal changing characteristics. As different body parts behave differently during walking, it is intuitive to model the spatio-temporal patterns of each part separately. However, existi…

Cited by 134PDFcodeScholar
2021

ATSO: Asynchronous Teacher-Student Optimization for Semi-Supervised Image Segmentation

CVPR 2021poster

Semi-supervised learning is a useful tool for image segmentation, mainly due to its ability in extracting knowledge from unlabeled data to assist learning from labeled data. This paper focuses on a popular pipeline known as self-learning, where we point out a weakness named lazy mimicking that refer…

Cited by 75PDFScholar
2021

Auto-Encoding Transformations in Reparameterized Lie Groups for Unsupervised Learning

AAAI 2021technical

Unsupervised training of deep representations has demonstrated remarkable potentials in mitigating the prohibitive expenses on annotating labeled data recently. Among them is predicting transformations as a pretext task to self-train representations, which has shown great potentials for unsupervised…

Cited by 5SourcePDFScholar
2021

BANG: Bridging Autoregressive and Non-autoregressive Generation with Large Scale Pretraining

ICML 2021spotlight

In this paper, we propose BANG, a new pretraining model to Bridge the gap between Autoregressive (AR) and Non-autoregressive (NAR) Generation. AR and NAR generation can be uniformly regarded as to what extent previous tokens can be attended, and BANG bridges AR and NAR generation through designing a…

2021

Conditional DETR for Fast Training Convergence

ICCV 2021poster

The recently-developed DETR approach applies the transformer encoder and decoder architecture to object detection and achieves promising performance. In this paper, we handle the critical issue, slow training convergence, and present a conditional cross-attention mechanism for fast DETR training. Ou…

Cited by 829PDFcodeScholar
2021

Contextual Similarity Aggregation with Self-attention for Visual Re-ranking

NeurIPS 2021poster

In content-based image retrieval, the first-round retrieval result by simple visual feature comparison may be unsatisfactory, which can be refined by visual re-ranking techniques. In image retrieval, it is observed that the contextual similarity among the top-ranked images is an important clue to di…

2021

Contrastive Transformation for Self-supervised Correspondence Learning

AAAI 2021technical

In this paper, we focus on the self-supervised learning of visual correspondence using unlabeled videos in the wild. Our method simultaneously considers intra- and inter-video representation associations for reliable correspondence estimation. The intra-video learning transforms the image contents a…

2021

Discovering Representation Sprachbund For Multilingual Pre-Training

EMNLP 2021finding

Multilingual pre-trained models have demonstrated their effectiveness in many multilingual NLP tasks and enabled zero-shot or few-shot transfer from high-resource languages to low-resource ones. However, due to significant typological differences and contradictions between some languages, such model…

2021

Dual Progressive Prototype Network for Generalized Zero-Shot Learning

NeurIPS 2021poster

Generalized Zero-Shot Learning (GZSL) aims to recognize new categories with auxiliary semantic information, e.g., category attributes. In this paper, we handle the critical issue of domain shift problem, i.e., confusion between seen and unseen categories, by progressively improving cross-domain tran…

Cited by 57SourcePDFScholar
2021

Fine-grained Semantic Alignment Network for Weakly Supervised Temporal Language Grounding

EMNLP 2021finding

Temporal language grounding (TLG) aims to localize a video segment in an untrimmed video based on a natural language description. To alleviate the expensive cost of manual annotations for temporal boundary labels,we are dedicated to the weakly supervised setting, where only video-level descriptions…

Cited by 20SourcePDFScholar
2021

Generating Diverse Structure for Image Inpainting With Hierarchical VQ-VAE

CVPR 2021poster

Given an incomplete image without additional constraint, image inpainting natively allows for multiple solutions as long as they appear plausible. Recently, multiple-solution inpainting methods have been proposed and shown the potential of generating diverse results. However, these methods have diff…

Cited by 274PDFcodeScholar
2021

IOT: Instance-wise Layer Reordering for Transformer Structures

ICLR 2021poster

With sequentially stacked self-attention, (optional) encoder-decoder attention, and feed-forward layers, Transformer achieves big success in natural language processing (NLP), and many variants have been proposed. Currently, almost all these models assume that the \emph{layer order} is fixed and kep…

2021

Improving Sign Language Translation With Monolingual Data by Sign Back-Translation

CVPR 2021poster

Despite existing pioneering works on sign language translation (SLT), there is a non-trivial obstacle, i.e., the limited quantity of parallel sign-text data. To tackle this parallel data bottleneck, we propose a sign back-translation (SignBT) approach, which incorporates massive spoken language text…

Cited by 249PDFScholar
2021

Instance Mining with Class Feature Banks for Weakly Supervised Object Detection

AAAI 2021technical

Recent progress on weakly supervised object detection (WSOD) is characterized by formulating WSOD as a Multiple Instance Learning (MIL) problem and taking online refinement with the selected region proposals from MIL. However, MIL inclines to select the most discriminative part rather than the entir…

Cited by 39SourcePDFScholar
2021

Instance-Wise Hard Negative Example Generation for Contrastive Learning in Unpaired Image-to-Image Translation

ICCV 2021poster

Contrastive learning shows great potential in unpaired image-to-image translation, but sometimes the translated results are in poor quality and the contents are not preserved consistently. In this paper, we uncover that the negative examples play a critical role in the performance of contrastive lea…

Cited by 104PDFScholar
2021

Joint Inductive and Transductive Learning for Video Object Segmentation

ICCV 2021poster

Semi-supervised video object segmentation is a task of segmenting the target object in a video sequence given only a mask annotation in the first frame. The limited information available makes it an extremely challenging task. Most previous best-performing methods adopt matching-based transductive r…

Cited by 122PDFcodeScholar
2021

Learning Deep Local Features With Multiple Dynamic Attentions for Large-Scale Image Retrieval

ICCV 2021poster

In image retrieval, learning local features with deep convolutional networks has been demonstrated effective to improve the performance. To discriminate deep local features, some research efforts turn to attention learning. However, existing attention-based methods only generate a single attention m…

Cited by 31PDFcodeScholar
2021

Probing Inter-modality: Visual Parsing with Self-Attention for Vision-and-Language Pre-training

NeurIPS 2021poster

Vision-Language Pre-training (VLP) aims to learn multi-modal representations from image-text pairs and serves for downstream vision-language tasks in a fine-tuning fashion. The dominant VLP models adopt a CNN-Transformer architecture, which embeds images with a CNN, and then aligns images and text w…

Cited by 92SourcePDFScholar
2021

Representing Videos As Discriminative Sub-Graphs for Action Recognition

CVPR 2021poster

Human actions are typically of combinatorial structures or patterns, i.e., subjects, objects, plus spatio-temporal interactions in between. Discovering such structures is therefore a rewarding way to reason about the dynamics of interactions and recognize the actions. In this paper, we introduce a n…

Cited by 34PDFScholar
2021

Revisiting Knowledge Distillation: An Inheritance and Exploration Framework

CVPR 2021poster

Knowledge Distillation (KD) is a popular technique to transfer knowledge from a teacher model or ensemble to a student model. Its success is generally attributed to the privileged information on similarities/consistency between the class distributions or intermediate feature representations of the t…

Cited by 41PDFcodeScholar
2021

SignBERT: Pre-Training of Hand-Model-Aware Representation for Sign Language Recognition

ICCV 2021poster

Hand gesture serves as a critical role in sign language. Current deep-learning-based sign language recognition (SLR) methods may suffer insufficient interpretability and overfitting due to limited sign data sources. In this paper, we introduce the first self-supervised pre-trainable SignBERT with in…

Cited by 104PDFScholar
2021

Task-Independent Knowledge Makes for Transferable Representations for Generalized Zero-Shot Learning

AAAI 2021technical

Generalized Zero-Shot Learning (GZSL) targets recognizing new categories by learning transferable image representations. Existing methods find that, by aligning image representations with corresponding semantic labels, the semantic-aligned representations can be transferred to unseen categories. How…

Cited by 18SourcePDFScholar
2021

TransVG: End-to-End Visual Grounding With Transformers

ICCV 2021poster

In this paper, we present a neat yet effective transformer-based framework for visual grounding, namely TransVG, to address the task of grounding a language query to the corresponding region onto an image. The state-of-the-art methods, including two-stage or one-stage ones, rely on a complex module…

Cited by 408PDFcodeScholar
2021

Transformer Meets Tracker: Exploiting Temporal Context for Robust Visual Tracking

CVPR 2021poster

In video object tracking, there exist rich temporal contexts among successive frames, which have been largely overlooked in existing trackers. In this work, we bridge the individual video frames and explore the temporal contexts across them via a transformer architecture for robust object tracking.…

Cited by 865PDFcodeScholar
2021

Unsupervised Pre-Training for Person Re-Identification

CVPR 2021poster

In this paper, we present a large scale unlabeled person re-identification (Re-ID) dataset "LUPerson" and make the first attempt of performing unsupervised pre-training for improving the generalization ability of the learned person Re-ID feature representation. This is to address the problem that al…

Cited by 226PDFcodeScholar
2021

Voxel R-CNN: Towards High Performance Voxel-based 3D Object Detection

AAAI 2021technical

Recent advances on 3D object detection heavily rely on how the 3D data are represented, i.e., voxel-based or point-based representation. Many existing high performance 3D detectors are point-based because this structure can better retain precise point positions. Nevertheless, point-level features le…

2020

Incorporating BERT into Neural Machine Translation

ICLR 2020poster

The recently proposed BERT (Devlin et al., 2019) has shown great power on a variety of natural language understanding tasks, such as text classification, reading comprehension, etc. However, how to effectively apply BERT to neural machine translation (NMT) lacks enough exploration. While BERT is mor…

Cited by 522SourcecodeScholar
2020

Promoting Stochasticity for Expressive Policies via a Simple and Efficient Regularization Method

NeurIPS 2020poster

Many recent reinforcement learning (RL) methods learn stochastic policies with entropy regularization for exploration and robustness. However, in continuous action spaces, integrating entropy regularization with expressive policies is challenging and usually requires complex inference procedures. To…

Cited by 8SourcePDFScholar
2020

Transformation GAN for Unsupervised Image Synthesis and Representation Learning

CVPR 2020poster

Generative Adversarial Networks (GAN) have shown promising performance in image synthesis and unsupervised learning (USL). In most cases, however, the representations extracted from unsupervised GAN are usually unsatisfactory in other computer vision tasks. By using conditional GAN (CGAN), this prob…

Cited by 31PDFScholar
2019

Relation Distillation Networks for Video Object Detection

ICCV 2019poster

It has been well recognized that modeling object-to-object relations would be helpful for object detection. Nevertheless, the problem is not trivial especially when exploring the interactions between objects to boost video object detectors. The difficulty originates from the aspect that reliable obj…

Cited by 280PDFScholar
2018

Affinity Derivation and Graph Merge for Instance Segmentation

ECCV 2018poster

We present an instance segmentation scheme based on pixel affinity information, which is the relationship of two pixels belonging to a same instance. In our scheme, we use two neural networks with similar structure. One is to predict pixel level semantic score and the other is designed to derive pix…

2018

Multi-Cue Correlation Filters for Robust Visual Tracking

CVPR 2018poster

In recent years, many tracking algorithms achieve impressive performance via fusing multiple types of features, however, most of them fail to fully explore the context among the adopted multiple features and the strength of them. In this paper, we propose an efficient multi-cue analysis framework fo…

2017

CVAE-GAN: Fine-Grained Image Generation Through Asymmetric Training

ICCV 2017poster

We present variational generative adversarial networks, a general learning framework that combines a variational auto-encoder with a generative adversarial network, for synthesizing images in fine-grained categories, such as faces of a specific person or objects in a category. Our approach models an…

Cited by 268PDFScholar
2016

Comparative Deep Learning of Hybrid Representations for Image Recommendations

CVPR 2016poster

In many image-related tasks, learning expressive and discriminative representations of images is essential, and deep learning has been studied for automating the learning of such representations. Some user-centric tasks, such as image recommendations, call for effective representations of not only i…

Cited by 162PDFScholar
2016

Jointly Modeling Embedding and Translation to Bridge Video and Language

CVPR 2016oral

Automatically describing video content with natural language is a fundamental challenge of computer vision. Recurrent Neural Networks (RNNs), which models sequence dynamics, has attracted increasing attention on visual interpretation. However, most existing approaches generate a word locally with th…

Cited by 716PDFScholar
2015

Semi-Supervised Domain Adaptation With Subspace Learning for Visual Recognition

CVPR 2015poster

In many real-world applications, we are often facing the problem of cross domain learning, i.e., to borrow the labeled data or transfer the already learnt knowledge from a source domain to a target domain. However, simply applying existing source data or knowledge may even hurt the performance, espe…

Cited by 276SourcePDFScholar
Houqiang Li — accepted AI-conference papers · AIConfPaper