← Search

Wengang Zhou

103 accepted papers

2026

D-Nav: End-to-End Dynamic UAV Navigation with Dual-Resolution Motion Awareness

RSS 2026poster

Autonomous navigation in dense, dynamic clutter remains a fundamental challenge for Unmanned Aerial Vehicles (UAVs) due to the heterogeneous obstacle scales and complex motion patterns. Existing methods often rely on fragile explicit tracking or noise-sensitive implicit flow estimation, both of whic…

Cited by 0SourceScholar
2026

Disentangling Length Bias in Preference Learning via Response-Conditioned Modeling

ICLR 2026poster

Reinforcement Learning from Human Feedback (RLHF) has achieved considerable success in aligning large language models (LLMs) by modeling human preferences with a learnable reward model and employing a reinforcement learning algorithm to maximize the reward model's scores. However, these reward model…

Cited by 0SourceScholar
2026

DocR1: Evidence Page-Guided GRPO for Multi-Page Document Understanding

AAAI 2026technical

Understanding multi-page documents poses a significant challenge for multimodal large language models (MLLMs), as it requires fine-grained visual comprehension and multi-hop reasoning across pages. While prior work has explored reinforcement learning (RL) for enhancing advanced reasoning in MLLMs, i

Cited by 0SourcePDFScholar
2026

Gait-Adaptive Perceptive Humanoid Locomotion With Real-Time Under-Base Terrain Reconstruction

RA-L 2026

For full-size humanoid robots, reliable locomotion on complex terrains—such as long staircases—remains challenging, even with recent advances in reinforcement-learning-based control. In such settings, limited perception, ambiguous terrain cues, and insufficient adaptation of gait timing can cause ev

Cited by 9SourcecodeScholar
2026

Geometry-Aware Dataset Condensation for Diffusion Model Training

ICML 2026poster

Dataset condensation aims to construct compact datasets from real data via synthesis or selection. However, existing approaches are ill-suited for diffusion model training: synthetic data generation often yields low-fidelity samples unsuitable for authentic modeling, while real subset selection typi…

Cited by 0SourceScholar
2026

Primary-Fine Decoupling for Action Generation in Robotic Imitation

ICLR 2026poster

Multi-modal distribution in robotic manipulation action sequences poses critical challenges for imitation learning. To this end, existing approaches often model the action space as either a discrete set of tokens or a continuous, latent-variable distribution. However, both approaches present trade-…

Cited by 0SourceScholar
2026

Rethinking Long-tailed Dataset Distillation: A Uni-Level Framework with Unbiased Recovery and Relabeling

AAAI 2026technical

Dataset distillation creates a small distilled set that enables efficient training by capturing key information from the full dataset. While existing dataset distillation methods perform well on balanced datasets, they struggle under long-tailed distributions, where imbalanced class frequencies indu

Cited by 0SourcePDFScholar
2026

Self-supervised Hierarchical Visual Reasoning with World Model

ICML 2026poster

3D open-world environments with adversarial opponents remain a core challenge for reinforcement learning due to their vast state spaces. Effective reasoning representations are essential in such settings. While existing self-supervised visual foresight reasoning approaches often suffer from multi-st…

Cited by 0SourceScholar
2026

Structural Action Transformer for 3D Dexterous Manipulation

CVPR 2026

Achieving human-level dexterity in robots via imitation learning from heterogeneous datasets is hindered by the challenge of cross-embodiment skill transfer, particularly for high-DoF robotic hands. Existing methods, often relying on 2D observations and temporal-centric action representation, strugg

Cited by 0SourcecodeScholar
2026

UAST: Unified Active Search and Tracking for Arbitrary Targets with UAVs

CVPR 2026

Active search and tracking of arbitrary targets by Unmanned Aerial Vehicles (UAVs) in cluttered environments remains a highly challenging problem. Existing methods either construct complex modular pipelines, leading to substantial computational costs, or adopt end-to-end controllers that often fail

Cited by 0SourcecodeScholar
2026

VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models

ICLR 2026poster

Visual reasoning is a core component of human intelligence and a critical capability for advanced multimodal models. Yet current reasoning evaluations of multimodal large language models (MLLMs) often rely on text descriptions and allow language-based reasoning shortcuts, failing to measure genuine…

Cited by 0SourcecodeScholar
2026

Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning

ICLR 2026poster

While significant research has focused on developing embodied reasoning capabilities using Vision-Language Models (VLMs) or integrating advanced VLMs into Vision-Language-Action (VLA) models for end-to-end robot control, few studies directly address the critical gap between upstream VLM-based reason…

Cited by 0SourcecodeScholar
2025

Active Perception Meets Rule-Guided RL: A Two-Phase Approach for Precise Object Navigation in Complex Environments

ICCV 2025poster

Object Goal Navigation (ObjectNav) in unknown environments presents significant challenges, particularly in Open-Vocabulary Mobile Manipulation (OVMM), where robots must efficiently explore large spaces, locate small objects, and accurately position themselves for subsequent manipulation. Existing a…

2025

Aligning Global Semantics and Local Textures in Generative Video Enhancement

ICCV 2025poster

Recent advances in video generation have demonstrated the utility of powerful diffusion models. One important direction among them is to enhance the visual quality of the AI-synthesized videos for artistic creation. Nevertheless, solely relying on the knowledge embedded in the pre-trained video diff…

2025

Controllable Style Arithmetic with Language Models

ACL 2025long

Language models have shown remarkable capabilities in text generation, but precisely controlling their linguistic style remains challenging. Existing methods either lack fine-grained control, require extensive computation, or introduce significant latency. We propose Style Arithmetic (SA), a novel p…

2025

DP-Habitat: Bridging the Gap Between Simulation and Reality for Visual Navigation in Dynamic Pedestrian Environments

ICRA 2025

Visual navigation in dynamic environments poses a considerable challenge, particularly in scenarios with diverse pedestrian behaviors. Traditional simulators primarily focus on static scenes, while existing dynamic pedestrian simulators often suffer limitations such as monotonous pedestrian models,

Cited by 0SourcecodeScholar
2025

DesignDiffusion: High-Quality Text-to-Design Image Generation with Diffusion Models

CVPR 2025poster

In this paper, we present DesignDiffusion, a simple yet effective framework for the novel task of synthesizing design images from textual descriptions. A primary challenge lies in generating accurate and style-consistent textual and visual content. Existing works in a related task of visual text gen…

2025

EG4D: Explicit Generation of 4D Object without Score Distillation

ICLR 2025poster

In recent years, the increasing demand for dynamic 3D assets in design and gaming applications has given rise to powerful generative pipelines capable of synthesizing high-quality 4D objects. Previous methods generally rely on score distillation sampling (SDS) algorithm to infer the unseen views a…

2025

Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling

NeurIPS 2025poster

Outcome‑reward reinforcement learning (RL) is a common—and increasingly significant—way to refine the step‑by‑step reasoning of multimodal large language models (MLLMs). In the multiple‑choice setting—a dominant format for multimodal reasoning benchmarks—the paradigm faces a significant yet often ov…

Cited by 0SourceScholar
2025

I2VGuard: Safeguarding Images against Misuse in Diffusion-based Image-to-Video Models

CVPR 2025poster

Recent advances in image-to-video generation have enabled animation of still images and offered pixel-level controllability. While these models hold great potential to transform single images into vivid and dynamic videos, they also carry risks of misuse that could impact privacy, security, and copy…

Cited by 0SourcePDFScholar
2025

Image as a World: Generating Interactive World from Single Image via Panoramic Video Generation

NeurIPS 2025poster

Generating an interactive visual world from a single image is both challenging and practically valuable, as single-view inputs are easy to acquire and align well with prompt-driven applications such as gaming and virtual reality. This paper introduces a novel unified framework, Image as a World (**I…

Cited by 0SourceScholar
2025

Incremental Transformer: Efficient Encoder for Incremented Text Over MRC and Conversation Tasks

COLING 2025main

Some encoder inputs such as conversation histories are frequently extended with short additional inputs like new responses. However, to obtain the real-time encoding of the extended input, existing Transformer-based encoders like BERT have to encode the whole extended input again without utilizing t…

Cited by 0SourcePDFScholar
2025

Leveraging Visual Captions for Enhanced Zero-Shot HOI Detection

ICASSP 2025accepted

Zero-shot Human-Object Interaction (HOI) detection aims to identify both seen and unseen HOI categories in an image. Most existing methods rely on semantic knowledge distilled from CLIP to find novel interactions but fail to fully exploit the powerful generalization ability of vision-language models…

Cited by 0SourceScholar
2025

Make-It-Animatable: An Efficient Framework for Authoring Animation-Ready 3D Characters

CVPR 2025highlight

3D characters are essential to modern creative industries, but making them animatable often demands extensive manual work in tasks like rigging and skinning. Existing automatic rigging tools face several limitations, including the necessity for manual annotations, rigid skeleton topologies, and limi…

2025

Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering

NeurIPS 2025poster

Multimodal large language models (MLLMs) have achieved remarkable progress in video understanding. However, hallucination, where the model generates plausible yet incorrect outputs, persists as a significant and under-addressed challenge in the video domain. Among existing solutions, activation engi…

Cited by 0SourceScholar
2025

Multi-Level Optimal Transport for Universal Cross-Tokenizer Knowledge Distillation on Language Models

AAAI 2025technical

Knowledge distillation (KD) has become a prevalent technique for compressing large language models (LLMs). Existing KD methods are constrained by the need for identical tokenizers (i.e., vocabularies) between teacher and student models, limiting their versatility in handling LLMs of different archit…

2025

OPTICAL: Leveraging Optimal Transport for Contribution Allocation in Dataset Distillation

CVPR 2025highlight

The demands for increasingly large-scale datasets pose substantial storage and computation challenges to building deep learning models. Dataset distillation methods, especially those via sample generation techniques, rise in response to condensing large original datasets into small synthetic ones wh…

Cited by 0SourcePDFScholar
2025

Optimizing Distributional Geometry Alignment with Optimal Transport for Generative Dataset Distillation

NeurIPS 2025poster

Dataset distillation seeks to synthesize a compact distilled dataset, enabling models trained on it to achieve performance comparable to models trained on the full dataset. Recent methods for large-scale datasets focus on matching global distributional statistics (e.g., mean and variance), but overl…

Cited by 0SourceScholar
2025

Robust Multimodal Large Language Models Against Modality Conflict

ICML 2025poster

Despite the impressive capabilities of multimodal large language models (MLLMs) in vision-language tasks, they are prone to hallucinations in real-world scenarios. This paper investigates the hallucination phenomenon in MLLMs from the perspective of modality conflict. Unlike existing works focusing…

Cited by 0SourcePDFScholar
2025

Self-Classification Enhancement and Correction for Weakly Supervised Object Detection

IJCAI 2025

In recent years, weakly supervised object detection (WSOD) has attracted much attention due to its low labeling cost. The success of recent WSOD models is often ascribed to the two-stage multi-class classification (MCC) task, i.e., multiple instance learning and online classification refinement. Des

Cited by 0SourcePDFScholar
2025

SmartEraser: Remove Anything from Images using Masked-Region Guidance

CVPR 2025poster

Object removal has so far been dominated by the mask-and-inpaint paradigm, where the masked region is excluded from the input, leaving models relying on unmasked areas to inpaint the missing region. However, this approach lacks contextual information for the masked area, often resulting in unstable…

Cited by 2SourcePDFScholar
2025

Uni-Sign: Toward Unified Sign Language Understanding at Scale

ICLR 2025poster

Sign language pre-training has gained increasing attention for its ability to enhance performance across various sign language understanding (SLU) tasks. However, existing methods often suffer from a gap between pre-training and fine-tuning, leading to suboptimal results. To address this, we propose…

2024

BoolQuestions: Does Dense Retrieval Understand Boolean Logic in Language?

EMNLP 2024finding

Dense retrieval, which aims to encode the semantic information of arbitrary text into dense vector representations or embeddings, has emerged as an effective and efficient paradigm for text retrieval, consequently becoming an essential component in various natural language processing systems. These…

2024

Forest2Seq: Revitalizing Order Prior for Sequential Indoor Scene Synthesis

ECCV 2024poster

"Synthesizing realistic 3D indoor scenes is a challenging task that traditionally relies on manual arrangement and annotation by expert designers. Recent advances in autoregressive models have automated this process, but they often lack semantic understanding of the relationships and hierarchies pre…

Cited by 7SourcePDFScholar
2024

Image2Sentence based Asymmetrical Zero-shot Composed Image Retrieval

ICLR 2024spotlight

The task of composed image retrieval (CIR) aims to retrieve images based on the query image and the text describing the users' intent. Existing methods have made great progress with the advanced large vision-language (VL) model in CIR task, however, they generally suffer from two main issues: lack…

Cited by 12SourcePDFScholar
2024

Instance-aware Exploration-Verification-Exploitation for Instance ImageGoal Navigation

CVPR 2024poster

As a new embodied vision task Instance ImageGoal Navigation (IIN) aims to navigate to a specified object depicted by a goal image in an unexplored environment. The main challenge of this task lies in identifying the target object from different viewpoints while rejecting similar distractors. Existin…

2024

Learning Spatial Adaptation and Temporal Coherence in Diffusion Models for Video Super-Resolution

CVPR 2024poster

Diffusion models are just at a tipping point for image super-resolution task. Nevertheless it is not trivial to capitalize on diffusion models for video super-resolution which necessitates not only the preservation of visual appearance from low-resolution to high-resolution videos but also the tempo…

Cited by 7SourcePDFScholar
2024

Revisiting Open-Set Panoptic Segmentation

AAAI 2024technical

In this paper, we focus on the open-set panoptic segmentation (OPS) task to circumvent the data explosion problem. Different from the close-set setting, OPS targets to detect both known and unknown categories, where the latter is not annotated during training. Different from existing work that only…

Cited by 0SourcePDFScholar
2024

SUF: Stabilized Unconstrained Fine-Tuning for Offline-to-Online Reinforcement Learning

AAAI 2024technical

Offline-to-online reinforcement learning (RL) provides a promising solution to improving suboptimal offline pre-trained policies through online fine-tuning. However, one efficient method, unconstrained fine-tuning, often suffers from severe policy collapse due to excessive distribution shift. To ens…

Cited by 3SourcePDFScholar
2024

Sinkhorn Distance Minimization for Knowledge Distillation

COLING 2024main

Knowledge distillation (KD) has been widely adopted to compress large language models (LLMs). Existing KD methods investigate various divergence measures including the Kullback-Leibler (KL), reverse Kullback-Leibler (RKL), and Jensen-Shannon (JS) divergences. However, due to limitations inherent in…

2024

TabPedia: Towards Comprehensive Visual Table Understanding with Concept Synergy

NeurIPS 2024poster

Tables contain factual and quantitative data accompanied by various structures and contents that pose challenges for machine comprehension. Previous methods generally design task-specific architectures and objectives for individual tasks, resulting in modal isolation and intricate workflows. In this…

2024

Trustworthy Alignment of Retrieval-Augmented Large Language Models via Reinforcement Learning

ICML 2024poster

Trustworthiness is an essential prerequisite for the real-world application of large language models. In this paper, we focus on the trustworthiness of language models with respect to retrieval augmentation. Despite being supported with external evidence, retrieval-augmented generation still suffers…

2023

$\mathcal{O}$-GNN: incorporating ring priors into molecular modeling

ICLR 2023poster

Cyclic compounds that contain at least one ring play an important role in drug design. Despite the recent success of molecular modeling with graph neural networks (GNNs), few models explicitly take rings in compounds into consideration, consequently limiting the expressiveness of the models. In this…

Cited by 0SourcePDFScholar
2023

AltFreezing for More General Video Face Forgery Detection

CVPR 2023highlight

Existing face forgery detection models try to discriminate fake images by detecting only spatial artifacts (e.g., generative artifacts, blending) or mainly temporal artifacts (e.g., flickering, discontinuity). They may experience significant performance degradation when facing out-domain artifacts.…

2023

AnchorFormer: Point Cloud Completion From Discriminative Nodes

CVPR 2023poster

Point cloud completion aims to recover the completed 3D shape of an object from its partial observation. A common strategy is to encode the observed points to a global feature vector and then predict the complete points through a generative process on this vector. Nevertheless, the results may suffe…

2023

BEST: BERT Pre-training for Sign Language Recognition with Coupling Tokenization

AAAI 2023technical

In this work, we are dedicated to leveraging the BERT pre-training success and modeling the domain-specific statistics to fertilize the sign language recognition~(SLR) model. Considering the dominance of hand and body in sign language expression, we organize them as pose triplet units and feed them…

Cited by 37SourcePDFScholar
2023

CLIP4HOI: Towards Adapting CLIP for Practical Zero-Shot HOI Detection

NeurIPS 2023poster

Zero-shot Human-Object Interaction (HOI) detection aims to identify both seen and unseen HOI categories. A strong zero-shot HOI detector is supposed to be not only capable of discriminating novel interactions but also robust to positional distribution discrepancy between seen and unseen categories w…

Cited by 21SourcePDFScholar
2023

Cyclic-Bootstrap Labeling for Weakly Supervised Object Detection

ICCV 2023poster

Recent progress in weakly supervised object detection is featured by a combination of multiple instance detection networks (MIDN) and ordinal online refinement. However, with only image-level annotation, MIDN inevitably assigns high scores to some unexpected region proposals when generating pseudo l…

Cited by 12PDFcodeScholar
2023

DIFFER:Decomposing Individual Reward for Fair Experience Replay in Multi-Agent Reinforcement Learning

NeurIPS 2023poster

Cooperative multi-agent reinforcement learning (MARL) is a challenging task, as agents must learn complex and diverse individual strategies from a shared team reward. However, existing methods struggle to distinguish and exploit important individual experiences, as they lack an effective way to deco…

Cited by 2SourcePDFScholar
2023

DIRE for Diffusion-Generated Image Detection

ICCV 2023poster

Diffusion models have shown remarkable success in visual synthesis, but have also raised concerns about potential abuse for malicious purposes. In this paper, we seek to build a detector for telling apart real images from diffusion-generated images. We find that existing detectors struggle to detect…

Cited by 221PDFcodeScholar
2023

Focus on Your Target: A Dual Teacher-Student Framework for Domain-Adaptive Semantic Segmentation

ICCV 2023poster

We study unsupervised domain adaptation (UDA) for semantic segmentation. Currently, a popular UDA framework lies in self-training which endows the model with two-fold abilities: (i) learning reliable semantics from the labeled images in the source domain, and (ii) adapting to the target domain via g…

Cited by 12PDFcodeScholar
2023

HandNeRF: Neural Radiance Fields for Animatable Interacting Hands

CVPR 2023poster

We propose a novel framework to reconstruct accurate appearance and geometry with neural radiance fields (NeRF) for interacting hands, enabling the rendering of photo-realistic images and videos for gesture animation from arbitrary views. Given multi-view images of a single hand or interacting hands…

Cited by 29SourcePDFScholar
2023

Learning robust representation for reinforcement learning with distractions by reward sequence prediction

UAI 2023poster

Reinforcement learning algorithms have achieved remarkable success in acquiring behavioral skills directly from pixel inputs. However, their application in real-world scenarios presents challenges due to their sensitivity to visual distractions (e.g., changes in viewpoint and light). A key factor co…

2023

Low-Light Video Enhancement with Synthetic Event Guidance

AAAI 2023technical

Low-light video enhancement (LLVE) is an important yet challenging task with many applications such as photographing and autonomous driving. Unlike single image low-light enhancement, most LLVE methods utilize temporal information from adjacent frames to restore the color and remove the noise of the…

Cited by 29SourcePDFScholar
2023

MA2CL:Masked Attentive Contrastive Learning for Multi-Agent Reinforcement Learning

IJCAI 2023poster

Recent approaches have utilized self-supervised auxiliary tasks as representation learning to improve the performance and sample efficiency of vision-based reinforcement learning algorithms in single-agent settings. However, in multi-agent reinforcement learning (MARL), these techniques face challen…

2023

Making Better Decision by Directly Planning in Continuous Control

ICLR 2023poster

By properly utilizing the learned environment model, model-based reinforcement learning methods can improve the sample efficiency for decision-making problems. Beyond using the learned environment model to train a policy, the success of MCTS-based methods shows that directly incorporating the learne…

2023

Masked Motion Predictors are Strong 3D Action Representation Learners

ICCV 2023poster

In 3D human action recognition, limited supervised data makes it challenging to fully tap into the modeling potential of powerful networks such as transformers. As a result, researchers have been actively investigating effective self-supervised pre-training strategies. In this work, we show that ins…

Cited by 46PDFcodeScholar
2023

Multi-Agent First Order Constrained Optimization in Policy Space

NeurIPS 2023poster

In the realm of multi-agent reinforcement learning (MARL), achieving high performance is crucial for a successful multi-agent system. Meanwhile, the ability to avoid unsafe actions is becoming an urgent and imperative problem to solve for real-life applications. Whereas, it is still challenging to…

Cited by 3SourcePDFScholar
2023

SimFIR: A Simple Framework for Fisheye Image Rectification with Self-supervised Representation Learning

ICCV 2023poster

In fisheye images, rich distinct distortion patterns are regularly distributed in the image plane. These distortion patterns are independent of the visual content and provide informative cues for rectification. To make the best of such rectification cues, we introduce SimFIR, a simple framework for…

Cited by 31PDFScholar
2023

State Sequences Prediction via Fourier Transform for Representation Learning

NeurIPS 2023spotlight

While deep reinforcement learning (RL) has been demonstrated effective in solving complex control tasks, sample efficiency remains a key challenge due to the large amounts of data required for remarkable performance. Existing research explores the application of representation learning for data-effi…

2022

CMD: Self-Supervised 3D Action Representation Learning with Cross-Modal Mutual Distillation

ECCV 2022poster

"In 3D action recognition, there exists rich complementary information between skeleton modalities. Nevertheless, how to model and utilize this information remains a challenging problem for self-supervised 3D action representation learning. In this work, we formulate the cross-modal interaction as a…

2022

CMT: Context-Matching-Guided Transformer for 3D Tracking in Point Clouds

ECCV 2022poster

"How to effectively match the target template features with the search area is the core problem in point-cloud-based 3D single object tracking. However, in the literature, most of the methods focus on devising sophisticated matching modules at point-level, while overlooking the rich spatial context…

Cited by 27SourcePDFScholar
2022

Domain-Agnostic Prior for Transfer Semantic Segmentation

CVPR 2022poster

Unsupervised domain adaptation (UDA) is an important topic in the computer vision community. The key difficulty lies in defining a common property between the source and target domains so that the source-domain features can align with the target-domain semantics. In this paper, we present a simple a…

Cited by 45PDFScholar
2022

Geometric Representation Learning for Document Image Rectification

ECCV 2022poster

"In document image rectification, there exist rich geometric constraints between the distorted image and the ground truth one. How- ever, such geometric constraints are largely ignored in existing advanced solutions, which limits the rectification performance. To this end, we present DocGeoNet for d…

2022

LDSA: Learning Dynamic Subtask Assignment in Cooperative Multi-Agent Reinforcement Learning

NeurIPS 2022accept

Cooperative multi-agent reinforcement learning (MARL) has made prominent progress in recent years. For training efficiency and scalability, most of the MARL algorithms make all agents share the same policy or value network. However, in many complex multi-agent tasks, different agents are expected to…

Cited by 45SourcePDFScholar
2022

Learning Token-Based Representation for Image Retrieval

AAAI 2022technical

In image retrieval, deep local features learned in a data-driven manner have been demonstrated effective to improve retrieval performance. To realize efficient retrieval on large image database, some approaches quantize deep local features with a large codebook and match images with aggregated match…

2022

TAPE: Task-Agnostic Prior Embedding for Image Restoration

ECCV 2022poster

"Learning a generalized prior for natural image restoration is an important yet challenging task. Early methods mostly involved handcrafted priors including normalized sparsity, â„“0 gradients, dark channel priors, etc.. Recently, deep neural networks have been used to learn various image priors but…

Cited by 65SourcePDFScholar
2022

Uformer: A General U-Shaped Transformer for Image Restoration

CVPR 2022poster

In this paper, we present Uformer, an effective and efficient Transformer-based architecture for image restoration, in which we build a hierarchical encoder-decoder network using the Transformer block. In Uformer, there are two core designs. First, we introduce a novel locally-enhanced window (LeWin…

Cited by 1992PDFcodeScholar
2021

ATSO: Asynchronous Teacher-Student Optimization for Semi-Supervised Image Segmentation

CVPR 2021poster

Semi-supervised learning is a useful tool for image segmentation, mainly due to its ability in extracting knowledge from unlabeled data to assist learning from labeled data. This paper focuses on a popular pipeline known as self-learning, where we point out a weakness named lazy mimicking that refer…

Cited by 75PDFScholar
2021

Contextual Similarity Aggregation with Self-attention for Visual Re-ranking

NeurIPS 2021poster

In content-based image retrieval, the first-round retrieval result by simple visual feature comparison may be unsatisfactory, which can be refined by visual re-ranking techniques. In image retrieval, it is observed that the contextual similarity among the top-ranked images is an important clue to di…

2021

Contrastive Transformation for Self-supervised Correspondence Learning

AAAI 2021technical

In this paper, we focus on the self-supervised learning of visual correspondence using unlabeled videos in the wild. Our method simultaneously considers intra- and inter-video representation associations for reliable correspondence estimation. The intra-video learning transforms the image contents a…

2021

Fine-grained Semantic Alignment Network for Weakly Supervised Temporal Language Grounding

EMNLP 2021finding

Temporal language grounding (TLG) aims to localize a video segment in an untrimmed video based on a natural language description. To alleviate the expensive cost of manual annotations for temporal boundary labels,we are dedicated to the weakly supervised setting, where only video-level descriptions…

Cited by 20SourcePDFScholar
2021

IOT: Instance-wise Layer Reordering for Transformer Structures

ICLR 2021poster

With sequentially stacked self-attention, (optional) encoder-decoder attention, and feed-forward layers, Transformer achieves big success in natural language processing (NLP), and many variants have been proposed. Currently, almost all these models assume that the \emph{layer order} is fixed and kep…

2021

Improving Sign Language Translation With Monolingual Data by Sign Back-Translation

CVPR 2021poster

Despite existing pioneering works on sign language translation (SLT), there is a non-trivial obstacle, i.e., the limited quantity of parallel sign-text data. To tackle this parallel data bottleneck, we propose a sign back-translation (SignBT) approach, which incorporates massive spoken language text…

Cited by 249PDFScholar
2021

Instance Mining with Class Feature Banks for Weakly Supervised Object Detection

AAAI 2021technical

Recent progress on weakly supervised object detection (WSOD) is characterized by formulating WSOD as a Multiple Instance Learning (MIL) problem and taking online refinement with the selected region proposals from MIL. However, MIL inclines to select the most discriminative part rather than the entir…

Cited by 39SourcePDFScholar
2021

Instance-Wise Hard Negative Example Generation for Contrastive Learning in Unpaired Image-to-Image Translation

ICCV 2021poster

Contrastive learning shows great potential in unpaired image-to-image translation, but sometimes the translated results are in poor quality and the contents are not preserved consistently. In this paper, we uncover that the negative examples play a critical role in the performance of contrastive lea…

Cited by 104PDFScholar
2021

Joint Inductive and Transductive Learning for Video Object Segmentation

ICCV 2021poster

Semi-supervised video object segmentation is a task of segmenting the target object in a video sequence given only a mask annotation in the first frame. The limited information available makes it an extremely challenging task. Most previous best-performing methods adopt matching-based transductive r…

Cited by 122PDFcodeScholar
2021

Learning Deep Local Features With Multiple Dynamic Attentions for Large-Scale Image Retrieval

ICCV 2021poster

In image retrieval, learning local features with deep convolutional networks has been demonstrated effective to improve the performance. To discriminate deep local features, some research efforts turn to attention learning. However, existing attention-based methods only generate a single attention m…

Cited by 31PDFcodeScholar
2021

SignBERT: Pre-Training of Hand-Model-Aware Representation for Sign Language Recognition

ICCV 2021poster

Hand gesture serves as a critical role in sign language. Current deep-learning-based sign language recognition (SLR) methods may suffer insufficient interpretability and overfitting due to limited sign data sources. In this paper, we introduce the first self-supervised pre-trainable SignBERT with in…

Cited by 104PDFScholar
2021

TransVG: End-to-End Visual Grounding With Transformers

ICCV 2021poster

In this paper, we present a neat yet effective transformer-based framework for visual grounding, namely TransVG, to address the task of grounding a language query to the corresponding region onto an image. The state-of-the-art methods, including two-stage or one-stage ones, rely on a complex module…

Cited by 408PDFcodeScholar
2021

Transformer Meets Tracker: Exploiting Temporal Context for Robust Visual Tracking

CVPR 2021poster

In video object tracking, there exist rich temporal contexts among successive frames, which have been largely overlooked in existing trackers. In this work, we bridge the individual video frames and explore the temporal contexts across them via a transformer architecture for robust object tracking.…

Cited by 865PDFcodeScholar
2021

Voxel R-CNN: Towards High Performance Voxel-based 3D Object Detection

AAAI 2021technical

Recent advances on 3D object detection heavily rely on how the 3D data are represented, i.e., voxel-based or point-based representation. Many existing high performance 3D detectors are point-based because this structure can better retain precise point positions. Nevertheless, point-level features le…

2020

Incorporating BERT into Neural Machine Translation

ICLR 2020poster

The recently proposed BERT (Devlin et al., 2019) has shown great power on a variety of natural language understanding tasks, such as text classification, reading comprehension, etc. However, how to effectively apply BERT to neural machine translation (NMT) lacks enough exploration. While BERT is mor…

Cited by 522SourcecodeScholar
2020

Transformation GAN for Unsupervised Image Synthesis and Representation Learning

CVPR 2020poster

Generative Adversarial Networks (GAN) have shown promising performance in image synthesis and unsupervised learning (USL). In most cases, however, the representations extracted from unsupervised GAN are usually unsatisfactory in other computer vision tasks. By using conditional GAN (CGAN), this prob…

Cited by 31PDFScholar
2020

Wavelet-Based Dual-Branch Network for Image Demoiréing

ECCV 2020poster

When smartphone cameras are used to take photos of digital screens, usually moire patterns result, severely degrading photo quality. In this paper, we design a wavelet-based dual-branch network (WDNet) with a spatial attention mechanism for image demoireing. Existing image restoration methods workin…

Cited by 126SourcePDFScholar
2019

Relation Distillation Networks for Video Object Detection

ICCV 2019poster

It has been well recognized that modeling object-to-object relations would be helpful for object detection. Nevertheless, the problem is not trivial especially when exploring the interactions between objects to boost video object detectors. The difficulty originates from the aspect that reliable obj…

Cited by 280PDFScholar
2018

Affinity Derivation and Graph Merge for Instance Segmentation

ECCV 2018poster

We present an instance segmentation scheme based on pixel affinity information, which is the relationship of two pixels belonging to a same instance. In our scheme, we use two neural networks with similar structure. One is to predict pixel level semantic score and the other is designed to derive pix…

2018

Multi-Cue Correlation Filters for Robust Visual Tracking

CVPR 2018poster

In recent years, many tracking algorithms achieve impressive performance via fusing multiple types of features, however, most of them fail to fully explore the context among the adopted multiple features and the strength of them. In this paper, we propose an efficient multi-cue analysis framework fo…

2016

Picking Deep Filter Responses for Fine-Grained Image Recognition

CVPR 2016poster

Recognizing fine-grained sub-categories such as birds and dogs is extremely challenging due to the highly localized and subtle differences in some specific parts. Most previous works rely on object/part level annotations to build part-based representation, which is demanding in practical application…

Cited by 401PDFScholar