← Search

Tao Mei

123 accepted papers

2026

Distillation Models are Good Samplers for Diffusion Reinforcement Learning

ICML 2026poster

We present DMSampler, a framework that accelerates diffusion reinforcement learning by using fast distillation models as its training-time sampling engine. It overcomes the key bottleneck of sampling from the policy model—typically requiring around 50 denoising steps—by employing a co-evolving disti…

Cited by 0SourceScholar
2026

EvoID: Reinforced Evolution for Identity-Preserving Video Generation

CVPR 2026

We present EvoID, a novel framework that reformulates Identity-Preserving Video Generation as a self-evolving process through Reinforcement Learning. Moving beyond the static paradigm of imitation learning, EvoID enables a generative model to actively learn and optimize the complex trade-offs betwee

Cited by 0SourceScholar
2026

FreeInpaint: Tuning-free Prompt Alignment and Visual Rationality Enhancement in Image Inpainting

AAAI 2026technical

Text-guided image inpainting endeavors to generate new content within specified regions of images using textual prompts from users. The primary challenge is to accurately align the inpainted areas with the user-provided prompts while maintaining a high degree of visual fidelity. While existing inpai

Cited by 0SourcePDFScholar
2026

In-Context Generation with Regional Constraints for Instructional Video Editing

ICML 2026poster

The In-context generation paradigm has demonstrated strong power in instructional image editing for better synthesis quality. Nevertheless, shaping such in-context learning for instructional video editing is not trivial. Without specifying editing regions, the results can suffer from the issue of in…

Cited by 0SourceScholar
2026

Multi-level Causal LLM-based Text-to-Motion Generation with Human Alignment

CVPR 2026

Although progress has been made in LLM-based text-driven motion generation, it still has the limitations of generating fine-grained and semantically consistent motions. These limitations stem from: 1) fine-grained motion quantization errors; 2) mismatches between causal reasoning language and non-ca

Cited by 0SourceScholar
2026

ReactID: Synchronizing Realistic Actions and Identity in Personalized Video Generation

ICLR 2026poster

Personalized video generation faces a fundamental trade-off between identity consistency and action realism: overly rigid identity preservation often leads to unnatural motion, while emphasis on action dynamics can compromise subject fidelity. This tension stems from three interrelated challenges: i…

Cited by 0SourceScholar
2026

Visual Autoregressive Modeling for Instruction-Guided Image Editing

ICLR 2026poster

Recent advances in diffusion models have brought remarkable visual fidelity to instruction-guided image editing. However, their global denoising process inherently entangles the edited region with the entire image context, leading to unintended spurious modifications and compromised adherence to edi…

Cited by 0SourcecodeScholar
2025

Aligning Global Semantics and Local Textures in Generative Video Enhancement

ICCV 2025poster

Recent advances in video generation have demonstrated the utility of powerful diffusion models. One important direction among them is to enhance the visual quality of the AI-synthesized videos for artistic creation. Nevertheless, solely relying on the knowledge embedded in the pre-trained video diff…

2025

Hierarchical Masked Autoregressive Models with Low-Resolution Token Pivots

ICML 2025poster

Autoregressive models have emerged as a powerful generative paradigm for visual generation. The current de-facto standard of next token prediction commonly operates over a single-scale sequence of dense image tokens, and is incapable of utilizing global context especially for early tokens prediction…

2025

Incorporating Visual Correspondence into Diffusion Model for Virtual Try-On

ICLR 2025poster

Diffusion models have shown preliminary success in virtual try-on (VTON) task. The typical dual-branch architecture comprises two UNets for implicit garment deformation and synthesized image generation respectively, and has emerged as the recipe for VTON task. Nevertheless, the problem remains chall…

2025

MotionPro: A Precise Motion Controller for Image-to-Video Generation

CVPR 2025poster

Animating images with interactive motion control has garnered popularity for image-to-video (I2V) generation. Modern approaches typically rely on large Gaussian kernels to extend motion trajectories as condition without explicitly defining movement region, leading to coarse motion control and failin…

2025

Ouroboros-Diffusion: Exploring Consistent Content Generation in Tuning-free Long Video Diffusion

AAAI 2025technical

The first-in-first-out (FIFO) video diffusion, built on a pre-trained text-to-video model, has recently emerged as an effective approach for tuning-free long video generation. This technique maintains a queue of video frames with progressively increasing noise, continuously producing clean frames at…

Cited by 0SourcePDFScholar
2025

Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose Interaction

CVPR 2025poster

Video virtual try-on aims to seamlessly dress a subject in a video with a specific garment. The primary challenge involves preserving the visual authenticity of the garment while dynamically adapting to the pose and physique of the subject. While existing methods have predominantly focused on image-…

Cited by 0SourcePDFScholar
2025

VTON-VLLM: Aligning Virtual Try-On Models with Human Preferences

NeurIPS 2025poster

Diffusion models have yielded remarkable success in virtual try-on (VTON) task, yet they often fall short of fully meeting user expectations regarding visual quality and detail preservation. To alleviate this issue, we curate a dataset of synthesized VTON images annotated with human judgments across…

Cited by 0SourcecodeScholar
2024

Boosting Diffusion Models with Moving Average Sampling in Frequency Domain

CVPR 2024poster

Diffusion models have recently brought a powerful revolution in image generation. Despite showing impressive generative capabilities most of these models rely on the current sample to denoise the next one possibly resulting in denoising instability. In this paper we reinterpret the iterative denoisi…

Cited by 20SourcePDFScholar
2024

DreamMesh: Jointly Manipulating and Texturing Triangle Meshes for Text-to-3D Generation

ECCV 2024poster

"Learning radiance fields (NeRF) with powerful 2D diffusion models has garnered popularity for text-to-3D generation. Nevertheless, the implicit 3D representations of NeRF lack explicit modeling of meshes and textures over surfaces, and such surface-undefined way may suffer from the issues, e.g., no…

2024

Improving Text-guided Object Inpainting with Semantic Pre-inpainting

ECCV 2024poster

"Recent years have witnessed the success of large text-to-image diffusion models and their remarkable potential to generate high-quality images. The further pursuit of enhancing the editability of images has sparked significant interest in the downstream task of inpainting a novel object described b…

2024

Improving Virtual Try-On with Garment-focused Diffusion Models

ECCV 2024poster

"Diffusion models have led to the revolutionizing of generative modeling in numerous image synthesis tasks. Nevertheless, it is not trivial to directly apply diffusion models for synthesizing an image of a target person wearing a given in-shop garment, i.e., image-based virtual try-on (VTON) task. T…

2024

Learning Spatial Adaptation and Temporal Coherence in Diffusion Models for Video Super-Resolution

CVPR 2024poster

Diffusion models are just at a tipping point for image super-resolution task. Nevertheless it is not trivial to capitalize on diffusion models for video super-resolution which necessitates not only the preservation of visual appearance from low-resolution to high-resolution videos but also the tempo…

Cited by 7SourcePDFScholar
2024

Prompt Refinement with Image Pivot for Text-to-Image Generation

ACL 2024long

For text-to-image generation, automatically refining user-provided natural language prompts into the keyword-enriched prompts favored by systems is essential for the user experience. Such a prompt refinement process is analogous to translating the prompt from “user languages” into “system languages”…

2024

SD-DiT: Unleashing the Power of Self-supervised Discrimination in Diffusion Transformer

CVPR 2024poster

Diffusion Transformer (DiT) has emerged as the new trend of generative diffusion models on image generation. In view of extremely slow convergence in typical DiT recent breakthroughs have been driven by mask strategy that significantly improves the training efficiency of DiT with additional intra-im…

Cited by 27SourcePDFScholar
2024

TRIP: Temporal Residual Learning with Image Noise Prior for Image-to-Video Diffusion Models

CVPR 2024poster

Recent advances in text-to-video generation have demonstrated the utility of powerful diffusion models. Nevertheless the problem is not trivial when shaping diffusion models to animate static image (i.e. image-to-video generation). The difficulty originates from the aspect that the diffusion process…

2024

VP3D: Unleashing 2D Visual Prompt for Text-to-3D Generation

CVPR 2024poster

Recent innovations on text-to-3D generation have featured Score Distillation Sampling (SDS) which enables the zero-shot learning of implicit 3D models (NeRF) by directly distilling prior knowledge from 2D diffusion models. However current SDS-based models still struggle with intricate text prompts a…

2024

VideoStudio: Generating Consistent-Content and Multi-Scene Videos

ECCV 2024poster

"The recent innovations and breakthroughs in diffusion models have significantly expanded the possibilities of generating high-quality videos for the given prompts. Most existing works tackle the single-scene scenario with only one video event occurring in a single background. Extending to generate…

2023

AnchorFormer: Point Cloud Completion From Discriminative Nodes

CVPR 2023poster

Point cloud completion aims to recover the completed 3D shape of an object from its partial observation. A common strategy is to encode the observed points to a global feature vector and then predict the complete points through a generative process on this vector. Nevertheless, the results may suffe…

2023

Learning Neural Implicit Surfaces with Object-Aware Radiance Fields

ICCV 2023poster

Recent progress on multi-view 3D object reconstruction has featured neural implicit surfaces via learning high-fidelity radiance fields. However, most approaches hinge on the visual hull derived from cost-expensive silhouette masks to obtain object surfaces. In this paper, we propose a novel Object-…

Cited by 2PDFScholar
2023

Learning To Generate Language-Supervised and Open-Vocabulary Scene Graph Using Pre-Trained Visual-Semantic Space

CVPR 2023poster

Scene graph generation (SGG) aims to abstract an image into a graph structure, by representing objects as graph nodes and their relations as labeled edges. However, two knotty obstacles limit the practicability of current SGG methods in real-world scenarios: 1) training SGG models requires time-cons…

2023

Modality-Agnostic Debiasing for Single Domain Generalization

CVPR 2023poster

Deep neural networks (DNNs) usually fail to generalize well to outside of distribution (OOD) data, especially in the extreme case of single domain generalization (single-DG) that transfers DNNs from single domain to multiple unseen domains. Existing single-DG techniques commonly devise various data-…

Cited by 29SourcePDFScholar
2023

ObjectFusion: Multi-modal 3D Object Detection with Object-Centric Fusion

ICCV 2023poster

Recent progress on multi-modal 3D object detection has featured BEV (Bird-Eye-View) based fusion, which effectively unifies both LiDAR point clouds and camera images in a shared BEV space. Nevertheless, it is not trivial to perform camera-to-BEV transformation due to the inherently ambiguous depth e…

Cited by 38PDFScholar
2023

PointClustering: Unsupervised Point Cloud Pre-Training Using Transformation Invariance in Clustering

CVPR 2023highlight

Feature invariance under different data transformations, i.e., transformation invariance, can be regarded as a type of self-supervision for representation learning. In this paper, we present PointClustering, a new unsupervised representation learning scheme that leverages transformation invariance f…

2023

Semantic-Conditional Diffusion Networks for Image Captioning

CVPR 2023poster

Recent advances on text-to-image generation have witnessed the rise of diffusion models which act as powerful generative models. Nevertheless, it is not trivial to exploit such latent variable models to capture the dependency among discrete words and meanwhile pursue complex visual-language alignmen…

2023

TRACE: 5D Temporal Regression of Avatars With Dynamic Cameras in 3D Environments

CVPR 2023poster

Although the estimation of 3D human pose and shape (HPS) is rapidly progressing, current methods still cannot reliably estimate moving humans in global coordinates, which is critical for many applications. This is particularly challenging when the camera is also moving, entangling human and camera m…

2022

CAViT: Contextual Alignment Vision Transformer for Video Object Re-identification

ECCV 2022poster

"Video object re-identification (reID) aims at re-identifying the same object under non-overlapping cameras by matching the video tracklets with cropped video frames. The key point is how to make full use of spatio-temporal interactions to extract more accurate representation. However, there are dil…

2022

Dynamic Temporal Filtering In Video Models

ECCV 2022poster

"Video temporal dynamics is conventionally modeled with 3D spatial-temporal kernel or its factorized version comprised of 2D spatial kernel and 1D temporal kernel. The modeling power, nevertheless, is limited by the fixed window size and static weights of a kernel along the temporal dimension. The p…

2022

Exploring Structure-Aware Transformer Over Interaction Proposals for Human-Object Interaction Detection

CVPR 2022poster

Recent high-performing Human-Object Interaction (HOI) detection techniques have been highly influenced by Transformer-based object detector (i.e., DETR). Nevertheless, most of them directly map parametric interaction queries into a set of HOI predictions through vanilla Transformer in a one-stage ma…

Cited by 93PDFcodeScholar
2022

Gait Recognition in the Wild With Dense 3D Representations and a Benchmark

CVPR 2022poster

Existing studies for gait recognition are dominated by 2D representations like the silhouette or skeleton of the human body in constrained scenes. However, humans live and walk in the unconstrained 3D space, so projecting the 3D human body onto the 2D plane will discard a lot of crucial information…

Cited by 192PDFcodeScholar
2022

Generalized One-shot Domain Adaptation of Generative Adversarial Networks

NeurIPS 2022accept

The adaptation of a Generative Adversarial Network (GAN) aims to transfer a pre-trained GAN to a target domain with limited training data. In this paper, we focus on the one-shot case, which is more challenging and rarely explored in previous works. We consider that the adaptation from a source doma…

2022

Out-of-Distribution Detection via Conditional Kernel Independence Model

NeurIPS 2022accept

Recently, various methods have been introduced to address the OOD detection problem with training outlier exposure. These methods usually count on discriminative softmax metric or energy method to screen OOD samples. In this paper, we probe an alternative hypothesis on OOD detection by constructing…

2022

Putting People in Their Place: Monocular Regression of 3D People in Depth

CVPR 2022poster

Given an image with multiple people, our goal is to directly regress the pose and shape of all the people as well as their relative depth. Inferring the depth of a person in an image, however, is fundamentally ambiguous without knowing their height. This is particularly problematic when the scene co…

Cited by 180PDFcodeScholar
2022

Responsive Listening Head Generation: A Benchmark Dataset and Baseline

ECCV 2022poster

"We present a new listening head generation benchmark, for synthesizing responsive feedbacks of a listener (e.g., nod, smile) during a face-to-face conversation. As the indispensable complement to talking heads generation, listening head generation has seldomly been studied in literature. Automatica…

Cited by 60SourcePDFScholar
2022

SPE-Net: Boosting Point Cloud Analysis via Rotation Robustness Enhancement

ECCV 2022poster

"In this paper, we propose a novel deep architecture tailored for 3D point cloud applications, named as SPE-Net. The embedded ""Selective Position Encoding (SPE)"" procedure relies on an attention mechanism that can effectively attend to the underlying rotation condition of the input. Such encoded r…

2022

Stand-Alone Inter-Frame Attention in Video Models

CVPR 2022poster

Motion, as the uniqueness of a video, has been critical to the development of video understanding models. Modern deep learning models leverage motion by either executing spatio-temporal 3D convolutions, factorizing 3D convolutions into spatial and temporal convolutions separately, or computing self-…

Cited by 62PDFcodeScholar
2022

Wave-ViT: Unifying Wavelet and Transformers for Visual Representation Learning

ECCV 2022poster

"Multi-scale Vision Transformer (ViT) has emerged as a powerful backbone for computer vision tasks, while the self-attention computation in Transformer scales quadratically w.r.t. the input patch number. Thus, existing solutions commonly employ down-sampling operations (e.g., average pooling) over k…

2021

A Style and Semantic Memory Mechanism for Domain Generalization

ICCV 2021poster

Mainstream state-of-the-art domain generalization algorithms tend to prioritize the assumption on semantic invariance across domains. Meanwhile, the inherent intra-domain style invariance is usually underappreciated and put on the shelf. In this paper, we reveal that leveraging intra-domain style in…

Cited by 53PDFScholar
2021

Action Unit Memory Network for Weakly Supervised Temporal Action Localization

CVPR 2021poster

Weakly supervised temporal action localization aims to detect and localize actions in untrimmed videos with only video-level labels during training. However, without frame-level annotations, it is challenging to achieve localization completeness and relieve background interference. In this paper, we…

Cited by 106PDFScholar
2021

Boosting Video Representation Learning With Multi-Faceted Integration

CVPR 2021poster

Video content is multifaceted, consisting of objects, scenes, interactions or actions. The existing datasets mostly label only one of the facets for model training, resulting in the video representation that biases to only one facet depending on the training dataset. There is no study yet on how to…

Cited by 13PDFScholar
2021

CM-NAS: Cross-Modality Neural Architecture Search for Visible-Infrared Person Re-Identification

ICCV 2021poster

Visible-Infrared person re-identification (VI-ReID) aims to match cross-modality pedestrian images, breaking through the limitation of single-modality person ReID in dark environment. In order to mitigate the impact of large modality discrepancy, existing works manually design various two-stream arc…

Cited by 159PDFcodeScholar
2021

Co-Optimization of Morphology and Actuation Parameters of Multi-Sectional FREEs for Trajectory Matching

RA-L 2021

Fiber Reinforced Elastomeric Enclosures (FREEs) have gained significant popularity as a form of soft artificial muscle for the diversified deformation behaviors upon pressurization. In particular, modular FREEs connected in series, demonstrate enhanced flexibility and reconfigurability, thus are ada

Cited by 1SourceScholar
2021

Condensing a Sequence to One Informative Frame for Video Recognition

ICCV 2021poster

Video is complex due to large variations in motion and rich content in fine-grained visual details. Abstracting useful information from such information-intensive media requires exhaustive computing resources. This paper studies a two-step alternative that first condenses the video sequence to an in…

Cited by 10PDFScholar
2021

Design of a deployable underwater robot for the recovery of autonomous underwater vehicles based on origami technique

ICRA 2021poster

The recovery of autonomous underwater vehicles (AUVs) has been a challenging mission due to the limited localization accuracy and movement capability of the AUVs. To overcome these limitations, we propose a novel design of a deployable underwater robot (DUR) for the recovery mission. Utilizing the o…

Cited by 0SourceScholar
2021

Dive Into Ambiguity: Latent Distribution Mining and Pairwise Uncertainty Estimation for Facial Expression Recognition

CVPR 2021poster

Due to the subjective annotation and the inherent inter-class similarity of facial expressions, one of key challenges in Facial Expression Recognition (FER) is the annotation ambiguity. In this paper, we proposes a solution, named DMUE, to address the problem of annotation ambiguity from two perspec…

Cited by 291PDFcodeScholar
2021

Explainable Person Re-Identification With Attribute-Guided Metric Distillation

ICCV 2021poster

Despite the great progress of person re-identification (ReID) with the adoption of Convolutional Neural Networks, current ReID models are opaque and only outputs a scalar distance between two persons. There are few methods providing users semantically understandable explanations for why two persons…

Cited by 57PDFcodeScholar
2021

Exploiting Relationship for Complex-scene Image Generation

AAAI 2021technical

The significant progress on Generative Adversarial Networks (GANs) has facilitated realistic single-object image generation based on language input. However, complex-scene generation (with various interactions among multiple objects) still suffers from messy layouts and object distortions, due to di…

2021

Group-aware Label Transfer for Domain Adaptive Person Re-identification

CVPR 2021poster

Unsupervised Domain Adaptive (UDA) person re-identification (ReID) aims at adapting the model trained on a labeled source-domain dataset to a target-domain dataset without any further annotations. Most successful UDA-ReID approaches combine clustering-based pseudo-label prediction with representatio…

Cited by 231PDFcodeScholar
2021

Improving Self-supervised Learning with Automated Unsupervised Outlier Arbitration

NeurIPS 2021poster

Our work reveals a structured shortcoming of the existing mainstream self-supervised learning methods. Whereas self-supervised learning frameworks usually take the prevailing perfect instance level invariance hypothesis for granted, we carefully investigate the pitfalls behind. Particularly, we argu…

2021

Monocular, One-Stage, Regression of Multiple 3D People

ICCV 2021poster

This paper focuses on the regression of multiple 3D people from a single RGB image. Existing approaches predominantly follow a multi-stage pipeline that first detects people in bounding boxes and then independently regresses their 3D body meshes. In contrast, we propose to Regress all meshes in a On…

Cited by 327PDFcodeScholar
2021

Motion-Focused Contrastive Learning of Video Representations

ICCV 2021poster

Motion, as the most distinct phenomenon in a video to involve the changes over time, has been unique and critical to the development of video representation learning. In this paper, we ask the question: how important is the motion particularly for self-supervised video representation learning. To th…

Cited by 47PDFcodeScholar
2021

Representing Videos As Discriminative Sub-Graphs for Action Recognition

CVPR 2021poster

Human actions are typically of combinatorial structures or patterns, i.e., subjects, objects, plus spatio-temporal interactions in between. Discovering such structures is therefore a rewarding way to reason about the dynamics of interactions and recognize the actions. In this paper, we introduce a n…

Cited by 34PDFScholar
2021

Scheduled Sampling in Vision-Language Pretraining with Decoupled Encoder-Decoder Network

AAAI 2021technical

Despite having impressive vision-language (VL) pretraining with BERT-based encoder for VL understanding, the pretraining of a universal encoder-decoder for both VL understanding and generation remains challenging. The difficulty originates from the inherently different peculiarities of the two disci…

2021

SeCo: Exploring Sequence Supervision for Unsupervised Representation Learning

AAAI 2021technical

A steady momentum of innovations and breakthroughs has convincingly pushed the limits of unsupervised image representation learning. Compared to static 2D images, video has one more dimension (time). The inherent supervision existing in such sequential structure offers a fertile ground for building…

2021

Weakly Supervised Semantic Segmentation for Large-Scale Point Cloud

AAAI 2021technical

Existing methods for large-scale point cloud semantic segmentation require expensive, tedious and error-prone manual point-wise annotation. Intuitively, weakly supervised training is a direct solution to reduce the labeling costs. However, for weakly supervised large-scale point cloud semantic segme…

2020

Classes Matter: A Fine-grained Adversarial Approach to Cross-domain Semantic Segmentation

ECCV 2020poster

Despite great progress in supervised semantic segmentation, a large performance drop is usually observed when deploying the model in the wild. Domain adaptation methods tackle the issue by aligning the source domain and the target domain. However, most existing methods attempt to perform the alignme…

2020

Edge-aware Graph Representation Learning and Reasoning for Face Parsing

ECCV 2020poster

Face parsing infers a pixel-wise label to each facial component, which has drawn much attention recently. Previous methods have shown their efficiency in face parsing, which however overlook the correlation among different face regions. The correlation is a critical clue about the facial appearance,…

2020

Exclusivity-Consistency Regularized Knowledge Distillation for Face Recognition

ECCV 2020poster

Knowledge distillation is an effective tool to compress large pre-trained Convolutional Neural Networks (CNNs) or their ensembles into models applicable to mobile and embedded devices. The success of which mainly comes from two aspects: the designed student network and the exploited knowledge. Howev…

2020

Exploring Category-Agnostic Clusters for Open-Set Domain Adaptation

CVPR 2020poster

Unsupervised domain adaptation has received significant attention in recent years. Most of existing works tackle the closed-set scenario, assuming that the source and target domains share the exactly same categories. In practice, nevertheless, a target domain often contains samples of classes unseen…

Cited by 94PDFScholar
2020

Joint Contrastive Learning with Infinite Possibilities

NeurIPS 2020spotlight

This paper explores useful modifications of the recent development in contrastive learning via novel probabilistic modeling. We derive a particular form of contrastive loss named Joint Contrastive Learning (JCL). JCL implicitly involves the simultaneous learning of an infinite number of query-key pa…

2020

Learning a Unified Sample Weighting Network for Object Detection

CVPR 2020poster

Region sampling or weighting is significantly important to the success of modern region-based object detectors. Unlike some previous works, which only focus on "hard" samples when optimizing the objective function, we argue that sample weighting should be data-dependent and task-dependent. The impor…

Cited by 45PDFcodeScholar
2020

Learning the Compositional Visual Coherence for Complementary Recommendations

IJCAI 2020poster

Complementary recommendations, which aim at providing users product suggestions that are supplementary and compatible with their obtained items, have become a hot topic in both academia and industry in recent years. Existing work mainly focused on modeling the co-purchased relations between two item…

Cited by 0SourcePDFScholar
2020

Learning to Localize Actions from Moments

ECCV 2020poster

With the knowledge of action moments (i.e., trimmed video clips that each contains an action instance), humans could routinely localize an action temporally in an untrimmed video. Nevertheless, most practical methods still require all training videos to be labeled with temporal annotations (action c…

2020

Look-Into-Object: Self-Supervised Structure Modeling for Object Recognition

CVPR 2020poster

Most object recognition approaches predominantly focus on learning discriminative visual patterns, while overlooking the holistic object structure. Though important, structure modeling usually requires significant manual annotations and therefore is labor-intensive. In this paper, we propose to "loo…

Cited by 99PDFcodeScholar
2020

Semi-Siamese Training for Shallow Face Learning

ECCV 2020poster

Most existing public face datasets, such as MS-Celeb-1M and VGGFace2, provide abundant information in both breadth (large number of IDs) and depth (sufficient number of samples) for training. However, in many real-world scenarios of face recognition, the training dataset is limited in depth, $ extit…

2020

Transferring and Regularizing Prediction for Semantic Segmentation

CVPR 2020poster

Semantic segmentation often requires a large set of images with pixel-level annotations. In the view of extremely expensive expert labeling, recent research has shown that the models trained on photo-realistic synthetic data (e.g., computer games) with computer-generated annotations can be adapted t…

Cited by 47PDFScholar
2019

Customizable Architecture Search for Semantic Segmentation

CVPR 2019poster

In this paper, we propose a Customizable Architecture Search (CAS) approach to automatically generate a network architecture for semantic image segmentation. The generated network consists of a sequence of stacked computation cells. A computation cell is represented as a directed acyclic graph, in w…

Cited by 179PDFScholar
2019

Destruction and Construction Learning for Fine-Grained Image Recognition

CVPR 2019poster

Delicate feature representation about object parts plays a critical role in fine-grained recognition. For example, experts can even distinguish fine-grained objects relying only on object parts according to professional knowledge. In this paper, we propose a novel "Destruction and Construction Learn…

Cited by 599PDFcodeScholar
2019

Gaussian Temporal Awareness Networks for Action Localization

CVPR 2019oral

Temporally localizing actions in a video is a fundamental challenge in video understanding. Most existing approaches have often drawn inspiration from image object detection and extended the advances, e.g., SSD and Faster R-CNN, to produce temporal locations of an action in a 1D sequence. Neverthele…

Cited by 438PDFScholar
2019

Human Mesh Recovery From Monocular Images via a Skeleton-Disentangled Representation

ICCV 2019poster

We describe an end-to-end method for recovering 3D human body mesh from single images and monocular videos. Different from the existing methods try to obtain all the complex 3D pose, shape, and camera parameters from one coupling feature, we propose a skeleton-disentangling based framework, which di…

Cited by 214PDFcodeScholar
2019

Learning Spatio-Temporal Representation With Local and Global Diffusion

CVPR 2019poster

Convolutional Neural Networks (CNN) have been regarded as a powerful class of models for visual recognition problems. Nevertheless, the convolutional filters in these networks are local operations while ignoring the large-range dependency. Such drawback becomes even worse particularly for video reco…

Cited by 236PDFScholar
2019

Relation Distillation Networks for Video Object Detection

ICCV 2019poster

It has been well recognized that modeling object-to-object relations would be helpful for object detection. Nevertheless, the problem is not trivial especially when exploring the interactions between objects to boost video object detectors. The difficulty originates from the aspect that reliable obj…

Cited by 280PDFScholar
2019

Sampling Wisely: Deep Image Embedding by Top-K Precision Optimization

ICCV 2019poster

Deep image embedding aims at learning a convolutional neural network (CNN) based mapping function that maps an image to a feature vector. The embedding quality is usually evaluated by the performance in image search tasks. Since very few users bother to open the second page search results, top-k pre…

Cited by 33PDFcodeScholar
2019

ScratchDet: Training Single-Shot Object Detectors From Scratch

CVPR 2019oral

Current state-of-the-art object objectors are fine-tuned from the off-the-shelf networks pretrained on large-scale classification dataset ImageNet, which incurs some additional problems: 1) The classification and detection have different degrees of sensitivity to translation, resulting in the learni…

Cited by 188PDFcodeScholar
2019

Social Relation Recognition From Videos via Multi-Scale Spatial-Temporal Reasoning

CVPR 2019poster

Discovering social relations, e.g., kinship, friendship, etc., from visual contents can make machines better interpret the behaviors and emotions of human beings. Existing studies mainly focus on recognizing social relations from still images while neglecting another important media--video. On one h…

Cited by 94PDFScholar
2019

Transferrable Prototypical Networks for Unsupervised Domain Adaptation

CVPR 2019oral

In this paper, we introduce a new idea for unsupervised domain adaptation via a remold of Prototypical Networks, which learn an embedding space and perform classification via a remold of the distances to the prototype of each class. Specifically, we present Transferrable Prototypical Networks (TPN)…

Cited by 461PDFScholar
2019

Unsupervised Person Image Generation With Semantic Parsing Transformation

CVPR 2019oral

In this paper, we address unsupervised pose-guided person image generation, which is known challenging due to non-rigid deformation. Unlike previous methods learning a rock-hard direct mapping between human bodies, we propose a new pathway to decompose the hard mapping into two more accessible subta…

Cited by 141PDFcodeScholar
2018

DA-GAN: Instance-Level Image Translation by Deep Attention Generative Adversarial Networks

CVPR 2018poster

Unsupervised image translation, which aims in translating two independent sets of images, is challenging in discovering the correct correspondences without paired data. Existing works build upon Generative Adversarial Networks (GANs) such that the distribution of the translated images are indistingu…

Cited by 182SourcePDFScholar
2018

Deep Attention Neural Tensor Network for Visual Question Answering

ECCV 2018poster

Visual question answering (VQA) has drawn great attention in cross-modal learning problems, which enables a machine to answer a natural language question given a reference image. Significant progress has been made by learning rich embedding features from images and questions by bilinear models, whil…

Cited by 83SourcePDFScholar
2018

Fully Convolutional Adaptation Networks for Semantic Segmentation

CVPR 2018poster

The recent advances in deep neural networks have convincingly demonstrated high capability in learning vision models on large datasets. Nevertheless, collecting expert labeled datasets especially with pixel-level annotations is an extremely expensive process. An appealing alternative is to render sy…

Cited by 429SourcePDFScholar
2018

Jointly Localizing and Describing Events for Dense Video Captioning

CVPR 2018poster

Automatically describing a video with natural language is regarded as a fundamental challenge in computer vision. The problem nevertheless is not trivial especially when a video contains multiple events to be worthy of mention, which often happens in real videos. A valid question is how to temporall…

2018

Part-Aligned Bilinear Representations for Person Re-Identification

ECCV 2018poster

Comparing the appearance of corresponding body parts is essential for person re-identification. As body parts are frequently misaligned between the detected human boxes, an image representation that can handle this misalignment is required. In this paper, we propose a network that learns a part-alig…

Cited by 671SourcePDFScholar
2018

Recurrent Tubelet Proposal and Recognition Networks for Action Detection

ECCV 2018poster

Detecting actions in videos is a challenging task as video is an information intensive media with complex variations. Existing approaches predominantly generate action proposals for each individual frame or fixed-length clip independently, while overlooking temporal context across them. Such tempora…

Cited by 146SourcePDFScholar
2017

Incorporating Copying Mechanism in Image Captioning for Learning Novel Objects

CVPR 2017poster

Image captioning often requires a large set of training image-sentence pairs. In practice, however, acquiring sufficient training pairs is always expensive, making the recent captioning models limited in their ability to describe objects outside of training corpora (i.e., novel objects). In this pap…

Cited by 181PDFcodeScholar
2017

Joint Detection and Recounting of Abnormal Events by Learning Deep Generic Knowledge

ICCV 2017poster

This paper addresses the problem of joint detection and recounting of abnormal events in videos. Recounting of abnormal events, i.e., explaining why they are judged to be abnormal, is an unexplored but critical task in video surveillance, because it helps human observers quickly judge if they are fa…

Cited by 295PDFScholar
2017

Learning Multi-Attention Convolutional Neural Network for Fine-Grained Image Recognition

ICCV 2017oral

Recognizing fine-grained categories (e.g., bird species) highly relies on discriminative part localization and part-based fine-grained feature learning. Existing approaches predominantly solve these challenges independently, while neglecting the fact that part localization (e.g., head of a bird) and…

Cited by 1166PDFcodeScholar
2017

Look Closer to See Better: Recurrent Attention Convolutional Neural Network for Fine-Grained Image Recognition

CVPR 2017oral

Recognizing fine-grained categories (e.g., bird species) is difficult due to the challenges of discriminative region localization and fine-grained feature learning. Existing approaches predominantly solve these challenges independently, while neglecting the fact that region detection and fine-graine…

Cited by 1640PDFScholar
2016

Jointly Modeling Embedding and Translation to Bridge Video and Language

CVPR 2016oral

Automatically describing video content with natural language is a fundamental challenge of computer vision. Recurrent Neural Networks (RNNs), which models sequence dynamics, has attracted increasing attention on visual interpretation. However, most existing approaches generate a word locally with th…

Cited by 716PDFScholar
2016

You Lead, We Exceed: Labor-Free Video Concept Learning by Jointly Exploiting Web Videos and Images

CVPR 2016spotlight

Video concept learning often requires a large set of training samples. In practice, however, acquiring noise-free training labels with sufficient positive examples is very expensive. A plausible solution for training data collection is by sampling from the vast quantities of images and videos on the…

Cited by 136PDFScholar
2015

Multi-Task Deep Visual-Semantic Embedding for Video Thumbnail Selection

CVPR 2015poster

Given the tremendous growth of online videos, video thumbnail, as the common visualization form of video content, is becoming increasingly important to influence user's browsing and searching experience. However, conventional methods for video thumbnail selection often fail to produce satisfying res…

Cited by 285SourcePDFScholar
2015

Relaxing From Vocabulary: Robust Weakly-Supervised Deep Learning for Vocabulary-Free Image Tagging

ICCV 2015poster

The development of deep learning has empowered machines with comparable capability of recognizing limited image categories to human beings. However, most existing approaches heavily rely on human-curated training data, which hinders the scalability to large and unlabeled vocabularies in image taggin…

Cited by 48PDFScholar
2015

Semi-Supervised Domain Adaptation With Subspace Learning for Visual Recognition

CVPR 2015poster

In many real-world applications, we are often facing the problem of cross domain learning, i.e., to borrow the labeled data or transfer the already learnt knowledge from a source domain to a target domain. However, simply applying existing source data or knowledge may even hurt the performance, espe…

Cited by 276SourcePDFScholar