← Search

Errui Ding

100 accepted papers

2026

GeoTikzBridge: Advancing Multimodal Code Generation for Geometric Perception and Reasoning

CVPR 2026

Multimodal Large Language Models (MLLMs) have recently demonstrated remarkable perceptual and reasoning abilities. However, they struggle to perceive fine-grained geometric structures, constraining their ability of geometric understanding and visual reasoning. To address this, we propose GeoTikzBrid

Cited by 0SourcecodeScholar
2025

AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers

CVPR 2025poster

Despite the recent progress of audio-driven video generation, existing methods mostly focus on driving facial movements, leading to non-coherent head and body dynamics. Moving forward, it is desirable yet challenging to generate holistic human videos with both accurate lip-sync and delicate co-speec…

Cited by 0SourcePDFScholar
2025

Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot Generalization

ICML 2025poster

Computer Vision (CV) has yet to fully achieve the zero-shot task generalization observed in Natural Language Processing (NLP), despite following many of the milestones established in NLP, such as large transformer models, extensive pre-training, and the auto-regression paradigm, among others. In thi…

2025

Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language Models

AAAI 2025technical

Face Anti-Spoofing (FAS) is essential for ensuring the security and reliability of facial recognition systems. Most existing FAS methods are formulated as binary classification tasks, providing confidence scores without interpretation. They exhibit limited generalization in out-of-domain scenarios,…

Cited by 0SourcePDFScholar
2025

MGMapNet: Multi-Granularity Representation Learning for End-to-End Vectorized HD Map Construction

ICLR 2025poster

The construction of vectorized high-definition map typically requires capturing both category and geometry information of map elements. Current state-of-the-art methods often adopt solely either point-level or instance-level representation, overlooking the strong intrinsic relationship between point…

Cited by 3SourcePDFScholar
2025

Re-HOLD: Video Hand Object Interaction Reenactment via adaptive Layout-instructed Diffusion Model

CVPR 2025poster

Current digital human studies focusing on lip-syncing and body movement are no longer sufficient to meet the growing industrial demand, while human video generation techniques that support interacting with real-world environments (e.g., objects) have not been well investigated. Despite human hand sy…

2025

Splatter-360: Generalizable 360 Gaussian Splatting for Wide-baseline Panoramic Images

CVPR 2025poster

Wide-baseline panoramic images are frequently used in applications like VR and simulations to minimize capturing labor costs and storage needs. However, synthesizing novel views from these panoramic images in real time remains a significant challenge, especially due to panoramic imagery's high resol…

2025

TexGarment: Consistent Garment UV Texture Generation via Efficient 3D Structure-Guided Diffusion Transformer

CVPR 2025poster

This paper introduces TexGarment, an efficient method for synthesizing high-quality, 3D-consistent garment textures in UV space. Traditional approaches based on 2D-to-3D mapping often suffer from 3D inconsistency, while methods learning from limited 3D data lack sufficient texture diversity. These l…

Cited by 0SourcePDFScholar
2025

TexGaussian: Generating High-quality PBR Material via Octree-based 3D Gaussian Splatting

CVPR 2025poster

Physically Based Rendering (PBR) materials play a crucial role in modern graphics, enabling photorealistic rendering across diverse environment maps. Developing an effective and efficient algorithm that is capable of automatically generating high-quality PBR materials rather than RGB texture for 3D…

2025

Uni$^2$Det: Unified and Universal Framework for Prompt-Guided Multi-dataset 3D Detection

ICLR 2025poster

We present Uni$^2$Det, a brand new framework for unified and universal multi-dataset training on 3D detection, enabling robust performance across diverse domains and generalization to unseen domains. Due to substantial disparities in data distribution and variations in taxonomy across diverse domain…

2025

VDG: Vision-Only Dynamic Gaussian for Driving Simulation

RA-L 2025

Recent advances in dynamic Gaussian splatting have significantly improved scene reconstruction and novel-view synthesis. However, existing methods often rely on pre-computed camera poses and Gaussian initialization using Structure from Motion (SfM) or other costly sensors, limiting their scalability

Cited by 23SourceScholar
2024

Decoupled Pseudo-labeling for Semi-Supervised Monocular 3D Object Detection

CVPR 2024poster

We delve into pseudo-labeling for semi-supervised monocular 3D object detection (SSM3OD) and discover two primary issues: a misalignment between the prediction quality of 3D and 2D attributes and the tendency of depth supervision derived from pseudo-labels to be noisy leading to significant optimiza…

Cited by 7SourcePDFScholar
2024

GGRt: Towards Generalizable 3D Gaussians without Pose Priors in Real-Time

ECCV 2024poster

"This paper presents GGRt, a novel approach to generalizable novel view synthesis that alleviates the need for real camera poses, complexity in processing high-resolution images, and lengthy optimization processes, thus facilitating stronger applicability of 3D Gaussian Splatting (3D-GS) in real-wor…

2024

Interactive 3D Object Detection with Prompts

ECCV 2024poster

"The evolution of 3D object detection hinges not only on advanced models but also on effective and efficient annotation strategies. Despite this progress, the labor-intensive nature of 3D object annotation remains a bottleneck, hindering further development in the field. This paper introduces a nove…

Cited by 0SourcePDFScholar
2024

KD-DETR: Knowledge Distillation for Detection Transformer with Consistent Distillation Points Sampling

CVPR 2024poster

DETR is a novel end-to-end transformer architecture object detector which significantly outperforms classic detectors when scaling up. In this paper we focus on the compression of DETR with knowledge distillation. While knowledge distillation has been well-studied in classic detectors there is a lac…

Cited by 8SourcePDFScholar
2024

LaMI-DETR: Open-Vocabulary Detection with Language Model Instruction

ECCV 2024poster

"Existing methods enhance open-vocabulary object detection by leveraging the robust open-vocabulary recognition capabilities of Vision-Language Models (VLMs), such as CLIP. However, two main challenges emerge: (1) A deficiency in concept representation, where the category names in CLIP’s text space…

2024

MS-DETR: Efficient DETR Training with Mixed Supervision

CVPR 2024poster

DETR accomplishes end-to-end object detection through iteratively generating multiple object candidates based on image features and promoting one candidate for each ground-truth object. The traditional training procedure using one-to-one supervision in the original DETR lacks direct supervision for…

2024

Multi-Domain Incremental Learning for Face Presentation Attack Detection

AAAI 2024technical

Previous face Presentation Attack Detection (PAD) methods aim to improve the effectiveness of cross-domain tasks. However, in real-world scenarios, the original training data of the pre-trained model is not available due to data privacy or other reasons. Under these constraints, general methods for…

Cited by 17SourcePDFScholar
2024

OPEN: Object-wise Position Embedding for Multi-view 3D Object Detection

ECCV 2024poster

"Accurate depth information is crucial for enhancing the performance of multi-view 3D object detection. Despite the success of some existing multi-view 3D detectors utilizing pixel-wise depth supervision, they overlook two significant phenomena: 1) the depth supervision obtained from LiDAR points is…

2024

Octopus: A Multi-modal LLM with Parallel Recognition and Sequential Understanding

NeurIPS 2024poster

A mainstream of Multi-modal Large Language Models (MLLMs) have two essential functions, i.e., visual recognition (e.g., grounding) and understanding (e.g., visual question answering). Presently, all these MLLMs integrate visual recognition and understanding in a same sequential manner in the LLM hea…

Cited by 1SourcePDFScholar
2024

OpenGaussian: Towards Point-Level 3D Gaussian-based Open Vocabulary Understanding

NeurIPS 2024poster

This paper introduces OpenGaussian, a method based on 3D Gaussian Splatting (3DGS) that possesses the capability for 3D point-level open vocabulary understanding. Our primary motivation stems from observing that existing 3DGS-based open vocabulary methods mainly focus on 2D pixel-level parsing. Thes…

2024

ReSyncer: Rewiring Style-based Generator for Unified Audio-Visually Synced Facial Performer

ECCV 2024oral

"Lip-syncing videos with given audio is the foundation for various applications including the creation of virtual presenters or performers. While recent studies explore high-fidelity lip-sync with different techniques, their task-orientated models either require long-term videos for clip-specific tr…

Cited by 4SourcePDFScholar
2024

ShowMaker: Creating High-Fidelity 2D Human Video via Fine-Grained Diffusion Modeling

NeurIPS 2024poster

Although significant progress has been made in human video generation, most previous studies focus on either human facial animation or full-body animation, which cannot be directly applied to produce realistic conversational human videos with frequent hand gestures and various facial movements simul…

Cited by 4SourcePDFScholar
2024

TexOct: Generating Textures of 3D Models with Octree-based Diffusion

CVPR 2024poster

This paper focuses on synthesizing high-quality and complete textures directly on the surface of 3D models within 3D space. 2D diffusion-based methods face challenges in generating 2D texture maps due to the infinite possibilities of UV mapping for a given 3D mesh. Utilizing point clouds helps circu…

Cited by 1SourcePDFScholar
2024

Towards Unified Multi-granularity Text Detection with Interactive Attention

ICML 2024spotlight

Existing OCR engines or document image analysis systems typically rely on training separate models for text detection in varying scenarios and granularities, leading to significant computational complexity and resource demands. In this paper, we introduce "Detect Any Text" (DAT), an advanced paradig…

Cited by 1SourcePDFScholar
2024

VRP-SAM: SAM with Visual Reference Prompt

CVPR 2024poster

In this paper we propose a novel Visual Reference Prompt (VRP) encoder that empowers the Segment Anything Model (SAM) to utilize annotated reference images as prompts for segmentation creating the VRP-SAM model. In essence VRP-SAM can utilize annotated reference images to comprehend specific objects…

2023

Ambiguity-Resistant Semi-Supervised Learning for Dense Object Detection

CVPR 2023poster

With basic Semi-Supervised Object Detection (SSOD) techniques, one-stage detectors generally obtain limited promotions compared with two-stage clusters. We experimentally find that the root lies in two kinds of ambiguities: (1) Selection ambiguity that selected pseudo labels are less accurate, since…

2023

CAPE: Camera View Position Embedding for Multi-View 3D Object Detection

CVPR 2023poster

In this paper, we address the problem of detecting 3D objects from multi-view images. Current query-based methods rely on global 3D position embeddings (PE) to learn the geometric correspondence between images and 3D space. We claim that directly interacting 2D image features with global 3D PE could…

Cited by 51SourcePDFScholar
2023

CFCG: Semi-Supervised Semantic Segmentation via Cross-Fusion and Contour Guidance Supervision

ICCV 2023poster

Current state-of-the-art semi-supervised semantic segmentation (SSSS) methods typically adopt pseudo labeling and consistency regularization between multiple learners with different perturbations. Although the performance is desirable, many issues remain: (1) supervisions from a single learner tend…

Cited by 16PDFScholar
2023

Cyclically Disentangled Feature Translation for Face Anti-spoofing

AAAI 2023technical

Current domain adaptation methods for face anti-spoofing leverage labeled source domain data and unlabeled target domain data to obtain a promising generalizable decision boundary. However, it is usually difficult for these methods to achieve a perfect domain-invariant liveness feature disentangleme…

2023

Delicate Textured Mesh Recovery from NeRF via Adaptive Surface Refinement

ICCV 2023poster

Neural Radiance Fields (NeRF) have constituted a remarkable breakthrough in image-based 3D reconstruction. However, their implicit volumetric representations differ significantly from the widely-adopted polygonal meshes and lack support from common 3D software and hardware, making their rendering…

Cited by 120PDFcodeScholar
2023

Forward Flow for Novel View Synthesis of Dynamic Scenes

ICCV 2023oral

This paper proposes a neural radiance field (NeRF) approach for novel view synthesis of dynamic scenes using forward warping. Existing methods often adopt a static NeRF to represent the canonical space, and render dynamic images at other time steps by mapping the sampled 3D points back to the canoni…

Cited by 48PDFcodeScholar
2023

Gradient-based Sampling for Class Imbalanced Semi-supervised Object Detection

ICCV 2023poster

Current semi-supervised object detection (SSOD) algorithms typically assume class balanced datasets (PASCAL VOC etc.) or slightly class imbalanced datasets (MSCOCO, etc). This assumption can be easily violated since real world datasets can be extremely class imbalanced in nature, thus making the per…

Cited by 13PDFcodeScholar
2023

Graph Contrastive Learning for Skeleton-based Action Recognition

ICLR 2023poster

In the field of skeleton-based action recognition, current top-performing graph convolutional networks (GCNs) exploit intra-sequence context to construct adaptive graphs for feature aggregation. However, we argue that such context is still $\textit{local}$ since the rich cross-sequence relations hav…

2023

Group DETR: Fast DETR Training with Group-Wise One-to-Many Assignment

ICCV 2023poster

Detection transformer (DETR) relies on one-to-one assignment, assigning one ground-truth object to one prediction, for end-to-end detection without NMS post-processing. It is known that one-to-many assignment, assigning one ground-truth object to multiple predictions, succeeds in detection methods s…

Cited by 160PDFcodeScholar
2023

Group Pose: A Simple Baseline for End-to-End Multi-Person Pose Estimation

ICCV 2023poster

In this paper, we study the problem of end-to-end multi-person pose estimation. State-of-the-art solutions adopt the DETR-like framework, and mainly develop the complex decoder, e.g., regarding pose estimation as keypoint box detection and combining with human detection in ED-Pose, hierarchically pr…

Cited by 41PDFcodeScholar
2023

HAP: Structure-Aware Masked Image Modeling for Human-Centric Perception

NeurIPS 2023poster

Model pre-training is essential in human-centric perception. In this paper, we first introduce masked image modeling (MIM) as a pre-training approach for this task. Upon revisiting the MIM training strategy, we reveal that human structure priors offer significant potential. Motivated by this insight…

2023

LMR: A Large-Scale Multi-Reference Dataset for Reference-Based Super-Resolution

ICCV 2023poster

It is widely agreed that reference-based super-resolution (RefSR) achieves superior results by referring to similar high quality images, compared to single image super-resolution (SISR). Intuitively, the more references, the better performance. However, previous RefSR methods have all focused on sin…

Cited by 23PDFcodeScholar
2023

PSVT: End-to-End Multi-Person 3D Pose and Shape Estimation With Progressive Video Transformers

CVPR 2023poster

Existing methods of multi-person video 3D human Pose and Shape Estimation (PSE) typically adopt a two-stage strategy, which first detects human instances in each frame and then performs single-person PSE with temporal model. However, the global spatio-temporal context among spatial instances can not…

Cited by 35SourcePDFScholar
2023

Robust Video Portrait Reenactment via Personalized Representation Quantization

AAAI 2023technical

While progress has been made in the field of portrait reenactment, the problem of how to produce high-fidelity and robust videos remains. Recent studies normally find it challenging to handle rarely seen target poses due to the limitation of source data. This paper proposes the Video Portrait via No…

Cited by 5SourcePDFScholar
2023

Semi-DETR: Semi-Supervised Object Detection With Detection Transformers

CVPR 2023poster

We analyze the DETR-based framework on semi-supervised object detection (SSOD) and observe that (1) the one-to-one assignment strategy generates incorrect matching when the pseudo ground-truth bounding box is inaccurate, leading to training inefficiency; (2) DETR-based detectors lack deterministic c…

Cited by 61SourcePDFScholar
2023

StereoDistill: Pick the Cream from LiDAR for Distilling Stereo-Based 3D Object Detection

AAAI 2023technical

In this paper, we propose a cross-modal distillation method named StereoDistill to narrow the gap between the stereo and LiDAR-based approaches via distilling the stereo detectors from the superior LiDAR model at the response level, which is usually overlooked in 3D object detection distillation. Th…

Cited by 10SourcePDFScholar
2023

StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training

ICLR 2023poster

In this paper, we present StrucTexTv2, an effective document image pre-training framework, by performing masked visual-textual prediction. It consists of two self-supervised pre-training tasks: masked image modeling and masked language modeling, based on text region-level image masking. The proposed…

2023

StyleSync: High-Fidelity Generalized and Personalized Lip Sync in Style-Based Generator

CVPR 2023poster

Despite recent advances in syncing lip movements with any audio waves, current methods still struggle to balance generation quality and the model's generalization ability. Previous studies either require long-term data for training or produce a similar movement pattern on all subjects with low quali…

Cited by 71SourcePDFScholar
2022

Action Quality Assessment with Temporal Parsing Transformer

ECCV 2022poster

"Action Quality Assessment(AQA) is important for action understanding and resolving the task poses unique challenges due to subtle visual differences. Existing state-of-the-art methods typically rely on the holistic video representations for score regression or ranking, which limits the generalizati…

Cited by 60SourcePDFScholar
2022

CODER: Coupled Diversity-Sensitive Momentum Contrastive Learning for Image-Text Retrieval

ECCV 2022poster

"Image-Text Retrieval (ITR) is challenging in bridging visual and lingual modalities. Contrastive learning has been adopted by most prior arts. Except for limited amount of negative image-text pairs, the capability of constrastive learning is restricted by manually weighting negative pairs as well a…

Cited by 37SourcePDFScholar
2022

Delving into Sequential Patches for Deepfake Detection

NeurIPS 2022accept

Recent advances in face forgery techniques produce nearly visually untraceable deepfake videos, which could be leveraged with malicious intentions. As a result, researchers have been devoted to deepfake detection. Previous studies have identified the importance of local low-level cues and temporal i…

Cited by 66SourcePDFScholar
2022

Diverse Learner: Exploring Diverse Supervision for Semi-Supervised Object Detection

ECCV 2022poster

"Current state-of-the-art semi-supervised object detection methods (SSOD) typically adopt the teacher-student framework featured with pseudo labeling and Exponential Moving Average (EMA). Although the performance is desirable, many remaining issues still need to be resolved, for example: (1) the tea…

Cited by 5SourcePDFScholar
2022

Expressive Talking Head Generation With Granular Audio-Visual Control

CVPR 2022poster

Generating expressive talking heads is essential for creating virtual humans. However, existing one- or few-shot methods focus on lip-sync and head motion, ignoring the emotional expressions that make talking faces realistic. In this paper, we propose the Granularly Controlled Audio-Visual Talking H…

Cited by 148PDFScholar
2022

Few-Shot Font Generation by Learning Fine-Grained Local Styles

CVPR 2022poster

Few-shot font generation (FFG), which aims to generate a new font with a few examples, is gaining increasing attention due to the significant reduction in labor cost. A typical FFG pipeline considers characters in a standard font library as content glyphs and transfers them to a new target font by e…

Cited by 80PDFcodeScholar
2022

GitNet: Geometric Prior-Based Transformation for Birds-Eye-View Segmentation

ECCV 2022poster

"Birds-eye-view (BEV) semantic segmentation is critical for autonomous driving for its powerful spatial representation ability. It is challenging to estimate the BEV semantic maps from monocular images due to the spatial gap, since it is implicitly required to realize both the perspective-to-BEV tra…

Cited by 35SourcePDFScholar
2022

Human-Object Interaction Detection via Disentangled Transformer

CVPR 2022poster

Human-Object Interaction Detection tackles the problem of joint localization and classification of human object interactions. Existing HOI transformers either adopt a single decoder for triplet prediction, or utilize two parallel decoders to detect individual objects and interactions separately, and…

Cited by 77PDFScholar
2022

Implicit Sample Extension for Unsupervised Person Re-Identification

CVPR 2022poster

Most existing unsupervised person re-identification (Re-ID) methods use clustering to generate pseudo labels for model training. Unfortunately, clustering sometimes mixes different true identities together or splits the same identity into two or more sub clusters. Training on these noisy clusters su…

Cited by 135PDFcodeScholar
2022

MixFormer: Mixing Features Across Windows and Dimensions

CVPR 2022oral

While local-window self-attention performs notably in vision tasks, it suffers from limited receptive field and weak modeling capability issues. This is mainly because it performs self-attention within non-overlapped windows and shares weights on the channel dimension. We propose MixFormer to find a…

Cited by 161PDFcodeScholar
2022

MobileFaceSwap: A Lightweight Framework for Video Face Swapping

AAAI 2022technical

Advanced face swapping methods have achieved appealing results. However, most of these methods have many parameters and computations, which makes it challenging to apply them in real-time applications or deploy them on edge devices like mobile phones. In this work, we propose a lightweight Identity-…

2022

Neural Color Operators for Sequential Image Retouching

ECCV 2022poster

"We propose a novel image retouching method by modeling the retouching process as performing a sequence of newly introduced trainable neural color operators. The neural color operator mimics the behavior of traditional color operators and learns pixelwise color transformation while its strength is c…

2022

Predict, Prevent, and Evaluate: Disentangled Text-Driven Image Manipulation Empowered by Pre-Trained Vision-Language Model

CVPR 2022poster

To achieve disentangled image manipulation, previous works depend heavily on manual annotation. Meanwhile, the available manipulations are limited to a pre-defined set the models were trained for. We propose a novel framework, i.e., Predict, Prevent, and Evaluate (PPE), for disentangled text-driven…

Cited by 49PDFcodeScholar
2022

RTFormer: Efficient Design for Real-Time Semantic Segmentation with Transformer

NeurIPS 2022accept

Recently, transformer-based networks have shown impressive results in semantic segmentation. Yet for real-time semantic segmentation, pure CNN-based approaches still dominate in this field, due to the time-consuming computation mechanism of transformer. We propose RTFormer, an efficient dual-resolut…

2022

Rope3D: The Roadside Perception Dataset for Autonomous Driving and Monocular 3D Object Detection Task

CVPR 2022poster

Concurrent perception datasets for autonomous driving are mainly limited to frontal view with sensors mounted on the vehicle. None of them is designed for the overlooked roadside perception tasks. On the other hand, the data captured from roadside cameras have strengths over frontal-view data, which…

Cited by 135PDFScholar
2022

Self-Guided Hard Negative Generation for Unsupervised Person Re-Identification

IJCAI 2022poster

Recent unsupervised person re-identification (reID) methods mostly apply pseudo labels from clustering algorithms as supervision signals. Despite great success, this fashion is very likely to aggregate different identities with similar appearances into the same cluster. In result, the hard negative…

Cited by 12SourcePDFScholar
2022

Singular Value Fine-tuning: Few-shot Segmentation requires Few-parameters Fine-tuning

NeurIPS 2022accept

Freezing the pre-trained backbone has become a standard paradigm to avoid overfitting in few-shot segmentation. In this paper, we rethink the paradigm and explore a new regime: {\em fine-tuning a small part of parameters in the backbone}. We present a solution to overcome the overfitting problem, le…

2022

StyleSwap: Style-Based Generator Empowers Robust Face Swapping

ECCV 2022poster

"Numerous attempts have been made to the task of person-agnostic face swapping given its wide applications. While existing methods mostly rely on tedious network and loss designs, they still struggle in the information balancing between the source and target faces, and tend to produce visible artifa…

2022

Towards Bidirectional Arbitrary Image Rescaling: Joint Optimization and Cycle Idempotence

CVPR 2022poster

Deep learning based single image super-resolution models have been widely studied and superb results are achieved in upscaling low-resolution images with fixed scale factor and downscaling degradation kernel. To improve real world applicability of such models, there are growing interests to develop…

Cited by 39PDFScholar
2022

ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval

CVPR 2022poster

Visual appearance is considered to be the most important cue to understand images for cross-modal retrieval, while sometimes the scene text appearing in images can provide valuable information to understand the visual semantics. Most of existing cross-modal retrieval approaches ignore the usage of s…

Cited by 79PDFScholar
2021

ASCNet: Self-Supervised Video Representation Learning With Appearance-Speed Consistency

ICCV 2021poster

We study self-supervised video representation learning, which is a challenging task due to 1) sufficient labels for supervision; 2) unstructured and noisy visual information. Existing methods mainly use contrastive loss with video clips as the instances and learn visual representation by discriminat…

Cited by 56PDFScholar
2021

AdaAttN: Revisit Attention Mechanism in Arbitrary Neural Style Transfer

ICCV 2021poster

Fast arbitrary neural style transfer has attracted widespread attention from academic, industrial and art communities due to its flexibility in enabling various applications. Existing solutions either attentively fuse deep style feature into deep content feature without considering feature distribut…

Cited by 444PDFcodeScholar
2021

DOLG: Single-Stage Image Retrieval With Deep Orthogonal Fusion of Local and Global Features

ICCV 2021poster

Image Retrieval is a fundamental task of obtaining images similar to the query one from a database. A common image retrieval practice is to firstly retrieve candidate images via similarity search using global image features and then re-rank the candidates by leveraging their local features. Previous…

Cited by 167PDFcodeScholar
2021

Drafting and Revision: Laplacian Pyramid Network for Fast High-Quality Artistic Style Transfer

CVPR 2021poster

Artistic style transfer aims at migrating the style from an example image to a content image. Currently, optimization-based methods have achieved great stylization quality, but expensive time cost restricts their practical applications. Meanwhile, feed-forward methods still fail to synthesize comple…

Cited by 119PDFcodeScholar
2021

Dual-stream Network for Visual Recognition

NeurIPS 2021poster

Transformers with remarkable global representation capacities achieve competitive results for visual tasks, but fail to consider high-level local pattern information in input images. In this paper, we present a generic Dual-stream Network (DS-Net) to fully explore the representation capacity of loc…

Cited by 71SourcePDFScholar
2021

Dynamic Class Queue for Large Scale Face Recognition in the Wild

CVPR 2021poster

Learning discriminative representation using large-scale face datasets in the wild is crucial for real-world applications, yet it remains challenging. The difficulties lie in many aspects and this work focus on computing resource constraint and long-tailed class distribution. Recently, classificatio…

Cited by 32PDFcodeScholar
2021

EC-DARTS: Inducing Equalized and Consistent Optimization Into DARTS

ICCV 2021poster

Based on the relaxed search space, differential architecture search (DARTS) is efficient in searching for a high-performance architecture. However, the unbalanced competition among operations that have different trainable parameters causes the model collapse. Besides, the inconsistent structures in…

Cited by 9PDFcodeScholar
2021

FaceController: Controllable Attribute Editing for Face in the Wild

AAAI 2021technical

Face attribute editing aims to generate faces with one or multiple desired face attributes manipulated while other details are preserved. Unlike prior works such as GAN inversion which has an expensive reverse mapping process, we propose a simple feed-forward network to generate high-fidelity manipu…

2021

MVFNet: Multi-View Fusion Network for Efficient Video Recognition

AAAI 2021technical

Conventionally, spatiotemporal modeling network and its complexity are the two most concentrated research topics in video action recognition. Existing state-of-the-art methods have achieved excellent accuracy regardless of the complexity meanwhile efficient spatiotemporal modeling solutions are slig…

2021

PGNet: Real-time Arbitrarily-Shaped Text Spotting with Point Gathering Network

AAAI 2021technical

The reading of arbitrarily-shaped text has received increasing research attention. However, existing text spotters are mostly built on two-stage frameworks or character-based methods, which suffer from either Non-Maximum Suppression (NMS), Region-of-Interest (RoI) operations, or character-level anno…

2021

Paint Transformer: Feed Forward Neural Painting With Stroke Prediction

ICCV 2021poster

Neural painting refers to the procedure of producing a series of strokes for a given image and non-photo-realistically recreating it using neural networks. While reinforcement learning (RL) based agents can generate a stroke sequence step by step for this task, it is not easy to train a stable RL ag…

Cited by 93PDFcodeScholar
2021

Revealing the Reciprocal Relations Between Self-Supervised Stereo and Monocular Depth Estimation

ICCV 2021poster

Current self-supervised depth estimation algorithms mainly focus on either stereo or monocular only, neglecting the reciprocal relations between them. In this paper, we propose a simple yet effective framework to improve both stereo and monocular depth estimation by leveraging the underlying complem…

Cited by 34PDFScholar
2021

The Devil Is in the Task: Exploiting Reciprocal Appearance-Localization Features for Monocular 3D Object Detection

ICCV 2021poster

Low-cost monocular 3D object detection plays a fundamental role in autonomous driving, whereas its accuracy is still far from satisfactory. Our objective is to dig into the 3D object detection task and reformulate it as the sub-tasks of object localization and appearance perception, which benefits t…

Cited by 58PDFScholar
2021

Unsupervised Multi-Source Domain Adaptation for Person Re-Identification

CVPR 2021poster

Unsupervised domain adaptation (UDA) methods for person re-identification (re-ID) aim at transferring re-ID knowledge from labeled source data to unlabeled target data. Among these methods, the pseudo-label-based branch has achieved great success, whereas most of them only use limited data from a si…

Cited by 112PDFScholar
2021

Weakly-Supervised Spatio-Temporal Anomaly Detection in Surveillance Video

IJCAI 2021poster

In this paper, we introduce a novel task, referred to as Weakly-Supervised Spatio-Temporal Anomaly Detection (WSSTAD) in surveillance video. Specifically, given an untrimmed video, WSSTAD aims to localize a spatio-temporal tube (i.e., a sequence of bounding boxes at consecutive times) that encloses…

Cited by 75SourcePDFScholar
2020

Associate-3Ddet: Perceptual-to-Conceptual Association for 3D Point Cloud Object Detection

CVPR 2020poster

Object detection from 3D point clouds remains a challenging task, though recent studies pushed the envelope with the deep learning techniques. Owing to the severe spatial occlusion and inherent variance of point density with the distance to sensors, appearance of a same object varies a lot in point…

Cited by 115PDFScholar
2020

Discriminative Sounding Objects Localization via Self-supervised Audiovisual Matching

NeurIPS 2020poster

Discriminatively localizing sounding objects in cocktail-party, i.e., mixed sound scenes, is commonplace for humans, but still challenging for machines. In this paper, we propose a two-stage learning framework to perform self-supervised class-aware sounding object localization. First, we propose to…

2020

Graph-PCNN: Two Stage Human Pose Estimation with Graph Pose Refinement

ECCV 2020poster

Recently, most of the state-of-the-art human pose estimation methods are based on heatmap regression. The final coordinates of keypoints are obtained by decoding heatmap directly. In this paper, we aim to find a better approach to get more accurate localization results. We mainly put forward two sug…

Cited by 117SourcePDFScholar
2020

Monocular 3D Object Detection via Feature Domain Adaptation

ECCV 2020poster

Monocular 3D object detection is a challenging task due to unreliable depth, resulting in a distinct performance gap between monocular and LiDAR-based approaches. In this paper, we propose a novel domain adaptation based monocular 3D object detection framework named DA-3Ddet, which adapts the featur…

Cited by 58SourcePDFScholar
2020

Segment as Points for Efficient Online Multi-Object Tracking and Segmentation

ECCV 2020poster

Current multi-object tracking and segmentation (MOTS) methods follow the tracking-by-detection paradigm and adopt convolutions for feature extraction. However, as affected by the inherent receptive field, convolution based feature extraction inevitably mixes up the foreground features and the backgr…

2020

Towards Accurate Scene Text Recognition With Semantic Reasoning Networks

CVPR 2020poster

Scene text image contains two levels of contents: visual texture and semantic information. Although the previous scene text recognition methods have made great progress over the past few years, the research on mining semantic information to assist text recognition attracts less attention, only RNN-l…

Cited by 424PDFScholar
2019

A Mutual Learning Method for Salient Object Detection With Intertwined Multi-Supervision

CVPR 2019poster

Though deep learning techniques have made great progress in salient object detection recently, the predicted saliency maps still suffer from incomplete predictions due to the internal complexity of objects and inaccurate boundaries caused by strides in convolution and pooling operations. To alleviat…

Cited by 289PDFcodeScholar
2019

ACFNet: Attentional Class Feature Network for Semantic Segmentation

ICCV 2019poster

Recent works have made great progress in semantic segmentation by exploiting richer context, most of which are designed from a spatial perspective. In contrast to previous works, we present the concept of class center which extracts the global context from a categorical perspective. This class-level…

Cited by 357PDFScholar
2019

BMN: Boundary-Matching Network for Temporal Action Proposal Generation

ICCV 2019poster

Temporal action proposal generation is an challenging and promising task which aims to locate temporal regions in real-world videos where action or event may occur. Current bottom-up proposal generation methods can generate proposals with precise boundary, but cannot efficiently generate adequately…

Cited by 791PDFcodeScholar
2019

Chinese Street View Text: Large-Scale Chinese Text Reading With Partially Supervised Learning

ICCV 2019poster

Most existing text reading benchmarks make it difficult to evaluate the performance of more advanced deep learning models in large vocabularies due to the limited amount of training data. To address this issue, we introduce a new large-scale text reading benchmark dataset named Chinese Street View T…

Cited by 63PDFScholar
2019

Image Inpainting With Learnable Bidirectional Attention Maps

ICCV 2019poster

Most convolutional network (CNN)-based inpainting methods adopt standard convolution to indistinguishably treat valid pixels and holes, making them limited in handling irregular holes and more likely to generate inpainting results with color discrepancy and blurriness. Partial convolution has been s…

Cited by 319PDFcodeScholar
2019

Look More Than Once: An Accurate Detector for Text of Arbitrary Shapes

CVPR 2019poster

Previous scene text detection methods have progressed substantially over the past years. However, limited by the receptive field of CNNs and the simple representations like rectangle bounding box or quadrangle adopted to describe text, previous methods may fall short when dealing with more challengi…

Cited by 328PDFScholar
2019

Perspective-Guided Convolution Networks for Crowd Counting

ICCV 2019poster

In this paper, we propose a novel perspective-guided convolution (PGC) for convolutional neural network (CNN) based crowd counting (i.e. PGCNet), which aims to overcome the dramatic intra-scene scale variations of people due to the perspective effect. While most state-of-the-arts adopt multi-scale o…

Cited by 245PDFcodeScholar
2019

STGAN: A Unified Selective Transfer Network for Arbitrary Image Attribute Editing

CVPR 2019poster

Arbitrary attribute editing generally can be tackled by incorporating encoder-decoder and generative adversarial networks. However, the bottleneck layer in encoder-decoder usually gives rise to blurry and low quality editing result. And adding skip connections improves image quality at the cost of w…

Cited by 427PDFcodeScholar
2018

Compact Generalized Non-local Network

NeurIPS 2018poster

The non-local module is designed for capturing long-range spatio-temporal dependencies in images and videos. Although having shown excellent performance, it lacks the mechanism to model the interactions between positions across channels, which are of vital importance in recognizing fine-grained obje…

2018

Fine-grained Video Categorization with Redundancy Reduction Attention

ECCV 2018poster

For fine-grained categorization tasks, videos could serve as a better source than static images as videos have a higher chance of containing discriminative patterns. Nevertheless, a video sequence could also contain a lot of redundant and irrelevant frames. How to locate critical information of inte…

Cited by 60SourcePDFScholar
2018

Multi-Attention Multi-Class Constraint for Fine-grained Image Recognition

ECCV 2018poster

Attention-based learning for fine-grained image recognition remains a challenging task, where most of the existing methods treat each object part in isolation, while neglecting the correlations among them. In addition, the multi-stage or multi-scale mechanisms involved make the existing methods less…

Cited by 501SourcePDFScholar
2017

WordSup: Exploiting Word Annotations for Character Based Text Detection

ICCV 2017poster

Imagery texts are usually organized as a hierarchy of several visual elements, i.e. characters, words, text lines and text blocks. Among these elements, character is the most basic one for various languages such as Western, Chinese, Japanese, mathematical expression and etc. It is natural and conven…

Cited by 250PDFScholar