← Search

Bingbing Ni

86 accepted papers

2026

Differentiable Stroke Planning with Dual Parameterization for Efficient and High-Fidelity Painting Creation

CVPR 2026

In stroke-based rendering, search methods often get trapped in local minima due to discrete stroke placement, while differentiable optimizers lack structural awareness and produce unstructured layouts. To bridge this gap, we propose a dual representation that couples discrete polylines with continuo

Cited by 0SourceScholar
2025

Easy-editable Image Vectorization with Multi-layer Multi-scale Distributed Visual Feature Embedding

CVPR 2025poster

Current parameterized image representations embed visual information along the semantic boundaries and struggle to express the internal detailed texture structures of image components, leading to a lack of content consistency after image editing and driving. To address these challenges, this work pr…

Cited by 0SourcePDFScholar
2025

InstantSticker: Realistic Decal Blending via Disentangled Object Reconstruction

AAAI 2025technical

We present InstantSticker, a disentangled reconstruction pipeline based on Image-Based Lighting (IBL), which focuses on highly realistic decal blending, simulates stickers attached to the reconstructed surface, and allows for instant editing and real-time rendering. To achieve stereoscopic impressio…

2025

Neural Block Compression: Variable Bitrates Feature Blocks for Texture Representation

AAAI 2025technical

The imperative for compression of material textures emerges from the critical demand for high-quality rendering, which necessitates sophisticated textures that, in turn, require substantial storage and memory resources. Thus, low-bitrate compression is crucial, especially in modern games demanding h…

Cited by 0SourcePDFScholar
2025

RAGDiffusion: Faithful Cloth Generation via External Knowledge Assimilation

ICCV 2025poster

Standard clothing asset generation involves restoring forward-facing flat-lay garment images displayed on a clear background by extracting clothing information from diverse real-world contexts, which presents significant challenges due to highly standardized structure sampling distributions and clot…

Cited by 0SourcePDFScholar
2025

ShoeFit: A New Dataset and Dual-image-stream DiT Framework for Virtual Footwear Try-On

NeurIPS 2025poster

Virtual footwear try-on (VFTON), a critical yet underexplored area in virtual try-on (VTON), aims to synthesize faithful try-on results given diverse footwear and model images while maintaining 3D consistency and texture authenticity. Unlike conventional garment-focused VTON methods, VFTON present…

Cited by 0SourceScholar
2025

SinGS: Animatable Single-Image Human Gaussian Splats with Kinematic Priors

CVPR 2025poster

Despite significant advances in accurately estimating geometry in contemporary single-image 3D human reconstruction, creating a high-quality, efficient, and animatable 3D avatar remains an open challenge. Two key obstacles persist: incomplete observation and inconsistent 3D priors. To address these…

2024

AnyFit: Controllable Virtual Try-on for Any Combination of Attire Across Any Scenario

NeurIPS 2024poster

While image-based virtual try-on has made significant strides, emerging approaches still fall short of delivering high-fidelity and robust fitting images across various scenarios, as their models suffer from issues of ill-fitted garment styles and quality degrading during the training process, not t…

Cited by 7SourcePDFScholar
2024

FocalDreamer: Text-Driven 3D Editing via Focal-Fusion Assembly

AAAI 2024technical

While text-3D editing has made significant strides in leveraging score distillation sampling, emerging approaches still fall short in delivering separable, precise and consistent outcomes that are vital to content creation. In response, we introduce FocalDreamer, a framework that merges base shape w…

Cited by 56SourcePDFScholar
2024

Intrinsic Phase-Preserving Networks for Depth Super Resolution

AAAI 2024technical

Depth map super-resolution (DSR) plays an indispensable role in 3D vision. We discover an non-trivial spectral phenomenon: the components of high-resolution (HR) and low-resolution (LR) depth maps manifest the same intrinsic phase, and the spectral phase of RGB is a superset of them, which suggests…

2024

Real-Time Neural BRDF with Spherically Distributed Primitives

CVPR 2024poster

We propose a neural reflectance model (NeuBRDF) that offers highly versatile material representation yet with light memory and neural computation consumption towards achieving real-time rendering. The results depicted in Fig. 1 rendered at full HD resolution on a contemporary desktop machine demonst…

Cited by 2SourcePDFScholar
2024

Towards High-fidelity Artistic Image Vectorization via Texture-Encapsulated Shape Parameterization

CVPR 2024poster

We develop a novel vectorized image representation scheme accommodating both shape/geometry and texture in a decoupled way particularly tailored for reconstruction and editing tasks of artistic/design images such as Emojis and Cliparts. In the heart of this representation is a set of sparsely and un…

Cited by 1SourcePDFScholar
2023

AudioEar: Single-View Ear Reconstruction for Personalized Spatial Audio

AAAI 2023technical

Spatial audio, which focuses on immersive 3D sound rendering, is widely applied in the acoustic industry. One of the key problems of current spatial audio rendering methods is the lack of personalization based on different anatomies of individuals, which is essential to produce accurate sound source…

2023

Boosting Point Clouds Rendering via Radiance Mapping

AAAI 2023technical

Recent years we have witnessed rapid development in NeRF-based image rendering due to its high quality. However, point clouds rendering is somehow less explored. Compared to NeRF-based rendering which suffers from dense spatial sampling, point clouds rendering is naturally less computation intensive…

2023

CiaoSR: Continuous Implicit Attention-in-Attention Network for Arbitrary-Scale Image Super-Resolution

CVPR 2023poster

Learning continuous image representations is recently gaining popularity for image super-resolution (SR) because of its ability to reconstruct high-resolution images with arbitrary scales from low-resolution inputs. Existing methods mostly ensemble nearby features to predict the new pixel at any que…

2023

Deep Arbitrary-Scale Image Super-Resolution via Scale-Equivariance Pursuit

CVPR 2023poster

The ability of scale-equivariance processing blocks plays a central role in arbitrary-scale image super-resolution tasks. Inspired by this crucial observation, this work proposes two novel scale-equivariant modules within a transformer-style framework to enhance arbitrary-scale image super-resolutio…

2023

Fast Fluid Simulation via Dynamic Multi-Scale Gridding

AAAI 2023technical

Recent works on learning-based frameworks for Lagrangian (i.e., particle-based) fluid simulation, though bypassing iterative pressure projection via efficient convolution operators, are still time-consuming due to excessive amount of particles. To address this challenge, we propose a dynamic multi-s…

Cited by 4SourcePDFScholar
2023

Frequency-Modulated Point Cloud Rendering With Easy Editing

CVPR 2023highlight

We develop an effective point cloud rendering pipeline for novel view synthesis, which enables high fidelity local detail reconstruction, real-time rendering and user-friendly editing. In the heart of our pipeline is an adaptive frequency modulation module called Adaptive Frequency Net (AFNet), whic…

2023

Generalized Deep 3D Shape Prior via Part-Discretized Diffusion Process

CVPR 2023poster

We develop a generalized 3D shape generation prior model, tailored for multiple 3D tasks including unconditional shape generation, point cloud completion, and cross-modality shape generation, etc. On one hand, to precisely capture local fine detailed shape information, a vector quantized variational…

2023

Learning Continuous Depth Representation via Geometric Spatial Aggregator

AAAI 2023technical

Depth map super-resolution (DSR) has been a fundamental task for 3D computer vision. While arbitrary scale DSR is a more realistic setting in this scenario, previous approaches predominantly suffer from the issue of inefficient real-numbered scale upsampling. To explicitly address this issue, we pro…

2023

Learning Shape Primitives via Implicit Convexity Regularization

ICCV 2023poster

Shape primitives decomposition has been an important and long-standing task in 3D shape analysis. Prior arts heavily rely on 3D point clouds or voxel data for shape primitives extraction, which are less practical in real-world scenarios. This paper proposes to learn shape primitives from multi-view…

Cited by 3PDFcodeScholar
2023

Movienet-PS: A Large-Scale Person Search Dataset in the Wild

ICASSP 2023accepted

Person search (PS) aims to jointly localize and identify a query person from natural, uncropped images. Existing works unintentionally adopt pedestrians (with similar poses and unchanging clothing) as the query and restrict the application scenarios in surveillance. This is due to that most PS datas…

Cited by 0SourceScholar
2023

Omni Aggregation Networks for Lightweight Image Super-Resolution

CVPR 2023poster

While lightweight ViT framework has made tremendous progress in image super-resolution, its uni-dimensional self-attention modeling, as well as homogeneous aggregation scheme, limit its effective receptive field (ERF) to include more comprehensive interactions from both spatial and channel dimension…

2023

Towards Interpreting and Utilizing Symmetry Property in Adversarial Examples

AAAI 2023technical

In this paper, we identify symmetry property in adversarial scenario by viewing adversarial attack in a fine-grained manner. A newly designed metric called attack proportion, is thus proposed to count the proportion of the adversarial examples misclassified between classes. We observe that the distr…

Cited by 2SourcePDFScholar
2022

Bi-volution: A Static and Dynamic Coupled Filter

AAAI 2022technical

Dynamic convolution has achieved significant gain in performance and computational complexity, thanks to its powerful representation capability given limited filter number/layers. However, SOTA dynamic convolution operators are sensitive to input noises (e.g., Gaussian noise, shot noise, e.t.c.) an…

2022

Contrastive Regression for Domain Adaptation on Gaze Estimation

CVPR 2022poster

Appearance-based Gaze Estimation leverages deep neural networks to regress the gaze direction from monocular images and achieve impressive performance. However, its success depends on expensive and cumbersome annotation capture. When lacking precise annotation, the large domain gap hinders the perfo…

Cited by 99PDFScholar
2022

Exploring Visual Context for Weakly Supervised Person Search

AAAI 2022technical

Person search has recently emerged as a challenging task that jointly addresses pedestrian detection and person re-identification. Existing approaches follow a fully supervised setting where both bounding box and identity annotations are available. However, annotating identities is labor-intensive,…

2022

HCSC: Hierarchical Contrastive Selective Coding

CVPR 2022poster

Hierarchical semantic structures naturally exist in an image dataset, in which several semantically relevant image clusters can be further integrated into a larger cluster with coarser-grained semantics. Capturing such structures with image representations can greatly benefit the semantic understand…

Cited by 101PDFcodeScholar
2022

ImplicitAtlas: Learning Deformable Shape Templates in Medical Imaging

CVPR 2022poster

Deep implicit shape models have become popular in the computer vision community at large but less so for biomedical applications. This is in part because large training databases do not exist and in part because biomedical annotations are often noisy. In this paper, we show that by introducing templ…

Cited by 35PDFScholar
2022

Object Wake-Up: 3D Object Rigging from a Single Image

ECCV 2022poster

"Given a single chair image, could we wake it up by reconstructing its 3D shape and skeleton, as well as animating its plausible articulations and motions, similar to that of human modeling? It is a new problem that not only goes beyond image-based object reconstruction but also involves articulated…

Cited by 7SourcePDFScholar
2022

RainNet: A Large-Scale Imagery Dataset and Benchmark for Spatial Precipitation Downscaling

NeurIPS 2022accept

AI-for-science approaches have been applied to solve scientific problems (e.g., nuclear fusion, ecology, genomics, meteorology) and have achieved highly promising results. Spatial precipitation downscaling is one of the most important meteorological problem and urgently requires the participation of…

2022

Representation-Agnostic Shape Fields

ICLR 2022poster

3D shape analysis has been widely explored in the era of deep learning. Numerous models have been developed for various 3D data representation formats, e.g., MeshCNN for meshes, PointNet for point clouds and VoxNet for voxels. In this study, we present Representation-Agnostic Shape Fields (RASF), a…

2021

3D Human Action Representation Learning via Cross-View Consistency Pursuit

CVPR 2021poster

In this work, we propose a Cross-view Contrastive Learning framework for unsupervised 3D skeleton-based action representation (CrosSCLR), by leveraging multi-view complementary supervision signal. CrosSCLR consists of both single-view contrastive learning (SkeletonCLR) and cross-view consistent know…

Cited by 240PDFcodeScholar
2021

Bilevel Online Adaptation for Out-of-Domain Human Mesh Reconstruction

CVPR 2021poster

This paper considers a new problem of adapting a pre-trained model of human mesh reconstruction to out-of-domain streaming videos. However, most previous methods based on the parametric SMPL model underperform in new domains with unexpected, domain-specific attributes, such as camera parameters, len…

Cited by 63PDFcodeScholar
2021

Context-Aware Image Inpainting with Learned Semantic Priors

IJCAI 2021poster

Recent advances in image inpainting have shown impressive results for generating plausible visual details on rather simple backgrounds. However, for complex scenes, it is still challenging to restore reasonable contents as the contextual information within the missing regions tends to be ambiguous.…

2021

Cross-Category Video Highlight Detection via Set-Based Learning

ICCV 2021poster

Autonomous highlight detection is crucial for enhancing the efficiency of video browsing on social media platforms. To attain this goal in a data-driven way, one may often face the situation where highlight annotations are not available on the target video category used in practice, while the superv…

Cited by 62PDFcodeScholar
2021

Joint Modeling of Visual Objects and Relations for Scene Graph Generation

NeurIPS 2021poster

An in-depth scene understanding usually requires recognizing all the objects and their relations in an image, encoded as a scene graph. Most existing approaches for scene graph generation first independently recognize each object and then predict their relations independently. Though these approache…

Cited by 16SourcePDFScholar
2021

Progressive Stage-Wise Learning for Unsupervised Feature Representation Enhancement

CVPR 2021poster

Unsupervised learning methods have recently shown their competitiveness against supervised training. Typically, these methods use a single objective to train the entire network. But one distinct advantage of unsupervised over supervised learning is that the former possesses more variety and freedom…

Cited by 6PDFScholar
2021

Self-supervised Graph-level Representation Learning with Local and Global Structure

ICML 2021spotlight

This paper studies unsupervised/self-supervised whole-graph representation learning, which is critical in many tasks such as molecule properties prediction in drug and material discovery. Existing methods mainly focus on preserving the local similarity structure between different graph instances but…

2021

Shape Self-Correction for Unsupervised Point Cloud Understanding

ICCV 2021poster

We develop a novel self-supervised learning method named Shape Self-Correction for point cloud analysis. Our method is motivated by the principle that a good shape representation should be able to find distorted parts of a shape and correct them. To learn strong shape representations in an unsupervi…

Cited by 58PDFScholar
2021

Skeleton2Mesh: Kinematics Prior Injected Unsupervised Human Mesh Recovery

ICCV 2021poster

In this paper, we decouple unsupervised human mesh recovery into the well-studied problems of unsupervised 3D pose estimation, and human mesh recovery from estimated 3D skeletons, focusing on the latter task. The challenges of the latter task are two folds: (1) pose failure (i.e., pose mismatching -…

Cited by 29PDFcodeScholar
2021

Sketch Generation with Drawing Process Guided by Vector Flow and Grayscale

AAAI 2021technical

We propose a novel image-to-pencil translation method that could not only generate high-quality pencil sketches but also offer the drawing process. Existing pencil sketch algorithms are based on texture rendering rather than the direct imitation of strokes, making them unable to show the drawing pro…

2021

Towards Alleviating the Modeling Ambiguity of Unsupervised Monocular 3D Human Pose Estimation

ICCV 2021poster

In this work, we study the ambiguity problem in the task of unsupervised 3D human pose estimation from 2D counterpart. On one hand, without explicit annotation, the scale of 3D pose is difficult to be accurately captured (scale ambiguity). On the other hand, one 2D pose might correspond to multiple…

Cited by 49PDFScholar
2020

CooGAN: A Memory-Efficient Framework for High-Resolution Facial Attribute Editing

ECCV 2020poster

In contrast to great success of memory-consuming face editing methods at a low resolution, to manipulate high-resolution (HR) facial images, \ie, typically larger than $768^2$ pixels, with very limited memory is still challenging. This is due to the reasons of 1) intractable huge demand of memory; 2…

2020

Cross-Domain Detection via Graph-Induced Prototype Alignment

CVPR 2020oral

Applying the knowledge of an object detector trained on a specific domain directly onto a new domain is risky, as the gap between two domains can severely degrade model's performance. Furthermore, since different instances commonly embody distinct modal information in object detection scenario, the…

Cited by 302PDFcodeScholar
2020

Deep Kinematics Analysis for Monocular 3D Human Pose Estimation

CVPR 2020poster

For monocular 3D pose estimation conditioned on 2D detection, noisy/unreliable input is a key obstacle in this task. Simple structure constraints attempting to tackle this problem, e.g., symmetry loss and joint angle limit, could only provide marginal improvements and are commonly treated as auxilia…

Cited by 233PDFScholar
2020

Hierarchical Style-based Networks for Motion Synthesis

ECCV 2020poster

Generating diverse and natural behaviors is one of the long-standing goals for creating intelligent characters in the animated world. In this paper, we propose an unsupervised method for generating long-range, diverse and plausible behaviors to achieve a specific goal location. Our proposed method l…

Cited by 35SourcePDFScholar
2020

Learning Black-Box Attackers with Transferable Priors and Query Feedback

NeurIPS 2020poster

This paper addresses the challenging black-box adversarial attack problem, where only classification confidence of a victim model is available. Inspired by consistency of visual saliency between different vision models, a surrogate model is expected to improve the attack performance via transferabil…

2020

Learning to Combine: Knowledge Aggregation for Multi-Source Domain Adaptation

ECCV 2020poster

Transferring knowledges learned from multiple source domains to target domain is a more practical and challenging task than conventional single-source domain adaptation. Furthermore, the increase of modalities brings more difficulty in aligning feature distributions among multiple domains. To mitiga…

2019

Dynamic Points Agglomeration for Hierarchical Point Sets Learning

ICCV 2019poster

Many previous works on point sets learning achieve excellent performance with hierarchical architecture. Their strategies towards points agglomeration, however, only perform points sampling and grouping in original Euclidean space in a fixed way. These heuristic and task-irrelevant strategies severe…

Cited by 135PDFScholar
2019

Modeling Point Clouds With Self-Attention and Gumbel Subset Sampling

CVPR 2019poster

Geometric deep learning is increasingly important thanks to the popularity of 3D sensors. Inspired by the recent advances in NLP domain, the self-attention transformer is introduced to consume the point clouds. We develop Point Attention Transformers (PATs), using a parameter-efficient Group Shuffle…

Cited by 519PDFScholar
2019

Variational Convolutional Neural Network Pruning

CVPR 2019poster

We propose a variational Bayesian scheme for pruning convolutional neural networks in channel level. This idea is motivated by the fact that deterministic value based pruning methods are inherently improper and unstable. In a nutshell, variational technique is introduced to estimate distribution of…

Cited by 455PDFScholar
2018

Crowd Counting via Adversarial Cross-Scale Consistency Pursuit

CVPR 2018poster

Crowd counting or density estimation is a challenging task in computer vision due to large scale variations, perspective distortions and serious occlusions, etc. Existing methods generally suffers from two issues: 1) the model averaging effects in multi-scale CNNs induced by the widely adopted L2 re…

Cited by 414SourcePDFScholar
2018

Deep Regression Tracking with Shrinkage Loss

ECCV 2018poster

Regression trackers directly learn a mapping from regularly dense samples of target objects to soft labels, which are usually generated by a Gaussian function, to estimate target positions. Due to the potential for fast-tracking and easy implementation, regression trackers have received increasing a…

2018

Fine-Grained Video Captioning for Sports Narrative

CVPR 2018poster

Despite recent emergence of video caption methods, how to generate fine-grained video descriptions (i.e., long and detailed commentary about individual movements of multiple subjects as well as their frequent interactions) is far from being solved, which however has great applications such as automa…

Cited by 76SourcePDFScholar
2018

Geometric Constrained Joint Lane Segmentation and Lane Boundary Detection

ECCV 2018poster

Lane detection is playing an indispensable role in advanced driver assistance systems. The existing approaches for lane detection can be categorized as lane area segmentation and lane boundary detection. Most of these methods abandon a great quantity of complementary information, such as geometric p…

2018

Multiple Granularity Group Interaction Prediction

CVPR 2018poster

Most human activity analysis works (i.e., recognition or prediction) only focus on a single granularity, i.e., either modelling global motion based on the coarse level movement such as human trajectories or forecasting future detailed action based on body parts’ movement such as skeleton motion. In…

Cited by 25SourcePDFScholar
2018

Pose Transferrable Person Re-Identification

CVPR 2018poster

Person re-identification (ReID) is an important task in the field of intelligent security. A key challenge is how to capture human pose variations, while existing benchmarks (i.e., Market1501, DukeMTMC-reID, CUHK03, etc.) do NOT provide sufficient pose coverage to train a robust ReID system. To add…

Cited by 456SourcePDFScholar
2017

Binary Coding for Partial Action Analysis With Limited Observation Ratios

CVPR 2017poster

Traditional action recognition methods aim to recognize actions with complete observations/executions. However, it is often difficult to capture fully executed actions due to occlusions, interruptions, etc. Meanwhile, action prediction/recognition in advance based on partial observations is essentia…

Cited by 34PDFScholar
2017

Performance Guaranteed Network Acceleration via High-Order Residual Quantization

ICCV 2017poster

Input binarization has shown to be an effective way for network acceleration. However, previous binarization scheme could be regarded as simple pixel-wise thresholding operations (i.e., order-one approximation) and suffers a big accuracy loss. In this paper, we propose a high-order binarization sche…

Cited by 137PDFScholar
2017

Zero-Shot Action Recognition With Error-Correcting Output Codes

CVPR 2017poster

Recently, zero-shot action recognition (ZSAR) has emerged with the explosive growth of action categories. In this paper, we explore ZSAR from a novel perspective by adopting the Error-Correcting Output Codes (dubbed ZSECOC). Our ZSECOC equips the conventional ECOC with the additional capability of Z…

Cited by 186PDFScholar
2016

Cascaded Interactional Targeting Network for Egocentric Video Analysis

CVPR 2016poster

Knowing how hands move and what object is being manipulated are two key sub-tasks for analyzing first-person (egocentric) action. However, lack of fully annotated hand data as well as imprecise foreground segmentation make either sub-task challenging. This work aims to explicitly address these two i…

Cited by 68PDFScholar
2016

Progressively Parsing Interactional Objects for Fine Grained Action Detection

CVPR 2016poster

Fine grained video action analysis often requires reliable detection and tracking of various interacting objects and human body parts, denoted as interactional object parsing. However, most of the previous methods based on either independent or joint object detection might suffer from high model com…

Cited by 95PDFcodeScholar
2016

Temporal Action Localization With Pyramid of Score Distribution Features

CVPR 2016spotlight

We investigate the feature design and classification architectures in temporal action localization. This application focuses on detecting and labeling actions in untrimmed videos, which brings more challenge than classifying pre-segmented videos. The major difficulty for action localization is the u…

Cited by 232PDFScholar
2015

Interaction Part Mining: A Mid-Level Approach for Fine-Grained Action Recognition

CVPR 2015poster

Modeling human-object interactions and manipulating motions lies in the heart of fine-grained action recognition. Previous methods heavily rely on explicit detection of the object being interacted, which requires intensive human labour on object annotation. To bypass this constraint and achieve bett…

Cited by 103SourcePDFScholar
2015

Motion Part Regularization: Improving Action Recognition via Trajectory Selection

CVPR 2015poster

Dense local motion features such as dense trajectories have been widely used in action recognition. For most actions, only a few local features (e.g., critical movements of the hand, arm, leg etc.) are responsible to the action label. Therefore, discovering important motion part will lead to a more…

Cited by 111SourcePDFScholar