← Search

Jianfei Cai

109 accepted papers

2026

An Empirical Study on How Video-LLMs Answer Video Questions

CVPR 2026

Taking advantage of large-scale data and pretrained language models, Video Large Language Models (Video-LLMs) have shown strong capabilities in answering video questions. However, most existing efforts focus on improving performance, with limited attention to understanding their internal mechanisms.

Cited by 0SourceScholar
2026

CoT Vectors: Transferring and Probing the Reasoning Mechanisms of LLMs

ICLR 2026poster

Chain-of-Thought (CoT) prompting has emerged as a powerful approach to enhancing the reasoning capabilities of Large Language Models (LLMs). However, existing implementations, such as in-context learning and fine-tuning, remain costly and inefficient. To improve CoT reasoning at a lower cost, and in…

Cited by 0SourceScholar
2026

Gradient-Aligned Calibration for Post-Training Quantization of Diffusion Models

ICLR 2026poster

Diffusion models have shown remarkable performance in image synthesis by progressively estimating a smooth transition from a Gaussian distribution of noise to a real image. Unfortunately, their practical deployment is limited by slow inference speed, high memory usage, and the computational demands…

Cited by 0SourceScholar
2026

Marginalized Generalized IoU (MGIoU): A Unified Objective Function for Optimizing Convex Parametric Shapes

AAAI 2026technical

Optimizing the similarity between parametric shapes is crucial for numerous computer vision tasks, where Intersection over Union (IoU) stands as the canonical measure. However, existing optimization methods exhibit significant shortcomings: regression-based losses like L1/L2 lack correlation with Io

Cited by 0SourcePDFScholar
2026

MotionCrafter: Dense Geometry and Motion Reconstruction with a 4D VAE

CVPR 2026

We present MotionCrafter, a framework that leverages video generators to jointly reconstruct 4D geometry and estimate dense motion from a monocular video. The key idea is a joint representation of dense 3D point maps and 3D scene flows in a shared coordinate system, together with a 4D VAE tailored t

Cited by 0SourceScholar
2026

Omni2Sound: Towards Unified Video-Text-to-Audio Generation

CVPR 2026

Training a unified model integrating video-to-audio (V2A), text-to-audio (T2A), and joint video-text-to-audio (VT2A) generation offers significant application flexibility, yet faces two unexplored foundational challenges: (1) the scarcity of high-quality audio captions with tight V-A-T alignment, le

Cited by 0SourceScholar
2026

Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos

CVPR 2026

We present Ov3R, a novel framework for open-vocabulary semantic 3D reconstruction from RGB video streams, designed to advance Spatial AI. The system features two key components: CLIP3R, a CLIP-informed 3D reconstruction module that predicts dense point maps from overlapping clips alongside object-le

Cited by 0SourceScholar
2026

PCGS: Progressive Compression of 3D Gaussian Splatting

AAAI 2026technical

3D Gaussian Splatting (3DGS) achieves impressive rendering fidelity and speed for novel view synthesis. However, its substantial data size poses a significant challenge for practical applications. While many compression techniques have been proposed, they fail to efficiently utilize existing bitstre

Cited by 0SourcePDFScholar
2026

PanFlow: Decoupled Motion Control for Panoramic Video Generation

AAAI 2026technical

Panoramic video generation has attracted growing attention due to its applications in virtual reality and immersive media. However, existing methods lack explicit motion control and struggle to generate scenes with large and complex motions. We propose PanFlow a novel approach that exploits the sphe

Cited by 0SourcePDFScholar
2026

Single-Rollout Hidden-State Dynamics for Training-Free RLVR Data Selection

ICML 2026poster

Reinforcement learning with verifiable rewards (RLVR) can yield large reasoning gains from very few training instances, yet its strong sensitivity to which instances are used makes data selection a central bottleneck. Most existing selection pipelines rely on training-time optimization signals and/o…

Cited by 0SourceScholar
2026

Unified Camera Positional Encoding for Controlled Video Generation

CVPR 2026

Transformers have emerged as a universal backbone across 3D perception, video generation, and world models for autonomous driving and embodied AI, where understanding camera geometry is essential for grounding visual observations in three-dimensional space. However, existing camera encoding methods

Cited by 0SourcecodeScholar
2026

VQ-VA World: Towards High-Quality Visual Question-Visual Answering

CVPR 2026

This paper studies Visual Question-Visual Answering (VQ-VA): generating an image, rather than text, in response to a visual question---an ability that has recently emerged in proprietary systems such as NanoBanana and GPT-Image. To also bring this capability to open-source models, we introduce VQ-VA

Cited by 0SourcecodeScholar
2026

Virtual Full-stack Scanning of Brain MRI via Imputing Any Quantised Code

CVPR 2026

Magnetic resonance imaging (MRI) is a powerful and versatile imaging technique, offering a wide spectrum of information about the anatomy by employing different acquisition modalities. However, in the clinical workflow, it is impractical to collect all relevant modalities due to the scan time and co

Cited by 0SourcecodeScholar
2026

Where and What Matters: Sensitivity-Aware Task Vectors for Many-Shot Multimodal In-Context Learning

AAAI 2026technical

Large Multimodal Models (LMMs) have shown promising in-context learning (ICL) capabilities, but scaling to many-shot settings remains difficult due to limited context length and high inference cost. To address these challenges, task-vector-based methods have been explored by inserting compact repres

Cited by 0SourcePDFScholar
2025

DrVideo: Document Retrieval Based Long Video Understanding

CVPR 2025poster

Most of the existing methods for video understanding primarily focus on videos only lasting tens of seconds, with limited exploration of techniques for handling long videos. The increased number of frames in long videos poses two main challenges: difficulty in locating key information and performing…

2025

Fast Feedforward 3D Gaussian Splatting Compression

ICLR 2025poster

With 3D Gaussian Splatting (3DGS) advancing real-time and high-fidelity rendering for novel view synthesis, storage requirements pose challenges for their widespread adoption. Although various compression techniques have been proposed, previous art suffers from a common limitation: for any existing…

2025

PaRa: Personalizing Text-to-Image Diffusion via Parameter Rank Reduction

ICLR 2025spotlight

Personalizing a large-scale pretrained Text-to-Image (T2I) diffusion model is chal- lenging as it typically struggles to make an appropriate trade-off between its training data distribution and the target distribution, i.e., learning a novel concept with only a few target images to achieve personali…

Cited by 0SourcePDFScholar
2025

PanSplat: 4K Panorama Synthesis with Feed-Forward Gaussian Splatting

CVPR 2025poster

With the advent of portable 360deg cameras, panorama has gained significant attention in applications like virtual reality (VR), virtual tours, robotics, and autonomous driving. As a result, wide-baseline panorama view synthesis has emerged as a vital task, where high resolution, fast inference, and…

2025

Point-Cache: Test-time Dynamic and Hierarchical Cache for Robust and Generalizable Point Cloud Analysis

CVPR 2025poster

This paper proposes a general solution to enable point cloud recognition models to handle distribution shifts at test time. Unlike prior methods, which rely heavily on training data (often inaccessible during online inference) and are limited to recognizing a fixed set of point cloud classes predefi…

2025

T-Stitch: Accelerating Sampling in Pre-Trained Diffusion Models with Trajectory Stitching

ICLR 2025poster

Sampling from diffusion probabilistic models (DPMs) is often expensive for high-quality image generation and typically requires many steps with a large model. In this paper, we introduce sampling Trajectory Stitching (T-Stitch), a simple yet efficient technique to improve the sampling efficiency wit…

2025

VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical Prior

ICCV 2025accepted

Video diffusion models (VDMs) have advanced significantly in recent years, enabling the generation of highly realistic videos and drawing the attention of the community in their potential as world simulators. However, despite their capabilities, VDMs often fail to produce physically plausible videos…

2024

Differentiable Convex Polyhedra Optimization from Multi-view Images

ECCV 2024poster

"This paper presents a novel approach for the differentiable rendering of convex polyhedra, addressing the limitations of recent methods that rely on implicit field supervision. Our technique introduces a strategy that combines non-differentiable computation of hyperplane intersection through dualit…

2024

Diffusion Model for Robust Multi-Sensor Fusion in 3D Object Detection and BEV Segmentation

ECCV 2024poster

"Diffusion models have recently gained prominence as powerful deep generative models, demonstrating unmatched performance across various domains. However, their potential in multi-sensor fusion remains largely unexplored. In this work, we introduce “DifFUSER”, a novel approach that leverages diffusi…

Cited by 2SourcePDFScholar
2024

Diversified and Personalized Multi-rater Medical Image Segmentation

CVPR 2024highlight

Annotation ambiguity due to inherent data uncertainties such as blurred boundaries in medical scans and different observer expertise and preferences has become a major obstacle for training deep-learning based medical image segmentation models. To address it the common practice is to gather multiple…

2024

GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI

NeurIPS 2024poster

Large Vision-Language Models (LVLMs) are capable of handling diverse data types such as imaging, text, and physiological signals, and can be applied in various fields. In the medical field, LVLMs have a high potential to offer substantial assistance for diagnosis and treatment. Before that, it is cr…

2024

Generative Region-Language Pretraining for Open-Ended Object Detection

CVPR 2024poster

In recent research significant attention has been devoted to the open-vocabulary object detection task aiming to generalize beyond the limited number of classes labeled during training and detect objects described by arbitrary category names at inference. Compared with conventional object detection…

2024

HAC: Hash-grid Assisted Context for 3D Gaussian Splatting Compression

ECCV 2024poster

"3D Gaussian Splatting (3DGS) has emerged as a promising framework for novel view synthesis, boasting rapid rendering speed with high fidelity. However, the substantial Gaussians and their associated attributes necessitate effective compression techniques. Nevertheless, the sparse and unorganized na…

2024

JRDB-PanoTrack: An Open-world Panoptic Segmentation and Tracking Robotic Dataset in Crowded Human Environments

CVPR 2024poster

Autonomous robot systems have attracted increasing research attention in recent years where environment understanding is a crucial step for robot navigation human-robot interaction and decision. Real-world robot systems usually collect visual data from multiple sensors and are required to recognize…

Cited by 2SourcePDFScholar
2024

MVSplat360: Feed-Forward 360 Scene Synthesis from Sparse Views

NeurIPS 2024poster

We introduce MVSplat360, a feed-forward approach for 360° novel view synthesis (NVS) of diverse real-world scenes, using only sparse observations. This setting is inherently ill-posed due to minimal overlap among input views and insufficient visual information provided, making it challenging for con…

2024

MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images

ECCV 2024oral

"We introduce , an efficient model that, given sparse multi-view images as input, predicts clean feed-forward 3D Gaussians. To accurately localize the Gaussian centers, we build a cost volume representation via plane sweeping, where the cross-view feature similarities stored in the cost volume can p…

2024

McGrids: Monte Carlo-Driven Adaptive Grids for Iso-Surface Extraction

ECCV 2024poster

"Iso-surface extraction from an implicit field is a fundamental process in various applications of computer vision and graphics. When dealing with geometric shapes with complicated geometric details, many existing algorithms suffer from high computational costs and memory usage. This paper proposes…

Cited by 0SourcePDFScholar
2024

Normal-GS: 3D Gaussian Splatting with Normal-Involved Rendering

NeurIPS 2024poster

Rendering and reconstruction are long-standing topics in computer vision and graphics. Achieving both high rendering quality and accurate geometry is a challenge. Recent advancements in 3D Gaussian Splatting (3DGS) have enabled high-fidelity novel view synthesis at real-time speeds. However, the noi…

Cited by 2SourcePDFScholar
2024

Point-PRC: A Prompt Learning Based Regulation Framework for Generalizable Point Cloud Analysis

NeurIPS 2024poster

This paper investigates the 3D domain generalization (3DDG) ability of large 3D models based on prevalent prompt learning. Recent works demonstrate the performances of 3D point cloud recognition can be boosted remarkably by parameter-efficient prompt tuning. However, we observe that the improvement…

2024

QLLM: Accurate and Efficient Low-Bitwidth Quantization for Large Language Models

ICLR 2024poster

Large Language Models (LLMs) have demonstrated unparalleled efficacy in natural language processing. However, their high computational demands and memory overheads hinder their broad deployment. To address this, two quantization strategies emerge, including Quantization-Aware Training (QAT) and Post…

2024

Sharpness-Aware Data Generation for Zero-shot Quantization

ICML 2024poster

Zero-shot quantization aims to learn a quantized model from a pre-trained full-precision model with no access to original real training data. The common idea in zero-shot quantization approaches is to generate synthetic data for quantizing the full-precision model. While it is well-known that deep n…

Cited by 0SourcePDFScholar
2024

Surface Reconstruction for 3D Gaussian Splatting via Local Structural Hints

ECCV 2024poster

"This paper presents a novel approach for surface mesh reconstruction from 3D Gaussian Splatting (3DGS) [?], a technique renowned for its efficiency in novel view synthesis but challenged for surface reconstruction. The key obstacle is the lack of geometry hints to regulate the optimization of milli…

2024

Taming Stable Diffusion for Text to 360 Panorama Image Generation

CVPR 2024highlight

Generative models e.g. Stable Diffusion have enabled the creation of photorealistic images from text prompts. Yet the generation of 360-degree panorama images from text remains a challenge particularly due to the dearth of paired text-panorama data and the domain gap between panorama and perspective…

2023

Accurate and Real-Time 3D Pedestrian Detection Using an Efficient Attentive Pillar Network

RA-L 2023

Efficiently and accurately detecting people from 3D point cloud data is of great importance in many robotic and autonomous driving applications. This fundamental perception task is still very challenging due to <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/199

Cited by 29SourcecodeScholar
2023

Dynamic Focus-Aware Positional Queries for Semantic Segmentation

CVPR 2023poster

The DETR-like segmentors have underpinned the most recent breakthroughs in semantic segmentation, which end-to-end train a set of queries representing the class prototypes or target segments. Recently, masked attention is proposed to restrict each query to only attend to the foreground regions predi…

2023

JRDB-Pose: A Large-Scale Dataset for Multi-Person Pose Estimation and Tracking

CVPR 2023poster

Autonomous robotic systems operating in human environments must understand their surroundings to make accurate and safe decisions. In crowded human scenes with close-up human-robot interaction and robot navigation, a deep understanding of surrounding people requires reasoning about human motion and…

Cited by 36SourcePDFScholar
2023

Learning Object-Language Alignments for Open-Vocabulary Object Detection

ICLR 2023poster

Existing object detection methods are bounded in a fixed-set vocabulary by costly labeled data. When dealing with novel categories, the model has to be retrained with more bounding box annotations. Natural language supervision is an attractive alternative for its annotation-free attributes and broad…

2023

MARLIN: Masked Autoencoder for Facial Video Representation LearnINg

CVPR 2023poster

This paper proposes a self-supervised approach to learn universal facial representations from videos, that can transfer across a variety of facial analysis tasks such as Facial Attribute Recognition (FAR), Facial Expression Recognition (FER), DeepFake Detection (DFD), and Lip Synchronization (LS). O…

2023

ObjectSDF++: Improved Object-Compositional Neural Implicit Surfaces

ICCV 2023poster

In recent years, neural implicit surface reconstruction has emerged as a popular paradigm for multi-view 3D reconstruction. Unlike traditional multi-view stereo approaches, the neural implicit surface-based methods leverage neural networks to represent 3D scenes as signed distance functions (SDFs).…

Cited by 39PDFcodeScholar
2023

Sensitivity-Aware Visual Parameter-Efficient Fine-Tuning

ICCV 2023oral

Visual Parameter-Efficient Fine-Tuning (PEFT) has become a powerful alternative for full fine-tuning so as to adapt pre-trained vision models to downstream tasks, which only tunes a small number of parameters while freezing the vast majority ones to ease storage burden and optimization difficulty. H…

Cited by 63PDFcodeScholar
2023

Vector Quantized Wasserstein Auto-Encoder

ICML 2023poster

Learning deep discrete latent presentations offers a promise of better symbolic and summarized abstractions that are more useful to subsequent downstream tasks. Inspired by the seminal Vector Quantized Variational Auto-Encoder (VQ-VAE), most of work in learning deep discrete representations has main…

Cited by 18SourcePDFScholar
2022

Bridging Global Context Interactions for High-Fidelity Image Completion

CVPR 2022poster

Bridging global context interactions correctly is important for high-fidelity image completion with large masks. Previous methods attempting this via deep or large receptive field (RF) convolutions cannot escape from the dominance of nearby interactions, which may be inferior. In this paper, we prop…

Cited by 120PDFcodeScholar
2022

Dual Adaptive Transformations for Weakly Supervised Point Cloud Segmentation

ECCV 2022poster

"Weakly supervised point cloud segmentation, i.e. semantically segmenting a point cloud with only a few labeled points in the whole 3D scene, is highly desirable due to the heavy burden of collecting abundant dense annotations for the model training. However, existing methods remain challenging to a…

Cited by 37SourcePDFScholar
2022

EcoFormer: Energy-Saving Attention with Linear Complexity

NeurIPS 2022accept

Transformer is a transformative framework for deep learning which models sequential data and has achieved remarkable performance on a wide range of tasks, but with high computational and energy cost. To improve its efficiency, a popular choice is to compress the models via binarization which constra…

2022

ExtrudeNet: Unsupervised Inverse Sketch-and-Extrude for Shape Parsing

ECCV 2022poster

"Sketch-and-extrude is a common and intuitive modeling process in computer aided design. This paper studies the problem of learning the shape given in the form of point clouds by “inverse” sketch-and-extrude. We present ExtrudeNet, an unsupervised end-to-end network for discovering sketch and extrud…

2022

GMFlow: Learning Optical Flow via Global Matching

CVPR 2022oral

Learning-based optical flow estimation has been dominated with the pipeline of cost volume with convolutions for flow regression, which is inherently limited to local correlations and thus is hard to address the long-standing challenge of large displacements. To alleviate this, the state-of-the-art…

Cited by 459PDFcodeScholar
2022

Less Is More: Pay Less Attention in Vision Transformers

AAAI 2022technical

Transformers have become one of the dominant architectures in deep learning, particularly as a powerful alternative to convolutional neural networks (CNNs) in computer vision. However, Transformer training and inference in previous works can be prohibitively expensive due to the quadratic complexity…

2022

MoVQ: Modulating Quantized Vectors for High-Fidelity Image Generation

NeurIPS 2022accept

Although two-stage Vector Quantized (VQ) generative models allow for synthesizing high-fidelity and high-resolution images, their quantization operator encodes similar patches within an image into the same index, resulting in a repeated artifact for similar adjacent regions using existing decoder ar…

Cited by 86SourcePDFScholar
2022

Multimodal Transformer with Variable-Length Memory for Vision-and-Language Navigation

ECCV 2022poster

"Vision-and-Language Navigation (VLN) is a task that an agent is required to follow a language instruction to navigate to the goal position, which relies on the ongoing interactions with the environment during moving. Recent Transformer-based VLN methods have made great progress benefiting from the…

2022

Object-Compositional Neural Implicit Surfaces

ECCV 2022poster

"The neural implicit representation has shown its effectiveness in novel view synthesis and high-quality 3D reconstruction from multi-view images. However, most approaches focus on holistic scene representation yet ignore individual objects inside it, thus limiting potential downstream applications.…

2022

Particle-based Adversarial Local Distribution Regularization

AISTATS 2022poster

Adversarial training defense (ATD) and virtual adversarial training (VAT) are the two most effective methods to improve model robustness against attacks and model generalization. While ATD is usually applied in robust machine learning, VAT is used in semi-supervised learning and domain adaption. In…

2022

ProposalCLIP: Unsupervised Open-Category Object Proposal Generation via Exploiting CLIP Cues

CVPR 2022poster

Object proposal generation is an important and fundamental task in computer vision. In this paper, we propose ProposalCLIP, a method towards unsupervised open-category object proposal generation. Unlike previous works which require a large number of bounding box annotations and/or can only generate…

Cited by 70PDFScholar
2022

Sem2NeRF: Converting Single-View Semantic Masks to Neural Radiance Fields

ECCV 2022poster

"Image translation and manipulation have gain increasing attention along with the rapid development of deep generative models. Although existing approaches have brought impressive results, they mainly operated in 2D space. In light of recent advances in NeRF-based 3D-aware generative models, we intr…

2021

A Unified 3D Human Motion Synthesis Model via Conditional Variational Auto-Encoder

ICCV 2021poster

We present a unified and flexible framework to address the generalized problem of 3D motion synthesis that covers the tasks of motion prediction, completion, interpolation, and spatial-temporal recovery. Since these tasks have different input constraints and various fidelity and diversity requiremen…

Cited by 80PDFScholar
2021

Auto-Parsing Network for Image Captioning and Visual Question Answering

ICCV 2021poster

We propose an Auto-Parsing Network (APN) to discover and exploit the input data's hidden tree structures for improving the effectiveness of the Transformer-based vision-language systems. Specifically, we impose a Probabilistic Graphical Model (PGM) parameterized by the attention operations on each s…

Cited by 44PDFScholar
2021

CSG-Stump: A Learning Friendly CSG-Like Representation for Interpretable Shape Parsing

ICCV 2021poster

Generating an interpretable and compact representation of 3D shapes from point clouds is an important and challenging problem. This paper presents CSG-Stump Net, an unsupervised end-to-end network for learning shapes from point clouds and discovering the underlying constituent modeling primitives an…

Cited by 50PDFcodeScholar
2021

Domain-Invariant Disentangled Network for Generalizable Object Detection

ICCV 2021poster

We address the problem of domain generalizable object detection, which aims to learn a domain-invariant detector from multiple "seen" domains so that it can generalize well to other "unseen" domains. The generalization ability is crucial in practical scenarios especially when it is difficult to coll…

Cited by 96PDFScholar
2021

High-Resolution Optical Flow From 1D Attention and Correlation

ICCV 2021poster

Optical flow is inherently a 2D search problem, and thus the computational complexity grows quadratically with respect to the search window, making large displacements matching infeasible for high-resolution images. In this paper, we take inspiration from Transformers and propose a new method for hi…

Cited by 98PDFcodeScholar
2021

RSG: A Simple but Effective Module for Learning Imbalanced Datasets

CVPR 2021poster

Imbalanced datasets widely exist in practice and are a great challenge for training deep neural models with a good generalization on infrequent classes. In this work, we propose a new rare-class sample generator (RSG) to solve this problem. RSG aims to generate some new samples for rare classes duri…

Cited by 128PDFcodeScholar
2021

Scalable Vision Transformers With Hierarchical Pooling

ICCV 2021poster

The recently proposed Visual image Transformers (ViT) with pure attention have achieved promising performance on image recognition tasks, such as image classification. However, the routine of the current ViT model is to maintain a full-length patch sequence during inference, which is redundant and l…

Cited by 186PDFcodeScholar
2020

End-to-End 3D Point Cloud Instance Segmentation Without Detection

CVPR 2020poster

3D instance segmentation plays a predominant role in environment perception of robotics and augmented reality. Many deep learning based methods have been presented recently for this task. These methods rely on either a detection branch to propose objects or a grouping step to assemble same-instance…

Cited by 41PDFScholar
2020

Exploring Bottom-Up and Top-Down Cues With Attentive Learning for Webly Supervised Object Detection

CVPR 2020poster

Fully supervised object detection has achieved great success in recent years. However, abundant bounding boxes annotations are needed for training a detector for novel classes. To reduce the human labeling effort, we propose a novel webly supervised object detection (WebSOD) method for novel classes…

Cited by 13PDFScholar
2020

Finding It at Another Side: A Viewpoint-Adapted Matching Encoder for Change Captioning

ECCV 2020poster

Change Captioning is a task that aims to describe the difference between images with natural language. Most existing methods treat this problem as a difference judgment without the existence of distractors such as viewpoint changes. However, in practice, viewpoint changes happen often and can overwh…

Cited by 54SourcePDFScholar
2020

Learning Progressive Joint Propagation for Human Motion Prediction

ECCV 2020poster

Despite the great progress in human motion prediction, it remains a challenging task due to the complicated structural dynamics of human behaviors. In this paper, we address this problem in three aspects. First, to capture the long-range spatial correlations and temporal dependencies, we apply a tra…

Cited by 197SourcePDFScholar
2020

Learning from the Scene and Borrowing from the Rich: Tackling the Long Tail in Scene Graph Generation

IJCAI 2020poster

Despite the huge progress in scene graph generation in recent years, its long-tail distribution in object relationships remains a challenging and pestering issue. Existing methods largely rely on either external knowledge or statistical bias information to alleviate this problem. In this paper, we t…

2020

Self-Supervised Relationship Probing

NeurIPS 2020poster

Structured representations of images that model visual relationships are beneficial for many vision and vision-language applications. However, current human-annotated visual relationship datasets suffer from the long-tailed predicate distribution problem which limits the potential of visual relation…

Cited by 20SourcePDFScholar
2020

Splitting vs. Merging: Mining Object Regions with Discrepancy and Intersection Loss for Weakly Supervised Semantic Segmentation

ECCV 2020poster

In this paper we focus on the task of weakly-supervised semantic segmentation supervised with image-level labels. Since the pixel-level annotation is not available in the training process, we rely on region mining models to estimate the pseudo-masks from the image-level labels. Thus, in order to imp…

Cited by 81SourcePDFScholar
2019

3D Hand Shape and Pose Estimation From a Single RGB Image

CVPR 2019oral

This work addresses a novel and challenging problem of estimating the full 3D hand shape and pose from a single RGB image. Most current methods in 3D hand analysis from monocular RGB images only focus on estimating the 3D locations of hand keypoints, which cannot fully express the 3D shape of hand.…

Cited by 565PDFScholar
2019

Exploiting Spatial-Temporal Relationships for 3D Pose Estimation via Graph Convolutional Networks

ICCV 2019poster

Despite great progress in 3D pose estimation from single-view images or videos, it remains a challenging task due to the substantial depth ambiguity and severe self-occlusions. Motivated by the effectiveness of incorporating spatial dependencies and temporal consistencies to alleviate these issues,…

Cited by 588PDFScholar
2019

Scene Graph Generation With External Knowledge and Image Reconstruction

CVPR 2019poster

Scene graph generation has received growing attention with the advancements in image understanding tasks such as object detection, attributes and relationship prediction, etc. However, existing datasets are biased in terms of object and relationship labels, or often come with noisy and missing annot…

Cited by 386PDFScholar
2019

Unpaired Image Captioning via Scene Graph Alignments

ICCV 2019poster

Most of current image captioning models heavily rely on paired image-caption datasets. However, getting large scale image-caption paired data is labor-intensive and time-consuming. In this paper, we present a scene graph-based approach for unpaired image captioning. Our framework comprises an image…

Cited by 209PDFScholar
2018

Deep Adaptive Attention for Joint Facial Action Unit Detection and Face Alignment

ECCV 2018poster

Facial action unit (AU) detection and face alignment are two highly correlated tasks since facial landmarks can provide precise AU locations to facilitate the extraction of meaningful local features for AU detection. Most existing AU detection works often treat face alignment as a preprocessing and…

Cited by 223SourcePDFScholar
2018

Generalized Robust Bayesian Committee Machine for Large-scale Gaussian Process Regression

ICML 2018oral

In order to scale standard Gaussian process (GP) regression to large-scale datasets, aggregation models employ factorized training process and then combine predictions from distributed experts. The state-of-the-art aggregation models, however, either provide inconsistent predictions or require time-…

2018

Look, Imagine and Match: Improving Textual-Visual Cross-Modal Retrieval With Generative Models

CVPR 2018poster

Textual-visual cross-modal retrieval has been a hot research topic in both computer vision and natural language processing communities. Learning appropriate representations for multi-modal data is crucial for the cross-modal retrieval performance. Unlike existing image-text retrieval approaches that…

Cited by 476SourcePDFScholar
2018

Shuffle-Then-Assemble: Learning Object-Agnostic Visual Relationship Features

ECCV 2018poster

Due to fact that it is prohibitively expensive to completely annotate visual relationships, ie, the (obj1, rel, obj2) triplets, relationship models are inevitably biased to object classes of limited pairwise patterns, leading to poor generalization to rare or unseen object combinations. Therefore, w…

2018

T2Net: Synthetic-to-Realistic Translation for Solving Single-Image Depth Estimation Tasks

ECCV 2018poster

Current methods for single-image depth estimation use training datasets with real image-depth pairs or stereo pairs, which are not easy to acquire. We propose a framework, trained on synthetic image-depth pairs and unpaired real images, that comprises an image translation network for enhancing reali…

2018

VQA-E: Explaining, Elaborating, and Enhancing Your Answers for Visual Questions

ECCV 2018poster

Most existing works in visual question answering (VQA) are dedicated to improving the accuracy of predicted answers, while disregarding the explanations. We argue that the explanation for an answer is of the same or even more importance compared with the answer itself, since it makes the question an…

Cited by 138SourcePDFScholar
2018

Weakly-supervised 3D Hand Pose Estimation from Monocular RGB Images

ECCV 2018poster

Compared with depth-based 3D hand pose estimation, it is more challenging to infer 3D hand pose from monocular RGB images, due to substantial depth ambiguity and the difficulty of obtaining fully-annotated training data. Different from existing learning-based monocular RGB-input approaches that requ…

Cited by 365SourcePDFScholar
2017

A Generative Model for Depth-Based Robust 3D Facial Pose Tracking

CVPR 2017poster

We consider the problem of depth-based robust 3D facial pose tracking under unconstrained scenarios with heavy occlusions and arbitrary facial expression variations. Unlike the previous depth-based discriminative or data-driven methods that require sophisticated training or manual intervention, we p…

Cited by 22PDFScholar
2017

MIML-FCN+: Multi-Instance Multi-Label Learning via Fully Convolutional Networks With Privileged Information

CVPR 2017poster

Multi-instance multi-label (MIML) learning has many interesting applications in computer visions, including multi-object recognition and automatic image tagging. In these applications, additional information such as bounding-boxes, image captions and descriptions is often available during training p…

Cited by 87PDFScholar
2016

Exploit Bounding Box Annotations for Multi-Label Object Recognition

CVPR 2016poster

Convolutional neural networks (CNNs) have shown great performance as general feature representations for object recognition applications. However, for multi-label images that contain multiple objects from different categories, scales and locations, global CNN features are not optimal. In this paper,…

Cited by 210PDFScholar
2016

Modality and Component Aware Feature Fusion For RGB-D Scene Classification

CVPR 2016accepted

While convolutional neural networks (CNN) have been excellent for object recognition, the greater spatial variability in scene images typically meant that the standard full-image CNN features are suboptimal for scene classification. In this paper, we investigate a framework allowing greater spatial…

Cited by 85SourcePDFScholar
2015

MMSS: Multi-Modal Sharable and Specific Feature Learning for RGB-D Object Recognition

ICCV 2015poster

Most of the feature-learning methods for RGB-D object recognition either learn features from color and depth modalities separately, or simply treat RGB-D as undifferentiated four-channel data, which cannot adequately exploit the relationship between different modalities. Motivated by the intuition t…

Cited by 118PDFScholar