← Search

Jianbing Shen

90 accepted papers

2026

DriveVLN: Towards Mapless Vision-and-Language Navigation in Autonomous Driving

CVPR 2026

Autonomous driving has made substantial progress recently, achieving reliable performance in most real-world environments. However, existing algorithms still depend heavily on high-definition maps, making them ineffective in mapless scenarios such as indoor parking lots. These limitations hinder sea

Cited by 0SourceScholar
2026

From Broad Exploration to Stable Synthesis: Entropy-Guided Optimization for Autoregressive Image Generation

ICLR 2026poster

Combining Chain-of-Thought (CoT) with Reinforcement Learning (RL) improves text-to-image (T2I) generation, yet the underlying interaction between CoT's exploration and RL's optimization remains unclear. We present a systematic entropy-based analysis that yields three key insights: (1) CoT expands th…

Cited by 0SourcecodeScholar
2026

HiCoGen: Hierarchical Compositional Text-to-Image Generation in Diffusion Models via Reinforcement Learning

CVPR 2026

Recent advances in diffusion models have demonstrated impressive capability in generating high-quality images for simple prompts. However, when confronted with complex prompts involving multiple objects and hierarchical structures, existing models struggle to accurately follow instructions, leading

Cited by 0SourcecodeScholar
2026

Less Is More: Vision Representation Compression for Efficient Video Generation with Large Language Models

AAAI 2026technical

Video generation using Large Language Models (LLMs) has shown promising potential, effectively leveraging the extensive LLM infrastructure to provide a unified framework for multimodal understanding and content generation. However, these methods face critical challenges, i.e., token redundancy and i

Cited by 0SourcePDFScholar
2026

Sim4Seg: Boosting Multimodal Multi-disease Medical Diagnosis Segmentation with Region-Aware Vision-Language Similarity Masks

AAAI 2026technical

Despite significant progress in pixel-level medical image analysis, existing medical image segmentation models rarely explore medical segmentation and diagnosis tasks jointly. However, it is crucial for patients that models can provide explainable diagnoses along with medical segmentation results. I

Cited by 0SourcePDFScholar
2026

Towards High-Fidelity 3D Portrait Generation with Rich Details by Cross-View Prior-Aware Diffusion

AAAI 2026technical

Recent diffusion-based Single-image 3D portrait generation methods typically employ 2D diffusion models to provide multi-view knowledge, which is then distilled into 3D representations. However, these methods usually struggle to produce high-fidelity 3D models, frequently yielding excessively blurre

Cited by 0SourcePDFScholar
2025

ALOcc: Adaptive Lifting-Based 3D Semantic Occupancy and Cost Volume-Based Flow Predictions

ICCV 2025poster

3D semantic occupancy and flow prediction are fundamental to spatiotemporal scene understanding. This paper proposes a vision-based framework with three targeted improvements. First, we introduce an occlusion-aware adaptive lifting mechanism incorporating depth denoising. This enhances the robustnes…

2025

DC-ControlNet: Decoupling Inter- and Intra-Element Conditions in Image Generation with Diffusion Models

ICCV 2025poster

In this paper, we introduce DC (Decouple)-ControlNet, a highly flexible and precisely controllable framework for multi-condition image generation. The core idea behind DC-ControlNet is to decouple control conditions, transforming global control into a hierarchical system that integrates distinct ele…

Cited by 0SourcePDFScholar
2025

DME-Driver: Integrating Human Decision Logic and 3D Scene Perception in Autonomous Driving

AAAI 2025technical

There are two crucial aspects of reliable autonomous driving systems: the reasoning behind decision-making and the precision of environmental perception. This paper introduces DME-Driver, a new autonomous driving system that enhances performance and robustness by fully leveraging the two crucial asp…

Cited by 25SourcePDFScholar
2025

Decoupling Fine Detail and Global Geometry for Compressed Depth Map Super-Resolution

CVPR 2025poster

Recovering high-quality depth maps from compressed sources has gained significant attention due to the limitations of consumer-grade depth cameras and the bandwidth restrictions during data transmission. However, current methods still suffer from two challenges. First, bit-depth compression produces…

2025

DrivingSphere: Building a High-fidelity 4D World for Closed-loop Simulation

CVPR 2025poster

Autonomous driving evaluation requires simulation environments that closely replicate actual road conditions, including real-world sensory data and responsive feedback loops. However, many existing simulations need to predict waypoints along fixed routes on public datasets or synthetic photorealisti…

2025

Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback

ACL 2025long

Existing Medical Large Vision-Language Models (Med-LVLMs), encapsulating extensive medical knowledge, demonstrate excellent capabilities in understanding medical images. However, there remain challenges in visual localization in medical images, which is crucial for abnormality detection and interpre…

Cited by 0SourcePDFScholar
2025

LOGICZSL: Exploring Logic-induced Representation for Compositional Zero-shot Learning

CVPR 2025poster

Compositional zero-shot learning (CZSL) aims to recognize unseen attribute-object compositions by learning the primitive concepts (*i.e.*, attribute and object) from the training set. While recent works achieve impressive results in CZSL by leveraging large vision-language models like CLIP, they ign…

2025

Language Prompt for Autonomous Driving

AAAI 2025technical

A new trend in the computer vision community is to capture objects of interest following flexible human command represented by a natural language prompt. However, the progress of using language prompts in driving scenarios is stuck in a bottleneck due to the scarcity of paired prompt-instance data.…

2025

MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration

ACL 2025finding

Recent advancements in medical Large Language Models (LLMs) have showcased their powerful reasoning and diagnostic capabilities. Despite their success, current unified multimodal medical LLMs face limitations in knowledge update costs, comprehensiveness, and flexibility. To address these challenges,…

2025

OLiDM: Object-aware LiDAR Diffusion Models for Autonomous Driving

AAAI 2025technical

To enhance autonomous driving, innovative approaches have been proposed to generate simulated LiDAR data. However, these methods often face challenges in producing high-quality and controllable foreground objects. To cater to the needs of object-aware tasks in 3D perception, we introduce OLiDM, a no…

Cited by 1SourcePDFScholar
2025

RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping

ICCV 2025poster

General robotic grasping systems require accurate object affordance perception in diverse open-world scenarios following human instructions. However, current studies suffer from the problem of lacking reasoning-based large-scale affordance prediction data, leading to considerable concern about open-…

Cited by 0SourcePDFScholar
2025

RLGF: Reinforcement Learning with Geometric Feedback for Autonomous Driving Video Generation

NeurIPS 2025poster

Synthetic data is crucial for advancing autonomous driving (AD) systems, yet current state-of-the-art video generation models, despite their visual realism, suffer from subtle geometric distortions that limit their utility for downstream perception tasks. We identify and quantify this critical issu…

Cited by 0SourceScholar
2025

Rethinking Temporal Fusion with a Unified Gradient Descent View for 3D Semantic Occupancy Prediction

CVPR 2025poster

We present GDFusion, a temporal fusion method for vision-based 3D semantic occupancy prediction (VisionOcc). GDFusion opens up the underexplored aspects of temporal fusion within the VisionOcc framework, with a focus on both temporal cues and fusion strategies. It systematically examines the entire…

2025

Self-Rewarding Large Vision-Language Models for Optimizing Prompts in Text-to-Image Generation

ACL 2025finding

Text-to-image models are powerful for producing high-quality images based on given text prompts, but crafting these prompts often requires specialized vocabulary. To address this, existing methods train rewriting models with supervision from large amounts of manually annotated data and trained aesth…

Cited by 0SourcePDFScholar
2025

Semantic Causality-Aware Vision-Based 3D Occupancy Prediction

ICCV 2025poster

Vision-based 3D semantic occupancy prediction is a critical task in 3D vision that integrates volumetric 3D reconstruction with semantic understanding. Existing methods, however, often rely on modular pipelines. These modules are typically optimized independently or use pre-configured inputs, leadin…

2025

Weak to Strong Generalization for Large Language Models with Multi-capabilities

ICLR 2025poster

As large language models (LLMs) grow in sophistication, some of their capabilities surpass human abilities, making it essential to ensure their alignment with human values and intentions, i.e., Superalignment. This superalignment challenge is particularly critical for complex tasks, as annotations p…

Cited by 88SourcePDFScholar
2024

DI-V2X: Learning Domain-Invariant Representation for Vehicle-Infrastructure Collaborative 3D Object Detection

AAAI 2024technical

Vehicle-to-Everything (V2X) collaborative perception has recently gained significant attention due to its capability to enhance scene understanding by integrating information from various agents, e.g., vehicles, and infrastructure. However, current works often treat the information from each agent e…

2024

Fine-Grained Distillation for Long Document Retrieval

AAAI 2024technical

Long document retrieval aims to fetch query-relevant documents from a large-scale collection, where knowledge distillation has become de facto to improve a retriever by mimicking a heterogeneous yet powerful cross-encoder. However, in contrast to passages or sentences, retrieval on long documents su…

Cited by 52SourcePDFScholar
2024

IS-Fusion: Instance-Scene Collaborative Fusion for Multimodal 3D Object Detection

CVPR 2024highlight

Bird's eye view (BEV) representation has emerged as a dominant solution for describing 3D space in autonomous driving scenarios. However objects in the BEV representation typically exhibit small sizes and the associated point cloud context is inherently sparse which leads to great challenges for rel…

2024

Leveraging Frame Affinity for sRGB-to-RAW Video De-rendering

CVPR 2024poster

Unprocessed RAW video has shown distinct advantages over sRGB video in video editing and computer vision tasks. However capturing RAW video is challenging due to limitations in bandwidth and storage. Various methods have been proposed to address similar issues in single image RAW capture through de-…

Cited by 2SourcePDFScholar
2024

RepVF: A Unified Vector Fields Representation for Multi-task 3D Perception

ECCV 2024poster

"Concurrent processing of multiple autonomous driving 3D perception tasks within the same spatiotemporal scene poses a significant challenge, in particular due to the computational inefficiencies and feature competition between tasks when using traditional multi-task learning approaches. This paper…

2024

TopoMLP: A Simple yet Strong Pipeline for Driving Topology Reasoning

ICLR 2024poster

Topology reasoning aims to comprehensively understand road scenes and present drivable routes in autonomous driving. It requires detecting road centerlines (lane) and traffic elements, further reasoning their topology relationship, \textit{i.e.}, lane-lane topology, and lane-traffic topology. In thi…

2024

Visual In-Context Learning for Large Vision-Language Models

ACL 2024findings

In Large Visual Language Models (LVLMs), the efficacy of In-Context Learning (ICL) remains limited by challenges in cross-modal interactions and representation disparities. To overcome these challenges, we introduce a novel Visual In-Context Learning (VICL) method comprising Visual Demonstration Ret…

Cited by 106SourcePDFScholar
2023

Exposing the Self-Supervised Space-Time Correspondence Learning via Graph Kernels

AAAI 2023technical

Self-supervised space-time correspondence learning is emerging as a promising way of leveraging unlabeled video. Currently, most methods adapt contrastive learning with mining negative samples or reconstruction adapted from the image domain, which requires dense affinity across multiple frames or op…

2023

LWSIS: LiDAR-Guided Weakly Supervised Instance Segmentation for Autonomous Driving

AAAI 2023technical

Image instance segmentation is a fundamental research topic in autonomous driving, which is crucial for scene understanding and road safety. Advanced learning-based approaches often rely on the costly 2D mask annotations for training. In this paper, we present a more artful framework, LiDAR-guided…

2023

OnlineRefer: A Simple Online Baseline for Referring Video Object Segmentation

ICCV 2023poster

Referring video object segmentation (RVOS) aims at segmenting an object in a video following human instruction. Current state-of-the-art methods fall into an offline pattern, in which each clip independently interacts with text embedding for cross-modal understanding. They usually present that the o…

Cited by 55PDFcodeScholar
2023

Referring Multi-Object Tracking

CVPR 2023poster

Existing referring understanding tasks tend to involve the detection of a single text-referred object. In this paper, we propose a new and general referring understanding task, termed referring multi-object tracking (RMOT). Its core idea is to employ a language expression as a semantic cue to guide…

2023

SSDA3D: Semi-supervised Domain Adaptation for 3D Object Detection from Point Cloud

AAAI 2023technical

LiDAR-based 3D object detection is an indispensable task in advanced autonomous driving systems. Though impressive detection results have been achieved by superior 3D detectors, they suffer from significant performance degeneration when facing unseen domains, such as different LiDAR configurations,…

2023

Self-Supervised Monocular Depth Estimation by Direction-aware Cumulative Convolution Network

ICCV 2023poster

Monocular depth estimation is known as an ill-posed task that objects in a 2D image usually do not contain sufficient information to predict their depth. Thus, it acts differently from other tasks (e.g., classification and segmentation) in many ways. In this paper, we find that self-supervised monoc…

Cited by 25PDFcodeScholar
2023

Weakly Supervised Monocular 3D Object Detection Using Multi-View Projection and Direction Consistency

CVPR 2023poster

Monocular 3D object detection has become a mainstream approach in automatic driving for its easy application. A prominent advantage is that it does not need LiDAR point clouds during the inference. However, most current methods still rely on 3D point cloud data for labeling the ground truths used in…

2022

BRNet: Exploring Comprehensive Features for Monocular Depth Estimation

ECCV 2022poster

"Self-supervised monocular depth estimation has achieved promising performance recently. A consensus is that high-resolution inputs often yield better results. However, we find that the performance gap between high and low resolutions lies in the inappropriate feature representation of the widely us…

2022

Counterfactual Cycle-Consistent Learning for Instruction Following and Generation in Vision-Language Navigation

CVPR 2022poster

Since the rise of vision-language navigation (VLN), great progress has been made in instruction following -- building a follower to navigate environments under the guidance of instructions. However, far less attention has been paid to the inverse task: instruction generation -- learning a speaker to…

Cited by 62PDFcodeScholar
2022

Learning Disentanglement with Decoupled Labels for Vision-Language Navigation

ECCV 2022poster

"Vision-and-Language Navigation (VLN) requires an agent to follow complex natural language instructions and perceive the visual environment for real-world navigation. Intuitively, we find that instruction disentanglement for each viewpoint along the agent’s path is critical for accurate navigation.…

2022

Modality Synergy Complement Learning with Cascaded Aggregation for Visible-Infrared Person Re-identification

ECCV 2022poster

"Visible-Infrared Re-Identification (VI-ReID) is challenging in image retrievals. The modality discrepancy will easily make huge intra-class variations. Most existing methods either bridge different modalities through modality-invariance or generate the intermediate modality for better performance.…

2022

Multi-Level Representation Learning With Semantic Alignment for Referring Video Object Segmentation

CVPR 2022poster

Referring video object segmentation (RVOS) is a challenging language-guided video grounding task, which requires comprehensively understanding the semantic information of both video content and language queries for object prediction. However, existing methods adopt multi-modal fusion at a frame-base…

Cited by 64PDFScholar
2022

ProposalContrast: Unsupervised Pre-training for LiDAR-Based 3D Object Detection

ECCV 2022poster

"Existing approaches for unsupervised point cloud pre-training are constrained to either scene-level or point/voxel-level instance discrimination. Scene-level methods tend to lose local details that are crucial for recognizing the road objects, while point/voxel-level methods inherently suffer from…

2022

Rethinking Clustering-Based Pseudo-Labeling for Unsupervised Meta-Learning

ECCV 2022poster

"The pioneering unsupervised meta-learning work is a clustering-based pseudo-labeling method, which is model-agnostic and can utilize supervised algorithms for learning from unlabeled data. However, it often suffers from label inconsistency and limited diversity, which leads to poor performance. In…

2022

Semi-Supervised 3D Object Detection with Proficient Teachers

ECCV 2022poster

"Dominated point cloud-based 3D object detectors in autonomous driving scenarios rely heavily on the huge amount of accurately labeled samples, however, 3D annotation in the point cloud is extremely tedious, expensive and time-consuming. To reduce the dependence on large supervision, semi-supervised…

2022

Tree Energy Loss: Towards Sparsely Annotated Semantic Segmentation

CVPR 2022poster

Sparsely annotated semantic segmentation (SASS) aims to train a segmentation network with coarse-grained (i.e.,point-, scribble-, and block-wise) supervisions, where only a small proportion of pixels are labeled in each image. In this paper, we propose a novel tree energy loss for SASS by providing…

Cited by 80PDFcodeScholar
2021

Cross-Modality Person Re-Identification via Modality Confusion and Center Aggregation

ICCV 2021poster

Cross-modality person re-identification is a challenging task due to large cross-modality discrepancy and intra-modality variations. Currently, most existing methods focus on learning modality-specific or modality-shareable features by using the identity supervision or modality label. Different from…

Cited by 213PDFScholar
2021

Full-Duplex Strategy for Video Object Segmentation

ICCV 2021poster

Appearance and motion are two important sources of information in video object segmentation (VOS). Previous methods mainly focus on using simplex solutions, lowering the upper bound of feature collaboration among and across these two cues. In this paper, we study a novel framework, termed the FSNet…

Cited by 180PDFcodeScholar
2021

Learning To Fuse Asymmetric Feature Maps in Siamese Trackers

CVPR 2021poster

Recently, Siamese-based trackers have achieved promising performance in visual tracking. Most recent Siamese-based trackers typically employ a depth-wise cross-correlation (DW-XCorr) to obtain multi-channel correlation information from the two feature maps (target and search region). However, DW-XCo…

Cited by 97PDFcodeScholar
2021

Structured Scene Memory for Vision-Language Navigation

CVPR 2021poster

Recently, numerous algorithms have been developed to tackle the problem of vision-language navigation (VLN), i.e., entailing an agent to navigate 3D environments through following linguistic instructions. However, current VLN agents simply store their past experiences/observations as latent states i…

Cited by 136PDFcodeScholar
2020

A Unified Object Motion and Affinity Model for Online Multi-Object Tracking

CVPR 2020poster

Current popular online multi-object tracking (MOT) solutions apply single object trackers (SOTs) to capture object motions, while often requiring an extra affinity network to associate objects, especially for the occluded ones. This brings extra computational overhead due to repetitive feature extra…

Cited by 139PDFcodeScholar
2020

Active Visual Information Gathering for Vision-Language Navigation

ECCV 2020poster

Vision-language navigation (VLN) is the task of entailing an agent to carry out navigational instructions inside photo-realistic environments. One of the key challenges in VLN is how to conduct a robust navigation by mitigating the uncertainty caused by ambiguous instructions and insufficient observ…

2020

CLNet: A Compact Latent Network for Fast Adjusting Siamese Trackers

ECCV 2020poster

In this paper, we provide a deep analysis for Siamese-based trackers and find that the one core reason for their failure on challenging cases can be attributed to the problem of {\it decisive samples missing} during offline training. Furthermore, we notice that the samples given in the first frame c…

2020

Dynamic Dual-Attentive Aggregation Learning for Visible-Infrared Person Re-Identification

ECCV 2020poster

Visible-infrared person re-identification (VI-ReID) is a challenging cross-modality pedestrian retrieval problem. Due to the large intra-class variations and cross-modality discrepancy with large amount of sample noise, it is difficult to learn discriminative part features. Existing VI-ReID methods…

2020

Hierarchical Human Parsing With Typed Part-Relation Reasoning

CVPR 2020poster

Human parsing is for pixel-wise human semantic understanding. As human bodies are underlying hierarchically structured, how to model human structures is the central theme in this task. Focusing on this, we seek to simultaneously exploit the representational capacity of deep graph networks and the hi…

Cited by 136PDFcodeScholar
2020

Learning Video Object Segmentation From Unlabeled Videos

CVPR 2020poster

We propose a new method for video object segmentation (VOS) that addresses object pattern learning from unlabeled videos, unlike most existing methods which rely heavily on extensive annotated data. We introduce a unified unsupervised/weakly supervised learning framework, called MuG, that comprehens…

Cited by 192PDFcodeScholar
2020

LiDAR-Based Online 3D Video Object Detection With Graph-Based Message Passing and Spatiotemporal Transformer Attention

CVPR 2020poster

Existing LiDAR-based 3D object detectors usually focus on the single-frame detection, while ignoring the spatiotemporal information in consecutive point cloud frames. In this paper, we propose an end-to-end online 3D video object detector that operates on point cloud sequences. The proposed model co…

Cited by 184PDFcodeScholar
2020

Multi-Mutual Consistency Induced Transfer Subspace Learning for Human Motion Segmentation

CVPR 2020poster

Human motion segmentation based on transfer subspace learning is a rising interest in action-related tasks. Although progress has been made, there are still several issues within the existing methods. First, existing methods transfer knowledge from source data to target tasks by learning domain-inva…

Cited by 43PDFScholar
2020

NETNet: Neighbor Erasing and Transferring Network for Better Single Shot Object Detection

CVPR 2020poster

Due to the advantages of real-time detection and improved performance, single-shot detectors have gained great attention recently. To solve the complex scale variations, single-shot detectors make scale-aware predictions based on multiple pyramid layers. However, the features in the pyramid are not…

Cited by 46PDFScholar
2020

Self-Learning With Rectification Strategy for Human Parsing

CVPR 2020poster

In this paper, we solve the sample shortage problem in the human parsing task. We begin with the self-learning strategy, which generates pseudo-labels for unlabeled data to retrain the model. However, directly using noisy pseudo-labels will cause error amplification and accumulation. Considering the…

Cited by 45PDFScholar
2020

Video Object Segmentation with Episodic Graph Memory Networks

ECCV 2020poster

How to make a segmentation model efficiently adapt to a specific video as well as online target appearance variations is a fun- damental issue in the field of video object segmentation. In this work, a graph memory network is developed to address the novel idea of “learning to update the segmentatio…

2020

Weakly Supervised 3D Object Detection from Lidar Point Cloud

ECCV 2020poster

It is laborious to manually label point cloud data for training high-quality 3D object detectors. This work proposes a weakly supervised approach for 3D object detection, only requiring a small set of weakly annotated scenes, associated with a few precisely labeled object instances. This is achieved…

2019

Adversarial Defense by Restricting the Hidden Space of Deep Neural Networks

ICCV 2019poster

Deep neural networks are vulnerable to adversarial attacks which can fool them by adding minuscule perturbations to the input images. The robustness of existing defenses suffers greatly under white-box attack settings, where an adversary has full knowledge about the network and can iterate several t…

Cited by 189PDFcodeScholar
2019

An Iterative and Cooperative Top-Down and Bottom-Up Inference Network for Salient Object Detection

CVPR 2019poster

This paper presents a salient object detection method that integrates both top-down and bottom-up saliency inference in an iterative and cooperative manner. The top-down process is used for coarse-to-fine saliency estimation, where high-level saliency is gradually integrated with finer lower-layer f…

Cited by 259PDFScholar
2019

Gaussian Affinity for Max-Margin Class Imbalanced Learning

ICCV 2019poster

Real-world object classes appear in imbalanced ratios. This poses a significant challenge for classifiers which get biased towards frequent classes. We hypothesize that improving the generalization capability of a classifier should improve learning on imbalanced datasets. Here, we introduce the firs…

Cited by 90PDFScholar
2019

Learning Compositional Neural Information Fusion for Human Parsing

ICCV 2019poster

This work proposes to combine neural networks with the compositional hierarchy of human bodies for efficient and complete human parsing. We formulate the approach as a neural information fusion framework. Our model assembles the information from three inference processes over the hierarchy: direct i…

Cited by 160PDFcodeScholar
2019

Learning Unsupervised Video Object Segmentation Through Visual Attention

CVPR 2019poster

This paper conducts a systematic study on the role of visual attention in Unsupervised Video Object Segmentation (UVOS) tasks. By elaborately annotating three popular video segmentation datasets (DAVIS, Youtube-Objects and SegTrack V2) with dynamic eye-tracking data in the UVOS setting, for the firs…

Cited by 277PDFcodeScholar
2019

Salient Object Detection With Pyramid Attention and Salient Edges

CVPR 2019poster

This paper presents a new method for detecting salient objects in images using convolutional neural networks (CNNs). The proposed network, named PAGE-Net, offers two key contributions. The first is the exploitation of an essential pyramid attention structure for salient object detection. This enable…

Cited by 602PDFScholar
2019

See More, Know More: Unsupervised Video Object Segmentation With Co-Attention Siamese Networks

CVPR 2019poster

We introduce a novel network, called as CO-attention Siamese Network (COSNet), to address the unsupervised video object segmentation task from a holistic view. We emphasize the importance of inherent correlation among video frames and incorporate a global co-attention mechanism to improve further th…

Cited by 598PDFcodeScholar
2019

Shifting More Attention to Video Salient Object Detection

CVPR 2019oral

The last decade has witnessed a growing interest in video salient object detection (VSOD). However, the research community long-term lacked a well-established VSOD dataset representative of real dynamic scenes with high-quality annotations. To address this issue, we elaborately collected a visual-at…

Cited by 561PDFcodeScholar
2019

Zero-Shot Video Object Segmentation via Attentive Graph Neural Networks

ICCV 2019oral

This work proposes a novel attentive graph neural network (AGNN) for zero-shot video object segmentation (ZVOS). The suggested AGNN recasts this task as a process of iterative information fusion over video graphs. Specifically, AGNN builds a fully connected graph to efficiently represent frames as n…

Cited by 353PDFcodeScholar
2018

Attentive Fashion Grammar Network for Fashion Landmark Detection and Clothing Category Classification

CVPR 2018poster

This paper proposes a knowledge-guided fashion network to solve the problem of visual fashion analysis, e.g., fashion landmark localization and clothing category classification. The suggested fashion model is leveraged with high-level human knowledge in this domain. We propose two important fashion…

Cited by 307SourcePDFScholar
2018

Hyperparameter Optimization for Tracking With Continuous Deep Q-Learning

CVPR 2018poster

Hyperparameters are numerical presets whose values are assigned prior to the commencement of the learning process. Selecting appropriate hyperparameters is critical for the accuracy of tracking algorithms, yet it is difficult to determine their optimal values, in particular, adaptive ones for each s…

Cited by 198SourcePDFScholar
2018

Learning Human-Object Interactions by Graph Parsing Neural Networks

ECCV 2018poster

This paper addresses the task of detecting and recognizing human-object interactions (HOI) in images and videos. We introduce the Graph Parsing Neural Network (GPNN), a framework that incorporates structural knowledge while being differentiable end-to-end. For a given scene, GPNN infers a parse grap…

2018

Pyramid Dilated Deeper ConvLSTM for Video Salient Object Detection

ECCV 2018poster

This paper proposes a fast video salient object detection model, based on a novel recurrent network architecture, named Pyramid Dilated Bidirectional ConvLSTM (PDB-ConvLSTM). A Pyramid Dilated Convolution (PDC) module is first designed for simultaneously extracting spatial features at multiple scale…

Cited by 593SourcePDFScholar
2018

Revisiting Video Saliency: A Large-Scale Benchmark and a New Model

CVPR 2018poster

In this work, we contribute to video saliency research in two ways. First, we introduce a new benchmark for predicting human eye movements during dynamic scene free-viewing, which is long-time urged in this field. Our dataset, named DHF1K~(Dynamic Human Fixation), consists of 1K high-quality, elabor…

2018

Salient Object Detection Driven by Fixation Prediction

CVPR 2018poster

Research in visual saliency has been focused on two major types of models namely fixation prediction and salient object detection. The relationship between the two, however, has been less explored. In this paper, we propose to employ the former model type to identify and segment salient objects in s…

2015

Linearization to Nonlinear Learning for Visual Tracking

ICCV 2015poster

Due to unavoidable appearance variations caused by occlusion, deformation, and other factors, classifiers for visual tracking are nonlinear as a necessity. Building on the theory of globally linear approximations to nonlinear functions, we introduce an elegant method that jointly learns a nonlinear…

Cited by 41PDFcodeScholar