← Search

Shengfeng He

80 accepted papers

2026

Cross-modal Proxy Evolving for OOD Detection with Vision-Language Models

AAAI 2026technical

Reliable zero-shot detection of out-of-distribution (OOD) inputs is critical for deploying vision-language models in open-world settings. However, the lack of labeled negatives in zero-shot OOD detection necessitates proxy signals that remain effective under distribution shift. Existing negative-lab

Cited by 0SourcePDFScholar
2026

FlashMesh: Faster and Better Autoregressive Mesh Synthesis via Structured Speculation

CVPR 2026

Autoregressive models can generate high-quality 3D meshes by sequentially producing vertices and faces, but their token-by-token decoding results in slow inference, limiting practical use in interactive and large-scale applications.We present FlashMesh, a fast and high-fidelity mesh generation frame

Cited by 0SourcecodeScholar
2026

Healthcare Insurance Fraud Detection via Continual Fiedler Vector Graph Model

ICLR 2026poster

Healthcare insurance fraud detection presents unique machine learning challenges: labeled data are scarce due to delayed verification processes, and fraudulent behaviors evolve rapidly, often manifesting in complex, graph-structured interactions. Existing methods struggle in such settings. Pretraini…

Cited by 0SourceScholar
2026

M$^3$E: Continual Vision-and-Language Navigation via Mixture of Macro and Micro Experts

ICLR 2026poster

Vision-and-Language Navigation (VLN) agents have shown strong capabilities in following natural language instructions. However, they often struggle to generalize across environments due to catastrophic forgetting, which limits their practical use in real-world settings where agents must continually…

Cited by 0SourceScholar
2026

NimbusGS: Unified 3D Scene Reconstruction under Hybrid Weather

CVPR 2026

We present NimbusGS, a unified framework for reconstructing high-quality 3D scenes from degraded multi-view inputs captured under diverse and mixed adverse weather conditions. Unlike existing methods that target specific weather types, NimbusGS addresses the broader challenge of generalization by mo

Cited by 0SourcecodeScholar
2026

PolarDepth: Monocular Transparent Object Depth from Polar-Physics Priors

ICML 2026poster

Depth estimation for transparent objects remains a fundamental challenge, as RGB-based cues often fail in regions affected by refraction and light transmission. Polarization provides physically grounded information related to surface orientation and material properties, offering reliable geometric c…

Cited by 0SourceScholar
2026

ProstaTD: Bridging Surgical Triplet from Classification to Fully Supervised Detection

ICLR 2026poster

Surgical triplet detection is a critical task in surgical video analysis, with significant implications for performance assessment and training novice surgeons. However, existing datasets like CholecT50 lack precise spatial bounding box annotations, rendering triplet classification at the image leve…

Cited by 0SourceScholar
2026

Self-Corrected Image Generation with Explainable Latent Rewards

CVPR 2026

Despite significant progress in text-to-image generation, aligning outputs with complex prompts remains challenging, particularly for fine-grained semantics and spatial relations. This difficulty stems from the feed-forward nature of generation, which requires anticipating alignment without fully un

Cited by 0SourceScholar
2025

Action Dubber: Timing Audible Actions via Inflectional Flow

ICML 2025poster

We introduce the task of Audible Action Temporal Localization, which aims to identify the spatio-temporal coordinates of audible movements. Unlike conventional tasks such as action recognition and temporal action localization, which broadly analyze video content, our task focuses on the distinct kin…

2025

Adversarial Attacks on Event-Based Pedestrian Detectors: A Physical Approach

AAAI 2025technical

Event cameras, known for their low latency and high dynamic range, show great potential in pedestrian detection applications. However, while recent research has primarily focused on improving detection accuracy, the robustness of event-based visual models against physical adversarial attacks has rec…

Cited by 0SourcePDFScholar
2025

Connecting Giants: Synergistic Knowledge Transfer of Large Multimodal Models for Few-Shot Learning

IJCAI 2025

Few-shot learning (FSL) addresses the challenge of classifying novel classes with limited training samples. While some methods leverage semantic knowledge from smaller-scale models to mitigate data scarcity, these approaches often introduce noise and bias due to the data’s inherent simplicity. In th

Cited by 0SourcePDFScholar
2025

Cross-Subject Mind Decoding from Inaccurate Representations

ICCV 2025poster

Decoding stimulus images from fMRI signals has advanced with pre-trained generative models. However, existing methods struggle with cross-subject mappings due to cognitive variability and subject-specific differences. This challenge arises from sequential errors, where unidirectional mappings genera…

Cited by 0SourcePDFScholar
2025

Dr. Tongue: Sign-Oriented Multi-label Detection for Remote Tongue Diagnosis

AAAI 2025technical

Tongue diagnosis is a vital tool in both Western and Traditional Chinese Medicine, providing key insights into a patient's health by analyzing tongue attributes. The COVID-19 pandemic has heightened the need for accurate remote medical assessments, emphasizing the importance of precise tongue attrib…

2025

Instance-Level Video Depth in Groups Beyond Occlusions

ICCV 2025poster

Depth estimation in dynamic, multi-object scenes remains a major challenge, especially under severe occlusions. Existing monocular models, including foundation models, struggle with instance-wise depth consistency due to their reliance on global regression. We tackle this problem from two key aspect…

Cited by 0SourcePDFScholar
2025

Instruct2See: Learning to Remove Any Obstructions Across Distributions

ICML 2025poster

Images are often obstructed by various obstacles due to capture limitations, hindering the observation of objects of interest. Most existing methods address occlusions from specific elements like fences or raindrops, but are constrained by the wide range of real-world obstructions, making comprehens…

2025

Jury-and-Judge Chain-of-Thought for Uncovering Toxic Data in 3D Visual Grounding

NeurIPS 2025poster

3D Visual Grounding (3DVG) faces persistent challenges due to coarse scene-level observations and logically inconsistent annotations, which introduce ambiguities that compromise data quality and hinder effective model supervision. To address these challenges, we introduce Refer-Judge, a novel framew…

Cited by 0SourcecodeScholar
2025

Keep the Balance: A Parameter-Efficient Symmetrical Framework for RGB+X Semantic Segmentation

CVPR 2025poster

Multimodal semantic segmentation is a critical challenge in computer vision, with early methods suffering from high computational costs and limited transferability due to full fine-tuning of RGB-based pre-trained parameters. Recent studies, while leveraging additional modalities as supplementary pro…

Cited by 0SourcePDFScholar
2025

Knowledge Bridger: Towards Training-Free Missing Modality Completion

CVPR 2025poster

Previous successful approaches to missing modality completion rely on carefully designed fusion techniques and extensive pre-training on complete data, which can limit their generalizability in out-of-domain (OOD) scenarios. In this study, we pose a new challenge: can we develop a missing modality c…

2025

MODfinity: Unsupervised Domain Adaptation with Multimodal Information Flow Intertwining

CVPR 2025poster

Multimodal unsupervised domain adaptation leverages unlabeled data in the target domain to enhance multimodal systems continuously. While current state-of-the-art methods encourage interaction between sub-models of different modalities through pseudo-labeling and feature-level exchange, varying samp…

Cited by 2SourcePDFScholar
2025

Neptune-X: Active X-to-Maritime Generation for Universal Maritime Object Detection

NeurIPS 2025spotlight

Maritime object detection is essential for navigation safety, surveillance, and autonomous operations, yet constrained by two key challenges: the scarcity of annotated maritime data and poor generalization across various maritime attributes (e.g., object category, viewpoint, location, and imaging en…

Cited by 0SourcecodeScholar
2025

NexusGS: Sparse View Synthesis with Epipolar Depth Priors in 3D Gaussian Splatting

CVPR 2025highlight

Neural Radiance Field (NeRF) and 3D Gaussian Splatting (3DGS) have noticeably advanced photo-realistic novel view synthesis using images from densely spaced camera viewpoints. However, these methods struggle in few-shot scenarios due to limited supervision. In this paper, we present NexusGS, a 3DGS-…

2025

Occlusion-Insensitive Talking Head Video Generation via Facelet Compensation

AAAI 2025technical

Talking head video generation involves animating a still face image using facial motion cues derived from a driving video to replicate target poses and expressions. Traditional methods often rely on the assumption that the relative positions of facial keypoints remain unchanged. However, this assump…

Cited by 0SourcePDFScholar
2025

OmniVTON: Training-Free Universal Virtual Try-On

ICCV 2025poster

Image-based Virtual Try-On (VTON) techniques rely on either supervised in-shop approaches, which ensure high fidelity but struggle with cross-domain generalization, or unsupervised in-the-wild methods, which improve adaptability but remain constrained by data biases and limited universality. A unifi…

2025

PersonaMagic: Stage-Regulated High-Fidelity Face Customization with Tandem Equilibrium

AAAI 2025technical

Personalized image generation has made significant strides in adapting content to novel concepts. However, a persistent challenge remains: balancing the accurate reconstruction of unseen concepts with the need for editability according to the prompt, especially when dealing with the complex nuances…

2025

Playing to the Strengths of High- and Low-Resolution Cues for Ultra-High Resolution Image Segmentation

RA-L 2025

In ultra-high resolution image segmentation task for robotic platforms like UAVs and autonomous vehicles, existing paradigms process a downsampled input image through a deep network and the original high-resolution image through a shallow network, then fusing their features for final segmentation. A

Cited by 1SourceScholar
2025

RecDreamer: Consistent Text-to-3D Generation via Uniform Score Distillation

ICLR 2025poster

Current text-to-3D generation methods based on score distillation often suffer from geometric inconsistencies, leading to repeated patterns across different poses of 3D assets. This issue, known as the Multi-Face Janus problem, arises because existing methods struggle to maintain consistency across…

Cited by 3SourcePDFScholar
2025

Registration is a Powerful Rotation-Invariance Learner for 3D Anomaly Detection

NeurIPS 2025poster

3D anomaly detection in point-cloud data is critical for industrial quality control, aiming to identify structural defects with high reliability. However, current memory bank-based methods often suffer from inconsistent feature transformations and limited discriminative capacity, particularly in cap…

Cited by 0SourceScholar
2025

Safe-Sora: Safe Text-to-Video Generation via Graphical Watermarking

NeurIPS 2025poster

The explosive growth of generative video models has amplified the demand for reliable copyright preservation of AI-generated content. Despite its popularity in image synthesis, invisible generative watermarking remains largely underexplored in video generation. To address this gap, we propose Safe-S…

Cited by 0SourceScholar
2025

StableGuard: Towards Unified Copyright Protection and Tamper Localization in Latent Diffusion Models

NeurIPS 2025poster

The advancement of diffusion models has enhanced the realism of AI-generated content but also raised concerns about misuse, necessitating robust copyright protection and tampering localization. Although recent methods have made progress toward unified solutions, their reliance on post hoc processing…

Cited by 0SourceScholar
2025

Stroke2Sketch: Harnessing Stroke Attributes for Training-Free Sketch Generation

ICCV 2025poster

Generating sketches guided by reference styles requires precise transfer of stroke attributes, such as line thickness, deformation, and texture sparsity, while preserving semantic structure and content fidelity. To this end, we propose Stroke2Sketch, a novel training-free framework that introduces c…

2025

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition

ICCV 2025poster

3D visual grounding aims to identify and localize objects in a 3D space based on textual descriptions. However, existing methods struggle with disentangling targets from anchors in complex multi-anchor queries and resolving inconsistencies in spatial descriptions caused by perspective variations.To…

2024

Beat-It: Beat-Synchronized Multi-Condition 3D Dance Generation

ECCV 2024oral

"Dance, as an art form, fundamentally hinges on the precise synchronization with musical beats. However, achieving aesthetically pleasing dance sequences from music is challenging, with existing methods often falling short in controllability and beat alignment. To address these shortcomings, this pa…

2024

Beyond Textual Constraints: Learning Novel Diffusion Conditions with Fewer Examples

CVPR 2024poster

In this paper we delve into a novel aspect of learning novel diffusion conditions with datasets an order of magnitude smaller. The rationale behind our approach is the elimination of textual constraints during the few-shot learning process. To that end we implement two optimization strategies. The f…

2024

D3still: Decoupled Differential Distillation for Asymmetric Image Retrieval

CVPR 2024poster

Existing methods for asymmetric image retrieval employ a rigid pairwise similarity constraint between the query network and the larger gallery network. However these one-to-one constraint approaches often fail to maintain retrieval order consistency especially when the query network has limited repr…

2024

Delving into Multimodal Prompting for Fine-Grained Visual Classification

AAAI 2024technical

Fine-grained visual classification (FGVC) involves categorizing fine subdivisions within a broader category, which poses challenges due to subtle inter-class discrepancies and large intra-class variations. However, prevailing approaches primarily focus on uni-modal visual concepts. Recent advancemen…

Cited by 28SourcePDFScholar
2024

Drag Your Noise: Interactive Point-based Editing via Diffusion Semantic Propagation

CVPR 2024poster

Point-based interactive editing serves as an essential tool to complement the controllability of existing generative models. A concurrent work DragDiffusion updates the diffusion latent map in response to user inputs causing global latent map alterations. This results in imprecise preservation of th…

2024

Learning with Unreliability: Fast Few-shot Voxel Radiance Fields with Relative Geometric Consistency

CVPR 2024poster

We propose a voxel-based optimization framework ReVoRF for few-shot radiance fields that strategically addresses the unreliability in pseudo novel view synthesis. Our method pivots on the insight that relative depth relationships within neighboring regions are more reliable than the absolute color v…

2024

OneRestore: A Universal Restoration Framework for Composite Degradation

ECCV 2024poster

"In real-world scenarios, image impairments often manifest as composite degradations, presenting a complex interplay of elements such as low light, haze, rain, and snow. Despite this reality, existing restoration methods typically target isolated degradation types, thereby falling short in environme…

2024

Rethinking Multi-view Representation Learning via Distilled Disentangling

CVPR 2024poster

Multi-view representation learning aims to derive robust representations that are both view-consistent and view-specific from diverse data sources. This paper presents an in-depth analysis of existing approaches in this domain highlighting a commonly overlooked aspect: the redundancy between view-co…

2023

CIRI: Curricular Inactivation for Residue-aware One-shot Video Inpainting

ICCV 2023poster

Video inpainting aims at filling in missing regions of a video. However, when dealing with dynamic scenes with camera or object movements, annotating the inpainting target becomes laborious and impractical. In this paper, we resolve the one-shot video inpainting problem in which only one annotated f…

Cited by 9PDFcodeScholar
2023

Curricular Contrastive Regularization for Physics-Aware Single Image Dehazing

CVPR 2023poster

Considering the ill-posed nature, contrastive regularization has been developed for single image dehazing, introducing the information from negative images as a lower bound. However, the contrastive samples are nonconsensual, as the negatives are usually represented distantly from the clear (i.e., p…

2023

Diffuse3D: Wide-Angle 3D Photography via Bilateral Diffusion

ICCV 2023poster

This paper aims to resolve the challenging problem of wide-angle novel view synthesis from a single image, a.k.a. wide-angle 3D photography. Existing approaches rely on local context and treat them equally to inpaint occluded RGB and depth regions, which fail to deal with large-region occlusion (i.e…

Cited by 8PDFcodeScholar
2023

Layout Generation as Intermediate Action Sequence Prediction

AAAI 2023technical

Layout generation plays a crucial role in graphic design intelligence. One important characteristic of the graphic layouts is that they usually follow certain design principles. For example, the principle of repetition emphasizes the reuse of similar visual elements throughout the design. To generat…

2023

RIGID: Recurrent GAN Inversion and Editing of Real Face Videos

ICCV 2023poster

GAN inversion is indispensable for applying the powerful editability of GAN to real images. However, existing methods invert video frames individually often leading to undesired inconsistent results over time. In this paper, we propose a unified recurrent framework, named Recurrent vIdeo GAN Inversi…

Cited by 8PDFcodeScholar
2023

Towards a Smaller Student: Capacity Dynamic Distillation for Efficient Image Retrieval

CVPR 2023poster

Previous Knowledge Distillation based efficient image retrieval methods employ a lightweight network as the student model for fast inference. However, the lightweight student model lacks adequate representation capacity for effective knowledge imitation during the most critical early training period…

Cited by 22SourcePDFScholar
2023

Where Is My Spot? Few-Shot Image Generation via Latent Subspace Optimization

CVPR 2023poster

Image generation relies on massive training data that can hardly produce diverse images of an unseen category according to a few examples. In this paper, we address this dilemma by projecting sparse few-shot samples into a continuous latent space that can potentially generate infinite unseen samples…

2022

A Simple Data Mixing Prior for Improving Self-Supervised Learning

CVPR 2022poster

Data mixing (e.g., Mixup, Cutmix, ResizeMix) is an essential component for advancing recognition models. In this paper, we focus on studying its effectiveness in the self-supervised setting. By noticing the mixed images that share the same source images are intrinsically related to each other, we he…

Cited by 48PDFcodeScholar
2022

Co-Advise: Cross Inductive Bias Distillation

CVPR 2022poster

The inductive bias of vision transformers is more relaxed that cannot work well with insufficient data. Knowledge distillation is thus introduced to assist the training of transformers. Unlike previous works, where merely heavy convolution-based teachers are provided, in this paper, we delve into th…

Cited by 82PDFcodeScholar
2022

Editing Out-of-Domain GAN Inversion via Differential Activations

ECCV 2022poster

"Despite the demonstrated editing capacity in the latent space of a pretrained GAN model, inverting real-world images is stuck in a dilemma that the reconstruction cannot be faithful to the original input. The main reason for this is that the distributions between training and real-world data are mi…

2022

Faithful Extreme Rescaling via Generative Prior Reciprocated Invertible Representations

CVPR 2022oral

This paper presents a Generative prior ReciprocAted Invertible rescaling Network (GRAIN) for generating faithful high-resolution (HR) images from low-resolution (LR) invertible images with an extreme upscaling factor (64x). Previous researches have leveraged the prior knowledge of a pretrained GAN m…

Cited by 15PDFcodeScholar
2022

High-Resolution Face Swapping via Latent Semantics Disentanglement

CVPR 2022poster

We present a novel high-resolution face swapping method using the inherent prior knowledge of a pre-trained GAN model. Although previous research can leverage generative priors to produce high-resolution results, their quality can suffer from the entangled semantics of the latent space. We explicitl…

Cited by 96PDFcodeScholar
2022

Shunted Self-Attention via Multi-Scale Token Aggregation

CVPR 2022oral

Recent Vision Transformer (ViT) models have demonstrated encouraging results across various computer vision tasks, thanks to its competence in modeling long-range dependencies of image patches or tokens via self-attention. These models, however, usually designate the similar receptive fields of each…

Cited by 335PDFcodeScholar
2021

Delving Deep Into Many-to-Many Attention for Few-Shot Video Object Segmentation

CVPR 2021poster

This paper tackles the task of Few-Shot Video Object Segmentation (FSVOS), i.e., segmenting objects in the query videos with certain class specified in a few labeled support images. The key is to model the relationship between the query videos and the support images for propagating the object inform…

Cited by 25PDFcodeScholar
2021

Discovering Interpretable Latent Space Directions of GANs Beyond Binary Attributes

CVPR 2021poster

Generative adversarial networks (GANs) learn to map noise latent vectors to high-fidelity image outputs. It is found that the input latent space shows semantic correlations with the output image space. Recent works aim to interpret the latent space and discover meaningful directions that correspond…

Cited by 64PDFcodeScholar
2021

From Contexts to Locality: Ultra-High Resolution Image Segmentation via Locality-Aware Contextual Correlation

ICCV 2021poster

Ultra-high resolution image segmentation has raised increasing interests in recent years due to its realistic applications. In this paper, we innovate the widely used high-resolution image segmentation pipeline, in which an ultra-high resolution image is partitioned into regular patches for local se…

Cited by 59PDFcodeScholar
2021

From Continuity to Editability: Inverting GANs With Consecutive Images

ICCV 2021poster

Existing GAN inversion methods are stuck in a paradox that the inverted codes can either achieve high-fidelity reconstruction, or retain the editing capability. Having only one of them clearly cannot realize real image editing. In this paper, we resolve this paradox by introducing consecutive images…

Cited by 46PDFcodeScholar
2021

Learning From the Master: Distilling Cross-Modal Advanced Knowledge for Lip Reading

CVPR 2021poster

Lip reading aims to predict the spoken sentences from silent lip videos. Due to the fact that such a vision task usually performs worse than its counterpart speech recognition, one potential scheme is to distill knowledge from a teacher pretrained by audio signals. However, the latent domain gap bet…

Cited by 86PDFScholar
2021

Projecting Your View Attentively: Monocular Road Scene Layout Estimation via Cross-View Transformation

CVPR 2021poster

HD map reconstruction is crucial for autonomous driving. LiDAR-based methods are limited due to the deployed expensive sensors and time-consuming computation. Camera-based methods usually need to separately perform road segmentation and view transformation, which often causes distortion and the abse…

Cited by 115PDFcodeScholar
2021

Reciprocal Transformations for Unsupervised Video Object Segmentation

CVPR 2021poster

Unsupervised video object segmentation (UVOS) aims at segmenting the primary objects in videos without any human intervention. Due to the lack of prior knowledge about the primary objects, identifying them from videos is the major challenge of UVOS. Previous methods often regard the moving objects a…

Cited by 107PDFcodeScholar
2021

Spatially-Invariant Style-Codes Controlled Makeup Transfer

CVPR 2021poster

Transferring makeup from the misaligned reference image is challenging. Previous methods overcome this barrier by computing pixel-wise correspondences between two images, which is inaccurate and computational-expensive. In this paper, we take a different perspective to break down the makeup transfer…

Cited by 78PDFcodeScholar
2020

Context-Aware and Scale-Insensitive Temporal Repetition Counting

CVPR 2020poster

Temporal repetition counting aims to estimate the number of cycles of a given repetitive action. Existing deep learning methods assume repetitive actions are performed in a fixed time-scale, which is invalid for the complex repetitive actions in real life. In this paper, we tailor a context-aware an…

Cited by 71PDFcodeScholar
2020

Don't Hit Me! Glass Detection in Real-World Scenes

CVPR 2020poster

Glass is very common in our daily life. Existing computer vision systems neglect it and thus may have severe consequences, e.g., a robot may crash into a glass wall. However, sensing the presence of glass is not straightforward. The key challenge is that arbitrary objects/scenes can appear behind th…

Cited by 165PDFScholar
2020

TENet: Triple Excitation Network for Video Salient Object Detection

ECCV 2020poster

In this paper, we propose a simple yet effective approach, named Triple Excitation Network, to reinforce the training of video salient object detection (VSOD) from three aspects, spatial, temporal, and online excitations. These excitation mechanisms are designed following the spirit of curriculum le…

Cited by 75SourcePDFScholar
2019

Context-Aware Spatio-Recurrent Curvilinear Structure Segmentation

CVPR 2019poster

Curvilinear structures are frequently observed in various images in different forms, such as blood vessels or neuronal boundaries in biomedical images. In this paper, we propose a novel curvilinear structure segmentation approach using context-aware spatio-recurrent networks. Instead of directly seg…

Cited by 26PDFScholar
2019

Self-Supervised Spatio-Temporal Representation Learning for Videos by Predicting Motion and Appearance Statistics

CVPR 2019poster

We address the problem of video representation learning without human-annotated labels. While previous efforts address the problem by designing novel self-supervised tasks using video data, the learned features are merely on a frame-by-frame basis, which are not applicable to many video analytic tas…

Cited by 259PDFcodeScholar
2019

Visualizing the Invisible: Occluded Vehicle Segmentation and Recovery

ICCV 2019poster

In this paper, we propose a novel iterative multi-task framework to complete the segmentation mask of an occluded vehicle and recover the appearance of its invisible parts. In particular, firstly, to improve the quality of the segmentation completion, we present two coupled discriminators that intro…

Cited by 45PDFScholar
2017

Delving Into Salient Object Subitizing and Detection

ICCV 2017poster

Subitizing (i.e., instant judgement on the number) and detection of salient objects are human inborn abilities. These two tasks influence each other in the human visual system. In this paper, we delve into the complementarity of these two tasks. We propose a multi-task deep neural network with weigh…

Cited by 68PDFScholar
2017

DeshadowNet: A Multi-Context Embedding Deep Network for Shadow Removal

CVPR 2017spotlight

Shadow removal is a challenging task as it requires the detection/annotation of shadows as well as semantic understanding of the scene. In this paper, we propose an automatic and end-to-end deep neural network (DeshadowNet) to tackle these problems in a unified manner. DeshadowNet is designed with a…

Cited by 367PDFcodeScholar
2016

Real-Time Salient Object Detection With a Minimum Spanning Tree

CVPR 2016spotlight

In this paper, we present a real-time salient object detection system based on the minimum spanning tree. Due to the fact that background regions are typically connected to the image boundaries, salient objects can be extracted by computing the distances to the boundaries. However, measuring the ima…

Cited by 340PDFScholar
2015

Oriented Object Proposals

ICCV 2015poster

In this paper, we propose a new approach to generate oriented object proposals (OOPs) to reduce the detection error caused by various orientations of the object. To this end, we propose to efficiently locate object regions according to pixelwise object probability, rather than measuring the objectne…

Cited by 36PDFScholar