← Search

Feng Zheng

79 accepted papers

2026

ActiveVLN: Towards Active Exploration Via Multi-Turn RL in Vision-And-Language Navigation

ICRA 2026poster

The Vision-and-Language Navigation (VLN) task requires an agent to follow natural language instructions and navigate through complex environments. Existing MLLM-based VLN methods primarily rely on imitation learning (IL) and often use DAgger for post-training to mitigate covariate shift. While effec…

2026

Leveraging Geometric Priors for Unaligned Scene Change Detection

ICRA 2026poster

Unaligned Scene Change Detection aims to detect scene changes between image pairs captured at different times without assuming viewpoint alignment. To handle viewpoint variations, current methods rely solely on 2D visual cues to establish cross-image correspondence to assist change detection. Howeve…

2026

MM-Snowball: Evaluating and Mitigating Hallucination Snowballing in Multimodal Multi-turn Dialogue

ICML 2026poster

Multimodal Large Language Models (MLLMs) demonstrate remarkable visual understanding, yet their reliability in interactive settings is severely undermined by {hallucination snowballing}: a phenomenon where initial errors amplify across conversational turns, leading to a collapse in coherence. This f…

Cited by 0SourceScholar
2026

R-AVST: Empowering Video-LLMs with Fine-Grained Spatio-Temporal Reasoning in Complex Audio-Visual Scenarios

AAAI 2026technical

Recently, rapid advancements have been made in multimodal large language models (MLLMs), especially in video understanding tasks. However, current research focuses on simple video scenarios, failing to reflect the complex and diverse nature of real-world audio-visual events in videos. To bridge this

Cited by 0SourcePDFScholar
2026

STABLE: Simulation-Ready Tabletop Layout Generation via a Semantics–Physics Dual System

ICML 2026poster

Generating simulation-ready tabletop scenes from task instructions is an intriguing and promising research direction in the field of Embodied AI. However, existing task-to-scene generation methods rely exclusively on large language models (LLMs) to predict scene layouts, inevitably yielding object c…

Cited by 0SourceScholar
2026

Scalable Object Relation Encoding for Better 3D Spatial Reasoning in Large Language Models

CVPR 2026

Spatial reasoning focuses on locating target objects based on spatial relations in 3D scenes, which plays a crucial role in developing intelligent embodied agents. Due to the limited availability of 3D scene-language paired data, it is challenging to train models with strong reasoning ability from s

Cited by 0SourcecodeScholar
2026

Spatially Guided Training for Vision-Language-Action Model

ICLR 2026poster

Large vision–language models (VLMs) excel at multimodal understanding but fall short when extended to embodied tasks, where instructions must be transformed into low-level motor actions. We introduce SP-VLA, a dual-system **V**ision–**L**anguage–**A**ction framework that leverages **S**patial **P**r…

Cited by 0SourceScholar
2026

Transferability of Adversarial Attacks in Video-based MLLMs: A Cross-modal Image-to-Video Approach

AAAI 2026technical

Video-based multimodal large language models (V-MLLMs) have shown vulnerability to adversarial examples in video-text multimodal tasks. However, the transferability of adversarial videos to unseen models—a common and practical real-world scenario—remains unexplored. In this paper, we pioneer an in

Cited by 0SourcePDFScholar
2025

An Information-theoretic Perspective of Hierarchical Clustering on Graphs

UAI 2025

The seminal work of \citep{dasgupta2016cost} has introduced a combinatorial cost function for hierarchical graph clustering that has inspired numerous follow-up studies adopting similar combinatorial approaches. In this paper, we investigate this problem from the \emph{information-theoretic} perspec

2025

LLplace: Embodied 3D Indoor Layout Synthesis Framework with Large Language Model

IROS 2025

Designing 3D indoor layouts is a crucial task with significant applications in embodied robot intelligence, virtual reality, and interior design. Existing methods for 3D layout design either rely on diffusion models, which utilize spatial relationship priors, or heavily leverage the inferential capa

Cited by 0SourceScholar
2025

LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos

CVPR 2025poster

Despite impressive advancements in video understanding, most efforts remain limited to coarse-grained or visual-only video tasks. However, real-world videos encompass omni-modal information (vision, audio, and speech) with a series of events forming a cohesive storyline. The lack of multi-modal vide…

2025

MMAD: A Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection

ICLR 2025poster

In the field of industrial inspection, Multimodal Large Language Models (MLLMs) have a high potential to renew the paradigms in practical applications due to their robust language capabilities and generalization abilities. However, despite their impressive problem-solving skills in many domains, MLL…

2025

MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial Reasoning

NeurIPS 2025spotlight

The ability of robots to interpret human instructions and execute manipulation tasks necessitates the availability of task-relevant tabletop scenes for training. However, traditional methods for creating these scenes rely on time-consuming manual layout design or purely randomized layouts, which are…

Cited by 0SourceScholar
2025

On the Generalization Ability of Next-Token-Prediction Pretraining

ICML 2025poster

Large language models (LLMs) have demonstrated remarkable potential in handling natural language processing (NLP) tasks and beyond. LLMs usually can be categorized as transformer decoder-only models (DOMs), utilizing Next-Token-Prediction (NTP) as their pre-training methodology. Despite their tremen…

Cited by 0SourcePDFScholar
2025

OptiScene: LLM-driven Indoor Scene Layout Generation via Scaled Human-aligned Data Synthesis and Multi-Stage Preference Optimization

NeurIPS 2025poster

Automatic indoor layout generation has attracted increasing attention due to its potential in interior design, virtual environment construction, and embodied AI. Existing methods fall into two categories: prompt-driven approaches that leverage proprietary LLM services (e.g., GPT APIs), and learning-…

Cited by 0SourceScholar
2025

Sample then Identify: A General Framework for Risk Control and Assessment in Multimodal Large Language Models

ICLR 2025spotlight

Multimodal Large Language Models (MLLMs) exhibit promising advancements across various tasks, yet they still encounter significant trustworthiness issues. Prior studies apply Split Conformal Prediction (SCP) in language modeling to construct prediction sets with statistical guarantees. However, thes…

Cited by 6SourcePDFScholar
2025

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors

EMNLP 2025

Recent advancements in large video-language models have revolutionized video understanding tasks. However, their efficiency is significantly constrained by processing high volumes of visual tokens. Existing token compression strategies apply a fixed compression ratio, ignoring the variability in sem

2024

Beyond Prototypes: Semantic Anchor Regularization for Better Representation Learning

AAAI 2024technical

One of the ultimate goals of representation learning is to achieve compactness within a class and well-separability between classes. Many outstanding metric-based and prototype-based methods following the Expectation-Maximization paradigm, have been proposed for this objective. However, they inevita…

2024

Block Image Compressive Sensing with Local and Global Information Interaction

AAAI 2024technical

Block image compressive sensing methods, which divide a single image into small blocks for efficient sampling and reconstruction, have achieved significant success. However, these methods process each block locally and thus disregard the global communication among different blocks in the reconstruct…

2024

Depth-Aware Concealed Crop Detection in Dense Agricultural Scenes

CVPR 2024poster

Concealed Object Detection (COD) aims to identify objects visually embedded in their background. Existing COD datasets and methods predominantly focus on animals or humans ignoring the agricultural domain which often contains numerous small and concealed crops with severe occlusions. In this paper w…

2024

Fine-grained Analysis of Stability and Generalization for Stochastic Bilevel Optimization

IJCAI 2024poster

Stochastic bilevel optimization (SBO) has been integrated into many machine learning paradigms recently including hyperparameter optimization, meta learning, reinforcement learning, etc. Along with the wide range of applications, there have been abundant studies on concerning the computing behavi…

Cited by 1SourcePDFScholar
2024

MS2SL: Multimodal Spoken Data-Driven Continuous Sign Language Production

ACL 2024findings

Sign language understanding has made significant strides; however, there is still no viable solution for generating sign sequences directlyfrom entire spoken content, e.g., text or speech. In this paper, we propose a unified framework for continuous sign language production, easing communication bet…

Cited by 3SourcePDFScholar
2024

Mutual Learning for Acoustic Matching and Dereverberation via Visual Scene-driven Diffusion

ECCV 2024poster

"Visual acoustic matching (VAM) is pivotal for enhancing the immersive experience, and the task of dereverberation is effective in improving audio intelligibility. Existing methods treat each task independently, overlooking the inherent reciprocity between them. Moreover, these methods depend on pai…

Cited by 1SourcePDFScholar
2024

Negative Label Guided OOD Detection with Pretrained Vision-Language Models

ICLR 2024spotlight

Out-of-distribution (OOD) detection aims at identifying samples from unknown classes, playing a crucial role in trustworthy models against errors on unexpected inputs. Extensive research has been dedicated to exploring OOD detection in the vision modality. {Vision-language models (VLMs) can lever…

2024

On the Noise Robustness of In-Context Learning for Text Generation

NeurIPS 2024poster

Large language models (LLMs) have shown impressive performance on downstream tasks by in-context learning (ICL), which heavily relies on the quality of demonstrations selected from a large set of annotated examples. Recent works claim that in-context learning is robust to noisy demonstrations in tex…

2024

Unlocking Memorization in Large Language Models with Dynamic Soft Prompting

EMNLP 2024main

Pretrained large language models (LLMs) have excelled in a variety of natural language processing (NLP) tasks, including summarization, question answering, and translation. However, LLMs pose significant security risks due to their tendency to memorize training data, leading to potential privacy bre…

2024

Unsupervised Continual Anomaly Detection with Contrastively-Learned Prompt

AAAI 2024technical

Unsupervised Anomaly Detection (UAD) with incremental training is crucial in industrial manufacturing, as unpredictable defects make obtaining sufficient labeled data infeasible. However, continual learning methods primarily rely on supervised annotations, while the application in UAD is limited due…

2023

Accelerating Vision-Language Pretraining With Free Language Modeling

CVPR 2023poster

The state of the arts in vision-language pretraining (VLP) achieves exemplary performance but suffers from high training costs resulting from slow convergence and long training time, especially on large-scale web datasets. An essential obstacle to training efficiency lies in the entangled prediction…

2023

Dense-Localizing Audio-Visual Events in Untrimmed Videos: A Large-Scale Benchmark and Baseline

CVPR 2023poster

Existing audio-visual event localization (AVE) handles manually trimmed videos with only a single instance in each of them. However, this setting is unrealistic as natural videos often contain numerous audio-visual events with different categories. To better adapt to real-life applications, in this…

2023

Detecting Out-of-distribution Data through In-distribution Class Prior

ICML 2023poster

Given a pre-trained in-distribution (ID) model, the inference-time out-of-distribution (OOD) detection aims to recognize OOD data during the inference stage. However, some representative methods share an unproven assumption that the probability that OOD data belong to every ID class should be the sa…

2023

Knowledge-Aware Prompt Tuning for Generalizable Vision-Language Models

ICCV 2023poster

Pre-trained vision-language models, e.g., CLIP, working with manually designed prompts have demonstrated great effectiveness in transfer learning. Recently, learnable prompts achieve state-of-the-art performance, which however are prone to overfit to seen classes while failing to generalize to unsee…

Cited by 37PDFScholar
2023

Learning Cross-Modal Affinity for Referring Video Object Segmentation Targeting Limited Samples

ICCV 2023poster

Referring video object segmentation (RVOS), as a supervised learning task, relies on sufficient annotated data for a given scene. However, in more realistic scenarios, only minimal annotations are available for a new scene, which poses significant challenges to existing RVOS methods. With this in mi…

Cited by 3PDFcodeScholar
2023

On the Stability and Generalization of Triplet Learning

AAAI 2023technical

Triplet learning, i.e. learning from triplet data, has attracted much attention in computer vision tasks with an extremely large number of categories, e.g., face recognition and person re-identification. Albeit with rapid progress in designing and applying triplet learning algorithms, there is a lac…

Cited by 5SourcePDFScholar
2023

Pushing the Limits of Fewshot Anomaly Detection in Industry Vision: Graphcore

ICLR 2023poster

In the area of few-shot anomaly detection (FSAD), efficient visual feature plays an essential role in the memory bank $\mathcal{M}$-based methods. However, these methods do not account for the relationship between the visual feature and its rotated visual feature, drastically limiting the anomaly de…

Cited by 79SourcePDFScholar
2023

Real3D-AD: A Dataset of Point Cloud Anomaly Detection

NeurIPS 2023poster

High-precision point cloud anomaly detection is the gold standard for identifying the defects of advancing machining and precision manufacturing. Despite some methodological advances in this area, the scarcity of datasets and the lack of a systematic benchmark hinder its development. We introduce Re…

2023

Set-level Guidance Attack: Boosting Adversarial Transferability of Vision-Language Pre-training Models

ICCV 2023oral

Vision-language pre-training (VLP) models have shown vulnerability to adversarial examples in multimodal tasks. Furthermore, malicious adversaries can be deliberately transferred to attack other black-box models. However, existing work has mainly focused on investigating white-box attacks. In this p…

Cited by 65PDFcodeScholar
2023

Skating-Mixer: Long-Term Sport Audio-Visual Modeling with MLPs

AAAI 2023technical

Figure skating scoring is challenging because it requires judging players’ technical moves as well as coordination with the background music. Most learning-based methods struggle for two reasons: 1) each move in figure skating changes quickly, hence simply applying traditional frame sampling will lo…

2023

Transferable Decoding with Visual Entities for Zero-Shot Image Captioning

ICCV 2023poster

Image-to-text generation aims to describe images using natural language. Recently, zero-shot image captioning based on pre-trained vision-language models (VLMs) and large language models (LLMs) has made significant progress. However, we have observed and empirically demonstrated that these methods a…

Cited by 53PDFcodeScholar
2022

Class-Aware Contrastive Semi-Supervised Learning

CVPR 2022poster

Pseudo-label-based semi-supervised learning (SSL) has achieved great success on raw data utilization. However, its training procedure suffers from confirmation bias due to the noise contained in self-generated artificial labels. Moreover, the model's judgment becomes noisier in real-world applicatio…

Cited by 139PDFcodeScholar
2022

Error-Based Knockoffs Inference for Controlled Feature Selection

AAAI 2022technical

Recently, the scheme of model-X knockoffs was proposed as a promising solution to address controlled feature selection under high-dimensional finite-sample settings. However, the procedure of model-X knockoffs depends heavily on the coefficient-based feature importance and only concerns the control…

Cited by 6SourcePDFScholar
2022

Generalized Brain Image Synthesis with Transferable Convolutional Sparse Coding Networks

ECCV 2022poster

"High inter-equipment variability and expensive examination costs of brain imaging remain key challenges in leveraging the heterogeneous scans effectively. Despite rapid growth in image-to-image translation with deep learning models, the target brain data may not always be achievable due to the spec…

Cited by 1SourcePDFScholar
2022

GuidedMix-Net: Semi-supervised Semantic Segmentation by Using Labeled Images as Reference

AAAI 2022technical

Semi-supervised learning is a challenging problem which aims to construct a model by learning from limited labeled examples. Numerous methods for this task focus on utilizing the predictions of unlabeled instances consistency alone to regularize networks. However, treating labeled and unlabeled data…

Cited by 26SourcePDFScholar
2022

Meta Distribution Alignment for Generalizable Person Re-Identification

CVPR 2022poster

Domain Generalizable (DG) person ReID is a challenging task which trains a model on source domains yet generalizes well on target domains. Existing methods use source domains to learn domain-invariant features, and assume those features are also irrelevant with target domains. However, they do not c…

Cited by 80PDFcodeScholar
2022

S2Contact: Graph-Based Network for 3D Hand-Object Contact Estimation with Semi-Supervised Learning

ECCV 2022poster

"Being able to reason about the physical contacts between hands and objects is crucial in understanding hand-object manipulation. However, despite the efforts in accurate 3D annotations in hand and object datasets, there still exist gaps in 3D hand and object reconstructions. Recent works leverage c…

Cited by 21SourcePDFScholar
2022

SoftPatch: Unsupervised Anomaly Detection with Noisy Data

NeurIPS 2022accept

Although mainstream unsupervised anomaly detection (AD) algorithms perform well in academic datasets, their performance is limited in practical application due to the ideal experimental setting of clean training data. Training with noisy data is an inevitable problem in real-world anomaly detection…

2022

Towards Generic 3D Tracking in RGBD Videos: Benchmark and Baseline

ECCV 2022poster

"Tracking in 3D scenes is gaining momentum because of its numerous applications in robotics, autonomous driving, and scene understanding. Currently, 3D tracking is limited to specific model-based approaches involving point clouds, which impedes 3D trackers from applying in natural 3D scenes. RGBD se…

2022

Unified Multivariate Gaussian Mixture for Efficient Neural Image Compression

CVPR 2022poster

Modeling latent variables with priors and hyperpriors is an essential problem in variational image compression. Formally, trade-off between rate and distortion is handled well if priors and hyperpriors precisely describe latent variables. Current practices only adopt univariate priors and process ea…

Cited by 73PDFcodeScholar
2022

VITA: A Multi-Source Vicinal Transfer Augmentation Method for Out-of-Distribution Generalization

AAAI 2022technical

Invariance to diverse types of image corruption, such as noise, blurring, or colour shifts, is essential to establish robust models in computer vision. Data augmentation has been the major approach in improving the robustness against common corruptions. However, the samples produced by popular augme…

Cited by 5SourcePDFScholar
2022

VLMixer: Unpaired Vision-Language Pre-training via Cross-Modal CutMix

ICML 2022spotlight

Existing vision-language pre-training (VLP) methods primarily rely on paired image-text datasets, which are either annotated by enormous human labors or crawled from the internet followed by elaborate data cleaning techniques. To reduce the dependency on well-aligned image-text pairs, it is promisin…

2021

A Unified Multi-Scenario Attacking Network for Visual Object Tracking

AAAI 2021technical

Existing methods of adversarial attacks successfully generate adversarial examples to confuse Deep Neural Networks (DNNs) of image classification and object detection, resulting in wrong predictions. However, these methods are difficult to attack models of video object tracking, because the tracking…

Cited by 19SourcePDFScholar
2021

Brain Image Synthesis With Unsupervised Multivariate Canonical CSCl4Net

CVPR 2021poster

Recent advances in neuroscience have highlighted the effectiveness of multi-modal medical data for investigating certain pathologies and understanding human cognition. However, obtaining full sets of different modalities is limited by various factors, such as long acquisition times, high examination…

Cited by 8PDFScholar
2021

DepthTrack: Unveiling the Power of RGBD Tracking

ICCV 2021poster

RGBD (RGB plus depth) object tracking is gaining momentum as RGBD sensors have become popular in many application fields such as robotics. However, the best RGBD trackers are extensions of the state-of-the-art deep RGB trackers. They are trained with RGB data and the depth channel is used as a sidek…

Cited by 92PDFcodeScholar
2021

Distributed Ranking with Communications: Approximation Analysis and Applications

AAAI 2021technical

Learning theory of distributed algorithms has recently attracted enormous attention in the machine learning community. However, most of existing works focus on learning problem with pointwise loss and does not consider the communication among local processors. In this paper, we propose a new distrib…

Cited by 1SourcePDFScholar
2021

Dual Distribution Alignment Network for Generalizable Person Re-Identification

AAAI 2021technical

Domain generalization (DG) offers a preferable real-world setting for Person Re-Identification (Re-ID), which trains a model using multiple source domain datasets and expects it to perform well in an unseen target domain without any model updating. Unfortunately, most DG approaches are designed expl…

Cited by 62SourcePDFScholar
2021

End-to-End Dense Video Captioning With Parallel Decoding

ICCV 2021poster

Dense video captioning aims to generate multiple associated captions with their temporal locations from the video. Previous methods follow a sophisticated "localize-then-describe" scheme, which heavily relies on numerous hand-crafted components. In this paper, we proposed a simple yet effective fram…

Cited by 238PDFcodeScholar
2021

FREE: Feature Refinement for Generalized Zero-Shot Learning

ICCV 2021poster

Generalized zero-shot learning (GZSL) has achieved significant progress, with many efforts dedicated to overcoming the problems of visual-semantic domain gaps and seen-unseen bias. However, most existing methods directly use feature extraction models trained on ImageNet alone, ignoring the cross-dat…

Cited by 244PDFcodeScholar
2021

Learning 3D Shape Feature for Texture-Insensitive Person Re-Identification

CVPR 2021poster

It is well acknowledged that person re-identification (person ReID) highly relies on visual texture information like clothing. Despite significant progress has been made in recent years, texture-confusing situations like clothing changing and persons wearing the same clothes receive little attention…

Cited by 147PDFScholar
2021

Norm-guided Adaptive Visual Embedding for Zero-Shot Sketch-Based Image Retrieval

IJCAI 2021poster

Zero-shot sketch-based image retrieval (ZS-SBIR), which aims to retrieve photos with sketches under the zero-shot scenario, has shown extraordinary talents in real-world applications. Most existing methods leverage language models to generate class-prototypes and use them to arrange the locations of…

Cited by 26SourcePDFScholar
2021

One for More: Selecting Generalizable Samples for Generalizable ReID Model

AAAI 2021technical

Current training objectives of existing person Re-IDentification (ReID) models only ensure that the loss of the model decreases on selected training batch, with no regards to the performance on samples outside the batch. It will inevitably cause the model to over-fit the data in the dominant positio…

Cited by 21SourcePDFScholar
2021

Seminar Learning for Click-Level Weakly Supervised Semantic Segmentation

ICCV 2021poster

Annotation burden has become one of the biggest barriers to semantic segmentation. Approaches based on click-level annotations have therefore attracted increasing attention due to their superior trade-off between supervision and annotation cost. In this paper, we propose seminar learning, a new lear…

Cited by 41PDFScholar
2020

Enabling Deep Residual Networks for Weakly Supervised Object Detection

ECCV 2020poster

Weakly supervised object detection (WSOD) has attracted extensive research attention due to its great flexibility of exploiting large-scale image-level annotation for detector training. Whilst deep residual networks such as ResNet and DenseNet have become the standard backbones for many computer vis…

Cited by 59SourcePDFScholar
2020

Hijacking Tracker: A Powerful Adversarial Attack on Visual Tracking

ICASSP 2020accepted

Visual object tracking has made important breakthroughs with the assistance of deep learning models. Unfortunately, recent research has clearly proved that deep learning models are vulnerable to malicious adversarial attacks, which mislead the models making wrong decisions by perturbing the input im…

Cited by 0SourceScholar
2020

Multi-task Additive Models for Robust Estimation and Automatic Structure Discovery

NeurIPS 2020poster

Additive models have attracted much attention for high-dimensional regression estimation and variable selection. However, the existing models are usually limited to the single-task learning framework under the mean squared error (MSE) criterion, where the utilization of variable structure depends he…

Cited by 17SourcePDFScholar
2020

Noise-Aware Fully Webly Supervised Object Detection

CVPR 2020poster

We investigate the emerging task of learning object detectors with sole image-level labels on the web without requiring any other supervision like precise annotations or additional images from well-annotated benchmark datasets. Such a task, termed as fully webly supervised object detection, is extre…

Cited by 40PDFScholar
2020

One-Shot Adversarial Attacks on Visual Tracking With Dual Attention

CVPR 2020poster

Almost all adversarial attacks in computer vision are aimed at pre-known object categories, which could be offline trained for generating perturbations. But as for visual object tracking, the tracked target categories are normally unknown in advance. However, the tracking algorithms also have potent…

Cited by 100PDFScholar
2020

Salience-Guided Cascaded Suppression Network for Person Re-Identification

CVPR 2020poster

Employing attention mechanisms to model both global and local features as a final pedestrian representation has become a trend for person re-identification (Re-ID) algorithms. A potential limitation of these methods is that they focus on the most salient features, but the re-identification of a pers…

Cited by 307PDFScholar
2020

Super-Resolution and Inpainting with Degraded and Upgraded Generative Adversarial Networks

IJCAI 2020poster

Image super-resolution (SR) and image inpainting are two topical problems in medical image processing. Existing methods for solving the problems are either tailored to recovering a high-resolution version of the low-resolution image or focus on filling missing values, thus inevitably giving rise to…

Cited by 0SourcePDFScholar
2020

Zero-Shot Object Detection via Learning an Embedding from Semantic Space to Visual Space

IJCAI 2020poster

Zero-shot object detection (ZSD) has received considerable attention from the community of computer vision in recent years. It aims to simultaneously locate and categorize previously unseen objects during inference. One crucial problem of ZSD is how to accurately predict the label of each object pro…

Cited by 0SourcePDFScholar
2019

Pyramidal Person Re-IDentification via Multi-Loss Dynamic Training

CVPR 2019poster

Most existing Re-IDentification (Re-ID) methods are highly dependent on precise bounding boxes that enable images to be aligned with each other. However, due to the challenging practical scenarios, current detection models often produce inaccurate bounding boxes, which inevitably degenerate the perf…

Cited by 502PDFcodeScholar
2018

PIRVS: An Advanced Visual-Inertial SLAM System with Flexible Sensor Fusion and Hardware Co-Design

ICRA 2018poster

In this paper, we present the PerceptIn Robotics Vision System (PIRVS), a visual-inertial computing hardware with embedded simultaneous localization and mapping (SLAM) algorithm. The PIRVS hardware is equipped with a multi-core processor, a global-shutter stereo camera, and an IMU with precise hardw…

Cited by 54SourceScholar
2018

Trifo-VIO: Robust and Efficient Stereo Visual Inertial Odometry Using Points and Lines

IROS 2018poster

In this paper, we present the Trifo Visual Inertial Odometry (Trifo-VIO), a tightly-coupled filtering-based stereo VIO system using both points and lines. Line features help improve system robustness in challenging scenarios when point features cannot be reliably detected or tracked, e.g. low-textur…

Cited by 70SourceScholar
2018

Unsupervised Deep Generative Adversarial Hashing Network

CVPR 2018poster

Unsupervised deep hash functions have not shown satisfactory improvements against the shallow alternatives, and usually, require supervised pretraining to avoid getting stuck in bad local minima. In this paper, we propose a deep unsupervised hashing function, called HashGAN, which outperforms unsupe…

Cited by 144SourcePDFScholar