← Search

Mingbo Zhao

6 accepted papers

2026

Where I Am & Where to Go: Egocentric Indoor Scene Perception with Agent Interaction for Remote Embodied Visual Grounding

ICRA 2026poster

Embodied Referring Expression Grounding (REVERIE) is a Vision-and-Language Navigation (VLN) task that better reflects real-world human instructions. Unlike conventional VLN, REVERIE is more challenging as agents must navigate in unseen environments and ground remote objects described by short, high-…

Cited by 0Scholar
2024

Attention Disturbance and Dual-Path Constraint Network for Occluded Person Re-identification

AAAI 2024technical

Occluded person re-identification (Re-ID) aims to address the potential occlusion problem when matching occluded or holistic pedestrians from different camera views. Many methods use the background as artificial occlusion and rely on attention networks to exclude noisy interference. However, the si…

Cited by 10SourcePDFScholar
2024

OSIC: A New One-Stage Image Captioner Coined

IJCAI 2024poster

Mainstream image captioning models are usually two-stage captioners, i.e., encoding the region features by a pre-trained detector and then feeding them into a language model to generate the captions. However, such a two-stage procedure will lead to a task-based information gap that decreases the per…

Cited by 7SourcePDFScholar
2023

Arbitrary Virtual Try-on Network: Characteristics Representation and Trade-off between Body and Clothing

ICLR 2023poster

Deep learning based virtual try-on system has achieved some encouraging progress recently, but there still remain several big challenges that need to be solved, such as trying on arbitrary clothes of all types, trying on the clothes from one category to another and generating image-realistic results…

Cited by 0SourcePDFScholar
2022

A Simple Approach to Automated Spectral Clustering

NeurIPS 2022accept

The performance of spectral clustering heavily relies on the quality of affinity matrix. A variety of affinity-matrix-construction (AMC) methods have been proposed but they have hyperparameters to determine beforehand, which requires strong experience and leads to difficulty in real applications, es…

2019

NDDR-CNN: Layerwise Feature Fusing in Multi-Task CNNs by Neural Discriminative Dimensionality Reduction

CVPR 2019poster

In this paper, we propose a novel Convolutional Neural Network (CNN) structure for general-purpose multi-task learning (MTL), which enables automatic feature fusing at every layer from different tasks. This is in contrast with the most widely used MTL CNN structures which empirically or heuristicall…

Cited by 347PDFcodeScholar