← Search

Mengshi Qi

21 accepted papers

2026

Active Exploring like a Pigeon: Reinforcing Spatial Reasoning via Agentic Vision-Language Models

ICML 2026poster

Enabling Vision-Language Models (VLMs) to perform spatial reasoning remains challenging. Existing approaches treat VLMs as passive observers, which is difficult for real-world applications. Moreover, reinforcement learning methods rely on sparse rewards, limiting their effectiveness for complex reas…

Cited by 0SourceScholar
2026

Improving Batch Normalization with Test-Time Adaptation for Robust Object Detection in Self-Driving

AAAI 2026technical

In open real-world autonomous driving scenarios, challenges such as sensor failure and extreme weather hinder the generalization of current autonomous driving perception models to these unseen domain, due to the domain shifts between the test and training data. As the parameter scale of autonomous d

Cited by 0SourcePDFScholar
2026

Robo-SGG: Exploiting Layout-Oriented Normalization and Restitution Can Improve Robust Scene Graph Generation

CVPR 2026

We propose Robo-SGG, a plug-and-play module for robust scene graph generation (SGG). Unlike standard SGG, robust SGG aims to perform inference on a diverse range of corrupted images, with the core challenge being the domain shift between clean and corrupted images. Existing SGG methods suffer from d

Cited by 0SourcecodeScholar
2025

Synergistic Tensor and Pipeline Parallelism

NeurIPS 2025poster

In the machine learning system, the hybrid model parallelism combining tensor parallelism (TP) and pipeline parallelism (PP) has become the dominant solution for distributed training of Large Language Models~(LLMs) and Multimodal LLMs (MLLMs). However, TP introduces significant collective communicat…

Cited by 0SourcecodeScholar
2025

T2SG: Traffic Topology Scene Graph for Topology Reasoning in Autonomous Driving

CVPR 2025poster

Understanding the traffic scenes and then generating high-definition (HD) maps present significant challenges in autonomous driving. In this paper, we defined a novel \underline T raffic \underline T opology \underline S cene \underline G raph (\text T ^2\text SG ), a unified scene graph explicitl…

2025

Towards Efficient Object Re-Identification with a Novel Cloud-Edge Collaborative Framework

AAAI 2025technical

Object re-identification (ReID) is committed to searching for objects of the same identity across cameras, and its real-world deployment is gradually increasing. Current ReID methods assume that the deployed system follows the centralized processing paradigm, i.e., all computations are conducted in…

Cited by 0SourcePDFScholar
2025

VIoTGPT: Learning to Schedule Vision Tools Towards Intelligent Video Internet of Things

AAAI 2025technical

Video Internet of Things (VIoT) has shown full potential in collecting an unprecedented volume of video data. How to schedule the domain-specific perceiving models and analyze the collected videos uniformly, efficiently, and especially intelligently to accomplish complicated tasks is challenging. To…

2024

Decomposed Vector-Quantized Variational Autoencoder for Human Grasp Generation

ECCV 2024poster

"Generating realistic human grasps is a crucial yet challenging task for applications involving object manipulation in computer graphics and robotics. Existing methods often struggle with generating fine-grained realistic human grasps that ensure all fingers effectively interact with objects, as the…

2024

SGFormer: Semantic Graph Transformer for Point Cloud-Based 3D Scene Graph Generation

AAAI 2024technical

In this paper, we propose a novel model called SGFormer, Semantic Graph TransFormer for point cloud-based 3D scene graph generation. The task aims to parse a point cloud-based scene into a semantic structural graph, with the core challenge of modeling the complex global structure. Existing methods b…

2024

Semi-Supervised Teacher-Reference-Student Architecture for Action Quality Assessment

ECCV 2024poster

"Existing action quality assessment (AQA) methods often require a large number of label annotations for fully supervised learning, which are laborious and expensive. In practice, the labeled data are difficult to obtain because the AQA annotation process requires domain-specific expertise. In this p…

2024

Weakly-Supervised Temporal Action Localization by Inferring Salient Snippet-Feature

AAAI 2024technical

Weakly-supervised temporal action localization aims to locate action regions and identify action categories in untrimmed videos simultaneously by taking only video-level labels as the supervision. Pseudo label generation is a promising strategy to solve the challenging problem, but the current metho…

2023

Disentangled Counterfactual Learning for Physical Audiovisual Commonsense Reasoning

NeurIPS 2023poster

In this paper, we propose a Disentangled Counterfactual Learning (DCL) approach for physical audiovisual commonsense reasoning. The task aims to infer objects’ physics commonsense based on both video and audio input, with the main challenge is how to imitate the reasoning ability of humans. Most of…

2023

Unsupervised Self-Driving Attention Prediction via Uncertainty Mining and Knowledge Embedding

ICCV 2023poster

Predicting attention regions of interest is an important yet challenging task for self-driving systems. Existing methodologies rely on large-scale labeled traffic datasets that are labor-intensive to obtain. Besides, the huge domain gap between natural scenes and traffic scenes in current datasets a…

Cited by 6PDFcodeScholar
2022

RGB-Depth Fusion GAN for Indoor Depth Completion

CVPR 2022poster

The raw depth image captured by the indoor depth sensor usually has an extensive range of missing depth values due to inherent limitations such as the inability to perceive transparent objects and limited distance range. The incomplete depth map burdens many downstream vision tasks, and a rising num…

Cited by 44PDFScholar
2019

Attentive Relational Networks for Mapping Images to Scene Graphs

CVPR 2019poster

Scene graph generation refers to the task of automatically mapping an image into a semantic structural graph, which requires correctly labeling each extracted object and their interaction relationships. Despite the recent success in object detection using deep learning techniques, inferring complex…

Cited by 199PDFScholar
2019

KE-GAN: Knowledge Embedded Generative Adversarial Networks for Semi-Supervised Scene Parsing

CVPR 2019poster

In recent years, scene parsing has captured increasing attention in computer vision. Previous works have demonstrated promising performance in this task. However, they mainly utilize holistic features, whilst neglecting the rich semantic knowledge and inter-object relationships in the scene. In addi…

Cited by 59PDFScholar
2018

stagNet: An Attentive Semantic RNN for Group Activity Recognition

ECCV 2018poster

Group activity recognition plays a fundamental role in a variety of applications, e.g. sports video analysis and intelligent surveillance. How to model the spatio-temporal contextual information in a scene still remains a crucial yet challenging issue. We propose a novel attentive semantic recurrent…

Cited by 180SourcePDFScholar