← Search

Mingliang Zhai

11 accepted papers

2026

ProSoftArena: Benchmarking Hierarchical Capabilities of Multi-modal Agents in Professional Software Environments

CVPR 2026

Multi-modal agents are making rapid progress on general computer-use tasks. However, existing benchmarks remain largely confined to web browsers and rudimentary applications, failing to capture the professional software workflows that dominate real-world scientific and industrial practices. To bridg

Cited by 0SourcecodeScholar
2025

World Knowledge-Enhanced Reasoning Using Instruction-Guided Interactor in Autonomous Driving

AAAI 2025technical

The Multi-modal Large Language Models (MLLMs) with extensive world knowledge have revitalized autonomous driving, particularly in reasoning tasks within perceivable regions. However, when faced with perception-limited areas (dynamic or static occlusion regions), MLLMs struggle to effectively integra…

Cited by 2SourcePDFScholar
2024

Compositional Substitutivity of Visual Reasoning for Visual Question Answering

ECCV 2024poster

"Compositional generalization has received much attention in vision-and-language and visual reasoning recently. Substitutivity, the capability to generalize to novel compositions with synonymous primitives such as words and visual entities, is an essential factor in evaluating the compositional gene…

2024

In-Context Compositional Generalization for Large Vision-Language Models

EMNLP 2024main

Recent work has revealed that in-context learning for large language models exhibits compositional generalization capacity, which can be enhanced by selecting in-context demonstrations similar to test cases to provide contextual information. However, how to exhibit in-context compositional generaliz…

Cited by 2SourcePDFScholar
2024

Self-Supervised Multi-Scale Hierarchical Refinement Method for Joint Learning of Optical Flow and Depth

ICASSP 2024accepted

Recurrently refining the optical flow based on a single high-resolution feature demonstrates high performance. We exploit the strength of this strategy to build a novel architecture for the joint learning of optical flow and depth. Our pro-posed architecture is improved to work in the case of traini…

Cited by 0SourceScholar
2023

Cross-Modal Optical Flow Estimation via Modality Compensation and Alignment

ICASSP 2023accepted

Cross-modal optical flow estimation aims to predict motion fields between two frames collected from different modalities, recently attracting intensive attention. However, a substantial yet challenging problem is how to match images across a large modal discrepancy. In this paper, we propose a modal…

Cited by 0SourceScholar
2023

Fast-StrucTexT: An Efficient Hourglass Transformer with Modality-guided Dynamic Token Merge for Document Understanding

IJCAI 2023poster

Transformers achieve promising performance in document understanding because of their high effectiveness and still suffer from quadratic computational complexity dependency on the sequence length. General efficient transformers are challenging to be directly adapted to model document. They are unabl…

Cited by 8SourcePDFScholar
2023

Learning Scene Flow from 3d Point Clouds with Cross-Transformer and Global Motion Cues

ICASSP 2023accepted

Scene flow estimation is critical for real-world vision problems such as autonomous driving and augmented reality. Due to the popularity of 3D LiDAR sensors, scene flow estimation from 3D point clouds arouses increasing attention. Existing methods usually use a flow embedding-based layer to find cor…

Cited by 0SourceScholar
2020

Multi-Task Learning in Autonomous Driving Scenarios Via Adaptive Feature Refinement Networks

ICASSP 2020accepted

Many deep learning applications benefit from multi-task learning with several related objectives. In autonomous driving scenarios, being able to accurately infer motion and spatial information is essential for scene understanding. In this paper, we combine an adaptive feature refinement module and a…

Cited by 0SourceScholar
2019

Ad-net: Attention Guided Network for Optical Flow Estimation Using Dilated Convolution

ICASSP 2019accepted

Variational models for optical flow estimation usually define an energy function that contains prior assumptions to explore rudimentary statistics of images. However, such methods cannot learn motion knowledge from the pre-prepared data and have many parameters that need to be set manually. Nowadays…

Cited by 0SourceScholar