← Search

Jiawei Zhu

6 accepted papers

2025

Enhancing Incomplete Multimodal Learning via Modal Complementary Recovering

ICASSP 2025accepted

Multimodal learning presents significant challenges arising from the unpredictable absence of modalities during both training and testing phases. Existing recovery methods struggle to leverage the available data, which can introduce additional noise during the recovery process and degrade performanc…

Cited by 0SourceScholar
2025

OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations

CVPR 2025poster

Document content extraction is a critical task in computer vision, underpinning the data needs of large language models (LLMs) and retrieval-augmented generation (RAG) systems. Despite recent progress, current document parsing methods have not been fairly and comprehensively evaluated due to the nar…

2025

Towards Robust Visual Question Answering via Prompt-Driven Geometric Harmonization

AAAI 2025technical

Visual Question Answering (VQA) has garnered significant attention as a crucial link between vision and language, aimed at generating accurate responses to visual queries. However, current VQA models still struggle with the challenges of minority class collapse and spurious semantic correlations pos…

Cited by 0SourcePDFScholar
2024

Biomimetic Crawling Robot Based on Dielectric Elastomer: Design, Modeling and Experiment

RA-L 2024

In this study, a bio-inspired crawling robot with multi-surface locomotion capability is developed. The robot is driven by Dielectric Elastomer Minimum Energy Structures (DEMES) and utilizes a three-dimensional scissor mechanism and electrostatic adhesion technology to achieve multi-surface crawling

Cited by 3SourceScholar
2024

Development of Negative-Pressure Artificial Muscles With Fiber Constraints and Pre-Stretched Soft Skin

RA-L 2024

Negative-pressure artificial muscles based on internal support and flexible skin provide a way to develop high-performance artificial muscles. However, the random and disordered skin wrinkles formed during the contraction may lead to uncertainty in the actuation behavior, and the flexible but inexte

Cited by 3SourceScholar
2024

Medical Vision-Language Representation Learning with Cross-Modal Multi-Teacher Contrastive Distillation

ICASSP 2024accepted

Medical vision-language representation learning has garnered considerable attention owing to its applicability to extracting generic representations from the image and text modality. However, it still remains challenging to acquire a more comprehensive understanding of intra- and inter-modal semanti…

Cited by 0SourceScholar