← Search

Hui Zhao

16 accepted papers

2025

Difference Bonds Consistency and Complementarity to Enhance Multimodal Representation Learning

ICASSP 2025accepted

In the field of multimodal representation learning, existing research has primarily focused on exploring modal consistency and modal complementarity, while overlooking the positive role of modal difference. Moreover, modal difference establishes a bonding relationship between modal consistency and m…

Cited by 0SourceScholar
2025

Efficient Modeling and Low Complexity Implementation of Rate Estimation in Versatile Video Coding

ICASSP 2025accepted

In Versatile Video Coding (VVC), Rate-Distortion Optimized Quantization (RDOQ) is a widely adopted technique to strike a balance between bit rate and distortion. However, the computational complexity introduced by RDOQ poses significant challenges for real-time applications. To address this issue, w…

Cited by 0SourceScholar
2025

Error Classification of Large Language Models on Math Word Problems: A Dynamically Adaptive Framework

EMNLP 2025

Large Language Models (LLMs) have demonstrated remarkable capabilities across various domains. Math Word Problems (MWPs) serve as a crucial benchmark for evaluating LLMs’ reasoning abilities. While most research primarily focuses on improving accuracy, it often neglects understanding and addressing

Cited by 0SourcePDFScholar
2025

Incorporate Global Information from Entire Datasets for Knowledge Tracing via Mini-Batch Input

ICASSP 2025accepted

Knowledge tracing uses students’ answer record data and the relationship between exercises to predict students’ future answering performance. However, in the deep learning model, the input manner of mini-batch may prevent the network from learning global information such as the difficulty of exercis…

Cited by 0SourceScholar
2025

Online Estimation of Table-Top Grown Strawberry Mass in Field Conditions with Occlusions

IROS 2025

Accurate mass estimation of table-top grown strawberries under field conditions remains challenging due to frequent occlusions and pose variations. This study proposes a vision-based pipeline integrating RGB-D sensing and deep learning to enable non-destructive, real-time and online mass estimation.

Cited by 0SourceScholar
2025

Rethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual Disparities

ICCV 2025poster

Recent Vision-and-Language Navigation (VLN) advancements are promising, but their idealized assumptions about robot movement and control fail to reflect physically embodied deployment challenges. To bridge this gap, we introduce VLN-PE, a physically realistic VLN platform supporting humanoid, quadru…

2025

Two Challenges, One Solution: Robust Multimodal Learning through Dynamic Modality Recognition and Enhancement

EMNLP 2025

Multimodal machine learning is often hindered by two critical challenges: modality missingness and modality imbalance. These challenges significantly degrade the performance of multimodal models. The majority of existing methods either require the availability of full-modality data during the traini

Cited by 0SourcePDFScholar
2024

A Novel Wide-Area Multiobject Detection System with High-Probability Region Searching

ICRA 2024poster

In recent years, wide-area visual surveillance systems have been widely applied in various industrial and transportation scenarios. These systems, however, face significant challenges when implementing multi-object detection due to conflicts arising from the need for high-resolution imaging, efficie…

Cited by 5SourceScholar
2024

Benchmarking Hallucination in Large Language Models Based on Unanswerable Math Word Problem

COLING 2024main

Large language models (LLMs) are highly effective in various natural language processing (NLP) tasks. However, they are susceptible to producing unreliable conjectures in ambiguous contexts called hallucination. This paper presents a new method for evaluating LLM hallucination in Question Answering…

2024

Better Late Than Never: Model-Agnostic Hallucination Post-Processing Framework Towards Clinical Text Summarization

ACL 2024findings

Clinical text summarization has proven successful in generating concise and coherent summaries. However, these summaries may include unintended text with hallucinations, which can mislead clinicians and patients. Existing methods for mitigating hallucinations can be categorized into task-specific an…

2024

BotanicGarden: A High-Quality Dataset for Robot Navigation in Unstructured Natural Environments

RA-L 2024

The rapid developments of mobile robotics and autonomous navigation over the years are largely empowered by public datasets for testing and upgrading, such as sensor odometry and SLAM tasks. Impressive demos and benchmark scores have arisen, which may suggest the maturity of existing navigation tech

Cited by 60SourcecodeScholar
2024

RSED: Zero-Shot Relation Triplet Extraction via Relation Selection and Entity Boundary Detection

ICASSP 2024accepted

Zero-shot relation triplet extraction (ZeroRTE) aims to extract relation triplets of unseen relation types from unstructured texts, with a core challenge of training models to recognize new relations without labeled data. The seminal work handles this task by leveraging pre-trained language models t…

Cited by 0SourceScholar
2024

Think Before You Act: A Two-Stage Framework for Mitigating Gender Bias Towards Vision-Language Tasks

NAACL 2024long

Gender bias in vision-language models (VLMs) can reinforce harmful stereotypes and discrimination. In this paper, we focus on mitigating gender bias towards vision-language tasks. We identify object hallucination as the essence of gender bias in VLMs. Existing VLMs tend to focus on salient or famili…