← Search

Die Hu

7 accepted papers

2026

Knowledge-Enhanced Image Captioning with Adaptive Graph-based Multimodal Alignment and LLM

AAAI 2026technical

Image captioning is crucial for multimodal understanding, bridging visual content and natural language. Despite recent advancements in Large Multimodal Models (LMMs), when faced with unseen entities or scenes in the open world, even when attempting to leverage learned knowledge, models still struggl

Cited by 0SourcePDFScholar
2026

Orion: Steering Personalized Web Agents via Global-Micro Profiling and Adaptive Intent Tracking

AAAI 2026technical

Recently, Large Language Models (LLMs) based Web Agents have shown significant potential in web understanding and interaction tasks. However, their personalization ability and user experience remain limited by the ambiguity and dynamic nature of user intent, struggling to model diverse user interest

Cited by 0SourcePDFScholar
2025

Emotion-aware Structural Enhancement Graph Auto-Encoder for Rumor Detection

ICASSP 2025accepted

Social media is a key channel for information dissemination, making effective rumor detection essential to mitigate misinformation’s societal impact. Although large language models excel in inference and text generation, they struggle with understanding propagation relationships and complex reasonin…

Cited by 0SourceScholar
2025

Semantic Reshuffling with LLM and Heterogeneous Graph Auto-Encoder for Enhanced Rumor Detection

COLING 2025main

Social media is crucial for information spread, necessitating effective rumor detection to curb misinformation’s societal effects. Current methods struggle against complex propagation influenced by bots, coordinated accounts, and echo chambers, which fragment information and increase risks of misjud…

Cited by 1SourcePDFScholar
2025

WebSurfer: Enhancing LLM Agents with Web-Wise Feedback for Web Navigation

ICASSP 2025accepted

As the Internet’s complexity and information volume surge, the need for efficient web automation becomes critical. Traditional web agents struggle with redundant web content, which disrupts their understanding of the environment. They also face inefficiencies in multi-task scenarios due to handcraft…

Cited by 0SourceScholar
2024

PJSCC: A Puncturing-Based Joint Source Channel Coding Scheme with Hierarchical Down-Sampling Layer

ICASSP 2024accepted

In this paper, we propose a puncturing-based joint source channel coding scheme with a hierarchical down-sampling layer (PJSCC). The proposed hierarchical down-sampling layer fully exploits both frequency and spatial priors. Moreover, to achieve adaptive compression ratio control, PJSCC utilizes a s…

Cited by 0SourceScholar
2024

Surface-Constrained Progressive Feature Preserving Point Cloud Compression

ICASSP 2024accepted

Current point cloud compression methods based on deep learning cannot guarantee that the reconstructed points are constrained to the surface, resulting in low reconstruction quality at low bitrates. Hence, this paper proposes an efficient deep learning-based point cloud geometry compression algorith…

Cited by 0SourceScholar