← Search

Hongxiang Li

20 accepted papers

2026

DanceTogether: Generating Interactive Multi-Person Video without Identity Drifting

ICLR 2026poster

Controllable video generation (CVG) has advanced rapidly, yet current systems falter when more than one actor must move, interact, and exchange positions under noisy control signals. We address this gap with DanceTogether, the first end-to-end diffusion framework that turns a single reference image…

Cited by 0SourceScholar
2026

GIR-Bench: Versatile Benchmark for Generating Images with Reasoning

ICLR 2026poster

Unified multimodal models integrate the reasoning capacity of large language models with both image understanding and generation, showing great promise for advanced multimodal intelligence. However, the community still lacks a rigorous reasoning-centric benchmark to systematically evaluate the align…

Cited by 0SourcecodeScholar
2026

Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation

AAAI 2026technical

Large vision-language models (LVLMs) have demonstrated impressive capabilities across diverse multimodal tasks, yet they remain highly susceptible to visual hallucinations (VH), often producing confident but inaccurate descriptions of visual content. Building on the insight that not all tokens and a

Cited by 0SourcePDFScholar
2025

DisPose: Disentangling Pose Guidance for Controllable Human Image Animation

ICLR 2025poster

Controllable human image animation aims to generate videos from reference images using driving videos. Due to the limited control signals provided by sparse guidance (e.g., skeleton pose), recent works have attempted to introduce additional dense conditions (e.g., depth map) to ensure motion alignme…

2025

Towards a Multimodal Large Language Model with Pixel-Level Insight for Biomedicine

AAAI 2025technical

In recent years, Multimodal Large Language Models (MLLM) have achieved notable advancements, demonstrating the feasibility of developing an intelligent biomedical assistant. However, current biomedical MLLMs predominantly focus on image-level understanding and restrict interactions to textual comman…

2024

Aligner²: Enhancing Joint Multiple Intent Detection and Slot Filling via Adjustive and Forced Cross-Task Alignment

AAAI 2024technical

Multi-intent spoken language understanding (SLU) has garnered growing attention due to its ability to handle multiple intent utterances, which closely mirrors practical scenarios. Unlike traditional SLU, each intent in multi-intent SLU corresponds to its designated scope for slots, which occurs in…

2024

Cyclical Contrastive Learning Based on Geodesic for Zero-shot Cross-lingual Spoken Language Understanding

ACL 2024findings

Owing to the scarcity of labeled training data, Spoken Language Understanding (SLU) is still a challenging task in low-resource languages. Therefore, zero-shot cross-lingual SLU attracts more and more attention. Contrastive learning is widely applied to explicitly align representations of similar se…

2024

Exploiting Auxiliary Caption for Video Grounding

AAAI 2024technical

Video grounding aims to locate a moment of interest matching the given query sentence from an untrimmed video. Previous works ignore the sparsity dilemma in video annotations, which fails to provide the context information between potential events and query sentences in the dataset. In this paper, w…

Cited by 17SourcePDFScholar
2024

KC-Prompt: End-To-End Knowledge-Complementary Prompting for Rehearsal-Free Continual Learning

ICASSP 2024accepted

Continuous learning requires adapting quickly to incoming tasks while avoiding catastrophic forgetting. Typical solutions resort to a rehearsal buffer to replay old data, which is intractable to apply in real-world scenarios with limited memory and inaccessible privacy. Recently, with the emergence…

Cited by 0SourceScholar
2024

KDProR: A Knowledge-Decoupling Probabilistic Framework for Video-Text Retrieval

ECCV 2024poster

"Existing video-text retrieval methods predominantly focus on designing diverse cross-modal interaction mechanisms between captions and videos. However, those approaches diverge from human learning paradigms, where humans possess the capability to seek and associate knowledge from an open set, rathe…

Cited by 8SourcePDFScholar
2024

Soul-Mix: Enhancing Multimodal Machine Translation with Manifold Mixup

ACL 2024long

Multimodal machine translation (MMT) aims to improve the performance of machine translation with the help of visual information, which has received widespread attention recently. It has been verified that visual information brings greater performance gains when the textual information is limited. Ho…

2024

Towards Multi-Intent Spoken Language Understanding via Hierarchical Attention and Optimal Transport

AAAI 2024technical

Multi-Intent spoken language understanding (SLU) can handle complicated utterances expressing multiple intents, which has attracted increasing attention from researchers. Although existing models have achieved promising performance, most of them still suffer from two leading problems: (1) each inten…

2024

Uncertainty-aware sign language video retrieval with probability distribution modeling

ECCV 2024poster

"Sign language video retrieval plays a key role in facilitating information access for the deaf community. Despite significant advances in video-text retrieval, the complexity and inherent uncertainty of sign language preclude direct applications of these techniques. Previous methods achieve mapping…

2023

Accelerating Multiple Intent Detection and Slot Filling via Targeted Knowledge Distillation

EMNLP 2023long findings

Recent non-autoregressive Spoken Language Understanding (SLU) models attracts increasing attention owing to the high inference speed. However, most of them still (1) suffer from the multi-modality problem since the prior knowledge about the reference is relatively poor during inference; (2) fail to…

Cited by 0SourceScholar
2023

G2L: Semantically Aligned and Uniform Video Grounding via Geodesic and Game Theory

ICCV 2023oral

The recent video grounding works attempt to introduce vanilla contrastive learning into video grounding. However, we claim that this naive solution is suboptimal. Contrastive learning requires two key properties: (1) alignment of features of similar samples, and (2) uniformity of the induced distrib…

Cited by 52PDFScholar
2023

ML-LMCL: Mutual Learning and Large-Margin Contrastive Learning for Improving ASR Robustness in Spoken Language Understanding

ACL 2023findings

Spoken language understanding (SLU) is a fundamental task in the task-oriented dialogue systems. However, the inevitable errors from automatic speech recognition (ASR) usually impair the understanding performance and lead to error propagation. Although there are some attempts to address this problem…

2023

SSVMR: Saliency-Based Self-Training for Video-Music Retrieval

ICASSP 2023accepted

With the rise of short videos, the demand for selecting appropriate background music (BGM) for a video has increased significantly, video-music retrieval (VMR) task gradually draws much attention by research community. As other cross-modal learning tasks, existing VMR approaches usually attempt to m…

Cited by 0SourceScholar
2023

Unify, Align and Refine: Multi-Level Semantic Alignment for Radiology Report Generation

ICCV 2023poster

Automatic radiology report generation has attracted enormous research interest due to its practical value in reducing the workload of radiologists. However, simultaneously establishing global correspondences between the image (e.g., Chest X-ray) and its related report and local alignments between im…

Cited by 42PDFScholar
2017

Hawkeye: Open source framework for field surveillance

IROS 2017poster

This paper introduces a generic framework for field surveillance using consumer rotorcrafts and ground vehicles. Building such an autonomous system comes with two key challenges in persistent perception and obstacle avoidance. We begin with explaining two core algorithms to solve the challenges: an…

Cited by 8SourceScholar