← Search

Jiahui Li

22 accepted papers

2026

Adaptive Video Distillation: Mitigating Oversaturation and Temporal Collapse in Few-Step Generation

CVPR 2026

Video generation has recently emerged as a central task in the field of generative AI. However, the substantial computational cost inherent in video synthesis makes model distillation a critical technique for efficient deployment. Despite its significance, there is a scarcity of methods specifically

Cited by 0SourceScholar
2026

Beyond Reassembly: Fractured Object Recovery with Missing Parts

CVPR 2026

We propose a novel learning-based task named fractured object recovery. Unlike the previous fractured object reassembly task that only aligns existing parts with overlaps, our task aims to recover the complete shape by not only reassembling irrelevant parts but also predicting missing parts. Our tas

Cited by 0SourceScholar
2026

IMPACT: Behavioral Intention-Aware Multimodal Trajectory Prediction With Adaptive Context Trimming

RA-L 2026

This paper presents a unified framework that jointly predicts behavioral intentions and vectorized occupancy, leveraging them as priors to dynamically prune context information during trajectory decoding, thereby enhancing prediction accuracy, interpretability, and efficiency. While most prior work

Cited by 4SourceScholar
2026

IMPACT: Behavioral Intention-Aware Multimodal Trajectory Prediction with Adaptive Context Trimming

ICRA 2026poster

This paper presents a unified framework that jointly predicts behavioral intentions and vectorized occupancy, leveraging them as priors to dynamically prune context information during trajectory decoding, thereby enhancing prediction accuracy, interpretability, and efficiency. While most prior work …

2026

Path-Decoupled Hyperbolic Flow Matching for Few-Shot Adaptation

ICML 2026poster

Recent advances in cross-modal few-shot adaptation treat visual-semantic alignment as a continuous feature transport problem via Flow Matching (FM). However, we argue that Euclidean-based FM overlooks fundamental limitations of flat geometry, where polynomial volume growth fails to accommodate diver…

Cited by 0SourceScholar
2026

Relation-R1: Progressively Cognitive Chain-of-Thought Guided Reinforcement Learning for Unified Relation Comprehension

AAAI 2026technical

Recent advances in multi-modal large language models (MLLMs) have significantly improved object-level grounding and region captioning. However, they remain limited in visual relation understanding, struggling even with binary relation detection, let alone N-ary relations involving multiple semantic

Cited by 0SourcePDFScholar
2025

Exploiting the Index Gradients for Optimization-Based Jailbreaking on Large Language Models

COLING 2025main

Despite the advancements in training Large Language Models (LLMs) with alignment techniques to enhance the safety of generated content, these models remain susceptible to jailbreak, an adversarial attack method that exposes security vulnerabilities in LLMs. Notably, the Greedy Coordinate Gradient (G…

2025

Judge and Improve: Towards a Better Reasoning of Knowledge Graphs with Large Language Models

EMNLP 2025

Graph Neural Networks (GNNs) have shown immense potential in improving the performance of large-scale models by effectively incorporating structured relational information. However, current approaches face two key challenges: (1) achieving robust semantic alignment between graph representations and

Cited by 0SourcePDFScholar
2025

Learning Causal Transition Matrix for Instance-dependent Label Noise

AAAI 2025technical

Noisy labels are both inevitable and problematic in machine learning methods, as they negatively impact models' generalization ability by causing overfitting. In the context of learning with noise, the transition matrix plays a crucial role in the design of statistically consistent algorithms. Howev…

Cited by 0SourcePDFScholar
2025

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution

EMNLP 2025

Reinforcement learning from human feedback (RLHF) offers a promising approach to aligning large language models (LLMs) with human preferences. Typically, a reward model is trained or supplied to act as a proxy for humans in evaluating generated responses during the reinforcement training phase. Howe

2025

RMP-YOLO: A Robust Motion Predictor for Partially Observable Scenarios Even if You Only Look Once

ICRA 2025

We introduce RMP-YOLO, a unified framework designed to provide robust motion predictions even with incomplete input data. Our key insight stems from the observation that complete and reliable historical trajectory data plays a pivotal role in ensuring accurate motion prediction. Therefore, we propos

Cited by 7SourcecodeScholar
2025

Spatial Frequency Interleaving Residual Autoencoder for Indoor Radio Map Reconstruction

ICASSP 2025accepted

Indoor radio maps with frequency domain data are difficult to reconstruct when only limited measurements at a few locations are available. Naive convolutional neural networks suffer from flawed structures in the frequency domain when predicting these radio maps, resulting in overly smoothed predicti…

Cited by 0SourceScholar
2025

Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse Rewards

CVPR 2025poster

Diffusion models have achieved remarkable success in text-to-image generation. However, their practical applications are hindered by the misalignment between generated images and corresponding text prompts. To tackle this issue, reinforcement learning (RL) has been considered for diffusion model fin…

2025

Which Demographics do LLMs Default to During Annotation?

ACL 2025long

Demographics and cultural background of annotators influence the labels they assign in text annotation – for instance, an elderly woman might find it offensive to read a message addressed to a “bro”, but a male teenager might find it appropriate. It is therefore important to acknowledge label variat…

Cited by 0SourcePDFScholar
2024

Autoregressive 3D Shape Completion via Sphere-Guided Disentangled Representation

ICASSP 2024accepted

This paper introduces a novel 3D shape completion method based on sphere-guided disentangled representation. Utilizing an autoregressive transformer-based model, our approach efficiently constructs object completion distributions given incomplete point clouds. To enhance completion modeling, we prop…

Cited by 0SourceScholar
2024

Optimizing Language Models with Fair and Stable Reward Composition in Reinforcement Learning

EMNLP 2024main

Reinforcement learning from human feedback (RLHF) and AI-generated feedback (RLAIF) have become prominent techniques that significantly enhance the functionality of pre-trained language models (LMs). These methods harness feedback, sourced either from humans or AI, as direct rewards or to shape rewa…

2023

Two Heads are Better Than One: A Simple Exploration Framework for Efficient Multi-Agent Reinforcement Learning

NeurIPS 2023poster

Exploration strategy plays an important role in reinforcement learning, especially in sparse-reward tasks. In cooperative multi-agent reinforcement learning~(MARL), designing a suitable exploration strategy is much more challenging due to the large state space and the complex interaction among agent…

Cited by 3SourcePDFScholar
2022

A Speech-driven Sign Language Avatar Animation System for Hearing Impaired Applications

IJCAI 2022poster

Sign language is the communication language used in hearing impaired community. Recently, the research of sign language production has made great progress but still need to cope with some critical challenges. In this paper, we propose a system-level scheme and push forward the implementation of sign…

Cited by 6SourcePDFScholar
2022

Deconfounded Value Decomposition for Multi-Agent Reinforcement Learning

ICML 2022spotlight

Value decomposition (VD) methods have been widely used in cooperative multi-agent reinforcement learning (MARL), where credit assignment plays an important role in guiding the agents’ decentralized execution. In this paper, we investigate VD from a novel perspective of causal inference. We first sho…

Cited by 23SourcePDFScholar
2022

Distributed Audio-Visual Parsing Based On Multimodal Transformer and Deep Joint Source Channel Coding

ICASSP 2022accepted

Audio-visual parsing (AVP) is a newly emerged multimodal perception task, which detects and classifies audio-visual events in video. However, most existing AVP networks only use a simple attention mechanism to guide audio-visual multimodal events, and are implemented in a single end. This makes it u…

Cited by 0SourceScholar
2021

SNR-Adaptive Deep Joint Source-Channel Coding for Wireless Image Transmission

ICASSP 2021accepted

Considering the problem of joint source-channel coding (JSCC) for multi-user transmission of images over noisy channels, an autoencoder-based novel deep joint source-channel coding scheme is proposed in this paper. In the proposed JSCC scheme, the decoder can estimate the signal-to-noise ratio (SNR)…

Cited by 0SourceScholar
2018

How to Mobilize Mmwave: A Joint Beam and Channel Tracking Approach

ICASSP 2018accepted

Maintaining reliable millimeter wave (mmWave) connections to many fast-moving mobiles is a key challenge in the theory and practice of 5G systems. In this paper, we develop a new algorithm that can jointly track the beam direction and channel coefficient of mm Wave propagation paths using phased ant…

Cited by 0SourceScholar