← Search

Linjun Li

17 accepted papers

2025

PACHAT: Persona-Aware Speech Assistant for Multi-party Dialogue

EMNLP 2025

Extensive research on LLM-based spoken dialogue systems has significantly advanced the development of intelligent voice assistants. However, the integration of role information within speech remains an underexplored area, limiting its application in real-world scenarios, particularly in multi-party

Cited by 0SourcePDFScholar
2024

Rethinking the Multimodal Correlation of Multimodal Sequential Learning via Generalizable Attentional Results Alignment

ACL 2024long

Transformer-based methods have gone mainstream in multimodal sequential learning. The intra and inter modality interactions are captured by the query-key associations of multi-head attention. In this way, the calculated multimodal contexts (attentional results) are expected to be relevant to the que…

Cited by 3SourcePDFScholar
2024

TransFace: Unit-Based Audio-Visual Speech Synthesizer for Talking Head Translation

ACL 2024findings

Direct speech-to-speech translation achieves high-quality results through the introduction of discrete units obtained from self-supervised learning. However, talking head translation, converting audio-visual speech (i.e., talking head video) from one language into another, still confronts several ch…

2023

3DRP-Net: 3D Relative Position-aware Network for 3D Visual Grounding

EMNLP 2023long main

3D visual grounding aims to localize the target object in a 3D point cloud by a free-form language description. Typically, the sentences describing the target object tend to provide information about its relative relation between other objects and its position within the whole scene. In this work, w…

Cited by 0SourceScholar
2023

AV-TranSpeech: Audio-Visual Robust Speech-to-Speech Translation

ACL 2023long

Direct speech-to-speech translation (S2ST) aims to convert speech from one language into another, and has demonstrated significant progress to date. Despite the recent success, current S2ST models still suffer from distinct degradation in noisy environments and fail to translate visual speech (i.e.,…

2023

Connecting Multi-modal Contrastive Representations

NeurIPS 2023poster

Multi-modal Contrastive Representation (MCR) learning aims to encode different modalities into a semantically aligned shared space. This paradigm shows remarkable generalization ability on numerous downstream tasks across various modalities. However, the reliance on massive high-quality data pairs l…

2023

Contrastive Token-Wise Meta-Learning for Unseen Performer Visual Temporal-Aligned Translation

ACL 2023findings

Visual temporal-aligned translation aims to transform the visual sequence into natural words, including important applicable tasks such as lipreading and fingerspelling recognition. However, various performance habits of specific words by different speakers or signers can lead to visual ambiguity, w…

Cited by 6SourcePDFScholar
2023

Distilling Coarse-to-Fine Semantic Matching Knowledge for Weakly Supervised 3D Visual Grounding

ICCV 2023poster

3D visual grounding involves finding a target object in a 3D scene that corresponds to a given sentence query. Although many approaches have been proposed and achieved impressive performance, they all require dense object-sentence pair annotations in 3D point clouds, which are both time-consuming an…

Cited by 19PDFcodeScholar
2023

Exploring Group Video Captioning with Efficient Relational Approximation

ICCV 2023poster

Current video captioning efforts most focus on describing a single video while the need for captioning videos in groups has increased considerably. In this study, we propose a new task, group video captioning, which aims to infer the desired content among a group of target videos and describe it wit…

Cited by 15PDFScholar
2023

MixSpeech: Cross-Modality Self-Learning with Audio-Visual Stream Mixup for Visual Speech Translation and Recognition

ICCV 2023poster

Multi-media communications facilitate global interaction among people. However, despite researchers exploring cross-lingual translation techniques such as machine translation and audio speech translation to overcome language barriers, there is still a shortage of cross-lingual studies on visual spee…

Cited by 25PDFcodeScholar
2023

OpenSR: Open-Modality Speech Recognition via Maintaining Multi-Modality Alignment

ACL 2023long

Speech Recognition builds a bridge between the multimedia streaming (audio-only, visual-only or audio-visual) and the corresponding text transcription. However, when training the specific model of new domain, it often gets stuck in the lack of new-domain utterances, especially the labeled visual utt…

2023

Semantic-conditioned Dual Adaptation for Cross-domain Query-based Visual Segmentation

ACL 2023findings

Visual segmentation from language queries has attracted significant research interest. Despite the effectiveness, existing works require expensive labeling and suffer severe degradation when deployed to an unseen domain. In this paper, we investigate a novel task Cross-domain Query-based Visual Segm…

2023

TAVT: Towards Transferable Audio-Visual Text Generation

ACL 2023long

Audio-visual text generation aims to understand multi-modality contents and translate them into texts. Although various transfer learning techniques of text generation have been proposed, they focused on uni-modal analysis (e.g. text-to-text, visual-to-text) and lack consideration of multi-modal con…

Cited by 17SourcePDFScholar
2023

Weakly-Supervised Spoken Video Grounding via Semantic Interaction Learning

ACL 2023long

The task of spoken video grounding aims to localize moments in videos that are relevant to descriptive spoken queries. However, extracting semantic information from speech and modeling the cross-modal correlation pose two critical challenges. Previous studies solve them by representing spoken querie…

2021

MPC-MPNet: Model-Predictive Motion Planning Networks for Fast, Near-Optimal Planning Under Kinodynamic Constraints

RA-L 2021

Kinodynamic Motion Planning (KMP) is to find a robot motion subject to concurrent kinematics and dynamics constraints. To date, quite a few methods solve KMP problems and those that exist struggle to find near-optimal solutions and exhibit high computational complexity as the planning space dimensio

Cited by 59SourceScholar
2020

Dynamically Constrained Motion Planning Networks for Non-Holonomic Robots

IROS 2020poster

Reliable real-time planning for robots is essential in today's rapidly expanding automated ecosystem. In such environments, traditional methods that plan by relaxing constraints become unreliable or slow-down for kinematically constrained robots. This paper describes the algorithm Dynamic Motion Pla…

Cited by 36SourceScholar
2020

SPOT: Selective Point Cloud Voting for Better Proposal in Point Cloud Object Detection

ECCV 2020poster

The sparsity of point clouds limits deep learning models on capturing long-range dependencies, which makes features extracted by the models ambiguous. In point cloud object detection, ambiguous features make it hard for detectors to locate object centers and finally lead to bad detection results. In…

Cited by 15SourcePDFScholar