← Search

Lu Xu

22 accepted papers

2026

DAM-VLA: A Dynamic Action Model-Based Vision-Language-Action Framework for Robot Manipulation

ICRA 2026poster

In dynamic environments such as warehouses, hospitals, and homes, robots must seamlessly transition between gross motion and precise manipulations to complete complex tasks. However, current Vision-Language-Action (VLA) frameworks, largely adapted from pre-trained Vision-Language Models (VLMs), ofte…

2026

DuPO: Enabling Reliable Self-Verification via Dual Preference Optimization

ICLR 2026poster

We present DuPO, a dual learning-based preference optimization framework that generates annotation-free feedback via the generalized duality. DuPO addresses two key limitations: Reinforcement Learning with Verifiable Rewards (RLVR)’s reliance on costly labels and applicability restricted to verifiab…

Cited by 0SourceScholar
2025

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition

EMNLP 2025

Recent advances in Automatic Speech Recognition (ASR) have been largely fueled by massive speech corpora. However, extending coverage to diverse languages with limited resources remains a formidable challenge. This paper introduces Speech Back-Translation, a a scalable pipeline that improves multili

2024

CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts

NeurIPS 2024poster

Recent advancements in Multimodal Large Language Models (LLMs) have focused primarily on scaling by increasing text-image pair data and enhancing LLMs to improve performance on multimodal tasks. However, these scaling approaches are computationally expensive and overlook the significance of efficien…

2024

HUGS: Holistic Urban 3D Scene Understanding via Gaussian Splatting

CVPR 2024poster

Holistic understanding of urban scenes based on RGB images is a challenging yet important problem. It encompasses understanding both the geometry and appearance to enable novel view synthesis parsing semantic labels and tracking moving objects. Despite considerable progress existing approaches often…

2024

Mask-Homo: Pseudo Plane Mask-Guided Unsupervised Multi-Homography Estimation

AAAI 2024technical

Homography estimation is a fundamental problem in computer vision. Previous works mainly focus on estimating either a single homography, or multiple homographies based on mesh grid division of the image. In practical scenarios, single homography is inadequate and often leads to a compromised result…

2024

Mitigating Data Scarcity in Semantic Parsing across Languages with the Multilingual Semantic Layer and its Dataset

ACL 2024findings

Data scarcity is a prevalent challenge in the era of Large Language Models (LLMs). The insatiable hunger of LLMs for large corpora becomes even more pronounced when dealing with non-English and low-resource languages. The issue is particularly exacerbated in Semantic Parsing (SP), i.e. the task of c…

2024

ProxyQA: An Alternative Framework for Evaluating Long-Form Text Generation with Large Language Models

ACL 2024long

Large Language Models (LLMs) have succeeded remarkably in understanding long-form contents. However, exploring their capability for generating long-form contents, such as reports and articles, has been relatively unexplored and inadequately assessed by existing benchmarks. The prevalent evaluation m…

2023

Class-Adaptive Self-Training for Relation Extraction with Incompletely Annotated Training Data

ACL 2023findings

Relation extraction (RE) aims to extract relations from sentences and documents. Existing relation extraction models typically rely on supervised machine learning. However, recent studies showed that many RE datasets are incompletely annotated. This is known as the false negative problem in which va…

2023

The Memory-Perturbation Equation: Understanding Model's Sensitivity to Data

NeurIPS 2023poster

Understanding model’s sensitivity to its training data is crucial but can also be challenging and costly, especially during training. To simplify such issues, we present the Memory-Perturbation Equation (MPE) which relates model's sensitivity to perturbation in its training data. Derived using Bayes…

2022

Revisiting DocRED - Addressing the False Negative Problem in Relation Extraction

EMNLP 2022main

The DocRED dataset is one of the most popular and widely used benchmarks for document-level relation extraction (RE). It adopts a recommend-revise annotation scheme so as to have a large-scale annotated dataset. However, we find that the annotation of DocRED is incomplete, i.e., false negative sampl…

2021

Efficient Deep Image Denoising via Class Specific Convolution

AAAI 2021technical

Deep neural networks have been widely used in image denoising during the past few years. Even though they achieve great success on this problem, they are computationally inefficient which makes them inappropriate to be implemented in mobile devices. In this paper, we propose an efficient deep neural…

2015

A low complexity iterative soft-decision feedback MMSE-PIC detection algorithm for massive MIMO

ICASSP 2015accepted

In MIMO applications, the minimum mean square error parallel interference cancellation (MMSE-PIC) based Soft-Input Soft-Output (SISO) detector has been widely adopted because of its low complexity and good bit error rate (BER) performance. In this paper, we firstly propose to use a Gaussian model ba…

Cited by 0SourceScholar
2015

Impulsive noise detection in PLC with smoothed L0-norm

ICASSP 2015accepted

Power-line communications (PLC) commonly employs orthogonal frequency-division multiplexing (OFDM) as the modulation technique, and impulsive noise has a significant negative impact on its performance. Using the property of null subcarriers in OFDM and the fact that impulsive noise is sparse, we for…

Cited by 0SourceScholar