← Search

Yu Lu

36 accepted papers

2026

BideDPO: Conditional Image Generation with Simultaneous Text and Condition Alignment

ICLR 2026poster

Conditional image generation augments text-to-image synthesis with structural, spatial, or stylistic priors and is used in many domains. However, current methods struggle to harmonize guidance from both sources when conflicts arise: 1) input-level conflict, where the semantics of the conditioning im…

Cited by 0SourcecodeScholar
2026

CoLoGen: Progressive Learning of Concept-Localization Duality for Unified Image Generation

CVPR 2026

Unified conditional image generation remains difficult because different tasks depend on fundamentally different internal representations. Some require conceptual understanding for semantic synthesis, while others rely on localization cues for spatial precision. Forcing these heterogeneous tasks to

Cited by 3SourcecodeScholar
2026

DuPO: Enabling Reliable Self-Verification via Dual Preference Optimization

ICLR 2026poster

We present DuPO, a dual learning-based preference optimization framework that generates annotation-free feedback via the generalized duality. DuPO addresses two key limitations: Reinforcement Learning with Verifiable Rewards (RLVR)’s reliance on costly labels and applicability restricted to verifiab…

Cited by 0SourceScholar
2026

GIFSplat: Generative Prior-Guided Iterative Feed-Forward 3D Gaussian Splatting from Sparse Views

CVPR 2026

Feed-forward 3D reconstruction offers substantial runtime advantages over per-scene optimization, which remains slow at inference and often fragile under sparse views. However, existing feed-forward methods still have potential for further performance gains, especially for out-of-domain data, and st

Cited by 0SourcecodeScholar
2026

Listening Through the Noise: Cauchy-Driven Diffusion Bridges for Robust Gastrointestinal Auscultation and Clinical Benchmarking

ICML 2026spotlight

Gastrointestinal (GI) motility assessment via bowel sounds (BS) offers a non-invasive alternative to resource-intensive clinical standards. However, the diagnostic utility of BS is often compromised by its spectral overlap with non-stationary speech interference. While generative models have advance…

Cited by 0SourceScholar
2025

AMSER: Accelerate Mobile Speech Emotion Recognition with Signal Compression

ICASSP 2025accepted

Speech-based interaction systems are widely used in mobile devices like smartphones. With advances in deep neural networks, tasks such as speech emotion recognition (SER) enhance these systems’ user-friendliness. However, deploying SER models on mobile devices is challenging due to their complexity…

Cited by 0SourceScholar
2025

CTGDiff: A Conditional Diffusion Model for Cardiotocography Signal Synthesis

ICASSP 2025accepted

The analysis of Cardiotocography (CTG) signals is often hindered by challenges such as limited data availability and label imbalance, which can undermine the performance of deep learning models. To address these issues, we present CTGDiff, a novel conditional diffusion model designed for generating…

Cited by 0SourceScholar
2025

Do Large Language Models Truly Grasp Addition? A Rule-Focused Diagnostic Using Two-Integer Arithmetic

EMNLP 2025

Large language models (LLMs) achieve impressive results on advanced mathematics benchmarks but sometimes fail on basic arithmetic tasks, raising the question of whether they have truly grasped fundamental arithmetic rules or are merely relying on pattern matching. To unravel this issue, we systemati

2025

Domain Generalized Medical Landmark Detection via Robust Boundary-Aware Pre-Training

AAAI 2025technical

In recent years, deep learning has revenue in automated medical landmark detection. Nonetheless, prevailing research in this field predominantly addresses single-center scenarios or domain adaptation settings. In practical environments, the acquisition of multi-center data faces privacy concerns, co…

2025

EFCWM-Mamba-YOLO: Real-Time Underwater Object Detection with Adaptive Feature Representation and Domain Adaptation

IROS 2025

Underwater object detection (UOD) is crucial for monitoring marine ecosystems, underwater robotics, environmental protection, and autonomous underwater vehicles (AUVs). Despite progress, many models struggle under real-world conditions due to poor visibility, dynamic lighting, and domain shifts. Tra

Cited by 0SourcecodeScholar
2025

EnAnchored-X2X: English-Anchored Optimization for Many-to-Many Translation

EMNLP 2025

Large language models (LLMs) have demonstrated strong machine translation capabilities for English-centric language pairs but underperform in direct non-English (x2x) translation. This work addresses this limitation through a synthetic data generation framework that leverages models’ established Eng

2025

Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer

NeurIPS 2025poster

Instruction-based image editing enables precise modifications via natural language prompts, but existing methods face a precision-efficiency tradeoff: fine-tuning demands massive datasets (>10M) and computational resources, while training-free approaches suffer from weak instruction comprehension.…

Cited by 0SourceScholar
2025

Extended LSTMs for Knowledge Tracing: Peeking Inside the Black Box (Student Abstract)

AAAI 2025technical

This paper proposes extended Long Short-Term Memory (LSTM) networks for the knowledge tracing task and employs explainable AI methods to address interpretability issues. Specifically, we developed an extended LSTM-based model to automatically diagnose students' knowledge states. We then leveraged th…

Cited by 0SourcePDFScholar
2025

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding

NeurIPS 2025poster

Long-form video understanding poses a significant challenge for video large language models (VideoLLMs) due to prohibitively high computational and memory demands. In this paper, We propose $\textbf{FlexSelect}$, a flexible and efficient token selection strategy for processing long videos. FlexSele…

Cited by 0SourcecodeScholar
2025

HarmonySet: A Comprehensive Dataset for Understanding Video-Music Semantic Alignment and Temporal Synchronization

CVPR 2025poster

This paper introduces HarmonySet, a comprehensive dataset designed to advance video-music understanding. HarmonySet consists of 48,328 diverse video-music pairs, annotated with detailed information on rhythmic synchronization, emotional alignment, thematic coherence, and cultural relevance. We propo…

2025

LLMs + Persona-Plug = Personalized LLMs

ACL 2025long

Personalization plays a critical role in numerous language tasks and applications, since users with the same requirements may prefer diverse outputs based on their interests. This has led to the development of various personalized approaches aimed at adapting large language models (LLMs) to generate…

2025

NOVA: An Iterative Planning Framework for Enhancing Scientific Innovation with Large Language Models

ACL 2025finding

Scientific innovation is pivotal for humanity, and harnessing large language models (LLMs) to generate research ideas could transform discovery. However, existing LLMs often produce simplistic and repetitive suggestions due to their limited ability in acquiring external knowledge for innovation. To…

2024

ActiveDC: Distribution Calibration for Active Finetuning

CVPR 2024poster

The pretraining-finetuning paradigm has gained popularity in various computer vision tasks. In this paradigm the emergence of active finetuning arises due to the abundance of large-scale data and costly annotation requirements. Active finetuning involves selecting a subset of data from an unlabeled…

Cited by 2SourcePDFScholar
2024

Automated Multi-level Preference for MLLMs

NeurIPS 2024poster

Current multimodal Large Language Models (MLLMs) suffer from ''hallucination'', occasionally generating responses that are not grounded in the input images. To tackle this challenge, one promising path is to utilize reinforcement learning from human feedback (RLHF), which steers MLLMs towards learni…

2024

FreeLong: Training-Free Long Video Generation with SpectralBlend Temporal Attention

NeurIPS 2024poster

Video diffusion models have made substantial progress in various video generation applications. However, training models for long video generation tasks require significant computational and data resources, posing a challenge to developing long video diffusion models. This paper investigates a strai…

Cited by 22SourcePDFScholar
2024

G-DIG: Towards Gradient-based DIverse and hiGh-quality Instruction Data Selection for Machine Translation

ACL 2024long

Large Language Models (LLMs) have demonstrated remarkable abilities in general scenarios. Instruction finetuning empowers them to align with humans in various tasks. Nevertheless, the Diversity and Quality of the instruction data remain two main challenges for instruction finetuning. With regard to…

2024

Model AI Assignments 2024

AAAI 2024technical

The Model AI Assignments session seeks to gather and dis- seminate the best assignment designs of the Artificial In- telligence (AI) Education community. Recognizing that as- signments form the core of student learning experience, we here present abstracts of five AI assignments from the 2024 sessi…

Cited by 0SourcePDFScholar
2024

Retaining Key Information under High Compression Ratios: Query-Guided Compressor for LLMs

ACL 2024long

The growing popularity of Large Language Models has sparked interest in context compression for Large Language Models (LLMs). However, the performance of previous methods degrades dramatically as compression ratios increase, sometimes even falling to the closed-book level. This decline can be attrib…

2024

Snore Sound Features Based on Percussive Enhancing and Positional Encoding Combined with Multi-Task Learning for Osahs Detection

ICASSP 2024accepted

Obstructive sleep apnea hypopnea syndrome (OSAHS) is a serious sleep disorder. As the typical symptom of OSAHS, snoring has been proved effective in OSAHS diagnosis and potential to replace the current laborious and expensive polysomnography. However, the lack of analysis on the characteristics of p…

Cited by 0SourceScholar
2024

Understanding the Therapeutic Relationship between Counselors and Clients in Online Text-based Counseling using LLMs

EMNLP 2024finding

Robust therapeutic relationships between counselors and clients are fundamental to counseling effectiveness. The assessment of therapeutic alliance is well-established in traditional face-to-face therapy but may not directly translate to text-based settings. With millions of individuals seeking supp…

Cited by 4SourcePDFScholar
2023

AUGUST: an Automatic Generation Understudy for Synthesizing Conversational Recommendation Datasets

ACL 2023findings

High-quality data is essential for conversational recommendation systems and serves as the cornerstone of the network architecture development and training strategy design. Existing works contribute heavy human efforts to manually labeling or designing and extending recommender dialogue templates. H…

2023

Develop AI Teaching and Learning Resources for Compulsory Education in China

AAAI 2023technical

Artificial intelligence course has been required to take for compulsory education students in China. However, not all teachers and schools are fully prepared and ready. This is partially because of the lack of adequate teaching and learning resources, which requires a major expenditure of time and e…

Cited by 9SourcePDFScholar
2023

Take a Closer Look at Multilinguality! Improve Multilingual Pre-Training Using Monolingual Corpora Only

EMNLP 2023long findings

Recent studies have revealed the remarkable cross-lingual capability of multilingual pre-trained language models (mPLMs), even when pre-trained without parallel corpora (mono-mPLMs). Intuitively, semantic alignments may be the reason behind such capability but remain under-explored. In this work, we…

Cited by 0SourceScholar
2022

CRIS: CLIP-Driven Referring Image Segmentation

CVPR 2022poster

Referring image segmentation aims to segment a referent via a natural linguistic expression. Due to the distinct data properties between text and image, it is challenging for a network to well align text and pixel-level features. Existing approaches use pretrained models to facilitate learning, yet…

Cited by 441PDFcodeScholar
2022

Learning Confidence for Transformer-based Neural Machine Translation

ACL 2022long

Confidence estimation aims to quantify the confidence of the model prediction, providing an expectation of success. A well-calibrated confidence estimate enables accurate failure prediction and proper risk measurement when given noisy samples and out-of-distribution data in real-world settings. Howe…

2021

Attention Calibration for Transformer in Neural Machine Translation

ACL 2021long

Attention mechanisms have achieved substantial improvements in neural machine translation by dynamically selecting relevant inputs for different predictions. However, recent studies have questioned the attention mechanisms’ capability for discovering decisive inputs. In this paper, we propose to cal…

2021

Efficient Speech Emotion Recognition Using Multi-Scale CNN and Attention

ICASSP 2021accepted

Emotion recognition from speech is a challenging task. Recent advances in deep learning have led bi-directional recurrent neural network (Bi-RNN) and attention mechanism as a standard method for speech emotion recognition, extracting and attending multi-modal features - audio and text, and then fusi…

Cited by 0SourceScholar
2020

GINet: Graph Interaction Network for Scene Parsing

ECCV 2020poster

Recently, context reasoning using image regions beyond local convolution has shown great potential for scene parsing. In this work, we explore how to incorperate the linguistic knowledge to promote context reasoning over image regions by proposing a Graph Interaction unit (GI unit) and a Semantic Co…