← Search

Dan Luo

16 accepted papers

2026

DART: Navigating Last-Mile Heterogeneity in Instant Delivery via Distribution-Adaptive Splines

IJCAI 2026

On-demand delivery platforms rely on Travel Time Estimation (TTE) to balance courier earnings and overdue risks. In collaboration with one of China's largest platforms, we address a critical "Fairness Gap" in TTE: current systems fail to capture complex delivery patterns in GNSS-denied environments,

Cited by 0Scholar
2026

PAGER: Proactive Monitoring Agent for Enterprise AI Assistant

AAAI 2026technical

We present a Proactive Monitoring Agent designed for large-scale customer data platforms, such as Adobe Experience Platform (AEP), to predict and prevent workflow disruptions before they impact business operations. Unlike existing reactive solutions that assist engineers only after failures occur, o

Cited by 0SourcePDFScholar
2026

SAFE: Semantic- and Frequency-Enhanced Curriculum for Cross-Domain Deepfake Detection

AAAI 2026technical

Driven by advances in GANs and diffusion models, deepfake content has reached an unprecedented level of photorealism, causing detectors to deteriorate once they leave their training domain. Most prior studies adopt CLIP as the backbone of an image-level binary classifier, yet overlook CLIP’s core st

Cited by 0SourcePDFScholar
2026

ST-DiffPlanner: A Safety-Enhanced Topology-Aware Diffusion Planner for Global Path Planning

ICRA 2026poster

In complex environments, traditional path planning methods rely on manually defined models, requiring tedious adjustments under varying scenarios or constraints. They also suffer from unstable time overhead and exponentially increasing computational costs as environmental complexity grows. Deep lear…

Cited by 0Scholar
2025

Federated Retrieval Augmented Generation for Multi-Product Question Answering

COLING 2025industry

Recent advancements in Large Language Models and Retrieval-Augmented Generation have boosted interest in domain-specific question-answering for enterprise products. However, AI Assistants often face challenges in multi-product QA settings, requiring accurate responses across diverse domains. Existin…

Cited by 3SourcePDFScholar
2025

Flow Matching Based Sequential Recommender Model

IJCAI 2025

Generative models, particularly diffusion model, have emerged as powerful tools for sequential recommendation. However, accurately modeling user preferences remains challenging due to the noise perturbations inherent in the forward and reverse processes of diffusion-based methods. Towards this end,

2025

How Does Topology Bias Distort Message Passing in Graph Recommender? A Dirichlet Energy Perspective

NeurIPS 2025poster

Graph-based recommender systems have achieved remarkable effectiveness by modeling high-order interactions between users and items. However, such approaches are significantly undermined by popularity bias, which distorts the interaction graph’s structure—referred to as topology bias. This leads to o…

Cited by 0SourcecodeScholar
2025

LOG-SLAM: Large-Scale Outdoor Gaussian SLAM for Dense Mapping and Loop Closure in Kilometer-Scale Scene Reconstruction

IROS 2025

The success of 3D Gaussian splatting in 3D reconstruction has recently led to efforts to integrate it with SLAM systems. However, most existing research has focused on indoor tracking and mapping, while outdoor Gaussian SLAM methods still heavily rely expensive LiDAR sensor. To address these challen

Cited by 0SourceScholar
2025

Mitigating Language Confusion through Inference-time Intervention

COLING 2025main

Although large language models (LLMs) trained on extensive multilingual corpora exhibit impressive language transfer, they often fail to respond in the user’s desired language due to corpus imbalances, an embarrassingly simple problem known as the language confusion. However, existing solutions like…

2025

Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis

ICASSP 2025accepted

Our research introduces an innovative framework for video-to-audio synthesis, which solves the problems of audio-video desynchronization and semantic loss in the audio. By incorporating a semantic alignment adapter and a temporal synchronization adapter, our method significantly improves semantic in…

Cited by 0SourceScholar
2025

SDD-SLAM: Semantic-Driven Dynamic SLAM With Gaussian Splatting

RA-L 2025

Recently, significant advancements have been made in 3D Gaussian Splatting SLAM for dynamic environments. However, most existing methods primarily address active dynamic objects, such as people and vehicles, and fail to account for the impact of passive dynamic objects on localization and mapping. T

Cited by 10SourceScholar
2025

Weak-to-Strong Honesty Alignment via Learning-to-Rank Supervision

ACL 2025finding

Honest alignment refers to the ability of a language model to truthfully convey its knowledge limitations by appropriately refusing to answer questions when it lacks sufficient information. Existing solutions, such as prompt engineering and fine-tuning, face limitations: the former provides only mar…

2024

FasterVD: On Acceleration of Video Diffusion Models

IJCAI 2024poster

Equipped with Denoising Diffusion Probabilistic Models, video content generation has gained significant research interest recently. However, diffusion pipelines call for intensive computation and model storage, which poses challenges for their wide and efficient deployment. In this work, we address…

Cited by 0SourcePDFScholar
2024

Generating Stereophonic Music with Single-Stage Language Models

ICASSP 2024accepted

The recent success of audio language models (LMs) has revolutionized the field of neural music generation. Among all audio LM approaches, MusicGen has demonstrated the success of a single-stage LMs based music generation framework, without needing to train multiple LMs. Despite its promising perform…

Cited by 0SourceScholar
2024

Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts

ICASSP 2024accepted

Zero-shot text-to-speech (TTS) synthesis aims to clone any unseen speaker’s voice without adaptation parameters. By quantizing speech waveform into discrete acoustic tokens and modeling these tokens with the language model, recent language model-based TTS models show zero-shot speaker adaptation cap…

Cited by 0SourceScholar
2022

Syntax-Based Graph Matching for Knowledge Base Question Answering

ICASSP 2022accepted

Semantic parsing is a mainstream method of knowledge base question answering task that first generates a set of logical forms according to question and knowledge base (KB), and then selects the most matching one to get answers. However, existing selection methods are usually based on word-level matc…

Cited by 0SourceScholar