← Search

Xingyu Zhang

14 accepted papers

2026

Beyond ReAct: A Planner-Centric Framework for Complex Tool-Augmented LLM Reasoning

AAAI 2026technical

Existing tool-augmented large language models (LLMs) encounter significant challenges when processing complex queries. Current frameworks such as ReAct are prone to local optimization traps due to their reliance on incremental decision-making processes. To address these limitations, we propose a nov

Cited by 0SourcePDFScholar
2026

PURIFICATION BEFORE FUSION: TOWARD MASK-FREE SPEECH ENHANCEMENT FOR ROBUST AUDIO-VISUAL SPEECH RECOGNITION

ICASSP 2026poster

Audio-visual speech recognition (AVSR) typically improves recognition accuracy in noisy environments by integrating noise-immune visual cues with audio signals. Nevertheless, high-noise audio inputs are prone to introducing adverse interference into the feature fusion process. To mitigate this, rece…

Cited by 0SourcePDFScholar
2026

Toward LoRA Copyright Protection with an Authorized Dual-Watermarking Framework

IJCAI 2026

Text-to-Image (T2I) diffusion models have been widely adopted due to their strong generative capabilities, while Low-Rank Adaptation (LoRA) has emerged as an efficient mechanism for customizing these models for diverse creative and commercial applications. This trend has fostered LoRA-centric servic

Cited by 0Scholar
2025

ComDrive: Comfort-Oriented End-to-End Autonomous Driving

IROS 2025

We propose ComDrive: the first comfort-oriented end-to-end autonomous driving system to generate temporally consistent and comfortable trajectories. Recent studies have demonstrated that imitation learning-based planners and learning-based trajectory scorers can effectively generate and select safet

Cited by 14SourcecodeScholar
2025

Don't Shake the Wheel: Momentum-Aware Planning in End-to-End Autonomous Driving

CVPR 2025poster

End-to-end autonomous driving frameworks enable seamless integration of perception and planning but often rely on one-shot trajectory prediction, which may lead to unstable control and vulnerability to occlusions in single-frame perception. To address this, we propose the Momentum-Aware Driving (Mom…

2025

EchoMimicV2: Towards Striking, Simplified, and Semi-Body Human Animation

CVPR 2025poster

Recent work on human animation usually involves audio, pose, or movement maps conditions, thereby achieves vivid animation quality. However, these methods often face practical challenges due to extra control conditions, cumbersome condition injection modules, or limitation to head region driving. He…

2025

GoalFlow: Goal-Driven Flow Matching for Multimodal Trajectories Generation in End-to-End Autonomous Driving

CVPR 2025poster

We propose GoalFlow, an end-to-end autonomous driving method for generating high-quality multimodal trajectories. In autonomous driving scenarios, there is rarely a single suitable trajectory. Recent methods have increasingly focused on modeling multimodal trajectory distributions. However, they suf…

2025

Igniting Creative Writing in Small Language Models: LLM-as-a-Judge versus Multi-Agent Refined Rewards

EMNLP 2025

Large Language Models (LLMs) have demonstrated remarkable creative writing capabilities, yet their substantial computational demands hinder widespread use. Enhancing Small Language Models (SLMs) offers a promising alternative, but current methods like Supervised Fine-Tuning (SFT) struggle with novel

Cited by 0SourcePDFScholar
2025

Learning Invariant Causal Mechanism from Vision-Language Models

ICML 2025poster

Contrastive Language-Image Pretraining (CLIP) has achieved remarkable success, but its performance can degrade when fine-tuned in out-of-distribution (OOD) scenarios. We model the prediction process using a Structural Causal Model (SCM) and show that the causal mechanism involving both invariant and…

Cited by 0SourcePDFScholar
2025

LipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech Recognition

ICASSP 2025accepted

Visual speech recognition (VSR), commonly known as lip reading, has garnered significant attention due to its wide-ranging practical applications. The advent of deep learning techniques and advancements in hardware capabilities have significantly enhanced the performance of lip reading models. Despi…

Cited by 0SourceScholar
2025

OccRWKV: Rethinking Efficient 3D Semantic Occupancy Prediction with Linear Complexity

ICRA 2025

3D semantic occupancy prediction networks have demonstrated remarkable capabilities in reconstructing the geometric and semantic structure of 3D scenes, providing crucial information for robot navigation and autonomous driving systems. However, due to their large overhead from dense network structur

Cited by 10SourcecodeScholar
2025

Towards Fair Graph Learning without Demographic Information

AISTATS 2025poster

Fair Graph Neural Networks (GNNs) have been extensively studied in graph-based applications. However, most approaches to fair GNNs assume the full availability of demographic information by default, which is often unrealistic due to legal restrictions or privacy concerns, leaving a noticeable gap in…

Cited by 0SourceScholar
2024

Landmark-Guided Cross-Speaker Lip Reading with Mutual Information Regularization

COLING 2024main

Lip reading, the process of interpreting silent speech from visual lip movements, has gained rising attention for its wide range of realistic applications. Deep learning approaches greatly improve current lip reading systems. However, lip reading in cross-speaker scenarios where the speaker identity…

Cited by 1SourcePDFScholar
2018

A Reliable Video Storage Architecture in Hybrid SLC/MLC Nand Flash

ICASSP 2018accepted

In this paper, we propose a reliable video storage architecture in hybrid SLC/MLC storage systems. In this architecture, the video stream is reconstructed as the key cluster and the non-key cluster according to the importance of video restoration. The key cluster is stored in SLC blocks to ensure th…

Cited by 0SourceScholar