← Search

Zeyu Yang

14 accepted papers

2026

SASST: Leveraging Syntax-Aware Chunking and LLMs for Simultaneous Speech Translation

AAAI 2026technical

This work proposes a grammar-based chunking strategy that segments input streams into semantically complete units by parsing dependency relations (e.g., noun phrase boundaries, verb-object structures) and punctuation features. The method ensures chunk coherence and minimizes semantic fragmentation.

Cited by 0SourcePDFScholar
2026

Scout Before You Attend: Sketch-and-Walk Sparse Attention for Efficient LLM Inference

ICML 2026poster

Self-attention dominates the computational and memory cost of long-context LLM inference across both prefill and decode phases. To address this challenge, we introduce **Sketch\&Walk** Attention, a training-free sparse attention method that determines sparsity with lightweight sketches and determini…

Cited by 0SourceScholar
2026

To Compress or Not? Pushing the Frontier of Lossless GenAI Model Weights Compression with Exponent Concentration

ICLR 2026poster

The scaling of Generative AI (GenAI) models into the hundreds of billions of parameters makes low-precision computation indispensable for efficient deployment. We argue that the fundamental solution lies in developing low-precision \emph{floating-point} formats, which inherently provide numerical st…

Cited by 0SourcecodeScholar
2025

Diffusion$^2$: Dynamic 3D Content Generation via Score Composition of Video and Multi-view Diffusion Models

ICLR 2025poster

Recent advancements in 3D generation are predominantly propelled by improvements in 3D-aware image diffusion models. These models are pretrained on Internet-scale image data and fine-tuned on massive 3D data, offering the capability of producing highly consistent multi-view images. However, due to t…

2025

LiDARDustX: A LiDAR Dataset for Dusty Unstructured Road Environments

ICRA 2025

Autonomous driving datasets are essential for validating the progress of intelligent vehicle algorithms, which include localization, perception, and prediction. However, existing datasets are predominantly focused on structured urban environments, which limits the exploration of unstructured and spe

Cited by 1SourcecodeScholar
2025

The Hidden Joules: Evaluating the Energy Consumption of Vision Backbones for Progress Towards More Efficient Model Inference

ICML 2025poster

Deep learning has achieved significant success but poses increasing concerns about energy consumption and sustainability. Despite these concerns, there is a lack of understanding of their energy efficiency during inference. In this study, we conduct a comprehensive analysis of the inference energy c…

Cited by 0SourcePDFScholar
2025

“Stupid robot, I want to speak to a human!” User Frustration Detection in Task-Oriented Dialog Systems

COLING 2025industry

Detecting user frustration in modern-day task-oriented dialog (TOD) systems is imperative for maintaining overall user satisfaction, engagement, and retention. However, most recent research is focused on sentiment and emotion detection in academic settings, thus failing to fully encapsulate implicat…

Cited by 0SourcePDFScholar
2024

ESVC: Combining Adaptive Style Fusion and Multi-Level Feature Disentanglement for Expressive Singing Voice Conversion

ICASSP 2024accepted

Nowadays, singing voice conversion (SVC) has made great strides in both naturalness and similarity for common SVC with a neutral expression. However, besides singer identity, emotional expression is also essential to convey the singer’s emotions and attitudes, but current SVC systems can not effecti…

Cited by 0SourceScholar
2024

Real-time Photorealistic Dynamic Scene Representation and Rendering with 4D Gaussian Splatting

ICLR 2024poster

Reconstructing dynamic 3D scenes from 2D images and generating diverse views over time is challenging due to scene complexity and temporal dynamics. Despite advancements in neural implicit models, limitations persist: (i) Inadequate Scene Structure: Existing methods struggle to reveal the spatial an…

2024

WoVoGen: World Volume-aware Diffusion for Controllable Multi-camera Driving Scene Generation

ECCV 2024poster

"Generating multi-camera street-view videos is critical for augmenting autonomous driving datasets, addressing the urgent demand for extensive and varied data. Due to the limitations in diversity and challenges in handling lighting conditions, traditional rendering-based methods are increasingly bei…

2022

DeepInteraction: 3D Object Detection via Modality Interaction

NeurIPS 2022accept

Existing top-performance 3D object detectors typically rely on the multi-modal fusion strategy. This design is however fundamentally restricted due to overlooking the modality-specific useful information and finally hampering the model performance. To address this limitation, in this work we introdu…

2022

Instinctive Real-time sEMG-based Control of Prosthetic Hand with Reduced Data Acquisition and Embedded Deep Learning Training

ICRA 2022poster

Achieving instinctive multi-grasp control of prosthetic hands typically still requires a large number of sensors, such as electromyography (EMG) electrodes mounted on a residual limb, that can be costly and time consuming to position, with their signals difficult to classify. Deep-learning-based EMG…

Cited by 14SourceScholar
2022

Virtual Reality Pre-Prosthetic Hand Training With Physics Simulation and Robotic Force Interaction

RA-L 2022

Virtual reality (VR) rehabilitation systems have been proposed to enable prosthetic hand users to perform training before receiving their prosthesis. Improving pre-prosthetic training to be more representative and better prepare the patient for prosthesis use is a crucial step forwards in rehabilita

Cited by 19SourceScholar