← Search

Yu Yuan

13 accepted papers

2026

GeoGramBench: Benchmarking the Geometric Program Reasoning in Modern LLMs

ICLR 2026poster

Geometric spatial reasoning forms the foundation of many applications in artificial intelligence, yet the ability of large language models (LLMs) to operate over geometric spatial information expressed in procedural code remains underexplored. In this paper, we address this gap by formalizing the \t…

Cited by 0SourcecodeScholar
2026

Navigating the Pareto Frontier of Alignment:Spectrum-Adaptive Fine-Tuning for LLMs

ICML 2026poster

Supervised Fine-Tuning (SFT) with Negative Log-Likelihood (NLL) remains the standard post-training paradigm for Large Language Models, yet it imposes an excessive penalty on low-probability target tokens. This focus forces the model to prioritize minimizing the loss of difficult samples over optimiz…

Cited by 0SourceScholar
2026

NewtonGen: Physics-consistent and Controllable Text-to-Video Generation via Neural Newtonian Dynamics

ICLR 2026poster

A primary bottleneck in large-scale text-to-video generation today is physical consistency and controllability. Despite recent advances, state-of-the-art models often produce unrealistic motions, such as objects falling upward, or abrupt changes in velocity and direction. Moreover, these models lack…

Cited by 0SourcecodeScholar
2026

SeeU: Seeing the Unseen World via 4D Dynamics-aware Generation

CVPR 2026

Images and videos are discrete 2D projections of the 4D world (3D space + time). Most visual understanding, prediction, and generation operate directly on 2D observations, leading to suboptimal performance. We propose SeeU, a novel approach that learns the continuous 4D dynamics and generate the uns

Cited by 0SourcecodeScholar
2025

CAD-Editor: A Locate-then-Infill Framework with Automated Training Data Synthesis for Text-Based CAD Editing

ICML 2025poster

Computer Aided Design (CAD) is indispensable across various industries. \emph{Text-based CAD editing}, which automates the modification of CAD models based on textual instructions, holds great potential but remains underexplored. Existing methods primarily focus on design variation generation or te…

Cited by 0SourcePDFScholar
2025

Generative Photography: Scene-Consistent Camera Control for Realistic Text-to-Image Synthesis

CVPR 2025highlight

Image generation today can produce somewhat realistic images from text prompts. However, if one asks the generator to synthesize a specific camera setting such as creating different fields of view using a 24mm lens versus a 70mm lens, the generator will not be able to interpret and generate scene-co…

2025

Learning Phase Distortion with Selective State Space Models for Video Turbulence Mitigation

CVPR 2025highlight

Atmospheric turbulence is a major source of image degradation in long-range imaging systems. Although numerous deep learning-based turbulence mitigation (TM) methods have been proposed, many are slow, memory-hungry, and do not generalize well. In the spatial domain, methods based on convolutional op…

2025

LlmFixer: Fix the Helpfulness of Defensive Large Language Models

EMNLP 2025

Defense strategies of large language models besides alignment are introduced to defend against jailbreak attacks, and they have managed to decrease the success rate of jailbreak attacks. However, these defense strategies weakened the helpfulness of large language models. In this work, we propose a u

Cited by 0SourcePDFScholar
2025

S3R-GS: Streamlining the Pipeline for Large-Scale Street Scene Reconstruction

ICCV 2025poster

Recently, 3D Gaussian Splatting (3DGS) has reshaped the field of photorealistic 3D reconstruction, achieving impressive rendering quality and speed. However, when applied to large-scale street scenes, existing methods suffer from rapidly escalating per-viewpoint reconstruction costs as scene size in…

2025

Taught Well Learned Ill: Towards Distillation-conditional Backdoor Attack

NeurIPS 2025poster

Knowledge distillation (KD) is a vital technique for deploying deep neural networks (DNNs) on resource-constrained devices by transferring knowledge from large teacher models to lightweight student models. While teacher models from third-party platforms may undergo security verification (e.g., backd…

Cited by 0SourcecodeScholar
2025

Text-to-CAD Generation Through Infusing Visual Feedback in Large Language Models

ICML 2025poster

Creating Computer-Aided Design (CAD) models requires significant expertise and effort. Text-to-CAD, which converts textual descriptions into CAD parametric sequences, is crucial in streamlining this process. Recent studies have utilized ground-truth parametric sequences, known as sequential signals…

Cited by 1SourcePDFScholar
2024

Do LLMs Overcome Shortcut Learning? An Evaluation of Shortcut Challenges in Large Language Models

EMNLP 2024main

Large Language Models (LLMs) have shown remarkable capabilities in various natural language processing tasks. However, LLMs may rely on dataset biases as shortcuts for prediction, which can significantly impair their robustness and generalization capabilities. This paper presents Shortcut Suite, a c…

2017

Face Album: Towards automatic photo management based on person identity on mobile phones

ICASSP 2017accepted

We implement a new photo management system `Face Album' on mobile phones, which organizes photos by person identity, as is shown in Fig. 1. We automatically group faces into clusters to release user workload. Our system is composed of two pools: a certain pool with reliable clusters consisting of fa…

Cited by 0SourceScholar