← Search

Yuxin Ma

11 accepted papers

2026

$\mu$pscaling small models: Principled warm starts and hyperparameter transfer

ICML 2026poster

Modern large-scale neural networks are often trained and released in multiple sizes to accommodate diverse inference budgets. To improve efficiency, recent work has explored *model upscaling*: initializing larger models from trained smaller ones in order to transfer knowledge and accelerate converge…

Cited by 0SourceScholar
2026

MAR: EFFICIENT LARGE LANGUAGE MODELS VIA MODULE-AWARE ARCHITECTURE REFINEMENT

ICASSP 2026poster

Large Language Models (LLMs) excel across diverse domains but suffer from high energy costs due to quadratic attention and dense Feed-Forward Network (FFN) operations. To address these issues, we propose Module-aware Architecture Refinement (MAR), a two-stage framework that integrates State Space Mo…

Cited by 0SourcePDFScholar
2025

Decouple to Reconstruct: High Quality UHD Restoration via Active Feature Disentanglement and Reversible Fusion

ICCV 2025poster

Ultra-high-definition (UHD) image restoration often faces computational bottlenecks and information loss due to its extremely high resolution. Existing studies based on Variational Autoencoders (VAE) improve efficiency by transferring the image restoration process from pixel space to latent space. H…

Cited by 0SourcePDFScholar
2025

ML4CFD Competition: Results and Retrospective Analysis

NeurIPS 2025poster

The integration of machine learning (ML) into the physical sciences is reshaping computational paradigms, offering the potential to accelerate demanding simulations such as computational fluid dynamics (CFD). Yet, persistent challenges in accuracy, generalization, and physical consistency hinder the…

Cited by 0SourceScholar
2025

Neural Fractional Attention Differential Equations

NeurIPS 2025poster

The integration of differential equations with neural networks has created powerful tools for modeling complex dynamics effectively across diverse machine learning applications. While standard integer-order neural ordinary differential equations (ODEs) have shown considerable success, they are limit…

Cited by 0SourcecodeScholar
2025

Nonlinear Laplacians: Tunable principal component analysis under directional prior information

NeurIPS 2025spotlight

We introduce a new family of algorithms for detecting and estimating a rank-one signal from a noisy observation under prior information about that signal's direction, focusing on examples where the signal is known to have entries biased to be positive. Given a matrix observation $\mathbf{Y}$, our al…

Cited by 0SourceScholar
2025

On Transferring Transferability: Towards a Theory for Size Generalization

NeurIPS 2025spotlight

Many modern learning tasks require models that can take inputs of varying sizes. Consequently, dimension-independent architectures have been proposed for domains where the inputs are graphs, sets, and point clouds. Recent work on graph neural networks has explored whether a model trained on low-dime…

Cited by 0SourcecodeScholar
2024

Spatial-Temporal-Decoupled Masked Pre-training for Spatiotemporal Forecasting

IJCAI 2024poster

Spatiotemporal forecasting techniques are significant for various domains such as transportation, energy, and weather. Accurate prediction of spatiotemporal series remains challenging due to the complex spatiotemporal heterogeneity. In particular, current end-to-end models are limited by input lengt…

2022

GeoRefine: Self-Supervised Online Depth Refinement for Accurate Dense Mapping

ECCV 2022poster

"We present a robust and accurate depth refinement system, named GeoRefine, for geometrically-consistent dense mapping from monocular sequences. GeoRefine consists of three modules: a hybrid SLAM module using learning-based priors, an online depth refinement module leveraging self-supervision, and a…

Cited by 11SourcePDFScholar
2021

Self-Supervised Vessel Segmentation via Adversarial Learning

ICCV 2021poster

Vessel segmentation is critically essential for diagnosinga series of diseases, e.g., coronary artery disease and retinal disease. However, annotating vessel segmentation maps of medical images is notoriously challenging due to the tiny and complex vessel structures, leading to insufficient availabl…

Cited by 61PDFcodeScholar