← Search

xiaoyu li

65 accepted papers

2026

AnyBand-Diff: A Unified Remote Sensing Image Generation and Band Repair Framework with Spectral Priors

ICML 2026poster

Existing diffusion models have made significant progress in generating realistic images. However, their direct adaptation to remote sensing imagery often disregards intrinsic physical laws. This oversight frequently leads to spectral distortion and radiometric inconsistency, severely limiting the sc…

Cited by 0SourceScholar
2026

CubeComposer: Spatio-Temporal Autoregressive 4K 360deg Video Generation from Perspective Video

CVPR 2026

Generating high-quality 360deg panoramic videos from perspective input is one of the crucial applications for virtual reality (VR), whereby high-resolution videos are especially important for immersive experience. Existing methods are constrained by computational limitations of vanilla diffusion mod

Cited by 0SourceScholar
2026

GenCompositor: Generative Video Compositing with Diffusion Transformer

ICLR 2026poster

Video compositing combines live-action footage to create video production, serving as a crucial technique in video creation and film production. Traditional pipelines require intensive labor efforts and expert collaboration, resulting in lengthy production cycles and high manpower costs. To address…

Cited by 0SourcecodeScholar
2026

HiFICL: High-Fidelity In-Context Learning for Multimodal Tasks

CVPR 2026

In-Context Learning (ICL) is a significant paradigm for Large Multimodal Models (LMMs), using a few in-context demonstrations (ICDs) for new task adaptation. However, its performance is sensitive to demonstration configurations and computationally expensive. Mathematically, the influence of these de

Cited by 0SourcecodeScholar
2026

IC-Custom: Diverse Image Customization via In-Context Learning

ICLR 2026poster

Image customization, a crucial technique for industrial media production, aims to generate content that is consistent with reference images. However, current approaches conventionally separate image customization into position-aware and position-free customization paradigms and lack a universal fram…

Cited by 0SourcecodeScholar
2026

Lost in Time? A Meta-Learning Framework for Time-Shift-Tolerant Physiological Signal Transformation

AAAI 2026technical

Translating non-invasive signals such as photoplethysmography (PPG) and ballistocardiography (BCG) into clinically meaningful signals like arterial blood pressure (ABP) is vital for continuous, low-cost healthcare monitoring. However, temporal misalignment in multimodal signal transformation impairs

Cited by 0SourcePDFScholar
2026

Micro-Macro Retrieval: Reducing Long-Form Hallucination in Large Language Models

ICLR 2026poster

Large Language Models (LLMs) achieve impressive performance across many tasks but remain prone to hallucination, especially in long-form generation where redundant retrieved contexts and lengthy reasoning chains amplify factual errors. Recent studies highlight a critical phenomenon: the closer key i…

Cited by 0SourceScholar
2026

MoCHA: Advanced Vision-Language Reasoning with MoE Connector and Hierarchical Group Attention

AAAI 2026technical

Vision large language models (VLLMs) are focusing primarily on handling complex and fine-grained visual information by incorporating advanced vision encoders and scaling up visual models. However, these approaches face high training and inference costs, as well as challenges in extracting visual det

Cited by 0SourcePDFScholar
2026

Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception

AAAI 2026technical

Spatio-temporal alignment is crucial for temporal modeling of end-to-end (E2E) perception in autonomous driving (AD), providing valuable structural and textural prior information. Existing methods typically rely on the attention mechanism to align objects across frames, simplifying the motion model

Cited by 0SourcePDFScholar
2026

STAMP: Multi-Pattern Attention-Aware Multiple Instance Learning for STAS Diagnosis in Multi-Center Histopathology Images

IJCAI 2026

Spread through air spaces (STAS) constitutes a novel invasive pattern in lung adenocarcinoma (LUAD), associated with tumor recurrence and diminished survival rates. However, large-scale STAS diagnosis in LUAD remains a labor-intensive endeavor, compounded by the propensity for oversight and misdiagn

Cited by 0Scholar
2026

Selecting Samples on Graphs: A Unified Dataset Pruning Framework for Lossless Training Acceleration

ICML 2026poster

The rapid growth of modern training datasets has significantly increased computational cost, motivating dataset pruning(DP) methods which retain only a subset of informative samples to reduce training cost. Existing pruning criteria typically rely on either intrinsic signals that assess samples inde…

Cited by 0SourceScholar
2026

ToonComposer: Streamlining Cartoon Production with Generative Post-Keyframing

ICLR 2026poster

Traditional cartoon and anime production involves keyframing, inbetweening, and colorization stages, which require intensive manual effort. Despite recent advances in AI, existing methods often handle these stages separately, leading to error accumulation and artifacts. For instance, inbetweening ap…

Cited by 0SourcecodeScholar
2026

UDCH: Unsupervised Dynamic Weighted Cluster-cooperative Hashing for Cross-modal Retreival

AAAI 2026technical

In cross-modal retrieval tasks, unsupervised hash code learning still faces key challenges, including the difficulty of modeling shared semantic structures across modalities and the inability to adaptively balance multiple supervision objectives during optimization. To address these issues, we propo

Cited by 0SourcePDFScholar
2026

VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control

CVPR 2026

Video world models aim to simulate dynamic, real-world environments, yet existing methods struggle to provide unified and precise control over camera and multi-object motion, as videos inherently capture dynamics in the projected 2D image plane. To bridge this gap, we introduce VerseCrafter, a geome

Cited by 0SourcecodeScholar
2025

Bypassing the Exponential Dependency: Looped Transformers Efficiently Learn In-context by Multi-step Gradient Descent

AISTATS 2025poster

In-context learning has been recognized as a key factor in the success of Large Language Models (LLMs). It refers to the model's ability to learn patterns on the fly from provided in-context examples in the prompt during inference. Previous studies have demonstrated that the Transformer architecture…

Cited by 0SourceScholar
2025

CAN-ST: Clustering Adaptive Normalization for Spatio-temporal OOD Learning

IJCAI 2025

Spatio-temporal data mining is crucial for decision-making and planning in diverse domains. However, in real-world scenarios, training and testing data are often not independent or identically distributed due to rapid changes in data distributions over time and space, resulting in spatio-temporal ou

Cited by 0SourcePDFScholar
2025

Circuit Complexity Bounds for RoPE-based Transformer Architecture

EMNLP 2025

Characterizing the expressive power of the Transformer architecture is critical to understanding its capacity limits and scaling law. Recent works provide the circuit complexity bounds to Transformer-like architecture. On the other hand, position embedding has emerged as a crucial technique in moder

Cited by 0SourcePDFScholar
2025

DFNeRF: Disentangled Facial Neural Radiance Fields for Text-based Editing of Free-view Talking Head

ICASSP 2025accepted

In this paper, we propose a text-based approach that can edit the speech content of a free-view talking head based on its transcript. The core of our method is to establish the relationship between phonemes and head attributes. To avoid discontinuities in head pose and facial expressions caused by e…

Cited by 0SourceScholar
2025

Decoding LLM Personality Measurement: Forced-Choice vs. Likert

ACL 2025finding

Recent research has focused on investigating the psychological characteristics of Large Language Models (LLMs), emphasizing the importance of comprehending their behavioral traits. Likert scale personality questionnaires have become the primary tool for assessing these characteristics in LLMs. Howev…

2025

DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos

CVPR 2025highlight

Estimating video depth in open-world scenarios is challenging due to the diversity of videos in appearance, content motion, camera movement, and length. We present DepthCrafter, an innovative method for generating temporally consistent long depth sequences with intricate details for open-world video…

2025

Deterministic Sparse Fourier Transform for Continuous Signals with Frequency Gap

ICML 2025poster

The Fourier transform is a fundamental tool in computer science and signal processing. In particular, when the signal is sparse in the frequency domain---having only $k$ distinct frequencies---sparse Fourier transform (SFT) algorithms can recover the signal in a sublinear time (proportional to the s…

Cited by 0SourcePDFScholar
2025

DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation

CVPR 2025poster

Sora-like video generation models have achieved remarkable progress with a Multi-Modal Diffusion Transformer (MM-DiT) architecture. However, the current video generation models predominantly focus on single-prompt, struggling to generate coherent scenes with multiple sequential prompts that better r…

2025

Dissecting Submission Limit in Desk-Rejections: A Mathematical Analysis of Fairness in AI Conference Policies

ICML 2025poster

As AI research surges in both impact and volume, conferences have imposed submission limits to maintain paper quality and alleviate organizational pressure. In this work, we examine the fairness of desk-rejection systems under submission limits and reveal that existing practices can result in subst…

Cited by 5SourcePDFScholar
2025

Efficient $k$-Sparse Band–Limited Interpolation with Improved Approximation Ratio

NeurIPS 2025poster

We consider the task of interpolating a $k$-sparse band–limited signal from a small collection of noisy time-domain samples. Exploiting a new analytic framework for hierarchical frequency decomposition that performs systematic noise cancellation, we give the first polynomial-time algorithm with a pr…

Cited by 0SourceScholar
2025

Fundamental Limits of Visual Autoregressive Transformers: Universal Approximation Abilities

ICML 2025poster

We investigate the fundamental limits of transformer-based foundation models, extending our analysis to include Visual Autoregressive (VAR) transformers. VAR represents a big step toward generating images using a novel, scalable, coarse-to-fine ``next-scale prediction'' framework. These models set a…

Cited by 0SourcePDFScholar
2025

GeometryCrafter: Consistent Geometry Estimation for Open-world Videos with Diffusion Priors

ICCV 2025poster

Despite remarkable advancements in video depth estimation, existing methods fall short in geometric fidelity due to their affine-invariant predictions, restricting their applicability in reconstruction and other metrically grounded downstream tasks. We propose a novel point map Variational Autoencod…

Cited by 0SourcePDFScholar
2025

Longitudinal Wrist PPG Analysis for Reliable Hypertension Risk Screening Using Deep Learning

ICASSP 2025accepted

Hypertension is a leading risk factor for cardiovascular diseases. Traditional blood pressure monitoring methods are cumbersome and inadequate for continuous tracking, prompting the development of PPG-based cuffless blood pressure monitoring wearables. This study leverages deep learning models, incl…

Cited by 0SourceScholar
2025

MagicMan: Generative Novel View Synthesis of Humans with 3D-Aware Diffusion and Iterative Refinement

AAAI 2025technical

Existing works in single-image human reconstruction suffer from weak generalizability due to insufficient training data or 3D inconsistencies for a lack of comprehensive multi-view knowledge. In this paper, we introduce MagicMan, a human-specific multi-view diffusion model to generate high-quality n…

Cited by 8SourcePDFScholar
2025

Mani-GS: Gaussian Splatting Manipulation with Triangular Mesh

CVPR 2025poster

Neural 3D representations, such as Neural Radiation Fields (NeRF), excel at producing photorealistic rendering results but lack the flexibility for manipulation and editing which is crucial for content creation. However, manipulating NeRF is not highly controllable and requires a long training and i…

Cited by 10SourcePDFScholar
2025

NRFlow: Towards Noise-Robust Generative Modeling via High-Order Mechanism

UAI 2025

Flow-based generative models have shown promise in various machine learning applications, but they often face challenges in handling noise and ensuring robustness in trajectory estimation. In this work, we propose NRFlow, a novel extension to flow-based generative modeling that incorporates second-o

Cited by 0SourcePDFScholar
2025

NVComposer: Boosting Generative Novel View Synthesis with Multiple Sparse and Unposed Images

CVPR 2025poster

Recent advancements in generative models have significantly improved novel view synthesis (NVS) from multi-view data. However, existing methods depend on external multi-view alignment processes, such as explicit pose estimation or pre-reconstruction, which limits their flexibility and accessibility,…

Cited by 1SourcePDFScholar
2025

Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision Content

CVPR 2025poster

Evaluating text-to-vision content hinges on two crucial aspects: **visual quality** and **alignment**. While significant progress has been made in developing objective models to assess these dimensions, the performance of such models heavily relies on the scale and quality of human annotations. Acco…

2025

SMILE: A Scale-aware Multiple Instance Learning Method for Multicenter STAS Lung Cancer Histopathology Diagnosis

IJCAI 2025

Spread through air spaces (STAS) represents a newly identified aggressive pattern in lung cancer, which is known to be associated with adverse prognostic factors and complex pathological features. Pathologists currently rely on time-consuming manual assessments, which are highly subjective and prone

2025

Spatio-temporal Prototype-based Hierarchical Learning for OD Demand Prediction

IJCAI 2025

Origin-Destination (OD) demand prediction is a pivotal yet highly challenging task in intelligent transportation systems, aiming to accurately forecast cross-region ridership flows within urban networks. While previous studies have focused on modeling node-to-node relationships, most of them neglect

Cited by 0SourcePDFScholar
2024

A Pre-convolved Representation for Plug-and-Play Neural Illumination Fields

AAAI 2024technical

Recent advances in implicit neural representation have demonstrated the ability to recover detailed geometry and material from multi-view images. However, the use of simplified lighting models such as environment maps to represent non-distant illumination, or using a network to fit indirect light mo…

Cited by 2SourcePDFScholar
2024

All Neural Kronecker Product Beamforming for Speech Extraction with Large-Scale Microphone Arrays

ICASSP 2024accepted

Existing frame-wise neural beamformers for speech extraction can obtain promising performance in relatively high signal-to-noise ratio (SNR) scenarios using small microphone arrays, while they still suffer from performance degradation in relatively low SNR environments, e.g., SNR<-5 dB. As an attemp…

Cited by 0SourceScholar
2024

CAMEL: Capturing Metaphorical Alignment with Context Disentangling for Multimodal Emotion Recognition

AAAI 2024technical

Understanding the emotional polarity of multimodal content with metaphorical characteristics, such as memes, poses a significant challenge in Multimodal Emotion Recognition (MER). Previous MER researches have overlooked the phenomenon of metaphorical alignment in multimedia content, which involves n…

Cited by 9SourcePDFScholar
2024

CV-VAE: A Compatible Video VAE for Latent Generative Video Models

NeurIPS 2024poster

Spatio-temporal compression of videos, utilizing networks such as Variational Autoencoders (VAE), plays a crucial role in OpenAI's SORA and numerous other video generative models. For instance, many LLM-like video models learn the distribution of discrete tokens derived from 3D VAEs within the VQVAE…

2024

ConTex-Human: Free-View Rendering of Human from a Single Image with Texture-Consistent Synthesis

CVPR 2024poster

In this work we propose a method to address the challenge of rendering a 3D human from a single image in a free-view manner. Some existing approaches could achieve this by using generalizable pixel-aligned implicit fields to reconstruct a textured mesh of a human or by employing a 2D diffusion model…

Cited by 6SourcePDFScholar
2024

Constraint Latent Space Matters: An Anti-anomalous Waveform Transformation Solution from Photoplethysmography to Arterial Blood Pressure

AAAI 2024technical

Arterial blood pressure (ABP) holds substantial promise for proactive cardiovascular health management. Notwithstanding its potential, the invasive nature of ABP measurements confines their utility primarily to clinical environments, limiting their applicability for continuous monitoring beyond medi…

Cited by 1SourcePDFScholar
2024

Fast-Poly: A Fast Polyhedral Algorithm for 3D Multi-Object Tracking

RA-L 2024

3D Multi-Object Tracking (MOT) captures stable and comprehensive motion states of surrounding obstacles, essential for robotic perception. However, current 3D trackers face issues with accuracy and latency consistency. In this letter, we propose Fast-Poly, a fast and effective filter-based method fo

Cited by 16SourceScholar
2024

Head360: Learning a Parametric 3D Full-Head for Free-View Synthesis in 360°

ECCV 2024poster

"Creating a 360◦ parametric model of a human head is a very challenging task. While recent advancements have demonstrated the efficacy of leveraging synthetic data for building such parametric head models, their performance remains inadequate in crucial areas such as expression-driven animation, hai…

2024

HiFi-123: Towards High-fidelity One Image to 3D Content Generation

ECCV 2024poster

"Recent advances in diffusion models have enabled 3D generation from a single image. However, current methods often produce suboptimal results for novel views, with blurred textures and deviations from the reference image, limiting their practical applications. In this paper, we introduce HiFi-123,…

Cited by 26SourcePDFScholar
2024

HumanRef: Single Image to 3D Human Generation via Reference-Guided Diffusion

CVPR 2024poster

Generating a 3D human model from a single reference image is challenging because it requires inferring textures and geometries in invisible views while maintaining consistency with the reference image. Previous methods utilizing 3D generative models are limited by the availability of 3D training dat…

2023

3D GAN Inversion With Facial Symmetry Prior

CVPR 2023poster

Recently, a surge of high-quality 3D-aware GANs have been proposed, which leverage the generative power of neural rendering. It is natural to associate 3D GANs with GAN inversion methods to project a real image into the generator's latent space, allowing free-view consistent synthesis and editing, r…

Cited by 45SourcePDFScholar
2023

CO-Net: Learning Multiple Point Cloud Tasks at Once with A Cohesive Network

ICCV 2023poster

We present CO-Net, a cohesive framework that optimizes multiple point cloud tasks collectively across heterogeneous dataset domains. CO-Net maintains the characteristics of high storage efficiency since models with the preponderance of shared parameters can be assembled into a single model. Specific…

Cited by 7PDFScholar
2023

Event Causality Extraction via Implicit Cause-Effect Interactions

EMNLP 2023long main

Event Causality Extraction (ECE) aims to extract the cause-effect event pairs from the given text, which requires the model to possess a strong reasoning ability to capture event causalities. However, existing works have not adequately exploited the interactions between the cause and effect event th…

Cited by 0SourceScholar
2023

Local Implicit Ray Function for Generalizable Radiance Field Representation

CVPR 2023poster

We propose LIRF (Local Implicit Ray Function), a generalizable neural rendering approach for novel view rendering. Current generalizable neural radiance fields (NeRF) methods sample a scene with a single ray per pixel and may therefore render blurred or aliased views when the input views and rendere…

Cited by 30SourcePDFScholar
2023

Narrative Order Aware Story Generation via Bidirectional Pretraining Model with Optimal Transport Reward

EMNLP 2023long findings

To create a captivating story, a writer often plans a sequence of logically coherent events and ingeniously manipulates the narrative order to generate flashback in place. However, existing storytelling systems suffer from both insufficient understanding of event correlations and inadequate awarenes…

Cited by 0SourceScholar
2023

Next3D: Generative Neural Texture Rasterization for 3D-Aware Head Avatars

CVPR 2023highlight

3D-aware generative adversarial networks (GANs) synthesize high-fidelity and multi-view-consistent facial images using only collections of single-view 2D imagery. Towards fine-grained control over facial attributes, recent efforts incorporate 3D Morphable Face Model (3DMM) to describe deformation in…

2023

Poly-MOT: A Polyhedral Framework For 3D Multi-Object Tracking

IROS 2023poster

3D Multi-object tracking (MOT) empowers mobile robots to accomplish well-informed motion planning and navigation tasks by providing motion trajectories of surrounding objects. However, existing 3D MOT methods typically employ a single similarity metric and physical model to perform data association…

Cited by 38SourcecodeScholar
2023

TOT:Topology-Aware Optimal Transport for Multimodal Hate Detection

AAAI 2023technical

Multimodal hate detection, which aims to identify the harmful content online such as memes, is crucial for building a wholesome internet environment. Previous work has made enlightening exploration in detecting explicit hate remarks. However, most of their approaches neglect the analysis of implicit…

Cited by 14SourcePDFScholar
2023

UV Volumes for Real-Time Rendering of Editable Free-View Human Performance

CVPR 2023poster

Neural volume rendering enables photo-realistic renderings of a human performer in free-view, a critical task in immersive VR/AR applications. But the practice is severely limited by high computational costs in the rendering process. To solve this problem, we propose the UV Volumes, a new approach t…

2022

Deblur-NeRF: Neural Radiance Fields From Blurry Images

CVPR 2022poster

Neural Radiance Field (NeRF) has gained considerable attention recently for 3D scene reconstruction and novel view synthesis due to its remarkable synthesis quality. However, image blurriness caused by defocus or motion, which often occurs when capturing scenes in the wild, significantly degrades it…

Cited by 289PDFcodeScholar
2022

FENeRF: Face Editing in Neural Radiance Fields

CVPR 2022poster

Previous portrait image generation methods roughly fall into two categories: 2D GANs and 3D-aware GANs. 2D GANs can generate high fidelity portraits but with low view consistency. 3D-aware GAN methods can maintain view consistency but their generated images are not locally editable. To overcome thes…

Cited by 169PDFcodeScholar
2022

Hallucinated Neural Radiance Fields in the Wild

CVPR 2022poster

Neural Radiance Fields (NeRF) has recently gained popularity for its impressive novel view synthesis ability. This paper studies the problem of hallucinated NeRF: i.e., recovering a realistic NeRF at a different time of day from a group of tourism images. Existing solutions adopt NeRF with a control…

Cited by 137PDFcodeScholar
2021

A Second look at Exponential and Cosine Step Sizes: Simplicity, Adaptivity, and Performance

ICML 2021spotlight

Stochastic Gradient Descent (SGD) is a popular tool in training large-scale machine learning models. Its performance, however, is highly variable, depending crucially on the choice of the step sizes. Accordingly, a variety of strategies for tuning the step sizes have been proposed, ranging from coor…

2020

GraphRQI: Classifying Driver Behaviors Using Graph Spectrums

ICRA 2020poster

We present a novel algorithm (GraphRQI) to identify driver behaviors from road-agent trajectories. Our approach assumes that the road-agents exhibit a range of driving traits, such as aggressive or conservative driving. Moreover, these traits affect the trajectories of nearby road-agents as well as…

Cited by 30SourceScholar