← Search

Xinli Xu

12 accepted papers

2026

ImpText: A Benchmark and Tool-Augmented Framework for Implicit Text Reasoning

ICML 2026poster

Multimodal Large Language Models (MLLMs) have demonstrated exceptional proficiency in standard text extraction, but they encounter significant challenges when confronting real-world implicit text. Such content typically contains malicious information, intentionally concealed through physical deforma…

Cited by 0SourceScholar
2025

ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive Feedback

NeurIPS 2025poster

With the rapid advancement of generative models, general-purpose generation has gained increasing attention as a promising approach to unify diverse tasks across modalities within a single system. Despite this progress, existing open-source frameworks often remain fragile and struggle to support com…

Cited by 0SourceScholar
2025

FlexGen: Flexible Multi-View Generation from Text and Image Inputs

ICCV 2025poster

In this work, we introduce FlexGen, a flexible framework designed to generate controllable and consistent multi-view images, conditioned on a single-view image, or a text prompt, or both. FlexGen tackles the challenges of controllable multi-view synthesis through additional conditioning on 3D-aware…

Cited by 0SourcePDFScholar
2025

GaussianProperty: Integrating Physical Properties to 3D Gaussians with LMMs

ICCV 2025poster

Estimating physical properties for visual data is a crucial task in computer vision, graphics, and robotics, underpinning applications such as augmented reality, physical simulation, and robotic grasping. However, this area remains under-explored due to the inherent ambiguities in physical property…

Cited by 0SourcePDFScholar
2025

Kiss3DGen: Repurposing Image Diffusion Models for 3D Asset Generation

CVPR 2025poster

Diffusion models have achieved great success in generating 2D images. However, the quality and generalizability of 3D content generation remain limited. State-of-the-art methods often require large-scale 3D assets for training, which are challenging to collect. In this work, we introduce Kiss3DGen (…

2025

Orchestrating Audio: Multi-Agent Framework for Long-Video Audio Synthesis

EMNLP 2025

Video-to-audio synthesis, which generates synchronized audio for visual content, critically enhances viewer immersion and narrative coherence in film and interactive media. However, video-to-audio dubbing for long-form content remains an unsolved challenge due to dynamic semantic shifts, audio diver

Cited by 0SourcePDFScholar
2025

PRM: Photometric Stereo based Large Reconstruction Model

ICCV 2025poster

We propose PRM, a novel photometric stereo based large reconstruction model to reconstruct high-quality meshes with fine-grained details. Previous large reconstruction models typically prepare training images under fixed and simple lighting, offering minimal photometric cues for precise reconstructi…

Cited by 0SourcePDFScholar
2025

Pancreatic Cystic Neoplasms Lesion Detection for Non-contrast CT Image via Teacher-student Model

ICASSP 2025accepted

Due to the low contrast between lesion features and surrounding tissues in non-contrast CT images, traditional detection methods often struggle to accurately identify and differentiate various types of cystic tumors. This limitation increases the risk of misdiagnosis and missed detection, thereby hi…

Cited by 0SourceScholar
2025

PreGenie: An Agentic Framework for High-quality Visual Presentation Generation

EMNLP 2025

Visual presentations are vital for effective communication. Early attempts to automate their creation using deep learning often faced issues such as poorly organized layouts, inaccurate text summarization, and a lack of image understanding, leading to mismatched visuals and text. These limitations r

Cited by 0SourcePDFScholar
2023

Sample-adaptive Augmentation for Point Cloud Recognition Against Real-world Corruptions

ICCV 2023poster

Robust 3D perception under corruption has become an essential task for the realm of 3D vision. While current data augmentation techniques usually perform random transformations on all point cloud objects in an offline way and ignore the structure of the samples, resulting in over-or-under enhancemen…

Cited by 8PDFcodeScholar
2022

FH-Net: A Fast Hierarchical Network for Scene Flow Estimation on Real-World Point Clouds

ECCV 2022poster

"Estimating scene flow from real-world point clouds is a fundamental task for practical 3D vision. Previous methods often rely on deep models to first extract expensive per-point features at full resolution, and then get the flow either from complex matching mechanism or feature decoding, suffering…

2022

MsSVT: Mixed-scale Sparse Voxel Transformer for 3D Object Detection on Point Clouds

NeurIPS 2022accept

3D object detection from the LiDAR point cloud is fundamental to autonomous driving. Large-scale outdoor scenes usually feature significant variance in instance scales, thus requiring features rich in long-range and fine-grained information to support accurate detection. Recent detectors leverage th…