← Search

Yawei Luo

29 accepted papers

2026

LangField4D: Learning Identity-Adaptive and Spatio-Temporal Continuous 4D Language Fields for Dynamic Scenes

CVPR 2026

Constructing a 4D language field that supports open-vocabulary queries is essential for semantic perception and interaction in dynamic environments. Existing 4D Gaussian-based approaches face two major challenges. First, the assumption of a static identity per Gaussian leads to semantic inconsistenc

Cited by 0SourceScholar
2026

Low-Rank Test-Time Training for Pre-Trained Point Cloud Models

CVPR 2026

Test-time training (TTT) enhances the robustness of pretrained models to out-of-distribution (OOD) data through auxiliary self-supervised tasks, without requiring labeled samples. However, existing TTT methods predominantly rely on decoder-based auxiliary objectives, which suffer from inefficient ad

Cited by 0SourceScholar
2026

Part-X-MLLM: Part-aware 3D Multimodal Large Language Model

ICLR 2026poster

We introduce Part-X-MLLM, a native 3D multimodal large language model that unifies diverse 3D tasks by formulating them as programs in a structured, executable grammar. Given an RGB point cloud and a natural language prompt, our model autoregressively generates a single, coherent token sequence enco…

Cited by 4SourcecodeScholar
2026

ParticleGS: Learning Neural Gaussian Particle Dynamics from Videos for Prior-free Physical Motion Extrapolation

CVPR 2026

The ability to extrapolate dynamic 3D scenes beyond the observed timeframe is fundamental to advancing physical world understanding and predictive modeling. Existing dynamic 3D reconstruction methods have achieved high-fidelity rendering of temporal interpolation, but typically lack physical consist

Cited by 0SourceScholar
2026

WorldMirror: Universal 3D World Reconstruction with Any-Prior Prompting

ICML 2026poster

We present WorldMirror, a unified feed-forward model for comprehensive 3D geometric prediction tasks. Unlike existing methods constrained to image-only inputs or customized for a specific task, our framework flexibly integrates diverse geometric priors, including camera poses, intrinsics, and depth …

Cited by 0SourceScholar
2025

DICS: Find Domain-Invariant and Class-Specific Features for Out-of-Distribution Generalization

ICASSP 2025accepted

While deep neural networks have made remarkable progress in various tasks, their performance typically deteriorates and faces insecurity when tested in out-of-distribution (OOD) scenarios. Many OOD methods focus on extracting domain-invariant features but neglect whether these features are unique to…

Cited by 0SourceScholar
2025

DecoupledESC: Enhancing Emotional Support Generation via Strategy-Response Decoupled Preference Optimization

EMNLP 2025

Recent advances in Emotional Support Conversation (ESC) have improved emotional support generation by fine-tuning Large Language Models (LLMs) via Supervised Fine-Tuning (SFT). However, common psychological errors still persist. While Direct Preference Optimization (DPO) shows promise in reducing su

2025

MASTER: Multi-Agent Security Through Exploration of Roles and Topological Structures - A Comprehensive Framework

EMNLP 2025

Large Language Models (LLMs)-based Multi-Agent Systems (MAS) exhibit remarkable problem-solving and task planning capabilities across diverse domains due to their specialized agentic roles and collaborative interactions. However, this also amplifies the severity of security risks under MAS attacks.

Cited by 0SourcePDFScholar
2025

MICAS: Multi-grained In-Context Adaptive Sampling for 3D Point Cloud Processing

CVPR 2025poster

Point cloud processing (PCP) encompasses tasks like reconstruction, denoising, registration, and segmentation, each often requiring specialized models to address unique task characteristics. While in-context learning (ICL) has shown promise across tasks by using a single model with task-specific dem…

Cited by 1SourcePDFScholar
2025

MaGS: Reconstructing and Simulating Dynamic 3D Objects with Mesh-adsorbed Gaussian Splatting

ICCV 2025poster

3D reconstruction and simulation, although interrelated, have distinct objectives: reconstruction requires a flexible 3D representation that can adapt to diverse scenes, while simulation needs a structured representation to model motion principles effectively. This paper introduces the Mesh-adsorbed…

Cited by 0SourcePDFScholar
2025

Optimized View and Geometry Distillation from Multi-view Diffuser

IJCAI 2025

Generating multi-view images from a single input view using image-conditioned diffusion models is a recent advancement and has shown considerable potential. However, issues such as the lack of consistency in synthesized views and over-smoothing in extracted geometry persist. Previous methods integra

2025

Ref-GS: Directional Factorization for 2D Gaussian Splatting

CVPR 2025poster

In this paper, we introduce Ref-GS, a novel approach for directional light factorization in 2D Gaussian splatting, which enables photorealistic view-dependent appearance rendering and precise geometry recovery. Ref-GS builds upon the deferred rendering of Gaussian splatting and applies directional e…

2025

SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding

CVPR 2025poster

Video-based Large Language Models (Video-LLMs) have witnessed substantial advancements in recent years, propelled by the advancement in multi-modal LLMs. Although these models have demonstrated proficiency in providing the overall description of videos, they struggle with fine-grained understanding,…

Cited by 1SourcePDFScholar
2025

TSD-SR: One-Step Diffusion with Target Score Distillation for Real-World Image Super-Resolution

CVPR 2025poster

Pre-trained text-to-image diffusion models are increasingly applied to real-world image super-resolution (Real-ISR) task. Given the iterative refinement nature of diffusion models, most existing approaches are computationally expensive. While methods such as SinSR and OSEDiff have emerged to condens…

2025

Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents

EMNLP 2025

Role-playing agents (RPAs) have attracted growing interest for their ability to simulate immersive and interactive characters. However, existing approaches primarily focus on static role profiles, overlooking the dynamic perceptual abilities inherent to humans. To bridge this gap, we introduce the c

Cited by 0SourcePDFScholar
2024

Balancing Humans and Machines: A Study on Integration Scale and Its Impact on Collaborative Performance

AAAI 2024technical

In the evolving artificial intelligence domain, hybrid human-machine systems have emerged as a transformative research area. While many studies have concentrated on individual human-machine interactions, there is a lack of focus on multi-human and multi-machine dynamics. This paper delves into these…

2024

Entangled View-Epipolar Information Aggregation for Generalizable Neural Radiance Fields

CVPR 2024poster

Generalizable NeRF can directly synthesize novel views across new scenes eliminating the need for scene-specific retraining in vanilla NeRF. A critical enabling factor in these approaches is the extraction of a generalizable 3D representation by aggregating source-view features. In this paper we pro…

2024

Epipolar-Free 3D Gaussian Splatting for Generalizable Novel View Synthesis

NeurIPS 2024poster

Generalizable 3D Gaussian splitting (3DGS) can reconstruct new scenes from sparse-view observations in a feed-forward inference manner, eliminating the need for scene-specific retraining required in conventional 3DGS. However, existing methods rely heavily on epipolar priors, which can be unreliable…

Cited by 0SourcePDFScholar
2023

Adaptive Patch Deformation for Textureless-Resilient Multi-View Stereo

CVPR 2023poster

In recent years, deep learning-based approaches have shown great strength in multi-view stereo because of their outstanding ability to extract robust visual features. However, most learning-based methods need to build the cost volume and increase the receptive field enormously to get a satisfactory…

2020

Adversarial Style Mining for One-Shot Unsupervised Domain Adaptation

NeurIPS 2020poster

We aim at the problem named One-Shot Unsupervised Domain Adaptation. Unlike traditional Unsupervised Domain Adaptation, it assumes that only one unlabeled target sample can be available when learning to adapt. This setting is realistic but more challenging, in which conventional adaptation approache…

2020

Copy and Paste GAN: Face Hallucination From Shaded Thumbnails

CVPR 2020oral

Existing face hallucination methods based on convolutional neural networks (CNN) have achieved impressive performance on low-resolution (LR) faces in a normal illumination condition. However, their performance degrades dramatically when LR faces are captured in low or non-uniform illumination condit…

Cited by 38PDFScholar
2020

Mesh-Guided Multi-View Stereo With Pyramid Architecture

CVPR 2020poster

Multi-view stereo (MVS) aims to reconstruct 3D geometry of the target scene by using only information from 2D images. Although much progress has been made, it still suffers from textureless regions. To overcome this difficulty, we propose a mesh-guided MVS method with pyramid architecture, which mak…

Cited by 38PDFcodeScholar
2019

P-MVSNet: Learning Patch-Wise Matching Confidence Aggregation for Multi-View Stereo

ICCV 2019poster

Learning-based methods are demonstrating their strong competitiveness in estimating depth for multi-view stereo reconstruction in recent years. Among them the approaches that generate cost volumes based on the plane-sweeping algorithm and then use them for feature matching have shown to be very prom…

Cited by 259PDFScholar
2019

Significance-Aware Information Bottleneck for Domain Adaptive Semantic Segmentation

ICCV 2019poster

For unsupervised domain adaptation problems, the strategy of aligning the two domains in latent feature space through adversarial learning has achieved much progress in image classification, but usually fails in semantic segmentation tasks in which the latent representations are overcomplex. In this…

Cited by 262PDFScholar
2019

Taking a Closer Look at Domain Shift: Category-Level Adversaries for Semantics Consistent Domain Adaptation

CVPR 2019oral

We consider the problem of unsupervised domain adaptation in semantic segmentation. The key in this campaign consists in reducing the domain shift, i.e., enforcing the data distributions of the two domains to be similar. A popular strategy is to align the marginal distribution in the feature space t…

Cited by 933PDFcodeScholar
2018

Macro-Micro Adversarial Network for Human Parsing

ECCV 2018poster

In human parsing, the pixel-wise classification loss has drawbacks in its low-level local inconsistency and high-level semantic inconsistency. The introduction of the adversarial network tackles the two problems using a single discriminator. However, the two types of parsing inconsistency are genera…