← Search

Yin Yang

30 accepted papers

2026

Analysing Satellite Imagery Classification Under Spatial Domain Shift Across Geographic Regions (Abstract Reprint)

AAAI 2026technical

Deep learning models are designed based on the i.i.d. assumption; consequently, they experience a significant performance drop due to the distribution shifts when deployed in real environments. Domain Generalisation (DG) aims to bridge the distribution shift between the source and target domains by

Cited by 0SourcePDFScholar
2026

AniMimic: Imitating 3D Animation from Video Priors

CVPR 2026

Creating realistic 3D animation remains a time-consuming and expertise-dependent process, requiring manual rigging, keyframing, and fine-tuning of complex motions. Meanwhile, video diffusion models have recently demonstrated remarkable 2D motion imagination, generating dynamic and visually coherent

Cited by 0SourceScholar
2026

ElastoGen: 4D Generative Elastodynamics

AAAI 2026technical

We present ElastoGen, a knowledge-driven AI model that generates physically accurate 4D elastodynamics. Unlike deep models that learn from video- or image-based observations, ElastoGen leverages the principles of physics and learns from established mathematical and optimization procedures. The core

Cited by 0SourcePDFScholar
2026

Physically Valid Biomolecular Interaction Modeling with Gauss-Seidel Projection

ICLR 2026poster

Biomolecular interaction modeling has been substantially advanced by foundation models, yet they often produce all-atom structures that violate basic steric feasibility. We address this limitation by enforcing physical validity as a strict constraint during both training and inference with a unified…

Cited by 0SourcecodeScholar
2026

Right-Side-Out: Learning Zero-Shot Sim-To-Real Garment Reversal

ICRA 2026poster

Turning garments right-side out is a challenging manipulation task: it is highly dynamic, entails rapid contact changes, and is subject to severe visual occlusion. We introduce Right-Side-Out, a zero-shot sim-to-real framework that effectively solves this challenge by exploiting task structures. We …

2026

SPARK: Sim-ready Part-level Articulated Reconstruction with VLM Knowledge

CVPR 2026

Articulated 3D objects are critical for embodied AI, robotics, and scene understanding, yet creating simulation-ready assets remains labor-intensive and requires expert modeling of part hierarchies and motion structures. We introduce SPARK, a framework for reconstructing physically consistent, kinem

Cited by 0SourcecodeScholar
2025

ARM: Appearance Reconstruction Model for Relightable 3D Generation

CVPR 2025highlight

Recent image-to-3D reconstruction models have greatly advanced geometry generation, but they still struggle to faithfully generate realistic appearance. To address this, we introduce ARM, a novel method that reconstructs high-quality 3D meshes and realistic appearance from sparse-view images. The co…

2025

Benchmarking Object Detectors under Real-World Distribution Shifts in Satellite Imagery

CVPR 2025poster

Object detectors have achieved remarkable performance in many applications; however, these deep learning models are typically designed under the i.i.d. assumption, meaning they are trained and evaluated on data sampled from the same (source) distribution. In real-world deployment, however, target di…

2025

Co-Evolving LLMs and Embedding Models via Density-Guided Preference Optimization for Text Clustering

EMNLP 2025

Large language models (LLMs) have shown strong potential in enhancing text clustering when combined with traditional embedding models. However, existing methods predominantly treat LLMs as static pseudo-oracles, i.e., unidirectionally querying them for similarity assessment or data augmentation, whi

Cited by 0SourcePDFScholar
2025

Embedded IPC: Fast and Intersection-Free Simulation in Reduced Subspace for Robot Manipulation

ICRA 2025

Physics-based simulation is essential for developing and evaluating robot manipulation policies, particularly in scenarios involving deformable objects and complex contact interactions. However, existing simulators often struggle to balance computational efficiency with numerical accuracy, especiall

Cited by 2SourceScholar
2025

GRIP: A General Robotic Incremental Potential Contact Simulation Dataset for Unified Deformable-Rigid Coupled Grasping

IROS 2025

Grasping is fundamental to robotic manipulation, and recent advances in large-scale grasping datasets have provided essential training data and evaluation benchmarks, accelerating the development of learning-based methods for robust object grasping. However, most existing datasets exclude deformable

Cited by 3SourcecodeScholar
2025

Gaussian Splashing: Unified Particles for Versatile Motion Synthesis and Rendering

CVPR 2025poster

We demonstrate the feasibility of integrating physics-based animations of solids and fluids with 3D Gaussian Splatting (3DGS) to create novel effects in virtual scenes reconstructed using 3DGS. Leveraging the coherence of the Gaussian Splatting and Position-Based Dynamics (PBD) in the underlying rep…

Cited by 10SourcePDFScholar
2025

High-fidelity 3D Object Generation from Single Image with RGBN-Volume Gaussian Reconstruction Model

CVPR 2025highlight

Recently single-view 3D generation via Gaussian splatting has emerged and developed quickly. They learn 3D Gaussians from 2D RGB images generated from pre-trained multi-view diffusion (MVD) models, and have shown a promising avenue for 3D generation through a single image. Despite the current progre…

Cited by 0SourcePDFScholar
2025

Invertible Fourier Neural Operators for Tackling Both Forward and Inverse Problems

AISTATS 2025poster

Fourier Neural Operator (FNO) is a powerful and popular operator learning method. However, FNO is mainly used in forward prediction, yet a great many applications rely on solving inverse problems. In this paper, we propose an invertible Fourier Neural Operator (iFNO) for jointly tackling the forwar…

Cited by 0SourcecodeScholar
2025

Real-time High-fidelity Gaussian Human Avatars with Position-based Interpolation of Spatially Distributed MLPs

CVPR 2025highlight

Many works have succeeded in reconstructing Gaussian human avatars from multi-view videos. However, they either struggle to capture pose-dependent appearance details with a single MLP, or rely on a computationally intensive neural network to reconstruct high-fidelity appearance but with rendering pe…

2025

WaveSpect: A Hybrid Approach to Synthetic Audio Detection via Waveform and Spectrogram Analysis

ICASSP 2025accepted

With the rapid advancement of synthetic speech technology, the challenges posed by audio deepfakes have become increasingly severe. Despite notable progress in synthetic speech detection, existing algorithms exhibit limited generalization to unknown attacks. To address these challenges, we propose W…

Cited by 0SourceScholar
2025

WonderPlay: Dynamic 3D Scene Generation from a Single Image and Actions

ICCV 2025poster

WonderPlay is a novel framework integrating physics simulation with video generation for generating action-conditioned dynamic 3D scenes from a single image. Our hybrid generative simulator first uses a physics solver to simulate coarse 3D dynamics, which subsequently conditions a video generator to…

Cited by 0SourcePDFScholar
2024

Atlas3D: Physically Constrained Self-Supporting Text-to-3D for Simulation and Fabrication

NeurIPS 2024poster

Existing diffusion-based text-to-3D generation methods primarily focus on producing visually realistic shapes and appearances, often neglecting the physical constraints necessary for downstream tasks. Generated models frequently fail to maintain balance when placed in physics-based simulations or 3D…

Cited by 5SourcePDFScholar
2024

PIE-NeRF: Physics-based Interactive Elastodynamics with NeRF

CVPR 2024poster

We show that physics-based simulations can be seamlessly integrated with NeRF to generate high-quality elastodynamics of real-world objects. Unlike existing methods we discretize nonlinear hyperelasticity in a meshless way obviating the necessity for intermediate auxiliary shape proxies like a tetra…

2024

PhysGaussian: Physics-Integrated 3D Gaussians for Generative Dynamics

CVPR 2024highlight

We introduce PhysGaussian a new method that seamlessly integrates physically grounded Newtonian dynamics within 3D Gaussians to achieve high-quality novel motion synthesis. Employing a customized Material Point Method (MPM) our approach enriches 3D Gaussian kernels with physically meaningful kinemat…

Cited by 178SourcePDFScholar
2022

Active Boundary Loss for Semantic Segmentation

AAAI 2022technical

This paper proposes a novel active boundary loss for semantic segmentation. It can progressively encourage the alignment between predicted boundaries and ground-truth boundaries during end-to-end training, which is not explicitly enforced in commonly used cross-entropy loss. Based on the predicted b…

2022

HoD-Net: High-Order Differentiable Deep Neural Networks and Applications

AAAI 2022technical

We introduce a deep architecture named HoD-Net to enable high-order differentiability for deep learning. HoD-Net is based on and generalizes the complex-step finite difference (CSFD) method. While similar to classic finite difference, CSFD approaches the derivative of a function from a higher-dimens…

Cited by 4SourcePDFScholar
2022

PlasticityNet: Learning to Simulate Metal, Sand, and Snow for Optimization Time Integration

NeurIPS 2022accept

In this paper, we propose a neural network-based approach for learning to represent the behavior of plastic solid materials ranging from rubber and metal to sand and snow. Unlike elastic forces such as spring forces, these plastic forces do not result from the positional gradient of any potential en…

Cited by 17SourcePDFScholar
2022

Pose Guided Image Generation from Misaligned Sources via Residual Flow Based Correction

AAAI 2022technical

Generating new images with desired properties (e.g. new view/poses) from source images has been enthusiastically pursued recently, due to its wide range of potential applications. One way to ensure high-quality generation is to use multiple sources with complementary information such as different vi…

Cited by 3SourcePDFScholar
2021

In-game Residential Home Planning via Visual Context-aware Global Relation Learning

AAAI 2021technical

In this paper, we propose an effective global relation learning algorithm to recommend an appropriate location of a building unit for in-game customization of residential home complex. Given a construction layout, we propose a visual context-aware graph generation network that learns the implicit gl…

Cited by 5SourcePDFScholar
2021

Location-Aware Single Image Reflection Removal

ICCV 2021poster

This paper proposes a novel location-aware deep-learning-based single image reflection removal method. Our network has a reflection detection module to regress a probabilistic reflection confidence map, taking multi-scale Laplacian features as inputs. This probabilistic map tells if a region is refl…

Cited by 110PDFcodeScholar
2021

Online 3D Bin Packing with Constrained Deep Reinforcement Learning

AAAI 2021technical

We solve a challenging yet practically useful variant of 3D Bin Packing Problem (3D-BPP). In our problem, the agent has limited information about the items to be packed into a single bin, and an item must be packed immediately after its arrival without buffering or readjusting. The item's placement…

2021

Unsupervised Image Generation With Infinite Generative Adversarial Networks

ICCV 2021poster

Image generation has been heavily investigated in computer vision, where one core research challenge is to generate images from arbitrarily complex distributions with little supervision. Generative Adversarial Networks (GANs) as an implicit approach have achieved great successes in this direction an…

Cited by 6PDFcodeScholar
2016

Contour-based 3D tongue motion visualization using ultrasound image sequences

ICASSP 2016accepted

This article describes a contour-based 3D tongue deformation visualization framework using B-mode ultrasound image sequences. A robust, automatic tracking algorithm characterizes tongue motion via a contour, which is then used to drive a generic 3D Finite Element Model (FEM). A novel contour-based 3…

Cited by 0SourceScholar