← Search

Yihao Luo

17 accepted papers

2026

EchoPOSE: 6D Pose Estimation of Sparse Echocardiograms for Left-Ventricular 3D Shape Reconstruction

CVPR 2026

3D echocardiography provides superior cardiac quantification to traditional 2D echocardiography, which suffers from geometric idealizations and imaging plane misalignment. However, despite its advantages, clinical adoption of 3D echo remains limited due to logistical and visualization challenges. We

Cited by 0SourceScholar
2026

Faithful Contouring: Near-Lossless 3D Voxel Representation Free from Iso-surface

CVPR 2026

Accurate and efficient voxelized representations of 3D meshes are the foundation of 3D reconstruction and generation. However, existing representations based on iso-surface heavily rely on water-tightening or rendering optimization, which inevitably compromise geometric fidelity. We propose Faithful

Cited by 0SourcecodeScholar
2025

LawDNet: Enhanced Audio-Driven Lip Synthesis via Local Affine Warping Deformation

ICASSP 2025accepted

In the domain of photorealistic talking head generation, the fidelity of audio-driven lip motion synthesis is essential for realistic virtual interactions. Existing methods face two key challenges: a lack of vivacity due to limited diversity in generated lip poses and noticeable anamorphose motions…

Cited by 0SourceScholar
2025

MeshAnything V2: Artist-Created Mesh Generation with Adjacent Mesh Tokenization

ICCV 2025poster

Meshes are the de facto 3D representation in the industry but are labor-intensive to produce. Recently, a line of research has focused on autoregressively generating meshes. This approach processes meshes into a sequence composed of vertices and then generates them vertex by vertex, similar to how a…

2025

Sparc3D: Sparse Representation and Construction for High-Resolution 3D Shapes Modeling

NeurIPS 2025poster

High-fidelity 3D object synthesis remains significantly more challenging than 2D image generation due to the unstructured nature of mesh data and the cubic complexity of dense volumetric grids. Existing two-stage pipelines—compressing meshes with a VAE (using either 2D or 3D supervision), followed b…

Cited by 0SourceScholar
2024

A Fourier Perspective of Feature Extraction and Adversarial Robustness

IJCAI 2024poster

Adversarial robustness and interpretability are longstanding challenges of computer vision. Deep neural networks are vulnerable to adversarial perturbations that are incomprehensible and imperceptible to humans. However, the opaqueness of networks prevents one from theoretically addressing adversari…

Cited by 1SourcePDFScholar
2023

An Application of Quantum Mechanics to Attention Methods in Computer Vision

ICASSP 2023accepted

This work proposes the quantum-state-based mapping (QSM) for machine learning. QSM uses wave functions that describe microscopic particle systems as mappings. By QSM, original inputs or features extracted by neural networks are processed as quantum states to train wave function parameters. QSM has a…

Cited by 0SourceScholar
2023

D-IF: Uncertainty-aware Human Digitization via Implicit Distribution Field

ICCV 2023poster

Realistic virtual humans play a crucial role in numerous industries, such as metaverse, intelligent healthcare, and self-driving simulation. But creating them on a large scale with high levels of realism remains a challenge. The utilization of deep implicit function sparks a new era of image-based 3…

Cited by 40PDFcodeScholar
2023

Frequency and Scale Perspectives of Feature Extraction

ICASSP 2023accepted

Convolutional neural networks (CNNs) have achieved superior performance but still lack clarity about the nature and properties of feature extraction. In this paper, by analyzing the sensitivity of neural networks to frequencies and scales, we find that neural networks not only have low- and mediumfr…

Cited by 0SourceScholar
2023

Training Robust Spiking Neural Networks on Neuromorphic Data with Spatiotemporal Fragments

ICASSP 2023accepted

Neuromorphic vision sensors (event cameras) are inherently suitable for spiking neural networks (SNNs) and provide novel neuromorphic vision data for this biomimetic model. Due to the spatiotemporal characteristics, novel data augmentations are required to process the unconventional visual signals o…

Cited by 0SourceScholar
2023

Training Robust Spiking Neural Networks with Viewpoint Transform and Spatiotemporal Stretching

ICASSP 2023accepted

Neuromorphic vision sensors (event cameras) simulate biological visual perception systems and have the advantages of high temporal resolution, less data redundancy, low power consumption, and large dynamic range. Since both events and spikes are modeled from neural signals, event cameras are inheren…

Cited by 0SourceScholar
2023

Training Stronger Spiking Neural Networks with Biomimetic Adaptive Internal Association Neurons

ICASSP 2023accepted

As the third generation of neural networks, spiking neural networks (SNNs) are dedicated to exploring more insightful neural mechanisms to achieve near-biological intelligence. Intuitively, biomimetic mechanisms are crucial to understanding and improving SNNs. For example, the associative long-term…

Cited by 0SourceScholar
2022

Dynamic Multi-Scale Loss Balance for Object Detection

ICASSP 2022accepted

It is a common paradigm in object detection frameworks to perform multi-scale detection. However, each scale is treated equally during training. In this paper, we carefully study the objective imbalance of multi-scale detector training. We argue that the loss in each scale is neither equally importa…

Cited by 0SourceScholar
2022

Kernel Estimation Network for Blind Super-Resolution

ICASSP 2022accepted

Existing super-resolution (SR) methods commonly assume that the degradation kernels are fixed and known (e.g., bicubic downsampling or single Gaussian blurring kernel). However, these methods suffer a severe performance drop when the real degradations deviate from this assumption. To address this is…

Cited by 0SourceScholar
2022

Multi-Scale Reinforcement Learning Strategy for Object Detection

ICASSP 2022accepted

Feature Pyramid Network (FPN) has become a common detection paradigm by improving multi-scale features with strong semantics. However, most FPN-based methods typically treat each feature map equally and sum the loss without distinction, which might lead to suboptimal overall performance. In this pap…

Cited by 0SourceScholar
2022

Multi-View Data Representation Via Deep Autoencoder-Like Nonnegative Matrix Factorization

ICASSP 2022accepted

Since a large proportion of real-world data is made of different representations or views, learning on data represented with multiple views (e.g., numerous types of features or modalities) has garnered considerable attention recently. Nonnegative matrix factorization (NMF) has been widely adopted fo…

Cited by 0SourceScholar
2019

Breast Cancer Image Classification on WSI with Spatial Correlations

ICASSP 2019accepted

As common cancer, breast cancer kills thousands of women every year. It’s significant to provide doctors computer-aided diagnosis (CAD) to ease their workload as well as improve detection quality. Patch-level CNNs are usually used to classify the breast tissue slice, and the CNNs classify each patch…

Cited by 0SourceScholar