← Search

Fengyu Yang

19 accepted papers

2026

Iris: Integrating Language into Diffusion-based Monocular Depth Estimation

CVPR 2026

Conventional monocular depth estimators suffer from visual ambiguities and nuisances. We demonstrate that language can improve the fidelity of estimates by providing additional information through text as a condition, thereby reducing the solution space for depth estimates. This conditional distribu

Cited by 0SourceScholar
2026

VideoSeg-R1:Reasoning Video Object Segmentation via Reinforcement Learning

AAAI 2026technical

Traditional video reasoning segmentation methods rely on supervised fine-tuning, which limits generalization to out-of-distribution scenarios and lacks explicit reasoning. To address this, we propose VideoSeg-R1, the first framework to introduce reinforcement learning into video reasoning segmentati

Cited by 0SourcePDFScholar
2025

Differentiation Through Black-Box Quadratic Programming Solvers

NeurIPS 2025poster

Differentiable optimization has attracted significant research interest, particularly for quadratic programming (QP). Existing approaches for differentiating the solution of a QP with respect to its defining parameters often rely on specific integrated solvers. This integration limits their applicab…

Cited by 0SourcecodeScholar
2025

Discretized Gaussian Representation for Tomographic Reconstruction

ICCV 2025poster

Computed Tomography (CT) enables detailed cross-sectional imaging but continues to face challenges in balancing reconstruction quality and computational efficiency. While deep learning-based methods have significantly improved image quality and noise reduction, they typically require large-scale tra…

2025

TextToucher: Fine-Grained Text-to-Touch Generation

AAAI 2025technical

Tactile sensation plays a crucial role in the development of multi-modal large models and embodied intelligence. To collect tactile data with minimal cost as possible, a series of studies have attempted to generate tactile images by vision-to-touch image translation. However, compared to text modali…

2025

Tri-Ergon: Fine-Grained Video-to-Audio Generation with Multi-Modal Conditions and LUFS Control

AAAI 2025technical

Video-to-audio (V2A) generation utilizes visual-only video features to produce realistic sounds that correspond to the scene. However, current V2A models often lack fine-grained control over the generated audio, especially in terms of loudness variation and the incorporation of multi-modal condition…

Cited by 2SourcePDFScholar
2024

APISR: Anime Production Inspired Real-World Anime Super-Resolution

CVPR 2024poster

While real-world anime super-resolution (SR) has gained increasing attention in the SR community existing methods still adopt techniques from the photorealistic domain. In this paper we analyze the anime production workflow and rethink how to use characteristics of it for the sake of the real-world…

2024

Binding Touch to Everything: Learning Unified Multimodal Tactile Representations

CVPR 2024poster

The ability to associate touch with other modalities has huge implications for humans and computational systems. However multimodal learning with touch remains challenging due to the expensive data collection process and non-standardized sensor outputs. We introduce UniTouch a unified tactile model…

Cited by 53SourcePDFScholar
2024

FreeMan: Towards Benchmarking 3D Human Pose Estimation under Real-World Conditions

CVPR 2024poster

Estimating the 3D structure of the human body from nat- ural scenes is a fundamental aspect of visual perception. 3D human pose estimation is a vital step in advancing fields like AIGC and human-robot interaction serving as a crucial tech- nique for understanding and interacting with human actions i…

2024

On the Viability of Monocular Depth Pre-training for Semantic Segmentation

ECCV 2024poster

"The question of whether pre-training on geometric tasks is viable for downstream transfer to semantic tasks is important for two reasons, one practical and the other scientific. If the answer is positive, we may be able to reduce pre-training costs and bias from human annotators significantly. If t…

2024

RSA: Resolving Scale Ambiguities in Monocular Depth Estimators through Language Descriptions

NeurIPS 2024poster

We propose a method for metric-scale monocular depth estimation. Inferring depth from a single image is an ill-posed problem due to the loss of scale from perspective projection during the image formation process. Any scale chosen is a bias, typically stemming from training on a dataset; hence, exis…

2024

WorDepth: Variational Language Prior for Monocular Depth Estimation

CVPR 2024poster

Three-dimensional (3D) reconstruction from a single image is an ill-posed problem with inherent ambiguities i.e. scale. Predicting a 3D scene from text description(s) is similarly ill-posed i.e. spatial arrangements of objects described. We investigate the question of whether two inherently ambiguou…

2022

Improving Emotional Speech Synthesis by Using SUS-Constrained VAE and Text Encoder Aggregation

ICASSP 2022accepted

Learning emotion embedding from reference audio is a straightforward approach for multi-emotion speech synthesis in encoder-decoder systems. But how to get better emotion embedding and how to inject it into TTS acoustic model more effectively are still under investigation. In this paper, we propose…

Cited by 0SourceScholar
2022

RBC: Rectifying the Biased Context in Continual Semantic Segmentation

ECCV 2022poster

"Recent years have witnessed a great development of Convolutional Neural Networks in semantic segmentation, where all classes of training images are simultaneously available. In practice, new images are usually made available in a consecutive manner, leading to a problem called Continual Semantic Se…

2022

Touch and Go: Learning from Human-Collected Vision and Touch

NeurIPS 2022accept

The ability to associate touch with sight is essential for tasks that require physically interacting with objects in the world. We propose a dataset with paired visual and tactile data called Touch and Go, in which human data collectors probe objects in natural environments using tactile sensors, wh…

Cited by 55SourcePDFScholar