← Search

Tiantian Feng

27 accepted papers

2026

MU-GeNeRF: Multi-view Uncertainty-guided Generalizable Neural Radiance Fields for Distractor-aware Scene

CVPR 2026

Generalizable Neural Radiance Fields (GeNeRF) enable high-quality scene reconstruction from a limited number of views and can generalize to unseen scenes. However, in real-world environments, transient distractors disrupt structural consistency across views, leading to deviated supervision signals a

Cited by 0SourcecodeScholar
2025

Convex Hull-based Algebraic Constraint for Visual Quadric SLAM

IROS 2025

Using Quadrics as the object representation has the benefits of both generality and closed-form projection derivation between image and world spaces. Although numerous constraints have been proposed for dual quadric reconstruction, we found that many of them are imprecise and provide minimal improve

Cited by 0SourcecodeScholar
2025

Creating a Lens of Chinese Culture: A Multimodal Dataset for Chinese Pun Rebus Art Understanding

ACL 2025finding

Large vision-language models (VLMs) have demonstrated remarkable abilities in understanding everyday content. However, their performance in the domain of art, particularly culturally rich art forms, remains less explored. As a pearl of human wisdom and creativity, art encapsulates complex cultural n…

2025

Data Efficient Child-Adult Speaker Diarization with Simulated Conversations

ICASSP 2025accepted

Automating child speech analysis is crucial for applications such as neurocognitive assessments. Speaker diarization, which identifies "who spoke when", is an essential component of the automated analysis. However, publicly available child-adult speaker diarization solutions are scarce due to privac…

Cited by 0SourceScholar
2025

Directional Source Separation for Robust Speech Recognition on Smart Glasses

ICASSP 2025accepted

Modern smart glasses leverage machine learning to offer real-time transcriptions, considerably enriching human communication experiences. However, such systems frequently encounter challenges related to environmental noises, leading to decreased speech recognition. To improve voice quality, this wor…

Cited by 15SourceScholar
2025

Enhancing Listened Speech Decoding from EEG via Parallel Phoneme Sequence Prediction

ICASSP 2025accepted

Brain-computer interfaces (BCI) offer numerous human-centered application possibilities, particularly affecting people with neurological disorders. Text or speech decoding from brain activities is a relevant domain that could augment the quality of life for people with impaired speech perception. We…

Cited by 0SourceScholar
2025

MutualVPR: A Mutual Learning Framework for Resolving Supervision Inconsistencies via Adaptive Clustering

NeurIPS 2025poster

Visual Place Recognition (VPR) enables robust localization through image retrieval based on learned descriptors. However, drastic appearance variations of images at the same place caused by viewpoint changes can lead to inconsistent supervision signals, thereby degrading descriptor learning. Exi…

Cited by 0SourceScholar
2025

Speech2rtMRI: Speech-Guided Diffusion Model for Real-time MRI Video of the Vocal Tract during Speech

ICASSP 2025accepted

Understanding speech production both visually and kinematically can inform second language learning system designs, as well as the creation of speaking characters in video games and animations. In this work, we introduce a data-driven method to visually represent articulator motion in Magnetic Reson…

Cited by 0SourceScholar
2025

UN3-Mapping: Uncertainty-Aware Neural Non-Projective Signed Distance Fields for 3D Mapping

RA-L 2025

Building accurate and reliable maps is a critical requirement for autonomous robots. In this paper, we propose UN3-Mapping, an implicit neural mapping method that enables high-quality 3D reconstruction with integrated uncertainty estimation. Our approach employs a hybrid representation: an implicit

Cited by 0SourcecodeScholar
2024

Audio-Visual Child-Adult Speaker Classification in Dyadic Interactions

ICASSP 2024accepted

Interactions involving children span a wide range of important domains from learning to clinical diagnostic and therapeutic contexts. Automated analyses of such interactions are motivated by the need to seek accurate insights and offer scale and robustness across diverse and wide-ranging conditions.…

Cited by 5SourceScholar
2024

Emotion-Aligned Contrastive Learning Between Images and Music

ICASSP 2024accepted

Traditional music search engines rely on retrieval methods that match natural language queries with music metadata. There have been increasing efforts to expand retrieval methods to consider the audio characteristics of music itself, using queries of various modalities including text, video, and spe…

Cited by 0SourceScholar
2024

Foundation Model Assisted Automatic Speech Emotion Recognition: Transcribing, Annotating, and Augmenting

ICASSP 2024accepted

Significant advances are being made in speech emotion recognition (SER) using deep learning models. Nonetheless, training SER systems remains challenging, requiring both time and costly resources. Like many other machine learning tasks, acquiring datasets for SER requires substantial data annotation…

Cited by 0SourceScholar
2024

LOG-LIO2: A LiDAR-Inertial Odometry With Efficient Uncertainty Analysis

RA-L 2024

Uncertainty in LiDAR measurements, stemming from factors such as range sensing, is crucial for LIO (LiDAR-Inertial Odometry) systems as it affects the accurate weighting in the loss function. While recent LIO systems address uncertainty related to range sensing, the impact of incident angle on uncer

Cited by 8SourcecodeScholar
2024

LOG-LIO: A LiDAR-Inertial Odometry With Efficient Local Geometric Information Estimation

RA-L 2024

Local geometric information, i.e., normal and distribution of points, is crucial for LiDAR-based simultaneous localization and mapping (SLAM) because it provides constraints for data association, which further determines the direction of optimization and ultimately affects the accuracy of localizati

Cited by 19SourcecodeScholar
2024

Learning Sequence Descriptor Based on Spatio-Temporal Attention for Visual Place Recognition

RA-L 2024

Visual Place Recognition (VPR) aims to retrieve frames from a geotagged database that are located at the same place as the query frame. To improve the robustness of VPR in perceptually aliasing scenarios, sequence-based VPR methods are proposed. These methods are either based on matching between fra

Cited by 12SourcecodeScholar
2024

N${3}$-Mapping: Normal Guided Neural Non-Projective Signed Distance Fields for Large-Scale 3D Mapping

RA-L 2024

Accurate and dense mapping in large-scale environments is essential for various robot applications. Recently, implicit neural signed distance fields (SDFs) have shown promising advances in this task. However, most existing approaches employ projective distances from range data as SDF supervision, in

Cited by 12SourcecodeScholar
2024

TRUST-SER: On The Trustworthiness Of Fine-Tuning Pre-Trained Speech Embeddings For Speech Emotion Recognition

ICASSP 2024accepted

Recent studies have explored using pre-trained embeddings for speech emotion recognition, achieving comparable performance to conventional methods that rely on low-level knowledge-inspired acoustic features. These embeddings are often generated from models trained on large-scale speech datasets usin…

Cited by 0SourceScholar
2023

Dynamic Object Classification of Low-Resolution Point Clouds: An LSTM-Based Ensemble Learning Approach

RA-L 2023

In unmanned vehicle perception, dynamic object classification is applied to classify objects accurately and timely, providing decision-making for obstacle avoidance and planning. Low-resolution LiDAR is one of the most important sensors for this task. Unfortunately, the existing approaches perform u

Cited by 2SourceScholar
2023

FedAudio: A Federated Learning Benchmark for Audio Tasks

ICASSP 2023accepted

Federated learning (FL) has gained substantial attention in recent years due to data privacy concerns related to the pervasiveness of consumer devices that continuously collect data from users. While a number of FL benchmarks have been developed to facilitate FL research, none of them include audio…

Cited by 33SourceScholar
2023

Toward Privacy-Enhancing Ambulatory-Based Well-Being Monitoring: Investigating User Re-Identification Risk in Multimodal Data

ICASSP 2023accepted

The sensitivity of data collected via ambulatory monitoring, which regularly involve the recording of speech signals and sensor information, can cause strong privacy concerns. We investigate user re-identification risk in a corpus of such data collected to observe the interplay between behavior, phy…

Cited by 0SourceScholar
2022

Enhancing Privacy Through Domain Adaptive Noise Injection For Speech Emotion Recognition

ICASSP 2022accepted

Speech Emotion Recognition (SER) techniques have gained considerable interest in many applications including smart virtual assistants and health state tracking. SER systems often acquire and transmit speech data collected at the client-side to remote cloud platforms for inference and decision making…

Cited by 18SourceScholar
2022

Scale Estimation with Dual Quadrics for Monocular Object SLAM

IROS 2022poster

The scale ambiguity problem is inherently unsolvable to monocular SLAM without the metric baseline between moving cameras. In this paper, we present a novel scale estimation approach based on an object-level SLAM system. To obtain the absolute scale of the reconstructed map, we formulate an optimiza…

Cited by 7SourceScholar
2021

Robust Dual Quadric Initialization for Forward-Translating Camera Movements

RA-L 2021

Herein, we present a novel approach for monocular dual quadric initialization that combines three-dimensional (3D) map points with two-dimensional (2D) object detection for forward-translating camera movements. The traditional approach using 2D detection bounding boxes in multiple views fails in str

Cited by 11SourceScholar
2020

Modeling Behavior as Mutual Dependency between Physiological Signals and Indoor Location in Large-Scale Wearable Sensor Study

ICASSP 2020accepted

Wearable sensors today can unobtrusively collect rich time-series of physiological states and human movement patterns over a prolonged period. Gaining a better understanding of how an individual's physiological responses vary in different workplace environments can be valuable in understanding human…

Cited by 0SourceScholar
2020

Modeling Behavioral Consistency in Large-Scale Wearable Recordings of Human Bio-Behavioral Signals

ICASSP 2020accepted

Continuously-worn wearable sensors provide an unprecedented opportunity to unobtrusively measure rich bio-behavioral time-series recordings in natural settings such as the workplace. These time-series data can be helpful in inferring broad patterns of behavior such as common routines and daily stres…

Cited by 0SourceScholar
2019

Discovering Optimal Variable-length Time Series Motifs in Large-scale Wearable Recordings of Human Bio-behavioral Signals

ICASSP 2019accepted

Continuously-worn wearable sensors produce copious amounts of rich bio-behavioral time series recordings. Exploring recurring patterns, often known as motifs, in wearable time series offers critical insights into understanding the nature of human behavior. Challenges in discovering motifs from weara…

Cited by 0SourceScholar
2019

Toward Robust Interpretable Human Movement Pattern Analysis in a Workplace Setting

ICASSP 2019accepted

Gaining a better understanding of how people move about and interact with their environment is an important piece of understanding human behavior. Careful analysis of individuals' deviations or variations in movement over time can provide an awareness about changes to their physical or mental state…

Cited by 0SourceScholar