← Search

Tian Tan

24 accepted papers

2026

Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking Tokens

ICML 2026poster

Large language models (LLMs) have demonstrated impressive reasoning capabilities by scaling test-time compute via long Chain-of-Thought (CoT). However, recent findings suggest that raw token counts are unreliable proxies for reasoning quality: increased generation length does not consistently correl…

Cited by 0SourceScholar
2024

AddBiomechanics Dataset: Capturing the Physics of Human Motion at Scale

ECCV 2024poster

"While reconstructing human poses in 3D from inexpensive sensors has advanced significantly in recent years, quantifying the dynamics of human motion, including the muscle-generated joint torques and external forces, remains a challenge. Prior attempts to estimate physics from reconstructed human po…

Cited by 4SourcePDFScholar
2024

Bridging The Domain Gap Arising from Text Description Differences for Stable Text-To-Image Generation

ICASSP 2024accepted

Generating high-quality images that conform to the semantics of captions has numerous potential applications. However, text-to-image generation is a challenging task due to its cross-modality nature. Current generative models are typically unstable, meaning that complex sentences can result in poor…

Cited by 0SourceScholar
2024

Connecting Speech Encoder and Large Language Model for ASR

ICASSP 2024accepted

The impressive capability and versatility of large language models (LLMs) have aroused increasing attention in automatic speech recognition (ASR), with several pioneering studies attempting to build integrated ASR models by connecting a speech encoder with an LLM. This paper presents a comparative s…

Cited by 0SourceScholar
2024

Extending Large Language Models for Speech and Audio Captioning

ICASSP 2024accepted

Multimodal large language models (LLMs) have shown promising visual perception abilities by connecting with image encoders, but their performance on auditory tasks has not yet been widely investigated. Meanwhile, automatic speech recognition (ASR) and automatic audio captioning (AAC) are often achie…

Cited by 0SourceScholar
2024

SALMONN: Towards Generic Hearing Abilities for Large Language Models

ICLR 2024poster

Hearing is arguably an essential ability of artificial intelligence (AI) agents in the physical world, which refers to the perception and understanding of general auditory information consisting of at least three types of sounds: speech, audio events, and music. In this paper, we propose SALMONN, a…

2024

video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

ICML 2024poster

Speech understanding as an element of the more generic video understanding using audio-visual large language models (av-LLMs) is a crucial yet understudied aspect. This paper proposes video-SALMONN, a single end-to-end av-LLM for video processing, which can understand not only visual frame sequences…

2023

Joint Discriminator and Transfer Based Fast Domain Adaptation For End-To-End Speech Recognition

ICASSP 2023accepted

Adapting End-to-End (E2E) models to unseen domains is still a big challenge since training E2E models requires lots of paired audio and text training data. We propose a novel domain adaptation framework for the E2E model, which only uses the text of the target domain. Moreover, the proposed methods…

Cited by 0SourceScholar
2023

Multi-Modality Deep Network for Extreme Learned Image Compression

AAAI 2023technical

Image-based single-modality compression learning approaches have demonstrated exceptionally powerful encoding and decoding capabilities in the past few years , but suffer from blur and severe semantics loss at extremely low bitrates. To address this issue, we propose a multimodal machine learning me…

Cited by 18SourcePDFScholar
2021

AISpeech-SJTU ASR System for the Accented English Speech Recognition Challenge

ICASSP 2021accepted

This paper describes the AISpeech-SJTU ASR system for the Interspeech-2020 Accented English Speech Recognition Challenge (AESRC). This task is challenging due to the diversity of pronunciation accuracy, intonation speed and pronunciation of some syllables. All participants were restricted to develop…

Cited by 0SourceScholar
2021

Formulation and Validation of an Intuitive Quality Measure for Antipodal Grasp Pose Evaluation

RA-L 2021

This letter describes a novel grasp quality measure that we developed for evaluating antipodal grasp poses in real-time. To quantify the grasp quality, we compute a set of object movement features from analyzing the interaction between the gripper and the object's projections in the image space. The

Cited by 6SourceScholar
2020

Adaptability Preserving Domain Decomposition for Stabilizing Sim2Real Reinforcement Learning

IROS 2020poster

In sim-to-real transfer of Reinforcement Learning (RL) policies for robot tasks, Domain Randomization (DR) is a widely used technique for improving adaptability. However, in DR there is a conflict between adaptability and training stability, and heavy DR tends to result in instability or even failur…

Cited by 6SourceScholar
2020

Generating Adjacency-Constrained Subgoals in Hierarchical Reinforcement Learning

NeurIPS 2020spotlight

Goal-conditioned hierarchical reinforcement learning (HRL) is a promising approach for scaling up reinforcement learning (RL) techniques. However, it often suffers from training inefficiency as the action space of the high-level, i.e., the goal space, is often large. Searching in a large goal space…

2018

Generative Adversarial Networks Based Data Augmentation for Noise Robust Speech Recognition

ICASSP 2018accepted

Data augmentation is an effective method to increase the size of training data and reduce the mismatch between training and testing for noise robust speech recognition. Different from the traditional approaches by directly adding noise to the original waveform, in this work we utilize generative adv…

Cited by 0SourceScholar
2018

Knowledge Transfer in Permutation Invariant Training for Single-Channel Multi-Talker Speech Recognition

ICASSP 2018accepted

This paper proposes a framework that combines teacher-student training and permutation invariant training (PIT) for single-channel multi-talker speech recognition. In contrast to most of conventional teacher-student training methods that aim at compressing the model, the proposed method distills kno…

Cited by 0SourceScholar
2016

Integrated adaptation with multi-factor joint-learning for far-field speech recognition

ICASSP 2016accepted

Although great progress has been made in automatic speech recognition (ASR), significant performance degradation still exists in distant talking scenarios due to significantly lower signal power. In this paper, a novel adaptation framework, named integrated adaptation with multi-factor joint-learnin…

Cited by 0SourceScholar
2016

Joint acoustic factor learning for robust deep neural network based automatic speech recognition

ICASSP 2016accepted

Deep neural networks (DNNs) for acoustic modeling have been shown to provide impressive results on many state-of-the-art automatic speech recognition (ASR) applications. However, DNN performance degrades due to mismatches in training and testing conditions and thus adaptation is necessary. In this p…

Cited by 0SourceScholar
2016

Speaker-aware training of LSTM-RNNS for acoustic modelling

ICASSP 2016accepted

Long Short-Term Memory (LSTM) is a particular type of recurrent neural network (RNN) that can model long term temporal dynamics. Recently it has been shown that LSTM-RNNs can achieve higher recognition accuracy than deep feed-forword neural networks (DNNs) in acoustic modelling. However, speaker ada…

Cited by 0SourceScholar