← Search

Kenan Dai

4 accepted papers

2026

TTS Can Speak in Any Style with Any Voice

ICLR 2026poster

This study proposes FlexiVoice, a text-to-speech (TTS) synthesis system capable of flexible style control with zero-shot voice cloning. The speaking style is controlled by a natural-language instruction and the voice timbre is provided by a speech reference in zero-shot manner. FlexiVoice is built w…

Cited by 0SourcecodeScholar
2021

Video Annotation for Visual Tracking via Selection and Refinement

ICCV 2021poster

Deep learning based visual trackers entail offline pre-training on large volumes of video datasets with accurate bounding box annotations that are labor-expensive to achieve. We present a new framework to facilitate bounding box annotations for video sequences, which investigates a selection-and-ref…

Cited by 11PDFcodeScholar
2020

High-Performance Long-Term Tracking With Meta-Updater

CVPR 2020oral

Long-term visual tracking has drawn increasing attention because it is much closer to practical applications than short-term tracking. Most top-ranked long-term trackers adopt the offline-trained Siamese architectures, thus,they cannot benefit from great progress of short-term trackers with online u…

Cited by 319PDFcodeScholar
2019

Visual Tracking via Adaptive Spatially-Regularized Correlation Filters

CVPR 2019oral

In this work, we propose a novel adaptive spatially-regularized correlation filters (ASRCF) model to simultaneously optimize the filter coefficients and the spatial regularization weight. First, this adaptive spatial regularization scheme could learn an effective spatial weight for a specific object…

Cited by 498PDFcodeScholar