← Search

Ye Tao

7 accepted papers

2025

GSV3D: Gaussian Splatting-based Geometric Distillation with Stable Video Diffusion for Single-Image 3D Object Generation

ICCV 2025poster

Image-based 3D generation has vast applications in robotics and gaming, where high-quality, diverse outputs and consistent 3D representations are crucial. However, existing methods have limitations: 3D diffusion models are limited by dataset scarcity and the absence of strong pre-trained priors, whi…

2025

SiQA: A Large Multi-Modal Question Answering Model for Structured Images Based on RAG

ICASSP 2025accepted

Existing Large Multimodal Models (LMMs) demonstrate excellent performance in handling visual tasks in everyday scenarios. However, they still face challenges in understanding structured images, such as flowcharts and organizational charts, which are characterized by text-rich and complex hierarchica…

Cited by 0SourceScholar
2024

A Fast and High-quality Text-to-Speech Method with Compressed Auxiliary Corpus and Limited Target Speaker Corpus

COLING 2024main

With an auxiliary corpus (non-target speaker corpus) for model pre-training, Text-to-Speech (TTS) methods can generate high-quality speech with a limited target speaker corpus. However, this approach comes with expensive training costs. To overcome the challenge, a high-quality TTS method is propose…

Cited by 0SourcePDFScholar
2023

General Category Network: Handwritten Mathematical Expression Recognition with Coarse-Grained Recognition Task

ICASSP 2023accepted

Handwritten Mathematical Expression Recognition (HMER) is an important task in pattern recognition. It is a challenging task due to symbols resembling each other in appearance("z/2", "B/β") and the complex mathematical syntax. The encoder-decoder architecture has been widely used in recent HMER meth…

Cited by 0SourceScholar
2015

Semi-supervised online learning for efficient classification of objects in 3D data streams

IROS 2015poster

We present a novel learning algorithm especially designed for challenging, large-scale classification problems in mobile robotics. Our method addresses two important aims: first it reduces the required amount of interaction with a human supervisor, which increases the level of autonomy of the learni…

Cited by 11SourceScholar