← Search

Yanqing Liu

13 accepted papers

2026

CAST: Modeling Visual State Transitions for Consistent Video Retrieval

ICML 2026poster

As video content creation shifts towards long-form narratives, retrieving and composing short clips into coherent storylines becomes a critical challenge. Standard retrieval formulations, however, perform context-agnostic retrieval, prioritizing local semantic alignment while neglecting procedural s…

Cited by 0SourceScholar
2026

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning

CVPR 2026

This paper provides a simplification on OpenVision's architecture and loss design for enhancing its training efficiency. Following the prior vision-language pretraining works CapPa and AIMv2, as well as modern multimodal designs like LLaVA, our changes are straightforward: we remove the text encoder

Cited by 0SourcecodeScholar
2026

STAR-1: Safer Alignment of Reasoning LLMs with 1K Data

AAAI 2026technical

This paper introduces STAR-1, a high-quality, just-1k-scale safety dataset specifically designed for large reasoning models (LRMs) like DeepSeek-R1. Built on three core principles --- diversity, deliberative reasoning, and rigorous filtering --- STAR-1 aims to address the critical needs for safety a

Cited by 0SourcePDFScholar
2025

Autoregressive Speech Synthesis without Vector Quantization

ACL 2025long

We present MELLE, a novel continuous-valued token based language modeling approach for text-to-speech synthesis (TTS). MELLE autoregressively generates continuous mel-spectrogram frames directly from text condition, bypassing the need for vector quantization, which is typically designed for audio co…

2025

OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning

ICCV 2025poster

OpenAI's CLIP models, released in early 2021, have long been the only viable choice for the research community in building multimodal foundation models. This dominance has only recently been challenged by a few alternatives like SigLIP. However, to the best of our knowledge, all these solutions are…

2024

Self-supervised Scale Recovery for Decoupled Visual-inertial Odometry

RA-L 2024

Accurate localization for intelligent robots remains a significant challenge, and self-supervised visual-inertial odometry (VIO) has emerged as a promising solution. However, existing self-supervised VIO works consider inertial information as the ordinary data input, losing its ability to recover ab

Cited by 3SourceScholar
2023

DREAM: Efficient Dataset Distillation by Representative Matching

ICCV 2023poster

Dataset distillation aims to synthesize small datasets with little information loss from original large-scale ones for reducing storage and training costs. Recent state-of-the-art methods mainly constrain the sample synthesis process by matching synthetic images and the original ones regarding gradi…

Cited by 91PDFcodeScholar
2023

Fast Extrinsic Calibration for Multiple Inertial Measurement Units in Visual-Inertial System

ICRA 2023poster

In this paper, we propose a fast extrinsic calibration method for fusing multiple inertial measurement units (MIMU) to improve visual-inertial odometry (VIO) localization accuracy. Currently, data fusion algorithms for MIMU highly depend on the number of inertial sensors. Based on the assumption tha…

Cited by 4SourceScholar
2023

Improving Contextual Spelling Correction by External Acoustics Attention and Semantic Aware Data Augmentation

ICASSP 2023accepted

We previously proposed contextual spelling correction (CSC) to correct the output of end-to-end (E2E) automatic speech recognition (ASR) models with contextual information such as name, place, etc. Although CSC has achieved reasonable improvement in the biasing problem, there are still two drawbacks…

Cited by 0SourceScholar
2022

Spatiotemporally Enhanced Photometric Loss for Self-Supervised Monocular Depth Estimation

IROS 2022poster

Recovering depth information from a single image is a long-standing challenge, and self-supervised depth estimation methods have gradually attracted attention due to not relying on high-cost ground truth. Constructing an accurate photometric loss based on photometric consistency is crucial for these…

Cited by 7SourceScholar
2021

AdaSpeech: Adaptive Text to Speech for Custom Voice

ICLR 2021poster

Custom voice, a specific text to speech (TTS) service in commercial speech platforms, aims to adapt a source TTS model to synthesize personal voice for a target speaker using few speech from her/him. Custom voice presents two unique challenges for TTS adaptation: 1) to support diverse customers, the…

2021

Camera Parameters Aware Motion Segmentation Network with Compensated Optical Flow

IROS 2021poster

Learning to distinguish independent moving objects from the observed optical flow with a moving camera remains challenging. In this work, we first present a novel camera pose compensation (CPC) scheme. With the help of ingenious geometric analysis, it breaks the observed optical flow into patterns t…

Cited by 2SourceScholar
2020

Richer Aggregated Features for Optical Flow Estimation with Edge-aware Refinement

IROS 2020poster

Recent CNN-based optical flow approaches have a separated structure of feature extraction and flow estimation. The core task of optical flow is finding the corresponding points while rich representation is just the key part of such matching problems. However, the prior work usually pays more attenti…

Cited by 1SourceScholar