← Search

Yu Fang

10 accepted papers

2026

SpikeTrack: A Spike-driven Framework for Efficient Visual Tracking

CVPR 2026

Spiking Neural Networks (SNNs) promise energy-efficient vision, but applying them to RGB visual tracking remains difficult: Existing SNN tracking frameworks either do not fully align with spike-driven computation or do not fully leverage neurons' spatiotemporal dynamics, leading to a trade-off betwe

Cited by 0SourcecodeScholar
2025

DCIM-AVSR: Efficient Audio-Visual Speech Recognition via Dual Conformer Interaction Module

ICASSP 2025accepted

Speech recognition is the technology that enables machines to interpret and process human speech, converting spoken language into text or commands. This technology is essential for applications such as virtual assistants, transcription services, and communication tools. The Audio-Visual Speech Recog…

Cited by 0SourceScholar
2025

DiVISe: Direct Visual-Input Speech Synthesis Preserving Speaker Characteristics And Intelligibility

NAACL 2025findings

Video-to-speech (V2S) synthesis, the task of generating speech directly from silent video input, is inherently more challenging than other speech synthesis tasks due to the need to accurately reconstruct both speech content and speaker characteristics from visual cues alone. Recently, audio-visual p…

2025

DreamGen: Unlocking Generalization in Robot Learning through Video World Models

CoRL 2025poster

In this work, we unlock new capabilities in robot learning from neural trajectories, synthetic robot data generated from video world models. Our proposed recipe is simple, but powerful: we take the most recent state-of-the-art video generative models (world models), adapt them to the target robot em…

Cited by 0SourcecodeScholar
2025

FLARE: Robot Learning with Implicit World Modeling

CoRL 2025poster

We introduce **F**uture **LA**tent **R**presentation Alignm**E**nt (**FLARE**), a novel framework that integrates predictive world modeling into robot policy learning. By aligning features from a diffusion transformer with latent embeddings of future observations, **FLARE** enables a diffusion trans…

Cited by 0SourceScholar
2025

ReBot: Scaling Robot Learning with Real-to-Sim-to-Real Robotic Video Synthesis

IROS 2025

Vision-language-action (VLA) models present a promising paradigm by training policies directly on real robot datasets like Open X-Embodiment. However, the high cost of real-world data collection hinders further data scaling, thereby restricting the generalizability of VLAs. In this paper, we introdu

Cited by 22SourceScholar
2025

Sim-and-Real Co-Training: A Simple Recipe for Vision-Based Robotic Manipulation

RSS 2025poster

Large real-world robot datasets hold great potential for developing generalist robot policies, but scaling real-world data collection is time-consuming, costly, and resource-intensive. Simulation offers a promising solution, with recent advances in generative AI and synthetic data generation tools e…

Cited by 4PDFScholar
2025

Social Robot Haru Assisting Dynamic Group Discussion with Autonomous Eye Gaze Behavior

IROS 2025

Due to recent advances in large language models and robotics, social robots will potentially play an important role in people’s daily lives soon, and are expected to improve dynamic multi-party group discussions in social scenarios. In this paper, we developed a system to assist dynamic group discus

Cited by 0SourceScholar
2022

Developing The Bottom-up Attentional System of A Social Robot

ICRA 2022poster

This paper describes the development of a 3- stage signalling framework to trigger a social robot's bottom- up reactive behavior inspired by a biological model. In the first stage, low-level firing of stimuli due to external sources is constructed through perception grounding. This is followed by a…

Cited by 7SourceScholar