← Search

Yang Bai

47 accepted papers

2026

240FPS Stereo Vision from Monocular Mixed Spikes

CVPR 2026

Stereo vision is fundamental for enabling machines to perceive and interact with the world. While monocular stereo methods offer hardware compactness, they struggle with generalization due to reliance on data-driven priors. Binocular and multi-view systems improve accuracy but incur higher hardware

Cited by 0SourcecodeScholar
2026

CoVAR: Co-Generation of Video and Action for Robotic Manipulation Via Multi-Modal Diffusion

ICRA 2026poster

We present a method to generate video–action pairs that follow text instructions, starting from an initial image observation and the robot’s joint states. Our approach automatically provides action labels for video diffusion mod- els, overcoming the common lack of action annotations and enabling the…

2026

DRAW2ACT: Turning Depth-Encoded Trajectories into Robotic Demonstration Videos

ICRA 2026poster

Video diffusion models provide powerful real-world simulators for embodied AI but remain limited in controllability for robotic manipulation. Recent works on trajectory-conditioned video generation address this gap but often rely on 2D trajectories or single modality conditioning, which restricts th…

2026

Interactive Person Retrieval via Multi-Turn Multimodal Conversation

ICML 2026poster

Traditional text-based person retrieval approaches typically rely on single-shot textual queries, which are generally incomplete or vague in real-world scenarios. Recently, chat-based person retrieval methods enable iterative query refinement via question-answering interactions between the system an…

Cited by 0SourceScholar
2026

Lightweight Learning From Actuation-Space Demonstrations via Flow Matching for Whole-Body Soft Robotic Grasping

RA-L 2026

Robotic grasping under uncertainty remains a fundamental challenge due to its uncertain and contact-rich nature. Traditional rigid robotic hands, with limited degrees of freedom and compliance, rely on complex model-based and heavy feedback controllers to manage such interactions. Soft robots, by co

Cited by 0SourceScholar
2026

Note2Chat: Improving LLMs for Multi-Turn Clinical History Taking Using Medical Notes

AAAI 2026technical

Effective clinical history taking is a foundational yet underexplored component of clinical reasoning. While large language models (LLMs) have shown promise on static benchmarks, they often fall short in dynamic, multi-turn diagnostic settings that require iterative questioning and hypothesis refine

Cited by 0SourcePDFScholar
2026

VideoWeaver: Multimodal Multi-View Video-to-Video Transfer for Embodied Agents

CVPR 2026

Recent progress in video-to-video (V2V) translation has enabled realistic resimulation of embodied AI demonstrations, a capability that allows pretrained robot policies to be transferable to new environments without additional data collection. However, prior works can only operate on a single view a

Cited by 0SourceScholar
2025

Chat-based Person Retrieval via Dialogue-Refined Cross-Modal Alignment

CVPR 2025poster

Traditional text-based person retrieval (TPR) relies on a single-shot text as query to retrieve the target person, assuming that the query completely captures the user's search intent. However, in real-world scenarios, it can be challenging to ensure the information completeness of such single-shot…

2025

Image-level Memorization Detection via Inversion-based Inference Perturbation

ICLR 2025poster

Recent studies have discovered that widely used text-to-image diffusion models can replicate training samples during image generation, a phenomenon known as memorization. Existing detection methods primarily focus on identifying memorized prompts. However, in real-world scenarios, image owners may n…

Cited by 0SourcePDFScholar
2025

Protecting Your Video Content: Disrupting Automated Video-based LLM Annotations

CVPR 2025poster

Recently, video-based large language models (video-based LLMs) have achieved impressive performance across various video comprehension tasks. However, this rapid advancement raises significant privacy and security concerns, particularly regarding the unauthorized use of personal video data in automa…

2025

RAMQA: A Unified Framework for Retrieval-Augmented Multi-Modal Question Answering

NAACL 2025findings

Multi-modal retrieval-augmented Question Answering (MRAQA), integrating text and images, has gained significant attention in information retrieval (IR) and natural language processing (NLP). Traditional ranking methods rely on small encoder-based language models, which are incompatible with modern d…

2025

RoboEnvision: A Long-Horizon Video Generation Model for Multi-Task Robot Manipulation

IROS 2025

We address the problem of generating long-horizon videos for robotic manipulation tasks. Text-to-video diffusion models have made significant progress in photorealism, language understanding, and motion generation but struggle with long-horizon robotic tasks. Recent works use video diffusion models

Cited by 11SourceScholar
2025

RoboSwap: A GAN-driven Video Diffusion Framework For Unsupervised Robot Arm Swapping

IROS 2025

Recent advancements in generative models have revolutionized video synthesis and editing. However, the scarcity of diverse, high-quality datasets continues to hinder video-conditioned robotic learning, limiting cross-platform generalization. In this work, we address the challenge of swapping a robot

Cited by 1SourceScholar
2025

SING: Spatial Context in Large Language Model for Next-Gen Wearables

ICML 2025poster

Integrating spatial context into large language models (LLMs) has the potential to revolutionize human-computer interaction, particularly in wearable devices. In this work, we present a novel system architecture that incorporates spatial speech understanding into LLMs, enabling contextually aware an…

Cited by 0SourcePDFScholar
2025

VQA4CIR: Boosting Composed Image Retrieval with Visual Question Answering

AAAI 2025technical

Albeit progress has been made in Composed Image Retrieval (CIR), we empirically find that a certain percentage of failure retrieval results are not consistent with their relative captions. To address this issue, this work provides a Visual Question Answering (VQA) perspective to boost the performanc…

2024

An Empirical Study of CLIP for Text-Based Person Search

AAAI 2024technical

Text-based Person Search (TBPS) aims to retrieve the person images using natural language descriptions. Recently, Contrastive Language Image Pretraining (CLIP), a universal large cross-modal vision-language pre-training model, has remarkably performed over various cross-modal downstream tasks due to…

2024

Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images

ICLR 2024poster

Large vision-language models (VLMs) such as GPT-4 have achieved exceptional performance across various multi-modal tasks. However, the deployment of VLMs necessitates substantial energy consumption and computational resources. Once attackers maliciously induce high energy consumption and latency tim…

2024

M3: A Multi-Task Mixed-Objective Learning Framework for Open-Domain Multi-Hop Dense Sentence Retrieval

COLING 2024main

In recent research, contrastive learning has proven to be a highly effective method for representation learning and is widely used for dense retrieval. However, we identify that relying solely on contrastive learning can lead to suboptimal retrieval performance. On the other hand, despite many retri…

2024

Sentence-level Prompts Benefit Composed Image Retrieval

ICLR 2024spotlight

Composed image retrieval (CIR) is the task of retrieving specific images by using a query that involves both a reference image and a relative caption. Most existing CIR models adopt the late-fusion strategy to combine visual and language features. Besides, several approaches have also been suggested…

2024

Solving General Noisy Inverse Problem via Posterior Sampling: A Policy Gradient Viewpoint

AISTATS 2024poster

Solving image inverse problems (e.g., super-resolution and inpainting) requires generating a high fidelity image that matches the given input (the low-resolution image or the masked image). By using the input image as guidance, we can leverage a pretrained diffusion generative model to solve a wide…

2024

What Makes Good Collaborative Views? Contrastive Mutual Information Maximization for Multi-Agent Perception

AAAI 2024technical

Multi-agent perception (MAP) allows autonomous systems to understand complex environments by interpreting data from multiple sources. This paper investigates intermediate collaboration for MAP with a specific focus on exploring "good" properties of collaborative view (i.e., post-collaboration featur…

2023

ATFormer: A Learned Performance Model with Transfer Learning Across Devices for Deep Learning Tensor Programs

EMNLP 2023long main

The training and inference efficiency of ever-larger deep neural networks highly rely on the performance of tensor operators on specific hardware platforms. Therefore, a compilation-based optimization flow with automatic tensor generation and parameter tuning is necessary for efficient model deploym…

Cited by 0SourceScholar
2023

An Empirical Study of Frame Selection for Text-to-Video Retrieval

EMNLP 2023long findings

Text-to-video retrieval (TVR) aims to find the most relevant video in a large video gallery given a query text. The intricate and abundant context of the video challenges the performance and efficiency of TVR. To handle the serialized video contexts, existing methods typically select a subset of fra…

Cited by 0SourceScholar
2023

AutoGraph: Optimizing DNN Computation Graph for Parallel GPU Kernel Execution

AAAI 2023technical

Deep learning frameworks optimize the computation graphs and intra-operator computations to boost the inference performance on GPUs, while inter-operator parallelism is usually ignored. In this paper, a unified framework, AutoGraph, is proposed to obtain highly optimized computation graphs in favo…

Cited by 6SourcePDFScholar
2023

Backdoor Defense via Adaptively Splitting Poisoned Dataset

CVPR 2023poster

Backdoor defenses have been studied to alleviate the threat of deep neural networks (DNNs) being backdoor attacked and thus maliciously altered. Since DNNs usually adopt some external training data from an untrusted third party, a robust backdoor defense strategy during the training stage is of impo…

2023

Guide and Select: A Transformer-Based Multimodal Fusion Method for Points of Interest Description Generation

ICASSP 2023accepted

The task of Points of Interest (POI) description generation aims to generate an objective and informative description for a given POI based on POI-related information. High-quality descriptions can better guide users and improve the performance of POI-related recommendation systems. A practical POI…

Cited by 0SourceScholar
2023

Learning Procedure-Aware Video Representation From Instructional Videos and Their Narrations

CVPR 2023poster

The abundance of instructional videos and their narrations over the Internet offers an exciting avenue for understanding procedural activities. In this work, we propose to learn video representation that encodes both action steps and their temporal ordering, based on a large-scale dataset of web ins…

2023

RaSa: Relation and Sensitivity Aware Representation Learning for Text-based Person Search

IJCAI 2023poster

Text-based person search aims to retrieve the specified person images given a textual description. The key to tackling such a challenging task is to learn powerful multi-modal representations. Towards this, we propose a Relation and Sensitivity aware representation learning method (RaSa), including…

2022

Action Quality Assessment with Temporal Parsing Transformer

ECCV 2022poster

"Action Quality Assessment(AQA) is important for action understanding and resolving the task poses unique challenges due to subtle visual differences. Existing state-of-the-art methods typically rely on the holistic video representations for score regression or ranking, which limits the generalizati…

Cited by 60SourcePDFScholar
2022

PCL: Proxy-Based Contrastive Learning for Domain Generalization

CVPR 2022poster

Domain generalization refers to the problem of training a model from a collection of different source domains that can directly generalize to the unseen target domains. A promising solution is contrastive learning, which attempts to learn domain-invariant representations by exploiting rich semantic…

Cited by 157PDFcodeScholar
2022

Untargeted Backdoor Watermark: Towards Harmless and Stealthy Dataset Copyright Protection

NeurIPS 2022accept

Deep neural networks (DNNs) have demonstrated their superiority in practice. Arguably, the rapid development of DNNs is largely benefited from high-quality (open-sourced) datasets, based on which researchers and developers can easily evaluate and improve their learning methods. Since the data collec…

2022

Watermark Vaccine: Adversarial Attacks to Prevent Watermark Removal

ECCV 2022poster

"As a common security tool, visible watermarking has been widely applied to protect copyrights of digital images. However, recent works have shown that visible watermarks can be removed by DNNs without damaging their host images. Such watermark-removal techniques pose a great threat to the ownership…

2021

Clustering Effect of Adversarial Robust Models

NeurIPS 2021spotlight

Adversarial robustness has received increasing attention along with the study of adversarial examples. So far, existing works show that robust models not only obtain robustness against various adversarial attacks but also boost the performance in some downstream tasks. However, the underlying mechan…

2021

Fast and Efficient DNN Deployment via Deep Gaussian Transfer Learning

ICCV 2021poster

Deep neural networks (DNNs) have been widely used recently while their hardware deployment optimizations are very time-consuming and the historical deployment knowledge is not utilized efficiently. In this paper, to accelerate the optimization process and find better deployment configurations, we pr…

Cited by 7PDFScholar
2021

Improving Adversarial Robustness via Channel-wise Activation Suppressing

ICLR 2021spotlight

The study of adversarial examples and their activations have attracted significant attention for secure and robust learning with deep neural networks (DNNs). Different from existing works, in this paper, we highlight two new characteristics of adversarial examples from the channel-wise activation p…

2020

Improving Query Efficiency of Black-box Adversarial Attack

ECCV 2020poster

Deep neural networks (DNNs) have demonstrated excellent performance on various tasks, however they are under the risk of adversarial examples that can be easily generated when the target model is accessible to an attacker (white-box setting). As plenty of machine learning models have been deployed v…

2020

Infobox-to-text Generation with Tree-like Planning based Attention Network

IJCAI 2020poster

We study the problem of infobox-to-text generation that aims to generate a textual description from a key-value table. Representing the input infobox as a sequence, previous neural methods using end-to-end models without order-planning suffer from the problems of incoherence and inadaptability to di…

Cited by 0SourcePDFScholar
2019

Hilbert-Based Generative Defense for Adversarial Examples

ICCV 2019poster

Adversarial perturbations of clean images are usually imperceptible for human eyes, but can confidently fool deep neural networks (DNNs) to make incorrect predictions. Such vulnerability of DNNs raises serious security concerns about their practicability in security-sensitive applications. To defend…

Cited by 62PDFScholar
2017

Adaptive trajectory tracking control for the ball-pendulum system with time-varying uncertainties

IROS 2017poster

An adaptive trajectory tracking problem for a spherical rolling robot driven by a 2DOF pendulum is considered in this paper. A feedback controller is proposed for the goal of tracking the trajectory for the full configuration of the spherical robot. To deal with time-varying uncertainty of the syste…

Cited by 17SourceScholar
2015

Dynamic model and motion planning for a pendulum-actuated spherical rolling robot

ICRA 2015poster

This paper deals with the dynamics and motion planning for a spherical rolling robot with a pendulum actuated by two motors. First a dynamic model for the rolling robot is established. In general, not all feasible kinematic trajectories of the rolling carrier are dynamically realizable. A notable ex…

Cited by 29SourceScholar
2015

Motion planning for a pendulum-driven rolling robot tracing spherical contact curves

IROS 2015poster

This paper deals with motion planning problems for spherical rolling robots driven by a two degree of freedom pendulum. A full dynamic model for this system is first introduced. Then, assuming that the contact path is specified and the sphere moves in pure rolling mode, the full dynamic model is red…

Cited by 9SourceScholar