← Search

Yu Cao

44 accepted papers

2026

Aligning Collaborative View Recovery and Tensorial Subspace Learning via Latent Representation for Incomplete Multi-View Clustering

ICLR 2026poster

Multi-view data usually suffer from partially missing views in open scenarios, which inevitably degrades clustering performance. The incomplete multi-view clustering (IMVC) has attracted increasing attention and achieved significant success. Although existing imputation-based IMVC methods perform we…

Cited by 0SourceScholar
2026

Enhancing Vision Transformers for Object Detection via Context-Aware Token Selection and Packing

ICLR 2026poster

In recent years, the long-range attention mechanism of vision transformers has driven significant performance breakthroughs across various computer vision tasks. However, these advancements come at the cost of inefficiency and substantial computational expense, especially when dealing with sparse da…

Cited by 0SourceScholar
2026

GUI-Spotlight: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding

ICML 2026poster

Multimodal large language models (MLLMs) have markedly expanded the competence of graphical user-interface (GUI) systems, propelling them beyond controlled simulations into complex, real-world environments across diverse platforms. However, practical usefulness is still bounded by the reliability of…

Cited by 0SourceScholar
2026

LiteVSR: Enabling Cross-Domain Fine-Grained Detail Generation in Light-Weight Transformers for Video Super-Resolution

ICML 2026poster

Large-scale pre-trained video generators offer powerful priors for Video Super-Resolution (VSR), yet adapting them remains computationally prohibitive. Full fine-tuning demands extensive resources, and ControlNet-style adapters lose their efficiency advantage under modern Diffusion Transformers (DiT…

Cited by 0SourceScholar
2026

PAPL-SLAM: Principal Axis-Anchored Monocular Point-Line SLAM

ICRA 2026poster

In point-line Simultaneous Localization and Mapping (SLAM) systems, the utilization of line structural information and the optimization of lines are two significant problems. The former is usually addressed through structural regularities, while the latter typically involves using minimal parameter …

2025

A Wearable Centaur Robot with Wheel-Legged Transformation for Enhanced Load-Carrying Assistance

IROS 2025

The execution of long-distance load-carrying tasks across multiple terrains remains a frequent requirement. These tasks often involve heavy loads, resulting in fatigue, decreased efficiency, and potential safety risks. To address this issue, this paper proposes a wearable centaur robot with wheel-le

Cited by 0SourceScholar
2025

AIM-Fair: Advancing Algorithmic Fairness via Selectively Fine-Tuning Biased Models with Contextual Synthetic Data

CVPR 2025poster

Recent advances in generative models have sparked research on improving model fairness with AI-generated data. However, existing methods often face limitations in the diversity and quality of synthetic data, leading to compromised fairness and overall model accuracy. Moreover, many approaches rely o…

2025

Hand1000: Generating Realistic Hands from Text with Only 1,000 Images

AAAI 2025technical

Text-to-image generation models have achieved remarkable advancements in recent years, aiming to produce realistic images from textual descriptions. However, these models often struggle with generating anatomically accurate representations of human hands. The resulting images frequently exhibit issu…

Cited by 4SourcePDFScholar
2025

Mozart: Modularized and Efficient MoE Training on 3.5D Wafer-Scale Chiplet Architectures

NeurIPS 2025spotlight

Mixture-of-Experts (MoE) architecture offers enhanced efficiency for Large Language Models (LLMs) with modularized computation, yet its inherent sparsity poses significant hardware deployment challenges, including memory locality issues, communication overhead, and inefficient computing resource uti…

Cited by 0SourceScholar
2025

Occult: Optimizing Collaborative Communications across Experts for Accelerated Parallel MoE Training and Inference

ICML 2025poster

Mixture-of-experts (MoE) architectures could achieve impressive computational efficiency with expert parallelism, which relies heavily on all-to-all communication across devices. Unfortunately, such communication overhead typically constitutes a significant portion of the total runtime, hampering th…

Cited by 0SourcePDFScholar
2025

Temporal Score Analysis for Understanding and Correcting Diffusion Artifacts

CVPR 2025poster

Visual artifacts remain a persistent challenge in diffusion models, even with training on massive datasets. Current solutions primarily rely on supervised detectors, yet lack understanding of why these artifacts occur in the first place. In our analysis, we identify three distinct phases in the diff…

Cited by 0SourcePDFScholar
2025

Uncertainty-Based Extensible Codebook for Discrete Federated Learning in Heterogeneous Data Silos

ICML 2025poster

Federated learning (FL), aimed at leveraging vast distributed datasets, confronts a crucial challenge: the heterogeneity of data across different silos. While previous studies have explored discrete representations to enhance model generalization across minor distributional shifts, these approaches…

2025

Variable Impedance Control for Floating-Base Supernumerary Robotic Leg in Walking Assistance

RA-L 2025

In human-robot systems, ensuring safety during force control in the presence of both internal and external disturbances is crucial. As a typical loosely coupled floating-base robot system, the supernumerary robotic leg (SRL) system is particularly susceptible to strong internal disturbances. To addr

Cited by 1SourceScholar
2024

Dual Semantic Fusion Hashing for Multi-Label Cross-Modal Retrieval

IJCAI 2024poster

Cross-modal hashing (CMH) has been widely used for multi-modal retrieval tasks due to its low storage cost and fast query speed. Although existing CMH methods achieve promising performance, most of them mainly rely on coarse-grained supervision information (\ie pairwise similarity matrix) to measure…

Cited by 4SourcePDFScholar
2024

Meta-Task Prompting Elicits Embeddings from Large Language Models

ACL 2024long

We introduce a new unsupervised text embedding method, Meta-Task Prompting with Explicit One-Word Limitation (MetaEOL), for generating high-quality sentence embeddings from Large Language Models (LLMs) without the need for model fine-tuning. Leveraging meta-task prompting, MetaEOL guides LLMs to pro…

2024

Transformer-Based Selective Super-resolution for Efficient Image Refinement

AAAI 2024technical

Conventional super-resolution methods suffer from two drawbacks: substantial computational cost in upscaling an entire large image, and the introduction of extraneous or potentially detrimental information for downstream computer vision tasks during the refinement of the background. To solve these i…

2024

Two-Stage Video Shadow Detection via Temporal-Spatial Adaption

ECCV 2024poster

"Video Shadow Detection (VSD) is an important computer vision task focusing on detecting and segmenting shadows throughout the entire video sequence. Despite their remarkable performance, existing VSD methods and datasets mainly focus on the dominant and isolated shadows. Consequently, VSD under com…

2023

Exploring the Optimal Choice for Generative Processes in Diffusion Models: Ordinary vs Stochastic Differential Equations

NeurIPS 2023poster

The diffusion model has shown remarkable success in computer vision, but it remains unclear whether the ODE-based probability flow or the SDE-based diffusion model is more superior and under what circumstances. Comparing the two is challenging due to dependencies on data distributions, score trainin…

Cited by 13SourcePDFScholar
2023

Stable Station Keeping of Autonomous Sailing Robots via the Switched Systems Approach for Ocean Observation

ICRA 2023poster

Ocean observation is an emerging field, and sailing robots have several promising features (e.g., long-range sailing, environmental friendliness, energy-saving and low-noise) to perform tasks. In this paper, we define an ocean observation mission in a restricted target area as a station keeping prob…

Cited by 7SourceScholar
2023

Unsupervised Dense Retrieval with Relevance-Aware Contrastive Pre-Training

ACL 2023findings

Dense retrievers have achieved impressive performance, but their demand for abundant training data limits their application scenarios. Contrastive pre-training, which constructs pseudo-positive examples from unlabeled data, has shown great potential to solve this problem. However, the pseudo-positiv…

2022

A Model-agnostic Data Manipulation Method for Persona-based Dialogue Generation

ACL 2022long

Towards building intelligent dialogue agents, there has been a growing interest in introducing explicit personas in generation models. However, with limited persona-based dialogue data at hand, it may be difficult to train a dialogue generation model well. We point out that the data challenges of th…

2022

Gradient-Based Novelty Detection Boosted by Self-Supervised Binary Classification

AAAI 2022technical

Novelty detection aims to automatically identify out-of-distribution (OOD) data, without any prior knowledge of them. It is a critical step in data monitoring, behavior analysis and other applications, helping enable continual learning in the field. Conventional methods of OOD detection perform mult…

Cited by 16SourcePDFScholar
2022

Interpretable Proof Generation via Iterative Backward Reasoning

NAACL 2022long

We present IBR, an Iterative Backward Reasoning model to solve the proof generation tasks on rule-based Question Answering (QA), where models are required to reason over a series of textual rules and facts to find out the related proof path and derive the final answer. We handle the limitations of e…

2022

Metabolic Efficiency Improvement of Human Walking by Shoulder Stress Reduction through Load Transfer Backpack

IROS 2022poster

The dynamic load attached to the load gravity imposes an excessive burden to human shoulders during load carriage, resulting in possible muscle injuries and additional physical exertion. This paper proposes an active suspension backpack, capable of transferring partial load from human shoulders to p…

Cited by 6SourceScholar
2022

On the Complementarity between Pre-Training and Random-Initialization for Resource-Rich Machine Translation

COLING 2022main

Pre-Training (PT) of text representations has been successfully applied to low-resource Neural Machine Translation (NMT). However, it usually fails to achieve notable gains (some- times, even worse) on resource-rich NMT on par with its Random-Initialization (RI) counterpart. We take the first step t…

2022

Phrase-level Textual Adversarial Attack with Label Preservation

NAACL 2022findings

Generating high-quality textual adversarial examples is critical for investigating the pitfalls of natural language processing (NLP) models and further promoting their robustness. Existing attacks are usually realized through word-level or sentence-level perturbations, which either limit the perturb…

2022

TASA: Deceiving Question Answering Models by Twin Answer Sentences Attack

EMNLP 2022main

We present Twin Answer Sentences Attack (TASA), an adversarial attack method for question answering (QA) models that produces fluent and grammatical adversarial contexts while maintaining gold answers. Despite phenomenal progress on general adversarial attacks, few works have investigated the vulner…

2021

A High-accuracy Framework for Vehicle Dynamic Modeling in Autonomous Driving

IROS 2021poster

Vehicle dynamic models are the key to bridge the gap between simulation and real road test in autonomous driving. An accurate vehicle model allows control algorithms in simulation being transferred to real road test with same quality. In this paper, we present a dynamic model residual correction fra…

Cited by 3SourceScholar
2021

DAGN: Discourse-Aware Graph Network for Logical Reasoning

NAACL 2021long

Recent QA with logical reasoning questions requires passage-level relations among the sentences. However, current approaches still focus on sentence-level relations interacting among tokens. In this work, we explore aggregating passage-level clues for solving logical reasoning QA by using discourse-…

2021

Experimental Validation of Unsteady Wave Induced Loads on a Stationary Remotely Operated Vehicle

ICRA 2021poster

Shallow water environments pose daunting scenarios for the operation of Unmanned Underwater Vehicles (UUVs), due to significantly larger wave disturbances being present in comparison to a typical deep sea situation. Performing inspection and maintenance tasks at close quarters in these conditions re…

Cited by 7SourceScholar
2021

Experimental Validation of Wave Induced Disturbances for Predictive Station Keeping of a Remotely Operated Vehicle

RA-L 2021

Predictive control methods can substantially improve the performance of Unmanned Underwater Vehicles (UUVs), particularly in shallow water environments or near the free surface where wave induced disturbance are of magnitude comparable to the vehicle characteristic inertia. To facilitate the adoptio

Cited by 40SourceScholar
2021

Improving Empathetic Response Generation by Recognizing Emotion Cause in Conversations

EMNLP 2021finding

Current approaches to empathetic response generation focus on learning a model to predict an emotion label and generate a response based on this label and have achieved promising results. However, the emotion cause, an essential factor for empathetic responding, is ignored. The emotion cause is a st…

Cited by 116SourcePDFScholar
2021

Reasoning Operational Decisions for Robots via Time Series Causal Inference

ICRA 2021poster

Justifying operational decisions for robots is a challenging task as the operator or the robot itself has to understand the underlying physical interaction between the robot and the environment to predict the potential outcome. It is desirable to understand how the decision influences the operationa…

Cited by 9SourceScholar
2021

Robust Underwater Visual SLAM Fusing Acoustic Sensing

ICRA 2021poster

In this paper, we propose an approach for robust visual Simultaneous Localisation and Mapping (SLAM) in underwater environments leveraging acoustic, inertial and altimeter/depth sensors. Underwater visual SLAM is challenging due to factors including poor visibility caused by suspended particles in w…

Cited by 61SourceScholar
2021

Towards Efficiently Diversifying Dialogue Generation Via Embedding Augmentation

ICASSP 2021accepted

Dialogue generation models face the challenge of producing generic and repetitive responses. Unlike previous augmentation methods that mostly focus on token manipulation and ignore the essential variety within a single sample using hard labels, we propose to promote the generation diversity of the n…

Cited by 0SourceScholar
2020

Deep Fashion3D: A Dataset and Benchmark for 3D Garment Reconstruction from Single Images

ECCV 2020poster

High-fidelity clothing reconstruction is the key to achieving photorealism in a wide range of applications including human digitization, virtual try-on, etc. Recent advances in learning-based approaches have accomplished unprecedented accuracy in recovering unclothed human shape and pose from single…

2020

Efficient and Modularized Training on FPGA for Real-time Applications

IJCAI 2020poster

Training of deep Convolution Neural Networks (CNNs) requires a tremendous amount of computation and memory and thus, GPUs are widely used to meet the computation demands of these complex training tasks. However, lacking the flexibility to exploit architectural optimizations, GPUs have poor energy ef…

2018

Towards a Wearable Cough Detector Based on Neural Networks

ICASSP 2018accepted

Persistent cough is a symptom common to a number of respiratory disorders; however, reliable monitoring of cough frequency and cough severity over an extended period of time can be a challenge. Traditional methods involve subjective evaluation by care providers or patient self-reports. As an alterna…

Cited by 0SourceScholar
2016

Ranking the parameters of deep neural networks using the fisher information

ICASSP 2016accepted

The large number of parameters in deep neural networks (DNNs) often makes them prohibitive for low-power devices, such as field-programmable gate arrays (FPGA). In this paper, we propose a method to determine the relative importance of all network parameters by measuring the amount of information th…

Cited by 0SourceScholar