← Search

Wei Yu

51 accepted papers

2026

Communication-efficient Multi-Agent Reinforcement Learning with Spatiotemporal Information Hub

AAAI 2026technical

Centralized training with decentralized execution (CTDE) is a framework for MARL with wide applications. In the CTDE paradigm, agents leverage global state information during training to mitigate the non-stationarity of the MARL environment, but must rely solely on partial observations during execut

Cited by 0SourcePDFScholar
2026

DialoGen: Towards Dialog Gesture Generation via Identity-Decoupled Style Guidance in Interactive Diffusion Model

AAAI 2026technical

We propose DialoGen, a novel framework for generating realistic gestures for both interlocutors in dialog scenarios, conditioned on conversational audios. Unlike most existing methods that focus solely on a single speaker, DialoGen simultaneously generates synchronized gestures for both participants

Cited by 0SourcePDFScholar
2026

HIVE-3D: Hierarchical Voxel Enhancement for High-Quality 3D Scene Generation

ICML 2026poster

Recently, a line of works can generate impressive 3D objects from a single image, but they are limited by restricted representation resolution, making them unsuitable for 3D scene generation. In this work, we introduce \name, a novel method for high-quality 3D scene generation based on hierarchical …

Cited by 0SourceScholar
2026

ORBIT: A Prognostic World Model for Ocular Reasoning Based on Imagined Trajectories

ICML 2026poster

The longitudinal management of blinding fundus diseases constitutes a Partially Observable Markov Decision Process (POMDP) necessitating a critical precision-risk trade-off between intervention and over-treatment, as true pathology is often obscured in static observations. However, existing paradigm…

Cited by 0SourceScholar
2026

Refining Few-Step Text-to-Multiview Diffusion via Reinforcement Learning

CVPR 2026

Text-to-multiview (T2MV) diffusion models have shown great promise in generating multiple views of a scene from a single text prompt. While few-step backbones enable real-time T2MV generation, they often compromise key aspects of generation quality, such as per-view fidelity and cross-view consisten

Cited by 0SourcecodeScholar
2025

AffPose: An Integrated RGB-Based Framework for Simultaneous Pose Estimation and Affordance Detection in Robotic Tool Manipulation

RA-L 2025

Enabling robots to perform tool manipulation like humans remains a great challenge. A semantic understanding of tool affordances and precise spatial localization is essential for this task. Conventional methods relying on RGB-D cameras for affordance detection and tool manipulation have been proven

Cited by 2SourceScholar
2025

CASleepNet: A Cross Attention-based multimodal fusion approach for sleep staging with EEG and EOG

ICASSP 2025accepted

Automatic sleep staging is crucial for sleep assessment and diagnosis. Signals of different modalities, such as electroencephalogram (EEG) and electrooculogram (EOG), are of crucial importance for sleep staging. Therefore, effective fusion of different modal signals is the key to improve sleep stagi…

Cited by 0SourceScholar
2025

Design and Performance Analysis of a Pipeline Crawling Robot Based on Spring-Roll Dielectric Elastomer Actuators

IROS 2025

With the increasing complexity of pipeline systems in various industrial and environmental applications, there is a critical need for flexible and efficient robotic solutions that can navigate and inspect confined spaces. This paper introduces a lightweight pipeline crawling robot based on spring-ro

Cited by 0SourceScholar
2025

Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

AAAI 2025technical

Multimodal large language models (MLLMs) have experienced significant advancements recently, but still struggle to recognize and interpret intricate details in high-resolution (HR) images effectively. While state-of-the-art (SOTA) MLLMs claim to process images at 4K resolution, existing MLLM benchma…

2025

EgoNet: An Unified Egocentric Active Speaker Detection Framework for both Camera Wearer and Visible Candidates

ICASSP 2025accepted

Active Speaker Detection (ASD) aims to determine whether each candidate in a video frame is speaking. The egocentric dataset Ego4D introduces unique challenges for this task, such as dynamic shooting angles that cause candidates to frequently leave the sight, leading to temporal discontinuities. Add…

Cited by 0SourceScholar
2025

EgoSim: Egocentric Exploration in Virtual Worlds with Multi-modal Conditioning

ICLR 2025poster

Recent advancements in video diffusion models have established a strong foundation for developing world models with practical applications. The next challenge lies in exploring how an agent can leverage these foundation models to understand, interact with, and plan within observed environments. This…

2025

Federated Graph Anomaly Detection Through Contrastive Learning with Global Negative Pairs

AAAI 2025technical

Anomaly detection on attributed graphs has applications in various domains such as finance and email spam detection, thus gaining substantial attention. Distributed scenarios can also involve issues related to anomaly detection in attribute graphs, such as in medical scenarios. However, most of the…

Cited by 0SourcePDFScholar
2025

MTGA: Multi-View Temporal Granularity Aligned Aggregation for Event-Based Lip-Reading

AAAI 2025technical

Lip-reading is to utilize the visual information of the speaker’s lip movements to recognize words and sentences. Existing event-based lip-reading solutions integrate different frame rate branches to learn spatio-temporal features of varying granularities. However, aggregating events into event fram…

2025

OTPNet: ODE-inspired Tuning-free Proximal Network for Remote Sensing Image Fusion

AAAI 2025technical

Remote sensing image fusion aims to reconstruct a high spatial and spectral resolution image by integrating the spatial and spectral information from multiple remote sensing sensor data. Despite the remarkable progress of deep learning-based fusion methods, most existing methods rely on manual netwo…

Cited by 0SourcePDFScholar
2025

Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose Interaction

CVPR 2025poster

Video virtual try-on aims to seamlessly dress a subject in a video with a specific garment. The primary challenge involves preserving the visual authenticity of the garment while dynamically adapting to the pose and physique of the subject. While existing methods have predominantly focused on image-…

Cited by 0SourcePDFScholar
2025

What Kind of Visual Tokens Do We Need? Training-Free Visual Token Pruning for Multi-Modal Large Language Models from the Perspective of Graph

AAAI 2025technical

Recent Multimodal Large Language Models(MLLMs) often use a large number of visual tokens to compensate their visual shortcoming, leading to excessive computation and obvious visual redundancy. In this paper, we investigate what kind of visual tokens are needed for MLLMs, and reveal that both foregro…

2024

E-GNN: An Enhanced Method for Multi-Object Tracking with Collective Motion Patterns

RA-L 2024

The long-term consistent visual tracking of large-scale moving swarms of animals or autonomous moving robots (AMR) is extremely challenging when the three factors are involved: 1) similar appearance of animals or AMR, 2) frequent and unpredictable occlusions, and 3) non-linear maneuvers. When facing

Cited by 3SourceScholar
2024

Joint Input and Output Coordination for Class-Incremental Learning

IJCAI 2024poster

Incremental learning is nontrivial due to severe catastrophic forgetting. Although storing a small amount of data on old tasks during incremental learning is a feasible solution, current strategies still do not 1) adequately address the class bias problem, and 2) alleviate the mutual interference be…

Cited by 2SourcePDFScholar
2024

Learning Scale-Aware Spatio-temporal Implicit Representation for Event-based Motion Deblurring

ICML 2024poster

Existing event-based motion deblurring methods mostly focus on restoring images with the same spatial and temporal scales as events. However, the unknown scales of images and events in the real world pose great challenges and have rarely been explored. To address this gap, we propose a novel Scale-A…

2024

Levenshtein Distance Embedding with Poisson Regression for DNA Storage

AAAI 2024technical

Efficient computation or approximation of Levenshtein distance, a widely-used metric for evaluating sequence similarity, has attracted significant attention with the emergence of DNA storage and other biological applications. Sequence embedding, which maps Levenshtein distance to a conventional dist…

Cited by 2SourcePDFScholar
2024

Rethinking Imbalance in Image Super-Resolution for Efficient Inference

NeurIPS 2024poster

Existing super-resolution (SR) methods optimize all model weights equally using $\mathcal{L}_1$ or $\mathcal{L}_2$ losses by uniformly sampling image patches without considering dataset imbalances or parameter redundancy, which limits their performance. To address this, we formulate the image SR tas…

Cited by 0SourcePDFScholar
2024

SBM: Smoothness-Based Minimization for Domain Generalization

ICASSP 2024accepted

In topical domain generalization (DG), trained models are asked to perform well on an unknown target domain with different data statistics. In order to improve domain generalization, adversarial learning has proven to be one of the most effective methods. Existing approaches, however, rely primarily…

Cited by 0SourceScholar
2024

Soft-Prompting with Graph-of-Thought for Multi-modal Representation Learning

COLING 2024main

The chain-of-thought technique has been received well in multi-modal tasks. It is a step-by-step linear reasoning process that adjusts the length of the chain to improve the performance of generated prompts. However, human thought processes are predominantly non-linear, as they encompass multiple as…

2024

Theoretical Modeling and Bio-inspired Trajectory Optimization of A Multiple-locomotion Origami Robot

IROS 2024poster

Recent research on mobile robots has focused on increasing their adaptability to unpredictable and unstructured environments using soft materials and structures. However, the determination of key design parameters and control over these compliant robots are predominantly iterated through experiments…

Cited by 1SourceScholar
2023

Enhanced Affine Formation Maneuver Control Using Historical Velocity Command (HVC)

RA-L 2023

Recent studies on the network of multi-vehicle systems have shown that the system performance can be improved comprehensively by actively using historical information without changing the network connectivity. Motivated by this observation, we aim to improve the performance of affine formation maneu

Cited by 6SourceScholar
2023

On the choice of Perception Loss Function for Learned Video Compression

NeurIPS 2023poster

We study causal, low-latency, sequential video compression when the output is subjected to both a mean squared-error (MSE) distortion loss as well as a perception loss to target realism. Motivated by prior approaches, we consider two different perception loss functions (PLFs). The first, PLF-JD, co…

Cited by 15SourcePDFScholar
2023

Parameter and Computation Efficient Transfer Learning for Vision-Language Pre-trained Models

NeurIPS 2023poster

With ever increasing parameters and computation, vision-language pre-trained (VLP) models exhibit prohibitive expenditure in downstream task adaption. Recent endeavors mainly focus on parameter efficient transfer learning (PETL) for VLP models by only updating a small number of parameters. However,…

Cited by 8SourcePDFScholar
2023

Scaling Law Analysis for Covariance Based Activity Detection in Cooperative Multi-Cell Massive Mimo

ICASSP 2023accepted

This paper studies the covariance based activity detection problem in a multi-cell massive multiple-input multiple-output (MIMO) system, where the active devices transmit their signature sequences to multiple base stations (BSs), and the BSs cooperatively detect the active devices based on the recei…

Cited by 0SourceScholar
2022

Attention-based Adversarial Partial Domain Adaptation

ICASSP 2022accepted

With the rapid development of vision-based deep learning (DL), it is an effective method to generate large-scale synthetic data to supplement real data to train the DL models for domain adaptation. However, previous vanilla domain adaptation methods generally assume the same label space, and such an…

Cited by 0SourceScholar
2022

Data-Driven Optimization for Zero-Delay Lossy Source Coding with Side Information

ICASSP 2022accepted

This paper proposes a data-driven architecture for zero-delay lossy source coding with side information (i.e., Wyner-Ziv coding) for sources with memory. The overall architecture involves designing suitable filters at the encoder and the decoder and performing fixed-rate scalar quantization followed…

Cited by 0SourceScholar
2022

Modular Action Concept Grounding in Semantic Video Prediction

CVPR 2022poster

Recent works in video prediction have mainly focused on passive forecasting and low-level action-conditional prediction, which sidesteps the learning of interaction between agents and objects. We introduce the task of semantic action-conditional video prediction, which uses semantic action labels to…

Cited by 15PDFScholar
2022

User Scheduling Using Graph Neural Networks for Reconfigurable Intelligent Surface Assisted Multiuser Downlink Communications

ICASSP 2022accepted

Reconfigurable intelligent surface (RIS) is capable of intelligently manipulating the phases of the incident electromagnetic wave to improve the wireless propagation environment between the base station (BS) and the users. This paper addresses the joint user scheduling, RIS configuration, and BS bea…

Cited by 0SourceScholar
2021

An Efficient Active Set Algorithm for Covariance Based Joint Data and Activity Detection for Massive Random Access with Massive MIMO

ICASSP 2021accepted

This paper proposes a computationally efficient algorithm to solve the joint data and activity detection problem for massive random access with massive multiple-input multiple-output (MIMO). The BS acquires the active devices and their data by detecting the transmitted preassigned nonorthogonal sign…

Cited by 0SourceScholar
2021

Deep Active Learning Approach to Adaptive Beamforming for mmWave Initial Alignment

ICASSP 2021accepted

This paper proposes a deep learning approach to the adaptive and sequential beamforming design problem for the initial access phase in a mmWave environment with a single-path channel model. In particular, for a single-user scenario where the problem is equivalent to designing the sequence of sensing…

Cited by 0SourceScholar
2020

Efficient and Information-Preserving Future Frame Prediction and Beyond

ICLR 2020poster

Applying resolution-preserving blocks is a common practice to maximize information preservation in video prediction, yet their high memory consumption greatly limits their application scenarios. We propose CrevNet, a Conditionally Reversible Network that uses reversible architectures to build a bije…

Cited by 142SourceScholar
2018

Sparse Activity Detection for Massive Connectivity in Cellular Networks: Multi-Cell Cooperation Vs Large-Scale Antenna Arrays

ICASSP 2018accepted

Sparse device activity detection for machine-type communications has attracted increasing attention in recent studies. However, most of the previous works focus on the single-cell case. This paper studies the impact of the inter-cell interference on the device activity detection problem with non-ort…

Cited by 0SourceScholar
2016

Coordinated uplink scheduling and beamforming for wireless cellular networks via sum-of-ratio programming and matching

ICASSP 2016accepted

This paper proposes a joint uplink user scheduling and beam-forming algorithm for a multiple-antenna wireless cellular network. We show that coordinated optimization across the cells can significantly alleviate intercell interference, thereby improving the cell-edge rates in a multicell network. Unl…

Cited by 0SourceScholar
2015

On Learning Optimized Reaction Diffusion Processes for Effective Image Restoration

CVPR 2015poster

For several decades, image restoration remains an active research topic in low-level computer vision and hence new approaches are constantly emerging. However, many recently proposed algorithms achieve state-of-the-art performance only at the expense of very high computation time, which clearly limi…

Cited by 391SourcePDFScholar