← Search

Shuang Wu

59 accepted papers

2026

All Patches Matter, More Patches Better: Enhance AI-Generated Image Detection via Panoptic Patch Learning

ICLR 2026poster

The rapid proliferation of AI-generated images (AIGIs) highlights the pressing demand for generalizable detection methods. In this paper, we establish two key principles for AIGI detection task through systematic analysis: **(1) All Patches Matter**, since the uniform generation process ensures that…

Cited by 0SourceScholar
2026

CaFe-TeleVision: A Coarse-To-Fine Teleoperation System with Immersive Situated Visualization for Enhanced Ergonomics

ICRA 2026poster

Teleoperation presents a promising paradigm for remote control and robot proprioceptive data collection. Despite recent progress, current teleoperation systems still suffer from limitations in efficiency and ergonomics, particularly in challenging scenarios. In this paper, we propose CaFe-TeleVision…

2026

CaFe-TeleVision: A Coarse-to-Fine Teleoperation System With Immersive Situated Visualization for Enhanced Ergonomics

RA-L 2026

Teleoperation presents a promising paradigm for remote control and robot proprioceptive data collection. Despite recent progress, current teleoperation systems still suffer from limitations in efficiency and ergonomics, particularly in challenging scenarios. In this paper, we propose CaFe-TeleVision

Cited by 0SourcecodeScholar
2026

DiffTrans: Differentiable Geometry-Materials Decomposition for Reconstructing Transparent Objects

ICLR 2026poster

Reconstructing transparent objects from a set of multi-view images is a challenging task due to the complicated nature and indeterminate behavior of light propagation. Typical methods are primarily tailored to specific scenarios, such as objects following a uniform topology, exhibiting ideal transpa…

Cited by 0SourceScholar
2026

Laplacian Kernelized Bandit

ICLR 2026poster

We study multi-user contextual bandits where users are related by a graph and their reward functions exhibit both non-linear behavior and graph homophily. We introduce a principled joint penalty for the collection of user reward functions $\\{f_u\\}$, combining a graph smoothness term based on RKHS…

Cited by 0SourceScholar
2026

OpenPyRo-A1: An Open Python-Based Low-Cost Bimanual Robot for Embodied AI

RA-L 2026

Many real-world tasks, such as assembly, cooking, and object handovers, require bi-manual coordination. Learning such skills via imitation remains challenging due to dataset scarcity, mainly caused by the high cost of bi-manual robotic platforms and barriers to entry in robotics software. To address

Cited by 1SourceScholar
2026

OpenPyRo-A1: An Open Python-Based Low-Cost Bimanual Robot for Embodied AI

ICRA 2026poster

Many real-world tasks, such as assembly, cooking, and object handovers, require bi-manual coordination. Learning such skills via imitation remains challenging due to dataset scarcity, mainly caused by the high cost of bi-manual robotic platforms and barriers to entry in robotics software. To address…

Cited by 0SourceScholar
2026

TEXTRIX: Latent Attribute Grid for Native Texture Generation and Beyond

CVPR 2026

Prevailing 3D texture generation methods, which often rely on multi-view fusion, are frequently hindered by inter-view inconsistencies and incomplete coverage of complex surfaces, limiting the fidelity and completeness of the generated content. To overcome these challenges, we introduce TEXTRIX, a n

Cited by 0SourcecodeScholar
2025

A Cross-Modal Densely Guided Knowledge Distillation Based on Modality Rebalancing Strategy for Enhanced Unimodal Emotion Recognition

IJCAI 2025

Multimodal emotion recognition has garnered significant attention for its ability to integrate data from multiple modalities to enhance performance. However, physiological signals like electroencephalogram are more challenging to acquire than visual data due to higher collection costs and complexity

Cited by 0SourcePDFScholar
2025

Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention

NeurIPS 2025poster

Generating high-resolution 3D shapes using volumetric representations such as Signed Distance Functions (SDFs) presents substantial computational and memory challenges. We introduce Direct3D-S2, a scalable 3D generation framework based on sparse volumes that achieves superior output quality with dra…

Cited by 0SourceScholar
2025

Dual Data Alignment Makes AI-Generated Image Detector Easier Generalizable

NeurIPS 2025spotlight

The rapid increase in AI-generated images (AIGIs) underscores the need for detection methods. Existing detectors are often trained on biased datasets, leading to overfitting on spurious correlations between non-causal image attributes and real/synthetic labels. While these biased features enhance p…

Cited by 0SourcecodeScholar
2025

High-quality Text-to-3D Character Generation with SparseCubes and Sparse Transformers.

ICLR 2025poster

Current state-of-the-art text-to-3D generation methods struggle to produce 3D models with fine details and delicate structures due to limitations in differentiable mesh representation techniques. This limitation is particularly pronounced in anime character generation, where intricate features such…

Cited by 0SourcePDFScholar
2025

Instruct Where the Model Fails: Generative Data Augmentation via Guided Self-contrastive Fine-tuning

AAAI 2025technical

Data augmentation is expected to bring about unseen features of training set, enhancing the model’s ability to generalize in situations where data is limited. Generative image models trained on large web-crawled datasets such as LAION are known to produce images with stereotypes and imperceptible bi…

Cited by 0SourcePDFScholar
2025

MARS: Unleashing the Power of Variance Reduction for Training Large Models

ICML 2025poster

Training deep neural networks--and more recently, large models--demands efficient and scalable optimizers. Adaptive gradient algorithms like Adam, AdamW, and their variants have been central to this task. Despite the development of numerous variance reduction algorithms in the past decade aimed at a…

2024

Augmenting Lane Perception and Topology Understanding with Standard Definition Navigation Maps

ICRA 2024poster

Autonomous driving has traditionally relied heavily on costly and labor-intensive High Definition (HD) maps, hindering scalability. In contrast, Standard Definition (SD) maps are more affordable and have worldwide coverage, offering a scalable alternative. In this work, we systematically explore the…

Cited by 34SourcecodeScholar
2024

Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion Transformer

NeurIPS 2024poster

Generating high-quality 3D assets from text and images has long been challenging, primarily due to the absence of scalable 3D representations capable of capturing intricate geometry distributions. In this work, we introduce Direct3D, a native 3D generative model scalable to in-the-wild input images,…

Cited by 35SourcePDFScholar
2024

Exposing the Deception: Uncovering More Forgery Clues for Deepfake Detection

AAAI 2024technical

Deepfake technology has given rise to a spectrum of novel and compelling applications. Unfortunately, the widespread proliferation of high-fidelity fake videos has led to pervasive confusion and deception, shattering our faith that seeing is believing. One aspect that has been overlooked so far is t…

2024

Field-VIO: Stereo Visual-Inertial Odometry Based on Quantitative Windows in Agricultural Open Fields

ICRA 2024poster

In agricultural open fields, accurate autonomous localization of robots requires long-term data correlation to reduce cumulative error. Our article presents a Stereo Visual-Inertial Odometry (VIO) system based on ORB-SLAM3 to address the malfunction of the Loop Closure Detection (LCD) methods in thi…

Cited by 1SourceScholar
2024

GMPC: Geometric Model Predictive Control for Wheeled Mobile Robot Trajectory Tracking

RA-L 2024

The configuration of most robotic systems lies in continuous transformation groups. However, in mobile robot trajectory tracking, many recent works still naively utilize optimization methods for elements in vector space without considering the manifold constraint of the robot configuration. In this

Cited by 27SourcecodeScholar
2024

Re-thinking Data Availability Attacks Against Deep Neural Networks

CVPR 2024poster

The unauthorized use of personal data for commercial purposes and the covert acquisition of private data for training machine learning models continue to raise concerns. To address these issues researchers have proposed availability attacks that aim to render data unexploitable. However many availab…

Cited by 5SourcePDFScholar
2024

Safe Table Tennis Swing Stroke with Low-Cost Hardware

ICRA 2024poster

Playing table tennis with a human player is a challenging robotic task due to its dynamic nature. Despite a number of researches being devoted to developing robotic table tennis systems, most of the works have demanding hardware requirements and ignore safety measures when generating the swing stoke…

Cited by 0SourceScholar
2024

UniVoxel: Fast Inverse Rendering by Unified Voxelization of Scene Representation

ECCV 2024poster

"Typical inverse rendering methods focus on learning implicit neural scene representations by modeling the geometry, materials and illumination separately, which entails significant computations for optimization. In this work we design a Unified Voxelization framework for explicit learning of scene…

2023

Action Recognition with Multi-stream Motion Modeling and Mutual Information Maximization

IJCAI 2023poster

Action recognition has long been a fundamental and intriguing problem in artificial intelligence. The task is challenging due to the high dimensionality nature of an action, as well as the subtle motion details to be considered. Current state-of-the-art approaches typically learn from articulated mo…

2023

Attack Can Benefit: An Adversarial Approach to Recognizing Facial Expressions under Noisy Annotations

AAAI 2023technical

The real-world Facial Expression Recognition (FER) datasets usually exhibit complex scenarios with coupled noise annotations and imbalanced classes distribution, which undoubtedly impede the development of FER methods. To address the aforementioned issues, in this paper, we propose a novel and flexi…

Cited by 14SourcePDFScholar
2023

Content-based Unrestricted Adversarial Attack

NeurIPS 2023poster

Unrestricted adversarial attacks typically manipulate the semantic content of an image (e.g., color or texture) to create adversarial examples that are both effective and photorealistic, demonstrating their ability to deceive human perception and deep neural networks with stealth and success. Howeve…

Cited by 86SourcePDFScholar
2023

Curriculum-based Co-design of Morphology and Control of Voxel-based Soft Robots

ICLR 2023poster

Co-design of morphology and control of a Voxel-based Soft Robot (VSR) is challenging due to the notorious bi-level optimization. In this paper, we present a Curriculum-based Co-design (CuCo) method for learning to design and control VSRs through an easy-to-difficult process. Specifically, we expand…

Cited by 10SourcePDFScholar
2023

Delving into the Adversarial Robustness of Federated Learning

AAAI 2023technical

In Federated Learning (FL), models are as fragile as centrally trained models against adversarial examples. However, the adversarial robustness of federated learning remains largely unexplored. This paper casts light on the challenge of adversarial robustness of federated learning. To facilitate a b…

Cited by 38SourcePDFScholar
2023

Learning Bifunctional Push-Grasping Synergistic Strategy for Goal-Agnostic and Goal-Oriented Tasks

IROS 2023poster

Both goal-agnostic and goal-oriented tasks have practical value for robotic grasping: goal-agnostic tasks target all objects in the workspace, while goal-oriented tasks aim at grasping pre-assigned goal objects. However, most current grasping methods are only better at coping with one task. In this…

Cited by 11SourcecodeScholar
2023

Online Map Vectorization for Autonomous Driving: A Rasterization Perspective

NeurIPS 2023poster

High-definition (HD) vectorized map is essential for autonomous driving, providing detailed and precise environmental information for advanced perception and planning. However, current map vectorization methods often exhibit deviations, and the existing evaluation metric for map vectorization lacks…

2023

PreCo: Enhancing Generalization in Co-Design of Modular Soft Robots via Brain-Body Pre-Training

CoRL 2023oral

Brain-body co-design, which involves the collaborative design of control strategies and morphologies, has emerged as a promising approach to enhance a robot's adaptability to its environment. However, the conventional co-design process often starts from scratch, lacking the utilization of prior know…

Cited by 9SourceScholar
2023

Quality-Similar Diversity via Population Based Reinforcement Learning

ICLR 2023poster

Diversity is a growing research topic in Reinforcement Learning (RL). Previous research on diversity has mainly focused on promoting diversity to encourage exploration and thereby improve quality (the cumulative reward), maximizing diversity subject to quality constraints, or jointly maximizing qual…

Cited by 22SourcePDFScholar
2023

Rethinking the Learning Paradigm for Dynamic Facial Expression Recognition

CVPR 2023poster

Dynamic Facial Expression Recognition (DFER) is a rapidly developing field that focuses on recognizing facial expressions in video format. Previous research has considered non-target frames as noisy frames, but we propose that it should be treated as a weakly supervised problem. We also identify the…

2022

Actor-Critic Policy Optimization in a Large-Scale Imperfect-Information Game

ICLR 2022poster

The deep policy gradient method has demonstrated promising results in many large-scale games, where the agent learns purely from its own experience. Yet, policy gradient methods with self-play suffer convergence problems to a Nash Equilibrium (NE) in multi-agent situations. Counterfactual regret min…

Cited by 33SourcePDFScholar
2022

Copy Motion From One to Another: Fake Motion Video Generation

IJCAI 2022poster

One compelling application of artificial intelligence is to generate a video of a target person performing arbitrary desired motion (from a source person). While the state-of-the-art methods are able to synthesize a video demonstrating similar broad stroke motion details, they are generally lacking…

2022

DENSE: Data-Free One-Shot Federated Learning

NeurIPS 2022accept

One-shot Federated Learning (FL) has recently emerged as a promising approach, which allows the central server to learn a model in a single communication round. Despite the low communication cost, existing one-shot FL methods are mostly impractical or face inherent limitations, \eg a public dataset…

2022

Detecting Camouflaged Object in Frequency Domain

CVPR 2022poster

Camouflaged object detection (COD) aims to identify objects that are perfectly embedded in their environment, which has various downstream applications in fields such as medicine, art, and agriculture. However, it is an extremely challenging task to spot camouflaged objects with the perception abili…

Cited by 217PDFcodeScholar
2022

Federated Learning Challenges and Opportunities: An Outlook

ICASSP 2022accepted

Federated learning (FL) has been developed as a promising framework to leverage the resources of edge devices, enhance customers’ privacy, comply with regulations, and reduce development costs. Although many methods and applications have been developed for FL, several critical challenges for practic…

Cited by 0SourceScholar
2022

Few-Shot Object Detection by Knowledge Distillation Using Bag-of-Visual-Words Representations

ECCV 2022poster

"While fine-tuning based methods for few-shot object detection have achieved remarkable progress, a crucial challenge that has not been addressed well is the potential class-specific overfitting on base classes and sample-specific overfitting on novel classes. In this work we design a novel knowledg…

Cited by 18SourcePDFScholar
2022

Greedy when Sure and Conservative when Uncertain about the Opponents

ICML 2022spotlight

We develop a new approach, named Greedy when Sure and Conservative when Uncertain (GSCU), to competing online against unknown and nonstationary opponents. GSCU improves in four aspects: 1) introduces a novel way of learning opponent policy embeddings offline; 2) trains offline a single best response…

2022

Multi-faceted Distillation of Base-Novel Commonality for Few-Shot Object Detection

ECCV 2022poster

"Most of existing methods for few-shot object detection follow the fine-tuning paradigm, which potentially assumes that the class-agnostic generalizable knowledge can be learned and transferred implicitly from base classes with abundant samples to novel classes with limited samples via such a two-st…

2022

Self-Aware Personalized Federated Learning

NeurIPS 2022accept

In the context of personalized federated learning (FL), the critical challenge is to balance local model improvement and global model tuning when the personal and global objectives may not be exactly aligned. Inspired by Bayesian hierarchical models, we develop a self-aware personalized FL method wh…

Cited by 28SourcePDFScholar
2022

Temporal Feature Alignment and Mutual Information Maximization for Video-Based Human Pose Estimation

CVPR 2022oral

Multi-frame human pose estimation has long been a compelling and fundamental problem in computer vision. This task is challenging due to fast motion and pose occlusion that frequently occur in videos. State-of-the-art methods strive to incorporate additional visual evidences from neighboring frames…

Cited by 82PDFcodeScholar
2022

Towards Practical Certifiable Patch Defense With Vision Transformer

CVPR 2022poster

Patch attacks, one of the most threatening forms of physical attack in adversarial examples, can lead networks to induce misclassification by modifying pixels arbitrarily in a continuous region. Certifiable patch defense can guarantee robustness that the classifier is not affected by patch attacks.…

Cited by 81PDFScholar
2022

Understanding Policy Gradient Algorithms: A Sensitivity-Based Approach

ICML 2022spotlight

The REINFORCE algorithm \cite{williams1992simple} is popular in policy gradient (PG) for solving reinforcement learning (RL) problems. Meanwhile, the theoretical form of PG is from \cite{sutton1999policy}. Although both formulae prescribe PG, their precise connections are not yet illustrated. Recent…

Cited by 10SourcePDFScholar
2022

Understanding and Mitigating Data Contamination in Deep Anomaly Detection: A Kernel-based Approach

IJCAI 2022poster

Deep anomaly detection has become popular for its capability of handling complex data. However, training a deep detector is fragile to data contamination due to overfitting. In this work, we study the performance of the anomaly detectors under data contamination and construct a data-efficient counte…

2021

Aggregated Multi-GANs for Controlled 3D Human Motion Prediction

AAAI 2021technical

Human motion prediction from historical pose sequence is at the core of many applications in machine intelligence. However, in current state-of-the-art methods, the predicted future motion is confined within the same activity. One can neither generate predictions that differ from the current activit…

2021

Deep Dual Consecutive Network for Human Pose Estimation

CVPR 2021poster

Multi-frame human pose estimation in complicated situations is challenging. Although state-of-the-art human joints detectors have demonstrated remarkable results for static images, their performances come short when we apply these models to video sequences. Prevalent shortcomings include the failure…

Cited by 166PDFcodeScholar
2021

MECT: Multi-Metadata Embedding based Cross-Transformer for Chinese Named Entity Recognition

ACL 2021long

Recently, word enhancement has become very popular for Chinese Named Entity Recognition (NER), reducing segmentation errors and increasing the semantic and boundary information of Chinese words. However, these methods tend to ignore the information of the Chinese character structure after integratin…

2021

Motion Prediction Using Trajectory Cues

ICCV 2021poster

Predicting human motion from a historical pose sequence is at the core of many applications in computer vision. Current state-of-the-art methods concentrate on learning motion contexts in the pose space, however, the high dimensionality and complex nature of human pose invoke inherent difficulties i…

Cited by 64PDFcodeScholar
2021

State-Aware Value Function Approximation with Attention Mechanism for Restless Multi-armed Bandits

IJCAI 2021poster

The restless multi-armed bandit (RMAB) problem is a generalization of the multi-armed bandit with non-stationary rewards. Its optimal solution is intractable due to exponentially large state and action spaces with respect to the number of arms. Existing approximation approaches, e.g., Whittle's inde…

Cited by 3SourcePDFScholar
2019

Convolution with even-sized kernels and symmetric padding

NeurIPS 2019poster

Compact convolutional neural networks gain efficiency mainly through depthwise convolutions, expanded channels and complex topologies, which contrarily aggravate the training process. Besides, 3x3 kernels dominate the spatial representation in these models, whereas even-sized kernels (2x2, 4x4) are…

2019

Towards Natural and Accurate Future Motion Prediction of Humans and Animals

CVPR 2019poster

Anticipating the future motions of 3D articulate objects is challenging due to its non-linear and highly stochastic nature. Current approaches typically represent the skeleton of an articulate object as a set of 3D joints, which unfortunately ignores the relationship between joints, and fails to enc…

Cited by 157PDFScholar
2016

Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin

ICML 2016poster

We show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech–two vastly different languages. Because it replaces entire pipelines of hand-engineered components with neural networks, end-to-end learning allows us to handle a diverse variety of s…