← Search

Qingyao Wu

24 accepted papers

2026

Dual-Adversarial Dynamic Variational Asset Pricing with Adaptive Spatio-Temporal Feature Clustering for Portfolio Recommendation

IJCAI 2026

Asset pricing and portfolio recommendation are two closely related fundamental tasks in quantitative investment, for which machine learning methods have attracted significant attention in both academia and industry. In particular, nonlinear asset pricing models based on deep learning architectures t

Cited by 0Scholar
2026

ManipEvalAgent: Promptable and Efficient Evaluation Framework for Robotic Manipulation Policies

ICLR 2026poster

In recent years, robotic manipulation policies have made substantial progress. However, evaluating these policies typically requires large-scale sampling in simulation benchmarks, leading to high time costs. Moreover, existing evaluation pipelines are usually fixed, do not account for user needs, an…

Cited by 0SourceScholar
2026

UAV-CB: A Complex-Background RGB-T Dataset and Local Frequency Bridge Network for UAV Detection

CVPR 2026

Detecting Unmanned Aerial Vehicles (UAVs) in low-altitude environments is essential for perception and defense systems but remains highly challenging due to complex backgrounds, camouflage, and multimodal interference. In real-world scenarios, UAVs are frequently visually blended with surrounding st

Cited by 0SourcecodeScholar
2025

Efficient Infrared Image Super-Resolution Reconstruction via Guided Filter Coefficients Estimation with Parallax Attention Mechanism

ICASSP 2025accepted

Due to the spectral range mismatch between the images, building an efficient infrared (IR) image super-resolution algorithm suitable for embedded devices remains a significant challenge. Given that visible images possess more abundant high-frequency information compared to infrared images, we utiliz…

Cited by 0SourceScholar
2025

IPVTON: Image-based 3D Virtual Try-on with Image Prompt Adapter

AAAI 2025technical

Given a pair of images depicting a person and a garment separately, image-based 3D virtual try-on methods aim to reconstruct a 3D human model that realistically portrays the person wearing the desired garment. In this paper, we present IPVTON, a novel image-based 3D virtual try-on framework. IPVTON…

Cited by 0SourcePDFScholar
2025

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios

NeurIPS 2025spotlight

Hand-Object Interaction (HOI) generation has significant application potential. However, current 3D HOI motion generation approaches heavily rely on predefined 3D object models and lab-captured motion data, limiting generalization capabilities. Meanwhile, HOI video generation methods prioritize pixe…

Cited by 0SourcecodeScholar
2024

Diverse and Stable 2D Diffusion Guided Text to 3D Generation with Noise Recalibration

AAAI 2024technical

In recent years, following the success of text guided image generation, text guided 3D generation has gained increasing attention among researchers. Dreamfusion is a notable approach that enhances generation quality by utilizing 2D text guided diffusion models and introducing SDS loss, a technique f…

2024

Spatial-Semantic Collaborative Cropping for User Generated Content

AAAI 2024technical

A large amount of User Generated Content (UGC) is uploaded to the Internet daily and displayed to people world-widely through the client side (mobile and PC). This requires the cropping algorithms to produce the aesthetic thumbnail within a specific aspect ratio on different devices. However, existi…

2024

Variance-Insensitive and Target-Preserving Mask Refinement for Interactive Image Segmentation

AAAI 2024technical

Point-based interactive image segmentation can ease the burden of mask annotation in applications such as semantic segmentation and image editing. However, fully extracting the target mask with limited user inputs remains challenging. We introduce a novel method, Variance-Insensitive and Target-Pres…

Cited by 3SourcePDFScholar
2023

Digging out Discrimination Information from Generated Samples for Robust Visual Question Answering

ACL 2023findings

Visual Question Answering (VQA) aims to answer a textual question based on a given image. Nevertheless, recent studies have shown that VQA models tend to capture the biases to answer the question, instead of using the reasoning ability, resulting in poor generalisation ability. To alleviate the issu…

Cited by 9SourcePDFScholar
2022

Self-Supervised Object Localization with Joint Graph Partition

AAAI 2022technical

Object localization aims to generate a tight bounding box for the target object, which is a challenging problem that has been deeply studied in recent years. Since collecting bounding-box labels is time-consuming and laborious, many researchers focus on weakly supervised object localization (WSOL).…

Cited by 19SourcePDFScholar
2021

Context Decoupling Augmentation for Weakly Supervised Semantic Segmentation

ICCV 2021poster

Data augmentation is vital for deep learning neural networks. By providing massive training samples, it helps to improve the generalization ability of the model. Weakly supervised semantic segmentation (WSSS) is a challenging problem that has been deeply studied in recent years, conventional data au…

Cited by 152PDFcodeScholar
2021

Debiased Visual Question Answering from Feature and Sample Perspectives

NeurIPS 2021poster

Visual question answering (VQA) is designed to examine the visual-textual reasoning ability of an intelligent agent. However, recent observations show that many VQA models may only capture the biases between questions and answers in a dataset rather than showing real reasoning abilities. For example…

2021

Self-Supervised 3D Skeleton Action Representation Learning With Motion Consistency and Continuity

ICCV 2021poster

Recently, self-supervised learning (SSL) has been proved very effective and it can help boost the performance in learning representations from unlabeled data in the image domain. Yet, very little is explored about its usefulness in 3D skeleton-based action recognition understanding. Directly applyin…

Cited by 80PDFScholar
2020

Fg2seq: Effectively Encoding Knowledge for End-To-End Task-Oriented Dialog

ICASSP 2020accepted

End-to-end Task-oriented spoken dialog systems typically require modeling two types of inputs, namely, the dialog history which is a sequence of utterances and the knowledge base (KB) associated with the dialog history. While modeling these inputs, current state-of-the-art models typically ignore th…

Cited by 0SourceScholar
2020

Human Interaction Learning on 3D Skeleton Point Clouds for Video Violence Recognition

ECCV 2020poster

This paper introduces a new method for recognizing violent behavior by learning contextual relationships between related people from human skeleton points. Unlike previous work, we first formulate 3D skeleton point clouds from human skeleton sequences extracted from videos and then perform interacti…

Cited by 95SourcePDFScholar
2019

Pyramid Graph Networks With Connection Attentions for Region-Based One-Shot Semantic Segmentation

ICCV 2019poster

One-shot image segmentation aims to undertake the segmentation task of a novel class with only one training image available. The difficulty lies in that image segmentation has structured data representations, which yields a many-to-many message passing problem. Previous methods often simplify it to…

Cited by 388PDFScholar
2018

Adversarial Learning with Local Coordinate Coding

ICML 2018oral

Generative adversarial networks (GANs) aim to generate realistic data from some prior distribution (e.g., Gaussian noises). However, such prior distribution is often independent of real data and thus may lose semantic information (e.g., geometric structure or content in images) of data. In practice,…

Cited by 44SourcePDFScholar
2018

Discrimination-aware Channel Pruning for Deep Neural Networks

NeurIPS 2018poster

Channel pruning is one of the predominant approaches for deep model compression. Existing pruning methods either train from scratch with sparsity constraints on channels, or minimize the reconstruction error between the pre-trained feature maps and the compressed ones. Both strategies suffer from s…

2017

A Self-Balanced Min-Cut Algorithm for Image Clustering

ICCV 2017poster

Many spectral clustering algorithms have been proposed and successfully applied to image data analysis such as content based image retrieval, image annotation, and image indexing. Conventional spectral clustering algorithms usually involve a two-stage process: eigendecomposition of similarity matrix…

Cited by 62PDFScholar