← Search

Ning Xu

74 accepted papers

2026

Alignment through Meta-Weighted Online Sampling: Bridging the Gap between Data Generation and Preference Optimization

ICLR 2026poster

Preference optimization is crucial for aligning large language models (LLMs) with human values and intentions. A significant challenge in this process is the distribution mismatch between pre-collected offline preference data and the evolving model policy. Existing methods attempt to reduce this gap…

Cited by 0SourcecodeScholar
2026

Class-Prior Perturbation-Robust Regularization for Imbalanced Unreliable Partial Label Learning

ICML 2026poster

Imbalanced Unreliable Partial Label Learning (I-UPLL) is a challenging weakly supervised learning setting in which severe class imbalance and unreliable candidate labels jointly degrade model performance. By revisiting existing approaches for imbalanced learning, we observe that most of them fundame…

Cited by 0SourceScholar
2026

LoGoSeg: Integrating Local and Global Features for Open-Vocabulary Semantic Segmentation

AAAI 2026technical

Open-vocabulary semantic segmentation (OVSS) extends traditional closed-set segmentation by enabling pixel-wise annotation for both seen and unseen categories using arbitrary textual descriptions. While existing methods leverage vision-language models (VLMs) like CLIP, their reliance on image-level

Cited by 0SourcePDFScholar
2026

RefineEvo: Planning-Guided Heuristic Evolution with Bidirectional Experience

ICML 2026poster

Automatic Heuristic Design (AHD) has emerged as a transformative approach for solving combinatorial optimization problems. While recent Large Language Model (LLM)-based methods have shown promise, they predominantly rely on fixed evolutionary operators and struggle to effectively accumulate and reus…

Cited by 0SourceScholar
2026

Thinking as Society: Multi-Social-Agent Self-Distillation for Multimodal Misinformation Detection

ICLR 2026poster

Multimodal Misinformation Detection (MMD) in realistic, mixed-sourced scenarios must incorporate robust reasoning capabilities to handle the social complexity and diverse types of forgeries. While MLLM-based agents are increasingly used for MMD task due to their powerful reasoning abilities, they su…

Cited by 0SourceScholar
2026

When Labelers Stay Silent: The Power of Ties in Cost-Effective Preference Learning

ICML 2026poster

Standard preference alignment relies on a binary forced-choice paradigm, assuming definitive preferences for all pairs. However, we find that indistinguishable pairs are prevalent even in standard benchmarks, where quality differences of two responses often fall below the labeler's discriminative re…

Cited by 0SourceScholar
2025

Bi-Level Knowledge Transfer for Multi-Task Multi-Agent Reinforcement Learning

NeurIPS 2025poster

Multi-Agent Reinforcement Learning (MARL) has achieved remarkable success in various real-world scenarios, but its high cost of online training makes it impractical to learn each task from scratch. To enable effective policy reuse, we consider the problem of zero-shot generalization from offline da…

Cited by 0SourceScholar
2025

Error Bounds Revisited, and How to Use Bayesian Statistics While Remaining a Frequentist

ICASSP 2025accepted

Signal processing makes extensive use of point estimators and accompanying error bounds. These work well up until the likelihood function has two or more high peaks. When it is important for an estimator to remain reliable, it becomes necessary to consider alternatives, such as set estimators. An ob…

Cited by 0SourceScholar
2025

Reduction-based Pseudo-label Generation for Instance-dependent Partial Label Learning

NeurIPS 2025poster

Instance-dependent Partial Label Learning (ID-PLL) aims to learn a multi-class predictive model given training instances annotated with candidate labels related to features, among which correct labels are hidden fixed but unknown. The previous works involve leveraging the identification capability o…

Cited by 0SourceScholar
2025

SMTPD: A New Benchmark for Temporal Prediction of Social Media Popularity

CVPR 2025poster

Social media popularity prediction task aims to predict the popularity of posts on social media platforms, which has a positive driving effect on application scenarios such as content optimization, digital marketing and online advertising. Though many studies have made significant progress, few of t…

2025

Uncertain Knowledge Graph Completion via Semi-Supervised Confidence Distribution Learning

NeurIPS 2025spotlight

Uncertain knowledge graphs (UKGs) associate each triple with a confidence score to provide more precise knowledge representations. Recently, since real-world UKGs suffer from the incompleteness, uncertain knowledge graph (UKG) completion attracts more attention, aiming to complete missing triples an…

Cited by 0SourceScholar
2025

VADIS: Investigating Inter-View Representation Biases for Multi-View Partial Multi-Label Learning

UAI 2025

Multi-view partial multi-label learning (MVPML) deals with training data where each example is represented by multiple feature vectors and associated with a set of candidate labels, only a subset of which are correct. The diverse representation biases present in different views complicate the annota

Cited by 0SourcePDFScholar
2024

Aligned Objective for Soft-Pseudo-Label Generation in Supervised Learning

ICML 2024poster

Soft pseudo-labels, generated by the softmax predictions of the trained networks, offer a probabilistic rather than binary form, and have been shown to improve the performance of deep neural networks in supervised learning. Most previous methods adopt classification loss to train a classifier as the…

Cited by 1SourcePDFScholar
2024

Correlation-Induced Label Prior for Semi-Supervised Multi-Label Learning

ICML 2024poster

Semi-supervised multi-label learning (SSMLL) aims to address the challenge of limited labeled data availability in multi-label learning (MLL) by leveraging unlabeled data to improve the model's performance. Due to the difficulty of estimating the reliable label correlation on minimal multi-labeled d…

Cited by 0SourcePDFScholar
2024

Learning with Partial-Label and Unlabeled Data: A Uniform Treatment for Supervision Redundancy and Insufficiency

ICML 2024spotlight

One major challenge in weakly supervised learning is learning from inexact supervision, ranging from partial labels (PLs) with *redundant* information to the extreme of unlabeled data with *insufficient* information. While recent work has made significant strides in specific inexact supervision cont…

Cited by 2SourcePDFScholar
2024

ULAREF: A Unified Label Refinement Framework for Learning with Inaccurate Supervision

ICML 2024spotlight

Learning with inaccurate supervision is often encountered in weakly supervised learning, and researchers have invested a considerable amount of time and effort in designing specialized algorithms for different forms of annotations in inaccurate supervision. In fact, different forms of these annotati…

Cited by 0SourcePDFScholar
2024

What Makes Partial-Label Learning Algorithms Effective?

NeurIPS 2024poster

A partial label (PL) specifies a set of candidate labels for an instance and partial-label learning (PLL) trains multi-class classifiers with PLs. Recently, many methods that incorporate techniques from other domains have shown strong potential. The expectation that stronger techniques would enhance…

Cited by 2SourcePDFScholar
2023

Decompositional Generation Process for Instance-Dependent Partial Label Learning

ICLR 2023top-25%

Partial label learning (PLL) is a typical weakly supervised learning problem, where each training example is associated with a set of candidate labels among which only one is true. Most existing PLL approaches assume that the incorrect labels in each training example are randomly picked as the candi…

2023

FREDIS: A Fusion Framework of Refinement and Disambiguation for Unreliable Partial Label Learning

ICML 2023poster

To reduce the difficulty of annotation, partial label learning (PLL) has been widely studied, where each example is ambiguously annotated with a set of candidate labels instead of the exact correct label. PLL assumes that the candidate label set contains the correct label, which induces disambiguati…

Cited by 7SourcePDFScholar
2023

FanoutNet: A Neuralized PCB Fanout Automation Method Using Deep Reinforcement Learning

AAAI 2023technical

In modern electronic manufacturing processes, multi-layer Printed Circuit Board (PCB) routing requires connecting more than hundreds of nets with perplexing topology under complex routing constraints and highly limited resources, so that takes intense effort and time of human engineers. PCB fanout a…

Cited by 4SourcePDFScholar
2023

Progressive Purification for Instance-Dependent Partial Label Learning

ICML 2023poster

Partial label learning (PLL) aims to train multiclass classifiers from the examples each annotated with a set of candidate labels where a fixed but unknown candidate label is correct. In the last few years, the instance-independent generation process of candidate labels has been extensively studied,…

Cited by 25SourcePDFScholar
2023

Towards Effective Visual Representations for Partial-Label Learning

CVPR 2023poster

Under partial-label learning (PLL) where, for each training instance, only a set of ambiguous candidate labels containing the unknown true label is accessible, contrastive learning has recently boosted the performance of PLL on vision tasks, attributed to representations learned by contrasting the s…

2022

Ambiguity-Induced Contrastive Learning for Instance-Dependent Partial Label Learning

IJCAI 2022poster

Partial label learning (PLL) learns from a typical weak supervision, where each training instance is labeled with a set of ambiguous candidate labels (CLs) instead of its exact ground-truth label. Most existing PLL works directly eliminate, rather than exploiting the label ambiguity, since they expl…

2022

Image Inpainting with Cascaded Modulation GAN and Object-Aware Training

ECCV 2022poster

"Recent image inpainting methods have made great progress but often struggle to generate plausible image structures when dealing with large holes in complex images. This is partially due to the lack of effective network structures that can capture both the long-range dependency and high-level semant…

2022

Learngene: From Open-World to Your Learning Task

AAAI 2022technical

Although deep learning has made significant progress on fixed large-scale datasets, it typically encounters challenges regarding improperly detecting unknown/unseen classes in the open-world scenario, over-parametrized, and overfitting small samples. Since biological systems can overcome the above d…

2022

One Positive Label is Sufficient: Single-Positive Multi-Label Learning with Label Enhancement

NeurIPS 2022accept

Multi-label learning (MLL) learns from the examples each associated with multiple labels simultaneously, where the high cost of annotating all relevant labels for each training example is challenging for real-world applications. To cope with the challenge, we investigate single-positive multi-label…

2022

SpaceEdit: Learning a Unified Editing Space for Open-Domain Image Color Editing

CVPR 2022poster

Recently, large pretrained models (e.g., BERT, StyleGAN, CLIP) show great knowledge transfer and generalization capability on various downstream tasks within their domains. Inspired by these efforts, in this paper we propose a unified model for open-domain image editing focusing on color and tone ad…

Cited by 19PDFScholar
2022

Transfer Learning and Prediction Consistency for Detecting Offensive Spans of Text

ACL 2022findings

Toxic span detection is the task of recognizing offensive spans in a text snippet. Although there has been prior work on classifying text snippets as offensive or not, the task of recognizing spans responsible for the toxicity of a text is not explored yet. In this work, we introduce a novel multi-t…

Cited by 5SourcePDFScholar
2022

Wavelet Knowledge Distillation: Towards Efficient Image-to-Image Translation

CVPR 2022poster

Remarkable achievements have been attained with Generative Adversarial Networks (GANs) in image-to-image translation. However, due to a tremendous amount of parameters, state-of-the-art GANs usually suffer from low efficiency and bulky memory usage. To tackle this challenge, firstly, this paper inve…

Cited by 104PDFScholar
2021

A Simple Baseline for Weakly-Supervised Scene Graph Generation

ICCV 2021poster

We investigate the weakly-supervised scene graph generation, which is a challenging task since no correspondence of label and object is provided. The previous work regards such correspondence as a latent variable which is iteratively updated via nested optimization of the scene graph generation obje…

Cited by 36PDFcodeScholar
2021

End-to-End Video Instance Segmentation via Spatial-Temporal Graph Neural Networks

ICCV 2021poster

Video instance segmentation is a challenging task that extends image instance segmentation to the video domain. Existing methods either rely only on single-frame information for the detection and segmentation subproblems or handle tracking as a separate post-processing step, which limit their capabi…

Cited by 38PDFcodeScholar
2021

Language-Guided Global Image Editing via Cross-Modal Cyclic Mechanism

ICCV 2021poster

Editing an image automatically via a linguistic request can significantly save laborious manual work and is friendly to photography novice. In this paper, we focus on the task of language-guided global image editing. Existing works suffer from imbalanced data distribution of real-world datasets and…

Cited by 29PDFScholar
2021

Learning by Planning: Language-Guided Global Image Editing

CVPR 2021poster

Recently, language-guided global image editing draws increasing attention with growing application potentials. However, previous GAN-based methods are not only confined to domain-specific, low-resolution data but also lacking in interpretability. To overcome the collective difficulties, we develop a…

Cited by 40PDFcodeScholar
2021

Mask Guided Matting via Progressive Refinement Network

CVPR 2021poster

We propose Mask Guided (MG) Matting, a robust matting framework that takes a general coarse mask as guidance. MG Matting leverages a network (PRN) design which encourages the matting model to provide self-guidance to progressively refine the uncertain regions through the decoding process. A series o…

Cited by 153PDFcodeScholar
2021

Self-Attention Graph Residual Convolutional Networks for Event Detection with dependency relations

EMNLP 2021finding

Event detection (ED) task aims to classify events by identifying key event trigger words embedded in a piece of text. Previous research have proved the validity of fusing syntactic dependency relations into Graph Convolutional Networks(GCN). While existing GCN-based methods explore latent node-to-no…

Cited by 29SourcePDFScholar
2020

A Study on the Transferability of Adversarial Attacks in Sound Event Classification

ICASSP 2020accepted

An adversarial attack is an algorithm that perturbs the input of a machine learning model in an intelligent way in order to change the output of the model. An important property of adversarial attacks is transferability. According to this property, it is possible to generate adversarial perturbation…

Cited by 15SourceScholar
2020

AOWS: Adaptive and Optimal Network Width Search With Latency Constraints

CVPR 2020oral

Neural architecture search (NAS) approaches aim at automatically finding novel CNN architectures that fit computational constraints while maintaining a good performance on the target platform. We introduce a novel efficient one-shot NAS approach to optimally search for channel numbers, given latency…

Cited by 37PDFcodeScholar
2020

Delving into the Cyclic Mechanism in Semi-supervised Video Object Segmentation

NeurIPS 2020poster

In this paper, we take attempt to incorporate the cyclic mechanism with the vision task of semi-supervised video object segmentation. By resorting to the accurate reference mask of the first frame, we try to mitigate the error propagation problem in most of current video object segmentation pipeline…

2020

GeoFusion: Geometric Consistency Informed Scene Estimation in Dense Clutter

RA-L 2020

We propose GeoFusion, a SLAM-based scene estimation method for building an object-level semantic map in dense clutter. In dense clutter, objects are often in close contact and severe occlusions, which brings more false detections and noisy pose estimates from existing perception methods. To solve th

Cited by 10SourceScholar
2020

Incorporating Reinforced Adversarial Learning in Autoregressive Image Generation

ECCV 2020poster

Autoregressive models recently achieved comparable results versus state-of-the-art Generative Adversarial Networks (GANs) with the help of Vector Quantized Variational AutoEncoders (VQ-VAE). However, autoregressive models have several limitations such as exposure bias and their training objective do…

Cited by 17SourcePDFScholar
2020

Minimizing FLOPs to Learn Efficient Sparse Representations

ICLR 2020poster

Deep representation learning has become one of the most widely adopted approaches for visual search, recommendation, and identification. Retrieval of such representations from a large database is however computationally challenging. Approximate methods based on learning compact representations, hav…

Cited by 77SourcecodeScholar
2020

Multiple Sound Sources Localization from Coarse to Fine

ECCV 2020poster

How to visually localize multiple sound sources in unconstrained videos is a formidable problem, especially when lack of the pairwise sound-object annotations. To solve this problem, we develop a two-stage audiovisual learning framework that disentangles audio and visual representations of different…

2019

An Internal Learning Approach to Video Inpainting

ICCV 2019poster

We propose a novel video inpainting algorithm that simultaneously hallucinates missing appearance and motion (optical flow) information, building upon the recent 'Deep Image Prior' (DIP) that exploits convolutional network architectures to enforce plausible texture in static images. In extending DIP…

Cited by 101PDFcodeScholar
2019

Controllable Artistic Text Style Transfer via Shape-Matching GAN

ICCV 2019oral

Artistic text style transfer is the task of migrating the style from a source image to the target text to create artistic typography. Recent style transfer methods have considered texture control to enhance usability. However, controlling the stylistic degree in terms of shape deformation remains an…

Cited by 129PDFcodeScholar
2019

EIGEN: Ecologically-Inspired GENetic Approach for Neural Network Structure Searching From Scratch

CVPR 2019poster

Designing the structure of neural networks is considered one of the most challenging tasks in deep learning, especially when there is few prior knowledge about the task domain. In this paper, we propose an Ecologically-Inspired GENetic (EIGEN) approach that uses the concept of succession, extinction…

Cited by 33PDFScholar
2019

End-To-End Time-Lapse Video Synthesis From a Single Outdoor Image

CVPR 2019poster

Time-lapse videos usually contain visually appealing content but are often difficult and costly to create. In this paper, we present an end-to-end solution to synthesize a time-lapse video from a single outdoor image using deep neural networks. Our key idea is to train a conditional generative adver…

Cited by 40PDFScholar
2019

Fast User-Guided Video Object Segmentation by Interaction-And-Propagation Networks

CVPR 2019poster

We present a deep learning method for the interactive video object segmentation. Our method is built upon two core operations, interaction and propagation, and each operation is conducted by Convolutional Neural Networks. The two networks are connected both internally and externally so that the netw…

Cited by 79PDFScholar
2019

Large-Scale Tag-Based Font Retrieval With Generative Feature Learning

ICCV 2019poster

Font selection is one of the most important steps in a design workflow. Traditional methods rely on ordered lists which require significant domain knowledge and are often difficult to use even for trained professionals. In this paper, we address the problem of large-scale tag-based font retrieval wh…

Cited by 36PDFScholar
2019

Temporal Structure Mining for Weakly Supervised Action Detection

ICCV 2019poster

Different from the fully-supervised action detection problem that is dependent on expensive frame-level annotations, weakly supervised action detection (WSAD) only needs video-level annotations, making it more practical for real-world applications. Existing WSAD methods detect action instances by sc…

Cited by 95PDFScholar
2018

WSNet: Compact and Efficient Networks Through Weight Sampling

ICML 2018oral

We present a new approach and a novel architecture, termed WSNet, for learning compact and efficient deep neural networks. Existing approaches conventionally learn full model parameters independently and then compress them via ad hoc processing such as model pruning or filter factorization. Alternat…

2018

WSNet: Learning Compact and Efficient Networks with Weight Sampling

ICLR 2018workshop

We present a new approach and a novel architecture, termed WSNet, for learning compact and efficient deep neural networks. Existing approaches conventionally learn full model parameters independently and then compress them via \emph{ad hoc} processing such as model pruning or filter factorization. A…

Cited by 0SourceScholar
2018

YouTube-VOS: Sequence-to-Sequence Video Object Segmentation

ECCV 2018poster

Learning long-term spatial-temporal features are critical for many video analysis tasks. However, existing video segmentation methods predominantly rely on static image segmentation techniques, and methods capturing temporal dependency for segmentation have to depend on pretrained optical flow model…

Cited by 594SourcePDFScholar