← Search

Jing Yang

112 accepted papers

2026

3DAlign-DAER: Dynamic Attention Policy and Efficient Retrieval Strategy for Fine-grained 3D-Text Alignment at Scale

AAAI 2026technical

Despite recent advancements in 3D-text cross-modal alignment, existing state-of-the-art methods still struggle to align fine-grained textual semantics with detailed geometric structures, and their alignment performance degrades significantly when scaling to large-scale 3D databases. To overcome this

Cited by 13SourcePDFScholar
2026

APT: Towards Universal Scene Graph Generation via Plug-in Adaptive Prompt Tuning

ICLR 2026poster

Scene Graph Generation (SGG) is pivotal for structured visual understanding, yet it remains hindered by a fundamental limitation: the reliance on fixed, frozen semantic representations from pre-trained language models. These semantic priors, while beneficial in other domains, are inherently misalign…

Cited by 0SourcecodeScholar
2026

Breaking the Computational Barrier: Provably Efficient Actor–Critic for Low-Rank MDPs

ICML 2026poster

Reinforcement learning (RL) is a fundamental framework for sequential decision-making, in which an agent learns an optimal policy through interactions with an unknown environment. In settings with function approximation, many existing RL algorithms achieve favorable sample complexity, but often rely…

Cited by 0SourceScholar
2026

DK-DDIL: Adaptive Knowledge Retention for Dynamic Domain-Incremental Learning in Medical Imaging

CVPR 2026

Large-scale foundation models pretrained on massive datasets have demonstrated strong generalization capabilities in medical image analysis. However, they are typically trained on static datasets and struggle to cope with the continuously evolving nature of clinical data, where new imaging devices,

Cited by 0SourceScholar
2026

DomainCQA: Crafting Knowledge-Intensive QA from Domain-Specific Charts

AAAI 2026technical

Chart Question Answering (CQA) evaluates Multimodal Large Language Models (MLLMs) on visual understanding and reasoning over chart data. However, existing benchmarks mostly test surface-level parsing, such as reading labels and legends, while overlooking deeper scientific reasoning. We propose Domai

Cited by 0SourcePDFScholar
2026

DreamSAC: Learning Hamiltonian World Models via Symmetry Exploration

CVPR 2026

Learned world models excel at interpolative generalization but fail at extrapolative generalization to novel physical properties. This limitation arises because they learn statistical correlations rather than the environment's underlying generative rules, such as physical invariances and conservatio

Cited by 0SourceScholar
2026

Dynamics-Aware Preference Optimization for Vision-Language Models

CVPR 2026

Preference-based finetuning of vision-language models (VLMs) is notoriously unstable, as trivially wrong negatives inject uninformative gradients that distort optimization and degrade calibration. This work revisits this issue through the lens of learning dynamics and identifies a core pathology, th

Cited by 0SourcecodeScholar
2026

Efficient Multi-objective Prompt Optimization via Pure-exploration Bandits

ICLR 2026poster

Prompt engineering has become central to eliciting the capabilities of large language models (LLMs). At its core lies prompt selection - efficiently identifying the most effective prompts. However, most prior investigations overlook a key challenge: the inherently multi-faceted nature of prompt perf…

Cited by 0SourceScholar
2026

Failure-Driven Workflow Refinement

ICML 2026spotlight

Workflow optimization for tool-using LLM agents is often cast as global search over candidate graphs, scored by a scalar metric. This collapses rich, multi-step failure traces into binary outcomes, obscuring recurring failure structure and making refinement inefficient. We reframe optimization as \e…

Cited by 0SourceScholar
2026

HiVA: Self-organized Hierarchical Variable Agent via Goal-driven Semantic-Topological Evolution

AAAI 2026technical

Autonomous agents play a crucial role in advancing Artificial General Intelligence, enabling problem decomposition and tool orchestration through Large Language Models (LLMs). However, existing paradigms face a critical trade-off. On one hand, reusable fixed workflows require manual reconfiguration

Cited by 0SourcePDFScholar
2026

Hugging Visual Prompt and Segmentation Tokens: Consistency Learning for Fine-Grained Visual Understanding in MLLMs

CVPR 2026

Recently, multimodal large language models (MLLMs) have achieved remarkable success in general multimodal tasks. Increasing attention has been given to leveraging MLLMs for fine-grained visual understanding, such as region-level captioning and pixel-level grounding. However, most existing approaches

Cited by 0SourceScholar
2026

ICTPolarReal: A Polarized Reflection and Material Dataset of Real World Objects

CVPR 2026

Accurately modeling how real-world materials reflect light remains a core challenge in inverse rendering, largely due to the scarcity of real measured reflectance data. Existing approaches rely heavily on synthetic datasets with simplified illumination and limited material realism, preventing models

Cited by 0SourceScholar
2026

Latent Space Robust Optimization of Neural Processes with Aligned Stratified Order-Statistic Loss Reduction

ICML 2026poster

Importance-Weighted Neural Processes (IWNPs) provide a principled framework for probabilistic meta-learning by using multi-particle latent representations to approximate the marginal log-likelihood of task data tightly. However, this work reveals that the standard optimization of IWNPs suffers from …

Cited by 0SourceScholar
2026

Multi-Metric Preference Alignment for Generative Speech Restoration

AAAI 2026technical

Recent generative models have significantly advanced speech restoration tasks, yet their training objectives often misalign with human perceptual preferences, resulting in suboptimal quality. While post-training alignment has proven effective in other generative domains like text and image generatio

Cited by 0SourcePDFScholar
2026

Position: The AI Imperative: Scaling High-Quality Peer Review in Machine Learning

ICML 2026oral

Peer review, the bedrock of scientific advancement in machine learning (ML), is strained by a crisis of scale. Exponential growth in manuscript submissions to premier ML venues such as NeurIPS, ICML, and ICLR is outpacing the finite capacity of qualified reviewers, leading to concerns about review q…

Cited by 0SourceScholar
2026

RLAP-CLIP: Continual Multimodal Learning with Prototype Adaptation and Difficulty-Aware Routing

ICLR 2026poster

Vision-language models, such as CLIP, achieve strong zero-shot performance through contrastive pre-training but face significant challenges in class-incremental image classification scenarios. When learning new tasks sequentially, current methods suffer from degradation in prototype quality due to p…

Cited by 0SourceScholar
2026

RaCoT: Plug-and-Play Contrastive Example Generation Mechanism for Enhanced LLM Reasoning Reliability

AAAI 2026technical

Retrieval-Augmented Generation (RAG) faces a core bottleneck with knowledge-sparse and semantically ambiguous long-tail queries, where retrieval noise distorts reasoning and necessitates costly post-processing. To tackle this, we propose RaCoT (Retrieval-aware Contrastive-of-Thought), a novel framew

Cited by 0SourcePDFScholar
2026

SOLAR for Offline MARL: Plateau-Triggered Potential Shaping under World-Model Uncertainty

ICML 2026poster

Reward shaping can accelerate reinforcement learning, but in sparse-reward \emph{offline} multi-agent RL it is often brittle: dense intrinsic rewards may alter the underlying Markov game, while world-model guidance can amplify model bias. We find that shaping becomes reliable when it is (i) activate…

Cited by 0SourceScholar
2026

Self-supervised Dynamic Heterogeneous Degradation Modeling for Unified Zero-Shot Image Restoration

CVPR 2026

Zero-shot image restoration provides a flexible way to handle diverse degradations without task-specific training. However, existing methods typically rely on stacked layers or pre-trained features to enhance degradation expression, while overlooking physically consistent priors. The insufficient de

Cited by 0SourcecodeScholar
2026

Text-Phase Synergy Network with Dual Priors for Unsupervised Cross-Domain Image Retrieval

CVPR 2026

This paper studies unsupervised cross-domain image retrieval (UCDIR), which aims to retrieve images of the same category across different domains without relying on labeled data. Existing methods typically utilize pseudo-labels, derived from clustering algorithms, as supervisory signals for intra-do

Cited by 0SourceScholar
2026

Top-Down Semantic Refinement for Image Captioning

AAAI 2026technical

Large Vision-Language Models (VLMs) face an inherent contradiction in image captioning: their powerful single-step generation capabilities often lead to a myopic decision-making process. This makes it difficult to maintain global narrative coherence while capturing rich details, a limitation that is

Cited by 0SourcePDFScholar
2026

Towards Multimodal Continual Knowledge Embedding with Modality Forgetting Modulation

AAAI 2026technical

The continuous emergence of new entities, relations, triples, and multimodal information drives the dynamic evolution of multimodal knowledge graph (MMKG). However, existing MMKG embedding models follow a static setting, where training from scratch for growing MMKG wastes learned knowledge, while fi

Cited by 0SourcePDFScholar
2026

Why Keep Your Doubts to Yourself? Trading Visual Uncertainties in Multi-Agent Bandit Systems

ICLR 2026poster

Vision-Language Models (VLMs) enable powerful multi-agent systems, but scaling them is economically unsustainable: coordinating heterogeneous agents under information asymmetry often spirals costs. Existing paradigms, such as Mixture-of-Agents and knowledge-based routers, rely on heuristic proxies t…

Cited by 0SourceScholar
2025

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis

UAI 2025

This paper investigates a hybrid learning framework for reinforcement learning (RL) in which the agent can leverage both an offline dataset and online interactions to learn the optimal policy. We present a unified algorithm and analysis and show that augmenting confidence-based online RL algorithms

Cited by 0SourcePDFScholar
2025

CCG: Rare-Label Prediction via Neural SEM–Driven Causal Game

EMNLP 2025

Multi-label classification (MLC) faces persistent challenges from label imbalance, spurious correlations, and distribution shifts, especially in rare label prediction. We propose the Causal Cooperative Game (CCG) framework, which models MLC as a multi-player cooperative process. CCG integrates expli

Cited by 0SourcePDFScholar
2025

D2-MLP: Dynamic Decomposed MLP Mixer for Medical Image Segmentation

ICASSP 2025accepted

Convolutional neural networks are widely used in various segmentation tasks in medical images. However, they are challenged to learn global features adaptively due to the inherent locality of convolutional operations. In contrast, MLP Mixers are proposed as a backbone to learn global information acr…

Cited by 0SourceScholar
2025

Data-adaptive Differentially Private Prompt Synthesis for In-Context Learning

ICLR 2025poster

Large Language Models (LLMs) rely on the contextual information embedded in examples/demonstrations to perform in-context learning (ICL). To mitigate the risk of LLMs potentially leaking private information contained in examples in the prompt, we introduce a novel data-adaptive differentially privat…

Cited by 1SourcePDFScholar
2025

Evaluating Cognitive-Behavioral Fixation via Multimodal User Viewing Patterns on Social Media

EMNLP 2025

Digital social media platforms frequently contribute to cognitive-behavioral fixation, a phenomenon in which users exhibit sustained and repetitive engagement with narrow content domains. While cognitive-behavioral fixation has been extensively studied in psychology, methods for computationally dete

2025

Exploiting Temporal State Space Sharing for Video Semantic Segmentation

CVPR 2025poster

Video semantic segmentation (VSS) plays a vital role in understanding the temporal evolution of scenes. Traditional methods often segment videos frame-by-frame or in a short temporal window, leading to limited temporal context, redundant computations, and heavy memory requirements. To this end, we i…

2025

FDDSGCN: Fractional Decoupling Dynamic Spatiotemporal Graph Convolutional Network for Traffic Forecasting

ICASSP 2025accepted

Urban traffic flow management faces increasing challenges due to accelerating urbanization. Traffic data collected from roadside sensors contain complex temporal and spatial dependencies that interact simultaneously. Although Graph Neural Networks and Recurrent Neural Networks have been successful i…

Cited by 0SourceScholar
2025

Fixing the Double Penalty in Data-Driven Weather Forecasting Through a Modified Spherical Harmonic Loss Function

ICML 2025poster

Recent advancements in data-driven weather forecasting models have delivered deterministic models that outperform the leading operational forecast systems based on traditional, physics-based models. However, these data-driven models are typically trained with a mean squared error loss function, whic…

2025

Gaussian Head & Shoulders: High Fidelity Neural Upper Body Avatars with Anchor Gaussian Guided Texture Warping

ICLR 2025poster

The ability to reconstruct realistic and controllable upper body avatars from casual monocular videos is critical for various applications in communication and entertainment. By equipping the most recent 3D Gaussian Splatting representation with head 3D morphable models (3DMM), existing methods mana…

Cited by 0SourcePDFScholar
2025

HUST: High-Fidelity Unbiased Skin Tone Estimation via Texture Quantization

ICCV 2025poster

Recent 3D facial reconstruction methods have made significant progress in shape estimation, but high-fidelity unbiased facial albedo estimation remains challenging. Existing methods rely on expensive light-stage captured data, and while they have made progress in either high-fidelity reconstruction…

2025

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias

ICML 2025poster

Language recognition tasks are fundamental in natural language processing (NLP) and have been widely used to benchmark the performance of large language models (LLMs). These tasks also play a crucial role in explaining the working mechanisms of transformers. In this work, we focus on two representat…

Cited by 0SourcePDFScholar
2025

HyperDiff: Masked Diffusion Model with High-efficient Transformer for Hyperspectral Image Cross-Scene Classification

ICASSP 2025accepted

Hyperspectral Image (HSI) cross-scene classification is a challenging task in remote sensing, particularly when real-time processing of Target Domain (TD) HSI is required, and data cannot be reused for training. While deep learning methods have shown promising results, the generalization ability of…

Cited by 0SourceScholar
2025

MGMapNet: Multi-Granularity Representation Learning for End-to-End Vectorized HD Map Construction

ICLR 2025poster

The construction of vectorized high-definition map typically requires capturing both category and geometry information of map elements. Current state-of-the-art methods often adopt solely either point-level or instance-level representation, overlooking the strong intrinsic relationship between point…

Cited by 3SourcePDFScholar
2025

Multi-Type Preference Learning: Empowering Preference-Based Reinforcement Learning with Equal Preferences

ICRA 2025

Preference-Based reinforcement learning (PBRL) learns directly from the preferences of human teachers regarding agent behaviors without needing meticulously designed reward functions. However, existing PBRL methods often learn primarily from explicit preferences, neglecting the possibility that teac

Cited by 1SourcecodeScholar
2025

Multiple Sclerosis Detection with Reinforcement Learning and Differential Evolution

ICASSP 2025accepted

Multiple Sclerosis (MS) disrupts nerve communication, potentially leading to permanent damage. Convolutional Neural Networks (CNNs) are commonly recommended to accelerate magnetic resonance imaging (MRI) analysis for MS. Traditional CNN-based methods often face challenges with feature selection, imb…

Cited by 0SourceScholar
2025

Occlusion-Aware 6D Pose Estimation with Depth-Guided Graph Encoding and Cross-Semantic Fusion for Robotic Grasping

ICRA 2025

Reliable 6D pose estimation is crucial for robotic tasks but presents significant challenges in environments with occlusion. Recent approaches tend to directly predict pose parameters of object with deep neural networks, lacking the modeling ability of non-adjacent and complex relationships of surfa

Cited by 3SourceScholar
2025

On the Learn-to-Optimize Capabilities of Transformers in In-Context Sparse Recovery

ICLR 2025poster

An intriguing property of the Transformer is its ability to perform in-context learning (ICL), where the Transformer can solve different inference tasks without parameter updating based on the contextual information provided by the corresponding input-output demonstration pairs. It has been theoreti…

Cited by 0SourcePDFScholar
2025

On the Training Convergence of Transformers for In-Context Classification of Gaussian Mixtures

ICML 2025poster

Although transformers have demonstrated impressive capabilities for in-context learning (ICL) in practice, theoretical understanding of the underlying mechanism that allows transformers to perform ICL is still in its infancy. This work aims to theoretically study the training dynamics of transformer…

Cited by 0SourcePDFScholar
2025

Position: The Artificial Intelligence and Machine Learning Community Should Adopt a More Transparent and Regulated Peer Review Process

ICML 2025poster

The rapid growth of submissions to top-tier Artificial Intelligence (AI) and Machine Learning (ML) conferences has prompted many venues to transition from closed to open review platforms. Some have fully embraced open peer reviews, allowing public visibility throughout the process, while others adop…

Cited by 0SourcePDFScholar
2025

Pretrained Reversible Generation as Unsupervised Visual Representation Learning

ICCV 2025poster

Recent generative models based on score matching and flow matching have significantly advanced generation tasks, but their potential in discriminative tasks remains underexplored. Previous approaches, such as generative classifiers, have not fully leveraged the capabilities of these models for discr…

2025

Singing Voice Conversion with Accompaniment Using Self-Supervised Representation-Based Melody Features

ICASSP 2025accepted

Melody preservation is crucial in singing voice conversion (SVC). However, in many scenarios, audio is often accompanied with background music (BGM), which can cause audio distortion and interfere with the extraction of melody and other key features, significantly degrading SVC performance. Previous…

Cited by 0SourceScholar
2025

Uncover Treasures in DCT: Advancing JPEG Quality Enhancement by Exploiting Latent Correlations

ICCV 2025poster

Joint Photographic Experts Group (JPEG) achieves data compression by quantizing Discrete Cosine Transform (DCT) coefficients, which inevitably introduces compression artifacts. Most existing JPEG quality enhancement methods operate in the pixel domain, suffering from the high computational costs of…

Cited by 0SourcePDFScholar
2024

Build a 50+ Hours Chinese Mandarin Corpus for Children's Speech Recognition

ICASSP 2024accepted

Children’s speech recognition plays an important role in the education research of children. The usual automatic speech recognition (ASR) systems are not satisfactory in terms of speech recognition for children, mainly due to the lack of child speech corpus. In recent years, there have been a large…

Cited by 0SourceScholar
2024

DiffPortrait3D: Controllable Diffusion for Zero-Shot Portrait View Synthesis

CVPR 2024highlight

We present DiffPortrait3D a conditional diffusion model that is capable of synthesizing 3D-consistent photo-realistic novel views from as few as a single in-the-wild portrait. Specifically given a single RGB input we aim to synthesize plausible but consistent facial details rendered from novel camer…

2024

Efficient Prompt Optimization Through the Lens of Best Arm Identification

NeurIPS 2024poster

The remarkable instruction-following capability of large language models (LLMs) has sparked a growing interest in automatically finding good prompts, i.e., prompt optimization. Most existing works follow the scheme of selecting from a pre-generated pool of candidate prompts. However, these designs m…

Cited by 7SourcePDFScholar
2024

Federated Online Prediction from Experts with Differential Privacy: Separations and Regret Speed-ups

NeurIPS 2024poster

We study the problems of differentially private federated online prediction from experts against both *stochastic adversaries* and *oblivious adversaries*. We aim to minimize the average regret on $m$ clients working in parallel over time horizon $T$ with explicit differential privacy (DP) guarantee…

Cited by 0SourcePDFScholar
2024

Federated Q-Learning: Linear Regret Speedup with Low Communication Cost

ICLR 2024poster

In this paper, we consider federated reinforcement learning for tabular episodic Markov Decision Processes (MDP) where, under the coordination of a central server, multiple agents collaboratively explore the environment and learn an optimal policy without sharing their raw data. While linear speedu…

Cited by 14SourcePDFScholar
2024

Improving Sample Efficiency of Model-Free Algorithms for Zero-Sum Markov Games

ICML 2024poster

The problem of two-player zero-sum Markov games has recently attracted increasing interests in theoretical studies of multi-agent reinforcement learning (RL). In particular, for finite-horizon episodic Markov decision processes (MDPs), it has been shown that model-based algorithms can find an $\epsi…

Cited by 1SourcePDFScholar
2024

Multi-View Midivae: Fusing Track- and Bar-View Representations for Long Multi-Track Symbolic Music Generation

ICASSP 2024accepted

Variational Autoencoders (VAEs) constitute a crucial component of neural symbolic music generation, among which some works have yielded outstanding results and attracted considerable attention. Nevertheless, previous VAEs still encounter issues with overly long feature sequences and generated result…

Cited by 0SourceScholar
2024

Non-asymptotic Convergence of Training Transformers for Next-token Prediction

NeurIPS 2024poster

Transformers have achieved extraordinary success in modern machine learning due to their excellent ability to handle sequential data, especially in next-token prediction (NTP) tasks. However, the theoretical understanding of their performance in NTP is limited, with existing studies focusing mainly…

Cited by 4SourcePDFScholar
2024

PMET: Precise Model Editing in a Transformer

AAAI 2024technical

Model editing techniques modify a minor proportion of knowledge in Large Language Models (LLMs) at a relatively low cost, which have demonstrated notable success. Existing methods assume Transformer Layer (TL) hidden states are values of key-value memories of the Feed-Forward Network (FFN). They usu…

2024

Provable Benefits of Multi-task RL under Non-Markovian Decision Making Processes

ICLR 2024poster

In multi-task reinforcement learning (RL) under Markov decision processes (MDPs), the presence of shared latent structures among multiple MDPs has been shown to yield significant benefits to the sample efficiency compared to single-task RL. In this paper, we investigate whether such a benefit can ex…

Cited by 1SourcePDFScholar
2024

Provably Efficient UCB-type Algorithms For Learning Predictive State Representations

ICLR 2024poster

The general sequential decision-making problem, which includes Markov decision processes (MDPs) and partially observable MDPs (POMDPs) as special cases, aims at maximizing a cumulative reward by making a sequence of decisions based on a history of observations and actions over time. Recent studies h…

Cited by 6SourcePDFScholar
2024

Rethinking Word-level Adversarial Attack: The Trade-off between Efficiency, Effectiveness, and Imperceptibility

COLING 2024main

Neural language models have demonstrated impressive performance in various tasks but remain vulnerable to word-level adversarial attacks. Word-level adversarial attacks can be formulated as a combinatorial optimization problem, and thus, an attack method can be decomposed into search space and searc…

Cited by 3SourcePDFScholar
2024

Smaller and Faster Robotic Grasp Detection Model via Knowledge Distillation and Unequal Feature Encoding

RA-L 2024

In order to achieve higher accuracy, the complexity of grasp detection network increases accordingly with complicated model structures and tremendous parameters. Although various light-weight strategies are adopted, directly designing the compact network can be sub-optimal and difficult to strike th

Cited by 12SourceScholar
2024

Transformers as Game Players: Provable In-context Game-playing Capabilities of Pre-trained Models

NeurIPS 2024poster

The in-context learning (ICL) capability of pre-trained models based on the transformer architecture has received growing interest in recent years. While theoretical understanding has been obtained for ICL in reinforcement learning (RL), the previous results are largely confined to the single-agent…

Cited by 1SourcePDFScholar
2023

ALIP: Adaptive Language-Image Pre-Training with Synthetic Caption

ICCV 2023poster

Contrastive Language-Image Pre-training (CLIP) has significantly boosted the performance of various vision-language tasks by scaling up the dataset with image-text pairs collected from the web. However, the presence of intrinsic noise and unmatched image-text pairs in web data can potentially affect…

Cited by 54PDFcodeScholar
2023

Contrastive Learning with Adversarial Examples for Alleviating Pathology of Language Model

ACL 2023long

Neural language models have achieved superior performance. However, these models also suffer from the pathology of overconfidence in the out-of-distribution examples, potentially making the model difficult to interpret and making the interpretation methods fail to provide faithful attributions. In t…

Cited by 4SourcePDFScholar
2023

Federated Linear Contextual Bandits with User-level Differential Privacy

ICML 2023poster

This paper studies federated linear contextual bandits under the notion of user-level differential privacy (DP). We first introduce a unified federated bandits framework that can accommodate various definitions of DP in the sequential decision-making setting. We then formally introduce user-level ce…

Cited by 18SourcePDFScholar
2023

Improved Sample Complexity for Reward-free Reinforcement Learning under Low-rank MDPs

ICLR 2023poster

In reward-free reinforcement learning (RL), an agent explores the environment first without any reward information, in order to achieve certain learning goals afterwards for any given reward. In this paper we focus on reward-free RL under low-rank MDP models, in which both the representation and lin…

Cited by 11SourcePDFScholar
2023

Light Sampling Field and BRDF Representation for Physically-based Neural Rendering

ICLR 2023poster

Physically-based rendering (PBR) is key for immersive rendering effects used widely in the industry to showcase detailed realistic scenes from computer graphics assets. A well-known caveat is that producing the same is computationally heavy and relies on complex capture devices. Inspired by the succ…

2023

Near-optimal Conservative Exploration in Reinforcement Learning under Episode-wise Constraints

ICML 2023poster

This paper investigates conservative exploration in reinforcement learning where the performance of the learning agent is guaranteed to be above a certain threshold throughout the learning process. It focuses on the tabular episodic Markov Decision Process (MDP) setting that has finite states and ac…

Cited by 3SourcePDFScholar
2023

Non-stationary Reinforcement Learning under General Function Approximation

ICML 2023poster

General function approximation is a powerful tool to handle large state and action spaces in a broad range of reinforcement learning (RL) scenarios. However, theoretical understanding of non-stationary MDPs with general function approximation is still limited. In this paper, we make the first such a…

Cited by 8SourcePDFScholar
2023

Provably Efficient Offline Reinforcement Learning with Perturbed Data Sources

ICML 2023poster

Existing theoretical studies on offline reinforcement learning (RL) mostly consider a dataset sampled directly from the target task. In practice, however, data often come from several heterogeneous but related sources. Motivated by this gap, this work aims at rigorously understanding offline RL with…

Cited by 4SourcePDFScholar
2023

Safe Exploration Incurs Nearly No Additional Sample Complexity for Reward-Free RL

ICLR 2023poster

Reward-free reinforcement learning (RF-RL), a recently introduced RL paradigm, relies on random action-taking to explore the unknown environment without any reward feedback information. While the primary goal of the exploration phase in RF-RL is to reduce the uncertainty in the estimated model with…

Cited by 6SourcePDFScholar
2023

Similarizing the Influence of Words with Contrastive Learning to Defend Word-level Adversarial Text Attack

ACL 2023findings

Neural language models are vulnerable to word-level adversarial text attacks, which generate adversarial examples by directly substituting discrete input words. Previous search methods for word-level attacks assume that the information in the important words is more influential on prediction than un…

Cited by 7SourcePDFScholar
2023

Uncertainty-Aware Few-Shot Class-Incremental Learning

ICASSP 2023accepted

In a real-world setting, machine needs to continuously recognize new categories without forgetting. However, the number of new categories may be small. For some difficult categories, even humans cannot recognize only based on few-shot examples. To address the above issues, an innovative uncertainty-…

Cited by 0SourceScholar
2023

Unicom: Universal and Compact Representation Learning for Image Retrieval

ICLR 2023poster

Modern image retrieval methods typically rely on fine-tuning pre-trained encoders to extract image-level descriptors. However, the most widely used models are pre-trained on ImageNet-1K with limited classes. The pre-trained feature representation is therefore not universal enough to generalize well…

2023

VERGNet: Visual Enhancement Guided Robotic Grasp Detection Under Low-Light Condition

RA-L 2023

Although existing grasp detection methods have achieved encouraging performance under well-light conditions, repetitive experiments have found that the detection performance would deteriorate drastically under low-light conditions. Although supplementary information can be provided by additional sen

Cited by 23SourceScholar
2022

Explainable Fact-Checking Through Question Answering

ICASSP 2022accepted

Misleading or false information has been creating chaos in some places around the world. To mitigate this issue, many researchers have proposed automated fact-checking methods to fight the spread of fake news. However, most methods cannot explain the reasoning behind their decisions, failing to buil…

Cited by 0SourceScholar
2022

Feature Generation and Hypothesis Verification for Reliable Face Anti-spoofing

AAAI 2022technical

Although existing face anti-spoofing (FAS) methods achieve high accuracy in intra-domain experiments, their effects drop severely in cross-domain scenarios because of poor generalization. Recently, multifarious techniques have been explored, such as domain generalization and representation disentang…

2022

Few-shot Named Entity Recognition with Entity-level Prototypical Network Enhanced by Dispersedly Distributed Prototypes

COLING 2022main

Few-shot named entity recognition (NER) enables us to build a NER system for a new domain using very few labeled examples. However, existing prototypical networks for this task suffer from roughly estimated label dependency and closely distributed prototypes, thus often causing misclassifications. T…

Cited by 36SourcePDFScholar
2022

Killing Two Birds With One Stone: Efficient and Robust Training of Face Recognition CNNs by Partial FC

CVPR 2022poster

Learning discriminative deep feature embeddings by using million-scale in-the-wild datasets and margin-based softmax loss is the current state-of-the-art approach for face recognition. However, the memory and computing cost of the Fully Connected (FC) layer linearly scales up to the number of identi…

Cited by 104PDFcodeScholar
2022

Multi-Channel Attentive Graph Convolutional Network with Sentiment Fusion for Multimodal Sentiment Analysis

ICASSP 2022accepted

Nowadays, with the explosive growth of multimodal reviews on social media platforms, multimodal sentiment analysis has recently gained popularity because of its high relevance to these social media posts. Although most previous studies design various fusion frameworks for learning an interactive rep…

Cited by 0SourceScholar
2022

PARSE: An Efficient Search Method for Black-box Adversarial Text Attacks

COLING 2022main

Neural networks are vulnerable to adversarial examples. The adversary can successfully attack a model even without knowing model architecture and parameters, i.e., under a black-box scenario. Previous works on word-level attacks widely use word importance ranking (WIR) methods and complex search met…

Cited by 9SourcePDFScholar
2022

Pre-training Strategies and Datasets for Facial Representation Learning

ECCV 2022poster

"What is the best way to learn a universal face representation? Recent work on Deep Learning in the area of face analysis has focused on supervised learning for specific tasks of interest (e.g. face recognition, facial landmark localization etc.) but has overlooked the overarching question of how to…

2022

Provable Benefit of Multitask Representation Learning in Reinforcement Learning

NeurIPS 2022accept

As representation learning becomes a powerful technique to reduce sample complexity in reinforcement learning (RL) in practice, theoretical understanding of its advantage is still limited. In this paper, we theoretically characterize the benefit of representation learning under the low-rank Markov d…

Cited by 28SourcePDFScholar
2021

(W)Earable Microphone Array and Ultrasonic Echo Localization for Coarse Indoor Environment Mapping

ICASSP 2021accepted

We present a microphone array structure for spherical sound incidence angle tracking that can be attached to headphones or directly integrated into earphones. We show that this microphone array together with an ultrasonic sound source, e.g., a home assistant speaker in the room, allows to estimate t…

Cited by 0SourceScholar
2021

Cross-Modal Knowledge Distillation For Fine-Grained One-Shot Classification

ICASSP 2021accepted

Few-shot learning can recognize a novel category based on only a few samples because it learns to learn from a lot of labeled samples during the training process. When data is insufficient, the performance is affected. And it is expensive to obtain a large-scale finegrained dataset with annotation.…

Cited by 0SourceScholar
2021

Heterogeneous Multi-player Multi-armed Bandits: Closing the Gap and Generalization

NeurIPS 2021poster

Despite the significant interests and many progresses in decentralized multi-player multi-armed bandits (MP-MAB) problems in recent years, the regret gap to the natural centralized lower bound in the heterogeneous MP-MAB setting remains open. In this paper, we propose BEACON -- Batched Exploration w…

2021

Knowledge distillation via softmax regression representation learning

ICLR 2021poster

This paper addresses the problem of model compression via knowledge distillation. We advocate for a method that optimizes the output feature of the penultimate layer of the student network and hence is directly related to representation learning. Previous distillation methods which typically impose…

2021

Looking Wider for Better Adaptive Representation in Few-Shot Learning

AAAI 2021technical

Building a good feature space is essential for the metric-based few-shot algorithms to recognize a novel class with only a few samples. The feature space is often built by Convolutional Neural Networks (CNNs). However, CNNs primarily focus on local information with the limited receptive field, and t…

Cited by 58SourcePDFScholar
2021

Variational Prototype Learning for Deep Face Recognition

CVPR 2021poster

Deep face recognition has achieved remarkable improvements due to the introduction of margin-based softmax loss, in which the prototype stored in the last linear layer represents the center of each class. In these methods, training samples are enforced to be close to positive prototypes and far apar…

Cited by 100PDFScholar
2020

Decentralized Multi-player Multi-armed Bandits with No Collision Information

AISTATS 2020poster

The decentralized stochastic multi-player multi-armed bandit (MP-MAB) problem, where the collision information is not available to the players, is studied in this paper. Building on the seminal work of Boursier and Perchet (2019), we propose error correction synchronization involving communication (…

Cited by 45SourcePDFScholar
2020

Training binary neural networks with real-to-binary convolutions

ICLR 2020poster

This paper shows how to train binary networks to within a few percent points (~3-5%) of the full precision counterpart. We first show how to build a strong baseline, which already achieves state-of-the-art accuracy, by combining recently proposed advances and carefully adjusting the optimization pro…

Cited by 298SourcecodeScholar
2018

To learn image super-resolution, use a GAN to learn how to do image degradation first

ECCV 2018poster

This paper is on image and face super-resolution. The vast majority of prior work for this problem focus on how to increase the resolution of low-resolution images which are artificially generated by simple bilinear down-sampling (or in a few cases by blurring followed by down-sampling). We show tha…