← Search

Xi Wang

92 accepted papers

2026

3DCS: Datasets and Benchmark for Evaluating Conformational Sensitivity in Molecular Representations

ICLR 2026poster

Molecular representations (MRs) that capture 3D conformations are critical for applications such as reaction prediction, drug design, and material discovery. Yet despite the rapid development of molecular representation models, there is no comprehensive benchmark to evaluate their treatment of 3D co…

Cited by 0SourcecodeScholar
2026

Attributing Response to Context: A Jensen–Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented Generation

ICLR 2026poster

Retrieval-Augmented Generation (RAG) leverages large language models (LLMs) combined with external contexts to enhance the accuracy and reliability of generated responses. However, reliably attributing generated content to specific context segments, context attribution, remains challenging due to th…

Cited by 0SourceScholar
2026

Autonomous Vehicle Path Planning by Searching with Differentiable Simulation

AAAI 2026technical

Planning allows an agent to safely refine its actions before executing them in the real world. In autonomous driving, this is crucial to avoid collisions and navigate in complex, dense traffic scenarios. One way to plan is to search for the best action sequence. However, this is challenging when all

Cited by 0SourcePDFScholar
2026

Bi-Stable Thin Soft Robot for In-Plane Locomotion in Narrow Space

ICRA 2026poster

Dielectric elastomer actuators (DEAs), also recognized as artificial muscle, have been widely developed for the soft locomotion robot. With the complaint skeleton and miniaturized dimension, they are well suited for the narrow space inspection. In this work, we propose a novel low profile (1.1mm) an…

2026

BotVA: Combating Social Bots via Variational Feature Augmentation and Adversarial Graph Learning

IJCAI 2026

Social bots threaten online platforms by spreading disinformation and manipulating public discourse. Graph neural networks have emerged as effective tools for bot detection by modeling user interactions, yet two fundamental challenges limit their practical deployment: severe class imbalance where bo

Cited by 0Scholar
2026

ChipMind: Retrieval-Augmented Reasoning for Long-Context Circuit Design Specifications

AAAI 2026technical

While Large Language Models (LLMs) demonstrate immense potential for automating integrated circuit (IC) development, their practical deployment is fundamentally limited by restricted context windows. Existing context-extension methods struggle to achieve effective semantic modeling and thorough mult

Cited by 0SourcePDFScholar
2026

Compute When Worth It: Risk Control for Reasoning on a Compute Budget

ICML 2026poster

Reasoning Large Language Models (LLMs) enable test-time scaling, with dataset-level accuracy improving as the token budget increases, motivating adaptive reasoning---spending tokens when they improve reliability and stopping early when additional computation is unlikely to help. However, setting the…

Cited by 0SourceScholar
2026

Decomposing Prompts, Composing Actions: A Multi-Granularity Prompting Approach for Incremental Action Learning

AAAI 2026technical

Continual learning for action recognition is a critical capability for next-generation Extended Reality (XR) systems. Yet it faces a severe real-world challenge: strict user privacy that prohibits data rehearsal. While recent prompt-based continual learning methods show promise, we argue their core

Cited by 0SourcePDFScholar
2026

EgoFlow: Gradient-Guided Flow Matching for Egocentric 6DoF Object Motion Generation

CVPR 2026

Understanding and predicting object motion from egocentric video is fundamental to embodied perception and interaction. However, generating physically consistent 6DoF trajectories remains challenging due to occlusions, fast motion, and the lack of explicit physical reasoning in existing generative m

Cited by 0SourcecodeScholar
2026

EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World

RSS 2026poster

Robot learning increasingly depends on large and diverse data, yet robot data collection remains expensive and difficult to scale. Egocentric human data offer a promising alternative by capturing rich manipulation behavior across everyday environments. However, existing human datasets are often limi…

Cited by 0SourceScholar
2026

FIXME: Towards End-to-End Benchmarking of LLM-Aided Design Verification

AAAI 2026technical

We introduce FIXME, the first end-to-end and large-scale benchmark for evaluating Large Language Models (LLMs) in hardware design functional verification (FV). Comprising 747 tasks derived from real-world hardware designs, FIXME spans five core FV sub-sets: specification comprehension, reference mod

Cited by 0SourcePDFScholar
2026

GaussianVLM: Scene-Centric 3D Vision-Language Models Using Language-Aligned Gaussian Splats for Embodied Reasoning and Beyond

ICRA 2026poster

As multimodal language models advance, their application to 3D scene understanding is a fast-growing frontier, driving the development of 3D Vision-Language Models (VLMs). Current methods show strong dependence on object detectors, introducing processing bottlenecks and limitations in taxonomic flex…

2026

Hallucination as a Computational Boundary: A Hierarchy of Inevitability and the Oracle Escape

AAAI 2026technical

The illusion phenomenon of large language models (LLMs) is the core obstacle to their reliable deployment. This article formalizes the large language model as a probabilistic Turing machine by constructing a "computational necessity hierarchy", and for the first time proves the illusions are inevita

Cited by 0SourcePDFScholar
2026

Pulp Motion: Framing-aware multimodal camera and human motion generation

ICLR 2026poster

Treating human motion and camera trajectory generation separately overlooks a core principle of cinematography: the tight interplay between actor performance and camera work in the screen space. In this paper, we are the first to cast this task as a text-conditioned joint generation, aiming to main…

Cited by 0SourceScholar
2026

RABot: Reinforcement-Guided Graph Augmentation for Imbalanced and Noisy Social Bot Detection

AAAI 2026technical

Social bot detection is pivotal for safeguarding the integrity of online information ecosystems. Although recent graph neural network (GNN) solutions achieve strong results, they remain hindered by two practical challenges: (i) severe class imbalance arising from the high cost of generating bots, an

Cited by 0SourcePDFScholar
2026

RapTB: Rooted Absorbed Trajectory Balance with Submodular Replay for Stable Autoregressive GFlowNet Training

ICML 2026poster

Generative Flow Networks (GFlowNets) enable fine-tuning large language models to approximate reward-proportional posteriors, but they remain prone to mode collapse, manifesting as prefix collapse and length bias. We attribute this to two factors: (i) weak credit assignment to early prefixes, and (ii…

Cited by 0SourceScholar
2026

Test-Time Perturbation Tuning with Delayed Feedback for Vision-Language-Action Models

CVPR 2026

Vision-Language-Action models (VLAs) achieve remarkable performance in sequential decision-making but remain fragile to subtle environmental shifts, such as small changes in object pose. We attribute this brittleness to trajectory overfitting, where VLAs over-attend to the spurious correlation betwe

Cited by 0SourcecodeScholar
2026

TokenPowerBench: Benchmarking the Power Consumption of LLM Inference

AAAI 2026technical

Large language model (LLM) services now answer billions of queries per day, and industry reports show that inference, not training, accounts for more than 90% of total power consumption. However, existing benchmarks focus on either training/fine-tuning or performance of inference and provide little

Cited by 0SourcePDFScholar
2025

AKiRa: Augmentation Kit on Rays for Optical Video Generation

CVPR 2025poster

Recent advances in text-conditioned video diffusion have greatly improved video quality. However, these methods offer limited or sometimes no control to users on camera aspects, including dynamic camera motion, zoom, distorted lens and focus shifts. These motion and optical aspects are crucial for a…

2025

Adaptive Retrieval-Augmented Generation for Conversational Systems

NAACL 2025findings

With the success of integrating large language models into the development of conversational systems, many studies have shown the effectiveness of retrieving and augmenting external knowledge for informative responses. While many existing studies agree on the necessity of Retrieval Augmented Generat…

2025

Analyzing values about gendered language reform in LLMs’ revisions

EMNLP 2025

Within the common LLM use case of text revision, we study LLMs’ revision of gendered role nouns (e.g., outdoorsperson/woman/man) and their justifications of such revisions. We evaluate their alignment with feminist and trans-inclusive language reforms for English. Drawing on insight from sociolingui

Cited by 0SourcePDFScholar
2025

Articulate3D: Holistic Understanding of 3D Scenes as Universal Scene Description

ICCV 2025poster

3D scene understanding is a long-standing challenge in computer vision and a key component in enabling mixed reality, wearable computing, and embodied AI. Providing a solution to these applications requires a multifaceted approach that covers scene-centric, object-centric, as well as interaction-cen…

Cited by 0SourcePDFScholar
2025

Bi-Stable Thin Soft Robot for in-Plane Locomotion in Narrow Space

RA-L 2025

Dielectric elastomer actuators (DEAs), also recognised as artificial muscle, have been widely developed for the soft locomotion robot. With the complaint skeleton and miniaturised dimension, they are well suited for the narrow space inspection. In this work, we propose a novel low profile (1.1mm) an

Cited by 3SourceScholar
2025

DCTMamba: Advancing JPEG Image Restoration Through Long-Sequence Modeling and Adaptive Frequency Strategy

AAAI 2025technical

Despite the advanced long-sequence modeling of Mamba, which has expanded its applications in image restoration, there remains a lack of exploration combining its strengths with the specific characteristics of JPEG image restoration, where high-frequency components are lost after the Discrete Cosine…

2025

Di[M]O: Distilling Masked Diffusion Models into One-step Generator

ICCV 2025poster

Masked Diffusion Models (MDMs) have emerged as a powerful generative modeling technique. Despite their remarkable results, they typically suffer from slow inference with several steps. In this paper, we propose Di\mathtt [M] O, a novel approach that distills masked diffusion models into a one-step g…

2025

Dual-Space Semantic Synergy Distillation for Continual Learning of Unlabeled Streams

NeurIPS 2025poster

Continual learning from unlabeled data streams while effectively combating catastrophic forgetting poses an intractable challenge. Traditional methods predominantly rely on visual clustering techniques to generate pseudo labels, which are frequently plagued by problems such as noise and suboptimal q…

Cited by 0SourceScholar
2025

Exploration-Driven Generative Interactive Environments

CVPR 2025poster

Modern world models require costly and time-consuming collection of large video datasets with action demonstrations by people or by environment-specific agents. To simplify training, we focus on using many virtual environments for inexpensive, automatically collected interaction data. Genie, a recen…

2025

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control

CVPR 2025poster

We present GEM, a Generalizable Ego-vision Multimodal world model that predicts future frames using a reference frame, sparse features, human poses, and ego-trajectories. Hence, our model has precise control over object dynamics, ego-agent motion and human poses. GEM generates paired RGB and depth o…

2025

LangHOPS: Language Grounded Hierarchical Open-Vocabulary Part Segmentation

NeurIPS 2025poster

We propose LangHOPS, the first Multimodal Large Language Model (MLLM)-based framework for open-vocabulary object–part instance segmentation. Given an image, LangHOPS can jointly detect and segment hierarchical object and part instances from open-vocabulary candidate categories. Unlike prior approach…

Cited by 0SourceScholar
2025

Leveraging Driver Field-of-View for Multimodal Ego-Trajectory Prediction

ICLR 2025poster

Understanding drivers’ decision-making is crucial for road safety. Although predicting the ego-vehicle’s path is valuable for driver-assistance systems, existing methods mainly focus on external factors like other vehicles’ motions, often neglecting the driver’s attention and intent. To address this…

2025

LineArt: A Knowledge-guided Training-free High-quality Appearance Transfer for Design Drawing with Diffusion Model

CVPR 2025poster

Image rendering from line drawings is vital in design and image generation technologies reduce costs, yet professional line drawings demand preserving complex details. Text prompts struggle with accuracy, and image translation struggles with consistency and fine-grained control. We present LineArt,…

Cited by 1SourcePDFScholar
2025

MADial-Bench: Towards Real-world Evaluation of Memory-Augmented Dialogue Generation

NAACL 2025long

Long-term memory is important for chatbots and dialogue systems (DS) to create consistent and human-like conversations, evidenced by numerous developed memory-augmented DS (MADS). To evaluate the effectiveness of such MADS, existing commonly used evaluation metrics, like retrieval accuracy and perpl…

2025

OMS: One More Step Noise Searching to Enhance Membership Inference Attacks for Diffusion Models

IJCAI 2025

The data-intensive nature of Diffusion models amplifies the risks of privacy infringements and copyright disputes, particularly when training on extensive unauthorized data scraped from the Internet. Membership Inference Attacks (MIA) aim to determine whether a data sample has been utilized by the t

Cited by 0SourcePDFScholar
2025

SWEA: Updating Factual Knowledge in Large Language Models via Subject Word Embedding Altering

AAAI 2025technical

The general capabilities of large language models (LLMs) make them the infrastructure for various AI applications, but updating their inner knowledge requires significant resources. Recent model editing is a promising technique for efficiently updating a small amount of knowledge of LLMs and has att…

2025

Stand on The Shoulders of Giants: Building JailExpert from Previous Attack Experience

EMNLP 2025

Large language models (LLMs) generate human-aligned content under certain safety constraints. However, the current known technique “jailbreak prompt” can circumvent safety-aligned measures and induce LLMs to output malicious content. Research on Jailbreaking can help identify vulnerabilities in LLMs

Cited by 0SourcePDFScholar
2025

StateSpaceDiffuser: Bringing Long Context to Diffusion World Models

NeurIPS 2025poster

World models have recently gained prominence for action-conditioned visual prediction in complex environments. However, relying on only a few recent observations causes them to lose long-term context. Consequently, within a few steps, the generated scenes drift from what was previously observed, und…

Cited by 0SourcecodeScholar
2025

Understanding Museum Exhibits using Vision-Language Reasoning

ICCV 2025poster

Museums serve as repositories of cultural heritage and historical artifacts from diverse epochs, civilizations, and regions, preserving well-documented collections that encapsulate vast knowledge, which, when systematically structured into large-scale datasets, can train specialized models. Visitors…

Cited by 0SourcePDFScholar
2025

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training

ICASSP 2025accepted

Style voice conversion aims to transform the speaking style of source speech into a desired style while keeping the original speaker’s identity. However, previous style voice conversion approaches primarily focus on well-defined domains such as emotional aspects, limiting their practical application…

Cited by 0SourceScholar
2024

An Empirical Investigation of Domain Adaptation Ability for Chinese Spelling Check Models

ICASSP 2024accepted

Chinese Spelling Check (CSC) is a meaningful task in the area of Natural Language Processing (NLP) which aims at detecting spelling errors in Chinese texts and then correcting these errors. However, CSC models are based on pretrained language models, which are trained on a general corpus. Consequent…

Cited by 0SourceScholar
2024

Bayesian Low-rank Adaptation for Large Language Models

ICLR 2024poster

Parameter-efficient fine-tuning (PEFT) has emerged as a new paradigm for cost-efficient fine-tuning of large language models (LLMs), with low-rank adaptation (LoRA) being a widely adopted choice. However, fine-tuned LLMs often become overconfident especially when fine-tuned on small datasets. Bayesi…

Cited by 75SourcePDFScholar
2024

Cut, Bond and Play: Volume-Preserved Reprogrammable Soft Pneumatic Actuators

RA-L 2024

Reprogrammable design enables soft actuators to change their performances after fabrication and obtain new deformation modes and functions. The reprogrammable design of soft pneumatic actuators is relatively difficult due to the intrinsic material being chemically inactive. This work proposes a new

Cited by 5SourceScholar
2024

Development of Negative-Pressure Artificial Muscles With Fiber Constraints and Pre-Stretched Soft Skin

RA-L 2024

Negative-pressure artificial muscles based on internal support and flexible skin provide a way to develop high-performance artificial muscles. However, the random and disordered skin wrinkles formed during the contraction may lead to uncertainty in the actuation behavior, and the flexible but inexte

Cited by 3SourceScholar
2024

E.T. the Exceptional Trajectory: Text-to-camera-trajectory generation with character awareness

ECCV 2024poster

"Stories and emotions in movies emerge through the effect of well-thought-out directing decisions, in particular camera placement and movement over time. Crafting compelling camera trajectories remains a complex iterative process, even for skilful artists. To tackle this, in this paper, we propose a…

Cited by 3SourcePDFScholar
2024

Extragradient Type Methods for Riemannian Variational Inequality Problems

AISTATS 2024poster

In this work, we consider monotone Riemannian Variational Inequality Problems (RVIPs), which encompass both Riemannian convex optimization and minimax optimization as particular cases. In Euclidean space, the last-iterates of both the extragradient (EG) and past extragradient (PEG) methods converge…

Cited by 7SourcePDFScholar
2024

Long-Tail Class Incremental Learning via Independent Sub-prototype Construction

CVPR 2024poster

Long-tail class incremental learning (LT-CIL) is designed to perpetually acquire novel knowledge from an imbalanced and perpetually evolving data stream while ensuring the retention of previously acquired knowledge. The existing method only re-balances data distribution and ignores exploring the pot…

Cited by 5SourcePDFScholar
2024

NeRFail: Neural Radiance Fields-Based Multiview Adversarial Attack

AAAI 2024technical

Adversarial attacks, i.e., generating adversarial perturbations with a small magnitude to deceive deep neural networks, are important for investigating and improving model trustworthiness. Traditionally, the topic was scoped within 2D images without considering 3D multiview information. Benefiting f…

2024

PALM: Predicting Actions through Language Models

ECCV 2024poster

"Understanding human activity is a crucial yet intricate task in egocentric vision, a field that focuses on capturing visual perspectives from the camera wearer’s viewpoint. Traditional methods heavily rely on representation learning that is trained on a large amount of video data. However, a major…

2024

Real Appearance Modeling for More General Deepfake Detection

ECCV 2024poster

"Recent studies in deepfake detection have shown promising results when detecting deepfakes of the same type as those present in training. However, their ability to generalize to unseen deepfakes remains limited. This work improves the generalizable deepfake detection from a simple principle: an ide…

Cited by 3SourcePDFScholar
2024

Source-Free Domain-Invariant Performance Prediction

ECCV 2024poster

"Accurately estimating model performance poses a significant challenge, particularly in scenarios where the source and target domains follow different data distributions. Most existing performance prediction methods heavily rely on the source data in their estimation process, limiting their applicab…

2024

SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation

NeurIPS 2024spotlight

We introduce SpreadsheetBench, a challenging spreadsheet manipulation benchmark exclusively derived from real-world scenarios, designed to immerse current large language models (LLMs) in the actual workflow of spreadsheet users. Unlike existing benchmarks that rely on synthesized queries and simpli…

Cited by 5SourcePDFScholar
2024

Stylespeech: Self-Supervised Style Enhancing with VQ-VAE-Based Pre-Training for Expressive Audiobook Speech Synthesis

ICASSP 2024accepted

The expressive quality of synthesized speech for audiobooks is limited by generalized model architecture and unbalanced style distribution in the training data. To address these issues, in this paper, we propose a self-supervised style enhancing method with VQ-VAE-based pre-training for expressive a…

Cited by 0SourceScholar
2024

Summarize the Past to Predict the Future: Natural Language Descriptions of Context Boost Multimodal Object Interaction Anticipation

CVPR 2024poster

We study object interaction anticipation in egocentric videos. This task requires an understanding of the spatio-temporal context formed by past actions on objects coined "action context". We propose TransFusion a multimodal transformer-based architecture for short-term object interaction anticipati…

Cited by 5SourcePDFScholar
2024

Transparent and Scrutable Recommendations Using Natural Language User Profiles

ACL 2024long

Recent state-of-the-art recommender systems predominantly rely on either implicit or explicit feedback from users to suggest new items. While effective in recommending novel options, many recommender systems often use uninterpretable embeddings to represent user preferences. This lack of transparenc…

2024

WANDR: Intention-guided Human Motion Generation

CVPR 2024poster

Synthesizing natural human motions that enable a 3D human avatar to walk and reach for arbitrary goals in 3D space remains an unsolved problem with many applications. Existing methods (data-driven or using reinforcement learning) are limited in terms of generalization and motion naturalness. A prima…

Cited by 12SourcePDFScholar
2024

What Do You See in Vehicle? Comprehensive Vision Solution for In-Vehicle Gaze Estimation

CVPR 2024poster

Driver's eye gaze holds a wealth of cognitive and intentional cues crucial for intelligent vehicles. Despite its significance research on in-vehicle gaze estimation remains limited due to the scarcity of comprehensive and well-annotated datasets in real driving scenarios. In this paper we present th…

Cited by 10SourcePDFScholar
2023

A Survey on Asking Clarification Questions Datasets in Conversational Systems

ACL 2023long

The ability to understand a user’s underlying needs is critical for conversational systems, especially with limited input from users in a conversation. Thus, in such a domain, Asking Clarification Questions (ACQs) to reveal users’ true intent from their queries or utterances arise as an essential ta…

2023

Convolutional Persistence as a Remedy to Neural Model Analysis

AISTATS 2023poster

While deep neural networks are proven to be effective learning systems, their analysis is complex due to the high-dimensionality of their weight space. Persistent topological properties can be used as an additional descriptor, providing insights on how the network weights evolve during training. In…

Cited by 2SourcePDFScholar
2023

GazeNeRF: 3D-Aware Gaze Redirection With Neural Radiance Fields

CVPR 2023poster

We propose GazeNeRF, a 3D-aware method for the task of gaze redirection. Existing gaze redirection methods operate on 2D images and struggle to generate 3D consistent results. Instead, we build on the intuition that the face region and eye balls are separate 3D structures that move in a coordinated…

2023

Improving Conversational Recommendation Systems via Bias Analysis and Language-Model-Enhanced Data Augmentation

EMNLP 2023long findings

Conversational Recommendation System (CRS) is a rapidly growing research area that has gained significant attention alongside advancements in language modelling techniques. However, the current state of conversational recommendation faces numerous challenges due to its relative novelty and limited e…

Cited by 0SourcecodeScholar
2023

JAWS: Just a Wild Shot for Cinematic Transfer in Neural Radiance Fields

CVPR 2023poster

This paper presents JAWS, an optimzation-driven approach that achieves the robust transfer of visual cinematic features from a reference in-the-wild video clip to a newly generated clip. To this end, we rely on an implicit-neural-representation (INR) in a way to compute a clip that shares the same c…

2023

OPT: One-shot Pose-Controllable Talking Head Generation

ICASSP 2023accepted

One-shot talking head generation produces lip-sync talking heads based on arbitrary audio and one source face. To guarantee the naturalness and realness, recent methods propose to achieve free pose control instead of simply editing mouth areas. However, existing methods do not preserve accurate iden…

Cited by 0SourceScholar
2023

Particle-based Variational Inference with Preconditioned Functional Gradient Flow

ICLR 2023poster

Particle-based variational inference (VI) minimizes the KL divergence between model samples and the target posterior with gradient flow estimates. With the popularity of Stein variational gradient descent (SVGD), the focus of particle-based VI algorithms has been on the properties of functions in Re…

Cited by 22SourcePDFScholar
2023

SpatialFormer: Semantic and Target Aware Attentions for Few-Shot Learning

AAAI 2023technical

Recent Few-Shot Learning (FSL) methods put emphasis on generating a discriminative embedding features to precisely measure the similarity between support and query sets. Current CNN-based cross-attention approaches generate discriminative representations via enhancing the mutually semantic similar r…

2022

Decoupling Classifier for Boosting Few-shot Object Detection and Instance Segmentation

NeurIPS 2022accept

This paper focus on few-shot object detection~(FSOD) and instance segmentation~(FSIS), which requires a model to quickly adapt to novel classes with a few labeled instances. The existing methods severely suffer from bias classification because of the missing label issue which naturally exists in an…

2022

Distributed Online Convex Optimization with Compressed Communication

NeurIPS 2022accept

We consider a distributed online convex optimization problem when streaming data are distributed among computing agents over a connected communication network. Since the data are high-dimensional or the network is large-scale, communication load can be a bottleneck for the efficiency of distributed…

Cited by 11SourcePDFScholar
2022

Improving Fastspeech TTS with Efficient Self-Attention and Compact Feed-Forward Network

ICASSP 2022accepted

FastSpeech, as a feed-forward transformer based TTS, can avoid the slow serial, autoregressive inference to generate the target mel-spectrogram in a parallel way. As a non-autoregressive TTS, the latency and computation load in inference is shifted from vocoder to transformer where the efficiency is…

Cited by 0SourceScholar
2022

JPEG Artifacts Removal via Contrastive Representation Learning

ECCV 2022poster

"To meet the needs of practical applications, current deep learning-based methods focus on using a single model to handle JPEG images with different compression qualities, while few of them consider the auxiliary effects of the compression quality information. Recently, several methods estimate qual…

2022

MetaASSIST: Robust Dialogue State Tracking with Meta Learning

EMNLP 2022main

Existing dialogue datasets contain lots of noise in their state annotations. Such noise can hurt model training and ultimately lead to poor generalization performance. A general framework named ASSIST has recently been proposed to train robust dialogue state tracking (DST) models. It introduces an a…

2022

Prosodyspeech: Towards Advanced Prosody Model for Neural Text-to-Speech

ICASSP 2022accepted

This paper proposes ProsodySpeech, a novel prosody model to enhance encoder-decoder neural Text-To-Speech (TTS), to generate high expressive and personalized speech even with very limited training data. First, we use a Prosody Extractor built from a large speech corpus with various speakers to gener…

Cited by 0SourceScholar
2022

SMASH: Improving SMAll Language Models’ Few-SHot Ability with Prompt-Based Distillation

EMNLP 2022finding

Large-scale language models coupled with prompts have shown remarkable performance on few-shot learning. However, through systematic experiments, we find that the few-shot performance of small language models is poor, and using prompts on them brings fewer improvements than on larger ones. In this p…

2022

Sparse Local Patch Transformer for Robust Face Alignment and Landmarks Inherent Relation Learning

CVPR 2022poster

Heatmap regression methods have dominated face alignment area in recent years while they ignore the inherent relation between different landmarks. In this paper, we propose a Sparse Local Patch Transformer (SLPT) for learning the inherent relation. The SLPT generates the representation of each singl…

Cited by 62PDFcodeScholar
2021

Improving Embedding-based Large-scale Retrieval via Label Enhancement

EMNLP 2021finding

Current embedding-based large-scale retrieval models are trained with 0-1 hard label that indicates whether a query is relevant to a document, ignoring rich information of the relevance degree. This paper proposes to improve embedding-based retrieval from the perspective of better characterizing the…

Cited by 6SourcePDFScholar
2021

QuadrupletBERT: An Efficient Model For Embedding-Based Large-Scale Retrieval

NAACL 2021long

The embedding-based large-scale query-document retrieval problem is a hot topic in the information retrieval (IR) field. Considering that pre-trained language models like BERT have achieved great success in a wide variety of NLP tasks, we present a QuadrupletBERT model for effective and efficient re…

Cited by 10SourcePDFScholar
2021

Self-Supervised 3D Hand Pose Estimation From Monocular RGB via Contrastive Learning

ICCV 2021poster

Encouraged by the success of contrastive learning on image classification tasks, we propose a new self-supervised method for the structured regression task of 3D hand pose estimation. Contrastive learning makes use of unlabeled data for the purpose of representation learning via a loss formulation t…

Cited by 85PDFScholar
2020

Relative Pose Estimation and Planar Reconstruction via Superpixel-Driven Multiple Homographies

IROS 2020poster

This paper proposes a novel method to simultaneously perform relative camera pose estimation and planar reconstruction of a scene from two RGB images. We start by extracting and matching superpixel information from both images and rely on a novel multi-model RANSAC approach to estimate multiple homo…

Cited by 8SourceScholar
2020

Relaxed Multivariate Bernoulli Distribution and Its Applications to Deep Generative Models

UAI 2020poster

Recent advances in variational auto-encoder (VAE) have demonstrated the possibility of approximating the intractable posterior distribution with a variational distribution parameterized by a neural network. To optimize the variational objective of VAE, the reparameterization trick is commonly applie…

2018

Optimized Contrast Enhancements to Improve Robustness of Visual Tracking in a SLAM Relocalisation Context

IROS 2018poster

Robustness of indirect SLAM techniques to light changing conditions remains a central issue in the robotics community. With the change in the illumination of a scene, feature points are either not extracted properly due to low contrasts, or not matched due to large differences in descriptors. In thi…

Cited by 3SourceScholar