← Search

Hao Yang

131 accepted papers

2026

BridgeDrive: Diffusion Bridge Policy for Closed-Loop Trajectory Planning in Autonomous Driving

ICLR 2026poster

Diffusion-based planners have shown great promise for autonomous driving due to their ability to capture multi-modal driving behaviors. However, guiding these models effectively in reactive, closed-loop environments remains a significant challenge. Simple conditioning often fails to provide sufficie…

Cited by 0SourcecodeScholar
2026

CARE-Bench: A Benchmark of Diverse Client Simulations Guided by Expert Principles for Evaluating LLMs in Psychological Counseling

AAAI 2026technical

The mismatch between the growing demand for psychological counseling and the limited availability of services has motivated research into the application of Large Language Models (LLMs) in this domain. Consequently, there is a need for a robust and unified benchmark to assess the counseling competen

Cited by 0SourcePDFScholar
2026

DAPointMamba: Domain Adaptive Point Mamba for Point Cloud Completion

AAAI 2026technical

Domain adaptive point cloud completion (DA PCC) aims to narrow the geometric and semantic discrepancies between the labeled source and unlabeled target domains. Existing methods either suffer from limited receptive fields or quadratic complexity due to using CNNs or vision Transformers. In this pape

Cited by 0SourcePDFScholar
2026

DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion Transformers

CVPR 2026

Diffusion models have recently motivated great success in many generation tasks like object removal. Nevertheless, existing image decomposition methods struggle to disentangle semi-transparent or transparent layer occlusions due to mask prior dependencies, static object assumptions, and the lack of

Cited by 0SourcecodeScholar
2026

Lumosaic: Hyperspectral Video via Active Illumination and Coded-Exposure Pixels

CVPR 2026

We present Lumosaic, a compact active hyperspectral video system designed for real-time capture of dynamic scenes. Our approach combines a programmable narrowband LED array with a coded-exposure-pixel (CEP) camera capable of high-speed pixel-wise exposure, thereby enabling joint encoding of scene in

Cited by 0SourceScholar
2026

MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX

AAAI 2026technical

We introduce MAVERIX (Multimodal Audio-Visual Evaluation and Recognition IndeX), a unified benchmark to probe video understanding in multimodal LLMs, encompassing video, audio, and text inputs with human performance baselines. Although recent advancements in audiovisual models have shown substantial

Cited by 0SourcePDFScholar
2026

OmniShow: Orchestrating Multimodal Conditions for Human-Object Interaction Video Generation

ICML 2026poster

In this work, we study **Human-Object Interaction Video Generation (HOIVG)**, which aims to synthesize high-quality HOI videos via text, reference image, audio, and pose conditions. To address the challenges of harmonious multimodal injection and heterogeneous data utility, we present **OmniShow**, …

Cited by 0SourceScholar
2026

PointDGRWKV: Generalizing RWKV-like Architecture to Unseen Domains for Point Cloud Classification

AAAI 2026technical

Domain Generalization (DG) has been recently explored to enhance the generalizability of Point Cloud Classification (PCC) models toward unseen domains. Prior works are based on convolutional networks, Transformer or Mamba architectures, either suffering from limited receptive fields or high computat

Cited by 0SourcePDFScholar
2026

PromptStereo: Zero-Shot Stereo Matching via Structure and Motion Prompts

CVPR 2026

Modern stereo matching methods have leveraged monocular depth foundation models to achieve superior zero-shot generalization performance. However, most existing methods primarily focus on extracting robust features for cost volume construction or disparity initialization. At the same time, the itera

Cited by 0SourcecodeScholar
2026

Veda: Scalable Video Diffusion via Distilled Sparse Attention

ICML 2026poster

Scaling Diffusion Transformers to generate high-resolution, long videos is constrained by the quadratic cost of self-attention, and existing sparse attention methods degrade under high sparsity. We show empirically that generation quality is determined not by the sparsity ratio itself, but by how we…

Cited by 0SourceScholar
2025

"I've Heard of You!": Generate Spoken Named Entity Recognition Data for Unseen Entities

ICASSP 2025accepted

Spoken named entity recognition (NER) aims to identify named entities from speech, playing an important role in speech processing. New named entities appear every day, however, annotating their Spoken NER data is costly. In this paper, we demonstrate that existing Spoken NER systems perform poorly w…

Cited by 0SourceScholar
2025

A novel multimodal personality prediction method based on pretrained models and graph relational transformer network

ICASSP 2025accepted

Multimodal personality analysis aims to identify and express human personality traits in videos. However, RNN and its variants have a limited ability to learn long-term temporal dependencies and existing methods neglect bimodal association features. Based on the fact that visual modalities play a do…

Cited by 0SourceScholar
2025

Alleviating Distribution Shift in Synthetic Data for Machine Translation Quality Estimation

ACL 2025long

Quality Estimation (QE) models evaluate the quality of machine translations without reference translations, serving as the reward models for the translation task.Due to the data scarcity, synthetic data generation has emerged as a promising solution.However, synthetic QE data often suffers from dist…

2025

An Evaluation Resource for Grounding Translation Errors

EMNLP 2025

Current fine-grained error analyses by LLMs gain more and more attention in machine translation, but these analyses do not ground the errors to the reasons why the annotated text spans are erroneous. If LLMs do not know such reasons, the corrections or refinements by LLMs will be untrustworthy.In th

2025

Audio Is the Achilles’ Heel: Red Teaming Audio Large Multimodal Models

NAACL 2025long

Large Multimodal Models (LMMs) have demonstrated the ability to interact with humans under real-world conditions by combining Large Language Models (LLMs) and modality encoders to align multimodal information (visual and auditory) with text. However, such models raise new safety challenges of whethe…

2025

Combining the Best of Both Worlds: A Method for Hybrid NMT and LLM Translation

ACL 2025finding

Large language model (LLM) shows promising performances in a variety of downstream tasks, such as machine translation (MT). However, using LLMs for translation suffers from high computational costs and significant latency. Based on our evaluation, in most cases, translations using LLMs are comparabl…

2025

DoCIA: An Online Document-Level Context Incorporation Agent for Speech Translation

ACL 2025finding

Document-level context is crucial for handling discourse challenges in text-to-text document-level machine translation (MT). Despite the increased discourse challenges introduced by noise from automatic speech recognition (ASR), the integration of document-level context in speech translation (ST) re…

2025

End-to-End Learnable Psychiatric Scale Guided Risky Post Screening for Depression Detection on Social Media

EMNLP 2025

Detecting depression through users’ social media posting history is crucial for enabling timely intervention; however, irrelevant content within these posts negatively impacts detection performance. Thus, it is crucial to extract pertinent content from users’ complex posting history. Current methods

Cited by 0SourcePDFScholar
2025

Enhancing Large Language Models for Document-Level Translation Post-Editing Using Monolingual Data

COLING 2025main

The translation capabilities of neural machine translation (NMT) models based on the encoder-decoder framework are extremely potent. Although Large Language Models (LLMs) have achieved remarkable results in many tasks, they have not reached state-of-the-art performance in NMT. However, traditional N…

2025

Enhancing Numerical Prediction of MLLMs with Soft Labeling

ICCV 2025poster

The optimality of using the de facto cross-entropy loss with one-hot target distribution (hard labeling) is questioned when training (Multimodal) Large Language Models (LLMs/MLLMs). Although it is reasonable for language token prediction, which is a typical multi-class classification problem in disc…

Cited by 0SourcePDFScholar
2025

Enhancing Speech Large Language Models with Prompt-Aware Mixture of Audio Encoders

EMNLP 2025

Connecting audio encoders with large language models (LLMs) allows the LLM to perform various audio understanding tasks, such as automatic speech recognition (ASR) and audio captioning (AC). Most research focuses on training an adapter layer to generate a unified audio feature for the LLM. However,

2025

From Observation to Understanding: Front-Door Adjustments with Uncertainty Calibration for Enhancing Egocentric Reasoning in LVLMs

ACL 2025finding

Recent progress in large vision-language models (LVLMs) has shown substantial potential across a broad spectrum of third-person tasks. However, adapting these LVLMs to egocentric scenarios remains challenging due to their third-person training bias. Existing methods that adapt LVLMs for first-person…

2025

Function-to-Style Guidance of LLMs for Code Translation

ICML 2025poster

Large language models (LLMs) have made significant strides in code translation tasks. However, ensuring both the correctness and readability of translated code remains a challenge, limiting their effective adoption in real-world software development. In this work, we propose F2STrans, a function-to…

Cited by 0SourcePDFScholar
2025

Generative Annotation for ASR Named Entity Correction

EMNLP 2025

End-to-end automatic speech recognition systems often fail to transcribe domain-speciffcnamed entities, causing catastrophic failuresin downstream tasks. Numerous fast and lightweight named entity correction (NEC) models have been proposed in recent years. These models, mainly leveraging phonetic-le

2025

Geometric Logit Decoupling for Energy-Based Graph Out-of-distribution Detection

NeurIPS 2025poster

GNNs have achieved remarkable performance across a range of tasks, but their reliability under distribution shifts remains a significant challenge. In particular, energy-based OOD detection methods—which compute energy scores from GNN logits—suffer from unstable performance due to a fundamental coup…

Cited by 0SourceScholar
2025

Goku: Flow Based Video Generative Foundation Models

CVPR 2025highlight

This paper introduces Goku, a state-of-the-art family of joint image-and-video generation models leveraging rectified flow Transformers to achieve industry-leading performance. We detail the foundational elements enabling high-quality visual generation, including the data curation pipeline, model ar…

Cited by 15SourcePDFScholar
2025

Imagination and Contemplation: A Balanced Framework for Semantic-Augmented Multimodal Machine Translation

EMNLP 2025

Multimodal Machine Translation (MMT) enhances textual translation through auxiliary inputs such as images, which is particularly effective in resolving linguistic ambiguities. However, visual information often introduces redundancy or noise, potentially impairing translation quality. To address this

2025

Invariant Deep Uplift Modeling for Incentive Assignment in Online Marketing via Probability of Necessity and Sufficiency

ICML 2025spotlight

In online platforms, incentives (\textit{e.g}., discounts, coupons) are used to boost user engagement and revenue. Uplift modeling methods are developed to estimate user responses from observational data, often incorporating distribution balancing to address selection bias. However, these methods ar…

Cited by 0SourcePDFScholar
2025

Investigating Numerical Translation with Large Language Models

ICASSP 2025accepted

The inaccurate translation of numbers can lead to significant security issues, ranging from financial setbacks to medical inaccuracies. While large language models (LLMs) have made significant advancements in machine translation, their capacity for translating numbers has not been thoroughly explore…

Cited by 0SourceScholar
2025

Large Language Model Should Understand Pinyin for Chinese ASR Error Correction

ICASSP 2025accepted

Large language models (LLMs) can enhance automatic speech recognition (ASR) systems through generative error correction (GEC). In this paper, we propose Pinyin-enhanced GEC (PY-GEC), which leverages Pinyin—the phonetic representation of Mandarin Chinese—as supplementary information to improve Chines…

Cited by 0SourceScholar
2025

Look Beyond Feeling: Unveiling Latent Needs from Implicit Expressions for Proactive Emotional Support

EMNLP 2025

In recent years, Large Language Models (LLMs) have made significant progress in emotional support dialogue. However, there are two major challenges for LLM-based support systems. First, users may be hesitant to fully disclose their emotions at the outset. Second, direct probing or excessive question

Cited by 0SourcePDFScholar
2025

M-Ped: Multi-Prompt Ensemble Decoding for Large Language Models

EMNLP 2025

With the widespread application of Large Language Models (LLMs) in the field of Natural Language Processing (NLP), enhancing their performance has become a research hotspot. This paper presents a novel multi-prompt ensemble decoding approach designed to bolster the generation quality of LLMs by leve

2025

Mamba-YOLO-World: Marrying YOLO-World with Mamba for Open-Vocabulary Detection

ICASSP 2025accepted

Open-vocabulary detection (OVD) aims to detect objects beyond a predefined set of categories. As a pioneering model incorporating the YOLO series into OVD, YOLO-World is well-suited for scenarios prioritizing speed and efficiency. However, its performance is hindered by its neck feature fusion mecha…

Cited by 0SourceScholar
2025

Met2Net: A Decoupled Two-Stage Spatio-Temporal Forecasting Model for Complex Meteorological Systems

ICCV 2025poster

The increasing frequency of extreme weather events due to global climate change urges accurate weather prediction. Recently, great advances are made by the end-to-end methods, thanks to deep learning techniques, but they face limitations of representation inconsistency in multivariable integration a…

2025

Multimodal Machine Translation with Text-Image In-depth Questioning

ACL 2025finding

Multimodal machine translation (MMT) integrates visual information to address ambiguity and contextual limitations in neural machine translation (NMT). Some empirical studies have revealed that many MMT models underutilize visual data during translation. They attempt to enhance cross-modal interacti…

2025

Optimizing Speech Multi-View Feature Fusion through Conditional Computation

ICASSP 2025accepted

Recent advancements have highlighted the efficacy of self-supervised learning (SSL) features in various speech-related tasks, providing lightweight and versatile multi-view speech representations. However, our study reveals that while SSL features expedite model convergence, they conflict with tradi…

Cited by 0SourceScholar
2025

PointDGMamba: Domain Generalization of Point Cloud Classification via Generalized State Space Model

AAAI 2025technical

Domain Generalization (DG) has been recently explored to improve the generalizability of point cloud classification (PCC) models toward unseen domains. However, they often suffer from limited receptive fields or quadratic complexity due to the use of convolution neural networks or vision Transformer…

2025

Reshaping Representation Space to Balance the Safety and Over-rejection in Large Audio Language Models

EMNLP 2025

Large Audio Language Models (LALMs) have extended the capabilities of Large Language Models (LLMs) by enabling audio-based human interactions. However, recent research has revealed that LALMs remain vulnerable to harmful queries due to insufficient safety-alignment. Despite advances in defence measu

Cited by 0SourcePDFScholar
2025

SRDC: Semantics-based Ransomware Detection and Classification with LLM-assisted Pre-training

AAAI 2025technical

In recent years, ransomware has emerged as a formidable data security threat, causing significant data privacy breaches that inflict substantial financial, reputational, and operational damages on society. Many studies employ dynamic feature analysis for ransomware detection. However, these methods…

2025

Scaling up Image Segmentation across Data and Tasks

CVPR 2025poster

Traditional segmentation models, while effective in isolated tasks, often fail to generalize to more complex and open-ended segmentation problems, such as free-form, open-vocabulary, and in-the-wild scenarios. To bridge this gap, we propose to scale up image segmentation across diverse datasets and…

Cited by 0SourcePDFScholar
2025

Stephanie: Step-by-Step Dialogues for Mimicking Human Interactions in Social Conversations

NAACL 2025findings

In the rapidly evolving field of natural language processing, dialogue systems primarily employ a single-step dialogue paradigm. Although this paradigm is commonly adopted, it lacks the depth and fluidity of human interactions and does not appear natural. We introduce a novel **Step**-by-Step Dialog…

Cited by 2SourcePDFScholar
2025

TableDreamer: Progressive and Weakness-guided Data Synthesis from Scratch for Table Instruction Tuning

ACL 2025finding

Despite the commendable progress of recent LLM-based data synthesis methods, they face two limitations in generating table instruction tuning data. First, they can not thoroughly explore the vast input space of table understanding tasks, leading to limited data diversity. Second, they ignore the wea…

2025

Taming Text-to-Image Synthesis for Novices: User-centric Prompt Generation via Multi-turn Guidance

EMNLP 2025

The emergence of text-to-image synthesis (TIS) models has significantly influenced digital image creation by producing high-quality visuals from written descriptions. Yet these models are sensitive on textual prompts, posing a challenge for novice users who may not be familiar with TIS prompt writin

2025

Test-Time Adaptation on Noisy Data via Model-Pruning-Based Filtering and Flatness-Aware Entropy Minimization

AAAI 2025technical

Test-time adaptation (TTA) deals with domain shifts during inference by training models based on only unlabeled test samples. Test samples may include noisy samples, which degrade domain adaptation. Existing methods rely on the model's output prediction to detect and filter noisy samples, and furthe…

2025

Two Intermediate Translations Are Better Than One: Fine-tuning LLMs for Document-level Translation Refinement

ACL 2025long

Recent research has shown that large language models (LLMs) can enhance translation quality through self-refinement. In this paper, we build on this idea by extending the refinement from sentence-level to document-level translation, specifically focusing on document-to-document (Doc2Doc) translation…

2025

UniCombine: Unified Multi-Conditional Combination with Diffusion Transformer

ICCV 2025poster

With the rapid development of diffusion models in image generation, the demand for more powerful and flexible controllable frameworks is increasing. Although existing methods can guide generation beyond text prompts, the challenge of effectively combining multiple conditional inputs while maintainin…

2025

VQA-Augmented Machine Translation with Cross-Modal Contrastive Learning

EMNLP 2025

Multimodal machine translation (MMT) aims to enhance translation quality by integrating visual information. However, existing methods often extract visual features using pre-trained models while learning text features from scratch, leading to representation imbalance. These methods are also prone to

Cited by 0SourcePDFScholar
2024

A Hybrid Model and Learning-Based Force Estimation Framework for Surgical Robots

IROS 2024poster

Haptic feedback to the surgeon during robotic surgery would enable safer and more immersive surgeries but estimating tissue interaction forces at the tips of robotically controlled surgical instruments has proven challenging. Few existing surgical robots can measure interaction forces directly and t…

Cited by 3SourcecodeScholar
2024

A Novel Paradigm Boosting Translation Capabilities of Large Language Models

NAACL 2024findings

This paper presents a study on strategies to enhance the translation capabilities of large language models (LLMs) in the context of machine translation (MT) tasks. The paper proposes a novel paradigm consisting of three stages: Secondary Pre-training using Extensive Monolingual Data, Continual Pre-t…

Cited by 17SourcePDFScholar
2024

An Empirical Study and Analysis of Text-to-Image Generation Using Large Language Model-Powered Textual Representation

ECCV 2024poster

"One critical prerequisite for faithful text-to-image generation is the accurate understanding of text inputs. Existing methods leverage the text encoder of the CLIP model to represent input prompts. However, the pre-trained CLIP model can merely encode English with a maximum token length of 77. Mor…

2024

AvatarVerse: High-Quality & Stable 3D Avatar Creation from Text and Pose

AAAI 2024technical

Creating expressive, diverse and high-quality 3D avatars from highly customized text descriptions and pose guidance is a challenging task, due to the intricacy of modeling and texturing in 3D that ensure details and various styles (realistic, fictional, etc). We present AvatarVerse, a stable pipelin…

2024

CB-Whisper: Contextual Biasing Whisper Using Open-Vocabulary Keyword-Spotting

COLING 2024main

End-to-end automatic speech recognition (ASR) systems often struggle to recognize rare name entities, such as personal names, organizations and terminologies that are not frequently encountered in the training data. This paper presents Contextual Biasing Whisper (CB-Whisper), a novel ASR system base…

2024

CHisIEC: An Information Extraction Corpus for Ancient Chinese History

COLING 2024main

Natural Language Processing (NLP) plays a pivotal role in the realm of Digital Humanities (DH) and serves as the cornerstone for advancing the structural analysis of historical and cultural heritage texts. This is particularly true for the domains of named entity recognition (NER) and relation extra…

2024

Clustering and Ranking: Diversity-preserved Instruction Selection through Expert-aligned Quality Estimation

EMNLP 2024main

With contributions from the open-source community, a vast amount of instruction tuning (IT) data has emerged. Given the significant resource allocation required by training and evaluating models, it is advantageous to have an efficient method for selecting high-quality IT data. However, existing met…

2024

Cross-Domain Audio Deepfake Detection: Dataset and Analysis

EMNLP 2024main

Audio deepfake detection (ADD) is essential for preventing the misuse of synthetic voices that may infringe on personal rights and privacy. Recent zero-shot text-to-speech (TTS) models pose higher risks as they can clone voices with a single utterance. However, the existing ADD datasets are outdated…

2024

Enhancing Hyperbolic Knowledge Graph Embeddings via Lorentz Transformations

ACL 2024findings

Knowledge Graph Embedding (KGE) is a powerful technique for predicting missing links in Knowledge Graphs (KGs) by learning the entities and relations. Hyperbolic space has emerged as a promising embedding space for KGs due to its ability to represent hierarchical data. Nevertheless, most existing hy…

2024

Evaluation Dataset for Lexical Translation Consistency in Chinese-to-English Document-level Translation

COLING 2024main

Lexical translation consistency is one of the most common discourse phenomena in Chinese-to-English document-level translation. To better evaluate the performance of lexical translation consistency, previous researches assumes that all repeated source words should be translated consistently. However…

Cited by 2SourcePDFScholar
2024

Event-based Few-shot Fine-grained Human Action Recognition

IROS 2024poster

Few-shot fine-grained human (FGH) action recognition is crucial in the context of human-robot interaction within open-set real-world environments. Existing works mainly focus on features extracted from RGB frames. However, their performances are drastically impacted in challenging scenarios, such as…

Cited by 1SourceScholar
2024

LDP: Language-driven Dual-Pixel Image Defocus Deblurring Network

CVPR 2024poster

Recovering sharp images from dual-pixel (DP) pairs with disparity-dependent blur is a challenging task. Existing blur map-based deblurring methods have demonstrated promising results. In this paper we propose to the best of our knowledge the first framework to introduce the contrastive language-imag…

Cited by 12SourcePDFScholar
2024

Moderate Message Passing Improves Calibration: A Universal Way to Mitigate Confidence Bias in Graph Neural Networks

AAAI 2024technical

Confidence calibration in Graph Neural Networks (GNNs) aims to align a model's predicted confidence with its actual accuracy. Recent studies have indicated that GNNs exhibit an under-confidence bias, which contrasts the over-confidence bias commonly observed in deep neural networks. However, our dee…

Cited by 2SourcePDFScholar
2024

Pmmwdeconv: Unsupervised Data-Consistent Blind Passive Millimeterwave Image Deconvolution with Global Context Priors

ICASSP 2024accepted

Passive millimeter-wave (PMMW) imaging is extensively employed in public security industries due to its privacy-safe and non-hazardous nature. Nevertheless, the quality of PMMW images is typically poor given blur and noise. Although the end-to-end learning-based image deconvolution methods have demo…

Cited by 0SourceScholar
2024

Submodular-based In-context Example Selection for LLMs-based Machine Translation

COLING 2024main

Large Language Models (LLMs) have demonstrated impressive performances across various NLP tasks with just a few prompts via in-context learning. Previous studies have emphasized the pivotal role of well-chosen examples in in-context learning, as opposed to randomly selected instances that exhibits u…

2024

THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models

CVPR 2024poster

Mitigating hallucinations in large vision-language models (LVLMs) remains an open problem. Recent benchmarks do not address hallucinations in open-ended free-form responses which we term "Type I hallucinations". Instead they focus on hallucinations responding to very specific question formats---typi…

Cited by 16SourcePDFScholar
2024

Towards Probing Speech-Specific Risks in Large Multimodal Models: A Taxonomy, Benchmark, and Insights

EMNLP 2024main

Large Multimodal Models (LMMs) have achieved great success recently, demonstrating a strong capability to understand multimodal information and to interact with human users. Despite the progress made, the challenge of detecting high-risk interactions in multimodal settings, and in particular in spee…

2024

Translate Meanings, Not Just Words: IdiomKB’s Role in Optimizing Idiomatic Translation with Language Models

AAAI 2024technical

To translate well, machine translation (MT) systems and general-purposed language models (LMs) need a deep understanding of both source and target languages and cultures. Therefore, idioms, with their non-compositional nature, pose particular challenges for Transformer-based systems, as literal tran…

2023

A Meta-Learning Approach to Predicting Performance and Data Requirements

CVPR 2023poster

We propose an approach to estimate the number of samples required for a model to reach a target performance. We find that the power law, the de facto principle to estimate model performance, leads to large error when using a small dataset (e.g., 5 samples per class) for extrapolation. This is becaus…

2023

Chain-of-Thought Reasoning in Tabular Language Models

EMNLP 2023long findings

Tabular mathematical reasoning task requires models to perform multi-step operations including information look-up and numerical calculation, based on heterogeneous data from tables and questions. Existing solutions tend to extend chain-of-thought (CoT) reasoning into powerful large language models…

Cited by 0SourceScholar
2023

ContraNeRF: Generalizable Neural Radiance Fields for Synthetic-to-Real Novel View Synthesis via Contrastive Learning

CVPR 2023poster

Although many recent works have investigated generalizable NeRF-based novel view synthesis for unseen scenes, they seldom consider the synthetic-to-real generalization, which is desired in many practical applications. In this work, we first investigate the effects of synthetic data in synthetic-to-r…

2023

Denoising Pre-training for Machine Translation Quality Estimation with Curriculum Learning

AAAI 2023technical

Quality estimation (QE) aims to assess the quality of machine translations when reference translations are unavailable. QE plays a crucial role in many real-world applications of machine translation. Because labeled QE data are usually limited in scale, recent research, such as DirectQE, pre-trains…

2023

FreeEnricher: Enriching Face Landmarks without Additional Cost

AAAI 2023technical

Recent years have witnessed significant growth of face alignment. Though dense facial landmark is highly demanded in various scenarios, e.g., cosmetic medicine and facial beautification, most works only consider sparse face alignment. To address this problem, we present a framework that can enrich l…

Cited by 3SourcePDFScholar
2023

From Trainable Negative Depth to Edge Heterophily in Graphs

NeurIPS 2023poster

Finding the proper depth $d$ of a graph convolutional network (GCN) that provides strong representation ability has drawn significant attention, yet nonetheless largely remains an open problem for the graph learning community. Although noteworthy progress has been made, the depth or the number of…

Cited by 27SourcePDFScholar
2023

Guided Recommendation for Model Fine-Tuning

CVPR 2023poster

Model selection is essential for reducing the search cost of the best pre-trained model over a large-scale model zoo for a downstream task. After analyzing recent hand-designed model selection criteria with 400+ ImageNet pre-trained models and 40 downstream tasks, we find that they can fail due to i…

2023

INarIG: Iterative Non-autoregressive Instruct Generation Model For Word-Level Auto Completion

EMNLP 2023long findings

Computer-aided translation (CAT) aims to enhance human translation efficiency and is still important in scenarios where machine translation cannot meet quality requirements. One fundamental task within this field is Word-Level Auto Completion (WLAC). WLAC predicts a target word given a source senten…

Cited by 0SourceScholar
2023

Improved Pseudo Data for Machine Translation Quality Estimation with Constrained Beam Search

EMNLP 2023long main

Machine translation (MT) quality estimation (QE) is a crucial task to estimate the quality of MT outputs when reference translations are unavailable. Many studies focus on generating pseudo data using large parallel corpus and achieve remarkable success in the supervised setting. However, pseudo dat…

Cited by 0SourcecodeScholar
2023

InterFormer: Real-time Interactive Image Segmentation

ICCV 2023poster

Interactive image segmentation enables annotators to efficiently perform pixel-level annotation for segmentation tasks. However, the existing interactive segmentation pipeline suffers from inefficient computations of interactive models because of the following two issues. First, annotators' later cl…

Cited by 25PDFcodeScholar
2023

Lexical Translation Inconsistency-Aware Document-Level Translation Repair

ACL 2023findings

Following the idea of “one translation per discourse”, in this paper we aim to improve translation consistency via document-level translation repair (DocRepair), i.e., automatic post-editing on translations of documents. To this end, we propose a lexical translation inconsistency-aware DocRepair to…

2023

Local and Global Logit Adjustments for Long-Tailed Learning

ICCV 2023poster

Multi-expert ensemble models for long-tailed learning typically either learn diverse generalists from the whole dataset or aggregate specialists on different subsets. However, the former is insufficient for tail classes due to the high imbalance factor of the entire dataset, while the latter may bri…

Cited by 25PDFScholar
2023

MaskCLIP: Masked Self-Distillation Advances Contrastive Language-Image Pretraining

CVPR 2023poster

This paper presents a simple yet effective framework MaskCLIP, which incorporates a newly proposed masked self-distillation into contrastive language-image pretraining. The core idea of masked self-distillation is to distill representation from a full image to the representation predicted from a mas…

2023

PRED: Pre-training via Semantic Rendering on LiDAR Point Clouds

NeurIPS 2023poster

Pre-training is crucial in 3D-related fields such as autonomous driving where point cloud annotation is costly and challenging. Many recent studies on point cloud pre-training, however, have overlooked the issue of incompleteness, where only a fraction of the points are captured by LiDAR, leading to…

2023

Probabilistic Masked Attention Networks for Explainable Sequential Recommendation

IJCAI 2023poster

Transformer-based models are powerful for modeling temporal dynamics of user preference in sequential recommendation. Most of the variants adopt the Softmax transformation in the self-attention layers to generate dense attention probabilities. However, real-world item sequences are often noisy, cont…

Cited by 11SourcePDFScholar
2023

Prompt Tuning for Unified Multimodal Pretrained Models

ACL 2023findings

Prompt tuning has become a new paradigm for model tuning and it has demonstrated success in natural language pretraining and even vision pretraining. The parameter-efficient prompt tuning methods that optimize soft embeddings while keeping the pretrained model frozen demonstrate advantages in low co…

2023

SmartSpanNER: Making SpanNER Robust in Low Resource Scenarios

EMNLP 2023long findings

Named Entity Recognition (NER) is one of the most fundamental tasks in natural language processing. Span-level prediction (SpanNER) is more naturally suitable for nested NER than sequence labeling (SeqLab). However, according to our experiments, the SpanNER method is more sensitive to the amount of…

Cited by 0SourceScholar
2023

Stochastic Feature Averaging for Learning with Long-Tailed Noisy Labels

IJCAI 2023poster

Deep neural networks have shown promising results on a wide variety of tasks using large-scale and well-annotated training datasets. However, data collected from real-world applications can suffer from two prevalent biases, i.e., long-tailed class distribution and label noise. Previous efforts on lo…

2023

SwiftAvatar: Efficient Auto-Creation of Parameterized Stylized Character on Arbitrary Avatar Engines

AAAI 2023technical

The creation of a parameterized stylized character involves careful selection of numerous parameters, also known as the "avatar vectors" that can be interpreted by the avatar engine. Existing unsupervised avatar vector estimation methods that auto-create avatars for users, however, often fail to wor…

2023

Text Style Transfer Back-Translation

ACL 2023long

Back Translation (BT) is widely used in the field of machine translation, as it has been proved effective for enhancing translation quality. However, BT mainly improves the translation of inputs that share a similar style (to be more specific, translation-liked inputs), since the source side of BT d…

2023

UCorrect: An Unsupervised Framework for Automatic Speech Recognition Error Correction

ICASSP 2023accepted

Error correction techniques have been used to refine the output sentences from automatic speech recognition (ASR) models and achieve a lower word error rate (WER). Previous works usually adopt end-to-end models and has strong dependency on Pseudo Paired Data and Original Paired Data. But when only p…

Cited by 0SourceScholar
2023

Your representations are in the network: composable and parallel adaptation for large scale models

NeurIPS 2023poster

We present a framework for transfer learning that efficiently adapts a large base-model by learning lightweight cross-attention modules attached to its intermediate activations. We name our approach InCA (Introspective-Cross-Attention) and show that it can efficiently survey a network’s representati…

Cited by 3SourcePDFScholar
2022

Capture Human Disagreement Distributions by Calibrated Networks for Natural Language Inference

ACL 2022findings

Natural Language Inference (NLI) datasets contain examples with highly ambiguous labels due to its subjectivity. Several recent efforts have been made to acknowledge and embrace the existence of ambiguity, and explore how to capture the human disagreement distribution. In contrast with directly lear…

Cited by 10SourcePDFScholar
2022

Face-Sensitive Image-to-Emotional-Text Cross-modal Translation for Multimodal Aspect-based Sentiment Analysis

EMNLP 2022main

Aspect-level multimodal sentiment analysis, which aims to identify the sentiment of the target aspect from multimodal data, recently has attracted extensive attention in the community of multimedia and natural language processing. Despite the recent success in textual aspect-based sentiment analysis…

Cited by 48SourcePDFScholar
2022

General Facial Representation Learning in a Visual-Linguistic Manner

CVPR 2022oral

How to learn a universal facial representation that boosts all face analysis tasks This paper takes one step toward this goal. In this paper, we study the transfer performance of pre-trained models on face analysis tasks and introduce a framework, called FaRL, for general facial representation learn…

Cited by 199PDFcodeScholar
2022

Large-Scale Pre-Training for Person Re-Identification With Noisy Labels

CVPR 2022poster

This paper aims to address the problem of pre-training for person re-identification (Re-ID) with noisy labels. To setup the pre-training task, we apply a simple online multi-object tracking system on raw videos of an existing unlabeled Re-ID dataset "LUPerson" and build the Noisy Labeled variant cal…

Cited by 81PDFcodeScholar
2022

Modeling Consistency Preference via Lexical Chains for Document-level Neural Machine Translation

EMNLP 2022main

In this paper we aim to relieve the issue of lexical translation inconsistency for document-level neural machine translation (NMT) by modeling consistency preference for lexical chains, which consist of repeated words in a source-side document and provide a representation of the lexical consistency…

2022

Neighbors Are Not Strangers: Improving Non-Autoregressive Translation under Low-Frequency Lexical Constraints

NAACL 2022long

Lexically constrained neural machine translation (NMT) draws much industrial attention for its practical usage in specific domains. However, current autoregressive approaches suffer from high latency. In this paper, we focus on non-autoregressive translation (NAT) for this problem for its efficiency…

2022

Normalization of Language Embeddings for Cross-Lingual Alignment

ICLR 2022poster

Learning a good transfer function to map the word vectors from two languages into a shared cross-lingual word vector space plays a crucial role in cross-lingual NLP. It is useful in translation tasks and important in allowing complex models built on a high-resource language like English to be direct…

2022

Omni-DETR: Omni-Supervised Object Detection With Transformers

CVPR 2022poster

We consider the problem of omni-supervised object detection, which can use unlabeled, fully labeled and weakly labeled annotations, such as image tags, counts, points, etc., for object detection. This is enabled by a unified architecture, Omni-DETR, based on the recent progress on student-teacher fr…

Cited by 63PDFcodeScholar
2022

Real-Time Neural Character Rendering with Pose-Guided Multiplane Images

ECCV 2022poster

"We propose pose-guided multiplane image (MPI) synthesis which can render an animatable character in real scenes with photorealistic quality. We use a portable camera rig to capture the multi-view images along with the driving signal for the moving subject. Our method generalizes the image-to-image…

2022

RedApt: An Adaptor for wav2vec 2 EncodingFaster and Smaller Speech Translation without Quality Compromise

EMNLP 2022finding

Pre-trained speech Transformers in speech translation (ST) have facilitated state-of-the-art (SotA) results; yet, using such encoders is computationally expensive. To improve this, we present a novel Reducer Adaptor block, RedApt, that could be seamlessly integrated within any Transformer-based spee…

2022

Rethinking Few-Shot Object Detection on a Multi-Domain Benchmark

ECCV 2022poster

"Most existing works on few-shot object detection (FSOD) focus on a setting where both pre-training and few-shot learning datasets are from a similar domain. However, few-shot algorithms are important in multiple domains; hence evaluation needs to reflect the broad applications. We propose a Multi-d…

2022

Self-supervised Rewiring of Pre-trained Speech Encoders:Towards Faster Fine-tuning with Less Labels in Speech Processing

EMNLP 2022finding

Pre-trained speech Transformers have facilitated great success across various speech processing tasks. However, fine-tuning these encoders for downstream tasks require sufficiently large training data to converge or to achieve state-of-the-art. In text domain this has been partly attributed to sub-o…

2022

Sentiment Word Aware Multimodal Refinement for Multimodal Sentiment Analysis with ASR Errors

ACL 2022findings

Multimodal sentiment analysis has attracted increasing attention and lots of models have been proposed. However, the performance of the state-of-the-art models decreases sharply when they are deployed in the real world. We find that the main reason is that real-world applications can only access the…

2021

ADNet: Leveraging Error-Bias Towards Normal Direction in Face Alignment

ICCV 2021poster

The recent progress of CNN has dramatically improved face alignment performance. However, few works have paid attention to the error-bias with respect to error distribution of facial landmarks. In this paper, we investigate the error-bias issue in face alignment, where the distributions of landmark…

Cited by 69PDFcodeScholar
2021

Adversarial Example Detection Using Latent Neighborhood Graph

ICCV 2021poster

Detection of adversarial examples with high accuracy is critical for the security of deployed deep neural network-based models. We present the first graph-based adversarial detection method that constructs a Latent Neighborhood Graph (LNG) around an input example to determine if the input example is…

Cited by 74PDFScholar
2021

Beating Attackers At Their Own Games: Adversarial Example Detection Using Adversarial Gradient Directions

AAAI 2021technical

Adversarial examples are input examples that are specifically crafted to deceive machine learning classifiers. State-of-the-art adversarial example detection methods characterize an input example as adversarial either by quantifying the magnitude of feature variations under multiple perturbations or…

Cited by 16SourcePDFScholar
2021

Integrating Subgraph-Aware Relation and Direction Reasoning for Question Answering

ICASSP 2021accepted

Question Answering (QA) models over Knowledge Bases (KBs) are capable of providing more precise answers by utilizing relation information among entities. Although effective, most of these models solely rely on fixed relation representations to obtain answers for different question-related KB subgrap…

Cited by 0SourceScholar
2021

Online Credit Payment Fraud Detection via Structure-Aware Hierarchical Recurrent Neural Network

IJCAI 2021poster

Online credit payment fraud detection plays a critical role in financial institutions due to the growing volume of fraudulent transactions. Recently, researchers have shown an increased interest in capturing users’ dynamic and evolving fraudulent tendencies from their behavior sequences. However, mo…

2021

Style-Based Point Generator With Adversarial Rendering for Point Cloud Completion

CVPR 2021poster

In this paper, we proposed a novel Style-based Point Generator with Adversarial Rendering (SpareNet) for point cloud completion. Firstly, we present the channel-attentive EdgeConv to fully exploit the local structures as well as the global shape in point features. Secondly, we observe that the conca…

Cited by 106PDFcodeScholar
2021

Unsupervised Pre-Training for Person Re-Identification

CVPR 2021poster

In this paper, we present a large scale unlabeled person re-identification (Re-ID) dataset "LUPerson" and make the first attempt of performing unsupervised pre-training for improving the generalization ability of the learned person Re-ID feature representation. This is to address the problem that al…

Cited by 226PDFcodeScholar
2020

Deep Merging: Vehicle Merging Controller Based on Deep Reinforcement Learning with Embedding Network

ICRA 2020poster

Vehicles at highway merging sections must make lane changes to join the highway. This lane change can generate congestion. To reduce congestion, vehicles should merge so as not to affect traffic flow as much as possible. In our study, we propose a vehicle controller called Deep Merging that uses dee…

Cited by 25SourceScholar
2020

FCEM: A Novel Fast Correlation Extract Model For Real Time Steganalysis Of VoIP Stream Via Multi-Head Attention

ICASSP 2020accepted

Extracting correlation features between codes-words with high computational efficiency is crucial to steganalysis of Voice over IP (VoIP) streams. In this paper, we utilized attention mechanisms, which have recently attracted enormous interests due to their highly parallelizable computation and flex…

Cited by 0SourceScholar
2020

Modelling Long-distance Node Relations for KBQA with Global Dynamic Graph

COLING 2020main

The structural information of Knowledge Bases (KBs) has proven effective to Question Answering (QA). Previous studies rely on deep graph neural networks (GNNs) to capture rich structural information, which may not model node relations in particularly long distance due to oversmoothing issue. To addr…

Cited by 13SourcePDFScholar
2020

Rethinking the Hyperparameters for Fine-tuning

ICLR 2020poster

Fine-tuning from pre-trained ImageNet models has become the de-facto standard for various computer vision tasks. Current practices for fine-tuning typically involve selecting an ad-hoc choice of hyperparameters and keeping them fixed to values normally used for training from scratch. This paper re-e…

Cited by 184SourcecodeScholar
2020

Salamanderbot: A soft-rigid composite continuum mobile robot to traverse complex environments

ICRA 2020poster

Soft robots are theoretically well-suited to rescue and exploration applications where their flexibility allows for the traversal of highly cluttered environments. However, most existing mobile soft robots are not fast or powerful enough to effectively traverse three dimensional environments. In thi…

Cited by 23SourceScholar
2017

MIML-FCN+: Multi-Instance Multi-Label Learning via Fully Convolutional Networks With Privileged Information

CVPR 2017poster

Multi-instance multi-label (MIML) learning has many interesting applications in computer visions, including multi-object recognition and automatic image tagging. In these applications, additional information such as bounding-boxes, image captions and descriptions is often available during training p…

Cited by 87PDFScholar
2016

Exploit Bounding Box Annotations for Multi-Label Object Recognition

CVPR 2016poster

Convolutional neural networks (CNNs) have shown great performance as general feature representations for object recognition applications. However, for multi-label images that contain multiple objects from different categories, scales and locations, global CNN features are not optimal. In this paper,…

Cited by 210PDFScholar