← Search

Xing Xu

31 accepted papers

2026

Cross-Tactile Sensor Representation Learning

ICML 2026poster

Visuo-tactile sensors have been widely adopted in robotic manipulation. However, inherent heterogeneity in sensor designs hinders the learning of unified tactile representations in cross-sensor scenarios. Existing methods that focus on reconstruction or task-specific supervision often fail to captur…

Cited by 0SourceScholar
2026

De-biased Natural Language Egocentric Task Verification via Prototypical Evidence Learning

AAAI 2026technical

Natural Language-based Egocentric Task Verification (NLETV) aims to verify the alignment between action sequences in egocentric videos and their corresponding textual descriptions. However, existing NLETV approaches are still facing two critical challenges: (1) These methods are designed for simul

Cited by 0SourcePDFScholar
2026

Ego3S: Select, Strengthen, and Synchronize for Efficient Egocentric Reasoning

ICML 2026poster

Egocentric reasoning fundamentally differs from third-person understanding in LVLMs. Third-person settings offer wide and stable contexts with consistent global regularities, allowing models to utilize broad statistical correlations. In contrast, egocentric scenes are highly dynamic and heterogeneou…

Cited by 0SourceScholar
2026

Hyper-Opinion Vagueness Quantification for Robust Multimodal Learning

AAAI 2026technical

Robust Multimodal Learning (RML) aims to address the issues of unreliable predictions of multimodal models. Nevertheless, previous RML works often struggle to distinguish between different categories that rely on identical intra-modal cues, making ambiguous predictions. We defined this degree of ``u

Cited by 0SourcePDFScholar
2026

Language-Grounded Decoupled Action Representation for Robotic Manipulation

CVPR 2026

The heterogeneity between high-level vision-language understanding and low-level action control remains a fundamental challenge in robotic manipulation. Although recent methods have advanced task-specific action alignment, they often struggle to generate robust and accurate actions for novel or sema

Cited by 0SourceScholar
2026

Multimodal Learning on Low-Quality Data with Conformal Predictive Self-Calibration

CVPR 2026

Multimodal learning often grapples with the challenge of low-quality data, which predominantly manifests as two facets: modality imbalance and noisy corruption. While these issues are often studied in isolation, we argue that they share a common root in the predictive uncertainty towards the reliabi

Cited by 0SourcecodeScholar
2025

From Observation to Understanding: Front-Door Adjustments with Uncertainty Calibration for Enhancing Egocentric Reasoning in LVLMs

ACL 2025finding

Recent progress in large vision-language models (LVLMs) has shown substantial potential across a broad spectrum of third-person tasks. However, adapting these LVLMs to egocentric scenarios remains challenging due to their third-person training bias. Existing methods that adapt LVLMs for first-person…

2025

PHGC: Procedural Heterogeneous Graph Completion for Natural Language Task Verification in Egocentric Videos

CVPR 2025poster

Natural Language-based Egocentric Task Verification (NLETV) aims to equip agents to determine if operation flows of procedural tasks in egocentric videos align with natural language instructions. Describing rules with natural language provides generalizable applications, but also raises cross-modal…

2025

ReCon: Enhancing True Correspondence Discrimination through Relation Consistency for Robust Noisy Correspondence Learning

CVPR 2025poster

Can we accurately identify the true correspondences from multimodal datasets containing mismatched data pairs? Existing methods primarily emphasize the similarity matching between the representations of objects across modalities, potentially neglecting the crucial relation consistency within modalit…

2025

TAU-106K: A New Dataset for Comprehensive Understanding of Traffic Accident

ICLR 2025poster

Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in general visual understanding tasks. However, their potential for high-level, fine-grained comprehension, such as anomaly understanding, remains unexplored. Focusing on traffic accidents, a critical and practical sce…

2024

Adaptive Uncertainty-Based Learning for Text-Based Person Retrieval

AAAI 2024technical

Text-based person retrieval aims at retrieving a specific pedestrian image from a gallery based on textual descriptions. The primary challenge is how to overcome the inherent heterogeneous modality gap in the situation of significant intra-class variation and minimal inter-class variation. Existing…

2024

Embracing Unimodal Aleatoric Uncertainty for Robust Multimodal Fusion

CVPR 2024poster

As a fundamental problem in multimodal learning multimodal fusion aims to compensate for the inherent limitations of a single modality. One challenge of multimodal fusion is that the unimodal data in their unique embedding space mostly contains potential noise which leads to corrupted cross-modal in…

Cited by 8SourcePDFScholar
2024

T-SciQ: Teaching Multimodal Chain-of-Thought Reasoning via Large Language Model Signals for Science Question Answering

AAAI 2024technical

Large Language Models (LLMs) have recently demonstrated exceptional performance in various Natural Language Processing (NLP) tasks. They have also shown the ability to perform chain-of-thought (CoT) reasoning to solve complex problems. Recent studies have explored CoT reasoning in complex multimodal…

2023

Alignment-Enriched Tuning for Patch-Level Pre-trained Document Image Models

AAAI 2023technical

Alignment between image and text has shown promising improvements on patch-level pre-trained document image models. However, investigating more effective or finer-grained alignment techniques during pre-training requires a large amount of computation cost and time. Thus, a question naturally arises:…

2023

ICL-D3IE: In-Context Learning with Diverse Demonstrations Updating for Document Information Extraction

ICCV 2023poster

Large language models (LLMs), such as GPT-3 and ChatGPT, have demonstrated remarkable results in various natural language processing (NLP) tasks with in-context learning, which involves inference based on a few demonstration examples. Despite their successes in NLP tasks, no investigation has been c…

Cited by 52PDFcodeScholar
2023

ImbSAM: A Closer Look at Sharpness-Aware Minimization in Class-Imbalanced Recognition

ICCV 2023poster

Class imbalance is a common challenge in real-world recognition tasks, where the majority of classes have few samples, also known as tail classes. We address this challenge with the perspective of generalization and empirically find that the promising Sharpness-Aware Minimization (SAM) fails to addr…

Cited by 17PDFcodeScholar
2023

LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models

EMNLP 2023long main

The success of large language models (LLMs), like GPT-4 and ChatGPT, has led to the development of numerous cost-effective and accessible alternatives that are created by finetuning open-access LLMs with task-specific data (e.g., ChatDoctor) or instruction data (e.g., Alpaca). Among the various fine…

Cited by 0SourcecodeScholar
2022

Semi-Supervised Video Paragraph Grounding With Contrastive Encoder

CVPR 2022poster

Video events grounding aims at retrieving the most relevant moments from an untrimmed video in terms of a given natural language query. Most previous works focus on Video Sentence Grounding (VSG), which localizes the moment with a sentence query. Recently, researchers extended this task to Video Par…

Cited by 35PDFScholar
2022

TVT: Three-Way Vision Transformer through Multi-Modal Hypersphere Learning for Zero-Shot Sketch-Based Image Retrieval

AAAI 2022technical

In this paper, we study the zero-shot sketch-based image retrieval (ZS-SBIR) task, which retrieves natural images related to sketch queries from unseen categories. In the literature, convolutional neural networks (CNNs) have become the de-facto standard and they are either trained end-to-end or used…

Cited by 49SourcePDFScholar
2021

Enhancing Audio-Visual Association with Self-Supervised Curriculum Learning

AAAI 2021technical

The recent success of audio-visual representations learning can be largely attributed to their pervasive concurrency property, which can be used as a self-supervision signal and extract correlation information. While most recent works focus on capturing the shared associations between the audio and…

Cited by 26SourcePDFScholar
2021

Feature Space Targeted Attacks by Statistic Alignment

IJCAI 2021poster

By adding human-imperceptible perturbations to images, DNNs can be easily fooled. As one of the mainstream methods, feature space targeted attacks perturb images by modulating their intermediate feature maps, for the discrepancy between the intermediate source and target features is minimized. Howev…

2021

From General to Specific: Informative Scene Graph Generation via Balance Adjustment

ICCV 2021poster

The scene graph generation (SGG) task aims to detect visual relationship triplets, i.e., subject, predicate, object, in an image, providing a structural vision layout for scene understanding. However, current models are stuck in common predicates, e.g., "on" and "at", rather than informative ones, e…

Cited by 105PDFcodeScholar
2021

Multi-Stage Aggregated Transformer Network for Temporal Language Localization in Videos

CVPR 2021poster

We address the problem of localizing a specific moment from an untrimmed video by a language sentence query. Generally, previous methods mainly exist two problems that are not fully solved: 1) How to effectively model the fine-grained visual-language alignment between video and language query? 2) Ho…

Cited by 98PDFScholar
2021

Partial Feature Selection and Alignment for Multi-Source Domain Adaptation

CVPR 2021poster

Multi-Source Domain Adaptation (MSDA), which dedicates to transfer the knowledge learned from multiple source domains to an unlabeled target domain, has drawn increasing attention in the research community. By assuming that the source and target domains share consistent key feature representations a…

Cited by 41PDFScholar
2021

PoseGTAC: Graph Transformer Encoder-Decoder with Atrous Convolution for 3D Human Pose Estimation

IJCAI 2021poster

Graph neural networks (GNNs) have been widely used in the 3D human pose estimation task, since the pose representation of a human body can be naturally modeled by the graph structure. Generally, most of the existing GNN-based models utilize the restricted receptive fields of filters and single-scale i…

Cited by 27SourcePDFScholar
2020

Universal Weighting Metric Learning for Cross-Modal Matching

CVPR 2020poster

Cross-modal matching has been a highlighted research topic in both vision and language areas. Learning appropriate mining strategy to sample and weight informative pairs is crucial for the cross-modal matching performance. However, most existing metric learning methods are developed for unimodal mat…

Cited by 113PDFcodeScholar
2020

What Machines See Is Not What They Get: Fooling Scene Text Recognition Models With Adversarial Text Images

CVPR 2020oral

The research on scene text recognition (STR) has made remarkable progress in recent years with the development of deep neural networks (DNNs). Recent studies on adversarial attack have verified that a DNN model designed for non-sequential tasks (e.g., classification, segmentation and retrieval) can…

Cited by 49PDFScholar
2019

Sequence-To-Sequence Domain Adaptation Network for Robust Text Image Recognition

CVPR 2019poster

Domain adaptation has shown promising advances for alleviating domain shift problem. However, recent visual domain adaptation works usually focus on non-sequential object recognition with a global coarse alignment, which is inadequate to transfer effective knowledge for sequence-like text images wit…

Cited by 163PDFScholar
2017

Matrix Tri-Factorization With Manifold Regularizations for Zero-Shot Learning

CVPR 2017poster

Zero-shot learning (ZSL) aims to recognize objects of unseen classes with available training data from another set of seen classes. Existing solutions are focused on exploring knowledge transfer via an intermediate semantic embedding (e.g.s, attributes) shared between seen and unseen classes. In thi…

Cited by 158PDFScholar
2016

Context adaptive thresholding and entropy coding for very low complexity JPEG transcoding

ICASSP 2016accepted

The ever increasing quantity of user generated photos, nearly all compressed using JPEG, has created a growing storage burden on photo storage and sharing services. This creates the need for compression techniques that take JPEG compressed images as inputs. In this paper we propose two novel very lo…

Cited by 0SourceScholar