← Search

Xin Huang

62 accepted papers

2026

AdaMCoT: Rethinking Cross-Lingual Factual Reasoning Through Adaptive Multilingual Chain-of-Thought

AAAI 2026technical

Large language models (LLMs) have shown impressive multilingual capabilities through pretraining on diverse corpora. While these models show strong reasoning abilities, their performance varies significantly across languages due to imbalanced training data distribution. Existing approaches using sam

Cited by 11SourcePDFScholar
2026

Differentially Private Cross-Silo Recommendation from Implicit Feedback

ICML 2026poster

Cross-silo recommendation from implicit feedback is a key task in modern recommender systems, where user-item interaction data are distributed across multiple parties and cannot be centrally collected. Unlike explicit feedback, which provides fully observed real-valued ratings, implicit feedback is …

Cited by 0SourceScholar
2026

NaTex: Seamless Texture Generation as Latent Color Diffusion

CVPR 2026

We present NaTex, a native texture generation framework that predicts texture color directly in 3D space. In contrast to previous approaches that rely on baking 2D multi-view images synthesized by geometry-conditioned Multi-View Diffusion models (MVDs), NaTex avoids several inherent limitations of t

Cited by 8SourcecodeScholar
2026

On the Wings of Imagination: Conflicting Script-based Multi-role Framework for Humor Caption Generation

ICLR 2026poster

Humor is a commonly used and high-level human language in daily life. However, humor generation is a challenging task for large language models (LLMs) in multi-modal contexts, but with many useful applications of funny caption generation for images, requiring visual understanding, humor reasoning, c…

Cited by 0SourceScholar
2025

Adaptive Prompt-Based Semantic Embedding with Inspire Potential of Implicit Knowledge for Cross-Modal Retrieval

AAAI 2025technical

In the era of big data, cross-modal retrieval is increasingly important in research and application. Given the latent complexity and non-intuitive nature of cross-modal relationships, leveraging external knowledge such as large models has become a popular approach to facilitate modality alignment. E…

2025

Bright-NeRF: Brightening Neural Radiance Field with Color Restoration from Low-Light RAW Images

AAAI 2025technical

Neural Radiance Fields (NeRF) have demonstrated prominent performance in novel view synthesis tasks. However, their input heavily relies on image acquisition under normal light conditions, making it challenging to learn accurate scene contents in low-light environments where images typically exhibit…

Cited by 0SourcePDFScholar
2025

Causal Composition Diffusion Model for Closed-loop Traffic Generation

CVPR 2025poster

Simulation is critical for safety evaluation in autonomous driving, particularly in capturing complex interactive behaviors. However, generating **realistic** and **controllable** traffic scenarios in long-tail situations remains a significant challenge. Existing generative models suffer from the co…

2025

Cohere3D: Exploiting Temporal Coherence for Unsupervised Representation Learning of Vision-Based Autonomous Driving

ICRA 2025

Multi-frame temporal inputs are important for vision-based autonomous driving. Observations from different angles enable the recovery of 3 D object states from 2 D images as long as we can identify the same instance from different input frames. However, the dynamic nature of driving scenes leads to

Cited by 3SourceScholar
2025

DriveGPT: Scaling Autoregressive Behavior Models for Driving

ICML 2025poster

We present DriveGPT, a scalable behavior model for autonomous driving. We model driving as a sequential decision-making task, and learn a transformer model to predict future agent states as tokens in an autoregressive fashion. We scale up our model parameters and training data by multiple orders of…

Cited by 1SourcePDFScholar
2025

Dynamic Seed-GrowthCM: A Dynamic Benefit-Oriented Algorithm for Core Maximization on Large Graphs

IJCAI 2025

The k-core has garnered significant attention in recent research as an effective measure of node importance within a graph. A k-core is defined as the maximal induced subgraph where each node has a degree of at least k. This paper addresses the core maximization problem: given a graph G, an integer

Cited by 0SourcePDFScholar
2025

GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning

NeurIPS 2025poster

With the rapid development of Large Vision Language Models, the focus of Graphical User Interface (GUI) agent tasks shifts from single-screen tasks to complex screen navigation challenges. However, real-world GUI environments, such as PC software and mobile Apps, are often complex and proprietary,…

Cited by 0SourceScholar
2025

Investigating and Scaling up Code-Switching for Multilingual Language Model Pre-Training

ACL 2025finding

Large language models (LLMs) exhibit remarkable multilingual capabilities despite the extreme language imbalance in the pre-training data. In this paper, we closely examine the reasons behind this phenomenon, focusing on the pre-training corpus. We find that the existence of code-switching, alternat…

2025

Large Language Models Are Cross-Lingual Knowledge-Free Reasoners

NAACL 2025long

Large Language Models have demonstrated impressive reasoning capabilities across multiple languages. However, the relationship between capabilities in different languages is less explored. In this work, we decompose the process of reasoning tasks into two separated components: knowledge retrieval an…

2025

LogRules: Enhancing Log Analysis Capability of Large Language Models through Rules

NAACL 2025findings

Currently, large language models (LLMs) have achieved impressive performance in natural language processing tasks. However, LLMs still exhibit many hallucinations when analyzing system logs, which is due to the implicit knowledge and rules in logs that LLMs cannot capture. Based on this, we propose…

Cited by 0SourcePDFScholar
2025

MGPose: Wide-Baseline Relative Camera Pose Estimation Using Matching-Guided Dual Channel-Attention

RA-L 2025

Relative camera pose estimation is a fundamental task in computer vision and robotics. In wide- baseline scenarios with limited visual overlap, traditional methods often perform poorly. Existing deep learning approaches are also hindered by irrelevant features and insufficient modeling of the relati

Cited by 0SourceScholar
2025

Make Information Diffusion Explainable: LLM-based Causal Framework for Diffusion Prediction

NeurIPS 2025poster

Information diffusion prediction, which aims to forecast future infected users during the information spreading process on social platforms, is a challenging and critical task for public opinion analysis. With the development of social platforms, mass communication has become increasingly widespread…

Cited by 0SourceScholar
2025

Material Anything: Generating Materials for Any 3D Object via Diffusion

CVPR 2025highlight

We present **Material Anything**, a fully-automated, unified diffusion framework designed to generate physically-based materials for 3D objects. Unlike existing methods that rely on complex pipelines or case-specific optimizations, Material Anything offers a robust, end-to-end solution adaptable to…

Cited by 4SourcePDFScholar
2025

Math-PUMA: Progressive Upward Multimodal Alignment to Enhance Mathematical Reasoning

AAAI 2025technical

Multimodal Large Language Models (MLLMs) excel in solving text-based mathematical problems, but they struggle with mathematical diagrams since they are primarily trained on natural scene images. For humans, visual aids generally enhance problem-solving, but MLLMs perform worse as information shifts…

2025

MoE-LPR: Multilingual Extension of Large Language Models Through Mixture-of-Experts with Language Priors Routing

AAAI 2025technical

Large Language Models (LLMs) are often English-centric due to the disproportionate distribution of languages in their pre-training data. Enhancing non-English language capabilities through post-pretraining often results in catastrophic forgetting of high-resource languages. Previous methods either a…

2025

Neural LightRig: Unlocking Accurate Object Normal and Material Estimation with Multi-Light Diffusion

CVPR 2025poster

Recovering the geometry and materials of objects from a single image is challenging due to its under-constrained nature. In this paper, we present Neural LightRig, a novel framework that boosts intrinsic estimation by leveraging auxiliary multi-lighting conditions from 2D diffusion priors. Specifica…

2025

PROFIT: A Specialized Optimizer for Deep Fine Tuning

NeurIPS 2025poster

The fine-tuning of pre-trained models has become ubiquitous in generative AI, computer vision, and robotics. Although much attention has been paid to improving the efficiency of fine-tuning model, there has been less scholarship around fine-tuning specifically for improved model performance. To reme…

Cited by 0SourceScholar
2025

PerReactor: Offline Personalised Multiple Appropriate Facial Reaction Generation

AAAI 2025technical

In dyadic human-human interactions, individuals may express multiple different facial reactions in response to the same/similar behaviours expressed by their conversational partners depending on their personalised behaviour patterns. As a result, frequently-employed reconstruction loss-based strateg…

2025

R-PRM: Reasoning-Driven Process Reward Modeling

EMNLP 2025

Process Reward Models (PRMs) have emerged as a promising solution to address the reasoning mistakes of large language models (LLMs). However, existing PRMs typically output evaluation scores directly, limiting both learning efficiency and evaluation accuracy. This limitation is further compounded by

2025

Reproducible Vision-Language Models Meet Concepts Out of Pre-Training

CVPR 2025poster

Contrastive Language-Image Pre-training (CLIP) models as a milestone of modern multimodal intelligence, its generalization mechanism grasped massive research interests in the community. While existing studies limited in the scope of pre-training knowledge, hardly underpinned its generalization to co…

Cited by 0SourcePDFScholar
2025

Understanding LLMs’ Cross-Lingual Context Retrieval: How Good It Is And Where It Comes From

EMNLP 2025

Cross-lingual context retrieval (extracting contextual information in one language based on requests in another) is a fundamental aspect of cross-lingual alignment, but the performance and mechanism of it for large language models (LLMs) remains unclear. In this paper, we evaluate the cross-lingual

2025

VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision

CoRL 2025poster

Human drivers rely on commonsense reasoning to navigate diverse and dynamic real-world scenarios. Existing end-to-end (E2E) autonomous driving (AD) models are typically optimized to mimic driving patterns observed in data, without capturing the underlying reasoning processes. This limitation constr…

Cited by 0SourceScholar
2025

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models

IJCAI 2025

Vision-Language Models (VLMs) are essential for multimodal tasks, especially compositional reasoning (CR) tasks, which require distinguishing fine-grained semantic differences between visual and textual embeddings. However, existing methods primarily fine-tune the model by generating text-based hard

2024

G2L-CariGAN: Caricature Generation from Global Structure to Local Features

AAAI 2024technical

Existing GAN-based approaches to caricature generation mainly focus on exaggerating a character’s global facial structure. This often leads to the failure in highlighting significant facial features such as big eyes and hook nose. To address this limitation, we propose a new approach termed as G2L-C…

Cited by 0SourcePDFScholar
2024

Getting More from Less: Large Language Models are Good Spontaneous Multilingual Learners

EMNLP 2024main

Recently, Large Language Models (LLMs) have shown impressive language capabilities, while most of them have very unbalanced performance across different languages. Multilingual alignment based on the translation parallel data is an effective method to enhance LLMs’ multilingual capabilities. In this…

2024

HumanNorm: Learning Normal Diffusion Model for High-quality and Realistic 3D Human Generation

CVPR 2024poster

Recent text-to-3D methods employing diffusion models have made significant advancements in 3D human generation. However these approaches face challenges due to the limitations of text-to-image diffusion models which lack an understanding of 3D structures. Consequently these methods struggle to achie…

2024

Privileged Prior Information Distillation for Image Matting

AAAI 2024technical

Performance of trimap-free image matting methods is limited when trying to decouple the deterministic and undetermined regions, especially in the scenes where foregrounds are semantically ambiguous, chromaless, or high transmittance. In this paper, we propose a novel framework named Privileged Prior…

Cited by 1SourcePDFScholar
2024

SeaEval for Multilingual Foundation Models: From Cross-Lingual Alignment to Cultural Reasoning

NAACL 2024long

We present SeaEval, a benchmark for multilingual foundation models. In addition to characterizing how these models understand and reason with natural language, we also investigate how well they comprehend cultural practices, nuances, and values. Alongside standard accuracy metrics, we investigate th…

2023

A Retrospect to Multi-prompt Learning across Vision and Language

ICCV 2023poster

The vision community is undergoing the unprecedented progress with the emergence of Vision-Language Pretraining Models (VLMs). Prompt learning plays as the holy grail of accessing VLMs since it enables their fast adaptation to downstream tasks with limited resources. Whereas existing research millin…

Cited by 7PDFcodeScholar
2023

Inverting the Imaging Process by Learning an Implicit Camera Model

CVPR 2023poster

Representing visual signals with implicit coordinate-based neural networks, as an effective replacement of the traditional discrete signal representation, has gained considerable popularity in computer vision and graphics. In contrast to existing implicit neural representations which focus on modell…

Cited by 14SourcePDFScholar
2023

Local Implicit Ray Function for Generalizable Radiance Field Representation

CVPR 2023poster

We propose LIRF (Local Implicit Ray Function), a generalizable neural rendering approach for novel view rendering. Current generalizable neural radiance fields (NeRF) methods sample a scene with a single ray per pixel and may therefore render blurred or aliased views when the input views and rendere…

Cited by 30SourcePDFScholar
2023

MixMAE: Mixed and Masked Autoencoder for Efficient Pretraining of Hierarchical Vision Transformers

CVPR 2023poster

In this paper, we propose Mixed and Masked AutoEncoder (MixMAE), a simple but efficient pretraining method that is applicable to various hierarchical Vision Transformers. Existing masked image modeling (MIM) methods for hierarchical Vision Transformers replace a random subset of input tokens with a…

2023

P4P: Conflict-Aware Motion Prediction for Planning in Autonomous Driving

IROS 2023poster

Motion prediction is crucial in enabling safe motion planning for autonomous vehicles in interactive scenarios. It allows the planner to identify potential conflicts with other traffic agents and generate safe plans. Existing motion predictors often focus on reducing prediction errors, yet it remain…

Cited by 4SourceScholar
2023

Ultra Real-Time Portrait Matting via Parallel Semantic Guidance

ICASSP 2023accepted

Most existing portrait matting models either require expensive auxiliary information or try to decompose the task into sub-tasks that are usually resource-hungry. These challenges limit its application on low-power computing devices. In this paper, we propose an ultra-light-weighted portrait matting…

Cited by 0SourceScholar
2022

"UniNet: Unified Architecture Search with Convolution, Transformer, and MLP"

ECCV 2022poster

"Recently, transformer and multi-layer perceptron (MLP) architectures have achieved impressive results on various vision tasks. However, how to effectively combine those operators to form high-performance hybrid visual architectures still remains a challenge. In this work, we study the learnable com…

2022

HYPER: Learned Hybrid Trajectory Prediction via Factored Inference and Adaptive Sampling

ICRA 2022poster

Modeling multi-modal high-level intent is important for ensuring diversity in trajectory prediction. Existing approaches explore the discrete nature of human intent before predicting continuous trajectories, to improve accuracy and support explainability. However, these approaches often assume the i…

Cited by 33SourceScholar
2022

InterSim: Interactive Traffic Simulation via Explicit Relation Modeling

IROS 2022poster

Interactive traffic simulation is crucial to autonomous driving systems by enabling testing for planners in a more scalable and safe way compared to real-world road testing. Existing approaches learn an agent model from large-scale driving data to simulate realistic traffic scenarios, yet it remains…

Cited by 35SourcecodeScholar
2022

M2I: From Factored Marginal Trajectory Prediction to Interactive Prediction

CVPR 2022poster

Predicting future motions of road participants is an important task for driving autonomously in urban scenes. Existing models excel at predicting marginal trajectories for single agents, yet it remains an open question to jointly predict scene compliant trajectories over multiple agents. The challen…

Cited by 122PDFScholar
2022

Pyramid-BERT: Reducing Complexity via Successive Core-set based Token Selection

ACL 2022long

Transformer-based language models such as BERT (CITATION) have achieved the state-of-the-art performance on various NLP tasks, but are computationally prohibitive. A recent line of works use various heuristics to successively shorten sequence length while transforming tokens through encoders, in tas…

Cited by 21SourcePDFScholar
2022

TIP: Task-Informed Motion Prediction for Intelligent Vehicles

IROS 2022poster

When predicting trajectories of road agents, motion predictors often approximate the future distribution by a limited number of samples. This constraint requires the predictors to generate samples that best support the task given task specifications. However, existing predictors are often optimized…

Cited by 15SourceScholar
2022

Trajectory Prediction with Linguistic Representations

ICRA 2022poster

Language allows humans to build mental models that interpret what is happening around them resulting in more accurate long-term predictions. We present a novel trajectory prediction model that uses linguistic intermediate representations to forecast trajectories, and is trained using trajectory samp…

Cited by 22SourceScholar
2021

An Anytime Algorithm for Chance Constrained Stochastic Shortest Path Problems and Its Application to Aircraft Routing

ICRA 2021poster

Aircraft routing problem is a crucial component for flight automation. Despite recent successes, challenges still remain when the environment is dynamic and uncertain. In this paper, we tackle the following two challenges. First, when the environment is uncertain, it is much safer if the route plann…

Cited by 23SourceScholar
2021

CARPAL: Confidence-Aware Intent Recognition for Parallel Autonomy

RA-L 2021

Predicting driver intentions is a difficult and crucial task for advanced driver assistance systems. Traditional confidence measures on predictions often ignore the way predicted trajectories affect downstream decisions for safe driving. In this letter, we propose a novel multi-task intent recogniti

Cited by 7SourceScholar
2021

Entity-level Cross-modal Learning Improves Multi-modal Machine Translation

EMNLP 2021finding

Multi-modal machine translation (MMT) aims at improving translation performance by incorporating visual information. Most of the studies leverage the visual information through integrating the global image features as auxiliary input or decoding by attending to relevant local regions of the image. H…

Cited by 12SourcePDFScholar
2021

Unseen Entity Handling in Complex Question Answering over Knowledge Base via Language Generation

EMNLP 2021finding

Complex question answering over knowledge base remains as a challenging task because it involves reasoning over multiple pieces of information, including intermediate entities/relations and other constraints. Previous methods simplify the SPARQL query of a question into such forms as a list or a gra…

Cited by 19SourcePDFScholar
2020

A Two-level Reinforcement Learning Algorithm for Ambiguous Mean-variance Portfolio Selection Problem

IJCAI 2020poster

Traditional modeling on the mean-variance portfolio selection often assumes a full knowledge on statistics of assets' returns. It is, however, not always the case in real financial markets. This paper deals with an ambiguous mean-variance portfolio selection problem with a mixture model on the retur…

Cited by 0SourcePDFScholar
2020

DiversityGAN: Diversity-Aware Vehicle Motion Prediction via Latent Semantic Sampling

RA-L 2020

Vehicle trajectory prediction is crucial for autonomous driving and advanced driver assistant systems. While existing approaches may sample from a predicted distribution of vehicle trajectories, they lack the ability to explore it - a key ability for evaluating safety from a planning and verificatio

Cited by 80SourceScholar
2020

Fast Risk Assessment for Autonomous Vehicles Using Learned Models of Agent Futures

RSS 2020poster

This paper presents fast non-sampling based methods to assess the risk of trajectories for autonomous vehicles when probabilistic predictions of other agents’ futures are generated by deep neural networks (DNNs). The presented methods address a wide range of representations for uncertain predictions…

2020

NMS by Representative Region: Towards Crowded Pedestrian Detection by Proposal Pairing

CVPR 2020poster

Although significant progress has been made in pedestrian detection recently, pedestrian detection in crowded scenes is still challenging. The heavy occlusion between pedestrians imposes great challenges to the standard Non-Maximum Suppression (NMS). A relative low threshold of intersection over uni…

Cited by 202PDFScholar
2020

Photon-Efficient 3D Imaging with A Non-Local Neural Network

ECCV 2020poster

Photon-efficient imaging has enabled a number of applications relying on single-photon sensors that can capture a 3D image with as few as one photon per pixel. In practice, however, measurements of low photon counts are often mixed with heavy background noise, which poses a great challenge for exist…

2019

Uncertainty-Aware Driver Trajectory Prediction at Urban Intersections

ICRA 2019poster

Predicting the motion of a driver’s vehicle is crucial for advanced driving systems, enabling detection of potential risks towards shared control between the driver and automation systems. In this paper, we propose a variational neural network approach that predicts future driver trajectory distribu…

Cited by 104SourceScholar