← Search

Xin Lin

42 accepted papers

2026

AirSim360: A Panoramic Simulation Platform within Drone View

CVPR 2026

The field of 360-degree omnidirectional understanding has been receiving increasing attention for advancing spatial intelligence. However, the lack of large-scale and diverse data remains a major limitation. In this work, we propose AirSim360, a simulation platform for omnidirectional data from aeri

Cited by 0SourcecodeScholar
2026

D$^2$GS: Depth-and-Density Guided Gaussian Splatting for Stable and Accurate Sparse-View Reconstruction

ICLR 2026poster

Recent advances in 3D Gaussian Splatting (3DGS) enable real-time, high-fidelity novel view synthesis (NVS) with explicit 3D representations. However, performance degradation and instability remain significant under sparse-view conditions. In this work, we identify two key failure modes under sparse-…

Cited by 0SourcecodeScholar
2026

DA$^{2}$: Depth Anything in Any Direction

ICLR 2026poster

Panorama has a full FoV (360$^\circ\times$180$^\circ$), offering a more complete visual description than perspective images. Thanks to this characteristic, panoramic depth estimation is gaining increasing traction in 3D vision. However, due to the scarcity of panoramic data, previous methods are oft…

Cited by 0SourcecodeScholar
2026

Depth Any Panoramas: A Foundation Model for Panoramic Depth Estimation

CVPR 2026

In this work, we present a panoramic metric depth foundation model that generalizes across diverse scene distances. We explore a data-in-the-loop paradigm from the view of both data construction and framework design. We collect a large-scale dataset by combining public datasets, high-quality synthet

Cited by 0SourcecodeScholar
2026

Enhancing Generalization of Depth Estimation Foundation Model via Weakly-Supervised Adaptation with Regularization

AAAI 2026technical

The emergence of foundation models has substantially advanced zero-shot generalization in monocular depth estimation (MDE), as exemplified by the Depth Anything series. However, given access to some data from downstream tasks, a natural question arises: can the performance of these models be further

Cited by 0SourcePDFScholar
2025

A Graph Interaction Framework on Relevance for Multimodal Named Entity Recognition with Multiple Images

COLING 2025main

Posts containing multiple images have significant research potential in Multimodal Named Entity Recognition nowadays. The previous methods determine whether the images are related to named entities in the text through similarity computation, such as using CLIP. However, it is not effective in some c…

Cited by 0SourcePDFScholar
2025

ChartM3: A Multi-Stage Code-Driven Pipeline for Constructing Multi-Dimensional and Multi-Step Visual Reasoning Data in Chart Comprehension

EMNLP 2025

Complex chart understanding tasks demand advanced visual recognition and reasoning capabilities from multimodal large language models (MLLMs). However, current research provides limited coverage of complex chart scenarios and computation-intensive reasoning tasks prevalent in real-world applications

Cited by 0SourcePDFScholar
2025

Distraction is All You Need for Multimodal Large Language Model Jailbreaking

CVPR 2025highlight

Multimodal Large Language Models (MLLMs) bridge the gap between visual and textual data, enabling a range of advanced applications. However, complex internal interactions among visual elements and their alignment with text can introduce vulnerabilities, which may be exploited to bypass safety mechan…

Cited by 1SourcePDFScholar
2025

Explore What LLM Does Not Know in Complex Question Answering

AAAI 2025technical

Complex question answering (QA) is a challenging task in artificial intelligence research which requires reasoning based on related knowledge. The retrieval-augmented generation (RAG) based on large language models (LLMs) have become one promising solution in QA. To facilitate RAG more effectively,…

2025

From Objectives to Questions: A Planning-based Framework for Educational Mathematical Question Generation

ACL 2025long

Automatically generating high-quality mathematical problems that align with educational objectives is a crucial task in NLP-based educational technology. Traditional generation methods focus primarily on textual quality, but they often overlook educational objectives. Moreover, these methods address…

Cited by 0SourcePDFScholar
2025

HQGS: High-Quality Novel View Synthesis with Gaussian Splatting in Degraded Scenes

ICLR 2025poster

3D Gaussian Splatting (3DGS) has shown promising results for Novel View Synthesis. However, while it is quite effective when based on high-quality images, its performance declines as image quality degrades, due to lack of resolution, motion blur, noise, compression artifacts, or other factors common…

2025

HYPERION: Fine-Grained Hypersphere Alignment for Robust Federated Graph Learning

NeurIPS 2025spotlight

Robust Federated Graph Learning (FGL) provides an effective decentralized framework for training Graph Neural Networks (GNNs) in noisy-label environments. However, the subtlety of noise during training presents formidable obstacles for developing robust FGL systems. Previous robust FL approaches nei…

Cited by 0SourceScholar
2025

LLM-OREF: An Open Relation Extraction Framework Based on Large Language Models

EMNLP 2025

The goal of open relation extraction (OpenRE) is to develop an RE model that can generalize to new relations not encountered during training. Existing studies primarily formulate OpenRE as a clustering task. They first cluster all test instances based on the similarity between the instances, and the

2025

RobuRCDet: Enhancing Robustness of Radar-Camera Fusion in Bird's Eye View for 3D Object Detection

ICLR 2025poster

While recent low-cost radar-camera approaches have shown promising results in multi-modal 3D object detection, both sensors face challenges from environmen- tal and intrinsic disturbances. Poor lighting or adverse weather conditions de- grade camera performance, while radar suffers from noise and po…

Cited by 1SourcePDFScholar
2025

SGC-Net: Stratified Granular Comparison Network for Open-Vocabulary HOI Detection

CVPR 2025poster

Recent open-vocabulary human-object interaction (OV-HOI) detection methods primarily rely on large language model (LLM) for generating auxiliary descriptions and leverage knowledge distilled from CLIP to detect unseen interaction categories. Despite their effectiveness, these methods face two challe…

2024

BDIQA: A New Dataset for Video Question Answering to Explore Cognitive Reasoning through Theory of Mind

AAAI 2024technical

As a foundational component of cognitive intelligence, theory of mind (ToM) can make AI more closely resemble human thought processes, thereby enhancing their interaction and collaboration with human. In particular, it can significantly improve a model's comprehension of videos in complex scenes. Ho…

2024

Decompose, Analyze and Rethink: Solving Intricate Problems with Human-like Reasoning Cycle

NeurIPS 2024oral

In this paper, we introduce DeAR (_Decompose-Analyze-Rethink_), a framework that iteratively builds a reasoning tree to tackle intricate problems within a single large language model (LLM). Unlike approaches that extend or search for rationales, DeAR is featured by 1) adopting a tree-based question…

Cited by 9SourcePDFScholar
2024

Hypernetwork-Assisted Parameter-Efficient Fine-Tuning with Meta-Knowledge Distillation for Domain Knowledge Disentanglement

NAACL 2024findings

Domain adaptation from labeled source domains to the target domain is important in practical summarization scenarios. However, the key challenge is domain knowledge disentanglement. In this work, we explore how to disentangle domain-invariant knowledge from source domains while learning specific kno…

Cited by 1SourcePDFScholar
2024

MNER-MI: A Multi-image Dataset for Multimodal Named Entity Recognition in Social Media

COLING 2024main

Recently, multimodal named entity recognition (MNER) has emerged as a vital research area within named entity recognition. However, current MNER datasets and methods are predominantly based on text and a single accompanying image, leaving a significant research gap in MNER scenarios involving multip…

2024

Restore Anything with Masks: Leveraging Mask Image Modeling for Blind All-in-One Image Restoration

ECCV 2024poster

"All-in-one image restoration aims to handle multiple degradation types using one model. This paper proposes a simple pipeline for all-in-one blind image restoration to Restore Anything with Masks (). We focus on the image content by utilizing Mask Image Modeling to extract intrinsic image informati…

2024

TD²-Net: Toward Denoising and Debiasing for Video Scene Graph Generation

AAAI 2024technical

Dynamic scene graph generation (SGG) focuses on detecting objects in a video and determining their pairwise relationships. Existing dynamic SGG methods usually suffer from several issues, including 1) Contextual noise, as some frames might contain occluded and blurred objects. 2) Label bias, primari…

Cited by 4SourcePDFScholar
2023

A Disentangled-Attention Based Framework with Persona-Aware Prompt Learning for Dialogue Generation

AAAI 2023technical

Endowing dialogue agents with personas is the key to delivering more human-like conversations. However, existing persona-grounded dialogue systems still lack informative details of human conversations and tend to reply with inconsistent and generic responses. One of the main underlying causes is tha…

Cited by 5SourcePDFScholar
2023

CCVO: Cascaded CNNs for Fast Monocular Visual Odometry Towards the Dynamic Environment

RA-L 2023

For AR applications, the present monocular VO methods can be further improved to provide more real-time and accurate self-localization in a dynamic environment with motion disturbance. This letter proposes the CCVO (Cascaded CNNs for Visual Odometry) which is a monocular VO approach to realize end-t

Cited by 10SourceScholar
2023

Pseudo-Query Generation For Semi-Supervised Visual Grounding With Knowledge Distillation

ICASSP 2023accepted

Visual grounding is a crucial multi-modal job for locating the objects that the referring queries refer to in images. In recent years, both fully-supervised and weakly-supervised algorithms rely on a large number of query annotations. However, collecting queries in natural language is labor-intensiv…

Cited by 0SourceScholar
2023

Towards a Holistic Understanding of Mathematical Questions with Contrastive Pre-training

AAAI 2023technical

Understanding mathematical questions effectively is a crucial task, which can benefit many applications, such as difficulty estimation. Researchers have drawn much attention to designing pre-training models for question representations due to the scarcity of human annotations (e.g., labeling difficu…

2023

Unsupervised Image Denoising in Real-World Scenarios via Self-Collaboration Parallel Generative Adversarial Branches

ICCV 2023poster

Deep learning methods have shown remarkable performance in image denoising, particularly when trained on large-scale paired datasets. However, acquiring such paired datasets for real-world scenarios poses a significant challenge. Although unsupervised approaches based on generative adversarial netwo…

Cited by 32PDFcodeScholar
2022

Curriculum Prompt Learning with Self-Training for Abstractive Dialogue Summarization

EMNLP 2022main

Succinctly summarizing dialogue is a task of growing interest, but inherent challenges, such as insufficient training data and low information density impede our ability to train abstractive models. In this work, we propose a novel curriculum-based prompt learning method with self-training to addres…

2022

HL-Net: Heterophily Learning Network for Scene Graph Generation

CVPR 2022poster

Scene graph generation (SGG) aims to detect objects and predict their pairwise relationships within an image. Current SGG methods typically utilize graph neural networks (GNNs) to acquire context information between objects/relationships. Despite their effectiveness, however, current SGG methods onl…

Cited by 68PDFcodeScholar
2022

RU-Net: Regularized Unrolling Network for Scene Graph Generation

CVPR 2022poster

Scene graph generation (SGG) aims to detect objects and predict the relationships between each pair of objects. Existing SGG methods usually suffer from several issues, including 1) ambiguous object representations, as graph neural network-based message passing (GMP) modules are typically sensitive…

Cited by 52PDFcodeScholar
2022

Shifting More Attention to Visual Backbone: Query-Modulated Refinement Networks for End-to-End Visual Grounding

CVPR 2022poster

Visual grounding focuses on establishing fine-grained alignment between vision and natural language, which has essential applications in multimodal reasoning systems. Existing methods use pre-trained query-agnostic visual backbones to extract visual feature maps independently without considering the…

Cited by 88PDFcodeScholar
2021

Cross-Modal Knowledge Distillation For Fine-Grained One-Shot Classification

ICASSP 2021accepted

Few-shot learning can recognize a novel category based on only a few samples because it learns to learn from a lot of labeled samples during the training process. When data is insufficient, the performance is affected. And it is expensive to obtain a large-scale finegrained dataset with annotation.…

Cited by 0SourceScholar
2021

HMS: A Hierarchical Solver with Dependency-Enhanced Understanding for Math Word Problem

AAAI 2021technical

Automatically solving math word problems is a crucial task for exploring the intelligence levels of machines in the general AI domain. It is highly challenging since it requires not only natural language understanding but also mathematical expression inference. Existing solutions usually explore seq…

2021

Looking Wider for Better Adaptive Representation in Few-Shot Learning

AAAI 2021technical

Building a good feature space is essential for the metric-based few-shot algorithms to recognize a novel class with only a few samples. The feature space is often built by Convolutional Neural Networks (CNNs). However, CNNs primarily focus on local information with the limited receptive field, and t…

Cited by 58SourcePDFScholar
2021

Wheel-Legged Robotic Limb to Assist Human With Load Carriage: An Application For Environmental Disinfection During COVID-19

RA-L 2021

During COVID-19, with a heavy sprayer filled with disinfectant, the risk of infection for epidemic prevention personnel has been increased by long-term environmental disinfection. In order to reduce the burden and save energy of human, this letter proposed a Wheel-Legged Robotic Limb (WRL) for the c

Cited by 20SourceScholar
2020

GPS-Net: Graph Property Sensing Network for Scene Graph Generation

CVPR 2020oral

Scene graph generation (SGG) aims to detect objects in an image along with their pairwise relationships. There are three key properties of scene graph that have been underexplored in recent works: namely, the edge direction information, the difference in priority between nodes, and the long-tailed d…

Cited by 321PDFcodeScholar
2016

A Convex Atomic-Norm Approach to Multiple Sequence Alignment and Motif Discovery

ICML 2016poster

Multiple Sequence Alignment and Motif Discovery, known as NP-hard problems, are two fundamental tasks in Bioinformatics. Existing approaches to these two problems are based on either local search methods such as Expectation Maximization (EM), Gibbs Sampling or greedy heuristic methods. In this work,…

Cited by 14SourcePDFScholar
2016

Iteratively reweighted tensor SVD for robust multi-dimensional harmonic retrieval

ICASSP 2016accepted

In this paper, parameter estimation for multi-dimensional sinusoids in additive impulsive noise is addressed. Our underlying idea is to minimize the ℓ <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">p</sub> -norm of the residual error tensor, where 1 <;…

Cited by 6SourceScholar
2015

A Convex Exemplar-based Approach to MAD-Bayes Dirichlet Process Mixture Models

ICML 2015poster

MAD-Bayes (MAP-based Asymptotic Derivations) has been recently proposed as a general technique to derive scalable algorithm for Bayesian Nonparametric models. However, the combinatorial nature of objective functions derived from MAD-Bayes results in hard optimization problem, for which current pract…

Cited by 8SourcePDFScholar