← Search

Yulong Wang

26 accepted papers

2026

DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object Retrieval

CVPR 2026

Vision foundation models have shown great promise for open-set 3D object retrieval (3DOR) through efficient adaptation to multi-view images. Leveraging semantically aligned latent space, previous work typically adapts the CLIP encoder to build view-based 3D descriptors. Despite CLIP's strong general

Cited by 0SourceScholar
2026

Hg-I2P: Bridging Modalities for Generalizable Image-to-Point-Cloud Registration via Heterogeneous Graphs

CVPR 2026

Image-to-point-cloud (I2P) registration aims to align 2D images with 3D point clouds by establishing reliable 2D-3D correspondences. The drastic modality gap between images and point clouds makes it challenging to learn features that are both discriminative and generalizable, leading to severe perfo

Cited by 0SourcecodeScholar
2026

Navigating the Flatlands: Dual Adaptive Sharpness-Aware Minimization for Domain Generalization

ICML 2026poster

Finding flat minima in the loss landscape is a key strategy for Domain Generalization (DG). However, its effectiveness is often limited by two crucial challenges. 1) Domain Shift: Existing methods like Sharpness-Aware Minimization (SAM) apply a uniform optimization strategy across all domains, overl…

Cited by 0SourceScholar
2025

Adversarial Training for Graph Convolutional Networks: Stability and Generalization Analysis

IJCAI 2025

Recently, numerous methods have been proposed to enhance the robustness of the Graph Convolutional Networks (GCNs) for their vulnerability against adversarial attacks. Despite their empirical success, a significant gap remains in understanding GCNs' adversarial robustness from the theoretical perspe

Cited by 0SourcePDFScholar
2025

Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object Retrieval

ICCV 2025poster

Open-set 3D object retrieval (3DOR) is an emerging task aiming to retrieve 3D objects of unseen categories beyond the training set. Existing methods typically utilize all modalities (i.e., voxels, point clouds, multi-view images) and train specific backbones before fusion. However, they still strugg…

2025

DisLoRA: Task-specific Low-Rank Adaptation via Orthogonal Basis from Singular Value Decomposition

EMNLP 2025

Parameter-efficient fine-tuning (PEFT) of large language models (LLMs) is critical for adapting to diverse downstream tasks with minimal computational cost. We propose **Di**rectional-**S**VD **Lo**w-**R**ank **A**daptation (DisLoRA), a novel PEFT framework that leverages singular value decompositio

Cited by 0SourcePDFScholar
2025

FilterTS: Comprehensive Frequency Filtering for Multivariate Time Series Forecasting

AAAI 2025technical

Multivariate time series forecasting is crucial across various industries, where accurate extraction of complex periodic and trend components can significantly enhance prediction performance. However, existing models often struggle to capture these intricate patterns. To address these challenges, we…

2025

From Individual to Universal: Regularized Multi-view Joint Representation for Multi-view Subspace-Preserving Recovery

IJCAI 2025

Recent years have witnessed an explosion of Multi- view Subspace Classification (MSCla) and Multi-view Subspace Clustering (MSClu) methods for various applications. However, their theoretical foundation have not been well explored and understood. In this paper, we investigate the multi-view subspace

Cited by 0SourcePDFScholar
2025

Lexical Diversity-aware Relevance Assessment for Retrieval-Augmented Generation

ACL 2025long

Retrieval-Augmented Generation (RAG) has proven effective in enhancing the factuality of LLMs’ generation, making them a focal point of research. However, previous RAG approaches overlook the lexical diversity of queries, hindering their ability to achieve a granular relevance assessment between que…

2025

Trajectory-Dependent Generalization Bounds for Pairwise Learning with φ-mixing Samples

IJCAI 2025

Recently, the mathematical tool from fractal geometry (i.e., fractal dimension) has been employed to investigate optimization trajectory-dependent generalization ability for some pointwise learning models with independent and identically distributed (i.i.d.) observations. This paper goes beyond the

Cited by 0SourcePDFScholar
2024

On the Robustness of Editing Large Language Models

EMNLP 2024main

Large language models (LLMs) have played a pivotal role in building communicative AI, yet they encounter the challenge of efficient updates. Model editing enables the manipulation of specific knowledge memories and the behavior of language generation without retraining. However, the robustness of mo…

2024

SocialCircle: Learning the Angle-based Social Interaction Representation for Pedestrian Trajectory Prediction

CVPR 2024poster

Analyzing and forecasting trajectories of agents like pedestrians and cars in complex scenes has become more and more significant in many intelligent systems and applications. The diversity and uncertainty in socially interactive behaviors among a rich variety of agents make this task more challengi…

2024

Superposed Atomic Representation for Robust High-Dimensional Data Recovery of Multiple Low-Dimensional Structures

AAAI 2024technical

This paper proposes a unified Superposed Atomic Representation (SAR) framework for high-dimensional data recovery with multiple low-dimensional structures. The data can be in various forms ranging from vectors to tensors. The goal of SAR is to recover different components from their sum, where each…

Cited by 0SourcePDFScholar
2024

Unsigned Orthogonal Distance Fields: An Accurate Neural Implicit Representation for Diverse 3D Shapes

CVPR 2024poster

Neural implicit representation of geometric shapes has witnessed considerable advancements in recent years. However common distance field based implicit representations specifically signed distance field (SDF) for watertight shapes or unsigned distance field (UDF) for arbitrary shapes routinely suff…

2022

ATF-3D: Semi-Supervised 3D Object Detection With Adaptive Thresholds Filtering Based on Confidence and Distance

RA-L 2022

Performance of current point cloud-based outdoor 3D object detection relies heavily on large-scale high-quality 3D annotations. However, such annotations are usually expensive to collect and outdoor scenes easily accumulate massive unlabeled data containing rich scenes. Semi-supervised learning is a

Cited by 12SourceScholar
2022

Enhancing Multi-modal Features Using Local Self-Attention for 3D Object Detection

ECCV 2022poster

"LiDAR and Camera sensors have complementary properties: LiDAR senses accurate positioning, while camera provides rich texture and color information. Fusing these two modalities can intuitively improve the performance of 3D detection. Most multi-modal fusion methods use networks to extract features…

Cited by 13SourcePDFScholar
2022

Error-Based Knockoffs Inference for Controlled Feature Selection

AAAI 2022technical

Recently, the scheme of model-X knockoffs was proposed as a promising solution to address controlled feature selection under high-dimensional finite-sample settings. However, the procedure of model-X knockoffs depends heavily on the coefficient-based feature importance and only concerns the control…

Cited by 6SourcePDFScholar
2022

Unsupervised Anomaly Detection for Container Cloud Via BILSTM-Based Variational Auto-Encoder

ICASSP 2022accepted

The appearance of container technology has profoundly changed the development and deployment of multi-tier distributed applications. However, the imperfect system resource isolation features and the kernel-sharing mechanism will introduce significant security risks to the container-based cloud. In t…

Cited by 0SourceScholar
2021

Distributed Ranking with Communications: Approximation Analysis and Applications

AAAI 2021technical

Learning theory of distributed algorithms has recently attracted enormous attention in the machine learning community. However, most of existing works focus on learning problem with pointwise loss and does not consider the communication among local processors. In this paper, we propose a new distrib…

Cited by 1SourcePDFScholar
2021

Question-Driven Span Labeling Model for Aspect–Opinion Pair Extraction

AAAI 2021technical

Aspect term extraction and opinion word extraction are two fundamental subtasks of aspect-based sentiment analysis. The internal relationship between aspect terms and opinion words is typically ignored, and information for the decision-making of buyers and sellers is insufficient. In this paper, we…

Cited by 80SourcePDFScholar
2020

6D Object Pose Regression via Supervised Learning on Point Clouds

ICRA 2020poster

This paper addresses the task of estimating the 6 degrees of freedom pose of a known 3D object from depth information represented by a point cloud. Deep features learned by convolutional neural networks from color information have been the dominant features to be used for inferring object poses, whi…

Cited by 111SourcecodeScholar
2018

Interpret Neural Networks by Identifying Critical Data Routing Paths

CVPR 2018poster

Interpretability of a deep neural network aims to explain the rationale behind its decisions and enable the users to understand the intelligent agents, which has become an important issue due to its importance in practical applications. To address this issue, we develop a Distillation Guided Routing…

Cited by 122SourcePDFScholar