← Search

Jiaming Zhang

52 accepted papers

2026

Algorithmic Recourse of In-Context Learning for Tabular Data

ICML 2026poster

As predictive models are increasingly deployed in high-stakes settings such as credit approval, there is a growing need for post-hoc methods that provide recourse to affected individuals. Many such models operate on tabular data, where features correspond to real-world attributes. Recently, in-conte…

Cited by 0SourceScholar
2026

Benign Overfitting in Adversarial Training for Vision Transformers

ICML 2026poster

Despite the remarkable success of Vision Transformers (ViTs) across a wide range of vision tasks, recent studies have revealed that they remain vulnerable to adversarial examples, much like Convolutional Neural Networks (CNNs). A common empirical defense strategy is adversarial training, yet the the…

Cited by 1SourceScholar
2026

Demystifying the Optimal Fair Classifier in Multi-Class Classification

ICML 2026poster

Ensuring fair and equitable treatment across diverse groups, particularly in multi-class classification tasks, poses a significant challenge due to the persistent biases inherent in machine learning models. Most existing bias mitigation techniques are tailored to binary settings, and the presence of…

Cited by 0SourceScholar
2026

Disrupting Hierarchical Reasoning: Adversarial Protection for Geographic Privacy in Multimodal Reasoning Models

ICLR 2026poster

Multi-modal large reasoning models (MLRMs) pose significant privacy risks by inferring precise geographic locations from personal images through hierarchical chain-of-thought reasoning. Existing privacy protection techniques, primarily designed for perception-based models, prove ineffective against…

Cited by 0SourceScholar
2026

GAHMN: A Generative Approach for High-Dimensional Mediation Analysis

AAAI 2026technical

High-dimensional mediation analysis (HMA) seeks to uncover complex causal mechanisms involving numerous mediators and plays a crucial role in scientific and social sciences. In this work, we introduce the Generative Adversarial High-dimensional Mediation Network (GAHMN), a novel, scalable structured

Cited by 0SourcePDFScholar
2026

Hallucinating 360°: Panoramic Street-View Generation Via Local Scenes Diffusion and Probabilistic Prompting

ICRA 2026poster

Panoramic perception holds significant potential for autonomous driving, enabling vehicles to acquire a comprehensive 360° surround view in a single shot. However, autonomous driving is a data-driven task. Complete panoramic data acquisition requires complex sampling systems and annotation pipelines…

2026

HybriDLA: Hybrid Generation for Document Layout Analysis

AAAI 2026technical

Conventional document layout analysis (DLA) traditionally depends on empirical priors or a fixed set of learnable queries executed in a single forward pass. While sufficient for early-generation documents with a small, predetermined number of regions, this paradigm struggles with contemporary docume

Cited by 0SourcePDFScholar
2026

Leveraging Machine Unlearning for Cost-Efficient Preference Alignment

ICML 2026poster

Despite advances in Preference Alignment (PA) for Large Language Models (LLMs), mainstream methods like reinforcement learning with human feedback face notable challenges. These approaches require high-quality datasets of positive preference examples, which are costly to obtain and computationally i…

Cited by 0SourceScholar
2026

More than the Sum: Panorama-Language Models for Adverse Omni-Scenes

CVPR 2026

Existing vision-language models (VLMs) are tailored for pinhole imagery, stitching multiple narrow field-of-view inputs to piece together a complete omni-scene understanding. Yet, such multi-view perception overlooks the holistic spatial and contextual relationships that a single panorama inherently

Cited by 0SourcecodeScholar
2026

MultiPriv: Benchmarking Individual-Level Privacy Reasoning in Vision-Language Models

ICML 2026poster

Modern Vision-Language Models (VLMs) pose significant individual-level privacy risks by linking fragmented multimodal data to identifiable individuals through hierarchical chain-of-thought reasoning. However, existing privacy benchmarks remain structurally insufficient for this threat, as they prima…

Cited by 0SourceScholar
2026

RHO: Robust Holistic OSM-Based Metric Cross-View Geo-Localization

CVPR 2026

Metric Cross-View Geo-Localization (MCVGL) aims to estimate the 3-DoF camera pose (position and heading) by matching ground and satellite images. In this work, instead of pinhole and satellite images, we study robust MCVGL using holistic panoramas and OpenStreetMap (OSM). To this end, we establish a

Cited by 0SourcecodeScholar
2026

SF-ODNav: Successor Feature Framework for Map-Less Target-Driven Outdoor Visual Navigation

ICRA 2026poster

Traditional deep reinforcement learning-based visual navigation techniques face challenges in dynamic and unstructured outdoor environments, particularly in the absence of high-resolution maps and GPS signals. This paper presents a deep reinforcement learning-based approach for target-driven visual …

Cited by 0Scholar
2026

SubspacePath Pruner: Inference-time Pruning via Probe-based Representation–Parameter Coupling

ICML 2026poster

Large-scale dedicated application of LLMs in diverse scenarios increasingly demands specialized model inference behavior under strict constraints of accuracy, latency, and memory. However, the heterogeneous and long-tailed nature of real-world specialized scenarios makes it difficult to obtain train…

Cited by 0SourceScholar
2026

TOFA: Training-Free One-Shot Federated Adaptation for Vision-Language Models

AAAI 2026technical

Efficient and lightweight adaptation of pre-trained Vision-Language Models (VLMs) to downstream tasks through collaborative interactions between local clients and a central server is a rapidly emerging research topic in federated learning. Existing adaptation algorithms are typically trained iterati

Cited by 0SourcePDFScholar
2026

VENOMREC: Cross-Modal Interactive Poisoning for Targeted Promotion in Multimodal LLM Recommender Systems

ICML 2026poster

Multimodal large language models (MLLMs) are pushing recommender systems (RecSys) toward content-grounded retrieval and ranking via cross-modal fusion. We find that while cross-modal consensus often mitigates conventional poisoning that manipulates interaction logs or perturbs a single modality, it …

Cited by 0SourceScholar
2026

XDomainBench: Diagnosing Reasoning Collapse in High-Dimensional Scientific Knowledge Composition

ICML 2026poster

Large Language Models (LLMs) are increasingly deployed for knowledge synthesis, yet their capacity for compositional generalization in scientific knowledge remains under-characterized. Existing benchmarks primarily focus on single-turn restricted scenarios, failing to capture the capability boundari…

Cited by 0SourceScholar
2025

An Image-Guided Robotic System for Transcranial Magnetic Stimulation: System Development and Experimental Evaluation

RA-L 2025

Transcranial magnetic stimulation is a noninvasive medical procedure that can modulate brain activity, and it is widely used in neuroscience, neurology research, and clinical practice. Compared to manual operators, robots may improve the outcome due to their superior accuracy and repeatability. Howe

Cited by 1SourceScholar
2025

Anyattack: Towards Large-scale Self-supervised Adversarial Attacks on Vision-language Models

CVPR 2025poster

Due to their multimodal capabilities, Vision-Language Models (VLMs) have found numerous impactful applications in real-world scenarios. However, recent studies have revealed that VLMs are vulnerable to image-based adversarial attacks. Traditional targeted adversarial attacks require specific targets…

Cited by 0SourcePDFScholar
2025

FedFACT: A Provable Framework for Controllable Group-Fairness Calibration in Federated Learning

NeurIPS 2025poster

With emerging application of Federated Learning (FL) in decision-making scenarios, it is imperative to regulate model fairness to prevent disparities across sensitive groups (e.g., female, male). Current research predominantly focuses on two concepts of group fairness within FL: *Global Fairness* (o…

Cited by 0SourceScholar
2025

Graph-based Document Structure Analysis

ICLR 2025poster

When reading a document, glancing at the spatial layout of a document is an initial step to understand it roughly. Traditional document layout analysis (DLA) methods, however, offer only a superficial parsing of documents, focusing on basic instance detection and often failing to capture the nuanced…

Cited by 0SourcePDFScholar
2025

Investigating and Enhancing Vision-Audio Capability in Omnimodal Large Language Models

ACL 2025finding

Omnimodal Large Language Models (OLLMs) have shown significant progress in integrating vision and text, but still struggle with integrating vision and audio, often exhibiting suboptimal performance when processing audio queries compared to text queries. This disparity is primarily due to insufficien…

2025

SAMBLE: Shape-Specific Point Cloud Sampling for an Optimal Trade-Off Between Local Detail and Global Uniformity

CVPR 2025poster

Driven by the increasing demand for accurate and efficient representation of 3D data in various domains, point cloud sampling has emerged as a pivotal research topic in 3D computer vision. Recently, learning-to-sample methods have garnered growing interest from the community, particularly for their…

Cited by 0SourcePDFScholar
2025

Scene-agnostic Pose Regression for Visual Localization

CVPR 2025poster

Absolute Pose Regression (APR) predicts 6D camera poses but lacks the adaptability to unknown environments without retraining, while Relative Pose Regression (RPR) generalizes better yet requires a large image retrieval database. Visual Odometry (VO) generalizes well in unseen environments but suffe…

Cited by 0SourcePDFScholar
2025

Situat3DChange: Situated 3D Change Understanding Dataset for Multimodal Large Language Model

NeurIPS 2025poster

Physical environments and circumstances are fundamentally dynamic, yet current 3D datasets and evaluation benchmarks tend to concentrate on either dynamic scenarios or dynamic situations in isolation, resulting in incomplete comprehension. To overcome these constraints, we introduce Situat3DChange,…

Cited by 0SourcecodeScholar
2025

Surge: On the Potential of Large Language Models as General-Purpose Surrogate Code Executors

EMNLP 2025

Neural surrogate models are powerful and efficient tools in data mining. Meanwhile, large language models (LLMs) have demonstrated remarkable capabilities in code-related tasks, such as generation and understanding. However, an equally important yet underexplored question is whether LLMs can serve a

2025

TAPT: Test-Time Adversarial Prompt Tuning for Robust Inference in Vision-Language Models

CVPR 2025poster

Large pre-trained Vision-Language Models (VLMs) such as CLIP have demonstrated excellent zero-shot generalizability across various downstream tasks. However, recent studies have shown that the inference performance of CLIP can be greatly degraded by small adversarial perturbations, especially its vi…

2025

Unlocking Constraints: Source-Free Occlusion-Aware Seamless Segmentation

ICCV 2025poster

Panoramic image processing is essential for omni-context perception, yet faces constraints like distortions, perspective occlusions, and limited annotations. Previous unsupervised domain adaptation methods transfer knowledge from labeled pinhole data to unlabeled panoramic images, but they require a…

2025

dARt Vinci: Egocentric Data Collection for Surgical Robot Learning at Scale

IROS 2025

Data scarcity has long been an issue in the robot learning community. Particularly, in safety-critical domains like surgical applications, obtaining high-quality data can be especially difficult. It poses challenges to researchers seeking to exploit recent advancements in reinforcement learning and

Cited by 4SourceScholar
2025

mmWalk: Towards Multi-modal Multi-view Walking Assistance

NeurIPS 2025poster

Walking assistance in extreme or complex environments remains a significant challenge for people with blindness or low vision (BLV), largely due to the lack of a holistic scene understanding. Motivated by the real-world needs of the BLV community, we build mmWalk, a simulated multi-modal dataset tha…

Cited by 0SourceScholar
2024

Adversarial Prompt Tuning for Vision-Language Models

ECCV 2024poster

"With the rapid advancement of multimodal learning, pre-trained Vision-Language Models (VLMs) such as CLIP have demonstrated remarkable capacities in bridging the gap between visual and language modalities. However, these models remain vulnerable to adversarial attacks, particularly in the image mod…

2024

CURE4Rec: A Benchmark for Recommendation Unlearning with Deeper Influence

NeurIPS 2024poster

With increasing privacy concerns in artificial intelligence, regulations have mandated the right to be forgotten, granting individuals the right to withdraw their data from models. Machine unlearning has emerged as a potential solution to enable selective forgetting in models, particularly in recomm…

2024

Elevating Skeleton-Based Action Recognition with Efficient Multi-Modality Self-Supervision

ICASSP 2024accepted

Self-supervised representation learning for human action recognition has developed rapidly in recent years. Most of the existing works are based on skeleton data while using a multi-modality setup. These works overlooked the differences in performance among modalities, which led to the propagation o…

Cited by 0SourceScholar
2024

MateRobot: Material Recognition in Wearable Robotics for People with Visual Impairments

ICRA 2024poster

People with Visual Impairments (PVI) typically recognize objects through haptic perception. Knowing objects and materials before touching is desired by the target users but under-explored in the field of human-centered robotics. To fill this gap, in this work, a wearable vision-based robotic system,…

Cited by 13SourcecodeScholar
2024

Navigating Open Set Scenarios for Skeleton-Based Action Recognition

AAAI 2024technical

In real-world scenarios, human actions often fall outside the distribution of training data, making it crucial for models to recognize known actions and reject unknown ones. However, using pure skeleton data in such open-set conditions poses challenges due to the lack of visual background cues and t…

2024

Occlusion-Aware Seamless Segmentation

ECCV 2024poster

"Panoramic images can broaden the Field of View (FoV), occlusion-aware prediction can deepen the understanding of the scene, and domain adaptation can transfer across viewing domains. In this work, we introduce a novel task, Occlusion-Aware Seamless Segmentation (OASS), which simultaneously tackles…

2024

One for All: A Universal Generator for Concept Unlearnability via Multi-Modal Alignment

ICML 2024poster

The abundance of free internet data offers unprecedented opportunities for researchers and developers, but it also poses privacy risks. Utilizing data without explicit consent raises critical challenges in protecting personal information.Unlearnable examples have emerged as a feasible protection app…

Cited by 3SourcePDFScholar
2024

Realtime Robust Shape Estimation of Deformable Linear Object

ICRA 2024poster

Realtime shape estimation of continuum objects and manipulators is essential for developing accurate planning and control paradigms. The existing methods that create dense point clouds from camera images, and/or use distinguishable markers on a deformable body have limitations in realtime tracking o…

Cited by 1SourceScholar
2024

Referring Atomic Video Action Recognition

ECCV 2024poster

"We introduce a new task called Referring Atomic Video Action Recognition (RAVAR), aimed at identifying atomic actions of a particular person based on a textual description and the video data of this person. This task differs from traditional action recognition and localization, where predictions ar…

2024

RoDLA: Benchmarking the Robustness of Document Layout Analysis Models

CVPR 2024poster

Before developing a Document Layout Analysis (DLA) model in real-world applications conducting comprehensive robustness testing is essential. However the robustness of DLA models remains underexplored in the literature. To address this we are the first to introduce a robustness benchmark for DLA mod…

Cited by 6SourcePDFScholar
2024

Skeleton-Based Human Action Recognition with Noisy Labels

IROS 2024poster

Understanding human actions from body poses is critical for assistive robots sharing space with humans in order to make informed and safe decisions about the next interaction. However, precise temporal localization and annotation of activity sequences is time-consuming and the resulting labels are o…

Cited by 5SourcecodeScholar
2023

Bi-Mapper: Holistic BEV Semantic Mapping for Autonomous Driving

RA-L 2023

A semantic map of the road scene, covering fundamental road elements, is an essential ingredient in autonomous driving systems. It provides important perception foundations for positioning and planning when rendered in the Bird's-Eye-View (BEV). Currently, the prior knowledge of hypothetical depth c

Cited by 24SourcecodeScholar
2023

Delivering Arbitrary-Modal Semantic Segmentation

CVPR 2023poster

Multimodal fusion can make semantic segmentation more robust. However, fusing an arbitrary number of modalities remains underexplored. To delve into this problem, we create the DeLiVER arbitrary-modal segmentation benchmark, covering Depth, LiDAR, multiple Views, Events, and RGB. Aside from this, we…

2023

ImageNet Pre-training Also Transfers Non-robustness

AAAI 2023technical

ImageNet pre-training has enabled state-of-the-art results on many tasks. In spite of its recognized contribution to generalization, we observed in this study that ImageNet pre-training also transfers adversarial non-robustness from pre-trained model into fine-tuned model in the downstream classific…

2023

Unlearnable Clusters: Towards Label-Agnostic Unlearnable Examples

CVPR 2023poster

There is a growing interest in developing unlearnable examples (UEs) against visual privacy leaks on the Internet. UEs are training samples added with invisible but unlearnable noise, which have been found can prevent unauthorized training of machine learning models. UEs typically are generated via…

2022

Bending Reality: Distortion-Aware Transformers for Adapting to Panoramic Semantic Segmentation

CVPR 2022poster

Panoramic images with their 360deg directional view encompass exhaustive information about the surrounding space, providing a rich foundation for scene understanding. To unfold this potential in the form of robust panoramic segmentation models, large quantities of expensive, pixel-wise annotations a…

Cited by 107PDFcodeScholar
2022

CPQNet: Contact Points Quality Network for Robotic Grasping

IROS 2022poster

In typical data-based grasping methods, a grasp based on parallel-jaw grippers is parameterized by the center of the gripper, the rotation angle, and the gripper opening width so as to predict the quality and pose of grasps at every pixel. In contrast, a grasp is represented using only two contact p…

Cited by 1SourceScholar
2022

TransDARC: Transformer-based Driver Activity Recognition with Latent Space Feature Calibration

IROS 2022poster

Traditional video-based human activity recognition has experienced remarkable progress linked to the rise of deep learning, but this effect was slower as it comes to the downstream task of driver behavior understanding. Understanding the situation inside the vehicle cabin is essential for Advanced D…

Cited by 39SourcecodeScholar
2021

Capturing Omni-Range Context for Omnidirectional Segmentation

CVPR 2021poster

Convolutional Networks (ConvNets) excel at semantic segmentation and have become a vital component for perception in autonomous driving. Enabling an all-encompassing view of street-scenes, omnidirectional cameras present themselves as a perfect fit in such systems. Most segmentation models for parsi…

Cited by 91PDFcodeScholar
2021

ISSAFE: Improving Semantic Segmentation in Accidents by Fusing Event-based Data

IROS 2021poster

Ensuring the safety of all traffic participants is a prerequisite for bringing intelligent vehicles closer to practical applications. The assistance system should not only achieve high accuracy under normal conditions, but obtain robust perception against extreme situations. However, traffic acciden…

Cited by 60SourcecodeScholar