← Search

Hongyang ZHANG

40 accepted papers

2026

InfoGeo: Information-Theoretic Object-Centric Learning for Cross-View Generalizable UAV Geo-Localization

ICML 2026poster

Cross-view geo-localization (CVGL) is fundamental for precise navigation in GPS-denied environments, aiming to match ground or UAV imagery with satellite views. While existing approaches rely on global feature alignment, they often suffer from substantial domain shifts induced by varying regional te…

Cited by 0SourceScholar
2026

Position-Aware Self-supervised Representation Learning for Cross-mode Radar Signal Recognition

ICASSP 2026poster

Radar signal recognition in open electromagnetic environments is challenging due to diverse operating modes and unseen radar types. Existing methods often overlook position relations in pulse sequences, limiting their ability to capture semantic dependencies over time. We propose RadarPos, a positio…

Cited by 0SourcePDFScholar
2026

Spatia: Video Generation with Updatable Spatial Memory

CVPR 2026

Existing video generation models struggle to maintain long-term spatial and temporal consistency due to the dense, high-dimensional nature of video signals. To overcome this limitation, we propose Spatia, a spatial memory-aware video generation framework that explicitly preserves a 3D scene point cl

Cited by 0SourcecodeScholar
2026

TrustGen: A Platform of Dynamic Benchmarking on the Trustworthiness of Generative Foundation Models

ICLR 2026poster

Generative foundation models (GenFMs), such as large language models and text-to-image systems, have demonstrated remarkable capabilities in various downstream applications. As they are increasingly deployed in high-stakes applications, assessing their trustworthiness has become both a critical nece…

Cited by 0SourceScholar
2026

WinQ: Accelerating Quantization-Aware Training of Large Language Models around Saddle Points

ICML 2026poster

Quantization-aware training is widely used for language model quantization in sub-4-bit precision, by training full-precision weights with gradients computed on the quantized model. The main bottleneck for this training approach is its slow convergence and plateauing of test performance, which gets …

Cited by 0SourceScholar
2025

BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs

CVPR 2025poster

Advancements in large Vision-Language Models have brought precise, accurate image captioning, vital for advancing multi-modal image understanding and processing. Yet these captions often carry lengthy, intertwined contexts that are difficult to parse and frequently overlook essential cues, posing a…

Cited by 0SourcePDFScholar
2025

EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test

NeurIPS 2025poster

The sequential nature of modern LLMs makes them expensive and slow, and speculative sam- pling has proven to be an effective solution to this problem. Methods like EAGLE perform autoregression at the feature level, reusing top- layer features from the target model to achieve better results than vani…

Cited by 0SourcecodeScholar
2025

PANDA: Patch-Aware Graph Network with Dual Alignment for Time Series Forecasting

ICASSP 2025accepted

Multivariate time series (MTS) forecasting aims to predict future patterns by extracting features from multivariate history. Predominant methods face challenges in learning spatial dependencies while capturing long-term trends and local details, leading to suboptimal performance in MTS forecasting.…

Cited by 0SourceScholar
2025

The Matrix: Infinite-Horizon World Generation with Real-Time Moving Control

NeurIPS 2025poster

We present The Matrix, a foundational realistic world simulator capable of generating infinitely long 720p high-fidelity real-scene video streams with real-time, responsive control in both first- and third-person perspectives. Trained on limited supervised data from video games like Forza Horizon 5…

Cited by 0SourceScholar
2024

A Large-Scale Human-Centric Benchmark for Referring Expression Comprehension in the LMM Era

NeurIPS 2024poster

Prior research in human-centric AI has primarily addressed single-modality tasks like pedestrian detection, action recognition, and pose estimation. However, the emergence of large multimodal models (LMMs) such as GPT-4V has redirected attention towards integrating language with visual content. Refe…

2024

A Resilient and Accessible Distribution-Preserving Watermark for Large Language Models

ICML 2024poster

Watermarking techniques offer a promising way to identify machine-generated content via embedding covert information into the contents generated from language models. A challenge in the domain lies in preserving the distribution of original generated content after watermarking. Our research extends…

2024

AnyTool: Self-Reflective, Hierarchical Agents for Large-Scale API Calls

ICML 2024poster

We introduce AnyTool, a large language model agent designed to revolutionize the utilization of a vast array of tools in addressing user queries. We utilize over 16,000 APIs from Rapid API, operating under the assumption that a subset of these APIs could potentially resolve the queries. AnyTool prim…

2024

EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees

EMNLP 2024main

Inference with modern Large Language Models (LLMs) is expensive and time-consuming, and speculative sampling has proven to be an effective solution. Most speculative sampling methods such as EAGLE use a static draft tree, implicitly assuming that the acceptance rate of draft tokens depends only on t…

2024

EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty

ICML 2024poster

Autoregressive decoding makes the inference of Large Language Models (LLMs) time-consuming. In this paper, we reconsider speculative sampling and derive two key observations. Firstly, autoregression at the feature (second-to-top-layer) level is more straightforward than at the token level. Secondly,…

2024

Lost Domain Generalization Is a Natural Consequence of Lack of Training Domains

AAAI 2024technical

We show a hardness result for the number of training domains required to achieve a small population error in the test domain. Although many domain generalization algorithms have been developed under various domain-invariance assumptions, there is significant evidence to indicate that out-of-distribu…

Cited by 3SourcePDFScholar
2024

Modelling and Analysis of Joint-to-End Variable Stiffness for Cable-Driven Hyper-Redundant Manipulator

IROS 2024poster

To ensure operational accuracy and flexibility in confined environments, the cable-driven hyper-redundant manipulator needs to take into account both compliance and stiffness. Although the cable-driven method enables the manipulator to have adjustable stiffness, the theoretical analyses and studies…

Cited by 0SourceScholar
2024

RAIN: Your Language Models Can Align Themselves without Finetuning

ICLR 2024poster

Large language models (LLMs) often demonstrate inconsistencies with human preferences. Previous research typically gathered human preference data and then aligned the pre-trained models using reinforcement learning or instruction tuning, a.k.a. the finetuning step. In contrast, aligning frozen LLMs…

2024

Unbiased Watermark for Large Language Models

ICLR 2024spotlight

The recent advancements in large language models (LLMs) have sparked a growing apprehension regarding the potential misuse. One approach to mitigating this risk is to incorporate watermarking techniques into LLMs, allowing for the tracking and attribution of model outputs. This study examines a cruc…

Cited by 129SourcePDFScholar
2023

Causal Balancing for Domain Generalization

ICLR 2023poster

While machine learning models rapidly advance the state-of-the-art on various real-world tasks, out-of-domain (OOD) generalization remains a challenging problem given the vulnerability of these models to spurious correlations. We propose a balanced mini-batch sampling strategy to transform a biased…

2023

Cooperation or Competition: Avoiding Player Domination for Multi-Target Robustness via Adaptive Budgets

CVPR 2023poster

Despite incredible advances, deep learning has been shown to be susceptible to adversarial attacks. Numerous approaches were proposed to train robust networks both empirically and certifiably. However, most of them defend against only a single type of attack, while recent work steps forward at defen…

Cited by 2SourcePDFScholar
2023

Nash Equilibria and Pitfalls of Adversarial Training in Adversarial Robustness Games

AISTATS 2023poster

Adversarial training is a standard technique for training adversarially robust models. In this paper, we study adversarial training as an alternating best-response strategy in a 2-player zero-sum game. We prove that even in a simple scenario of a linear classifier and a statistical model that abstra…

Cited by 12SourcePDFScholar
2023

Understanding the Impact of Adversarial Robustness on Accuracy Disparity

ICML 2023poster

While it has long been empirically observed that adversarial robustness may be at odds with standard accuracy and may have further disparate impacts on different classes, it remains an open question to what extent such observations hold and how the class imbalance plays a role within. In this paper,…

2022

Boosting Barely Robust Learners: A New Perspective on Adversarial Robustness

NeurIPS 2022accept

We present an oracle-efficient algorithm for boosting the adversarial robustness of barely robust learners. Barely robust learning algorithms learn predictors that are adversarially robust only on a small fraction $\beta \ll 1$ of the data distribution. Our proposed notion of barely robust learning…

Cited by 3SourcePDFScholar
2022

Building Robust Ensembles via Margin Boosting

ICML 2022spotlight

In the context of adversarial robustness, a single model does not usually have enough power to defend against all possible adversarial attacks, and as a result, has sub-optimal robustness. Consequently, an emerging line of work has focused on learning an ensemble of neural networks to defend against…

2022

Certified Error Control of Candidate Set Pruning for Two-Stage Relevance Ranking

EMNLP 2022main

In information retrieval (IR), candidate set pruning has been commonly used to speed up two-stage relevance ranking. However, such an approach lacks accurate error control and often trades accuracy against computational efficiency in an empirical fashion, missing theoretical guarantees. In this pape…

2020

A Closer Look at Accuracy vs. Robustness

NeurIPS 2020poster

Current methods for training robust networks lead to a drop in test accuracy, which has led prior works to posit that a robustness-accuracy tradeoff may be inevitable in deep learning. We take a closer look at this phenomenon and first show that real image datasets are actually separated. With this…

2020

Design and Interpretation of Universal Adversarial Patches in Face Detection

ECCV 2020poster

We consider universal adversarial patches for faces --- small visual elements whose addition to a face image reliably destroys the performance of face detectors. Unlike previous work that mostly focused on the algorithmic design of adversarial examples in terms of improving the success rate as an at…

Cited by 51SourcePDFScholar
2020

On the Generalization Effects of Linear Transformations in Data Augmentation

ICML 2020poster

Data augmentation is a powerful technique to improve performance in applications such as image and text classification tasks. Yet, there is little rigorous understanding of why and how various augmentations work. In this work, we consider a family of linear transformations and study their effects on…

2019

Deep Neural Networks with Multi-Branch Architectures Are Intrinsically Less Non-Convex

AISTATS 2019poster

Several recently proposed architectures of neural networks such as ResNeXt, Inception, Xception, SqueezeNet and Wide ResNet are based on the designing idea of having multiple branches and have demonstrated improved performance in many applications. We show that one cause for such success is due to t…

Cited by 48SourcePDFScholar
2019

Efficient Symmetric Norm Regression via Linear Sketching

NeurIPS 2019poster

We provide efficient algorithms for overconstrained linear regression problems with size $n \times d$ when the loss function is a symmetric norm (a norm invariant under sign-flips and coordinate-permutations). An important class of symmetric norms are Orlicz norms, where for a function $G$ and a ve…

Cited by 29SourcePDFScholar
2019

Optimal Analysis of Subset-Selection Based L_p Low-Rank Approximation

NeurIPS 2019poster

We show that for the problem of $\ell_p$ rank-$k$ approximation of any given matrix over $R^{n\times m}$ and $C^{n\times m}$, the algorithm of column subset selection enjoys approximation ratio $(k+1)^{1/p}$ for $1\le p\le 2$ and $(k+1)^{1-1/p}$ for $p\ge 2$. This improves upon the previous $O(k+1)$…

Cited by 21SourcePDFScholar
2019

Recovery Guarantees For Quadratic Tensors With Sparse Observations

AISTATS 2019poster

We consider the tensor completion problem of predicting the missing entries of a tensor. The commonly used CP model has a triple product form, but an alternate family of quadratic models which are the sum of pairwise products instead of a triple product have emerged from applications such as recomme…

Cited by 3SourcePDFScholar
2019

Theoretically Principled Trade-off between Robustness and Accuracy

ICML 2019oral

We identify a trade-off between robustness and accuracy that serves as a guiding principle in the design of defenses against adversarial examples. Although this problem has been widely studied empirically, much remains unknown concerning the theory underlying this trade-off. In this work, we decompo…

2017

Differentially Private Clustering in High-Dimensional Euclidean Spaces

ICML 2017poster

We study the problem of clustering sensitive data while preserving the privacy of individuals represented in the dataset, which has broad applications in practical machine learning and data analysis tasks. Although the problem has been widely studied in the context of low-dimensional, discrete space…

Cited by 106SourcePDFScholar
2017

Noise-Tolerant Interactive Learning Using Pairwise Comparisons

NeurIPS 2017poster

We study the problem of interactively learning a binary classifier using noisy labeling and pairwise comparison oracles, where the comparison oracle answers which one in the given two instances is more likely to be positive. Learning from such oracles has multiple applications where obtaining direct…

Cited by 43SourcePDFScholar
2017

Sample and Computationally Efficient Learning Algorithms under S-Concave Distributions

NeurIPS 2017poster

We provide new results for noise-tolerant and sample-efficient learning algorithms under $s$-concave distributions. The new class of $s$-concave distributions is a broad and natural generalization of log-concavity, and includes many important additional distributions, e.g., the Pareto distribution a…

Cited by 39SourcePDFScholar