← Search

Nanyang Ye

34 accepted papers

2026

Anchor-Final Self-Supervision Drives Hallucination-Aware Optimization in Large Vision-Language Models

ICML 2026poster

Hallucinations in large vision-language models (LVLMs) remain a critical challenge, with models often generate tokens that fail to align with visual evidence. To address this issue, we propose AFS: Anchor-Final Self-Supervision, a novel framework for hallucination-aware optimization in LVLMs. By lev…

Cited by 0SourceScholar
2026

ImpQuant: Fine-Grained Importance-Aware Quantization for Large Vision-Language Models

ICML 2026poster

Large Vision–Language Models (LVLMs) have demonstrated remarkable capabilities across diverse multimodal tasks, yet their high inference costs necessitate low-bit deployment. Existing post-training quantization (PTQ) pipelines primarily adopt methodologies from text-only LLMs by treating multimodal …

Cited by 0SourceScholar
2026

LABO: LLM-Accelerated Bayesian Optimization through Broad Exploration and Selective Experimentation

ICML 2026poster

The high cost and data scarcity in scientific exploration have motivated the use of large language models (LLMs) as knowledge-driven components in Bayesian optimization (BO). However, existing approaches typically embed LLMs directly into the sampling or surrogate modeling pipeline, without fully le…

Cited by 0SourceScholar
2026

Logit-Attention Divergence: Mitigating Position Bias in Multi-Image Retrieval via Attention-Guided Calibration

ICML 2026poster

Multimodal Large Language Models (MLLMs) have shown strong performance in multi-image cross-modal retrieval, yet suffer from severe position bias, where predictions are dominated by input order rather than semantic relevance. Through empirical analysis, we identify a phenomenon termed Logit-Attentio…

Cited by 0SourceScholar
2026

SPUR: Scale-Partitioned Uncertainty Rectification for Robust UAV-on-UAV Interception

ICML 2026poster

Robust aerial target detection for autonomous UAV-on-UAV pursuit is severely hindered by continuous scale drift, long-tailed scale imbalance, and flight-induced visual noise, rendering standard empirical risk minimization strategies poorly aligned with real-world deployment. To address these challen…

Cited by 0SourceScholar
2026

Scaling by Diversified Experience for Vision-Language-Action Models

ICML 2026poster

Vision-Language-Action models face significant challenges in real-world deployment due to the entanglement of high-level reasoning with low-level control, and the instability of policy optimization. In this paper, we introduce SyVLA, a robust VLA model trained with diversified experiences. We propos…

Cited by 0SourceScholar
2026

Unleashing LLMs in Bayesian Optimization: Preference-Guided Framework for Scientific Discovery

ICLR 2026poster

Scientific discovery is increasingly constrained by costly experiments and limited budgets, making efficient optimization essential for AI for science. Bayesian Optimization (BO), while widely adopted for balancing exploration and exploitation, suffers from slow cold-start performance and poor scala…

Cited by 0SourceScholar
2025

$\Delta \mathrm{Energy}$: Optimizing Energy Change During Vision-Language Alignment Improves both OOD Detection and OOD Generalization

NeurIPS 2025poster

Recent approaches for vision-language models (VLMs) have shown remarkable success in achieving fast downstream adaptation. When applied to real-world downstream tasks, VLMs inevitably encounter both the in-distribution (ID) data and out-of-distribution (OOD) data. The OOD datasets often include bot…

Cited by 0SourceScholar
2025

ConcealGS: Concealing Invisible Copyright Information in 3D Gaussian Splatting

ICASSP 2025accepted

As 3D Gaussian Splatting (3D-GS) emerges as a promising technique for 3D reconstruction and novel view synthesis, offering superior rendering quality and efficiency, it becomes crucial to ensure secure transmission and copyright protection of 3D assets in anticipation of widespread distribution. Whi…

Cited by 0SourceScholar
2025

Enhancing Nursing and Elderly Care with Large Language Models: An AI-Driven Framework

COLING 2025main

This paper explores the application of large language models (LLMs) in nursing and elderly care, focusing on AI-driven patient monitoring and interaction. We introduce a novel Chinese nursing dataset and implement incremental pre-training (IPT) and supervised fine-tuning (SFT) techniques to enhance…

Cited by 2SourcePDFScholar
2025

Generalizable Multi-Camera 3D Object Detection from a Single Source via Fourier Cross-View Learning

ICML 2025poster

Improving the generalization of multi-camera 3D object detection is essential for safe autonomous driving in the real world. In this paper, we consider a realistic yet more challenging scenario, which aims to improve the generalization when only single source data available for training, as gatherin…

Cited by 0SourcePDFScholar
2025

Less is More: Masking Elements in Image Condition Features Avoids Content Leakages in Style Transfer Diffusion Models

ICLR 2025poster

Given a style-reference image as the additional image condition, text-to-image diffusion models have demonstrated impressive capabilities in generating images that possess the content of text prompts while adopting the visual style of the reference image. However, current state-of-the-art methods of…

2025

Less is More: an Attention-free Sequence Prediction Modeling for Offline Embodied Learning

NeurIPS 2025poster

Offline reinforcement learning (offline RL) is increasingly approached as a sequence modeling task, with methods leveraging advanced architectures like Transformers to capture trajectory dependencies. Despite significant progress, the mechanisms underlying their effectiveness and limitations remain…

Cited by 0SourcecodeScholar
2025

OODD: Test-time Out-of-Distribution Detection with Dynamic Dictionary

CVPR 2025poster

Out-of-distribution (OOD) detection remains challenging for deep learning models, particularly when test-time OOD samples differ significantly from training outliers. We propose OODD, a novel test-time OOD detection method that dynamically maintains and updates an OOD dictionary without fine-tuning.…

2025

Tri-AutoAug: Single Domain Generalization for Bird's-Eye-View 3D Object Detection Through Pixel-2D-3D Features

ICRA 2025

With the increasing popularity of autonomous driving based on the Bird's-Eye-View (BEV) representation, improving the generalization of such detection models is key for safe real-world applications. However, a realistic yet challenging scenario: Single Domain Generalization (SDG) for BEV, is still u

Cited by 0SourceScholar
2024

CRoFT: Robust Fine-Tuning with Concurrent Optimization for OOD Generalization and Open-Set OOD Detection

ICML 2024poster

Recent vision-language pre-trained models (VL-PTMs) have shown remarkable success in open-vocabulary tasks. However, downstream use cases often involve further fine-tuning of VL-PTMs, which may distort their general knowledge and impair their ability to handle distribution shifts. In real-world scen…

2024

Domain Invariant Learning for Gaussian Processes and Bayesian Exploration

AAAI 2024technical

Out-of-distribution (OOD) generalization has long been a challenging problem that remains largely unsolved. Gaussian processes (GP), as popular probabilistic model classes, especially in the small data regime, presume strong OOD generalization abilities. Surprisingly, their OOD generalization abilit…

2024

G-NAS: Generalizable Neural Architecture Search for Single Domain Generalization Object Detection

AAAI 2024technical

In this paper, we focus on a realistic yet challenging task, Single Domain Generalization Object Detection (S-DGOD), where only one source domain's data can be used for training object detectors, but have to generalize multiple distinct target domains. In S-DGOD, both high-capacity fitting and gener…

2024

MiniConGTS: A Near Ultimate Minimalist Contrastive Grid Tagging Scheme for Aspect Sentiment Triplet Extraction

EMNLP 2024main

Aspect Sentiment Triplet Extraction (ASTE) aims to co-extract the sentiment triplets in a given corpus. Existing approaches within the pretraining-finetuning paradigm tend to either meticulously craft complex tagging schemes and classification heads, or incorporate external semantic augmentation to…

2024

PNAS-MOT: Multi-Modal Object Tracking With Pareto Neural Architecture Search

RA-L 2024

Multiple object tracking is a critical task in autonomous driving. Existing works primarily focus on the heuristic design of neural networks to obtain high accuracy. As tracking accuracy improves, however, neural networks become increasingly complex, posing challenges for their practical application

Cited by 21SourcecodeScholar
2023

Bayesian Cross-Modal Alignment Learning for Few-Shot Out-of-Distribution Generalization

AAAI 2023technical

Recent advances in large pre-trained models showed promising results in few-shot learning. However, their generalization ability on two-dimensional Out-of-Distribution (OoD) data, i.e., correlation shift and diversity shift, has not been thoroughly investigated. Researches have shown that even with…

2023

Certifiable Out-of-Distribution Generalization

AAAI 2023technical

Machine learning methods suffer from test-time performance degeneration when faced with out-of-distribution (OoD) data whose distribution is not necessarily the same as training data distribution. Although a plethora of algorithms have been proposed to mitigate this issue, it has been demonstrated t…

2022

OoD-Bench: Quantifying and Understanding Two Dimensions of Out-of-Distribution Generalization

CVPR 2022oral

Deep learning has achieved tremendous success with independent and identically distributed (i.i.d.) data. However, the performance of neural networks often degenerates drastically when encountering out-of-distribution (OoD) data, i.e., when training and test data are sampled from different distribut…

Cited by 125PDFcodeScholar
2022

OoDHDR-Codec: Out-of-Distribution Generalization for HDR Image Compression

AAAI 2022technical

Recently, deep learning has been proven to be a promising approach in standard dynamic range (SDR) image compression. However, due to the wide luminance distribution of high dynamic range (HDR) images and the lack of large standard datasets, developing a deep model for HDR image compression is much…

2022

Regularization Penalty Optimization for Addressing Data Quality Variance in OoD Algorithms

AAAI 2022technical

Due to the poor generalization performance of traditional empirical risk minimization (ERM) in the case of distributional shift, Out-of-Distribution (OoD) generalization algorithms receive increasing attention. However, OoD generalization algorithms overlook the great variance in the quality of trai…

Cited by 6SourcePDFScholar
2021

Amata: An Annealing Mechanism for Adversarial Training Acceleration

AAAI 2021technical

Despite the empirical success in various domains, it has been revealed that deep neural networks are vulnerable to maliciously perturbed input data that much degrade their performance. This is known as adversarial attacks. To counter adversarial attacks, adversarial training formulated as a form of…

2021

DecAug: Out-of-Distribution Generalization via Decomposed Feature Representation and Semantic Augmentation

AAAI 2021technical

While deep learning demonstrates its strong ability to handle independent and identically distributed (IID) data, it often suffers from out-of-distribution (OoD) generalization, where the test data come from another distribution (w.r.t. the training one). Designing a general OoD generalization frame…

Cited by 86SourcePDFScholar
2021

NAS-OoD: Neural Architecture Search for Out-of-Distribution Generalization

ICCV 2021poster

Recent advances on Out-of-Distribution (OoD) generalization reveal the robustness of deep learning models against distribution shifts. However, existing works focus on OoD algorithms, such as invariant risk minimization, domain generalization, or stable learning, without considering the influence of…

Cited by 56PDFScholar
2021

Semi-supervised Vein Segmentation of Ultrasound Images for Autonomous Venipuncture

IROS 2021poster

Venipuncture is an indispensable procedure for both diagnosis and treatment. In this paper, unlike existing solutions that fully or partially rely on professional assistance, a compact robotic system integrating both novel hardware and software developments is introduced. The hardware consists of a…

Cited by 7SourceScholar
2019

Predicting Visible Image Differences Under Varying Display Brightness and Viewing Distance

CVPR 2019poster

Numerous applications require a robust metric that can predict whether image differences are visible or not. However, the accuracy of existing white-box visibility metrics, such as HDR-VDP, is often not good enough. CNN-based black-box visibility metrics have proven to be more accurate, but they can…

Cited by 26PDFScholar
2017

Langevin Dynamics with Continuous Tempering for Training Deep Neural Networks

NeurIPS 2017poster

Minimizing non-convex and high-dimensional objective functions is challenging, especially when training modern deep neural networks. In this paper, a novel approach is proposed which divides the training process into two consecutive phases to obtain better generalization performance: Bayesian sampl…

Cited by 27SourcePDFScholar