← Search

DanDan Guo

13 accepted papers

2026

A Guardrail for Safety Preservation: When Safety-Sensitive Subspace Meets Harmful-Resistant Null-Space

ICLR 2026poster

Large language models (LLMs) have achieved remarkable success in diverse tasks, yet their safety alignment remains fragile during adaptation. Even when fine-tuning on benign data or with low-rank adaptation, pre-trained safety behaviors are easily degraded, leading to harmful responses in the fine-…

Cited by 0SourceScholar
2026

Imitating the Truth: Attention-aware Truth-Guided Enhancement for Hallucination Mitigation in Large Vision-Language Models

ICLR 2026poster

Large Vision-Language Models (LVLMs) achieve impressive multimodal reasoning but remain prone to hallucinations, generating content inconsistent with visual evidence. Existing mitigation methods often rely on auxiliary modules or coarse decoding-time adjustments, overlooking the fine-grained dynamic…

Cited by 0SourceScholar
2026

LLM as an Algorithmist: Enhancing Anomaly Detectors via Programmatic Synthesis

ICLR 2026poster

Existing anomaly detection (AD) methods for tabular data usually rely on some assumptions about anomaly patterns, leading to inconsistent performance in real-world scenarios. While Large Language Models (LLMs) show remarkable reasoning capabilities, their direct application to tabular AD is impeded…

Cited by 0SourcecodeScholar
2026

Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling

ICML 2026oral

Reward models learned from human preferences are central to aligning large language models (LLMs) via reinforcement learning from human feedback, yet they are often vulnerable to reward hacking due to noisy annotations and systematic biases such as response length or style. We propose Bayesian Non-N…

Cited by 0SourceScholar
2025

APLOT: Robust Reward Modeling via Adaptive Preference Learning with Optimal Transport

EMNLP 2025

The reward model (RM) plays a crucial role in aligning Large Language Models (LLMs) with human preferences through Reinforcement Learning, where the Bradley-Terry (BT) objective has been recognized as simple yet powerful, specifically for pairwise preference learning. However, BT-based RMs often str

2025

Balancing Two Classifiers via A Simplex ETF Structure for Model Calibration

CVPR 2025poster

In recent years, deep neural networks (DNNs) have demonstrated state-of-the-art performance across various domains. However, despite their success, they often face calibration issues, particularly in safety-critical applications such as autonomous driving and healthcare, where unreliable predictions…

2025

Beyond Words: Augmenting Discriminative Richness via Diffusions in Unsupervised Prompt Learning

CVPR 2025poster

Fine-tuning vision-language models (VLMs) with large amounts of unlabeled data has recently garnered significant interest. However, a key challenge remains the lack of high-quality pseudo-labeled data. Current pseudo-labeling strategies often struggle with mismatches between semantic and visual info…

2025

FedAWA: Adaptive Optimization of Aggregation Weights in Federated Learning Using Client Vectors

CVPR 2025poster

Federated Learning (FL) has emerged as a promising framework for distributed machine learning, enabling collaborative model training without sharing local data, thereby preserving privacy and enhancing security. However, data heterogeneity resulting from differences across user behaviors, preference…

2023

A Simple Yet Effective Subsequence-Enhanced Approach for Cross-Domain NER

AAAI 2023technical

Cross-domain named entity recognition (NER), aiming to address the limitation of labeled resources in the target domain, is a challenging yet important task. Most existing studies alleviate the data discrepancy across different domains at the coarse level via combing NER with language modelings or i…

2020

Recurrent Hierarchical Topic-Guided RNN for Language Generation

ICML 2020poster

To simultaneously capture syntax and global semantics from a text corpus, we propose a new larger-context recurrent neural network (RNN) based language model, which extracts recurrent hierarchical semantic structure via a dynamic deep topic model to guide natural language generation. Moving beyond a…

2018

WHAI: Weibull Hybrid Autoencoding Inference for Deep Topic Modeling

ICLR 2018poster

To train an inference network jointly with a deep generative topic model, making it both scalable to big corpora and fast in out-of-sample prediction, we develop Weibull hybrid autoencoding inference (WHAI) for deep latent Dirichlet allocation, which infers posterior samples via a hybrid of stochast…

Cited by 126SourcePDFScholar