← Search

Mansi Phute

4 accepted papers

2025

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety

EMNLP 2025

As large language models (LLMs) see wider real-world use, understanding and mitigating their unsafe behaviors is critical. Interpretation techniques can reveal causes of unsafe outputs and guide safety, but such connections with safety are often overlooked in prior surveys. We present the first surv

Cited by 0SourcePDFScholar
2025

LLM Attributor: Interactive Visual Attribution for LLM Generation

AAAI 2025technical

While large language models (LLMs) have shown remarkable capability to generate convincing text across diverse domains, concerns around its potential risks have highlighted the importance of understanding the rationale behind text generation. We present LLM ATTRIBUTOR, a Python library that provides…

2025

RenderBender: A Survey on Adversarial Attacks Using Differentiable Rendering

IJCAI 2025

Differentiable rendering techniques like Gaussian Splatting and Neural Radiance Fields have become powerful tools for generating high-fidelity models of 3D objects and scenes. Their ability to produce both physically plausible and differentiable models of scenes are key ingredient needed to produce

2024

Semi-Truths: A Large-Scale Dataset of AI-Augmented Images for Evaluating Robustness of AI-Generated Image detectors

NeurIPS 2024poster

Text-to-image diffusion models have impactful applications in art, design, and entertainment, yet these technologies also pose significant risks by enabling the creation and dissemination of misinformation. Although recent advancements have produced AI-generated image detectors that claim robustness…