AAAI 2025technical2 citations

ScriptSmith: A Unified LLM Framework for Enhancing IT Operations via Automated Bash Script Generation, Assessment, and Refinement

Pooja Aggarwal, Oishik Chatterjee, Ting Dai, Suranjana Samanta, Prateeti Mohapatra, Debanjana Kar, Ruchi Mahindru, Steve Barbier

Abstract

In the rapidly evolving landscape of site reliability engineering (SRE), the demand for efficient and effective solutions to manage and resolve issues in site and cloud applications is paramount. This paper presents an innovative approach to action automation using large language models (LLMs) for script generation, assessment, and refinement. By leveraging the capabilities of LLMs, we aim to significantly reduce the human effort involved in writing and debugging scripts, thereby enhancing the productivity of SRE teams. Our experiments focus on Bash scripts, a commonly used tool in SRE, and involve the CodeSift dataset of 100 tasks and the InterCode dataset of 153 tasks. The results show that LLMs can automatically assess and refine scripts efficiently, reducing the need for script validation in an execution environment. Results demonstrate that the framework shows an overall improvement of 7-10% in script generation.

BibTeX
@article{Aggarwal_Chatterjee_Dai_Samanta_Mohapatra_Kar_Mahindru_Barbier_Postea_Blancett_de Magalhaes_2025, title={ScriptSmith: A Unified LLM Framework for Enhancing IT Operations via Automated Bash Script Generation, Assessment, and Refinement}, volume={39}, url={https://ojs.aaai.org/index.php/AAAI/article/view/35147}, DOI={10.1609/aaai.v39i28.35147}, abstractNote={In the rapidly evolving landscape of site reliability engineering (SRE), the demand for efficient and effective solutions to manage and resolve issues in site and cloud applications is paramount. This paper presents an innovative approach to action automation using large language models (LLMs) for script generation, assessment, and refinement. By leveraging the capabilities of LLMs, we aim to significantly reduce the human effort involved in writing and debugging scripts, thereby enhancing the productivity of SRE teams. Our experiments focus on Bash scripts, a commonly used tool in SRE, and involve the CodeSift dataset of 100 tasks and the InterCode dataset of 153 tasks. The results show that LLMs can automatically assess and refine scripts efficiently, reducing the need for script validation in an execution environment. Results demonstrate that the framework shows an overall improvement of 7-10% in script generation.}, number={28}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, author={Aggarwal, Pooja and Chatterjee, Oishik and Dai, Ting and Samanta, Suranjana and Mohapatra, Prateeti and Kar, Debanjana and Mahindru, Ruchi and Barbier, Steve and Postea, Eugen and Blancett, Brad and de Magalhaes, Arthur}, year={2025}, month={Apr.}, pages={28829-28835} }
ScriptSmith: A Unified LLM Framework for Enhancing IT Operations via Automated Bash Script Generation, Assessment, and Refinement · AAAI 2025