TY - GEN
T1 - PromptOps
T2 - 32nd Asia-Pacific Software Engineering Conference, APSEC 2025
AU - Sontesadisai, Chommakorn
AU - Sae-Ngow, Chalisa
AU - Rudeerudchanawong, Jirateep
AU - Dangsungnoen, Lapatrada
AU - Ragkhitwetsagul, Chaiyong
AU - Racharak, Teeradaj
AU - Sunetnanta, Thanwadee
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Large Language Models (LLMs) are increasingly utilized in a wide range of natural language processing tasks. Despite their growing adoption, concerns regarding their trustworthiness, i.e., reliability and validity across diverse applications, still remain. This paper introduces a novel visual-based LLM testing tool called PromptOps using the principles of metamorphic testing to assess LLMs beyond traditional accuracy metrics. The tool evaluates LLMs on critical properties such as robustness, fairness, and logical consistency. The tool enables users to design custom test cases via visual programming, define specific prompts, and automatically generate diverse test scenarios. PromptOps fosters greater transparency for model developers by identifying areas for improvement in both performance and fairness. The video demonstration of the PromptOps tool is available at https://youtu.be/M6TbvPIt9kE, and the tool is available at https://github.com/MUICT-SERU/PromptOps.
AB - Large Language Models (LLMs) are increasingly utilized in a wide range of natural language processing tasks. Despite their growing adoption, concerns regarding their trustworthiness, i.e., reliability and validity across diverse applications, still remain. This paper introduces a novel visual-based LLM testing tool called PromptOps using the principles of metamorphic testing to assess LLMs beyond traditional accuracy metrics. The tool evaluates LLMs on critical properties such as robustness, fairness, and logical consistency. The tool enables users to design custom test cases via visual programming, define specific prompts, and automatically generate diverse test scenarios. PromptOps fosters greater transparency for model developers by identifying areas for improvement in both performance and fairness. The video demonstration of the PromptOps tool is available at https://youtu.be/M6TbvPIt9kE, and the tool is available at https://github.com/MUICT-SERU/PromptOps.
KW - LLMs
KW - Metamorphic testing
KW - Trustworthiness
UR - https://www.scopus.com/pages/publications/105035196380
U2 - 10.1109/APSEC66846.2025.00117
DO - 10.1109/APSEC66846.2025.00117
M3 - Conference contribution
AN - SCOPUS:105035196380
T3 - Proceedings - Asia-Pacific Software Engineering Conference, APSEC
SP - 1005
EP - 1008
BT - Proceedings - 2025 32nd Asia-Pacific Software Engineering Conference, APSEC 2025
A2 - Zhang, Tao
A2 - Luo, Xiapu
A2 - Keung, Jacky
A2 - Choi, Eunjong
PB - IEEE Computer Society
Y2 - 2 December 2025 through 5 December 2025
ER -