Skip to main navigation Skip to search Skip to main content

PromptOps: Automated Tool for Testing Trustworthiness of LLMs

  • Mahidol University
  • Tohoku University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Large Language Models (LLMs) are increasingly utilized in a wide range of natural language processing tasks. Despite their growing adoption, concerns regarding their trustworthiness, i.e., reliability and validity across diverse applications, still remain. This paper introduces a novel visual-based LLM testing tool called PromptOps using the principles of metamorphic testing to assess LLMs beyond traditional accuracy metrics. The tool evaluates LLMs on critical properties such as robustness, fairness, and logical consistency. The tool enables users to design custom test cases via visual programming, define specific prompts, and automatically generate diverse test scenarios. PromptOps fosters greater transparency for model developers by identifying areas for improvement in both performance and fairness. The video demonstration of the PromptOps tool is available at https://youtu.be/M6TbvPIt9kE, and the tool is available at https://github.com/MUICT-SERU/PromptOps.

Original languageEnglish
Title of host publicationProceedings - 2025 32nd Asia-Pacific Software Engineering Conference, APSEC 2025
EditorsTao Zhang, Xiapu Luo, Jacky Keung, Eunjong Choi
PublisherIEEE Computer Society
Pages1005-1008
Number of pages4
ISBN (Electronic)9798331566531
DOIs
Publication statusPublished - 2025
Event32nd Asia-Pacific Software Engineering Conference, APSEC 2025 - Macau, China
Duration: 2 Dec 20255 Dec 2025

Publication series

NameProceedings - Asia-Pacific Software Engineering Conference, APSEC
ISSN (Print)1530-1362

Conference

Conference32nd Asia-Pacific Software Engineering Conference, APSEC 2025
Country/TerritoryChina
CityMacau
Period2/12/255/12/25

Keywords

  • LLMs
  • Metamorphic testing
  • Trustworthiness

Fingerprint

Dive into the research topics of 'PromptOps: Automated Tool for Testing Trustworthiness of LLMs'. Together they form a unique fingerprint.

Cite this