TY - GEN
T1 - Phebe
T2 - 14th International Conference on Advances in Information Technology, IAIT 2026
AU - Anurakboonying, Guntawit
AU - Proongpattanaskul, Ratchaphon
AU - Sawangwongchinsri, Chalantorn
AU - Phienthrakul, Tanasanee
N1 - Publisher Copyright:
© 2026 Copyright held by the owner/author(s).
PY - 2026/6/16
Y1 - 2026/6/16
N2 - Most data visualization tools require technical expertise, creating significant accessibility barriers for non-technical users. We present Phebe, a web-based platform that enables users to upload tabular datasets, perform AI-guided data cleaning, and generate interactive visualizations by expressing analytical needs in plain English. The system is built around a multi-agent architecture powered by large language models (LLMs), comprising an Orchestrator Agent for intent classification, a Text-to-SQL Agent with execution-based self-correction, an Analysis Agent for statistical computations, and a Chart Recommendation Agent for automated visual encoding selection. An end-to-end pipeline integrates CSV ingestion with robust validation, a six-category data quality subsystem, SQL generation against user-uploaded datasets stored in MySQL, and a D3.js-based rendering layer supporting seven chart types, a drag-and-drop dashboard, and PDF export. Evaluated on the BIRD-Mini Text-to-SQL benchmark (500 queries), Phebe achieves 61.6% execution accuracy - an 8 percentage point improvement over the GPT baseline - with a 0% SQL syntax error rate attributable to the self-correction loop. A user survey further confirms that 85% of participants (n = 27) rated their ability to explore complex datasets without writing SQL or code at 4 or 5 out of 5, validating the platform's goal of democratizing data exploration.
AB - Most data visualization tools require technical expertise, creating significant accessibility barriers for non-technical users. We present Phebe, a web-based platform that enables users to upload tabular datasets, perform AI-guided data cleaning, and generate interactive visualizations by expressing analytical needs in plain English. The system is built around a multi-agent architecture powered by large language models (LLMs), comprising an Orchestrator Agent for intent classification, a Text-to-SQL Agent with execution-based self-correction, an Analysis Agent for statistical computations, and a Chart Recommendation Agent for automated visual encoding selection. An end-to-end pipeline integrates CSV ingestion with robust validation, a six-category data quality subsystem, SQL generation against user-uploaded datasets stored in MySQL, and a D3.js-based rendering layer supporting seven chart types, a drag-and-drop dashboard, and PDF export. Evaluated on the BIRD-Mini Text-to-SQL benchmark (500 queries), Phebe achieves 61.6% execution accuracy - an 8 percentage point improvement over the GPT baseline - with a 0% SQL syntax error rate attributable to the self-correction loop. A user survey further confirms that 85% of participants (n = 27) rated their ability to explore complex datasets without writing SQL or code at 4 or 5 out of 5, validating the platform's goal of democratizing data exploration.
KW - Data Cleansing
KW - Data Visualization
KW - Large Language Models
KW - Multi-Agent Systems
KW - Natural Language Interfaces
KW - Text-to-SQL
UR - https://www.scopus.com/pages/publications/105045273429
U2 - 10.1145/3816713.3819784
DO - 10.1145/3816713.3819784
M3 - Conference contribution
AN - SCOPUS:105045273429
T3 - IAIT 2026 - 14th International Conference on Advances in Information Technology
BT - IAIT 2026 - 14th International Conference on Advances in Information Technology
PB - Association for Computing Machinery, Inc
Y2 - 17 June 2026 through 19 June 2026
ER -