TY - GEN
T1 - Why Visualize Data When Coding? Preliminary Categories for Coding in Jupyter Notebooks
AU - Settewong, Tasha
AU - Ritta, Natanon
AU - Kula, Raula Gaikovina
AU - Ragkhitwetsagul, Chaiyong
AU - Sunetnanta, Thanwadee
AU - Matsumoto, Kenichi
N1 - Publisher Copyright:
© 2022 IEEE.
PY - 2022
Y1 - 2022
N2 - Data visualization becomes a crucial component in data analytics, especially data exploration, understanding, and analysis. Effective data visualization impacts decision-making and aids in discovering and understanding relationships. It leads to benefits in data-intensive software development tasks e.g., feature engineering in machine learning-based software projects. However, it is unknown how visualizations are used in competitive programming. The idea of this paper is to report early results on what visualizations are prevalent in competitive programming. Grandmasters are the highest level reached in competitions (novice, expert, master, and grandmaster). Analyzing the visualizations of 7 high-rank competitors (i.e., Grandmaster) in Kaggle, we identify and present a catalog of visualizations used to both tell a story from the data, as well as explain the process and pipelines involved to explain their coding solutions. Our taxonomy includes nine types from over 821 visualizations in 68 instances of Jupyter notebooks. Furthermore, most visualizations are for data analysis for distribution (DA Distribution), and frequency (DA Frequency) are most used. We envision that this catalog can be useful to better understand different situations in which to employ these visualizations.
AB - Data visualization becomes a crucial component in data analytics, especially data exploration, understanding, and analysis. Effective data visualization impacts decision-making and aids in discovering and understanding relationships. It leads to benefits in data-intensive software development tasks e.g., feature engineering in machine learning-based software projects. However, it is unknown how visualizations are used in competitive programming. The idea of this paper is to report early results on what visualizations are prevalent in competitive programming. Grandmasters are the highest level reached in competitions (novice, expert, master, and grandmaster). Analyzing the visualizations of 7 high-rank competitors (i.e., Grandmaster) in Kaggle, we identify and present a catalog of visualizations used to both tell a story from the data, as well as explain the process and pipelines involved to explain their coding solutions. Our taxonomy includes nine types from over 821 visualizations in 68 instances of Jupyter notebooks. Furthermore, most visualizations are for data analysis for distribution (DA Distribution), and frequency (DA Frequency) are most used. We envision that this catalog can be useful to better understand different situations in which to employ these visualizations.
KW - data analysis
KW - data visualization
KW - machine learning competition
UR - https://www.scopus.com/pages/publications/85149180580
U2 - 10.1109/APSEC57359.2022.00063
DO - 10.1109/APSEC57359.2022.00063
M3 - Conference contribution
AN - SCOPUS:85149180580
T3 - Proceedings - Asia-Pacific Software Engineering Conference, APSEC
SP - 462
EP - 466
BT - Proceedings - 2022 29th Asia-Pacific Software Engineering Conference, APSEC 2022
PB - IEEE Computer Society
T2 - 29th Asia-Pacific Software Engineering Conference, APSEC 2022
Y2 - 6 December 2022 through 9 December 2022
ER -