TY - GEN
T1 - Early Detection At-Risk Students using Machine Learning
AU - Pongpaichet, Siripen
AU - Jankapor, Sawarin
AU - Janchai, Sarun
AU - Tongsanit, Todsaporn
N1 - Publisher Copyright:
© 2020 IEEE.
PY - 2020/10/21
Y1 - 2020/10/21
N2 - Machine Learning is one of the most popular technologies using in many industries, especially to analyze the data and find key insight or new knowledge. In education industry, many studies have applied machine learning techniques for various purposes. One important area is to early detect at-risk students by using data from various sources such as log data from learning management systems (LMSs), class attendances, and actual score from both formative and summative assessments. We present a comparative study aiming to find the most important features and the best classification algorithms to classify at-risk students based on they behaviors. The data are collected from Moodle system [1], printing services system, and students grad system at one of the faculty in the university. The experiment results are evaluated in terms of overall accuracy, precision, and recall. The random forest with oversampling on minority class shows the best result. The performances of the models is better when we have more data in each week of the semester. During week 5, the model can detect about 74 percent of at-risk students.
AB - Machine Learning is one of the most popular technologies using in many industries, especially to analyze the data and find key insight or new knowledge. In education industry, many studies have applied machine learning techniques for various purposes. One important area is to early detect at-risk students by using data from various sources such as log data from learning management systems (LMSs), class attendances, and actual score from both formative and summative assessments. We present a comparative study aiming to find the most important features and the best classification algorithms to classify at-risk students based on they behaviors. The data are collected from Moodle system [1], printing services system, and students grad system at one of the faculty in the university. The experiment results are evaluated in terms of overall accuracy, precision, and recall. The random forest with oversampling on minority class shows the best result. The performances of the models is better when we have more data in each week of the semester. During week 5, the model can detect about 74 percent of at-risk students.
UR - https://www.scopus.com/pages/publications/85098943040
U2 - 10.1109/ICTC49870.2020.9289185
DO - 10.1109/ICTC49870.2020.9289185
M3 - Conference contribution
AN - SCOPUS:85098943040
T3 - International Conference on ICT Convergence
SP - 283
EP - 287
BT - ICTC 2020 - 11th International Conference on ICT Convergence
PB - IEEE Computer Society
T2 - 11th International Conference on Information and Communication Technology Convergence, ICTC 2020
Y2 - 21 October 2020 through 23 October 2020
ER -