Abstract
In machine learning processes, parameter settings affect model accuracy. Text-based emotion detection requires stable and accurate models, making parameter choices, such as the random state, increasingly important. Previous studies usually set the random state to 42, claiming that this should be the best for obtaining good accuracy. This study examined random state settings, experimenting with values from 1 to 720 and observing the results in accuracy. In addition, a dataset was employed for emotion detection using the Random Forest (RF) classifier with two vectorizers, TF-IDF and Count. The results show that different random state settings affect model accuracy. In the training subset, the TF-IDF vectorizer offered higher and more stable accuracy than the Count vectorizer. However, the Count vectorized achieved higher accuracy on both the validation and test sets.
| Original language | English |
|---|---|
| Pages (from-to) | 33247-33252 |
| Number of pages | 6 |
| Journal | Engineering, Technology and Applied Science Research |
| Volume | 16 |
| Issue number | 2 |
| DOIs | |
| Publication status | Published - 2026 |
Keywords
- Count
- TF-IDF
- emotion detection
- random forest
- random state
- vectorizers
Fingerprint
Dive into the research topics of 'A Comparative Study of TF-IDF and Count Vectorizer under Random State Changes in a Random Forest Classifier for Emotion Detection'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver