Next Article in Journal
DPDRC, a Novel Machine Learning Method about the Decision Process for Dimensionality Reduction before Clustering
Previous Article in Journal
A Natural Language Interface to Relational Databases Using an Online Analytic Processing Hypercube
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Chinese Comma Disambiguation in Math Word Problems Using SMOTE and Random Forests

1
School of Information Technology in Education, South China Normal University, Guangzhou 510631, China
2
School of Educational Information Technology, Central China Normal University, Wuhan 430079, China
*
Authors to whom correspondence should be addressed.
AI 2021, 2(4), 738-755; https://doi.org/10.3390/ai2040044
Submission received: 23 November 2021 / Revised: 15 December 2021 / Accepted: 17 December 2021 / Published: 20 December 2021

Abstract

Natural language understanding technologies play an essential role in automatically solving math word problems. In the process of machine understanding Chinese math word problems, comma disambiguation, which is associated with a class imbalance binary learning problem, is addressed as a valuable instrument to transform the problem statement of math word problems into structured representation. Aiming to resolve this problem, we employed the synthetic minority oversampling technique (SMOTE) and random forests to comma classification after their hyperparameters were jointly optimized. We propose a strict measure to evaluate the performance of deployed comma classification models on comma disambiguation in math word problems. To verify the effectiveness of random forest classifiers with SMOTE on comma disambiguation, we conducted two-stage experiments on two datasets with a collection of evaluation measures. Experimental results showed that random forest classifiers were significantly superior to baseline methods in Chinese comma disambiguation. The SMOTE algorithm with optimized hyperparameter settings based on the categorical distribution of different datasets is preferable, instead of with its default values. For practitioners, we suggest that hyperparameters of a classification models be optimized again after parameter settings of SMOTE have been changed.
Keywords: comma disambiguation; feature engineering; hyperparameter tuning; imbalanced learning; natural language understanding; random forests comma disambiguation; feature engineering; hyperparameter tuning; imbalanced learning; natural language understanding; random forests

Share and Cite

MDPI and ACS Style

Huang, J.; Liu, Q.; Zheng, Y.; Wu, L. Chinese Comma Disambiguation in Math Word Problems Using SMOTE and Random Forests. AI 2021, 2, 738-755. https://doi.org/10.3390/ai2040044

AMA Style

Huang J, Liu Q, Zheng Y, Wu L. Chinese Comma Disambiguation in Math Word Problems Using SMOTE and Random Forests. AI. 2021; 2(4):738-755. https://doi.org/10.3390/ai2040044

Chicago/Turabian Style

Huang, Jingxiu, Qingtang Liu, Yunxiang Zheng, and Linjing Wu. 2021. "Chinese Comma Disambiguation in Math Word Problems Using SMOTE and Random Forests" AI 2, no. 4: 738-755. https://doi.org/10.3390/ai2040044

APA Style

Huang, J., Liu, Q., Zheng, Y., & Wu, L. (2021). Chinese Comma Disambiguation in Math Word Problems Using SMOTE and Random Forests. AI, 2(4), 738-755. https://doi.org/10.3390/ai2040044

Article Metrics

Back to TopTop