ABSTRACT
Diabetes mellitus (DM) is a diverse group of metabolic disorders that is frequently associated with a high disease burden in developing countries such as Nigeria. It also needs continuous blood glucose monitoring and self-management. Multicollinearity is one of the problems usually encountered by Economists and Statisticians when predicting a dependent variable from the set of independent variables that are significantly or highly correlated especially when the traditional method of regression analysis is used. This research is aimed to determine the effect of multicollinearity in predicting diabetes mellitus using statistical neural network. In this research, 100 patients were considered from Ahmadu Bello University Teaching Hospital who have undergone diabetes screening test and 29 risk factors were used. Variance inflation factor (VIF) detected multicollinearity among some risk factors and Principal component technique (PCA) was employed to remove it. Levenberg-Marquardt (LM) algorithm was used to train the statistical neural network for the original and principal components data. The results show that when five (5) hidden neurons architecture is used, the model achieved 99.0% and 93.9% accuracy for training the original and reduced data respectively. Similarly, when the number of hidden neurons is increased to ten (10), the model achieved 98.7% and 94.4% accuracy for training the original and reduced data respectively. The research therefore concludes that unlike traditional econometrics and statistical models, statistical neural networks estimation process is not negatively affected by the presence of multicollinearity in the data but get better when more information are utilised because it gives better estimates when the whole data is considered as inputs than when the reduced data is used.