A COMPARATIVE ANALYSIS OF FORECASTING ACCURACY BETWEEN MACHINE LEARNING MODELS AND OLS REGRESSION: EMPIRICAL EVIDENCE FROM THE VIETNAMESE STOCK MARKET
DOI:
https://doi.org/10.62985/j.huit_ojs.vol26.no3E.463Từ khóa:
Artificial neural networks, machine learning, OLS regression, stock return prediction, Vietnamese stock market.Tóm tắt
This research investigates and compares the predictive performance of stock return forecasting between the Ordinary Least Squares (OLS) regression model and machine learning approaches in the context of the volatile Vietnamese stock market. Using a panel dataset combined with time-series data of listed firms on the Vietnamese stock exchange from 2015 to 2024, the study contrasts the OLS model with three advanced machine learning algorithms, including Artificial Neural Networks (ANN), Random Forest, and XGBoost. Predictive performance is primarily assessed using Root Mean Squared Error (RMSE), alongside Mean Absolute Error (MAE) and out-of-sample R-squared (R²OOS) as robustness measures. The empirical results demonstrate that machine learning models significantly outperform OLS in capturing complex nonlinear relationships in stock returns. Among them, the ANN model achieves the lowest RMSE, indicating the highest predictive accuracy, and generates superior long–short portfolio returns compared to the other models. Furthermore, the Diebold–Mariano test confirms that the differences in predictive accuracy between machine learning models and OLS are statistically significant. Although OLS retains advantages in terms of simplicity and interpretability, machine learning models exhibit clear superiority in predictive performance and quantitative risk management. This study provides important empirical evidence from an emerging market such as Vietnam and offers practical implications for investors and policymakers in optimizing asset allocation decisions.
Tài liệu tham khảo
[1] Q. He, J. Liu, S. Wang, and J. Yu, “The impact of COVID-19 on stock markets,” Econ. Polit. Stud., vol. 8, no. 3, pp. 275–288, Jul. 2020, doi: https://doi.org/10.1080/20954816.2020.1757570
[2] P. O. Takyi and I. Bentum-Ennin, “The impact of COVID-19 on stock market performance in Africa: A Bayesian structural time series approach,” J. Econ. Bus., vol. 115, p. 105968, May 2021, doi: https://doi.org/10.1016/j.jeconbus.2020.105968
[3] D. K. Pandey and V. Kumari, “An event study on the impacts of Covid-19 on the global stock markets,” Int. J. Financ. Mark. Deriv., vol. 8, no. 2, pp. 148-168, 2021, doi: https://doi.org/10.1504/IJFMD.2021.115871
[4] D. V. Hung, N. T. M. Hue, and V. T. Duong, “The Impact of COVID-19 on Stock Market Returns in Vietnam,” J. Risk Financ. Manag., vol. 14, no. 9, art. no. 441, Sep. 2021, doi: https://doi.org/10.3390/jrfm14090441
[5] S. Agusta, F. Rakhman, J. H. Mustakini, and S. Wijayana, “Enhancing the accuracy of stock return movement prediction in Indonesia through recent fundamental value incorporation in multilayer perceptron,” Asian J. Account. Res., vol. 9, no. 4, pp. 358–377, Aug. 2024, doi: https://doi.org/10.1108/AJAR-01-2024-0006
[6] T. Bittmann, “A practical guide from Ordinary Least Squares to causal machine learning,” Int. Food Agribus. Manag. Rev., vol. 28, no. 2, pp. 457–486, Jun. 2025, doi: https://doi.org/10.22434/ifamr.1087
[7] J. Yao, “A Fusion Method Integrated Econometrics and Deep Learning to Improve the Interpretability of Prediction: Evidence From Chinese Carbon Emissions Forecast Based on OLS-CNN Model,” Comput. Econ., vol. 66, no. 4, pp. 2987–3006, Oct. 2025, doi: https://doi.org/10.1007/s10614-024-10793-0
[8] M. Messmer and F. Audrino, “The Lasso and the Factor Zoo-Predicting Expected Returns in the Cross-Section,” Forecasting, vol. 4, no. 4, pp. 969–1003, Nov. 2022, doi: https://doi.org/10.3390/forecast4040053
[9] Y. Wang, X. Hao, and C. Wu, “Forecasting stock returns: A time-dependent weighted least squares approach,” J. Financ. Mark., vol. 53, art. no. 100568, Mar. 2021, doi: https://doi.org/10.1016/j.finmar.2020.100568
[10] A. Y. Polat, “Investor bias, risk and price volatility,” J. Econ. Stud., vol. 50, no. 7, pp. 1317–1335, Oct. 2023, doi: https://doi.org/10.1108/JES-04-2022-0211
[11] D. M. Khan, M. Ali, Z. Ahmad, S. Manzoor, and S. Hussain, “A New Efficient Redescending M-Estimator for Robust Fitting of Linear Regression Models in the Presence of Outliers,” Math. Probl. Eng., vol. 2021, no. 1, pp. 1–11, art. no. 3090537, Nov. 2021, doi: https://doi.org/10.1155/2021/3090537
[12] M. Maiti, Applied Financial Econometrics: Theory, Method and Applications. Singapore: Springer Singapore, 2021. doi: https://doi.org/10.1007/978-981-16-4063-6
[13] N. Caravaggio, R. Lagravinese, and G. Resce, “The determinants of health expenditure: a machine learning approach,” Empir. Econ., vol. 70, no. 2, art. no. 30, Feb. 2026, doi: https://doi.org/10.1007/s00181-025-02854-6
[14] H. Niu, K. Xu, and W. Wang, “A hybrid stock price index forecasting model based on variational mode decomposition and LSTM network,” Appl. Intell., vol. 50, no. 12, pp. 4296–4309, Dec. 2020, doi: https://doi.org/10.1007/s10489-020-01814-0
[15] F. D. Freitas, A. F. De Souza, and A. R. De Almeida, “Prediction-based portfolio optimization model using neural networks,” Neurocomputing, vol. 72, no. 10–12, pp. 2155–2170, Jun. 2009, doi: https://doi.org/10.1016/j.neucom.2008.08.019
[16] S. Gu et al., "Empirical Asset Pricing via Machine Learning," The Review of Financial Studies, vol. 33, no. 5, pp. 2223–2273, Feb. 2020, doi: https://doi.org/10.1093/rfs/hhaa009
[17] D. Bernaciak and J. E. Griffin, “A loss discounting framework for model averaging and selection in time series models,” Int. J. Forecast., vol. 40, no. 4, pp. 1721–1733, Oct. 2024, doi: https://doi.org/10.1016/j.ijforecast.2024.03.001
[18] V. DeMiguel, L. Garlappi, and R. Uppal, “Optimal Versus Naive Diversification: How Inefficient is the 1/ N Portfolio Strategy?,” Rev. Financ. Stud., vol. 22, no. 5, pp. 1915–1953, May 2009, doi: https://doi.org/10.1093/rfs/hhm075
[19] H. Sebastião and P. Godinho, “Forecasting and trading cryptocurrencies with machine learning under changing market conditions,” Financ. Innov., vol. 7, no. 1, art. no. 3, Jan. 2021, doi: https://doi.org/10.1186/s40854-020-00217-x
[20] H. Haider et al., "Effective ways to build and evaluate individual survival distributions," Journal of Machine Learning Research, vol. 21, pp. 1–63, 2020. https://www.jmlr.org/papers/volume21/18-772/18-772.pdf
[21] D. Wolff and U. Neugebauer, “Tree-based machine learning approaches for equity market predictions,” J. Asset Manag., vol. 20, no. 4, pp. 273–288, Jul. 2019, doi: https://doi.org/10.1057/s41260-019-00125-5
[22] V. Lalwani and V. V. Meshram, “The cross-section of Indian stock returns: evidence using machine learning,” Appl. Econ., vol. 54, no. 16, pp. 1814–1828, Apr. 2022, doi: https://doi.org/10.1080/00036846.2021.1982132
[23] E. F. Fama, “Efficient Capital Markets: A Review of Theory and Empirical Work,” J. Finance, vol. 25, no. 2, pp. 383-417, May 1970, doi: https://doi.org/10.2307/2325486
[24] B. G. Malkiel, "The Efficient Market Hypothesis and Its Critics," Journal of Economic Perspectives, vol. 17, no. 1, pp. 59–82, Apr. 2003, doi: https://doi.org/10.1257/089533003321164958
[25] A. W. Lo, "The Adaptive Markets Hypothesis," The Journal of Portfolio Management, vol. 30, no. 5, pp. 15–29, 2004, doi: https://doi.org/10.3905/jpm.2004.442611
[26] C. Brooks, Introductory Econometrics for Finance, 4th ed. Cambridge University Press, 2019. doi: https://doi.org/10.1017/9781108524872
[27] K. Fukuda, “Distribution switching in financial time series,” Math. Comput. Simul., vol. 79, no. 5, pp. 1711–1720, Jan. 2009, doi: https://doi.org/10.1016/j.matcom.2008.08.012
[28] S. Athey and G. W. Imbens, “Machine learning methods that economists should know about,” Annual Review of Economics, vol. 11, pp. 685–725, Aug. 2019, doi: https://doi.org/10.1146/annurev-economics-080217-053433.
[29] W. H. Crown, “Real-World Evidence, Causal Inference, and Machine Learning,” Value in Health, vol. 22, no. 5, pp. 587–592, May 2019, doi: https://doi.org/10.1016/j.jval.2019.03.001
[30] Y. Peng and J. G. De Moraes Souza, “Chaos, overfitting and equilibrium: To what extent can machine learning beat the financial market?,” Int. Rev. Financ. Anal., vol. 95, art. no. 103474, Oct. 2024, doi: https://doi.org/10.1016/j.irfa.2024.103474
[31] M. Bagnara, “Asset Pricing and Machine Learning: A critical review,” J. Econ. Surv., vol. 38, no. 1, pp. 27–56, Feb. 2024, doi: https://doi.org/10.1111/joes.12532
[32] S. Giglio, B. Kelly, and D. Xiu, “Factor models, machine learning, and asset pricing,” Annual Review of Financial Economics, vol. 14, pp. 337–368, 2022, doi: https://doi.org/10.1146/annurev-financial-101521-104735.
[33] R. Priel and L. Rokach, “Machine learning-based stock picking using value investing and quality features,” Neural Comput. Appl., vol. 36, no. 20, pp. 11963–11986, Jul. 2024, doi: https://doi.org/10.1007/s00521-024-09700-3
[34] A. S. Al-Jawarneh, A. R. M. Alsayed, and H. N. Ayyoub, “Optimizing modelling accuracy using variational mode decomposition and elastic net regression: Evidence in stock market prediction,” Array, vol. 28, art. no. 100603, Dec. 2025, doi: https://doi.org/10.1016/j.array.2025.100603
[35] R. Gupta, J. Nel, and C. Pierdzioch, “Investor Confidence and Forecastability of US Stock Market Realized Volatility: Evidence from Machine Learning,” J. Behav. Finance, vol. 24, no. 1, pp. 111–122, Jan. 2023, doi: https://doi.org/10.1080/15427560.2021.1949719
[36] F. Moreno-Pino and S. Zohren, “DeepVol: Volatility forecasting from high-frequency data with dilated causal convolutions,” Quantitative Finance, vol. 24, no. 8, pp. 1105–1127, 2024, doi: https://doi.org/10.1080/14697688.2024.2387222.
[37] D. Chun, J. Kang, and J. Kim, “Forecasting returns with machine learning and optimizing global portfolios: evidence from the Korean and U.S. stock markets,” Financ. Innov., vol. 10, no. 1, art. no. 124, Dec. 2024, doi: https://doi.org/10.1186/s40854-024-00648-w
[38] G. Ji, J. Yu, K. Hu, J. Xie, and X. Ji, “An adaptive feature selection schema using improved technical indicators for predicting stock price movements,” Expert Syst. Appl., vol. 200, art. no. 116941, Aug. 2022, doi: https://doi.org/10.1016/j.eswa.2022.116941
[39] D. A. Dickey and W. A. Fuller, “Distribution of the Estimators for Autoregressive Time Series with a Unit Root,” J. Am. Stat. Assoc., vol. 74, no. 366a, pp. 427–431, Jun. 1979, doi: https://doi.org/10.1080/01621459.1979.10482531
[40] P. C. B. Phillips and P. Perron, “Testing for a unit root in time series regression,” Biometrika, vol. 75, no. 2, pp. 335–346, 1988, doi: https://doi.org/10.1093/biomet/75.2.335
[41] D. N. Gujarati, Basic Econometrics, 4th ed. New York, NY, USA: McGraw-Hill, 2003.
[42] E. F. Fama and K. R. French, “The Cross‐Section of Expected Stock Returns,” J. Finance, vol. 47, no. 2, pp. 427–465, Jun. 1992, doi: https://doi.org/10.1111/j.1540-6261.1992.tb04398.x
[43] N. Jegadeesh, “Evidence of Predictable Behavior of Security Returns,” J. Finance, vol. 45, no. 3, pp. 881–898, Jul. 1990, doi: https://doi.org/10.1111/j.1540-6261.1990.tb05110.x
[44] S. Basu, “Investment Performance of Common Stocks in Relation to Their Price-Earnings Ratios: A Test of the Efficient Market Hypothesis,” J. Finance, vol. 32, no. 3, pp. 663-682, Jun. 1977, doi: https://doi.org/10.2307/2326304
[45] E. F. Fama and K. R. French, “Dividend yields and expected stock returns,” Journal of Financial Economics, vol. 22, no. 1, pp. 3–25, 1988, doi: https://doi.org/10.1016/0304-405X(88)90020-7.
[46] E. F. Fama and K. R. French, “A five-factor asset pricing model,” J. Financ. Econ., vol. 116, no. 1, pp. 1–22, Apr. 2015, doi: https://doi.org/10.1016/j.jfineco.2014.10.010
[47] K. Hou, C. Xue, and L. Zhang, “Digesting Anomalies: An Investment Approach,” Rev. Financ. Stud., vol. 28, no. 3, pp. 650–705, Mar. 2015, doi: https://doi.org/10.1093/rfs/hhu068
[48] R. Novy-Marx, “The other side of value: The gross profitability premium,” J. Financ. Econ., vol. 108, no. 1, pp. 1–28, Apr. 2013, doi: https://doi.org/10.1016/j.jfineco.2013.01.003
[49] J. Lakonishok, A. Shleifer, and R. W. Vishny, “Contrarian Investment, Extrapolation, and Risk,” J. Finance, vol. 49, no. 5, pp. 1541–1578, Dec. 1994, doi: https://doi.org/10.1111/j.1540-6261.1994.tb04772.x
[50] M. J. Cooper, H. Gulen, and M. J. Schill, "Asset Growth and the Cross-Section of Stock Returns," The Journal of Finance, vol. 63, no. 4, pp. 1609–1651, 2008, doi: https://doi.org/10.1111/j.1540-6261.2008.01370.x
[51] W. F. Sharpe, “Capital Asset Prices: A Theory of Market Equilibrium Under Conditions of Risk*,” J. Finance, vol. 19, no. 3, pp. 425–442, Sep. 1964, doi: https://doi.org/10.1111/j.1540-6261.1964.tb02865.x
[52] L. C. Bhandari, “Debt/Equity Ratio and Expected Common Stock Returns: Empirical Evidence,” J. Finance, vol. 43, no. 2, pp. 507–528, Jun. 1988, doi: https://doi.org/10.1111/j.1540-6261.1988.tb03952.x
[53] A. Ang, R. J. Hodrick, Y. Xing, and X. Zhang, “The Cross‐Section of Volatility and Expected Returns,” J. Finance, vol. 61, no. 1, pp. 259–299, Feb. 2006, doi: https://doi.org/10.1111/j.1540-6261.2006.00836.x
[54] N. Jegadeesh and S. Titman, “Returns to Buying Winners and Selling Losers: Implications for Stock Market Efficiency,” J. Finance, vol. 48, no. 1, pp. 65–91, Mar. 1993, doi: https://doi.org/10.1111/j.1540-6261.1993.tb04702.x
[55] M. J. Brennan, T. Chordia, and A. Subrahmanyam, “Alternative factor specifications, security characteristics, and the cross-section of expected stock returns,” Journal of Financial Economics, vol. 49, no. 3, pp. 345–373, 1998, doi: https://doi.org/10.1016/S0304-405X(98)00028-2.
[56] V. T. Datar, N. Y. Naik, and R. Radcliffe, “Liquidity and stock returns: An alternative test,” J. Financ. Mark., vol. 1, no. 2, pp. 203–219, Aug. 1998, doi: https://doi.org/10.1016/S1386-4181(97)00004-9
[57] Y. Amihud, "Illiquidity and stock returns: cross-section and time-series effects," Journal of Financial Markets, vol. 5, no. 1, pp. 31–56, 2002, doi: https://doi.org/10.1016/S1386-4181(01)00024-6
[58] L. Breiman, “Random Forests,” Mach. Learn., vol. 45, no. 1, pp. 5–32, Oct. 2001, doi: https://doi.org/10.1023/A:1010933404324
[59] T. Chen and C. Guestrin, “XGBoost: A Scalable Tree Boosting System,” in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco California USA: ACM, Aug. 2016, pp. 785–794. doi: https://doi.org/10.1145/2939672.2939785
[60] T. Hastie, R. Tibshirani, and J. Friedman, The Elements of Statistical Learning. New York, NY, USA: Springer, 2009, doi: https://doi.org/10.1007/978-0-387-84858-7
[61] J. Snoek, H. Larochelle, and R. P. Adams, “Practical Bayesian optimization of machine learning algorithms,” in Advances in Neural Information Processing Systems 25, 2012, pp. 2951–2959. doi: https://doi.org/10.5555/2999325.2999464.
[62] J. Bergstra and Y. Bengio, “Random search for hyper-parameter optimization,” J. Mach. Learn. Res., vol. 13, no. 10, pp. 281–305, 2012.
[63] J. D. Hamilton, Time Series Analysis. Princeton University Press, 2020. doi: https://doi.org/10.2307/j.ctv14jx6sm


