Multilevel Stacking with Domain-Adversarial Neural Network and Enriched Metadata for Improved Renewable Energy Output Forecasting
DOI:
https://doi.org/10.64922/jteas.v1i1.86Keywords:
Multilevel Stacking, Renewable Energy Forecasting, Domain Adversarial Neural Network, Enriched Metadata, Walk-Forward ValidationAbstract
This study proposes a multilevel stacking approach that extends Wolpert's (1992) stacked generalisation, aiming to minimise forecasting error, reduce systematic bias, and enforce temporal domain invariance. It incorporates walk-forward out-of-fold and step-based walk-forward validation, enriched metadata, and a domain adversarial neural network (DANN). The base learners XGB and LSTM function as residual and latent encoders for the Random Forest with DANN meta-learner. DANN is integrated in meta-learning to minimise task regression loss and maximize domain discrimination loss through a gradient reversal layer. This fusion improves domain transition stability, as the model’s high explanatory power with ^2 values exceed 0.98 across validation and test regimes. The MASE values below 0.07 on both regimes indicate strong predictive performance relative to a naïve benchmark, and NRMSE values around 0.025 confirm that the proposed stacking framework effectively captures temporal dependencies. Moreover, the model achieves a near-symmetric residual distribution, with a skewness of 0.37 on the validation set and -0.01 on the test set, suggesting that adversarial feature learning helps reduce asymmetric prediction errors and stabilise meta-learning. This proposed framework demonstrates empirically that stacking should not be evaluated solely on MAE/RMSE but on residual distribution behaviour under domain transition. Furthermore, SHAP analysis provides empirical evidence that enriched metadata significantly enhances the learning capacity, allowing the meta-learner to capture complex dependencies.
References
Ben-David, S., Blitzer, J., Crammer, K., Kulesza, A., Pereira, F., & Vaughan, J. W. (2010). A theory of learning from different domains. Machine Learning, 79(1–2), 151–175. https://doi.org/10.1007/s10994-009-5152-4
Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32. https://doi.org/10.1023/A:1010933404324
Cawood, P., & Van Zyl, T. (2022). Evaluating state-of-the-art, forecasting ensembles, and meta-learning strategies for model fusion. Forecasting, 4(3), 732–751. https://doi.org/10.3390/forecast4030040
Cerqueira, V., Torgo, L., & Mozetič, I. (2020). Evaluating time series forecasting models: An empirical study on performance estimation methods. Machine Learning, 109(11), 1997–2028. https://doi.org/10.1007/s10994-020-05910-7
Cerqueira, V., Torgo, L., Mozetič, I., Omlin, C., & Bifet, A. (2022). A survey of predictive modelling under concept drift. ACM Computing Surveys, 55(6), 1–37. https://doi.org/10.1145/3475668
Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785–794). ACM. https://doi.org/10.1145/2939672.2939785
Fatima, S. S. W., & Rahimi, A. (2024). A review of time-series forecasting algorithms for industrial manufacturing systems. Machines, 12(6), 380. https://doi.org/10.3390/machines12060380
Friedman, J. H. (2001). Greedy function approximation: A gradient boosting machine. Annals of Statistics, 29(5), 1189–1232. https://doi.org/10.1214/aos/1013203451
Gama, J., Žliobaitė, I., Bifet, A., Pechenizkiy, M., & Bouchachia, A. (2014). A survey on concept drift adaptation. ACM Computing Surveys, 46(4), 1–37. https://doi.org/10.1145/2523813
Ganin, Y., & Lempitsky, V. (2015, June). Unsupervised domain adaptation by backpropagation. In International Conference on Machine Learning (pp. 1180–1189). PMLR. https://doi.org/10.48550/arXiv.1409.7495
Ganin, Y., Ustinova, E., Ajakan, H., Germain, P., Larochelle, H., Laviolette, F., … Lempitsky, V. (2016). Domain-adversarial training of neural networks. Journal of Machine Learning Research, 17(59), 1–35. https://doi.org/10.5555/2946645.2946704
Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural Computation, 9(8), 1735–1780. https://doi.org/10.1162/neco.1997.9.8.1735
Huang, H., Zhu, Q., Zhu, X., & Zhang, J. (2023). An adaptive, data-driven stacking ensemble learning framework for the short-term forecasting of renewable energy generation. Energies, 16(4), 1963. https://doi.org/10.3390/en16041963
Johnson, M. (2024). Deep learning methods for time series forecasting: A comparative review. Journal of Computer Technology and Software, 3(3). https://doi.org/10.5281/zenodo.12669948
Khan, W., Walker, S., & Zeiler, W. (2022). Improved solar photovoltaic energy generation forecast using a deep learning-based ensemble stacking approach. Energy, 240, 122812. https://doi.org/10.1016/j.energy.2021.122812
Kontopoulou, V. I., Panagopoulos, A. D., Kakkos, I., & Matsopoulos, G. K. (2023). A review of ARIMA vs. machine learning approaches for time series forecasting in data-driven networks. Future Internet, 15(8), 255. https://doi.org/10.3390/fi15080255
Krauss, C., Do, X. A., & Huck, N. (2017). Deep neural networks, gradient-boosted trees, random forests: Statistical arbitrage on the S&P 500. European Journal of Operational Research, 259(2), 689–702. https://doi.org/10.1016/j.ejor.2016.10.031
Kuncheva, L. I., Matthews, C. E., Arnaiz-González, Á., & Rodríguez, J. J. (2020). Feature selection from high-dimensional data with very low sample size: A cautionary tale (arXiv:2008.12025). arXiv. https://doi.org/10.48550/arXiv.2008.12025
Lessmann, S., Baesens, B., Seow, H. V., & Thomas, L. C. (2015). Benchmarking state-of-the-art classification algorithms for credit scoring: An update of research. European Journal of Operational Research, 247(1), 124–136. https://doi.org/10.1016/j.ejor.2015.05.030
Li, N., Yu, Y., & Zhou, Z. H. (2012, September). Diversity regularized ensemble pruning. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases (pp. 330–345). Springer Berlin Heidelberg. https://doi.org/10.1007/978-3-642-33460-3_27
Middlehurst, M., Large, J., Flynn, M., Lines, J., Bostrom, A., & Bagnall, A. (2021). HIVE-COTE 2.0: A new meta ensemble for time series classification. Machine Learning, 110(11), 3211–3243. https://doi.org/10.1007/s10994-021-06057-9
Tashman, L. J. (2000). Out-of-sample tests of forecasting accuracy: An analysis and review. International Journal of Forecasting, 16(4), 437–450. https://doi.org/10.1016/S0169-2070(00)00065-0
Ting, K. M., & Witten, I. H. (1999). Issues in stacked generalization. Journal of Artificial Intelligence Research, 10, 271–289. https://doi.org/10.5555/1622859.1622868
Tsymbal, A. (2004). The problem of concept drift: Definitions and related work. Computer Science Department, Trinity College Dublin, 106(2), 58.
Wolpert, D. H. (1992). Stacked generalization. Neural Networks, 5(2), 241–259. https://doi.org/10.1016/S0893-6080(05)80023-1
Yang, A. (2025). Big data-driven corporate financial forecasting and decision support: A study of CNN-LSTM machine learning models. Frontiers in Applied Mathematics and Statistics, 11, 1566078. https://doi.org/10.3389/fams.2025.1566078
Zhang, X., Zhou, X., Lin, M., & Sun, J. (2018). ShuffleNet: An extremely efficient convolutional neural network for mobile devices. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 6848–6856). https://doi.org/10.1109/CVPR.2018.00716
Zhou, K., Yang, Y., Qiao, Y., & Xiang, T. (2021). Domain adaptive ensemble learning. IEEE Transactions on Image Processing, 30, 8008–8018. https://doi.org/10.1109/TIP.2021.3112012