Imputation techniques for missing rainfall data in the Indian Sundarbans
DOI:
https://doi.org/10.70917/jcc-2026-011Keywords:
Indian Sundarbans, Missing Data Imputation, Grid Data; Rain Gauge, IDW, k-nearest neighboursAbstract
Climatic station data is crucial for understanding the meteorological characteristics of the Indian Sundarbans, a World Heritage site, where over 70% of the population is engaged in agriculture. However, due to the exiguity of rain gauge stations and insufficient data from the operational stations, quantifying the climate change information and prediction at the ground level is challenging. Moreover, studies on imputing missing rainfall values, particularly in Indian Sundarbans, are limited. The present study experimented various imputation algorithms such as hot deck, k-Nearest Neighbour, Linear Regression and Inverse Distance Weighted. These were applied to the nearest IMD rain gauge stations in the study area. We have also considered the k-NN technique, as applied to IMD gridded data (0.25◦ × 0.25◦). In our analysis, we artificially excluded 6%, 16%, and 25% of the data points at random. The results obtained from the various imputation methods were then compared with the actual observed data, using agreement indices such as R2, MAE, RMSE, and MAPE. The findings reveal that the k-NN applied to IMD gridded data is the most effective approach, achieving an R2 value exceeding 0.9.
References
Addi, M., Gyasi-Agyei, Y., Obuobie, E., & Amekudzi, L. K.. Evaluation of imputation techniques for infilling missing daily rainfall records on river basins in Ghana. Hydrological Sciences Journal, 67(4), 613–627 (2022). https://doi.org/10.1080/02626667.2022.2030868
Aieb A, Madani K, Scarpa M, Bonaccorso B, and Lefsih K 2019 A new approach for processing climate missing databases applied to daily rainfall data in Soummam watershed, Algeria. Heliyon, 5:1–27. https://doi.org/10.1016/j.heliyon.2019.e01247
Amin Burhanuddin S N Z, Deni S and Mohamed Ramli N 2017 Imputation of missing rainfall data using the revised normal ratio method. Advanced Science Letters, 23:10981–10985, 11. DOI: 10.1166/asl.2017.10203
Batista G E A P A and Monard M C 2003 An analysis of four missing data treatment methods for supervised learning. Applied Artificial Intelligence, 17:519-533. https://doi.org/10.1080/ 713827181
Chen, FW., Liu, CW. Estimation of the spatial rainfall distribution using inverse distance weighting (IDW) in the middle of Taiwan. Paddy Water Environ 10, 209–222 (2012). https://doi.org/10. 1007/s10333-012-0319-1
Chow, V. T. (1971). Hand Book of Applied Hydrology. McGraw-Hill Book Company.
Das P, Banik P, and Rath K C 2021 Precipitation extremes and anomalies of the Indian Sundarban 1984-2018. MAUSAM, 72(4):847-858. DOI: 10.54302/mausam.v72i4.3552
De Silva R, Dayawansa N and Ratnasiri M 2007 A comparison of methods used in estimating missing rainfall data. Journal of Agricultural Sciences, 3:101–108, 05. DOI: 10.4038/jas.v3i2.8107
Fouad K M, Ismail M M Azar A T, and M. M. Arafa 2021 Advanced methods for missing values imputation based on similarity learning. Peer J Computer Science, 7:1–38. https://doi.org/ 10.7717/peerj-cs.619
Ghosh A, Schmidt S, Fickert T, and Nusser M 2015 The Indian Sundar- ban mangrove forests: history, utilization, conservation strategies, and local perception. Diversity, 7(2):149–169. https://doi.org/ 10.3390/d7020149
Khosravi G, Nafarzadegan A R, Nohegar A, Fathizad H, and Malekian A 2015 A modified distance-weighted approach for filling annual precipitation gaps: Application to different climates of Iran. Theoretical and Applied Climatology, 119(07) 33–42. https://doi.org/10.1007/s00704-014-1091-5
Lai W 2019 A Study on Sequential k-Nearest Neighbour SKNN Imputation for Treating Missing Rainfall Data. International Journal of Advanced Trends in Computer Science and Engineering, 8:363–368, 06. DOI: 10.30534/ijatcse/2019/05832019
Lettenmaier D 2024 Role of Rainfall in Hydrological Modeling and Water Resource Planning. J Geogr Nat Disasters. 14:319. DOI: 10.35841/2167-0587.24.14.319
Liu Z G, Liu Y, Dezert J and Pan Q 2015 Classification of incomplete data based on belief functions and k-nearest neighbours. Knowledge-Based Systems, 89:113–125. DOI: 10.1016/j.knosys.2015. 06.022
Lo Presti R, Barca E and Passarella G 2009 A methodology for treating missing data applied to daily rainfall data in the Candelaro River Basin (Italy). Environmental Monitoring and Assessment, 160:1–22, 01. DOI: 10.1007/s10661-008-0653-3
Meher J, Das L, Akhter J, Benestad R, and Mezghani A 2017 Performance of CMIP3 and CMIP5 GCMS to simulate observed rainfall characteristics over the Western Himalayan Region. Journal of Climate, 30(7777–7799), 09. https://doi.org/10.1175/JCLI-D-16-0774.1
Meher J and Das L 2019 Gridded data as a source of missing data replacement in station records. Journal of Earth System Science, 128:1–14, 04. https://doi.org/10.1007/s12040-019-1079-8
Pramanik M 2016 Assessment of the Impacts of Sea Level Rise on Mangrove Dynamics in the Indian part of Sundarbans using geospatial techniques. Journal of Biodiversity, Bioprospecting and Development, 14:117–127, 01. doi: 10.4172/2376-0214.1000155
Pratama I, Permanasari A E, Ardiyanto I, and Indrayani R. 2016 A review of missing values handling methods on time-series data. International Conference on Information Technology Systems and Innovation (ICITSI), pages 1–6. doi: 10.1109/ICITSI.2016.7858189.
Ramos Calzado P, G´omez Camacho J, Perez-Bernal F and Pita M 2008 A novel approach to precipitation series completion in climatological datasets: Application to Andalusia. International Journal of Climatology, 28:1525 – 1534, 09. https://doi.org/10.1002/joc.1657
Rahman M G & Islam M Z 2013 kDMI: A Novel Method for Missing Values Imputation Using Two Levels of Horizontal Partitioning in a Data set. International Conference on Advanced Data Mining and Applications. https://doi.org/10.1007/978-3-642-53917-6_23
Rahman S A, Yuxiao H, Claassen J, Heintzman N and Kleinberg S 2015 Combining Fourier and Lagged k-Nearest Neighbor Imputation for Biomedical Time Series Data. Journal of Biomedical Informatics, 58:198–207, 10. https://doi.org/10.1016/j.jbi.2015.10.004Get rights and content
Rodríguez, R.; Pastorini, M.; Etcheverry, L.; Chreties, C.; Fossati, M.; Castro, A.; Gorgoglione, A. Water-Quality Data Imputation with a High Percentage of Missing Values: A Machine Learning Approach. Sustainability 2021, 13, 6318. https://doi.org/10.3390/su13116318
Sahana M, Ahmed R and Sajjad H 2016 Analyzing land surface temperature distribution in response to land use/land cover change using split window algorithm and spectral radiance model in Sundarban Biosphere Reserve, India. Modelling Earth Systems and Environment, 2:1–11, 05. https://doi.org/10.1007/s40808-016-0135-5
Sattari M T, Rezazadeh-Joudi A and Kusiak A 2016 Assessment of different methods for estimation of missing data in precipitation studies. Hydrology Research, 48(4) 032–1044, 09. https://doi.org/ 10.2166/nh.2016.364
Teegavarapu R and Chandramouli V 2005 Improved weighting methods, Deterministic and Stochastic Data-Driven Models for Estimation of Missing Precipitation Records. Journal of Hydrology, 312:191–206, 10. https://doi.org/10.1016/j.jhydrol.2005.02.015
WWF 2017 Sundarban in a Global Perspective: Long-Term Adaptation and Development -Discussion Paper (English).
Xia Y, Fabian P, Stohl A and Winterhalter M 1999 Forest climatology: estimation of missing values for Bavaria, Germany. Agricultural and Forest Meteorology, 96:131–144.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Priyanka Das, Prof. Arup Bose, Prof. Pabitra Banik, Dr. Aditi Sarkar, Prof. Mike A Powell, Prof. Krishna Chandra Rath (Author)

This work is licensed under a Creative Commons Attribution 4.0 International License.