In a technology-driven world, data is created every second, and data science has grown exponentially and is an invaluable discipline to incorporate into many fields, including environmental science. Data science can provide valuable information from data, enhancing the understanding of systems and environments. This publication is intended to give an overview of the different facets of data science and how data science can be leveraged into and applied to soil and water health. This publication will benefit policymakers, environmental consultants, agricultural managers, soil scientists, water scientists, and conservationists.
Introduction
Data science is generally understood as the practice of analyzing data to gain useful information and uncover new patterns and relationships.1 By providing practical insights, data science supports a wide range of operations across industries. Data science combines multiple fields of expertise, including statistical science, computer science, and specific field knowledge, such as water or soil science (as used as an example in this paper).2,3 Due to its interdisciplinary nature, data science is versatile and can be applied to almost any industry.3 The process typically involves collecting, cleaning, processing, analyzing, and interpreting either spatial or non-spatial data.4,5,6 Through these steps, data scientists can identify patterns and trends, infer data behavior or anomalies, and ultimately support informed decision-making.4
The United States Department of Agriculture (USDA) defines soil health as “the continued capacity of soil to function as a vital living ecosystem that sustains plants, animals, and humans”.7 Many factors, such as soil structure, organic matter content, pH, cation exchange capacity (CEC), hydrologic group, bulk density and nutrient-carrying capacity, etc., influence soil health.8,9 For the sustainable production of different crops and vegetation, soil health requirements vary. Similarly, water quality, which is a measurement of how suitable water is for a specific use, can also vary.10 In other words, the conditions for water quality may change depending on the intended use.11 For example, drinking water has standards different from stormwater and water used for irrigation and recreational purposes.11 Conditions such as turbidity,10 dissolved oxygen,10,11 nutrients,10,11 fecal coliform bacteria11 and hardness10,11 influence water quality.
Soil health and water quality are intrinsically linked, as healthy soils help to keep water sources clean. For example, soils with high organic matter content and high aggregate stability can absorb and retain more water than soils with low organic matter content and low aggregate stability.12 This helps to reduce the amount of water lost via runoff during heavy rain events.12 When runoff is reduced, less sediments and nutrients get transported to waterways, helping to keep the water sources clean.12 Soil also acts as a filter and can remove contaminants from water as the water moves through the soil.12
As the world faces challenges associated with climate change, environmental degradation, and a growing population, understanding the complexity of soil health and water quality is more important than ever.13 Doing so will help policymakers, landowners, and environmentalists make educated decisions to protect the valuable resources that the world depends on. This need for decision making is underscored by the vast amount of data available for soil health and water quality. For example, there are approximately 19,000 different soil series globally with around 17,000 found in the United States alone.14 The latest spatial soil map of the United States, a 10-meter spatial resolution gSSURGO database, contains close to 200 soil attributes of importance and the USDA-NRCS gSSURGO Value Added Look Table highlights 60 attributes commonly used by general users.15 Similarly, while most decisions about water quality focus on a core set of water quality parameters, the USGS, National Water Information System provides over 800 different water related parameters in the Parameter Code Definition spreadsheet.16 This represents the enormity of soil health and water quality spatial and nonspatial data availability. Therefore, data science is one discipline in particular that can be integrated and used to better understand soil and water health. In fact, data science is already being applied to some aspects of environmental science, such as agriculture, and as a result, concepts such as “smart farming” have been developed.
Data Types
Data science incorporates and relies on many data types, each offering unique insights. Commonly used data types for understanding soil and water health can be categorized into four groups: sensor-based data, remote sensing data, data collected from traditional field measurements, and geospatial data.
Sensor-Based Data
Sensor-based data involves using sensors to measure a parameter and collect data accurately.17 Commonly used sensors, such as time domain reflectance (TDR), frequency domain reflectance (FDR), and electrical conductivity sensors, utilize electrodes to produce an electric current.18 By producing an electric current, the sensors can measure the electrical conductivity, resistivity, and/or permittivity within a medium.18 This allows a sensor to directly or indirectly measure a desired parameter, such as soil moisture and salinity.18 Additionally, Internet of Things (IoT) devices can be used in conjunction with sensors. IoT devices allow sensors to share data and communicate with other devices through the internet without human interactions.19 For example, water sensors can be installed in a crop field and, through IoT devices, can communicate to irrigation systems when dry conditions persist so that the irrigation system turns on, keeping plants from becoming water-stressed.20 Many parameters can be measured with sensors, including soil temperature, soil moisture, soil pH, air temperature, humidity, rainfall, and wind speed.17 Figure 1 below shows examples of in situ (a,b) and mobile (c,d) sensors that are collecting soil moisture data in the field.

Figure 1. Examples of sensors used for collecting data include installing Time Domain Reflectometry (TDR) sensors integrated with Internet of Things (IoT) technology for collecting soil moisture and temperature data in fields as part of the Precision Sustainable Agriculture Project (a, b) as well as using handheld sensors such as the FieldScout TDR 150 soil moisture meter (c, d). Photo credit: Payton Davis.
Remote Sensing Data
Remote sensing technologies have advanced over the years and have proven to be an essential resource for gathering information about the Earth’s surface.21 Remote sensing technologies collect spectral, spatial, temporal, and radiometric data from sensors mounted on satellites, drones, and aircraft through electromagnetic radiation (Figure 2).22,23 Sensors used for remote sensing can be classified into one of two types: active and passive.23 Active sensors produce and utilize their own electromagnetic radiation to collect data, while passive sensors detect and record natural radiation.23,24 Active sensor technology such as Synthetic Aperture Radar (SAR), and LiDAR, helps with mapping land surfaces, vegetation, and detecting objects; while passive sensors, such as Landsat and ECOSTRESS are used to characterized and monitor soil and water health dynamics and can help identify environmental stress such as water and heat stress on plants and wildfires.25,26,27 Additionally, remote sensing technologies detect many types of electromagnetic radiation, allowing for many different kinds of data to be obtained, leading to a wide range of uses and applications.23 For example, remote sensing has been used to obtain data on soil and terrain mapping,28 land cover,29 water quality30 and biophysical plant features.24

Figure 2. Examples of remote sensing data collected from a drone using a multispectral camera in a cover crop research field in Clemson, SC. Figure 2(a) shows an image that captures reflected near-infrared radiation. Figure 2(b) shows an image taken using a Normalized Difference Vegetation Index camera that captures the near-infrared and visible light reflected by plants. Photo credit: Coleman Scroggs.
Traditional Field Measurements
Traditional field measurements remain essential for validating models and data collected through sensors and remote sensing technologies.31,32 Traditional field measurements for soil and water health include collecting and monitoring water and soil samples for biological, physical, and chemical properties in situ (such as infiltration measurements (Figure 3), canopy cover assessments, and topographic surveys), as well as brought back to a laboratory for analysis (such as nutrient content determinations and microbiological assessment). Data science can use and combine these data types (sensor, remote sensing, and traditional field measurements) to develop a comprehensive understanding of soil and water health, enabling more informed decisions to be made, including decisions about management practices. When collecting data, it is also important to record information about the data and what methods and devices were used to collect the data. This information is called metadata.33 Metadata is a crucial data collection functionality, allowing users to share information, compare data, and find points of interest within the dataset.33

Figure 3. Examples of measuring infiltration in the field using tools such as a mini disk infiltrometer (a) and a double ring infiltrometer (b). Photo credit: Payton Davis.
Geospatial Data
Geospatial data refers to information that is specific to a location or coordinates.34 In other words, any data that is tied to a specific point on earth is considered geospatial data.35 Geospatial data is not refined to one type of data but rather encompasses a variety of data types including but limited to images, videos, vector, raster, digital, and non-digital data .35 In the environmental industry, there are many open access databases that can be used to retrieve geospatial data and have been created for stakeholders’ efficient usage. Some examples include:
The STATSGO2 database provides general information on soil properties and provides interpretation for water management and environmental assessment.36 The data provided expands across the United States and is primarily vector and tabular datasets that can be downloaded as shapefiles.36
The SSURGO database provides soil data that has been collected from the National Cooperative Soil Survey (NCSS), which is a nationwide partnership between federal, regional, state, and local agencies.37 Over the past century, the agencies working under NCSS have surveyed and analyzed many soil samples, providing information such as available water capacity, landscape position, crop yields, and limitations to development.37 The datasets found in SSURGO include map and tubular data which can be downloaded as shapefiles.37 The gSSURGO database was developed from the SSURGO database and provides information on soil properties collected from NCSS.38,39 However, gSSURGO contains raster datasets and is formatted as an Environmental Systems Research Institute, Inc. file database.38,39 More information on the gSSURGO database is available here.
Additionally, the Environmental Protection Agency (EPA) provides many geospatial datasets on water quality parameters which can be found here.40 The EPA updates these datasets using state-submitted data as it becomes available.40 One database worth mentioning is the EPA Total Maximum Daily Loads 303D spatial database.40 This database is updated every two years and contains information on impaired waterways and the maximum amount of pollutants a waterway can sustain daily and still meet water quality standards.40
Data Analysis
Data type is just one aspect of data science. The data needs to be analyzed effectively to derive meaningful insights and advance the understanding of complex systems. Three common approaches to data analysis include big data processing, machine learning algorithms, and predictive modeling.
Big Data Processing
Big data processing involves handling extremely large and complex data sets traditionally characterized by the data’s variety, volume, and velocity.41
- Data variety: the different types of data included in the data set.41
- Data volume: the amount of data being processed.41
- Data velocity: the speed at which data is generated, how the data is constantly changing, and the need to process and analyze the data instantly to provide real-time solutions.41
Big data processing involves three main techniques: batch processing, real-time processing, and hybrid computation.41
- Batch processing: analyzes static data that is already stored.41 Once batch processing has started, no new data can be incorporated into the analysis.
- Real-time processing: analyzes small data sets to give real-time results.41
- Hybrid computation: combines batch processing and real-time processing so that both static and new data can be processed.41
A variety of technologies and tools are available to perform these analyses, each suited to different types of data and processing needs.41 Some examples include using MapReduce for batch processing or using Storm (The Apache Software Foundation, Wilmington, DE) to analyze real-time data.41
Machine Learning Algorithms
Machine learning algorithms refer to encoded procedures that can sort through data42 and use data to develop models.43 The data used for developing the machine learning algorithms is called “training data,” as the data is essentially training the model.43 As more data becomes available, the model can adjust and improve its capabilities.43 Therefore, it is important to have a large and consistent data set available when using machine learning algorithms.43 Once the machine learning algorithms have been trained, the model can be used to make predictions, understand relationships between variables, identify differences, and even discover new materials.43 Just as there are many ways to approach a problem, many different algorithms exist and can be used in machine learning.44 For a review of common machine learning algorithms, please see Machine Learning Algorithms- A Review.44 Figure 4 below illustrates the workflow for developing a model using machine learning. Machine learning must be performed first to train the model, after which the model can be used.

Figure 4. Simplified workflow for the development of a machine learning model adapted from Zhong et al..43
Predictive Modeling
As defined by Kuhn and Johnson (2013), predictive modeling is “the process of developing a mathematical tool or model that generates an accurate prediction”.45 In other words, predictive modeling is used to forecast an unknown outcome.45 Such models can be developed by using machine learning algorithms, as mentioned above, or statistical analysis using packages such as JMP or R.45 In both cases, various models, including regression and time series models, can be developed. While predictive modeling does not explain why an event does or does not occur, predictive modeling can provide valuable insights into what to expect in various scenarios.45 Predictive modeling may be particularly useful when identifying how climate change impacts soil and water health.
By using and integrating analytical approaches such as the ones explained above, patterns between multiple variables can be uncovered, and data can be transformed into actionable insights. The choice of which analytical method to use depends on the project’s specific requirements, the research question, the nature of the data, and the desired outcomes.
Roles and Applications of Data Science for Soil and Water Health
Soil is a complex system influenced by many parameters, including its physical structure, chemical composition, biological activity, and environmental conditions.8 Additionally, soil characteristics can also change with depth.46 These parameters can interact in complex ways, creating a dynamic and interdependent network of variables.8 For instance, soil water content can affect microbial activity,47 which influences nutrient availability in soils and plant growth.48 Similarly, soil pH can impact minerals’ solubility, affecting soil nutrient availability.49 The relationships between soil parameters can also be non-linear.50,51,52 This, combined with spatial variability of soil properties and the vast amount of data available, makes it challenging to accurately model and predict soil health.53
The complex relationships within soil health, therefore, require data science. Data science can handle large, diverse datasets and use advanced analytical methods to uncover patterns, correlations, and causal relationships within the soil.54 By leveraging data science into soil health, researchers and stakeholders can gain deeper insights into soil health, optimize agricultural practices, and develop more sustainable land management strategies. For example, soil spectroscopy datasets have been used to develop a model that predicts soil organic carbon, cation exchange capacity, clay content, sand content, soil pH, and total nitrogen from spectral data.55 Understanding these soil parameters can help identify soil health and be used to determine the suitability of the soil for different land uses and crops.56,57 For example, knowing that a soil has a low pH can inform a producer what to plant, as some crops, such as blueberries, grow better in acidic soils.58 Traditionally, to determine the soil pH, a producer would have to take a soil sample and test the pH themselves or send the soil sample to the lab for analysis. For large fields, this can be time-consuming and costly. By utilizing data science, information such as soil pH can be obtained from models developed from soil spectroscopy datasets. Data science can also be applied to model the fate and transport of soil contaminants, which can help in remediation processes.59 Such models can use real-time data from soil sensors with the combination of different data analyses to give environmentalists a deep understanding of how contaminants move through the soil.59
Additionally, data science has been applied to agricultural data to develop the concept of “smart farming,” which enhances agricultural production and sustainability.17 Some examples include using models developed by data science to predict yields, detect crop diseases and pests, and identify nutrient deficiencies.17 Developing and using models like these can help save money, time, and resources, as information can be gained without constantly taking samples and conducting laboratory analyses.
Similar to soil health, water quality is influenced by many factors and water quality values can fluctuate throughout the vertical profiles of bodies of waters such as rivers, ponds and lakes.60 Climatic conditions, topography, weather, and land use (including urbanization) are some factors influencing water quality.61 Due to this complexity, and the enormous amount of data that can be obtained, data science has also proved valuable in understanding water quality and health. For example, water quality data sets have been used to develop models that predict the biochemical oxygen demand,62 an important indicator of water quality that can impair water use.63 Another example is using data science to create predictive runoff models.64 Han & Morrison developed a model that uses rainfall, soil moisture, baseflow and surface temperature to predict the hourly runoff at the Russian River basin in California.64 Such models provide insights crucial for water management and resource planning.64 Data science has also been used to develop models that optimize irrigation scheduling by using data such as soil moisture, soil temperature, weather parameters, and environmental conditions.65 By optimizing irrigation scheduling, farming systems and natural resources can become more sustainable.
Challenges with Data Science
While data science is a powerful discipline that can be incorporated to better understand soil and water health, it comes with several challenges.66 A primary concern is the quality and reliability of data.43 The effectiveness and performance of any model are directly tied to the quality of its data, making it crucial to ensure that the data used is accurate, complete, and of high quality.67 This can be particularly challenging in environmental monitoring, where even sensor-collected data can be prone to inaccuracies. For instance, TDR sensors used for soil moisture measurements can produce measurement errors due to factors like air pockets,68 bent probes,69 and improper probe installation.70 Additionally, there is no one way to collect and analyze water samples; instead, multiple techniques and protocols exist, with some suited to different water quality and environmental conditions. The instruments used to measure environmental parameters can also change over the years, resulting in different precision and accuracy levels. This variety can lead to significant variation within the data, affecting the data quality and creating discrepancies among data sets.43
Data security presents another significant challenge. Information leakage and cyber-attacks can occur as technology is involved in many aspects of data science, including data collection, storage, and analysis.71 Both can lead to data loss and data manipulation, which endangers data integrity.71
Interpreting results, a major goal of data science, can also be a challenge. Some relationships that models depict may not make sense in the real world and thus require further evaluation.72 This can be the outcome of not using the appropriate analytical approaches or could result from having irrelevant and/or incomplete data in the model.72 It is also essential to acknowledge that correlation does not imply causation.73 Just because two variables appear to be correlated does not mean that one variable directly causes a particular result in another variable.73
Additionally, there is a growing mistrust of science by the public.74,75 This creates a challenge as the public will be less likely to trust the results generated by data science. Mistrust in science stems from many factors, including religious and political beliefs, level of education, socioeconomics, and failure to communicate science effectively.75 Effectively communicating science is particularly difficult as the general population often does not have a scientific background and may not understand the scientific process and results.76 Therefore, when a study releases new results that differ from previous results, it is hard for the general public to understand why, as the scientific process is often not explained. Emotion is also a driving force in the mistrust of science.77 Furman (2020) explains that trust is based on expectations and results.77 One expects something to happen if they do what is recommended. When someone does what is recommended, and the result differs from what they expect, they feel betrayed and lack the motivation to trust science.77 For example, a farmer may read a scientific article demonstrating that using cover crops improves certain soil properties and increases cash crop yields. If the farmer uses cover crops that year and sees a decrease in cash crop yield, the farmer may feel betrayed and not trust science because the farmer expected that using cover crops would result in increased yields. However, many times, especially with environmental science, it may take many years to see results, or recommendations may change due to the dynamic nature of the environment. For example, in the cover crop scenario, it may have been an unusually dry year. The lack of rain could have contributed to the decrease in yields.
Furthermore, the computational demands of processing large datasets and training sophisticated models can be substantial.43 This can require a lot of time and often requires access to powerful computing resources, which can be expensive.78 These challenges highlight the need for careful consideration and robust methodologies in applying data science to soil and water health.
Conclusion
Data science is an expanding field that enables the analysis of vast and diverse datasets, significantly enhancing our understanding of complex systems. The application of data science to understand soil and water health represents a significant advancement in the environmental field. Data science has enabled researchers to develop predictive models for soil properties, create advanced runoff models, and optimize irrigation scheduling. These applications demonstrate the power of data science in using and transforming data into valuable knowledge that can be acted upon. There are, however, challenges to be aware of, including ensuring data quality and security and increasing the general public’s trust in science. As methodologies improve and current challenges are addressed, data science will continue to play a crucial role in ensuring the health and sustainability of resources such as soil and water.
References Cited
- Zhu, Y., & Xiong, Y. (2015). Defining data science—Beyond the study of the rules of the natural world as reflected by data. arXiv. https://arxiv.org/abs/1501.05039
- Oliver, J. C., & McNeil, T. (2021). Undergraduate data science degrees emphasize computer science and statistics but fall short in ethics training and domain-specific context. PeerJ Computer Science, 7, e441. https://doi.org/10.7717/peerj-cs.441
- Martinez, I., Viles, E., & Olaizola, I. G. (2021). Data science methodologies: Current challenges and future approaches. Big Data Research, 24, 100183. https://doi.org/10.1016/j.bdr.2020.100183
- Provost, F., & Fawcett, T. (2013). Data science and its relationship to big data and data-driven decision making. Big Data, 1(1), 51–59. https://doi.org/10.1089/big.2013.1508
- Chai, C. P. (2020). The importance of data cleaning: Three visualization examples. CHANCE, 33(1), 4–9. https://doi.org/10.1080/09332480.2020.1726112
- Ozdemir, S. (2016). Principles of data science. Packt Publishing.
- U.S. Department of Agriculture, Natural Resources Conservation Service. (2024). Soil health. https://www.nrcs.usda.gov/conservation-basics/natural-resource-concerns/soils/soil-health
- Kumar, K. S. A., & Karthika, K. S. (2020). Abiotic and biotic factors influencing soil health and/or soil degradation. In B. Giri & A. Varma (Eds.), Soil health (pp. 145–161). Springer. https://doi.org/10.1007/978-3-030-44364-1_9
- Kibblewhite, M. G., Ritz, K., & Swift, M. J. (2008). Soil health in agricultural systems. Philosophical Transactions of the Royal Society B: Biological Sciences, 363(1492), 685–701. https://doi.org/10.1098/rstb.2007.2178
- Omer, N. H. (2019). Water quality parameters. In K. Summers (Ed.), Water quality—Science, assessments and policy. IntechOpen. https://doi.org/10.5772/intechopen.89657
- Thatai, S., Verma, R., Khurana, P., Goel, P., & Kumar, D. (2018). Water quality standards, its pollution and treatment methods. In M. Naushad (Ed.), A new generation material graphene: Applications in water technology (pp. 21–42). Springer. https://doi.org/10.1007/978-3-319-75484-0_2
- Keesstra, S., Sannigrahi, S., López-Vicente, M., Pulido, M., Novara, A., Visser, S., & Kalantari, Z. (2021). The role of soils in regulation and provision of blue and green water. Philosophical Transactions of the Royal Society B: Biological Sciences, 376(1834), 20200175. https://doi.org/10.1098/rstb.2020.0175
- Gomiero, T. (2016). Soil degradation, land scarcity and food security: Reviewing a complex challenge. Sustainability, 8(3), 281. https://doi.org/10.3390/su8030281
- Soil Survey Staff, Natural Resources Conservation Service, U.S. Department of Agriculture. (n.d.). Soil series classification database. Retrieved May 1, 2025, from https://www.nrcs.usda.gov/resources/data-and-reports/soil-series-classification-database-sc
- National Cooperative Soil Survey. (n.d.). NCSS soil characterization database (Lab Data Mart). Retrieved May 1, 2025, from https://ncsslabdatamart.sc.egov.usda.gov/
- U.S. Geological Survey. (n.d.). USGS Water Data for the Nation help. Retrieved May 1, 2025, from https://help.waterdata.usgs.gov/parameter_cd?group_cd=PHY
- Johnson, N., Kumar, M. S., & Dhannia, T. (2020). A study on the significance of smart IoT sensors and data science in digital agriculture. In 2020 Advanced Computing and Communication Technologies for High Performance Applications (ACCTHPA) (pp. 80–88). IEEE. https://doi.org/10.1109/ACCTHPA49271.2020.9213207
- Kuang, B., Mahmood, H. S., Quraishi, M. Z., Hoogmoed, W. B., Mouazen, A. M., & van Henten, E. J. (2012). Sensing soil properties in the laboratory, in situ, and on-line: A review. Advances in Agronomy, 114, 155–223. https://doi.org/10.1016/B978-0-12-394275-3.00003-1
- Gazis, A. (2021). What is IoT? The Internet of Things explained. Academia Letters, Article 1003. https://doi.org/10.20935/AL1003
- Muangprathub, J., Boonnam, N., Kajornkasirat, S., Lekbangpong, N., Wanichsombat, A., & Nillaor, P. (2019). IoT and agriculture data analysis for smart farm. Computers and Electronics in Agriculture, 156, 467–474. https://doi.org/10.1016/j.compag.2018.12.011
- Liu, P. (2015). A survey of remote-sensing big data. Frontiers in Environmental Science, 3, Article 45. https://doi.org/10.3389/fenvs.2015.00045
- Salaria, M., & Suresh, A. (1990). Remote sensing in agriculture. Smart and Sustainable Agricultural Technology, 408.
- Navalgund, R. R., Jayaraman, V., & Roy, P. S. (2007). Remote sensing applications: An overview. Current Science, 93(12), 1747–1766. http://www.jstor.org/stable/24102069
- Wójtowicz, M., Wójtowicz, A., & Piekarczyk, J. (2016). Application of remote sensing methods in agriculture. Communications in Biometry and Crop Science, 11, 31–50.
- Janga, B., Asamani, G. P., Sun, Z., & Cristea, N. (2023). A review of practical AI for remote sensing in Earth sciences. Remote Sensing, 15(16), 4112. https://doi.org/10.3390/rs15164112
- Fisher, J. B., Lee, B., Purdy, A. J., Halverson, G. H., Dohlen, M. B., Cawse-Nicholson, K., Wang, A., Anderson, R. G., Aragon, B., Arain, M. A., & Baldocchi, D. D. (2020). ECOSTRESS: NASA’s next generation mission to measure evapotranspiration from the International Space Station. Water Resources Research, 56(4), e2019WR026058. https://doi.org/10.1029/2019WR026058
- Wulder, M. A., Roy, D. P., Radeloff, V. C., Loveland, T. R., Anderson, M. C., Johnson, D. M., Healey, S., Zhu, Z., Scambos, T. A., Pahlevan, N., & Hansen, M. (2022). Fifty years of Landsat science and impacts. Remote Sensing of Environment, 280, 113195. https://doi.org/10.1016/j.rse.2022.113195
- Mulder, V. L., de Bruin, S., Schaepman, M. E., & Mayr, T. R. (2011). The use of remote sensing in soil and terrain mapping—A review. Geoderma, 162(1–2), 1–19. https://doi.org/10.1016/j.geoderma.2010.12.018
- Boyle, S. A., Kennedy, C. M., Torres, J., Colman, K., Pérez-Estigarribia, P. E., & de la Sancha, N. U. (2014). High-resolution satellite imagery is an important yet underutilized resource in conservation biology. PLOS ONE, 9(1), e86908. https://doi.org/10.1371/journal.pone.0086908
- Ritchie, J. C., Zimba, P. V., & Everitt, J. H. (2003). Remote sensing techniques to assess water quality. Photogrammetric Engineering & Remote Sensing, 69(6), 695–704. https://doi.org/10.14358/PERS.69.6.695
- Zobeck, T. M., Sterk, G., Funk, R., Rajot, J. L., Stout, J. E., & Van Pelt, R. S. (2003). Measurement and data analysis methods for field-scale wind erosion studies and model validation. Earth Surface Processes and Landforms, 28(11), 1163–1188. https://doi.org/10.1002/esp.1033
- Wu, B., Yan, N., Xiong, J., Bastiaanssen, W. G. M., Zhu, W., & Stein, A. (2012). Validation of ETWatch using field measurements at diverse landscapes: A case study in Hai Basin of China. Journal of Hydrology, 436–437, 67–80. https://doi.org/10.1016/j.jhydrol.2012.02.043
- Riley, J. (2017). Understanding metadata: What is metadata, and what is it for? National Information Standards Organization.
- Praveen, P., Babu, C. J., & Rama, B. (2016). Big data environment for geospatial data analysis. In 2016 International Conference on Communication and Electronics Systems (ICCES) (pp. 1–6). IEEE. https://doi.org/10.1109/CESYS.2016.7889816
- Selmy, S. A., Kucher, D. E., Yang, Y., & García-Navarro, F. J. (2024). Geospatial data: Acquisition, applications, and challenges. In Exploring remote sensing—Methods and applications. IntechOpen. https://doi.org/10.5772/intechopen.1006635
- U.S. Department of Agriculture, Natural Resources Conservation Service. (n.d.). Description of STATSGO2 database. Retrieved May 1, 2025, from https://www.nrcs.usda.gov/resources/data-and-reports/description-of-statsgo2-database
- U.S. Department of Agriculture, Natural Resources Conservation Service. (n.d.). Soil Survey Geographic Database (SSURGO). Retrieved May 1, 2025, from https://www.nrcs.usda.gov/resources/data-and-reports/soil-survey-geographic-database-ssurgo
- U.S. Department of Agriculture, Natural Resources Conservation Service. (n.d.). Gridded Soil Survey Geographic (gSSURGO) database. Retrieved May 1, 2025, from https://www.nrcs.usda.gov/resources/data-and-reports/gridded-soil-survey-geographic-gssurgo-database
- U.S. Department of Agriculture, Natural Resources Conservation Service. (n.d.). gSSURGO factsheet. Retrieved May 1, 2025, from https://www.nrcs.usda.gov/sites/default/files/2022-08/gSSURGO_Factsheet.pdf
- U.S. Environmental Protection Agency. (n.d.). WATERS geospatial data downloads. Retrieved May 1, 2025, from https://www.epa.gov/waterdata/waters-geospatial-data-downloads
- Casado, R., & Younas, M. (2015). Emerging trends and technologies in big data processing. Concurrency and Computation: Practice and Experience, 27(8), 2078–2091. https://doi.org/10.1002/cpe.3398
- Gillespie, T. (2014). The relevance of algorithms. In T. Gillespie, P. J. Boczkowski, & K. A. Foot (Eds.), Media technologies: Essays on communication, materiality, and society. MIT Press. https://doi.org/10.7551/mitpress/9042.003.0013
- Zhong, S., Zhang, K., Bagheri, M., Burken, J. G., Gu, A., Li, B., Ma, X., Marrone, B. L., Ren, Z. J., Schrier, J., Shi, W., Tan, H., Wang, Y., Wang, X., Wong, B., Xiao, X., Yu, X., Zhu, J.-J., & Zhang, H. (2021). Machine learning: New ideas and tools. Environmental Science & Technology, 55(19), 12741–12754. https://doi.org/10.1021/acs.est.1c01339
- Mahesh, B. (2020). Machine learning algorithms—A review. International Journal of Science and Research, 9(1), 381–386. https://doi.org/10.21275/ART20203995
- Kuhn, M., & Johnson, K. (2013). Applied predictive modeling. Springer.
- Wang, H. M., Wang, W. J., Chen, H., Zhang, Z., Mao, Z., & Zu, Y. G. (2014). Temporal changes of soil physic-chemical properties at different soil depths during larch afforestation by multivariate analysis of covariance. Ecology and Evolution, 4(7), 1039–1048. https://doi.org/10.1002/ece3.947
- Geisseler, D., Horwath, W. R., & Scow, K. M. (2011). Soil moisture and plant residue addition interact in their effect on extracellular enzyme activity. Pedobiologia, 54(2), 71–78. https://doi.org/10.1016/j.pedobi.2010.10.001
- Miransari, M. (2013). Soil microbes and the availability of soil nutrients. Acta Physiologiae Plantarum, 35, 3075–3084. https://doi.org/10.1007/s11738-013-1338-2
- Gondal, A. H., Hussain, I., Ijaz, A. B., Zafar, A., Ch, B. I., Zafar, H., Sohail, M. D., Niazi, H., Touseef, M., Khan, A. A., & Tariq, M. (2021). Influence of soil pH and microbes on mineral solubility and plant nutrition: A review. International Journal of Agriculture and Biological Sciences, 5(1), 71–81.
- Beresnev, I. A., & Wen, K.-L. (1996). Nonlinear soil response—A reality? Bulletin of the Seismological Society of America, 86(6), 1964–1978. https://doi.org/10.1785/BSSA0860061964
- Atkinson, J. H. (2000). Non-linear soil stiffness in routine design. Géotechnique, 50(5), 487–508. https://doi.org/10.1680/geot.2000.50.5.487
- Wen, Y., Su, L. M., Qin, W. C., Fu, L., He, J., & Zhao, Y. H. (2012). Linear and non-linear relationships between soil sorption and hydrophobicity: Model, validation and influencing factors. Chemosphere, 86(6), 634–640. https://doi.org/10.1016/j.chemosphere.2011.11.001
- Lin, H., Wheeler, D., Bell, J., & Wilding, L. (2005). Assessment of soil spatial variability at multiple scales. Ecological Modelling, 182(3–4), 271–290. https://doi.org/10.1016/j.ecolmodel.2004.04.006
- Cielen, D., & Meysman, A. (2016). Introducing data science: Big data, machine learning, and more, using Python tools. Manning.
- Padarian, J., Minasny, B., & McBratney, A. B. (2019). Using deep learning to predict soil properties from regional spectral data. Geoderma Regional, 16, e00198. https://doi.org/10.1016/j.geodrs.2018.e00198
- Karlen, D. L., Ditzler, C. A., & Andrews, S. S. (2003). Soil quality: Why and how? Geoderma, 114(3–4), 145–156. https://doi.org/10.1016/S0016-7061(03)00039-9
- Singh, B. T., Devi, K. N., Kumar, B. Y., Bishworjit, N., Singh, L. N. K., & Singh, H. A. (2013). Characterization and evaluation for crop suitability in lateritic soils. African Journal of Agricultural Research, 8(37), 4628–4636. https://doi.org/10.5897/AJAR12013.7309
- Yang, H., Wu, Y., Zhang, C., Wu, W., Lyu, L., & Li, W. (2022). Growth and physiological characteristics of four blueberry cultivars under different high soil pH treatments. Environmental and Experimental Botany, 197, 104842. https://doi.org/10.1016/j.envexpbot.2022.104842
- Fan, Y., Wang, X., Funk, T., Rashid, I., Herman, B., Bompoti, N., Mahmud, M. S., Chrysochoou, M., Yang, M., Vadas, T. M., & Lei, Y. (2022). A critical review for real-time continuous soil monitoring: Advantages, challenges, and perspectives. Environmental Science & Technology, 56(19), 13546–13564. https://doi.org/10.1021/acs.est.2c03562
- Mukhopadhyay, G., Mondal, D., Biswas, P., & Dewanji, A. (2004). Water quality monitoring of tropical ponds: Location and depth effect in two case studies. Acta Hydrochimica et Hydrobiologica, 32(2), 138–148. https://doi.org/10.1002/aheh.200300524
- Anh, N. T., Nhan, N. T., Schmalz, B., & Le Luu, T. (2023). Influences of key factors on river water quality in urban and rural areas: A review. Case Studies in Chemical and Environmental Engineering, 8, 100424. https://doi.org/10.1016/j.cscee.2023.100424
- Li, X., & Song, J. (2015). A new ANN-Markov chain methodology for water quality prediction. In 2015 International Joint Conference on Neural Networks (IJCNN) (pp. 1–6). IEEE. https://doi.org/10.1109/IJCNN.2015.7280320
- Vigiak, O., Grizzetti, B., Udias-Moinelo, A., Zanni, M., Dorati, C., Bouraoui, F., & Pistocchi, A. (2019). Predicting biochemical oxygen demand in European freshwater bodies. Science of the Total Environment, 666, 1089–1105. https://doi.org/10.1016/j.scitotenv.2019.02.252
- Han, H., & Morrison, R. R. (2022). Data-driven approaches for runoff prediction using distributed data. Stochastic Environmental Research and Risk Assessment, 36, 2153–2171. https://doi.org/10.1007/s00477-021-01993-3
- Vianny, D. M. M., John, A., Mohan, S. K., Sarlan, A., & Ahmadian, A. (2022). Water optimization technique for precision irrigation system using IoT and machine learning. Sustainable Energy Technologies and Assessments, 52, 102307. https://doi.org/10.1016/j.seta.2022.102307
- Wu, Y., Zhang, Z., Kou, G., Zhang, H., Chao, X., Li, C.-C., Dong, Y., & Herrera, F. (2021). Distributed linguistic representations in decision making: Taxonomy, key elements and applications, and challenges in data science and explainable artificial intelligence. Information Fusion, 65, 165–178. https://doi.org/10.1016/j.inffus.2020.08.018
- Gan, T. Y., Dlamini, E. M., & Biftu, G. F. (1997). Effects of model complexity and structure, data quality, and objective functions on hydrologic modeling. Journal of Hydrology, 192(1–4), 81–103. https://doi.org/10.1016/S0022-1694(96)03114-9
- Pelletier, M. G., Karthikeyan, S., Green, T. R., Schwartz, R. C., Wanjura, J. D., & Holt, G. A. (2012). Soil moisture sensing via swept frequency based microwave sensors. Sensors, 12(1), 753–767. https://doi.org/10.3390/s120100753
- Wen, M. M., Liu, G., Horton, R., & Noborio, K. (2018). An in situ probe-spacing-correction thermo-TDR sensor to measure soil water content accurately. European Journal of Soil Science, 69(6), 1030–1034. https://doi.org/10.1111/ejss.12718
- Skierucha, W. (2000). Accuracy of soil moisture measurement by TDR technique. International Agrophysics, 14(4), 417-426.
- Gupta, M., Abdelsalam, M., Khorsandroo, S., & Mittal, S. (2020). Security and privacy in smart farming: Challenges and opportunities. IEEE Access, 8, 34564–34584. https://doi.org/10.1109/ACCESS.2020.2975142
- Sarker, I. H. (2021). Data science and analytics: An overview from data-driven smart computing, decision-making and applications perspective. SN Computer Science, 2(5), 377. https://doi.org/10.1007/s42979-021-00765-8
- Barrowman, N. (2014). Correlation, causation, and confusion. The New Atlantis, 43, 23–44. https://www.jstor.org/stable/43551404
- Kennedy, B., & Tyson, A. (2023). Americans’ trust in scientists, positive views of science continue to decline. Pew Research Center.
- Kabat, G. C. (2017). Taking distrust of science seriously: To overcome public distrust in science, scientists need to stop pretending that there is a scientific consensus on controversial issues when there is not. EMBO Reports, 18(7), 1052–1055. https://doi.org/10.15252/embr.201744294
- Wilbur, E. M. (2019). Where we fail: An examination of scientific miscommunication. University of Washington.
- Furman, K. (2020). Emotions and distrust in science. International Journal of Philosophical Studies, 28(5), 713–730. https://doi.org/10.1080/09672559.2020.1846281
- Kim, M., Zimmermann, T., DeLine, R., & Begel, A. (2018). Data scientists in software teams: State of the art and challenges. IEEE Transactions on Software Engineering, 44(11), 1024–1038. https://doi.org/10.1109/TSE.2017.2754374
