With the advancements in the medical sector world-wide, the use of machine learning has been in use. In order to use these machine learning models for prediction and diagnosis of certain diseases one of the main components is datasets. The need for high-quality datasets in healthcare prediction models is critical due to the data-driven nature of machine learning. These models rely on comprehensive, accurate, and representative datasets to make reliable predictions that can impact real-world patient outcomes. This paper provides an insight about the different components in the datasets present for diseases such as osteoporosis, heart disease, diabetes, respiratory, syncytial virus, interactive thyroid, Parkinson’s and sepsis. Also, a comparative study on the parameters of the datasets in the Indian perspective and globally are also discussed.
Copyrights © 2026