Early childhood education is increasingly expected to support children’s foundational scientific understanding and sustainability-oriented dispositions, yet Earth and Space Science (ESS) remains weakly assessed in early years settings. Globally, existing early childhood assessments tend to measure broad developmental domains, general science process skills, environmental attitudes, or integrated STEM engagement, leaving limited evidence on domain-specific and psychometrically tested tools for assessing young children’s ESS competence. This study developed and validated a teacher-rated observational instrument for measuring NGSS-derived ESS competencies among children aged 4 to 7 years in Indonesian kindergartens. The instrument operationalises ESS learning as observable classroom behaviours related to human-environment relations, resource use, local weather, and climate-related preparedness. Data were collected from 331 children in 16 kindergartens using an 18-item scale completed by trained teachers. Descriptive statistics, internal consistency estimation, Principal Component Analysis, and Confirmatory Factor Analysis were used to examine the instrument’s empirical structure and construct validity. The analysis supported a two-factor structure consisting of Environmental Impact and Resource Use and Weather and Climate Preparedness. This distinction suggests that young children’s ESS competence may involve two related but different forms of learning: reasoning about human-environment and resource relations, and recognising or responding to weather-related conditions and risks. The two-factor CFA model demonstrated acceptable fit (χ² = 302.461, df = 109, RMSEA = .073, CFI = .969, TLI = .957, SRMR = .036) and performed better than the one-factor alternative. By translating NGSS-derived ESS content into observable and psychometrically tested classroom indicators, the study contributes to global early childhood science assessment while offering evidence from a climate-vulnerable, non-Western context. Further research should examine inter-rater reliability, measurement invariance, predictive validity, and cross-cultural applicability.
Copyrights © 2026