Investigating How the Fractions Skill Score and Brier Divergence Skill Score Reflect Forecast Error
Abstract:
<jats:title>Abstract</jats:title> <jats:p>Meaningful scores for forecast verification are essential for developing reliable forecasts, and there has been much effort to develop scores that align well with human perceptions of forecast quality. Whilst many of these scores have intuitive interpretations, relatively little is known about how these scores rank different forecasts, and how scores reflect forecast error. We theoretically explore the behaviour of two scores that fall within the ‘neighbourhood’ paradigm of spatial verification; the Fractions Skill Score (FSS) and Brier Divergence Skill Score (BDnSS). We investigate how each score ranks forecasts with two types of error; errors in the mean frequency (corresponding to intensity or shape errors) and errors in the standard deviation (corresponding to errors in spatial structure, such as blurring or excess noise). We find that under many situations the FSS assigns higher scores to forecasts that over-predict mean frequency, thus theoretically confirming the need to use the FSS with percentile thresholds. Both scores generally assign higher scores to forecasts with lower neighbourhood standard deviation, a reflection of the ‘double penalty’ problem; however, we observe that size of this effect is larger for the BDnSS than the FSS, showing that the FSS under some situations is less susceptible to the double penalty problem than the BDnSS.</jats:p>Seasonal forecasting using the GenCast probabilistic machine learning model
Abstract:
Machine-learnt weather prediction (MLWP) models are now well established as being competitive with conventional numerical weather prediction (NWP) models in the medium range. However, there is still much uncertainty as to how this performance extends to longer timescales, where interactions with slower components of the earth system become important. We take GenCast, a state-of-the-art probabilistic MLWP model, and apply it to the task of seasonal forecasting with prescribed sea surface temperature (SST), by providing anomalies persisted over climatology (GenCast-Persisted) or forcing with observed SSTs (GenCastForced). The forecasts are compared to the European Centre for Medium-Range Weather Forecasts seasonal forecasting system, SEAS5. Our results indicate that, despite being trained at short timescales, GenCast-Persisted produces much of the correct precipitation patterns in response to El Ni˜no and La Ni˜na events, with several erroneous patterns in GenCast-Persisted corrected with GenCast-Forced. The uncertainty in precipitation response, as represented by the ensemble, compares favourably to SEAS5. Whilst SEAS5 achieves superior skill in the tropics for 2-metre temperature and mean sea level pressure (MSLP), GenCast-Persisted achieves higher skill in some areas in higher latitudes, including mountainous areas, with notable improvements for MSLP in particular; this is reflected in a slightly higher correlation with the observed NAO index. Reliability diagrams indicate that GenCast-Persisted has little skill relative to climatology, whilst GenCast-Forced produces forecasts with reliability comparable to SEAS5. These results provide an indication of the potential of MLWP models similar to GenCast for the ‘full’ seasonal forecasting problem, where the atmospheric model is coupled to ocean, land and cryosphere models.How to derive skill from the fractions skill score
Abstract:
The fractions skill score (FSS) is a widely used metric for assessing forecast skill, with applications ranging from precipitation to volcanic ash forecasts. By evaluating the fraction of grid squares exceeding a threshold in a neighborhood, the intuition is that it can avoid the pitfalls of pixelwise comparisons and identify length scales at which a forecast has skill. The FSS is typically interpreted relative to a “useful” criterion, where a forecast is considered skillful if its score exceeds a simple reference score. However, the typical reference score used is problematic, since it is not derived in a way that provides obvious meaning, does not scale with neighborhood size, and may not be exceeded by forecasts that have skill. We, therefore, provide a new method to determine forecast skill from the FSS, by deriving an expression for the FSS achieved by a random forecast, which provides a more robust and meaningful reference score to compare with. Through illustrative examples, we show that this new method considerably changes the length scales at which a forecast would be regarded as skillful and reveals subtleties in how the FSS should be interpreted.