Tuesday, February 17, 2015
Plot of the Global Ocean Anomaly from 1880 to 2014
I computed a similar forecast for the global ocean temperature anomaly using data from Jan 1880 to the Aug 2013 and got similar results. The monthly data used, 20 year averages, fit and error bounds are shown in red and the additional data to the end of 2014 is in blue.
The actual anomaly data looks more like a random walk in this case and there is a more pronounced decrease in the projected temperature anomaly. It remains to be seen how reliable the prediction is. The global ocean anomaly appears to be much more stable than the global land anomaly.
Wednesday, February 11, 2015
Comparison of the 2014 Global Land Anomalies with a 2013 Prediction
The observed global land temperature anomalies agree quite well with the a prediction made in October 2013. The observed values are within the error bounds of the computed curve. The forecast was based on the data in red and the blue points are for the observed data after the prediction. The solid red curve is a smoothed version of the curve obtained by taking 20 year averages. Combining the deviations from the 20 year with the 20 year average produced a much smoother curve. The doted lines are the fit to smoothed curve and the error bounds.
Supplemental (Feb 11): One might ask why the "corrected" 20 year average works so well and the answer might be that there is a feedback mechanism involved. The global temperatures are somewhat volatile but there also appears to be some resistance to change which is what one would expect if the temperatures are near an equilibrium point. Canceling the deviations from the 20 year average is a mathematical way of adding some resistance to change which results in a stiffened curve.
Saturday, December 13, 2014
Systematic Error
The problem that we encountered fitting the distribution for the standard deviation is a systematic error due to the fact that the Theory of Errors does not work out exactly because the estimates of the moments are off. Bessel's correction is needed for the standard deviation and we saw that a second correction was needed for the standard deviation's own distribution function. One can plot the residuals for the original plot, the corrected plot and plot with the standard deviation calculated using its known mean.
The original plot shows evidence of a shift of the standard deviation distribution to left while the correction has greater standard deviations towards the sides and lower standard deviations near the peak as already mentioned. The plot using the known mean of the standard deviation shows a non-uniform distribution with greater fluctuations at the center of the plot. The probability distribution is responsible for reducing the counts on the sides.
Wednesday, December 10, 2014
Error Distribution for the Standard Deviation 2
The two extra degrees of freedom in the standard deviation discussed in the last blog appear to be due to the use of the average of the xk in the datasets for the computation of the observed values. If one uses the known value, μx, instead the observed and calculated curves fit quite well with n as the number of degrees of freedom. The observed values shift to the left when the averages are used to estimate the distribution mean.
Monday, December 8, 2014
The Error Distribution for the Standard Deviation
One can also look more closely at the error distribution for the standard deviation by generating m datasets of n random numbers for a known probability distribution and compute a standard deviation for each dataset.
The computed distribution assuming the standard deviation of the sum is reduced by the square root of n gives a calculated distribution that is slightly offset from the observed distribution.
Using Bessel's correction and assuming two additional degrees of freedom gives a better fit to the observed distribution. The observed distribution appears to have a slightly different shape with a small reduction in the peak value and slightly larger values at the sides sides indicating a broadening of the peak. The sum of the probabilities for histogram intervals in both cases is 1 as expected. One has to use a large number of datasets to notice the differences.
Bessel's Correction
Student published a paper on The Probable Error of a Mean in 1908 in the journal Biometrika. In it he discusses the definition of mean and standard deviation and the distribution of their errors. It is based on Airy's Theory of Errors which contains a discussion of the error for the sum of two random numbers. So what has been presented in these blogs makes use of Airy's approach to the problem. The use of n-1 instead of n in the formula for the standard deviation is known as Bessel's correction.
It is necessary to make two assumptions in order to derive formulas for the mean, μ, and standard deviation, σ. Assuming that a set of random numberss has a mean value and a purely random component imposes a constraint on possible values for them. If we take the average data set, xk, it will be equal to μ plus σ times the average of the z-values. We have two unknowns and the average of the z-values which is subject to some error. The same is true for the mean square of the xk. Both assumptions yield and equation with an unknown sum involving z-values, Eqns (2) and (3) below.
For a normal distribution of errors the expected values of the sum of the z-values and their squares are approximately 0 and n. Making these substitutions we get the usual formulas for the mean and standard deviations. Near the peak the sum of the squares is approximately n-1 instead of n and we get Bessel's correction for the standard deviation.
Using random numbers to check of these two formulas we see that the average for Bessel's correction is closer to chosen value for the standard deviation but its variation is slightly larger.
Friday, December 5, 2014
Least Squares Estimates of the Mean and Standard Deviation and a Sample Calculation
Determining the mean and standard deviation a set of measurements is very difficult to do exactly. One has to make an additional assumption in order to get an estimate of the mean but that adds a little more error to the measurement errors. The assumption used in least squares is that variance V, the sum of the square of the errors, is a minimum. This gives the average as the best estimate of μ. Once we have μ we can then estimate the errors which in turn the variance and standard deviation, σ.
I generated 2000 sets of 20 random numbers to test the formulas used in statistics and the procedure gives the mean, standard deviation and z-values. We also find that the root mean square of the z-values for each set of numbers is exactly 1. Using n-1 in the denominator of the standard deviation formula causes this to deviate slightly. Setting the rms z value equal to 1 is another starting point for the determination of μ and σ.
Each set of numbers will have an mean and standard deviation and there is a little variation among the results for the 2000 data sets but the over-all average is close to the chosen values for μ and σ. The rms variation in the mean is approximately equal to the standard deviation of x divided by the square root of n, the number of values in each data set.
Using the theory of errors from in the last couple of blogs we can calculate the first two moments and standard deviation for the data sets. The σ standard deviation is very close to the μ standard deviation divided by √2 as predicted.
Here is a comparison of the 2000 σ standard deviations with the two estimates of σ.
Subscribe to:
Posts (Atom)

















