It is nearly a decade since Arsenal managed to wrest the Premiership title away from the likes of the two Manchester clubs or their near rivals, Chelsea. So despite a inauspicious start, culminating in a home defeat to Aston Villa that led to knee-jerk calls for Wenger to go, they now find themselves handily placed atop the six game table. A run of five consecutive wins has sent Arsenal four points clear of Chelsea, five clear of Manchester City and eight points ahead of the most recent Manchester side to lift the title.
Goal scoring has been at the forefront of Arsenal's rise back to the top and although it is obviously early days, they are currently on pace to score 80+ goals. That's also the average seasonal goal scoring total achieved by the champions over the the last decade. Sides finishing in fourth average 67 goals a season, rising to 70 for third 76 for second. Simplistically, a side finishing fourth, Arsenal's most common finishing spot since their last title in 2003/04, would ideally be looking to improve their goal scoring by around 23% to achieve a more typical value of Champions.
Defensively, Arsenal has allowed just over a goal a game, on pace for 38+ goals over the season. Typically, fourth placed sides allow around a goal a game over the season, compared to 26 for the Champions. That represents a defensive improvement of nearly 30% between the average defensive performance of the fourth placed side to attain the average record of the Champions.
So increasing goals scored may be the easier half of the deal for a top four side to aspire to rise to the top of the pile.
These crude, early season projections hint at an attack that may have reached a level worthy of a championship tilt, but 13 scoring events hardly inspires confidence in any prediction. However, shot totals do increase the body of evidence. This campaign, Arsenal has made 87 attempts on goal, 15 of which were headers and 72 shots, in scoring their 13 goals. 27 efforts have been blocked and 25 have been off target. So the raw numbers are impressive, if likely unsustainable.
Shot location models can be used to estimate an overall goal expectation for an average team presented with all of Arsenal's attempts so far this season and Arsenal's actual goal tally can then be used to see how effective the Gunners have been in their first six games of 2013/14. A record of 13 goals, when an average side would likely only score 6.5 goals is undoubtedly impressive, as are 35 shots on target against an average expectation of around 26.
However, such favourable comparisons are hardly unexpected for a regular top four side, such as Arsenal over the last decade and the yardstick of a fictitious average side also provides an opaque standard. Therefore, to create a more relevant conclusion from the shooting data, I took all of Arsenal's shots from a recent season, 2010/11, ran a regression using those actual outcomes to see how likely that earlier side were to score from any shooting location on the field and then inputted the shot co-ordinates so far for Arsenal from 2013/14.
By creating a baseline model based around Arsenal (2010/11), a side that scored 72 goals in finishing fourth, we can see how likely it is that a side of that quality might score the 13 goals and hit the target 35 times from the 87 opportunities created by the Arsenal 2013/14 vintage.
Likely Goal and Accuracy Outcomes for Arsenal's 87 Attempts in 2013/14.
The 2010/11 team of van Persie, Arshavin, Nasri, Chamakh, Fabregas and Walcott, had they been presented with the 87 2013/14 chances, would most likely score between 7 and 8 goals. The distribution from simulating thousands of 2013/14 seasons, using shot co-ordinates from the current campaign, but the conversion and accuracy expectation of the 2010/11 side, indicates that such a combination would score the present Arsenal's total of 13 goals around 2.8% of the time. Just over 5% of the time, 13 or more goals would be the result. Also, Arsenal's 2010/11 side would possibly hit the target on the 35 occasions achieved already this season, 2% of the time. Equaling or exceeding this total in 5% of the trials.
Once again, this time when matched against an earlier incarnation of themselves, rather than an average baseline, the present Arsenal team appear to be at least worthy of their current position.
These simulations, which (imperfectly) compare the attacking component of the Arsenal side that finished 4th three seasons ago with 72 goals and 68 points, to the current achievements of Ozil, Giroud, Ramsey and Podolski, can be interpreted variously.
For instance, the pessimist may point out that there appears to be a 5% chance that the 4th placed also-rans from 2010/11 could have produced an as good, if not better record than the one posted in the six matches during August and September by Wenger's current team. So the first six games may just be a lucky, short term streak from a side with similar offensive capabilities to the one that fell short two completed seasons ago.
Or alternatively, the optimistic Arsenal supporter may consider the present record of 13 goals so far to the right of the likely range of outcomes that could have occurred if the opportunities had fallen to the 2010/11 side, that it is reasonable to conclude that Wenger's has a more potent strike force at his disposal than he had when van Persie was the focus.
Shot models can create numerous scenarios, but the subsequent interpretation can be much more subjective.
Monday, 30 September 2013
Sunday, 29 September 2013
Innovative or Just Very Lucky?
The role of random variation, sometimes labelled as luck, is increasingly being recognised in the interpretation of footballing stats. The record of a player's on-field actions over a single season will always be a combination of his true capabilities and a measure of randomness, that sometimes inflates his figures and at other times reduces them. The very good will often be not quite as good as one outstanding season among many merely good ones appears to indicate and a hugely disappointing year from a journeyman may prove to be partly down to bad luck and won't be precisely repeated if he is given the benefit of the doubt.
However, the temptation to always presume that any deviation from the expected level of performance is always down to solely random variation is to assume the existence a uniformity of approach and application of the talent available to a coach, that may not prevail throughout the league.
To celebrate London hosting another regular season NFL game this weekend, I'll draw an example from gridiron. Defended passes are more frequent occurrences than the more valuable full bloodied interception, although both disciplines require a similar skill set and are therefore, reasonably closely correlated. It is analogous to a situation in football, where the much more numerous final third touches in one season appear to better predict goals scored in a subsequent season.
Over the last ten completed years, an NFL side could have seen between 60 to 130 defended passes in a year and between 6 to 30 actual interception made by their defense. You can use the relationship between passes defended and interceptions to produce an expected number of interceptions in a single season. This derived interception total can then be used instead of a side's actual interception total in year N to predict interceptions made in year N+1. In around 65% of the cases the expected figure is a better predictor of future picks.
At the start of the 2012 season, passed defended from 2011 suggested that New England's defense would grab 17 regular season picks. They actually caught 20, so they exceeded the prediction and it is tempting to say they were slightly better than average (the average number of interceptions over the last decade is 16 per season) and lucky. In the previous season, the same thing happened, they beat the prediction from a model that, overall improves the reliability of simply using previous year totals across all 32 NFL teams. And the next....and the next....
How NWE Out Performed A Predictive Model for Defensive Interceptions.
Over the ten year period, NWE outperformed the model (essentially a predictive regression) in nine years. The most likely number of seasons a side would expect to out perform the prediction is, unsurprisingly, five.
For a side to out perform nine times out of ten, if the regression models reality could happen by chance around once every 500 team seasons. We have looked at ten years for 32 teams, so to have found one team who went 9-1 against the model is certainly unusual.
However things get worse for the model, because Chicago beat predictions in eight of the ten seasons (about a 2% chance of happening by chance alone). So now we have two sides that recorded more interceptions than predicted by this model. 26 of the 32 teams over the ten seasons have over (or under) performing years that are within 2 season of the five season average. So the model works well for them.
But for NWE, Chicago, as well as Green Bay, Atlanta, Tennessee and (appropriately) Tampa it under estimates, consistently their intercepting prowess or alternatively, we have to assume these six teams were good and lucky (very, very lucky as a group). As ever, everyone is able to set their own level of confidence in each an every possible scenario or explanation.
An alternative solution is that although improving on a naive use of previous interception totals, this model still omits all possible causes for a side's intercepting abilities. Defensive scheme is one glaring omission (Tampa 2 type zone defenses invite interceptions to be thrown, whereas bounty hunting blitzes, recently favoured by New Orleans prize the opponent above the ball). It also omits the coaching input from defensive gurus, such as Belichick at NWE (who wasn't afraid to use wide receivers on the defensive side of the ball and wasn't above secretly taping their opponents (aka cheating), although this usually referred to the offensive side of the Pats game).
In short, models don't always capture everything and the missing bits may be what sets some teams, coaches and players above the rest or at least fails to identify those that may demonstrate a different tactical approach. A sides relationship to a measurement that can well describe the majority of the league may be branding them lucky, rather than recognising the flaws of a deficient model. Invoking luck to cover the unexpectedly different performance with undue haste should be resisted at least until we see if the "luck" is sustainable and therefore may have a causative agent.
However, the temptation to always presume that any deviation from the expected level of performance is always down to solely random variation is to assume the existence a uniformity of approach and application of the talent available to a coach, that may not prevail throughout the league.
To celebrate London hosting another regular season NFL game this weekend, I'll draw an example from gridiron. Defended passes are more frequent occurrences than the more valuable full bloodied interception, although both disciplines require a similar skill set and are therefore, reasonably closely correlated. It is analogous to a situation in football, where the much more numerous final third touches in one season appear to better predict goals scored in a subsequent season.
Over the last ten completed years, an NFL side could have seen between 60 to 130 defended passes in a year and between 6 to 30 actual interception made by their defense. You can use the relationship between passes defended and interceptions to produce an expected number of interceptions in a single season. This derived interception total can then be used instead of a side's actual interception total in year N to predict interceptions made in year N+1. In around 65% of the cases the expected figure is a better predictor of future picks.
At the start of the 2012 season, passed defended from 2011 suggested that New England's defense would grab 17 regular season picks. They actually caught 20, so they exceeded the prediction and it is tempting to say they were slightly better than average (the average number of interceptions over the last decade is 16 per season) and lucky. In the previous season, the same thing happened, they beat the prediction from a model that, overall improves the reliability of simply using previous year totals across all 32 NFL teams. And the next....and the next....
How NWE Out Performed A Predictive Model for Defensive Interceptions.
| Year | Actual Interceptions. | Predicted Interceptions | Lucky? |
| 2012 | 20 | 17 | Yes |
| 2011 | 23 | 13 | Yes |
| 2010 | 25 | 17 | Yes |
| 2009 | 18 | 15 | Yes |
| 2008 | 14 | 13 | Yes |
| 2007 | 19 | 15 | Yes |
| 2006 | 22 | 15 | Yes |
| 2005 | 10 | 14 | No |
| 2004 | 20 | 15 | Yes |
| 2003 | 29 | 26 | Yes |
Over the ten year period, NWE outperformed the model (essentially a predictive regression) in nine years. The most likely number of seasons a side would expect to out perform the prediction is, unsurprisingly, five.
For a side to out perform nine times out of ten, if the regression models reality could happen by chance around once every 500 team seasons. We have looked at ten years for 32 teams, so to have found one team who went 9-1 against the model is certainly unusual.
However things get worse for the model, because Chicago beat predictions in eight of the ten seasons (about a 2% chance of happening by chance alone). So now we have two sides that recorded more interceptions than predicted by this model. 26 of the 32 teams over the ten seasons have over (or under) performing years that are within 2 season of the five season average. So the model works well for them.
But for NWE, Chicago, as well as Green Bay, Atlanta, Tennessee and (appropriately) Tampa it under estimates, consistently their intercepting prowess or alternatively, we have to assume these six teams were good and lucky (very, very lucky as a group). As ever, everyone is able to set their own level of confidence in each an every possible scenario or explanation.
An alternative solution is that although improving on a naive use of previous interception totals, this model still omits all possible causes for a side's intercepting abilities. Defensive scheme is one glaring omission (Tampa 2 type zone defenses invite interceptions to be thrown, whereas bounty hunting blitzes, recently favoured by New Orleans prize the opponent above the ball). It also omits the coaching input from defensive gurus, such as Belichick at NWE (who wasn't afraid to use wide receivers on the defensive side of the ball and wasn't above secretly taping their opponents (aka cheating), although this usually referred to the offensive side of the Pats game).
In short, models don't always capture everything and the missing bits may be what sets some teams, coaches and players above the rest or at least fails to identify those that may demonstrate a different tactical approach. A sides relationship to a measurement that can well describe the majority of the league may be branding them lucky, rather than recognising the flaws of a deficient model. Invoking luck to cover the unexpectedly different performance with undue haste should be resisted at least until we see if the "luck" is sustainable and therefore may have a causative agent.
Saturday, 28 September 2013
Liverpool's Shooting So Far.
Following on from Friday's look at the likely outcomes of the goal attempts made by Robin van Persie had they been made by a league average player, here is a similar plot for all 63 goal attempts made by Liverpool.
Van Persie, has so far out performed the model by scoring slightly more goals than you would expect from an average player. However, such is the limited sample size for the Dutchman, it isn't possible to use data from the four matches he has played so far to demonstrate without doubt that he is an above average striker.
An inferior striker (if we assume from extensive evidence over the longer period of van Persie's career that he is indeed above average) could quite easily have scored the three goal total already achieved by van Persie in 2013/14. So in the absence of a larger cv, an elevated strike rate compared to a generic shooting model shouldn't be regarded as evidence of better than average talent. Data evaluations of playing talent will always come with levels of uncertainty, rather than cast-iron conclusions.
Shot data accumulates more rapidly for teams than for individual players and following Liverpool's 1-0 home defeat to Southampton, the Merseysiders had executed 63 attempts in scoring 5 times. The overall, model predicted, goal expectancy from all 63 shots is total just under five goals. Therefore, it is no surprise to see that the most likely individual goals tally recorded by an average side if they had been given those 63 opportunities is also five.
There is a 20% chance of an average team scoring five goals and a shade of odds on that a side would score at least five goals given Liverpool's opportunities. So we can tentatively say that over the first 63 chances created by Liverpool, there has been little surprise in their conversion rate. Everything that has happened once the chances presented themselves could reasonably have been duplicated by an averagely competent converting side.
Liverpool's shooting accuracy, however, is more extreme. The model predicts an average of just under 20 of the 63 goal attempts to have been on target and Liverpool so far have hit well in excess of that prediction with 27, making them the joint second most accurate shooting side in the EPL in 2013/14.
An over-performance of 40% in hitting the target 27 times compared to an expected 19.6 does appear outstanding and can easily give the impression that we are looking at a real and possibly sustainable effect. The temptation is to look to causes and explanations. But first we perhaps should see how unusual such a rate is for our baseline, average side to have achieved in 63 trials.
In simulations assuming an average level of all round competence, 27 shots are seen to hit the target around 1.5% of the time and at least 27 shots were recorded about 5% of the time. Certainly unusual, but not within the bounds that could be considered significant. On the evidence of 63 shots, Liverpool may be more accurate than an average side, but by quoting the 40% improvement over average, (especially if sample size is omitted), an inflated expectation of their true ability is almost certainly being created.
So, in statistical terms, there is a justifiable reason to suppose that, despite an impressive accuracy rate, Liverpool may be little better than average in reality.
Previous seasons and repetition of this inflated accuracy by broadly similar Liverpool teams of the recent past, is one route to adding weight to any opinion regarding Liverpool's shooting accuracy. But intimate knowledge of the shooting model that has been used is another. The model I've used includes many of the readily collectible variables, such as shot location, shot type, but it doesn't include such things as shot power, which are both subjective and virtually impossible to collect in any great numbers.
In the limited data I have, the power of the shot impacts negatively upon the accuracy, and yet increased power doesn't appear to statistically significantly improve conversion rate compared to normally struck efforts. (Placement is the obvious missing link). Therefore, (with the caveat that this is very limited data) you can construct a scenario, where reducing the power of a shot, doesn't reduce the conversion rate, but increases the accuracy and a shot extra saved is an extra possibility of a further shot attempted from an additional rebound. The exact profile seen at Liverpool this year.
Models can tell us much about how teams perform, as long as we aren't too dogmatic about conclusions. Ultimately, they just provide information on how likely it is that a real, data based assessment is going to coincide with an unobtainable, all encompassing knowledge of a side's true ability. Random variation can turn world beaters, short term, into average, run of the mill sides and vice versa. But equally (as in the proposed effect of shot strength), seemingly unusual results can be an early indication of a model depleted of minor, yet important variables.
Van Persie, has so far out performed the model by scoring slightly more goals than you would expect from an average player. However, such is the limited sample size for the Dutchman, it isn't possible to use data from the four matches he has played so far to demonstrate without doubt that he is an above average striker.
An inferior striker (if we assume from extensive evidence over the longer period of van Persie's career that he is indeed above average) could quite easily have scored the three goal total already achieved by van Persie in 2013/14. So in the absence of a larger cv, an elevated strike rate compared to a generic shooting model shouldn't be regarded as evidence of better than average talent. Data evaluations of playing talent will always come with levels of uncertainty, rather than cast-iron conclusions.
Shot data accumulates more rapidly for teams than for individual players and following Liverpool's 1-0 home defeat to Southampton, the Merseysiders had executed 63 attempts in scoring 5 times. The overall, model predicted, goal expectancy from all 63 shots is total just under five goals. Therefore, it is no surprise to see that the most likely individual goals tally recorded by an average side if they had been given those 63 opportunities is also five.
There is a 20% chance of an average team scoring five goals and a shade of odds on that a side would score at least five goals given Liverpool's opportunities. So we can tentatively say that over the first 63 chances created by Liverpool, there has been little surprise in their conversion rate. Everything that has happened once the chances presented themselves could reasonably have been duplicated by an averagely competent converting side.
Liverpool's shooting accuracy, however, is more extreme. The model predicts an average of just under 20 of the 63 goal attempts to have been on target and Liverpool so far have hit well in excess of that prediction with 27, making them the joint second most accurate shooting side in the EPL in 2013/14.
An over-performance of 40% in hitting the target 27 times compared to an expected 19.6 does appear outstanding and can easily give the impression that we are looking at a real and possibly sustainable effect. The temptation is to look to causes and explanations. But first we perhaps should see how unusual such a rate is for our baseline, average side to have achieved in 63 trials.
In simulations assuming an average level of all round competence, 27 shots are seen to hit the target around 1.5% of the time and at least 27 shots were recorded about 5% of the time. Certainly unusual, but not within the bounds that could be considered significant. On the evidence of 63 shots, Liverpool may be more accurate than an average side, but by quoting the 40% improvement over average, (especially if sample size is omitted), an inflated expectation of their true ability is almost certainly being created.
So, in statistical terms, there is a justifiable reason to suppose that, despite an impressive accuracy rate, Liverpool may be little better than average in reality.
Previous seasons and repetition of this inflated accuracy by broadly similar Liverpool teams of the recent past, is one route to adding weight to any opinion regarding Liverpool's shooting accuracy. But intimate knowledge of the shooting model that has been used is another. The model I've used includes many of the readily collectible variables, such as shot location, shot type, but it doesn't include such things as shot power, which are both subjective and virtually impossible to collect in any great numbers.
In the limited data I have, the power of the shot impacts negatively upon the accuracy, and yet increased power doesn't appear to statistically significantly improve conversion rate compared to normally struck efforts. (Placement is the obvious missing link). Therefore, (with the caveat that this is very limited data) you can construct a scenario, where reducing the power of a shot, doesn't reduce the conversion rate, but increases the accuracy and a shot extra saved is an extra possibility of a further shot attempted from an additional rebound. The exact profile seen at Liverpool this year.
Models can tell us much about how teams perform, as long as we aren't too dogmatic about conclusions. Ultimately, they just provide information on how likely it is that a real, data based assessment is going to coincide with an unobtainable, all encompassing knowledge of a side's true ability. Random variation can turn world beaters, short term, into average, run of the mill sides and vice versa. But equally (as in the proposed effect of shot strength), seemingly unusual results can be an early indication of a model depleted of minor, yet important variables.
Subscribe to:
Posts (Atom)

