Pages

Monday, 12 August 2013

Towards A Better With or Without You.

In the previous post, I looked at the the growing trend to attempt to evaluate the impact of a single player by referring to game results when he takes part in a match or is wholly absent either through injury,non selection or suspension. Superficially the methodology appears sound, if rather crude. However, on closer inspection the pitfalls are both numerous and largely insurmountable.

By taking such an approach we trying to demonstrate how good a single player is by comparing the difference in team performance in games where he is absent, but replaced by another (not necessarily the same) player, who may play against widely diverse opponents, surrounded by a similar, but often varied collection of colleagues. And the difference is measured in that rarest of footballing commodities, namely goals. Often the exercise runs over a single season, a period that is often even insufficiently large to accurately demonstrate the existence of such universals as the home field advantage enjoyed by a side.

The twin terrors of opponent strength and small sample size can be illustrated with a contrived, extreme example. Often universal problems that occur to varying degrees in reality can be highlighted by use of the unlikely, but possible scenario.

A side wins the league in a canter, winning every match, bar one where an early red card led to a narrow defeat. For the final match, the ever present star is rested to the stands and watches as his team is the early beneficiary of a red card decision and run out easy 6-0 winners.

A raw with/without percentage comparison will give his side a better win or success rate or goal difference in that final game compared to the previous 37 matches when he was a confirmed starter. A yardstick derived from one match is (hopefully) obviously inadequate, but similar problems arise when drawing conclusions even from 10 or 15 game samples.

The fundamental idea is fine, but the current application of the method is awash with noise, leading to little worthwhile conclusions.

Spurious, title related, pop image.
To demonstrate how this approach may be improved it may pay to look to the NFL, where parity is slightly more keenly experienced than in the EPL and player contribution to useful, game defining events is more readily apparent because of the play by play nature of the contest.

An NFL offense takes part in around 60 individual plays per game over the 16 match regular season. None of the 11 players on the field during each play is a passive spectator, as can sometimes be the case in football. Every NFL player has a designated contribution to make on a single play, even if it is merely diversionary and takes place far from the area target by the ball. Furthermore, the success of each play can be quantified, most readily, if slightly unsatisfactory, in terms of raw yardage gained or lost or alternatively, in terms of first downs achieved.

So by looking at the NFL, the concerns around quality of opposition and small sample size is partly addressed. Further, by looking at unique lineups, rather than plays that took place with or without certain players, we can avoid the problem whereby the remaining 10 players may also change identity. And finally if we use the least used lineups to compare against the most used lineup, we can begin to see the likely ranges of the difference in performance at a team level. This may give a ball park figure, when scaled down to individual player levels, of the kind of differences we are (naively) attempting to quantify in a sport with similar levels of professionalism, such as soccer.

More frequent use of the same 11 players on offensive plays, appears to be quality driven. The less man games a side lost to injury the more they use the same players on a single play and on average the more successful they were measured against their recent achievements. So it is not too big a leap to suggest that unique formations that occurred most are likely to represent the near cream of a side's playing staff that year and much less used combinations are more likely to contain lesser players. Visually this also appears to be the case.

Average Performances of Most and Least Used Lineups in the NFL.

Offensive Lineup. % of Total Plays Involving Lineup. % of  Total Yards Gained. % of First Downs Gained.
Most Used. 6.3 7.0 6.9
Least Used. 10.8 9.7 10.0

(Most Used lineups took part in 2100 plays, the least were used in 3800 plays over the 2012 season).

Above I've averaged the outcomes of plays made by the least and most used lineups for every NFL team last season. Success has been quantified both in yardage and first down terms and the percentage of plays each composite unit was involved in is included to give context to the opportunities given to each group.

Fuller context is missing, but the broad picture hopefully matches reality. The most used, presumably more star studded offensive lineups produce slightly more of a side's overall yardage and first downs than their share of the plays would suggest. The least used lineups took 3800 snaps or 10.8% of the total experienced by NFL offenses in 2012 and produced slightly below that % of first downs or raw yardage. 

Arguably, a single season of matches for every NFL side, using nearly 6,000 individually quantified on field plays has managed to show a (small) difference in performance levels between what may generally be described as the generic "best" 11 man lineup and the less favoured ones.

In hindsight, attempting the same for individual soccer players, from individual teams, using just 38 total trials, decided by rare events, should possibly be considered a tad optimistic. The quality of players in the EPL is undoubtedly high, but the difference in quality between interchangeable colleagues in the same side is likely to be very small. Expecting this difference to show itself in match results over a relatively limited timescale, especially with the attendant, unaddressed baggage, is unrealistic.

At the very least we should be looking at using more frequently seen individual or team events, such as successful passes or chances created rather than placing so much faith in match outcomes and accept that we are trying to measure differences that, in individuals, at least is very likely to be swamped by noise. 

Friday, 9 August 2013

A Decade of Steven Gerrard As A Liverpool Mainstay/Liability.

Much of the current internet football discussion revolves around the transfer window and the destination of big name players to pastures new. Understandably, speculation about the likely impact of a newly acquired player on the fortunes of his new  employers is rife. However, it is also entertaining to ponder about the size of the hole the departing player will leave at his former club and how reliant they were on his talents.

The most common approach uses the seemingly simple, yet powerful tool of examining the record of a side when the player was in the lineup to those matches where he was absent. Various twists on the original format are usually used and team match performance is judged either through a combination of goals scored and allowed, points accrued or a weighting of wins, losses and defeats.

A player who helps his team to a more impressive win, loss, draw record is the obvious candidate to be the causative agent for that improved record. The first player identified as a force for providing better results when he was on the pitch compared to those recorded when he was absent was Patrick Vieira in his time at Arsenal. Such was the combative nature of his play, especially when faced with natural championship rivals, such as Manchester United and Roy Keane in particular, that the connection became widely accepted and has often been used for other, high profile players such as Xavi at Barca and latterly Suarez at Liverpool.

A method that uses the relatively rare currency of goals or wins and draws to evaluate a side's performance when one certain individual is either absent or present is fraught with problems. Most obviously, quality of opposition, but also the quality of the remaining 10 or 11 players is largely ignored. Standout players and their colleagues may be rested en masse against inferior opponents or arbitrarily suspended or injured against the best.

However, by far the biggest problem in relying on this kind of with/without analysis revolves around sample size. Small game samples, especially when success or failure is decided by rare events, such as goals often lead to headline grabbing figures, usually expressed in percentages that appear to compare like with like when they do no such thing.

To use the example of Steven Gerrard, a formidable one club man with Liverpool. If we look at the season on season record of his team Liverpool when Gerrard played a part in the match and when he didn't, we find that in ten of those 15 seasons the "Gerrardless" Reds had a better season long success rate than when he played. In just five seasons was Gerrard, (apparently) a force for improved results.

So to decide if Gerrard is an undroppable icon who more often hurts Liverpool's chances, it may pay to look at the actual numbers rather than the percenatge figures.

In 2007-08, Liverpool's record when Gerrard was on the pitch during a game was 29 wins, 16 draws and 7 losses for a success rate of 0.712 compared to 4,2,1 (0.714) when he wasn't. Liverpool were marginally better without Gerrard, but we are comparing the outcome of 53 games to that from just 7.

To illustrate the potential volatility, one more or less loss added to a Liverpool with Gerrard sees their success rate fall to 70% or rise to 72.5%. The same alteration applied to Liverpool deprived of Gerrard sees the high rise to 83% and the lows fall to 62%. There is little change in the former percentage figures, but much larger swings are possible in the latter, when Stevie G is absent.

High profile players are invariably the target for such type of analysis, so they will invariably play often. So the stick we are using to beat them with or the carrot as a reward will almost always be prone to wild fluctuations as illustrated by the effect of a single altered result in the smaller sample size above. In short, the so called "player absent" yardstick should come complete with massive levels of uncertainty because they are based on very small sample sizes, yet, much like many percentage based player figures used across the web, they never do.

Even if we extend the sample size by taking in multiple years, we can still slice and dice the data to suit any narrative by manipulating the figures thrown out by the smaller sized sample.

Steven Gerrard Is An Essential Player For Liverpool.

Time Period. Success Rate with Gerrard. Success Rate without Gerrard. Were Liverpool "Better" With Steven Gerrard ?
Debut to 2012/13 0.643 0.626 Yes
Debut to 2011/12 0.646 0.627 Yes
Debut to 2010/11 0.650 0.628 Yes
Debut to 2009/10 0.657 0.628 Yes
Debut to 2008/09 0.665 0.634 Yes
Debut to 2007/08 0.653 0.628 Yes
Debut to 2006/07 0.644 0.625 Yes
Debut to 2005/06 0.645 0.624 Yes
Debut to 2004/05 0.625 0.621 Yes
Debut to 2003/04 0.634 0.619 Yes


Steven Gerrard Is A Liability For Liverpool.

Six Season Period. Success Rate with Gerrard. Success Rate without Gerrard. Were Liverpool "Better" With Steven Gerrard ?
2013-2007 0.476 0.626 No
2012-2006 0.500 0.626 No
2011-2005 0.535 0.627 No
2010-2004 0.561 0.628 No
2009-2003 0.582 0.634 No
2008-2002 0.591 0.628 No
2007-2001 0.601 0.625 No
2006-2000 0.627 0.624 Yes
2005-1999 0.625 0.621 Yes

For anyone considering using the currently flawed methodology to bolster a subjective opinion piece, feel free to use either of the two tables above. I've calculated Liverpool's success rate with or without Gerrard over differing, multiple timescales since his debut back in the 1990's. The data has been diced and expanded to cover multiple seasons (to enhance it's credibility). One table can be used to illustrate his importance to Liverpool and the second to show how much better they are without him. 

Seemingly powerful data backing diametrically opposed viewpoints for one of the most consistent, high profile and talented players of his generation. What chance for accurate or even meaningful assessment of lesser lights only partway into their careers when they are put under a similar spotlight using similar methods?

Evaluating player contribution is an obvious interest for many parties, but some current methods have the potential to greatly, and surreptitiously deceive.

Just to be clear, this is about abusing stats, not about Steven Gerrard, who is of course a magnificent player.

Tuesday, 6 August 2013

Predicting The Rare From The Commonplace In Football.

Each sport has an event or series of events that have a disproportionate effect upon the outcome of the match. The ability to create chances in soccer is an obviously vital factor in determining match result. They are the precursors to goals and goals are the ultimate arbiter of the successful, defeated or stalemated in soccer.

The ability to score goals, as Blackpool most recently demonstrated, fulfills only half of requirement to be successful. Defence is also important and the Seasiders 55 goals equaled the tally set by Tottenham, but saw them finish 15 places below Spurs because of their 78 scores conceded. The general case is still fairly strong, the more goals a team scores, then the higher up the league you tend to finish, but for the complete picture, we also have to look at defensive performance as well..

Past performance is often an indicator of future achievement. Managers and players invariably come and go, bringing changing skills and tactical development, but a sizable rump of the previous team often remains and previous performance levels still explain at least part of what we may see in the future.

In the previous post I looked at the balancing act between the limitations of using direct comparisons between significant events and instead gathering more copious amounts of data by moving a stage or two back in the process or incorporating more numerous actions that require similar skills to the perform the key acts that we wish to project. Relatively rare goal scoring may be better predicted by examining the more frequent assists from where they originate and assists themselves may also have a more accurate predictive ancestor.

The ball controlling nature of the NFL makes for a much better testing ground for the use of more numerous, secondary events as better predictors of rare, but important, game changing events. Numbers of possessions is almost always equally shared between sides in the NFL. So the effective use a side makes of that possession is a decisive factor in determining the result.

Turnovers, whereby one offense hands the ball over to the other defense (and hence onto the opposing offense) without scoring are extremely difficult to overcome in a single match. Drives are time consuming and teams can ill afford to pass up scoring opportunities through their own carelessness and hand an "extra" one to an opponent.

Interceptions are the most straightforward of turnovers. The quarterback throws to an intended receiver, but a safety or cornerback, most usually, intervenes and catches the ball instead. Possession lost, often along with the game if the process is repeated and the side ends the match with a negative turnover differential.

Interceptions in the NFL are rarer than goals are in soccer. A typical side picks about 16 errant throws a season or an average of one a game, split between game chasing desperation throws and game changing miscues. But their effect on the game result is often so profound that it is extremely useful to have as accurate an estimation of future performance as is possible.

The tried and trusted route of relying on previous performance to predict future outcome doesn't provide a strong relationship. We've already noted that teams change from year to year, but coupled with the small sample size, interception numbers provide scant clues to future intercepting potential. So can improve the strength of the predictive relationship if we instead look to a precursor to interceptions that require similar skills and, crucially occur in much greater numbers?

A pass thrown presents the opportunity for the receiver to make the play, the ball to fall incomplete, the ball to be successfully intercepted or the defender, by his presence to get close enough to the intended receiver to knock the ball away. The latter, a so called defensed pass is a close relative to the full interception and a side records around 90 such plays a season compared to just 16 for interceptions.

Do Defensed Passes Better Predict Future Interceptions?

NFL Team. Actual Def. Interceptions in 2011. Predicted Def. Interceptions in 2011. Actual Def. Interceptions in 2012.
Arizona. 10 18 22
Atlanta. 19 17 20
Baltimore. 15 22 13
Buffalo. 20 15 12
Carolina. 14 15 11
Chicago. 20 14 24
Cincinnati. 10 16 14
Cleveland. 9 15 17
Dallas. 15 13 7
Denver. 9 13 16
Detroit. 21 16 11
Green Bay. 31 20 18
Houston. 17 20 15
Indinapolis. 8 11 12
Jacksonville. 17 14 12
Kansas City. 20 18 7
Miami. 16 14 10
Minnesota. 8 11 10
New England. 23 14 20
New Orleans. 9 19 15
NY Giants. 20 18 21
NY Jets. 19 15 11
Oakland. 18 18 11
Philadelphia. 15 14 8
Pittsburgh. 11 16 10
San Diego. 17 14 14
San Francisco. 23 21 14
Seattle. 22 17 18
St Louis. 12 14 17
Tampa Bay. 14 13 18
Tennessee. 11 16 19
Washington. 13 16 21

What we lose by no longer comparing like with like, may be compensated for by a substantially larger pool of data that describes an associated skill. The relationship between defensed passes and interceptions is relatively strong and in the expected direction, so they are likely to be the product of similar player talent. Therefore using data from 2006 to 2010, I calculated the number of interceptions each side would have expected to make in 2011 based on their number of defensed passes by their defense.

For example, Seattle's defense claimed 22 interceptions in 2011's regular season, but based on the number of passes they managed to physically touch and disrupt during that season, the league average regression from 2006-2010 suggests that they probably over achieved by around 5 interceptions. A league average team managing Seattle's 105 defensed passes would have more likely just grabbed 17 picks. The full list for each team in 2011 is in the table above. 

The discrepancy between the actual figure and the pass defense predicted figures for 2011 can be explained in two ways. Either sides which differed greatly from the prediction got lucky (in a good or bad way) or they were in reality better or worse than the league average. Or much more likely a combination of the two.

As with all recorded stats the skill component tends to persist across time, but the random luck does not. Sides that persistently under or over-perform may be exhibiting their true deviation from the league average skill levels, but the temptation is to see any variation from the average to be fully resulting from different levels of real talent and disregard random variation as even a partial cause. 

For the purpose of this rudementary trial, we can use a single season to see if the underlying talent needed to intercept a pass is better represented in more numerously defensed passes than in relatively rare real interceptions. If there is more signal than noise in an average of 90 defensed passes than there is in 16 interceptions, this should show up in how well each set of 2011 figures predict figures for 2012 for each team. 

Formally, neither actual nor predicted 2011 interceptions are strongly related to defensive interception performance in 2012. Team churn and rarity of the event almost guarantees that. However, the amount of interceptions predicted from passes defensed in 2011 is closer to the actual 2012 figure in 22 of the 32 cases. One honourable tie leaves previous actual performance the better indicator of year N+1 interception performance on just nine occasions. 

If you wanted to predict interceptions in 2012 by looking at interceptions in 2011, you were not looking in the best place.

A couple of new signings can partially change a side's likely stats.
The important events in a sport are often just special cases of a more general talent. A player who can pass accurately in the final third in the Eredivisie, will as likely also be able to provide a key pass that creates a chance for a colleague. Therefore, final third passing ability because of it's relative frequency may well be a much better indicator of a player's ability to create chances than the limited occasions on which he has actually done so over previous seasons.

Correlations will never approach the levels of certainty to which visual evidence often tricks us into believing is possible, but exposing noise among the signal and vice versa is always a satisfying advancement of the raw, often deceptive figures.