Modelling Test Cricket Batting Performances
/INTRODUCTION
Over 10 years ago I wrote a piece based on a paper by Brendon Brewer from the University of New South Wales that provided a way to model batter’s final scores in Test innings.
In that paper he assumes that, for most batters, each Test innings plays out over two phases: an initial phase, before the batsman is said to be "set" and during which he is more susceptible to dismissal, and a second and final phase during which his probability of dismissal slowly declines as each run is scored and asymptotes towards some fixed value.
THE MODEL
The model in the paper characterises each batter using four parameters:
μ1: which we can think of as a measure of the batter's ability in the 1st phase of their innings
μ2: which we can think of as a measure of the batter's ability in the 2nd phase of their innings
τ: which is the transition mid-point such that, when the batter’s score is τ they are halfway between transitioning from μ1 to μ2. A smaller τ means the batter gets their eye in earlier.
L: which controls the abruptness of the transition between the two phases. A small L gives a sharp change near τ and a large L spreads the change across more runs. About 10% of the transition is complete at τ - 2.2L and 90% at τ + 2.2L, so the central transition spans about 4.4L runs.
THE APP
I’ve created an app that implements this model and allows me to do a variety of things with it.
Modelling Actual Players
One is to download the actual batting careers of different players and then compare and contrast them across these four parameters.
To demonstrate this feature, let’s compare the batting careers of Steve Waugh, Adam Gilchrist, Virat Kohli, and Chris Martin.
Firstly, on the right is a table showing the estimates of the four key parameters and a 95% confidence interval for each parameter. (Note that we are taking a Bayesian approach in the modelling, which results in a distribution being created for each parameter.)
With career data being relatively scarce, even for players with relatively long tenures, most of the parameters, μ2 aside, have quite wide confidence intervals.
Nonetheless we can see that:
Waugh, Gilchrist, and Kohli have broadly similar expected scores when set
Waugh and Gilchrist have lower expected scores before they are set than does Kohli, but they also transition into being set relatively more rapidly
Martin has a very low expected score before he’s set but transitions relatively quickly (in terms of runs) into his set phase but even there only has an expected score of around 13.
We can use the means of these parameters to create score distributions for all four players.
We can see Waugh’s and Gilchrist’s relative susceptibility early in their innings and how their distributions differ markedly from Kohli’s.
Martin’s distribution is one of a kind. His Test average is 2.4 and high score is 12*, so it is debatable if he ever achieved “set” status.
The table at right provides the full career batting statistics for the four.
It shows, for example, that 10% of Waugh’s completed innings were ducks as were 12% of Gilchrist’s but only 8% of Kohli’s. Martin managed 69%.
As well, 30% of Waugh’s completed innings were scores under 10 as were 35% of Gilchrist’s but only 27% of Kohli’s. Martin, in particular, excelled here, managing a perfect 100 record.
The last column in this table uses the distribution of the four parameters generated for each player to create 10,000 careers of the same length, that is with the same number of innings, deploying the following methodology:
A parameter vector, μ1, μ2, τ, L is selected from the player’s posterior distribution.
For every innings, a latent dismissal score X is drawn from the fitted score distribution. This is the score the player would eventually make if the innings continued until dismissal.
A potential censoring score C is sampled from the player’s actual historical not-out scores.
The model independently determines whether the innings is exposed to censoring.
The innings becomes not out only if:
censoring is activated; and
C ≤ X, meaning the match ends before the simulated dismissal.
If both conditions hold, recorded score = C and the innings is recorded as not out, otherwise recorded score = X and the innings is recorded as completed.
The censoring activation probability is set equal to the proportion of not out scores in the player’s actual history.
What’s interesting about that column is the range of different career averages that emerge from the simulations. In Waugh’s case the range is about 20 runs, which would see him move from being remembered as a moderately successful batter with a 41 average to even more of a titan with an average of 62.
Gilchrist’s alternative careers span an even greater range, from about 35 to 62, while Kohli’s span a range sized similarly to Waugh’s but pitched about three to five runs lower and running from 37 to 57.
Overlaying Selection Rules
We’ve seen that there can be significant variability in the output of a career, even for talented players when we generate a set of careers of equal length, but what happens if we overlay some rules that serve to shorten careers with periods of underperformance?
For this analysis we’ll use Steve Smith, whose modelled and actual details appear below.
Only 6% of his completed innings have been ducks and only 26% have produced less than 10 runs, a rate that is slightly better than Kohli’s.
What happens if we generate 1,000 careers for him but invoke two different possible career-shortening rules
If he averages less than 20 in his first 20 innings, we shorten his career by 20 innings
If he averages less than 20 in any 20 consecutive innings after that, his career is ended
The summary results appear below.
In 33 careers (or about 3%) he experiences a career-shortening event, in the most extreme case ending his career after just 22 innings. In his worst career he finishes averaging just 26.55. His average career is shortened by about 3 innings. On the whole, no significant damage except in a rare few careers.
To finish, let’s imagine we were even less patient with Smith and ended his career if he went 15 innings averaging below 25.
In this case, almost 60% of his careers involve some shortening, the shortest just 15 innings and the worst in terms of average just 14.57. His average career is only about two-thirds of the longest. That’s a lot of lost joy.
Now I wouldn’t claim that selectors ever behave in such a strict rule-based fashion and it might well be that talent, once spotted, is given ample opportunity to prove itself, but I do think that seeing the impact of imposing such rules does give some pause for thought.
MAKING THE APP LIVE
I am investigating the possibility of making my app available live here on the site, but there are a number of practical and technical issues I will need to resolve before I can do that.
In the meantime if there are any players you’d like to see analysed or different career-shortening rules you’d like tested, send me an email via the address in the navigation sidebar.
