Predicting Ultimate Men's AFL Career Length

THE QUESTION

The analysis in this blog was designed to answer a single, seemingly simple, question:

“Given only what was known immediately after a player's 1st, 10th, 25th, 50th or 100th game, how far was his AFL career ultimately likely to run?”

THESE DAYS, THERE’S CENSORSHIP EVERYWHERE YOU GO

To measure survival we need to reliably measure death or, in the case of an AFL player’s career, retirement.

We know with certainty how many games a retired player finished with. For a current player, however, their current total is only the lower bound on their final total. Treating that total as final would systematically label the most durable active players as if they had already stopped playing.

The statistical solution to this problem is to recognise that the career lengths for current players are what’s called right censored. A player who is active at the analysis endpoint contributes everything observed up to that endpoint, but the model is “told” that the career ending likely lies somewhere beyond that point.

For players who are inactive at the analysis endpoint, we can use all the data we have about their entire observed career.

The model therefore learns from both completed and incomplete careers without pretending that incomplete careers are complete.

WHEN TO START AND END

The start-point of the analysis was chosen to be 1 January 2012, which is the earliest season for which the AFL website has some of the requisite player data.

The end-point of the analysis was chosen to be 31 December 2025 (as the analysis was commenced part way through the 2026 season). A player was recorded as having a completed career if a FootyWire retired-player record was linked to him by name and date of birth, or if he had missed two complete seasons and did not appear on the current-list extraction for 2025. A player who did appear on a current list was treated as having a right-censored career. If neither rule resolved the status, it was assumed that the player was still active.

PLAYER STATUS

It total, the careers of 1,825 players were available for analysis whose career statuses broke down as follows:

  • Have a FootyWire retired-player record: 1,117

  • Have been inactive for two full seasons: 71

  • On a FootyWire current player list: 631

  • Unresolved: 6

(Note that, throughout this analysis, data anomalies will be thrown up, some of which could be resolved via manual intervention. I’ve refrained from doing this as I’m looking to create a repeatable process that won’t require further manual interventions in the future. None of the identified anomalies have a substantial impact on the outcome.)

We should acknowledge that this heuristic doesn’t identify a guaranteed end of a career as some players might come out of retirement, recover from a long-term injury, or move to another competition.

WHEN TO MAKE CAREER-LENGTH PREDICTIONS

To begin, let me define two terms that we’ll use throughout this blog:

  • Landmark: a given number of completed games. We use landmarks of 1, 10, 25, 50, and 100 games in this blog.

  • Thresholds: a given minimum number of games potentially obtained in the future. We use thresholds of 50, 100, 150, 200, and 250 games in this blog.

We estimate the remaining number of games in a player’s career at five landmarks: immediately after games 1, 10, 25, 50 and 100. Each threshold creates a new prediction problem conditional on having reached that a given landmark. The game-100 model, for example, is a model for established 100-game players, not a model that can be applied to debutants, used to estimate the probability of attaining a future threshold.

The prediction target is final official AFL premiership-season career games with primary outputs the probabilities of finishing with at least 50, 100, 150, 200 and 250 games. A secondary output is a restricted estimate of expected final games, capped at 286 because that was the longest completed career with fully observed history in the dataset.

The table below records how many player records are available for each landmark model.

Note that only player careers for which data is available from Game 1 onwards were allowed into the landmark models. Players whose careers began before 2012 would otherwise appear with missing early-career ratings, positions and selection patterns. Excluding these players with left-censored histories sacrifices sample size but prevents an older player's mid-career record from masquerading as a debut or early-career record.

We lose 550 player careers as a result of this decision.

WHAT DO THE MODELS KNOW ABOUT A PLAYER?

Models were built of increasing complexity using progressively more player data.

When a time-varying player data element was included in a model, it was recalculated at each landmark game using games 1 through to that landmark game and nothing later. This ensured that no information about the player’s future career, if any, was allowed to leak into the data available for a model meant to be fitted at a given landmark.

The player predictive data used in one or more of the models came from the AFL website via the R package fitzRoy

  • Age at the landmark, age at debut, debut cohort, and the rules era in which the landmark occurred.

  • Average and recent rating, rating variability and recent change, and a short-term trajectory calculated only from games leading up to a landmark.

  • Average and recent time on ground (note that values outside 0 to 100 percent were treated as missing).

  • Most common on-field position and the number of distinct positions observed up to the landmark.

  • Games per observed season, games in the current landmark season, gaps between games, interruptions and recent selection spacing.

  • Team played for in the landmark game, and the number of teams represented so far.

Current list and retired status came from FootyWire current-player and retired-player pages.

The algorithm employed various strategies for matching player data across sources.

MODEL 1: THE SIMPLEST

The first model we build - which we’ll call the Simple model - is the simplest, and provides answers to the following: among players who had reached a given landmark, what estimated proportion continued to each threshold? The Kaplan-Meier model used for this purpose handles censoring but ignores all differences between players. It establishes the predictive value of achieving a given landmark in estimating how likely it is that a future threshold be achieved.

This model tells us, for example, that a debutant has only about a 48% chance of playing 100 or more games in their entire career but, if he reaches 50 games, is now a 79% chance of playing 100 or more games in their career. More generally, the more games a player plays the higher his chances of reaching future thresholds.

This benchmark model is valuable because it is easy to explain and hard to overfit. But, for example, it assumes that a 20-year-old high-rating player and a 25-year-old low-rating player at the same landmark have identical chances of reaching a given threshold.

Our next model investigates how much more age alone can improve predictions.

MODEL 2: DOES AGE WEARY THEM?

The Simple model described above gives every player at the same landmark the same probability estimates for acheiving future thresholds. Here we ask how much two facts known at the landmark - the player's age and his age at debut - explain the difference between player’s remaining career length by building what we’ll call the Age Enhanced model.

The Age Enhanced model is a Cox proportional-hazards model with flexible curves for age at the landmark being estimated and the debut age. It estimates the chance that a career ends at each subsequent career-game total, conditional on the player having reached a given landmark. Flexible curves allow the relationship with age to bend rather than forcing each additional year to have the same effect. We evaluate this model at each of the remaining future thresholds.

An example makes the change from the Simple model concrete. Consider two hypothetical players who have both reached game 25. The first is 20 and debuted at 18, and the second is 24 and debuted at 22. The Simple model gives them the same probabilities of attaining all future thresholds because it knows only that each has played 25 games. The Age Enhanced model estimates that the younger player has probabilities of about 84 percent, 71 percent and 54 percent of reaching 100, 150 and 200 games, respectively. The corresponding estimates for the older player are about 57 percent, 32 percent and 12 percent. These show what age adjustment does: it shifts the player's entire continuation curve rather than adding or subtracting a fixed number of games.

The term proportional hazards describes the model's main structural assumption, which is that the underlying risk a career ends changes as the career progresses, but the relative difference associated with a player's age profile is assumed to operate proportionally across future game totals.

Because age enters through flexible curves, there is no single coefficient that can be read as the effect of being one year older. Predicted probabilities, such as those in the example, provide a more useful interpretation of the fitted age relationship.

The Age Enhanced model improved on the Simple model, confirming that players with the same number of games can have very different prospects of threshold attainment because they reached that point at different ages.

So, age is in as a predictor. Next we ask, among players of the same age, what other player information improves the forecasts?

MODEL 3: ADDING PERFORMANCE, POSITION AND OPPORTUNITY

We next fitted a Penalised Cox Proportional-Hazards model, which is “an extension of the standard Cox proportional-hazards survival model that adds a penalty term to the partial likelihood function. This penalty shrinks regression coefficients toward zero or sets some to zero entirely” (Google).

It retained the flexible age curves from the Age Enhanced model (which were nonlinear natural splines with three degrees of freedom for age at the landmark and age at debut) and added:

  • Player Position (as a categorical term and unpenalised) - the player’s most frequently recorded position up to the landmark game (excluding ‘unspecified’ and ‘interchange’)

  • Average AFL Player Rating - mean rating across every game through to the landmark. Captures the player’s underlying performance level across the period

  • Average Time on Ground - mean Time on Ground across every game through to the landmark

  • Time on Ground in Last 5 Games

  • Games per Observed Season - landmark games divided by the number of seasons in which the player has appeared

  • Games in Landmark Season - games played in the season containing the landmark, up to the landmark game

  • Rating variability and Recent Movement, measured by

    • AFL Player Rating SD - standard deviation of all match ratings from debut through the landmark. A higher value means the player’s ratings have varied more from match to match

    • Rating Recent Change - mean rating over the most recent five games minus the mean over the preceding five games. Positive values indicate recent improvement relative to the previous block (eg when fitted at Game 25, uses Games 21 to 25 versus Games 16 to 20)

    • Rating Trend - slope from a linear regression of rating against career-game number over the most recent ten games. It measures the direction and rate of the recent trajectory

  • Career Interruptions, measured by

    • Longest within-season Gap in days

    • Number of within-season Gaps of at least 35 days

  • Number of Teams - number of teams played for up to and including the landmark game

Except for Player Position and the two age-related variables, these variables entered through ridge regularisation, which shrinks unstable coefficients towards zero instead of allowing a noisy variable to dominate.

The table below descibes the importance and interpretation of each of these variables.

(Note that a hazard ratio below 1 indicates a lower estimated risk of the career ending at the next career-game step and therefore a longer predicted career. A ratio above 1 indicates a shorter predicted career.)

Note that an 'unspecified' position at game 1 carried a very large estimated career-ending hazard.

Also note that ridge regularisation shrinks coefficients towards zero but does not normally remove variables entirely. Consequently, all 14 predictors remained in every fitted Penalised Cox model, even when an individual coefficient was small or statistically uninformative


MODEL 4: DO THE VARIABLES ACT ALONE?

The penalised Cox model requires the analyst to specify in what form predictors enter the model. The next model, the Random Survival Forest model, is designed to determine whether complex combinations of player characteristics might improve prediction even when some of those characteristics appeared unimportant when included on their own.

This modelling algorithm builds many survival trees, each with access to its own subset of the total player data drawn with replacement, and each dividing those players into groups with different career-ending game counts using a random subset of the available predictors. The next step is to create a cumulative hazard curve for each tree node using the known future career game totals for the players in that node.

The details behind this modelling technique are quite technical. To learn more about the algorithm see, for example, this paper. The main thing to understand is that each of the constructed trees is capable of providing an estimate of:

  • the probability of reaching 50, 100, 150, 200 or 250 total games

  • the expected number of additional games, obtained from the area under the survival curve

  • quantiles of the predictive final career-total distribution

These individual tree estimates can then be aggregated for a given player who has been run through all of the trees to land at a terminal node in each.

The forest used all of the predictors used in the Cox Proportional-Hazards plus a handful of new variables:

  • Number of Player Positions (excluding ‘unspecified’ and ‘interchange’) - number of distinct positions played up to and including the landmark match

  • Landmark Rating - the AFL Player Rating in the landmark match

  • Landmark Team - the team played for in the landmark match

  • Recent Rating - the average rating across the five most recent games up to and including the landmark game

  • Recent Gap - the days between the landmark appearance and the immediately preceding appearance

One of the benefits of fitting forests is that overfitting is less likely, so adding new, even highly-correlated variables, is less problematic. Also, repeated tree splits allow one predictor's contribution to depend on another without imposing constraints such as proportional hazards.

The forest's importance rankings provided a coherent progression:

  • At game 1, debut age and number of positions played contributed most predictive value

  • At game 10, age remained most important while average and recent rating began to matter

  • At games 25 and 50, games per season and recent rating moved upward in the order of importance

  • At game 100, age was still most important, followed by debut age, recent rating, recent rating change and average rating.

The forest does not produce conventional coefficients. Its importance values measure how much a variable improves prediction, including through nonlinear relationships and interactions. It does not show whether the variable increases or decreases expected career length and should not be used for causal interpretations.

MODEL 5: DOES MODEL AVERAGING HELP?

Averaging models can reduce the dependence on the sets of assumptions associated with each of them individually. As a final model, an Ensemble model was created by averaging the survival probabilities estimated from the Simple model (Model 1), Age model (Model 2), the penalised Cox proportional hazards model (Model 3) and the random survival forest (Model 4).

All five candidate models - the four base models and the Ensemble - were tested on players from later debut cohorts than those used for fitting as per the protocol described below:

To be specific about how each fold was used, consider Validation Fold A. In using that Fold, each candidate model was fitted separately at each landmark using eligible players who debuted between 2012 and 2014. Its predictions were then evaluated on eligible players who debuted in 2015 or 2016 and who had reached the same landmark, using only landmark–threshold comparisons with sufficient observed outcomes (that is, with at least five players known to have reached the threshold and five retired players known not to have reached the threshold).

Two performance metrics were used to evaluate the out-of-sample fit of the models

  • the censoring-adjusted Brier score, which is the mean squared error of a probability forecast after correcting for careers whose threshold outcome was not yet known (see this paper for a discussion of this method). On this metric, lower is better

  • the time-dependent area under the receiver operating characteristic curve measured ranking ability (see this paper for a discussion of this method). Here, higher is better.

The Brier score and AUC results for each of the five models evaluated at each of the five landmarks are summarised in the plots below.

Figure 1. Censoring-adjusted Brier score by model and landmark. Lower values indicate more accurate probability forecasts. The Simple model is consistently weakest.

Figure 2. Time-dependent AUC by model and landmark. Higher values indicate better ranking of players who did and did not reach the relevant career-game LANDMARK. AUC was secondary to the Brier score in model selection

The tables below summarise the choice of best model for each of the five landmarks using, firstly, the average Brier Score, then the average time-dependent AUC.

Using Brier Score as the metric, the Penalised Cox model is selected for models built after games 1, 10, 25 and 100, while the Age Enhanced model is selected for models built after game 50.

Instead using AUC as the metric, the Penalised Cox model is selected for models built after games 1, 50 and 100, while the Ensemble model is selected for models built after games 10 and 25.

Using the Penalised Cox model for all landmarks and thresholds would be the pragmatic choice, but for the remainder of this blog we’ll choose the preferred model based on the Brier Score table.

HOW WELL-CALIBRATED ARE THE PREFERRED MODELS?

In the previous section, models’ time-dependent AUC results were reported. One interpretation of the AUC metric is that it measures a model’s discrimination, that is, its ability to reliably attach higher probabilities to players that turn out to reach a specified future threshold compared to players that do not. This section turns to the separate question of model calibration.

Calibration asks whether the estimated probabilities turn out to have approximately the correct numerical magnitude - for example, among players given a 60 percent probability of reaching some threshold, do roughly 60 percent do so? Discrimination and calibration measure different, desirable qualities of a model that are not necessarily positively correlated. A model can, for example, tend to rank players correctly while assigning probabilities that are systematically too high, too low, or too conservative (ie near 50%).

Figure 3. Later-cohort calibration for the selected model at thresholds through 200 games. The diagonal represents perfect agreement between predicted and observed continuation.

Calibration was broadly reasonable where the validation data contained enough observed threshold achievers and careers completed before the threshold to estimate it reliably. The evidence was weaker for estimates relating to achieving 200 career games because relatively few players could be classified conclusively on both sides of that threshold, especially amongst players who debuted most recently.

No combination of prediction landmark and historical validation fold contained at least five players known to have reached 250 games and five players whose completed careers ended below 250 games. The 250-game probabilities can still be estimated from the fitted survival curves, but their calibration cannot be assessed reliably. As such, those estimates should be treated as exploratory forecasts rather than historically validated probabilities.

WHICH VARIABLES ARE MOST IMPORTANT TO THE AGE ENHANCED AND PENALISED COX MODELS?

Coefficient size is not a reliable measure of importance, especially when predictors have different units, age is represented by several spline terms, and position is represented by several contrasts. To obtain some rough estimates of importance we perform a grouped permutation analysis on the same later-cohort validation folds used to compare the models earlier. Within each test fold, one complete predictor group was shuffled 20 times while the fitted model was left unchanged. Importance was measured as the increase in IPCW Brier score and the reduction in time-dependent AUC caused by that shuffle. Positive values mean that disrupting the predictor made out-of-sample prediction worse.

The results are summarised in the table below.

(Note that the Age Enhanced model is preferred only for the Game 50 landmark and the Penalised Cox for all other landmarks.)

The Age Enhanced model has only two predictors: landmark age and age at debut. Landmark age dominated at every landmark. Debut age added information once players had accumulated games, but at game 1 it added nothing separately because age at the debut match and debut age have the same value for every player.

The Penalised Cox results tell a consistent story. Landmark age remained the strongest predictor at every landmark. Position and mean rating supplied the main additional information at game 1, although the position result partly reflects the unusually adverse signal carried by an unspecified position. From games 10 to 50, mean rating became the most important non-age predictor. At game 100, recent rating change and recent-five time on ground were the leading non-age variables in the permutation analysis. Note that these estimates are relatively uncertain because only three validation comparisons were eligible..

Finally, note that these values measure how much performance deteriorated when one predictor was reassigned among players while all other predictors were retained. This procedure measures both the variable’s direct predictive contribution and information it shares with other variables and so can tend to understate the predictive value of variables that are highly correlated with one or more other variables, because those variables have not been reassigned among players. It can also produce combinations of predictor values that were uncommon in the observed data, resulting in imprecise estimates of the effects. So, the importance values shown here should be treated as comparative diagnostics rather than precise measures of independent effects.

FOR WHICH PLAYER TYPES WERE PREDICTIONS LEAST ACCURATE?

A single overall Brier score or AUC value for a model can hide its uneven performance for some subsets of players. To investigate this, we recalculated the metrics for each of the selected models by position and debut-age groups, retaining only cells with sufficient membership numbers to justify the effort (ie at least five known threshold achievers and five known non-achievers),

The table below summarises the results.

(A reminder that the Brier score measures probability error, with lower values indicating better predictions. Time-dependent AUC measures discrimination - that is, the ability to assign higher probabilities to players who reached the threshold than to players whose completed careers ended below it - with higher values indicating better discrimination.)

Figure 4. Supported subgroup Brier scores. Lower is better. Sparse positions and older-debut groups disappear at long horizons when minimum case/control support is not met.

It’s interesting to hypothesise about the bases fro some of these results.

For example, why do players who debuted before age 19.5 have lower average probability error and stronger discrimination in the eligible validation comparisons? At a given landmark, all players have played the same number of career games, so this difference cannot be attributed to younger debutants having provided more data on which to make predictions than have older debutants at the same landmark. Instead, age at debut may act as a proxy for factors such as early club confidence, perceived remarkable talent, or draft investment.

Note, however, that the middle and oldest debut-age groups had similar average probability error values to each other, which warns against hypothesising an overly simplistic age story.

Looking next at positions, half backs had the highest average probability error among the position groups that had sufficient observed outcomes to justify model-fitting. Inside midfielders showed the highest average discrimination, although this result was supported by only five eligible landmark–threshold comparisons, compared with 10 to 12 for the other reported position groups and therefore covered a narrower range of validation settings and should be interpreted cautiously.

ESTIMATING REMAINING CAREER GAMES FOR CURRENT PLAYERS

The final piece of analysis we’ll do in this blog is to apply the preferred models to all of the players who were in a named squad in 2026. The output from that analysis is available in this Excel spreadsheet and also in this Excel spreadsheet. Eagle-eyed fans of particular teams will no doubt find possible or even blatant errors in this data. Please let me know if you do and, even if you don’t, what you think of the analysis.

In interpreting the numbers in those spreadsheets bear in mind that each current-player estimate contains two kinds of seperately estimable uncertainty:

  • A predictive interval describes the range of final career totals that a player could plausibly realise under the fitted survival distribution. It represents uncertainty about the player’s eventual outcome, including future circumstances not known at the prediction point, such as injury, selection, form, team circumstances and chance. Predictive intervals can remain wide even when the underlying survival probabilities are estimated relatively precisely.

  • A bootstrap confidence interval describes uncertainty in an estimated threshold probability - for example, the estimated probability that a player will reach 150 games. To assess this uncertainty, we repeatedly resampled historical players and refitted the player’s latest-landmark model 50 times. The middle 95 per cent of the resulting probability estimates formed an approximate 95 per cent confidence interval. Across current players, the median width of the interval for the probability of reaching 150 games was 17.8 percentage points, and the widest was 62.2 percentage points. A wide interval indicates that the estimate is sensitive to which historical players are included in the fitting sample, often because the player’s combination of characteristics has limited historical support.

The expected career total is a restricted estimate because the fitted survival curve is not extended beyond the range supported by the historical data. The predictive limits are subject to the same restriction. When an upper limit extends beyond 286 career games, the output reports that it exceeds the observed support rather than assigning a precise total unsupported by the data.

Also note that 50 bootstrap refits are enough to provide a useful indication of model uncertainty, but they are relatively few for estimating the tails of a 95% interval. The reported bootstrap bounds should therefore not be interpreted as highly precise.

Only after the endpoint, predictors, candidates and validation rules were fixed were the models refitted and applied to current players. The model used for each player was the one corresponding to their most recently achieved landmark. The data contains 694 active players with at least one exact landmark history. Their latest usable landmark is game 1 for 90 players, game 10 for 106, game 25 for 97, game 50 for 145 and game 100 for 256.

A player's record is deliberately an as-of-landmark forecast. If a player has played 83 games, his latest score uses only information through game 50, not games 51 to 83. This preserves the meaning established in validation and prevents current totals from leaking into a model that was tested at an earlier decision point.

The 250-game probability is included because it is strategically interesting, but it has not passed the later-cohort validation support rule. It should be treated as exploratory, especially for very young players whose predicted careers extend beyond observed historical support.

COMING CLEAN ABOUT DIRTY DATA

The most difficult data problem was not match performance, but was instead determining whether a player had truly finished his career. FootyWire data was searched separately for current-list and retired-player records for every team and the raw pages were cached. The extraction produced 810 current rows and 1,684 retired rows. Two current records and one retired record lacked a date of birth, reducing the safest linkage information for those records.

Of the 1,825 players in the 2025 development endpoint, 1,699 matched FootyWire exactly on normalised name and date of birth. Another 22 required a unique normalised-name match and 27 required a unique surname-plus-date-of-birth match. These secondary rules were used only when the result was unique. They mainly addressed shortened given names, full given names, initials, suffixes such as 'Jnr', and source-specific punctuation.

Seventy-one players did not have a usable current or retired linkage but had missed two complete seasons. They were labelled as completed careers under the fallback rule, with the source retained so that any future sensitivity analysis can exclude them. This group includes well-known long-career players as well as short-career players, which suggests the failure was largely a coverage/linkage issue rather than a single type of career.

Other completeness problems affected eligibility for modelling inclusion rather than career status. Careers already under way when the rating panel began were excluded because their early games were missing. Some landmark rows had uncertain position data because no reliable starting-position record had yet been observed. Nine time-on-ground values lay outside 0 to 100 and were treated as missing.

CONCLUSION AND NEXT STEPS

We started with a deceptively simple question that, in essence, was “can we estimate how many remaining games a player has left in his career?”. The Simple model supplied a baseline, the Age Enhanced model supplied the first model that allowed individual player distinction, and then AFL Player Ratings, time on ground, and selection continuity data added additional predictive information used in the Penalised Cox model. The Random Survival Forest model tested flexible interactions but was not ultimately selected for use. The Ensemble model tested whether averaging the four underlying models improved predictive validity.

A best model - usually the penalised Cox but the Age Enhanced model for one landmark - was chosen for modelling at each landmark event The resulting forecasts are most credible for the 50-, 100- and 150-game thresholds and for the game-25 and game-50 landmarks. They are less certain for middle-aged debutants, backs and half backs, sparse positions, game-100 long horizons, and 200- or 250-game outcomes.

Doubtless, a club with access to data about the reasons for player absences and career endings could substantially improve the models described here and, in fact, built separate models for career endings due to injury, age, or other causes