Projecting MLS Performance based on MLS Next Pro Data — American Soccer Analysis



Alex Freeman, FB, 2023 MLS Debut – Orlando City

Freeman obviously took the world by storm, going from Next Pro prospect to World Cup starter in basically a year and a half. Our projection had him as a slightly above average MLS fullback, and you can see all the g+ components actually track quite well. Similarly, every non g+ metric fell into his P25-P75 (the most likely 50% of outcomes) except for xG, xA, shots, key passes, and receiving g+. The model simply didn’t know he’d become American Dani Alves under Oscar Pareja, but this is a pretty good call regardless.

Adri Mehmeti, DM, 2026 MLS Debut – New York Red Bulls

Mehmeti has been one of the surprise young stars of the 2026 season after being one of the best sixes in the entire MLSNP pool. His strong passing was projected to translate, and it has done so better than expected. Where it’s missed has been on the interrupting and receiving side. Mehmeti has been inserted into a much more attacking role, frequently making runs beyond the back line into the penalty area, and less the metronomic hub breaking up play he was with the Baby Bulls.

Alonso Coello, DM, 2023 MLS Debut – Toronto FC

Coello is in some ways the opposite of Mehmeti. His projection looks like a perfectly cromulent, well rounded defensive midfielder, that’s a useful player. His passing and carrying as a progressive hub held up somewhat better than expected, but where it mostly missed was on the interrupting g+ side. Coello’s first season with the first team was a lot of defending the penalty area, racking up interrupting g+

Benjamin Cremaschi, CM, 2023 MLS Debut – Inter Miami

I’m quite happy with how accurate the Cremaschi prediction was. A somewhat below average midfielder who struggles to progress the ball enough goes up to the first team and is a below average midfielder who struggles to progress the ball enough. What the model didn’t know is that Cremaschi would get to play with the greatest creative passer of all time and become a willing box runner to receive Messi passes to juice that receiving all the way up.

Cavan Sullivan, AM, 2024 MLS Debut – Philadelphia Union

This is a good edge case example for what this model does tell you, and what it doesn’t. 17 of Sullivan’s 19 metrics fall in his P25-P75 band, with take ons and pass completion falling outside. But it all looks wrong to the eye. However, Sullivan has one of the widest projection bands in the entire class for a few important reasons. 

  1. His extremely young MLSNP age basically makes him a total wildcard. 

  2. The model stops tracking your MLSNP production once you make your MLS debut (more on this later), so he only had 500 MLSNP minutes, while it took him two more years to hit 1000 MLS minutes, during which he played another 2000 MLSNP minutes (and became dominant in the NP league).

  3. Because of 2, and his growth from 14 to 16 from a passing AM into a risk taking dynamic wide attacker, the profile looks extremely different. So it’s just wrong.

  4. There are exactly three MLSNP AMs promoted to the first team listed as an AM, Dado Valenzuela, Cavan Sullivan, and Kenji Mboma Dem, and only Valenzuela made it all the way to 1000 minutes, so it makes transitions awkward to map.

Regardless, Sullivan has developed into a really effective wide attacker both as a passer and off ball mover. Good for him.

Ousseni Bouda, W, 2024 MLS Debut – San Jose Earthquakes

Some things translated better than expected, but the projection captures the shape of the player he’s become quite well!

Jacen Russell-Rowe, ST, 2022 MLS Debut – Columbus Crew

JRR was one of the original MLSNP graduates that had people’s ears perking up at how viable it was to generate first teamers, and the projection model handles him really well. Basically dead on what his first 1000 MLS minutes looked like. 

It’s a fun exercise to go through and compare some of the noteworthy names back against it, it certainly doesn’t hit on everyone, but it feels pretty good given how little data we still have.

The biggest limitation is that we simply don’t have enough transitions between the two leagues. To try and be cute about this, I used a technique called K-fold cross validation. Normally, we’d train the model on the first four MLSNP seasons and then test it on the most recent one. The problem is that’s a test set of approximately 50 players in 2026, and there’s just going to be huge noise there. So instead, we run the model where one fifth of the player-transition set gets to be the test dataset, and the four remaining are the training set. Then we can take the average across those five folds. This gives us a more stable view of what changes to the model are helping and hurting. I chose five folds somewhat arbitrarily (there are five seasons), but have gotten some feedback since that with such a small player-transition set, we could have done “leave-one-out” cross validation. This would be effectively training the model on 199 players to predict one, and then doing that 200 times. A quick unoptimized trial of this improved the GK model by 2%, and the player model by 1%, perhaps some ground to cover there. Nonetheless, I would bet that in 10 years this model has tightened up considerably.

Similarly, some of these metrics have considerable season-to-season variability even for players who play the exact same role for the exact same club. A perhaps more accurate way to evaluate the model would be error above regular, run of the mill, variance. We see in many of the largest misses, it’s goalkeepers whose shot stopping value swung massively season over season, and strikers who maintained an unbelievably high level of receiving value.

Another thing I quite dislike is that the model stops tracking the MLSNP performance once the player makes their MLS debut. This is similar to how a transfer model would work, given you cannot play for two teams at the same time, but you can do that with a youth team and a first team! In fact, across the 197 graduating players only 15% don’t come back down, 50% play another five 90’s, and 10% play another 20 90’s with their respective reserve team. The vast majority of players continue to play MLSNP over the time it takes them to accrue 1000 senior minutes (about a calendar year on average), so you have this weird, blended period. I tried including these minutes into the model and it didn’t seem to help across the board (MSE reduction of 0% on the CV folds, 6% on the 25/26 holdout), though I’m sure in specific use cases it would and there’s a more clever way to handle this.

The model includes some information about how the players in their position with the first team are doing. I think there are probably significantly more insights to dig into about the specifics of the transition that could make a model much better, but given the limited data now I didn’t spend much time fleshing out tiny slices. If you have thoughts, I’m always open to hear more.

This feels like a good place to point out that the model is only predicting the first 1000 MLS minutes of a player, not how likely they are to figure it out from there. These are likely to be the least productive minutes of a player’s career. For example, Julian Hall was a frequently requested backtest candidate that I left out, because his first 1000 MLS minutes were actually quite weak and the model gets that mostly right. But minutes 1000 to 3000 have been excellent.

One thing I don’t account for at all is that generally talent evaluators are largely quite good at their jobs. There aren’t droves of USMNT players in your local park that were just incorrectly evaluated, so the players that do get called up are generally the best players. I suspect in actuality, the transitions for the bottom 50% of MLSNP players would be worse than this model suggests.

To actually generate the projections, I used six potential models for each metric.

  • B0, Direct Translation of MLSNP Performance: Take whatever the player was doing in MLS Next Pro and write it down unchanged as the MLS prediction. No adjustment at all. Every “+42%” in the project means “42% less error than B0’s guess”.

  • B1, League-Position Translation Factor: Work out, across all past call-ups, how much a player’s numbers change graduating from MLSNP then apply that same haircut to everyone in the position. One translation factor per position per metric.

  • M1, Ridge Regression from B1: Starting with the league-position translation factor, use a ridge regression model based on the two clubs, the player metrics, and the players age.

  • M2, LightGBM: Gradient boosted model with the entire context block we mentioned at the beginning, with a quantile objective. More on this later.

  • GKRidge, a GK version of B1: There are so few GK transitions that it’s basically just B1 with much harder pushing towards a flat translation factor.

  • TabPFN, a prior fitted transformer: We met our friend TabPFN in the goalkeeper cross claiming model and it did a good job with a small dataset there, and given we’ve got an even smaller one here I figured I’d give it a shot. 

We tried out these 6 models on all of the metrics, with GKRidge being exclusive to goalkeepers (duh) and picked whichever one was best for a given metric. You can see their comparison in terms of MSE% reduction in the graphics below and the model selected for each metric as well. For outfielders, LightGBM is the model used for 16 of 19 metrics for outfielders, with a flat translation factor for shots, goals added shooting, and progressive carries. Surprisingly, TabPFN improves upon Light GBM for only a few metrics (and narrowly so) and is much worse for most. I’m not really sure why, but I’m far from an expert on these foundational transformer models so if you know, hit me up. As such, because TabPFN is a bit annoying to implement due to the necessity for a GPU, unless it was better than everything, I didn’t want to use it for any of the metrics. 

For goalkeepers, TabPFN does a better job with the extremely small dataset here. For the bottom four metrics (crosses faced, claims, handling, fielding), it actually performs very poorly, but those metrics are A) somewhat less important to me or B) very small in magnitude, so the errors are extremely small anyways. For simplicity’s sake it’s TabPFN for everything for GKs.

We will be happy to hear your thoughts

Leave a reply

Som2ny Network
Logo
Register New Account
Compare items
  • Total (0)
Compare
0
Shopping cart