Odd Lots - Why Soccer Analytics Works Like Volatility Arbitrage Trading
Episode Date: July 16, 2026American sports fans have long been comfortable talking in the language of stats and analytics. Soccer embraced the 'moneyball' revolution later; the sport was once perceived as too complex to model a...nalytically — there were too many players on the pitch, the game's progression was too random and chaotic to reliably predict. That's no longer the case, and soccer watchers are well aware of stats like xG (Expected Goals) and each match is an opportunity for a team to mine data, whether its tracking data, on-ball data, or even analyzing body poses and movement. Today, we speak with two soccer analytics veterans, Mike Treacy (head of risk at Apex Fintech Solutions) and Joris Bekkers (a soccer analytics consultant). Treacy's background includes a stint in analytics for a Premier League team and he's currently advising the MLS team Austin FC while Bekkers has built software that analyzes raw soccer data and he's worked with the US Soccer Federation. We talk to them about how VAR has affected the sport, how data analytics can capture ineffable things like hustle, how European leagues and the MLS differ in their analytics strategy, and why chess and soccer are not so dissimilar. Read more:The Lawyer Taking On StubHub Over World Cup Ticket SalesPolymarket Partners With Crypto Firm During World Cup Only http://Bloomberg.com subscribers can get the Odd Lots newsletter in their inbox each week, plus unlimited access to the site and app. Subscribe at bloomberg.com/subscriptions/oddlots Subscribe to the Odd Lots NewsletterJoin the conversation: discord.gg/oddlotsSee omnystudio.com/listener for privacy information.
Transcript
Discussion (0)
The Big Take podcast from Bloomberg News keeps you on top of the biggest stories of the day.
My fellow Americans, this is Liberation Day.
Stories that move markets.
Chair Powell opened the door to this first interest rate cut.
Impact politics, change businesses.
This is a really stunning development for the AI world and how you think about your bottom line.
Listen to the big take from Bloomberg News every weekday afternoon on the IHeart Radio app, Apple Podcasts, or wherever you're at
you get your podcasts.
Bloomberg Audio Studios.
Podcasts Radio News.
Hello and welcome to another episode of the Odd Lots podcast.
I'm Joe Wisenthal.
And I'm Tracy Alley.
Tracy, I have a question.
I know you spent a lot of your youth overseas.
My youth.
Oh, dear.
Did you ever go to many baseball games as a kid?
Yeah.
So I was in Chicago for a few years.
So I went to the Cubs games.
And then I was in Japan.
And the baseball scene in Japan is amazing.
like the best vibes of a live sports event that I've ever, ever witnessed or encountered.
I was just talking to someone about this last night.
I've always wanted to go to a Japanese baseball game.
Highly recommended.
That's actually like, if we ever do a live show in Tokyo, let's try to schedule it during baseball season.
Because that is like, I sort of think that's a bucket list thing.
But anyway, the reason I asked this question is I have this really vague memory as a child going to see Detroit Tigers games with my grandfather.
Like probably when I was like, maybe these.
memories are probably from when I was younger than like six or five. But there used to be a non-trivial
number of people who would go to the games and they would keep score and they would write down
every single it bat and the outcome of every single one. I'm pretty sure I've seen that,
not in person, but like maybe in movies or something like that. It doesn't really happen anymore.
Like I never see it. There may be a few like old timers who still as a hobby or habit do that.
But it was like a non-insignificant number of people. And it's interesting to me, you know,
I've been to a couple of soccer games this year.
There is no equivalent way you could do that, right?
Because it's like baseball is filled with all of these discrete events.
The pitcher, who is the pitcher?
Who is the batter?
Hit, not a hit.
Single, not a strikeout walk, et cetera.
Like, what would even be the equivalent in soccer?
So this has been a long-running debate in soccer.
And I remember when Moneyball came out and, you know, sports analytics became a big thing,
especially for baseball because, as you point out, it's these sort of discrete events that
have a lot of statistics embedded in them. A lot of people were saying that soccer, you could never
use data analytics in the same way for soccer. Like it's too chaotic, it's too fluid, there's too
many variables, there's not enough goals. That complaint comes up a lot whenever we talk about,
are we going to say soccer or football, by the way? You know what? Let's just say soccer.
Okay. All right. I'll try. We can say football. I actually don't feel strongly about this one.
But that said, we do see soccer analytics on the rise, right? Like now we get to,
all these stories about like tiny clubs that are using data to source, you know, specific players
in very moneyball style ways. We obviously have prediction markets where people are doing a lot
of sports betting. And so all the analytics seem to be like becoming more important. And I will
just say, I've read this crazy stat right before we came on from the law firm Morgan Lewis.
And they were saying in the 26 FIFA World Cup, match-based data is going to be,
something like, so it's 104 matches generating more than 90 petabytes of data.
Yeah.
Which is a 45-fold increase over the volume produced in the last World Cup in 2022.
That's stunning.
So this raises a question, and I know we're going to get to this in the conversation.
And it almost is like a philosophical question, which is, okay, we see the game of soccer.
It's like very fluid, right?
There's no, we just said there's fewer discrete events.
Like, in theory is something that's fluid, a series of, like, microscopic discrete events.
You know, could you get a million frames per second and actually turn it into discrete events?
Like, that is sort of like an interesting question to me.
A parallel that I think of in this conversation is chess is like all discrete events.
And that was seen as like a very difficult thing to crack.
But then, like, over 25 years ago, they cracked.
Yeah, of course.
They cracked it.
But then you would think, okay.
Well, like, what is the opposite of chess, which would be language?
And you say, okay, you can't.
That's fluid.
That's all over the place.
And yet computers seem to be understanding language pretty good.
My inner Luddite says there must still be a secret, like, unmodelable thing.
I can't talk when it comes to like football, the beautiful game.
But you're right that technology may prove me wrong very quickly.
Well, as you mentioned, you know, soccer analytics growing.
And obviously the interest now is for obvious reasons.
there's all the betting on the World Cup.
People talk about the XG of like, I don't know, a situation or a player, the expected goals.
Now we have body posing analytics as well, which kind of blows my mind.
And when our producer Kale introduced me, I hadn't seen them before earlier this year,
like the momentum charts that show like, you know, how dominant a team is in any given moment.
You see it going waves and stuff like that.
So there clearly is a lot of work being done.
I don't know how much it works.
I don't know how much like momentum charts consistently predict who is going to be the winner.
But just this question of like the ability to model deeply fluid things, the beautiful game.
Like it's an art, right?
Like can you can computers actually model this?
And then like if they can, what does that say about the ability to model a bunch of other things?
Strikes me as just like a very relevant question right now because it's the World Cup,
but also relevant question in general about where.
For markets and finance.
For markets and finance.
Exactly.
So I'm curious to know like how it works.
Yeah, let's talk about it.
How the moneyball revolution, where we are in the arc with soccer.
Anyway, I'm really excited to say.
We really do have two perfect guests because they do sit right in this space and also
at the intersection of everything that we're talking about.
We're going to be speaking with Joris Becker's.
He is a soccer analytics consultant, a professional soccer analytics consultant for the last
decade.
As well as Mike Tracy, he's the head of risk at Apex FinTech Solutions.
He used to be a volatility arbitrage trader at peak six.
And he also does soccer analytics for the club FC.
So literally the two perfect guests to talk about some of these questions that are arising right now.
So Mike and yours, thank you both so much for coming on the podcast.
Thank you for having us.
This is a great introduction.
Yeah, thank you.
You know, Mike, let me start with you question.
A lot of people, I think, have, for good reason, an intuitive feel that, like, sports
analytics and trading markets are adjacent ideas that like we know this. We know that the,
a lot of the prop shops and market makers do both and so forth. How would you articulate based
on your career, the sort of overlap between the skills and techniques you've developed as an
arbitrage trader, a volatility trader, and the overlap of skills that they're like, okay,
this thing that we'll get into that we call Salker Analytics.
Well, soccer is unique because I'd say every aspect of the game is distribution.
Okay.
When you look at on pitch, the performance by a team, when you look at performance by an
individual player, when you look at the seasonal outcomes and this unique component
that is promotion relegation.
So your finances are highly variant year over year.
And so there's a lot of overlap between volatility trading and working in soccer in that you are making highly levered bets on often imperfect information to them.
That's not necessarily predictive like other sports, such as baseball.
Can I ask a very basic question, which is what is the point of soccer analytics?
So if we go back to the market's analogy, we talk about price discovery.
right? Like you're trying to find the right price for a particular asset. With soccer, are you trying
to price players, improve the training, make more successful predictive bets? I imagine it's a bunch of
different things. Yeah, I'd say everything. I think you said what are the things you're trying to do?
And I think from an investor operator perspective, it's what are you trying to not do? You're not trying to
get relegated and you're not trying to spend 30 million pounds on a player that is going to be
terrible and you'll have to get rid of in two years. Yours, why don't you talk to, from your
perspective, I mentioned your soccer analytics pro. Actually, our producer, by the way,
just put it in the chat, the Bloomberg style guy, this football, which is injured. Maybe we'll just
switch to, yeah, let's stay in a company style. You're a football analytics pro. Tracy asks,
like, what is the point? Is it more on the team side of things in terms of identity?
identifying talent, et cetera, or is it more on, I don't know, I guess, like the sort of the betting
side, like what is the problem you and your professional capacity are trying to solve?
So it highly depends on the organization, right?
Some organizations might want to, like Mike said, make good hires, make sure that they
don't lose millions of pounds.
Some organizations might use soccer analytics to try and improve their strategy, so they
might look at in-game data and try to figure out if there's some optimization that they can make
based on where the passenger should go or where players should be, what play styles make more sense
or what yields more expected goals, if you will. And then I guess the third one is entertainment.
There are a lot of apps out there, websites out there that provide fans with some sort of information.
And even back in the, let's say, in the 90s, they would have these overlays on the TV where they show
number of corners, number of yellow cards, and possession percentage.
I remember that.
It doesn't mean anything, but it's interesting.
So you can only imagine that you can go much deeper with that and people will still be engaged.
Actually, just say more on that.
You said it doesn't mean anything.
Is that something that has always been understood?
Or is this something in 2026?
We could say that a few of these stats that they like to put on the TV were just extremely like low signal.
data point. Yes. So when I say they don't mean anything, I don't think those were necessarily
predictive of the outcome of the match. Got it. Corners might because it's an indicator of
which team it has the overhand in a game, but I'm quite sure most of the others, like possession
percentage, don't mean all that much when I'm trying to look at game outcomes. Wait, so we touched
on this in the intro, but like the perceived wisdom in the sort of early 2000s was that soccer was
far too complicated to be given the baseball moneyball style treatment. What actually changed to get
us to this point where we're not just talking about things like expected goals, but we're also
talking about like body movements and posturing and things like that. You know, baseball had money
ball in 2003 and soccer had two things in the 2010. So in 2013, Chris Anderson and David Sally
published this book called The Numbers Game that distilled soccer.
down to more of a weakest link game.
And then the other thing that happened was,
I would say we had a bit of this revolution on Twitter,
which you kind of spoke about recently, Joe,
where ideas were incubated.
And I give Michael Cayley a lot of credit for this.
He's done the pod.
But every Saturday in the Premier League,
you'd have six matches,
and then you would wait 30 minutes,
and Kaylee would just post a stream.
of every expected goals chart.
And it started to stimulate this discussion of, does this represent what should have happened,
what should have been the outcome?
And I think it started to lead us down this path of where is XG flawed?
It only registers when you have a shot.
And so there is much more to the game than simply which shots occur.
There's possession and the threat of each possession.
There's match momentum and there's changing styles.
So it's evolved from on-ball data to tracking data to now we're going to that deeper layer of body pose and what movements can prevent or create opportunities to score.
I think what goes hand in hand with that is when I started 10 years ago, soccer football was perceived as too complex.
There were 22 players.
There was a ball.
People were doing interesting analytics research and applied research already in basketball.
And what was said was, well, there's five players on each team.
So it's a lot less complex.
There's more games.
And we have more data.
And there's more scoring.
So you could derive metrics a lot easier.
And I think the evolution in just how AI was applied in data availability.
So going from, like Mike said, only on ball events where we know which player made the pass, which player made the shot.
But we always have to say, well, we don't know where all the other players are because those are not recorded.
And now we have the ability to make highly complicated artificial intel or neural nets with this positional tracking data where at 10 frames per second or 25 frames per second, we know where all the players are and we know where the ball is.
And then additionally, at some point, we will get full body pose.
So we know where like the whole skeleton of all players are basically.
Like that sort of went hand in hand and it felt like it was inevitable, but also the game is just more complex.
So it needed more compute.
It needed more data.
Needed more knowledge.
Median women are looking for more to themselves, their businesses, their elected leaders, and the world are out of them.
And that's why we're thrilled to introduce the Honest Talk podcast.
I'm Jennifer Stewart.
And I'm Catherine Clark.
And in this podcast, we interview Canada's most inspiring women.
Entrepreneurs, artists, athletes, politicians, and newsmakers,
all at different stages of their journey.
So if you're looking to connect, then we hope you'll join us.
Listen to the Honest Talk podcast on Iheart Radio or wherever you listen to your podcasts.
The Big Take podcast from Bloomberg News keeps you on top of the biggest stories of the day.
My fellow Americans, this is Liberation Day.
Stories that move markets.
Chair Powell opened the door to this first interest rate cut.
Impact politics, change businesses.
This is a really stunning development for the AI world and how you think about your bottom line.
Listen to the big take from Bloomberg News every weekday afternoon on the IHeart radio app, Apple Podcasts, or wherever you get your podcasts.
Tracy, Mike mentioned that, you know, just looking at a shot on goal, so you got to tell you so much.
for example, it doesn't tell you if the referees are going to revisit a call made a minute before the shot
halfway on the other half of the screen and take away the goal.
This was going to be my question, which we were talking about before the podcast recording.
But like, how do you factor in, let's say, the occasional randomness of the game and perhaps some erratic,
trying to be diplomatic here, erratic decision making by referees and footballing bodies?
I think it's control the controllables.
That's a fixed parameter within the game is that uncertainty that you can't control.
And it's hard to predict.
There is human error, obviously, and it's impossible to isolate when that's going to happen in a match and how it will.
You just sort of play the game.
You also don't look at individual events per se.
You might from an analysis perspective, but if you're building a basic predictor,
model for football outcomes, you might just take all the scoreboard results from the last
10 years and then all those things cancel out. Like if one team has a red card somewhere,
you wouldn't even know from these models, but you can build some interesting, let's say,
rudimentary predictive models with just the outcomes of games. Out of curiosity, has VAR changed football
analytics at all or presented new opportunities for data? It presents new opportunities for data
because that is an event that happens.
If a player is fractionally offside because his hand is ahead of the last defender,
that doesn't take away from the fact that he still got into a good opportunistic position,
possessed the ball, turned, and struck the ball in the net.
And often that data is nullified because of the var,
but in theory, it's something that you should consider in your dataset.
Within the raw data itself, there's tons of data that you could add or censor out.
You know, yours mentioned red cards, and a lot of our models, we censor that out of our data set.
There's an infamous match from three years ago between Chelsea and Tottenham, where Tottenham went down to nine men.
and their response was to play a very high line,
meaning they put all of their defenders up near midfield
and tried to catch Chelsea off sides.
And as one would expect, Chelsea proceeded to score multiple goals later in the match.
And one player in particular, Nicholas Jackson,
had three goals in that one match.
and when you look at his seasonal outputs for that entire season,
that three goals was probably around 20% of his total goals.
So when something like that happens,
when you have an irregular game state,
it is wise to sort of manipulate your data and remove that
to give you a full picture of what does this game look like at an equal game state.
Oh, that's interesting.
So it's not like that game in particular was a rich,
fountain of data, it's important to sort of like recognize that the data from this particular
game is not going to be particularly predictive about other games. And therefore, you know,
a guy who scores three goals of that match is probably not going to continue to score three goals
again for the rest of this. It was a rich fountain of bad data. Okay. And so at the end of the
season, when you see people analyzing the player, they often analyze their season. And what do they do
per 90 minutes of football.
And so you have this highly skewed data set by some really poor data.
And so there's a lot of data mining involved in the process of building a model when you
evaluate a team and individual players.
So here's a question I have, and you're talking about using neural networks.
And it's like eventually, like, you know, we'll have the compute to like have the position
of every player's body flashed maybe, you know, hundreds of images, frames per second, and so forth,
and then you feed it all into a model, and then we learn something about who's more likely to win.
But one of the things that happens in a lot of other domains, and here I'm thinking about, like,
chess or go, for example, is that you can have these models that are extraordinary,
but they don't speak English or they don't speak any human language.
And so the transmission of like what was learned from these events is something usable by, say, a coach who's thinking about strategy or a general manager who's thinking about player selection.
Talk to us about like when you think about these machine learning models that can't really communicate their findings in any way in a language that humans can understand, how you sort of bridge that gap to where this is useful information for a team or a manager.
This is generally be understood, I think, in the last couple of years or the last eight years, as what the biggest problem.
You can build these models, these neural nets already exist.
The main thing is the translation, like you said, from model outputs to coach.
And I think in the last couple of years, most things have found is that they need an expert analyst and data analysts to do this conversion or this translation step where the coach isn't fed the data directly.
Sure.
The coach is fed just the information that the analyst finds from the data.
And that could be, in the end, that generally still boils down to having video clips.
So you can use the data to find video and then you can show the video to the coach,
which then from bottom up you have this approach.
That makes sense.
But the part about, okay, we're going to use the data to find video.
So all of this makes sense.
You have the data specialist who translates.
You have the video so that there's something tangible.
But talk to us about that specific.
step where the data analyst sees some sort of model output and then is able to use that to find
the relevant clip to show the coach something potentially instructive because that seems like the
hard part to me. Yeah, so the model output could be many things. It could be outputs from a
classification model that says in this given 10 second or 15 or 20 second window, this team played in
this sort of buildup or they had this type of structure. It could also be a little bit more advanced where you can
simply say, well, in this instance, we had a high probability of conceding a goal,
and that could be just from an expected goal shot if you're looking at only event data,
but you could also have model outputs from something we call an expected possession value
model or an expected tread model where you measure the chance that the team is going to
score in, let's say, the next 30 seconds or the next possession, and then you can find this
fights there and either measure when your team is likely to concede or measure when the other team
is likely to score. And you can use those kind of signals to boil it down to video.
Yeah, this is something I wanted to ask, actually. So you mentioned speed just then.
Like, what is the actual latency that we're talking about since we're using all these market
analogies? Like, are we talking about an insight that's actionable within seconds? Like,
you're going to sub a player on after your model spit something out, like live during the game?
Or is it more realistically that you're reviewing the model and the apps?
analytics after a game and sort of tweaking, I guess, when things have calmed down.
Apparently, most of this does not have been live. So most of it happens either pre-match or post-match,
but you can still do it. There's enough live data to make these instances worthwhile.
Yeah, I'd say most of those types of adjustments based on data typically happen at halftime,
or if you're in the World Cup during a higher duration break. And in the derivatives world, I always
think of expected outcome versus realized outcome.
So what is the market implying?
What do you expect?
And then what is actually happening?
And that is sort of the nexus of how teams prepare for matches.
They come up with their own expected outcome.
How is our opposition going to play?
How are we going to play?
And then that live in-game is your realized outcome.
And so you're receiving that data and it's being transmitted to analysts who can then communicate it down to the bench to discuss with the manager who can then make those changes at that half time.
It's interesting the game of two halves, Joe, is now the game of four quarters, offering up more opportunities to make model-based adjustments.
That's right.
We used to have two discrete events in a game, the first half of the second, and now we have at least four.
So I guess that, yeah, that creates more data as well as maybe more ad revenue.
Adjustment opportunities.
As well as more ad revenue.
You know, obviously in baseball, at least according to Michael Lewis, right, that there was this period where, you know, you had the old time scouts and they're like, oh, this guy has good hustle, right?
This guy has a good heart.
There was like no data behind any of it.
He may just sort of had the, you know, the swagger of someone who looked.
like maybe a star player and then it's like, oh, no, but look at his like on base percentage
or that's or his VARP or whatever and then they get promoted. But it was like a culture thing.
Has there been a similar cultural clash within sort of soccer scouting where what the data says
about what constitutes a player does not map to traditional intuitions?
I think a lot of clubs look at it from from two perspectives. I don't think there's this old
school scouts versus the data guy mentality. I think it's a very collaborative. I think what a lot,
you know, what a lot of clubs do is they have an individual scout who will go watch a player
and give his assessment and his rating. And then they will have their own internal model with data.
And they will look at the deltas between those two different models. And if something seems off,
you often have collaboration between the data person and the scout,
and they figure out who's right, who's wrong.
And then to your point about hustle, let's call it,
there are Swiss Army knife ways to quantify that in soccer, in certain instances.
If you look at game state, say the game state, you know,
game state is essentially zero.
plus one minus one, plus two, minus two, plus three, minus three.
And so say you're in a plus three minus three game state and win probability for one team is
98%.
You can manipulate the data and only look at how players are performing in that game state.
You're out of the game.
Are you still competing?
Do you still care?
Are you in the right position?
Now, that may not align with the old school schedule.
saying, you know, this guy has grit. But again, there's Swiss Army knife ways to give some
type of indication. You know, I'd like to say soccer data creates questions. It doesn't give us
answers. Wait, just to better understand this, can you give us like an analytics framework?
If you were trying to judge the best, this is the loaded question, but it comes up on every
discussion, but you're trying to judge the best soccer player either of all time or currently.
Like, what would the analytics framework for that actually look like?
Jude Bellingham?
Okay.
Let me say more.
Yeah, yeah.
Like, really, seriously, this is a great question.
Like, walk us through what the math says about Jude Bellingham and how you would derive that.
Well, Jude Bellingham can play four or five different positions, right?
Most players, they have the number nine and their role is number nine.
But if you think of the game as this book with different chapters within the story,
Jew Bellingham can perform whatever task he needs to perform at every single chapter throughout the book.
And he does it at the highest level. He could play any position on the pitch besides goalkeeper.
And just sorry, like, how is this established? Like someone could say, oh, this guy, they say this in baseball too. He's a good all-around player, whatever. We can see.
But like, what is the data that actually establishes the Jude Bellingham?
can play at high levels in a wide range of position.
How do you derive that conclusion quantitatively?
So from a data lens, we think of it in two fronts, in possession and out of possession.
Okay.
So in possession is how you're progressing the ball into threatening areas and obviously
creating high probability opportunities.
Then out of possession is how are you preventing a team from moving the ball into
threatening areas?
since it's a very tricky question in sports because you're trying to quantify the value of an event that does not happen.
So let's say, Joe, you have the ball out on the wing.
Tracy, you're right in front of the box.
Jude Bellingham.
He would be both of us.
He is always moving into that passing lane.
He is right in between you guys at the right moment.
Jude guy.
So annoying.
And we have the ability to quantify the value of those movements and the closure of these.
lanes as players move out on the pitch and then you'll obviously get to the next phase where it's
how do they do this with their feet.
You've heard the chaos. Now you can see it.
Hello.
My gosh.
Watch all your favorite podcasts from start to finish right inside the free IHare radio app.
Gets the blood going up.
Catch every laugh and eye roll on shows like the Tom Green Farmcast.
Park the bus. Now with full video.
We're building a ramp for Tony Hawk.
It's the same hosts and the same chemistry with all the hilarious moments you've been missing right on your screen.
Cheers.
Open the free I-HIR radio app.
Search video podcasts and tap watch.
The Big Take podcast from Bloomberg News keeps you on top of the biggest stories of the day.
My fellow Americans, this is Liberation Day.
Stories that move markets.
Chair Powell opened the door to this first interest rate cut.
impact politics, change businesses.
This is a really stunning development for the AI world
and how you think about your bottom line.
Listen to the big take from Bloomberg News every weekday afternoon
on the IHeart radio app, Apple Podcasts,
or wherever you get your podcasts.
We started talking about this sort of translation
from what the model says to a coach or something like that,
and that still seems tricky.
For better, someone who's betting on sports,
that might be totally irrelevant.
They're just like, the model says this, this team is better.
The line doesn't match up with this.
Therefore, we're going to bet on this team.
I don't know why the model says, my model says this team is better, but it does.
So the translation is unnecessary if you're just betting.
Like, are we at the stage where we get in close to the stage where you could, for example,
feed a model the first five minutes of a game, scoreless zero zero, and to all the players,
like, oh, this looks like a competitive match.
It's going to be good, both sides.
But models are able to detect something.
that we can't articulate that says, oh, no, this team is playing.
And even though it looks like a tie game and a competitive one,
actually for reasons that we can't put into English, put into language,
this team looks like they're going to win the game.
They have a 70% chance.
Is that a thing or is that a phenomenon or is that a realistic thing to expect?
There are in-game with probability models,
which start with just a pre-game team strength.
So both teams have some value, perhaps you can think of an.
ELO rating.
Okay.
And they boiled down to a win, a draw, and a loss percentage.
And then those percentages can in-game be updated, but they won't swing all that
much because you have this prior information.
So maybe after the first minute, one team has some XG, or maybe after the first
five minutes, let's say.
Some team has created some high probability chances and that team might be the underdog team.
This was highly unexpected, I guess, since the ELO rating said that the other team would be
have the overhand. So you can slightly update your beliefs there. I don't think it's like changing the
needle or moving the needle all that much, given that it's just five minutes of information. But you can
definitely update your beliefs throughout the game given chances created, expected possession value,
as we just talked about, or momentum or something in between depending on the data you have
available to you. Mike, I wanted to ask you, given your involvement with Austin FC, we know that there are
obviously differences between major league soccer and European leagues. And some of those are, you know,
things like salary caps, designated players, roster rules.
Lack of relegation.
Lack of, yeah, that's a big one. Does that actually make soccer analytics and MLS like a little
bit cleaner in some ways in the sense that like in the European leagues, the money is, again,
trying to be diplomatic here, but it's very free flowing. There's a bit of rule stretching.
going on at times when it comes to salary restrictions and things like that.
Like compare MLS analytics versus European analytics for us.
So in European analytics, you have essentially your recruitment and then first team analysis.
In MLS, you have recruitment, first team analysis, but you have this third vector that I call
portfolio management, right?
This cap structure in the MLS is put in place with the intention of creating parity.
And as you mentioned, it's highly complicated.
The best way for you to understand it would be like me saying, Tracy, I'm going to give you
$10 million.
You can spend $2 million on NVIDIA, $7 million on Walmart, and $1 million on a speculative
biotech stock.
And so you have to think of each player from a relative value perspective based on where they slot in your cap structure.
Has that worked?
I mean, with MLS.
So it's like, I understand you.
This new league that's, you know, you don't want some really rich team to win all the time.
And then, I don't know, only Miami wins.
And then fans lose interest in the rest of the country.
I'm just, I don't know what the actual.
I think my understanding is Austin in particular is a very big.
fan base. But has that worked in practice to sort of create an equal level of like fan affinity
that's geographically distributed? Well, at Austin, we're only nine months into our project here.
And we just had some turnover with our sporting department. And so I would say that we have yet
to integrate and prove this portfolio management theory and the impact on success and points in the
table. But as a market participant, when I read these rules, my brain immediately goes to
the markets and portfolio allocation. Okay, so we have to watch Austin FC as a test case for the
portfolio management thesis in football. But wait, just to be clear, so this portfolio management
thesis, which you obviously have a strong intuition for having traded volatility for a long time,
this is the framework that you're bringing to your work in Austin.
Yeah, correct.
Every player has a designated slot.
So a player that could come in as a DP could be terrible.
But if he's going to be a designated player.
Oh, yeah.
Okay.
But that same player, if he is going to be a senior minimum salary player,
could be in the top percentile of talent for that specific slot.
And then within the league itself, you have different roster construction strategies.
You can either have three designated players or you could opt for two designated players
and for U-22 players.
And the way the league works is they have a salary cap, but it's a salary cap, but it's
It's more of a salary cap charge.
So whilst Leonel Messi may be making over $20 million in salary, his salary cap charge as a designated player is going to be much lower than that.
It could be $750,000.
Wait, I don't understand that.
Yeah, it's confusing.
There's a salary cap charge based on these roster designations.
So every slot has a dollar charge associated with it.
So you have to work within that framework of what is the charge for each player and your finite amount of capital you're allowed to spend.
So the designated players are the ones you're allowed to pay more.
Got it.
Because this is the legacy of like L.A. Galaxy getting David Beckham and wanting to pay him.
Right.
Lots and lots and lots and lots of money.
This is helpful.
Okay.
So one of the criticisms of modern football, I guess, is that it's dominated by.
the wealthiest clubs.
Like whoever has the most money
can buy the best players, certainly in Europe.
And so the big just kind of get bigger.
And I could certainly see an argument
where if analytics becomes more important
to actually playing the game,
then whoever has the most resources,
the most compute,
the most engineers, I guess nowadays,
is going to have an edge here.
But on the other hand,
we have had technology before
that has a sort of democratizing effect.
That's the money ball story.
Yeah, we have open source models, all of that.
And I think there happened some instances of the smaller clubs, actually,
like using this technology to perform better.
But which way are we going to go?
Is it the big get bigger or maybe we see some smaller clubs level the playing field here?
I hope it's the latter.
Yeah, same.
So you mentioned open source models.
I actually build open source software to help these smaller clubs.
I guess it's also helping the bigger clubs.
But it allows them to build these craft neural.
nets and these expected possession value models, and then also load these tracking data,
this tracking data, which has been a big task just in general because of the data structure
and the data size and the software that I work on helps clubs, any club basically get started.
So I hope by doing that it will help the smaller clubs, help data providers, I guess,
provide better insights to level the playing fields.
But I'm not sure, not sure that it is working out necessarily.
Because getting the knowledge, like we talked about before, actually distilling the knowledge from the data is still, I think, one of the bottlenecks.
Oh, yeah.
So this reminds me, I wanted to ask as well, like, who actually owns the data here?
Where does the data come from?
It depends.
And I think some of it is a gray area.
There are a lot of data providers.
Some have licenses.
Some don't.
Yeah, it's a big question mark in some instances.
I'll tell you a story about when I first started building models for AFC Bournemouth back in the Premier League in 2016, I had no idea where I could find the data.
And so I went on Upwork and I posted an ad and I just said, I need a developer who has worked with European football betters and a gentleman named.
Dmitri in Ukraine replied to me and he said, yeah, I've worked with many professional betters.
And I said, give me all the data that you could find. And so it was very skunk works, but Dmitri was able to
find a way for me to get access to all of the on-ball data across the world at the time to start
building my models. What is the data? So it's like, okay, you find some data provider or
Dmitri and Ukraine collects data.
Is this numerical data?
Is this a series of frames?
Like, when you say, okay, you need to go out and get the data, what form are you getting
it?
What does that mean?
Like, what does it look like?
So it's varied throughout the years, but the most prevalent generic data out there
is on ball data.
Okay.
And on ball data, if you're just to put on your Excel hat, you think about
what it looks like in an Excel spreadsheet, has every single event tagged. So you have, it just goes down
a series of events, pass, pass, pass, pass, dribble, shot, goal, pass, pass, dribble, tackle. It has the
event. It has the player or players involved. It has the X, Y, coordinate on a pitch, and it has the time.
And so when you have these raw data points, you can conditionally put together a mosaic of what is happening on the pitch, because you can measure the events, the speed, and the location.
Tracking data is a bit more nuanced.
I'll let Gioris touch on that.
So you can still imagine an Excel spreadsheet, but I don't think your Excel spreadsheet would like.
it very much if you try to load in this data. It's basically a player identifier, a team identifier,
and then X and Y coordinates for all players at 10 or 25 frames per second. So that means, I guess,
25 rows for a single frame. We never really touch, let's say, the raw data in the sense that
we don't get the pictures of the game and then try to figure out ourselves where the coordinates are
or what the coordinates are. So there are a lot of data providers out there that either
put cameras in the stadium or use the broadcast footage to extract this data.
So they will have different models.
One of the models would be first identify where all the pitch markings are.
So they have a way to understand like where all the players are relative to the pitch markings.
Then they know where all the players are and you can convert all of that into coordinates.
They know where the goals are obviously.
And then you'll get that in a single file for a single game.
Some providers might give you one file per minute.
they might give you one file per half.
And then that's just the raw data with the identifiers in the quarter.
You might get an additional file that has all the metadata that says this identifier belongs to this player,
this identifier belongs to this team.
If you're lucky in your tracking data, you get an identifier that says which team is actually
on the ball because that's highly relevant, but sometimes it's not included.
And you have to figure it out yourself by calculating the distance to the ball for each player.
And then obviously you have skeletal data, which is, I guess, 27 times more dense.
or maybe even more because you have 27 coordinates,
one for each body pose points or body points.
So you might have one for your left shoulder,
for the tip of your nose, for your left ear, for your right ear,
for your right foot, for your ankle.
And that gets into the millions and millions of data points.
So when you talked about in the introduction about these petabytes of data,
I assume it's going to be mostly skeletal data
because that data is incredibly rich.
How do players actually feel about all of this?
Because, you know, if people were worried,
watching me do my job and monitoring like my neck movement. They are. That's the thing they are,
Tracy. All right. But no one's, no one's modeling like what I'm doing with my hands or my feet at all
hours of this particular recording. Like I would have mixed feelings about it, right? Like,
do you have any color on how players actually feel about, I guess, the rise of statistical
analysis in football? I think they find it useful. I think something that yours has worked on this
specifically is eyesight.
What can they see and what they cannot see?
And so when you look at the data and you think about opportunity cost decision making
and you make a suboptimal decision,
when you go and talk to the player,
they might just simply tell you I couldn't see it.
And then you move on to the next.
So it's very complex.
I find most players to embrace the data.
There's curiosity around it.
but, you know, they likewise know the limitations.
You know, we started just talking about, oh, soccer is fluid.
It's a beautiful game.
It's an art.
I totally agree.
But, like, people actually did say this about chess 30 or 40 years ago.
And there was actually some people who held out to believe computers will never be able to
be humans in chess because of this an art.
And it almost seems hilarious that, like, this view held on for as long as it did.
But it turns out that no chess is just a calculation problem.
And when you have enough data and compute, you can.
solve the game, kind of. Is soccer in the end, like, is it just a series of lots and lots of
discrete events that our eyes are not capable of? But at the end of the day, with sufficient
compute and data collection, is it just like chess and that it's just a lot of
micro-binary decisions that can all be summed up? And is this therefore how life is?
When you work with this data as long as I have and with this tracking data specifically,
in the beginning, it seems very overwhelming because it's just an infinite stream of coordinates.
And so I struggled with that a little bit in the beginning.
I was wondering what I should do with this.
How are you going to model any of this?
But I totally agree with you.
You can disparate this.
So you can district it into, let's say, very minor events where we use the on-ball event data.
And if we align that with the positional tracking data, we might know every single
moment where a player may surpass, and then you might know the next moment where a player makes a
reception. So that could be discretized into one event that might last two and a half seconds or three
seconds. The next step would then be the player received the ball, they make an onball action
that also lasts two and a half seconds. That would then be your next discretized event. And then
you have all these small events, which are micro-movements by players, counter-movements by defenders,
and if you go these, let's say, sequences one at a time,
you can still make aggregated metrics from this,
but using all the tracking data that you have at your disposal,
but it's actually still understandable for you.
So you might say, well, this player made this many dribbles,
and it gained the team this much in terms of added value to score on a goal.
And you can do the reverse for defenders,
where you can say, well, this defender was always close to the ball,
so he was helping not consider.
concede a goal. And if you go back to the analysis part where you want your video analysts to look
at this, they also do this, just discreptized. Discretizing it like this will help significantly.
So I have one more question, which is it is obviously World Cup season, which means it is also
a cell side analyst publishing World Cup prediction notes and research season. And in my experience,
they tend not to be very good.
Like, often they will publish that England or the USA are going to win whatever World Cup.
Does the numerous strategists predict Japan?
Not to my knowledge.
But, you know, they often get it wrong.
If we think about statistical modeling, data, I mean, the sell side firms, they should be pretty good at this.
And yet, do you have a take on why they seem to struggle with soccer predictions every four years?
Honestly, for them, I have no idea what data they're using, whether they're using an ELO model, whether Nomura has on-ball event data, tracking data.
There's different ways that you can come up with these predictions.
But I think, you know, I like to stay in your lane.
And I think cell site analysts should stick to cell site analyzing.
Isn't there simply the age-old problem of if you have a good model, you wouldn't publish it?
You would just beat the bookies.
Right.
Yeah.
There's often criticism of people who publish things for a living.
Mike and yours, thank you so much for coming on Odd Lots, learned a lot there,
and I really appreciate your time and enjoy the rest of the World Cup.
Thank you.
Thank you.
Tracy, I have to admit, I find it a little depressing that probably most things in life
are probably just computation things.
You know what I'm saying?
It's like, I want to think that, like, there is something called art and beauty and intuition and something.
It's probably just computers.
all the way down, binary events that can be chunked and analyzed by microchips.
Oh, you know, the question I would have asked. Yeah. But we ran out of time was the idea of, like,
good arts law, which is like once you have a measure, like the measures. Yeah, because you see this
criticism of sports analytics, which is like players start focusing on their stats. Coaches are
focusing on the player stats. And then they get the players with the good stats. And then they just
focus on improving their stats more, but the stats don't necessarily translate into like wins all
the time. Yeah, you know, I was thinking something that Mike said in the beginning, which is that
like a particularly like an English Premier League team, there's multiple things they could be
optimizing for. So they could be optimizing for profit. Right. They could be optimizing for avoiding
relegation. They could be optimizing for avoiding relegation. They could be optimizing for wins.
Those are all distinct things. But you see this when a sport gets over-optimized.
It didn't come up.
But good example is like basketball when I was younger.
It was like the game was fun because there were lots of slam dunks.
And then everyone realized that three-point attempts were not being taken enough.
And so suddenly the game is dominated by three-pointers, which may be a better way to play.
But it's not necessarily a more fun fan experience than watching a dunk.
So you think like, okay, could the game get better formally but it becomes less entertaining?
There's a lot of criticism.
I mean, I think that's a possibility because people are already talking about convergence.
and like actual football play style.
Totally.
And you see this, like, you know,
there was a lot of criticism in the, like,
people were very critical of how Paraguay played, right?
Like, play to just survive into penalty kicks
than hope that the variance of the penalty kick period
allows you to be in France.
But it's like, no, that's like the game theory optimal
or just the game optimal play
if you're considered to be the weaker team.
So it does feel like there's all different things,
you know, again, going just to the core of the question,
stats like, what are you solving for?
Solving for winning is very different from solving for profit is very different from solving for
a gambler and there are different answers to each one.
Yeah, you win, but no one's paying for it or happy about it.
This is plausible.
Although for now, people are paying crazy amounts of money still to go see a soccer game.
All right.
Shall we leave it there?
Let's leave it there.
This has been another episode of the All Thoughts podcast.
I'm Tracy Allo.
You can follow me at Tracy Alloway.
And I'm Joe Wisenthall.
You can follow me at the stalwart.
Follow our producers, Carmen Rodriguez, at Carmen Armand, Dashel Bennett at Dashbot, Kale Brooks at Kail Brooks, and Kevin Lazzano at Kevin Lloyd Lazzano.
And for more Oddlots content, go to Bloomberg.com slash oddlots for the daily newsletter and all of our episodes.
And you can chat about all of these topics 24-7 in our Discord.
Discord.g.g. slash oddlots.
And if you enjoy Oddlots, if you like it when we talk about stochastic soccer modeling, then please leave us a positive review on your favorite podcast platform.
And remember, if you are a Bloomberg subscriber, you.
you can listen to all of our episodes, absolutely ad-free.
All you need to do is find the Bloomberg channel on Apple Podcasts
and follow the instructions there.
Thanks for listening.
The Big Take podcast from Bloomberg News
keeps you on top of the biggest stories of the day.
My fellow Americans, this is Liberation Day.
Stories that move markets.
Chair Powell opened the door to this first interest rate cut.
Impact politics, change businesses,
This is a really stunning development for the AI world and how you think about your bottom line.
Listen to the big take from Bloomberg News every weekday afternoon on the IHeart radio app, Apple Podcasts, or wherever you get your podcasts.
