Signals and Threads - Wrestling the world into rows with Eric Mannes

Episode Date: September 2, 2026

"Alternative data" is Wall Street's name for information that doesn't arrive from an exchange: satellite photos of parking lots, credit card panels, SEC filings. Eric Mannes has spent over a decade at... Jane Street, first as a commodities trader and now helping lead the firm's alternative data team. In this episode, Eric and Ron talk about what it takes to turn messy external data into datasets a trading strategy can rely on. Along the way, they cover the day oil futures settled at a negative price and the systems that broke as a result; the years when the commodities desk's risk system was one very large Excel spreadsheet; the hard question of what a company even is; and why better ML models raise the value of careful data engineering. You can find the transcript for this episode on our website. Some links to topics that came up in the discussion: Figgie Adverse selection 2020 Russia–Saudi Arabia oil price war Dual-listed company "LAMBDA: The ultimate Excel worksheet function" Trino The Bitter Lesson Learn more about Jane Street’s internship program.

Transcript
Discussion (0)
Starting point is 00:00:00 It's my pleasure to introduce Eric Maness. Eric has been at Jane Street for about a decade. And he's had a really interesting range of roles here, spanning trading and technology. And these days, thinking a lot about alt data, which is a lot about what we're going to talk about today. So thanks for joining me. Thanks for having me. So maybe to start with, I just like to hear a little bit more about how you got to Jane Street in the first place. I studied math in college, and I looked at what many of my peers did.
Starting point is 00:00:25 And a lot of them were going into academic research, perhaps. in Boston. A lot of them were going into tech, often in San Francisco, and some of them were going into finance, usually in New York City. And so I thought, like, well, I have three summers in undergrad, and I'll just spend one summer trying each of these. And at the end, I was like, Jane Street is by far the most interesting of these and the best fit for me. So I'll do that. So what about the trading experience struck you is so interesting? So I was a trading intern, and I was really struck by, like, how much I was learning. But I mean, I guess one of the interesting things about trading is, like, by and large, it's not a thing that anybody knows, right?
Starting point is 00:01:14 Like, we hire software engineers, and they've often done software engineering in other places. But, like, the vast majority of people who come here have no experience in the trading world. Right. And so there was nothing really outside that had, you know, given me. like practice trading or any depth in thinking about markets. And I came here, I was struck by like how much time people at Gene Street like spent teaching the interns, like about like the fundamentals of how markets work, about concepts like adverse selection, how to do research well. And people use the word anti-inductive environment of like when you discover a law, When you discover some effect in physics, it's not like that effect goes away.
Starting point is 00:02:03 In finance, like the things that you find, the patterns that you find often get discovered and competed away over time. And the things that you, like, knew stop being true or as important. And so you, like, have to keep discovering new things. Right. And then what remains looks ever more and more like a random walk. This is why actually our signals and noise ratios are kind of miserable because there's this whole system whose job is to like extract signal from the system and have it look more like a kind of pure representation of our remaining uncertainty, which is sort of by its nature looks more random. Right. And when you act in this system, other people like respond and change their behavior in ways that like weren't visible in your training data. Yeah, like kind of an amazing thing about like the world of trading is it's this interesting like.
Starting point is 00:02:59 compromise and incentive system that rewards people for learning things about the world, but does it in a way that essentially forces them to give some of that information up. So you have this kind of information discovery system where all of these people trading and competing with each other feeds information into the markets. And then the markets now are a kind of more reliable gauge for what things are worth, which tells you all sorts of things about what's happening in the world. I think that felt very real on commodities and in my current work on all data, where like we are looking at data from the real world
Starting point is 00:03:33 to understand it better and incorporate that information into market prices. So you said that like you were struck by the amount of effort that people put into teaching and I think teaching really is an important thing here. And actually I think one of the interesting like facts about Janstreet, which is maybe not obvious from the outside, is like the insane amount of effort that goes into the internship because we think like getting great people is incredibly important
Starting point is 00:03:54 and we've kind of put almost a pathological amount of work into making the internship a really great experience. But I'm kind of curious, like, how do you feel like that played out? Like, what were examples of the kinds of things that happened in the internship that you felt like taught you more about how to think about markets and kind of what the job was like? One important part of the internship was, like, mock trading or trading at low stakes. You can talk about good decision-making in the abstract.
Starting point is 00:04:26 No replacement for actually trying to do it live. Yeah. There's no replacement for actually trying to do it live. Doing it in a context where you kind of commit to something and there are consequences to what you choose, I think, provides, like, useful feedback. Right. I think this goes back to the fact that very early on it changed you, we felt like people, the early people here learned an enormous amount from like being on the trading floors. And in this kind of live trading environment, like it's kind of the nature of what trading is and how it works and the basics, the kind of basics of adverse selection and information flow and all of that. But like you can't actually send everyone back to the trading floors. Many of them don't exist anymore. So like what does it actually look like being in a mock trading session? Yeah. This has improved a lot over time. But like, one of the first of the first thing, but like, one of the first of the first thing. that we do is, like, we do have, like, a simulated stock exchange where interns can trade. And we'll have these, like, one or two hour, like, scenarios that kind of distill, like, some important trading concept or thing that happens in real life into a, like, time-boxed toy model of it. So, what's an example?
Starting point is 00:05:43 we have a trading card game, not trading card game, well, it is trading and it is a card game. Close enough. Called Figgy, where you know, you and three other players each have some private information about what the different cards in the game are worth. And you do trades based on that information. and you pay attention to what other people are doing, what information that they have that is probably causing them to behave in that way and think about how you should update what you are doing
Starting point is 00:06:25 or the trades that should happen when you and someone else have different information or reasons to value the same thing differently. It also gets people, like used to market-making language and like putting themselves out there and being willing to buy or sell spades for seven chips. And I think like getting people comfortable with making decisions under uncertainty in these like kind of controlled low stakes environments. I think is good practice for doing it in the real world without like the psychological pressure that would normally get him the way. If you read the book, Ender's game? I have. Right.
Starting point is 00:07:25 Classic. Yeah, like we can't actually do that in, uh, like have people do like mock trading and then slowly at some point and discover the mock trading is actually real trading. Right. I think for licensing legal supervisory reasons. It's problematic, but it is easier to teach intuitions about trading in this game-like environment where you can try things and see the outcome that isn't quite like the heat of battle. Right. And I guess like a game, like there is an adversarial component to it, right? It's not like the markets aren't a zero-sum game. They're like a positive some game. In the end, people all, you know, people all in benefit from participating in the markets,
Starting point is 00:08:10 But like small positive sum, right? And so if like player A is like winning a lot, then player B who they're trading with is losing a lot. And like that gets reflected in the kind of game structure as well, I imagine. Yeah. Identifying when you are doing a positive sum trade with someone versus situations where you and someone who is just as smart as you disagree about some fact of the world. Right. One of you is wrong. One of you is wrong.
Starting point is 00:08:38 and you should think hard about like why you think what you do. Right. And why you're not the one who's making the mistake. That's right. Okay. So you learned a lot through this internship and you decided actually this is the thing that you wanted to do. What was it like going from having been an intern to like actually being here as a trader? It was like even more interesting than I expected because like there were all of these like real life trading decisions and strategies that like.
Starting point is 00:09:08 we couldn't get into the details of for IP reasons. But I felt like I was learning a lot during the internship. And then once I got to Jane Street, we could talk about like the trades that we were actually doing as a firm and the decisions that we were making as a firm. And so I started off on the commodities desk learning to trade oil and natural gas ETFs and in index ETS, which are about 50% like oil and energy related products. And looking at the trades that we were doing and thinking about like, did they make sense? Were they good trades? Why were we doing them? If they were good trades, why didn't we do more of them? And if they were bad trades, why did we do so many? All of this happened, like, sitting next to, you know, a mentor who
Starting point is 00:10:12 was an experienced trader who had been here a long time. And, you know, a bunch of trades would happen. And then he would say, like, okay, well, what do you think we should do here? How should we change our behavior? I would think for a bit and say something. And he would say, well, okay, we're doing something else. And then work for a bit. through with me like, okay, well, you should also take into account like this, that, and the other thing. And over time, you know, through working through all of these examples and, like, getting this feedback, you learn to make better decisions in these contexts. How long do you feel like it took until you were able to make positive suggestions that, like,
Starting point is 00:10:58 actually would influence the trading that we do? I don't know. A couple months. I suppose, like, in the small. I think in the small, like, pretty quickly. Because there's just stuff that like other people don't have time to pay attention to. Right. You don't have to like better than everybody else to add value.
Starting point is 00:11:13 You can become, you know, Jane Street's expert on this one ETF and it's trading dynamics or like think really hard about something that no one else at the firm has yet and have some interesting suggestions. And over time, like the scope of what you are comfortable doing on your own increases. At first it's like, hey, I think we should make this small change. What do you think? And you'd get comfortable with doing that and eventually get more comfortable like making the bigger decisions on your own. But like reaching out and gathering more input when you weren't sure. You said that you could become like, Jean Street's expert on some particular corner. Can you give me like an example of like some little corner of the world that at some point you became like the local expert in?
Starting point is 00:12:03 Oh, man. We're going back a lot of years. I think at the time, our trading systems were pretty slow. And we were rolling out a new generation of trading systems that were considerably faster by like two orders of magnitude or something. But still slower than the frontier and slower than we are today. someone had to think about like how to optimize our use of these systems and analyze like where we were gaining from speed or what opportunities we were missing and I don't know who should do that. Eric's new. He is time. I think another thing is like I spent a lot of time. I spent a lot of time. I am thinking about, like, futures and their multipliers and the risk that you have when you trade a future and how that's different from, like, when you trade a stock.
Starting point is 00:13:19 Yeah, this seems like a dumb corner of the world, but also one that I've spent a weirdly large amount of time over the years thinking about, where it's like you get used to trading equities. And it's like, this is great. When I want to think about the risk of the trade, I can, you know, a starting. point is the cash flow. What's the cash flow of trade? It's just the product of like the price and the size. This is so nice. And then like turns out not so much in futures. Right. Maybe just what is a future? A future is like, yeah, in agreement in the future, hence the name, right? You exchange some money, right? Based on some other like benchmark that this thing is tied to. So like, you know, the S&P 500 futures are tied to like the closing value of the basket of S&P 500 stocks at a particular point in time, I think.
Starting point is 00:14:01 Or is it, no, I think it's actually the opening prices. Okay. On the third Friday of the month. Amazing. But on which exchange? Yeah. Yeah. And when you buy a like crude oil future, you are, you agree to pay some money for either
Starting point is 00:14:25 like the value of some, you know, index related to oil prices or, you know, a thousand barrels of oil delivered to you in Cushing, Oklahoma. And so-called physical delivery. Yeah, physical delivery. Or like, you know, live cattle where actually sometimes the cattle are delivered, um, dead. But, uh, oh, you know, tens of thousands of pounds of beef.
Starting point is 00:14:54 We, you know, are generally financial players who did not want. to take delivery of a thousand barrels of oil in Cushing, Oklahoma. I've never been there. And, like, I don't know where I would store, you know, the 20,000 pounds of beef. Like, when you buy a future, like, you put up some margin with the exchange. And, you know, if you, you. are long and the future goes up, the exchange credits some money to you. When the future goes down, the exchange removes some of your money and possibly asks you to put up some more margin
Starting point is 00:15:46 so that they know that you'll be able to pay off losing bets. Another thing you can trade is like the spread of two futures. You get long one and short another. So, like the CLU6, CLV6 spread. If you buy that on the CME, you get longer crude oil features deliverable in September and shorter crude oil deliverable in October. And if you trade like S&P 500 future spreads, you have the opposite sign convention.
Starting point is 00:16:23 But in either case, like your position, sorry, the thing you trade, like, decomposes into two different positions rather than like the one thing you've bought. Right. And I guess the point here is that all of this is like one level more abstract than like the underlying things are being traded. So like, you know, instead of trading the outright thing, you're trading this future on the thing. And then there's a bunch of conventions. And some of the conventions are about like it has a name that says what the two legs of the spread are. But like which one is positive and which one is negative depends on the future. Right. And you can't just like look at the
Starting point is 00:16:58 notation for sure and know the answer. And then where we started, there's other crazy thing called the multiplier, which actually tells you the real size of this. Right. Like when you pull up the ticker for some future, say crude oil, you'll see a price for like one barrel of crude oil or one pound of live cattle. But actually like when you buy the contract and take delivery, you get a thousand barrels. And the like notional about. is a thousand times of like the quoted price. And different sources will, depending on where you're getting your pricing from, you will get different prices and different multipliers and sometimes like different currencies.
Starting point is 00:17:47 Like someone will quote in dollars and another person will quote it in cents. It's a small matter of a factor of 100. Yeah. And as long as you always keep track of which prices correspond to which multipliers and never mix up the two, you won't accidentally trade 100 times less or much worse, 100 times more than you intended. Right. And so this matters kind of from a risk perspective. Like ideally we'd like to have relatively simple systems that are kind of like stepped back from the details of the trading strategy, which kind of put the kind of outside envelope on the risk that the system you're taking.
Starting point is 00:18:28 And that's, like, way easier if you actually know what's going on. But what you're pointing to is this all this whole, like, what we call a metadata problem, of, like, actually understanding the details of how these contracts actually work. And it's just made much more complicated by the fact that, like, the world is full of huge amounts of unnecessary complexity of, like, different conventions and different details of how these things are quoted in different contexts. Yeah. And like you have assumptions that are baked into your trading and risk systems.
Starting point is 00:19:01 Like this price will always be positive. That turn out to not be true. On April 20th, 2020, the price of the like almost expiring crude oil future on the CME settled at a negative price. And, you know, you'd always think like, well, Good. They're not. They're worth something. They're worth something. You should pay for them. You should pay a positive amount of money for them. And that breaks a lot of assumptions. But how can it happen? How can a crude oil future be worth negative? Like goods normally are good. Like what happens? Yeah. So that crude oil contract was for oil delivered to a particular place at a particular time. And, And at the time in Cushing, Oklahoma, storage was getting full. There wasn't much place to put the crude. Supply of oil everywhere was high.
Starting point is 00:20:06 And demand, for obvious reasons, was considerably lower than usual. If you remember. If you remember. April 2020, yeah. And that added up to like, well, if you. were long, what would you do with it? And this isn't quite true because the, like, for one day, the settlement price was a negative number. And the, like, final settlement price of this contract was positive. Like, in the end, people were willing to pay positive dollars for crude.
Starting point is 00:20:48 But on April 20th, there just were not enough people willing to step in and say, I will pay money for this thing. And when all of your systems had assumed this price was positive for a long time, yes. Maybe your trading systems do not have a great day. Right. Maybe the people who would normally provide to these moves were unable to send orders because, like, their order entry system didn't know what to do with a negative 100% return. also like if you were modeling other products based on the price of crude oil on that day, how are you modeling oil companies or everything else?
Starting point is 00:21:32 Should you be using like that negative crude future? I'm not sure. Right. I guess just like zooming out for a second, this like points to like a thing about trading, which is some of trading is thinking hard about the technology stack and the math and the modeling. and some of which is just diving into a bunch of grotty detail of like how does the world actually work? What are the mechanics of the markets and the systems and like the amount of storage space there is in Cushing, Oklahoma? And like all of those things can play in.
Starting point is 00:22:01 And this is like a general fact about trading is like lots of different kinds of things about the world can flow into the trading process. It's not like, like I think sometimes people who hear about James Street think it's all about like, you know, high tech, high performance, smart math models and stuff. And like that stuff is all real and part of it. But there's also a lot of detailed engagement with how the world actually works that also plays into being an effective trader. Right. And I think that's true across like all of our businesses, but it is very true in commodities where like the things that you are buying and selling are connected to the physical world. Oil and wheat and things that like people actually. physically produce and consume.
Starting point is 00:22:49 And occasionally deliver to you. And occasionally deliver to you. Yeah. So in addition to thinking about all of these kind of trading-focused concepts on the desk, you also spend a bunch of time thinking about technology on the desk. Can you say a little bit more about what that was like? Right. At the time, I think how Jane Street approached developing technology for trading
Starting point is 00:23:11 was having a bit of a phase shift. we went from having developers who primarily worked on like the core infrastructure that undergrated our trading the trading systems, the market data, bookings pipeline, that sort of thing, while like traders like people on our commodities desk and our domestic ETFs desk would plug in to those systems but do like less systems development themselves. And at that time, I guess we finally finally. felt that we had enough breathing room to say, like, well, what if we built commodity-specific systems to help address the unique problems with, like, trading commodities?
Starting point is 00:23:55 And this is, like, a thing that happened not, like, all at once, but, like, at different desks at different points of time, some desks were more technical and their needs, some were less. And so this kind of, like, desk-dev phenomenon, like, unrolled, like, you know, over a period of years as we kind of got enough capacity. And now, like, every desk across the firm has embedded dust devs because you just kind of need it everywhere. Right. But at the time, it felt like, well, this is working here. We should try this elsewhere, too.
Starting point is 00:24:21 And when we were starting to build out the commodities dust dev team, we needed someone from trading to help think about it. And I was relatively technical, having done my tech internship for one of my free summer. And having learned OCamill and O Camel Boot Camp, I got involved with that. And that involved things like thinking about how we like keep track of our risk or exposure to different commodities. If you're looking at your positions in energy products, like you could look at it at a very high level or you could look at just oil and related products. You could look at like West Texas intermediate crude oil alone.
Starting point is 00:25:03 Or you could look at like the, you know, November contract of this one commodity. and you need to be able to, like, shift from level to level all at once and be able to dig into exactly where your exposure to these was coming from. And we'd been doing this before. It was in large Microsoft Excel spreadsheet, U.S. Commodities Risk.XLSM. And you'd, like, press a button, and it would pull in, you know, data about our trades and positions and crunch some numbers. One, sometimes it was slow. Two, it was something that had built up over time and was built by people who, you know, were very good at using Microsoft Excel and very knowledgeable about commodities. But we're not, by and large, like, software engineers.
Starting point is 00:26:00 And kind of, like, building it up again from first principles and making sure that you got the foundation right. to something that we could really trust, let us build, you know, more and more ambitious things on top of that than we were able to build when it was an Excel spreadsheet that ran slowly and worked except occasionally when it didn't. Right. I feel like people who don't use Excel much probably underestimate how good it is. And then, like, when you've used Excel a lot, you have a really good sense of, like, oh, my God, what the limitations are.
Starting point is 00:26:35 Like, Excel in many ways is kind of great. It gives you great ways of visually inspecting and understanding the data. It's extremely flexible. It's actually quite easy to debug in important ways. But yeah, there are profound performance limitations. Like what else actually was bad about doing it in Excel? What were the downsides? Yeah.
Starting point is 00:26:53 I should say before all of this that I love Microsoft Excel and I don't use it very much anymore. But, you know, we like functional programming here. Excel formulas, spreadsheet formulas are, I think, the world's most widely used functional programming language. For sure. I think Excel has lambda's now. It does. And let bindings.
Starting point is 00:27:15 And let bindings, yeah. I think thanks to Simon Peyton Jones and some other people at Microsoft to help build it. Yeah. When you write, like, Python code, usually you have the, like, code in front of you, the logic in front of you, but the data is, like, hidden away until you, like, write something. that visualizes it for you. And with Excel, like, the data is front and center, and the logic is hidden in the back. And, like, that's pretty useful. People are making decisions about what to do based on the data that they see. So what were the problems here? Performance was
Starting point is 00:28:00 sometimes fine, sometimes at critical moments, quite bad. You aren't just like writing the cell formulas yourself. You're writing visual basic to construct the spreadsheet that you use. This is like cell metaprogramming. You're writing a program that writes a program that's like laid out in the sheet. Yeah, you press a button that like runs a nested, visual basic macro is that like initialize a sheet with all of our positions and trades and the code the code on top to analyze them. And visual basic is okay. No, I mean, it's like it was impressive what we were able to do with it. But having a type safe language where the compiler can catch a lot of your mistakes where you can write robust tests and interfaces
Starting point is 00:29:05 is an advantage that many programming languages have that is like hard to do an individual basic also like version control yeah hard to do version control yeah Excel is not optimized for that I once wrote an Excel grep tool that that yeah that yeah was just like a firm utility for searching through spreadsheets because like opening them up one at a time and in Excel and control effing first through the cell formulas and then through the visual basic is impractical. No way to do. Yeah. And like if you need to do like large migrations of many spreadsheets at once, which we sometimes did, it helped to be able to search them. maybe related to the lack of version control is that things just build up.
Starting point is 00:30:03 Someone, like, put some throw away cell formulas for analysis in, like, one corner of one sheet, and then it's just there forever. Yep. So you get to a point where, like, no one fully understands what it's doing, though some people know a lot, and are pretty sure. And it's hard to, like, rebuild that model of how the spreadsheet works
Starting point is 00:30:34 and, like, become, like, confident that there are no issues. So that itself, I feel like, makes the process of replacing it hard, right? Yeah. Like, there's this spreadsheet. It's an artifact. It has a lot of embedded domain knowledge that, like, actually no human knows all of anymore. And then, like, what was your role in trying to help bridge the gap? between like the spreadsheet that did a thing and then like the piece of software that we were going to build to replace it. Right.
Starting point is 00:31:03 So a lot of what I did was translating between the domain understanding that like the other traders and I had and a description of the domain and the problem that like a software engineer can work with and like build a robust system around. Rather than someone saying,
Starting point is 00:31:26 well, want this and it should do that and that and the other thing, like, distilling that down into like what we're fundamentally trying to do, which is, you know, model the risk exposures of our commodities desk and how we are trying to do it, the input sources and how they should fit together, what the basic atoms are. And we built up model of the thing we're building, the flow chart of how all this data, like, fit together and turned into, like, the end result. And then people were able to just confidently build some part of it, knowing that it would fit together into, like, the whole correctly. I imagine it's, like, much better for the software engineers. If they
Starting point is 00:32:19 don't just have like a spec, a narrow spec of a thing to build, but also they like understand the background of the domain in a way that's like a little more kind of actionable and fills out their understanding of the world. Yeah. And information flows the other way as well. I think the domain experts like who are not themselves, software engineers, like don't understand how good things can be if we build the right software for it. Like I have suffered for years like, I have suffered for years, doing things in this way because there weren't tools to help me do it better. And I've gotten so used to that I don't realize I need something better. Right. Like one of the characteristics of many Jane Street traders is high pain thresholds, right? People are like willing to deal with
Starting point is 00:33:09 like pretty messy things and pretty unpleasant things. If like, you know, I have to like, you know, knock my head on the table three times and slot my cheek in order to like cause money to come out of the machine. They'll just like do that all over and over. Yeah. And someone who makes a good desk dev at Jane Street can say like, no, there's no way you should be doing it that way. If we, you know, built this tool, it would be much easier to do these studies that are
Starting point is 00:33:41 either difficult to do now or that no one would bother doing now because we don't have the right language for working with them. Right. So part of what you need to convey to the traders is, like, a better cost model, like, understanding what's possible and understanding that, like, some improvements are, like, maybe way cheaper than they imagine and some are more expensive than they imagine. It's not just the costs. You can have, like, great ideas for what we as a, like, desk or as a company should be doing differently. It's like, I am someone with, like, this particular set of skills and knowledge and my job is to help Jane Street. you know, do the best trades.
Starting point is 00:34:21 All right. So you're no longer on a trading desk. Today you help lead our alt data effort. Can you see more about like, like what is alt data and what is it at Jane Street? Yeah. So out in the world on Wall Street, alternative data is like data that you use for the investment process that isn't normally used in the investment process. It's an unstable naming convention.
Starting point is 00:34:50 Right. Like, if enough people use it, does it, like, drop its alt, like, moniker? The canonical example of alternative data is, like, a satellite photo of a Walmart parking lot. You take a photo of the parking lot. You see how many cars are there. If there are a lot of cars, they have a lot of customers, and, you know, they're doing well. If parking lot's empty, that's a bad sign. And in contrast, there's like traditional data, orders and trades on exchanges, basic, like, information about, like, instruments, stocks.
Starting point is 00:35:29 Imagine, like, Warren Buffett, like reading an annual report, thinking about the balance sheet and income statement. That's traditional data. Okay. And at Jane Street, a lot of data is alternative. The satellites are alternative. The, like, company filings with the SEC are alternative. The things that are traditional are, like, the, like, real-time market data from exchanges, you know, the basic instrument metadata that we need in order to, like, run our, like, trading and risk systems and booking and so on. Right. Back to the multipliers.
Starting point is 00:36:12 Yeah. And it's based on, you know, the path-dependent evolution of Jane Streets trading, you know, starting as a market maker and then expanding into adjacent areas. Right. In fact, it's reflected in like team structure. We have a market data team that does that kind of like live data from exchanges and from other places. And we have a metadata team. And then we have this alt data team, which in some sense is like kind of picking up the rest of that. of that kind of value that you get from other sources of information. And then how did you get involved in this space? How did I get involved in this space? So at Jane Street, you can do whatever you want as long as it's the right thing to do. Your job is not to maintain this one system or to trade this one product.
Starting point is 00:37:06 Your job is to help Jane Street, like, maximize it. it's P&L over long term. What that means is that if you see a problem that you think is important to solve, what's stopping you from solving it? I mean, probably you should talk to other people and make sure that, like, this is a real problem and that it is worth fixing, and it is worth your time fixing.
Starting point is 00:37:41 And at some point, I was thinking about, like, a bunch of different data-related things at Jane Street and realized that, like, we were failing to find trades because our data was too hard to work with. Because, like, external data specifically was too difficult to work with. Like, our infrastructure for streaming in real time and surfacing. for like historical studies of like the real time market data was great and had lots of engineers and specialized infrastructure while the data infrastructure for all that other data was kind of weaker. It was annoying to just get data in the walls of Dean Street, both like from a contractual perspective. There was no all data team.
Starting point is 00:38:37 There was no all data team. So how did it happen? Like where was the work done? Who was doing the work? Like, what was the shape of all that? Yeah. I think when we started, it was like me and, you know, a small number of devs, including Jacob, who was on the podcast earlier.
Starting point is 00:38:56 Mm-hmm. And a very small number of data engineers. And we had three areas we were focusing on. But I think you're going ahead. I thought, like, before that, I was sort of thinking, like, Like the early version of this, I assume, was just like kind of crowdsourced by the death. Oh, sure, sure. Before all, like how, sure.
Starting point is 00:39:18 We were using external data before we had an all data team. And the work of building the like data pipelines and iterating that data would be done by whoever was around and available all the time. So someone on the commodities desk would say, like, we want this data set, and our data vendor management team would go out and buy that data set. And a dev or a trader or a trading desk operations engineer would, like, bring in the data and put it somewhere. Maybe in some corner of a large NFS tree, our shared file system, maybe in a database, but which one?
Starting point is 00:40:06 And like the result would be, I don't know, like 60% as good as it could be. And I guess this process was happening like some on the commodities desk. Yeah, some on the commodities desk. And then whatever, like very different desks, reinventing the wheel, doing slightly different versions of it. By people whose job was not primarily like thinking about how to build data pipelines well, who were not thinking about what could be shared between those processes. who sometimes, but not always thinking about like, how should people around the firm,
Starting point is 00:40:42 not just like in the use case that I'm thinking about right now, be using this data. And thesis was that we could have a centralized team, try to do that. Think about a centralized team to source some of these weirder alternative data sets to build better tools for building these data pipelines. and to maintain some of these data pipelines, especially like ones that kind of had for my impact
Starting point is 00:41:13 that would be used in a lot of different places. So how did you go from like the recognition of like, man, this is kind of a mess. Like people are generating datasets in messy ways, don't have great ways of distributing them, don't have good technology for serving the data, don't have good ways of cleaning the data, don't have good shared practices.
Starting point is 00:41:30 And so like we have less good data that is like just harder to convert into valuable traits. and is therefore worth less money. So, like, that seems like a key observation. How did you go from that observation to, like, actually doing something about it? So, like, sure, you've identified a problem, and many people at the firm would agree that it was bad in all these ways.
Starting point is 00:41:53 But, like, why? There are lots of things we can make better. Like, why this one? Part of it was, like, we'd experiment. we'd do something on a small scale and the results would work well, so we'd do it more. We'd, like, find one particularly basic but annoyingly difficult to use dataset and make it easier to use, like, think very hard about it, and people would start using that new thing. And lo and behold, like, there were effects that they were.
Starting point is 00:42:33 were not able to find when the data was really difficult to work with, that they were suddenly able to find when the data was easy to work with. And when someone had thought really hard about those edge cases and like put in the work of understanding like the data model and so on. So basically incrementally go do some things, make it better, demonstrate that it's worth money. You could see the P&L coming out the other side. And after a while, it just kind of becomes clear to everyone that, like, oh, we could use, like, a lot more of this.
Starting point is 00:43:08 You could use a lot more of this. And so, like, we hired our first data engineer at Jane Street ever in 2023. And the feedback has been like, oh, we could really use a lot more data engineering help, actually. the more people work with data engineers and see like the results from good data engineering work, the more they are excited to like apply that toolkit to other problems. I guess by informing people about how much easier it is to like build good quality data sets than it was before and be like results that were. able to get in doing so they want to do it more. So induced to demand.
Starting point is 00:44:02 Induced demand. Yeah. So what, just to like step back. So like we went in 2020, from like having our first data engineer, like how many do we have now? Let's say, 20 something.
Starting point is 00:44:15 Got it. So like a much bigger effort. Like as it's grown out, like, I'm kind of curious like what's the structure of the team, right? You mentioned data engineers, right? But I don't think all data is. just data engineer. So like how does the team break down into different kinds of people with different kinds of expertise? One group is data strategy and that's not engineering at all. But their
Starting point is 00:44:37 job is to understand what data Jane Street needs and what data is out there in the world. And then if there's a match, like, you know, get the data for us so that we can do good data engineering to it. And so internally, that looks like talking to our different trading desks and understanding like what problems they are trying to solve and what problems they could solve with better data. The external part involves getting out there and showing up and... There's like another case where you have to understand like grotty details of the real world. Yeah. We go to conferences where we'll like talk to data vendors. So there are these like data catalog companies, one example of which is new data. And the data catalog companies,
Starting point is 00:45:33 like, they have analysts who learn about the different kinds of data that are being offered to the financial industry and, like, document them somewhere. And data consumers like us, you know, pay for a subscription to that catalog. And they'll also, like, organize these conferences where data buyers show up, data sellers show up. They'll have like speed dating between the buyers and sellers where you'll like read all these profiles
Starting point is 00:46:08 and you'll say like, I want to talk to these ones. And then they match you and you have like 10, 15 minute dates in a row where they explain their product to you and you ask the questions that you have and try to figure out whether like this is some data that would make sense.
Starting point is 00:46:32 Do you get through there with like a lot of chardonnay as you're? Yeah. Usually it's drinks after. And, you know, you go there and some of them, you say, well, this was interesting, but I don't think there's a match here. Other others you say, like, we should talk more. can I get, you know, your number, your business group.
Starting point is 00:46:56 So that's one way of meeting. So this is like a weird two-sided market, right? There's like buyers and sellers. Like, are the, who are the, who are they? Like, first of all, who are the buyers? Like, it's obviously people like us, trading firms? Is it all trading firms who are the buyers? Yeah.
Starting point is 00:47:11 So I think there are a lot of trading firms that don't look like us. There are like hedge funds, whether like big multistrats or small asset managers, banks, participants who want data for their like investment decisions and analysis or research. Some of them are expanding into like other
Starting point is 00:47:41 kinds of buyers, like private equity or whatever. The sellers vary. There are satellite providers, say, who have satellite data. So you can see the Walmart parking lot. So you can see the Walmart parking lot. So there are like big companies that own lots of data sets over lots of verticals,
Starting point is 00:48:04 the like Bloomberg, S&P, LSEG refinatives of the world. There are companies that like focus on a specific kind of data, consumer transaction data. there are companies that happen to have data from their business. Either like, you know, they're a satellite company and they have satellites, or they're like a company that like provides goods or services and they have exhaust data and they think, hmm, like maybe someone would want to buy this. Do you actually have a sense of like how much of this? big market for data is, has trading of various forms or financial companies for various
Starting point is 00:48:56 forms on the buying side and how much is other stuff? Some of this data is probably also used for like targeting ads. Sure. And like one important difference between the trading use case and the ad use case is that we don't want to know like who the individual consumer. are. Right. We want to know, like, oh, like, people are spending more at Lou Lemon. People are, you know, buying more Ritz crackers. But, like, part from, you know, demographic information that might help us weight the data, or maybe information that would let us see, like, oh, people who buy this also buy that, we don't want to know.
Starting point is 00:49:43 Like who the actual individual is. Well, with advertising, I think it's the point. right for so for if you're doing this kind of like micro targeting yeah if you're doing the micro targeting okay so let's take a step back so we just said like one part of this is this data strategy which is kind of thinking about the sources and the six understanding places where data can come from and places where data can be used internally and doing this like kind of internal and external kind of matchmaking process yeah like some of it is through that uh through catalogs data sets that are um advertised and exist and some of it is from thinking, you know, from first principles, like, what question am I trying to answer and what data could help me answer that question and who might have that data, whether they
Starting point is 00:50:27 are selling it right now or not? Mm-hmm. Sometimes I guess there's, like, someone who you think might have data who isn't selling it, but, like, you can go talk to them and maybe there's a product they can make. Yes. In addition to, like, data sourcing, like, we have our, you know, data infrastructure via, like, tools for ingesting and transforming data for like monitoring or data pipelines, the building blocks that people who are building data pipelines actually use.
Starting point is 00:50:59 Got it. And this is what looks maybe more like a kind of straight ahead software engineering role? Yeah. That is more like software engineering. Your product is like software and systems and services. And the people who do. that are largely people we hired through our, you know, standard software engineering pipelines. And then we have data engineering where, like, they're also engineers. But the product is, like,
Starting point is 00:51:29 a good data set. What they do is, like, they build data pipelines. They bring data into chain streets walls and they transform it into the, like, form that we need in order to, like, use it in research and trading. But they also have to learn a lot when doing so. They have to understand like how does this data set work? How does this data domain work? Like what is this data actually modeling? They have to like understand how we're going to use the data, understand how the data is produced, and synthesizing all of that, like figure out like, okay, what should Jane Street be doing with this data? So you said a lot of what they need to understand,
Starting point is 00:52:19 but I'd love to hear a little more about what they actually do, right? Part of what you said is writing a pipeline, and a pipeline itself is just a kind of program, right? You've written and transforms and regularizes the data into the shape that you need. But like, to what degree is that just like a mechanical matter of like parsing and understanding, you know, what is the shape in which the bits have been laid out?
Starting point is 00:52:41 And to what degree do you need stuff that's like more, kind of in the domain and like, I know, what's the shape of that work that you need to do that involves understanding more about the data? Some of it is mechanical, but like we build tools to automate as much of that as possible. Part of it is thinking about what does this field actually mean? Maybe like we could go into an example. Like what's an example of like some data source that has some like stuff that is on its base hard to interpret and then you go out and like understand more about the domain and can produce an easier to use version of that data?
Starting point is 00:53:16 So company financials are like not alternative at all. They're the opposite of alternative. Like, you know, in the U.S. at the moment, companies, public companies publish their financial statements every quarter. And, you know, it's published via like filings with the SEC. maybe a snippet is shared via press release at around the same time. And you have all of this data about, you know, from like lines from a company's balance sheet or income statement.
Starting point is 00:53:53 And you want to turn that into some like structured form that we can do like quantitative studies on top of. So that's like when you say structure versus some structure, there's like like a PDF or a big pile of text and you want to convert that into like morally like rows in the table. Yeah. And you know, there are companies that will do this, that sell like tabular versions of company financials. Let's start off with what even is a company? How do you know like what a company is? How do you know which financial statements are about the same company or about the same like financial period of a company. Some companies like have these like
Starting point is 00:54:41 weird structures, I think they're called like dual listed companies where there's company A and there's company B and neither one is a parent company or the other but like they have the same management, the same board often. And there is some, I don't know, know, inter-company agreement that they have such that, like, owning them is economically equivalent. Okay. And, like... So you have two things that are, like, formally different companies? Formally different companies.
Starting point is 00:55:13 And then you practice economically, like, are the same thing. Yeah. And so you think about educases like that. And how do you, like, you receive data about something? Is it one company? Is it both companies together? What is it exactly? And because you're trading on this.
Starting point is 00:55:32 you associate it with like securities. Does it apply to all the securities of like both companies? Right. And this just kind of goes back into the metadata problem. There are all sorts of weird company structures and different names for things and different naming conventions. And sometimes you talk about the Ison, which is like the name of the thing internationally. And sometimes there's the C-Dol, which differs depending on which clearing system you're in.
Starting point is 00:55:59 And which, like, does the data apply to the company or, like, a share class or a specific Icen or a specific listing? Like, there's like a million ways of expressing the names of things. And there's a bunch of, like, complex structural stuff that, like, is kind of real, but also maybe you want to collapse. Yeah. And then companies publish multiple versions of their financial statements. they'll put out one and then they might like revise it or they'll like change their, you know, fiscal calendar or accounting in ways that make like a not completely comparable to be. Throughout this, you have to think through like what data was knowable at what point in time
Starting point is 00:56:50 and like what features do I actually need for our trading. Are there other like the, I mean the kind of metadata one is one that's kind of come up a lot in this conversation. Are there, like, good non-meditated examples where you have to, like, you know, you get some data source. It has a bunch of numbers in it, but like, what actually do these numbers mean? And, like, how do you go from, like, oh, yeah, they told you some stuff to, like, understanding, like, where there might be errors in the numbers or where, like, you know, it might, you know, mean different things for different companies or different records or different examples or whatever. Like, what's, it's like, just to get a sense of, like, how the kind of domain specific stuff. shows up, like, beyond the kind of traditional metadata problems and into, like, the wilder world of the different kinds of data you encounter. So, like, we talked about the satellite photos of Walmart parking lots. And... And, but that's, like, an old...
Starting point is 00:57:45 It's, like, such an old example. Like, people used to talk about, when people talk about, like, the classic example of, like, does it count as insider trading if, like, you're flying in a plane and you look down and you see, like, you know, I guess people, you know, that you see, like, you know, something like this. Like, does that, does that one, is that actually a good one? I don't think so. I might get, like, a bunch of LinkedIn messages from vendors after this saying,
Starting point is 00:58:08 you're totally wrong. I have these great photos of parking lots that I'd like to sell you. One problem is that, I don't know, it's kind of annoying. You have to get a lot of photos of a lot of parking lots. It's noisy. Maybe you get not that many photos. per store per, you know, day or week. Walmart has 20% of its shopping online.
Starting point is 00:58:35 Other companies have more. And parking lots are not going to help you when, you know, measuring the behaviors of New York City shoppers. And it's also just like really indirect. Like the thing you're getting at is like how many people showed up, which is different from how much they're spending.
Starting point is 00:58:55 And like, you kind of think from first principles, like, who knows what people are spending at Walmart stores? Well, Walmart does, but they're not going to tell you until the end of the quarter. Then the consumers do. So, I don't know, you can ask them. Although one by one, it's going to take a while. Yeah, it's going to take a while. But, like, you know, their bank knows.
Starting point is 00:59:23 the person, the bank that gave them that credit card or, you know, the Visa or MasterCard network, or like the company that provides the, you know, plumbing to a host of like credit unions, for their like credit and debit card programs, like, they know. And maybe they're willing to share like anonymized or aggregated. statistics about like what spend they're seeing. And so you get some information about the transactions that some panel of cards are making. And then where do you go from there? It's not a random sample all consumer transaction in the United States. Transactions belonging to like one credit card program, maybe the people who shop and have cards from, like, credit unions are different from
Starting point is 01:00:31 the transactions from people who, like, have a J. Sapphire Reserve. Maybe, like, the demographics of the panel are just not the same as the broader United States. Maybe people, like, drop in or drop out of the panel. And, you know, if there are more consumers in this sample, does that mean that people are spending more at Walmart? Or does it mean there are more consumers in the sample? Sure. And so, like, there are interesting modeling questions about, like, well, we have this information and how do we, like, make, like, make. good predictions based on it adjusting for all of those factors. And is that like trying to kind of model out and adjust and maybe like cross compare different
Starting point is 01:01:26 sources? Is that part of the data engineering work or is that part of what happens like on the trading desk after you get like the cleaned up data put in front of you? It's a mix of it too, I think. But like some of this like really just does happen on the trading desk. But I think the groundwork for doing all of this comes from like a deep understanding of how these data sets work and what information they provide you when, which the data engineer typically builds up a understanding of. One thing I'd worry about it, you mentioned there's all this like aggregation and anonymization that's done by the upstream vendors.
Starting point is 01:02:07 You could imagine like that could be done wrong or just done in ways it introduces weird biases into the data that's hard for you to see. Like is that a problem that you guys end up having to wrestle with? Oh, I mean, so, yes, vendors like do weird things with their data all the time. Sometimes they'll, like, you know, have data missing. Sometimes they will say we compute this column one way and then later tell you, who made a mistake, all those values were wrong. here's what you should have used and said.
Starting point is 01:02:51 Sometimes, you know, sometimes associating, like, a transaction with the, like, stock of the, I don't know, the merchant is difficult. And they improve these taggings over time. So sometimes they'll, like, say, you know, these transactions that, you know, we said or one thing are actually, like, from this other company or other brand. And, Haki, one thing that shows up if you have this issue where they're going back and fixing stuff in the past is all these causality questions, right? Right. On one hand, you know, isn't that helpful? They're fixing the data set. Yeah, it was wrong now. It's right. On the other hand, if you are like doing a study at a firm like Jane Street and you're trying to figure out like if I had been subscribing to this data set, like would I be able to do? good trades. Should I purchase the dataset?
Starting point is 01:03:49 If you're trying to make that decision and you are looking at the data that is not actually what you would have received, but instead, you know, the corrected, more accurate version later, you're not actually studying like what trades you would have done if you've been subscribing to the dataset. You're simulating what you would have done if you had a time machine. Right. And although it's subtle, right? it depends on the nature of the correction, right?
Starting point is 01:04:17 It's sort of, there's a question of, like, in some sense, what could have been known at the time? And also, like, what does it mean for, like, what could have been known in the future? There could be some data set where, like, the underlying data, like, has real information. And then they correct a completely mechanical thing. And then it's okay.
Starting point is 01:04:31 But, like, it's very easy when you're doing that thing to smuggle information into the past. And, like, it becomes very subtle words. If you just, like, timestamp what you had at the time, then you know, well, at least it was definitely possible to have it at the time because at the time I had it. Yes. So what we love is when vendors say, like, this is exactly what I delivered at this moment in time. And, you know, here is the corrected version that I delivered at this other point in time. And not everyone does that. And so what we try to do is save all the data that we have access to, along with the time we received it. So that no matter what transformations we do later, we can always reconstruct, like, what we would have known when.
Starting point is 01:05:19 Right. And no matter what the vendor does with their timestamps, we at least always capture our timestamps of when we got it, which, by the way, is a very close analogy of what we do with market data, right? Market data also comes with timestamps from the exchange, and we look at them with a certain amount of, like, suspicion. Because, like, all sorts of weird stuff can happen. You can have some weird clock synchronization.
Starting point is 01:05:36 The times can be off in various ways. But we capture stuff when it lands on our network. And again, like we sort of, there's still information in their timestamps, but like there's a way in which we trust ours more than we trust theirs. Yeah. And knowing like, okay, I receive this packet at this nanosecond in time gives us more confidence, you know, making decisions based on our historical studies than like, I receive this packet at about this moment in time. Right. Yeah.
Starting point is 01:06:08 So to switch gears for a sec. Yeah. One thing that has changed a lot in the last handful of years. In fact, a lot of that change overlapped with a period of, you know, building up the all data team is, boy, golly, has AI moved a lot. And I think it's kind of, there's kind of two different ways within the J&Ching context where it's changed. One is we have a lot more ML models, neural net models,
Starting point is 01:06:35 that are being used as part of the trading process for consuming data and transforming it. And then also we have these like super smart LLMs that are like useful productivity tools. And I feel like both of these have lots of potential to change the alt data process. And I'm kind of curious how the impact of those has landed along the last few years. So I think there are two kinds of impacts of AI here. One is that like we have better models for consuming this data and making predictions. about the world, finding patterns. And, you know, that means that, like, the value of all of our data, of good quality
Starting point is 01:07:19 data sets has just, like, gone up. And with ELMs in particular, like, it's easier to make use of text data to, like, extract information or features from text that, like, would have been manuals, annoying, impractical, a nest of rejects beforehand. Yeah, they're like an amazing tool for extracting structured data from unstructured text. Yeah. Although also there's like a huge causality problem, right? Because if you use like a really up-to-date LLM to like that knows stuff about old data
Starting point is 01:08:00 and trying to use it to featureize some old data, like you're going to get very confusing things to happen because the LLM is smart and knows things about the future then. Yeah, the LM knows what. happened to, like, Enron. The LLM's training data, like, maybe it includes the financial statement or a future one. Right. So that presents some difficulties. So then there's, like, AI as a tool where, like, a lot of the same trends and
Starting point is 01:08:31 limitations that you see in software engineering also high to data engineering. They make it easier to produce code to do analyses. And in the hands of, like, good data engineers who know what they're trying to get from these tools and can tell whether the results are good or bad, they're really helpful. And, you know, our data engineers use LLMs. Also, they make it easier than ever to generate, like, bad. code right um never has the gap between merely doing something and doing something well been larger in terms of cost right it's so easy to just do something but like sometimes the results are very bad yeah and it's easy to do something and think i have solved the problem when in fact you
Starting point is 01:09:25 have like failed to understand the problem so you have the normal ups and downs of using AI tools like they're they're amazing and also that have lots of pitfalls um and then you talked about AI as a kind of feature extraction, but there's also like neural nets as models, right? There's new ways of extracting value out of the data that you get. Like, has that changed the process? Has it changed like the value of having alternative data? Has it changed the pitfalls of using it? It certainly changed the value of having data alternative or otherwise. Like it's generally more valuable. Maybe like some.
Starting point is 01:10:06 shapes and quantities of data are, like, more amenable to these models than others. So one I'd wonder about is, like, the value of some of the kind of what in other contexts we'd call future engineering. Like, I think one of the, there's the whole, like, bitter lesson idea that lots of ways that you try and, like, express priors into the data itself or the shape of the model or whatever, maybe become less useful as the models get bigger. and more powerful and stuff because some of the kind of regularization and featureization can kind of be done inside of the model. So there's like that. And then there's also a question of like how just in general how much like data cleaning and data preparation like is that like more
Starting point is 01:10:49 useful in the neural net context? Is it less? I think the value of cleaning the data is greater. I think like I don't know if your data is not point in time and secretly leaking information from the future into data that should be in the past, these stronger models are going to be much better at picking up on those effects that are not actually, like, tradable. I think there are certain kinds of outliers and weird patches of data that, like, some models, like, handle well and other models, like, handle poorly in their training process.
Starting point is 01:11:33 and getting the quality right is, you know, more important for those models. So, like, internally, that models will be able to do more. You will have to, like, handcraft features less, and it will be better at figuring out, like, which ones are important or not. But, like, the value of getting the data that you do pass in right is pretty high. And I think then it becomes like, well, our LLMs with the skills and harness and whatever we build about them going to be able to completely automate that. I think that's similar to will we replace all software engineers with LLMs too. Right. The current horizon and at least what you can see me that seems like an enormous amount of like human judgment is incredibly.
Starting point is 01:12:32 incredibly important. Yeah. Like in some ways, it feels like human judgment just gets more important as like, it becomes the thing that unlocks. Like, you can, like, do all the stuff really fast if you can figure out if it's good. So, like, the verification bottleneck becomes really important. And it's a kind of very critically human part of that. And taste for what should be built.
Starting point is 01:12:50 Well. Absolutely. Yeah. Speaking of what should be built, I feel like the alt-data world here has been, like, you know, a kind of Jane Street-flavored software engineering. adventure in that like, boy, we've made a lot of our own exciting custom stuff. And we've also used some like external outside world stuff. And I'm curious like how you guys have thought about and navigated the question of like,
Starting point is 01:13:14 where does it make sense to buy some external product for like helping to manage this? After all, we're not the only people who manage these kind of all data pipelines. And where we have decided to make our own things and like, has that been good, has it been bad? Like, how do you think about the choices there? Yeah. Some of our software stack, like, is very much pulled from the rest of the world. Like, a lot of data engineers outside Jane Street use DBT, which stands for data build tool. Data build tool is like a tool for, like, describing and orchestrating transformations, like, in your data warehouse, usually, like, using SQL queries, but in, like, a code reviewable, testable,
Starting point is 01:14:01 composable, good software engineering practices way. Sounds great. So I have seen 300 line postgres views and functions in my time at Jane Street. And we don't want any of that. Like DBT lets you break down the data transformation problem into smaller components and test the properties of the data at each stage so that you can rely on it. Mm-hmm. Our, like, data warehouse's query engine is Trino.
Starting point is 01:14:36 There are lots of Python tools for, like, processing tabular data. But other things, you know, we do build ourselves. Where, like, if we were starting from scratch and we're entirely cloud-based, maybe we would do something more off the shelf. But one thing that is pretty cool is that Jane Street is, like, good at building and operating, like, on-prem data centers. and we had to be in order to, like, run low latency trading systems before we were doing any of that alternative data stuff.
Starting point is 01:15:13 That's right. Right. We need our own data centers in part for, like, the kind of physical location thing where you need to, like, have your boxes close to the exchange. And also just some of the, like, underlying abstractions, like clouds. No good at giving you multicast, right? There's all, like, you can't easily deploy FPGAs. There's all sorts of things that you want in terms of, like, control of the things.
Starting point is 01:15:31 physical architecture that you get in your own data center that you just kind of can't get in the cloud. And like the cloud gives you a bunch of other benefits. And these days we do a lot of both. Right. If you do have the ability to do things well on prem, solution that works is maybe not the one that, uh, you would use if you were like just using BigQuery or snowflake or, um, what have you. Right. And I guess part of this is like, you know, came up in a previous, uh, one of these conversations was we built our own data warehouse, right? So we have this thing called Superstore. And we didn't build it for all data, or at least not primarily for all data, but like, that's
Starting point is 01:16:09 like a tool that Alt Data now uses. Yes. And it's great. It's great having all of our data in one place and being able to handle the volumes of trading data and other data that we use. It's useful for being able to join it to all of our other services internally. and the cost model is different. I don't want to imagine how much we would spend in credits or slots
Starting point is 01:16:40 if all of the queries that we are running now were built in like a cloud warehouse. So why is that like how is the cloud cost model different from the kind of cost model that we get from this thing that we built ourselves? So, like, the cloud cost model will be based on, you know, either the amount of data you read or the amount of time on, like, time on, like, in their units of compute, they use. And what we can do is just think about, like, the cost of acquiring the hardware and operating the hardware, which, all in all is like significantly lower. Like they have like large, uh, gross margins.
Starting point is 01:17:36 Um, is it, isn't it the overall margins are, are larger? Is it more that like the incremental cost model? Yeah. You're also not thinking about like the cost of like trying this one new thing, this one new query. And if you just sort of take the things that you get from cloud vendors and say, let's, let's look at those prices and then apply it to our workloads, you'd be like, wow, I could get a lot of hardware for that cost.
Starting point is 01:17:59 So just like the way the math and practice works out is that like it would incentivize us to spend a lot of time figuring how to optimize and do less work on the system than it does when we actually go out and buy the physical hardware. So like economic theory and like architecture aside, it's just like that's how the pricing model seems to work. Yeah. If we were using one of these like other products,
Starting point is 01:18:24 we would spend a tremendous amount of time. like optimizing our spend and and like trying to shape our workloads to like fit their cost model. And we do think about the performance of Trino and our data warehouse. We end up spending less time on that. Right. Because because overall the cost in fact are lower. Yeah.
Starting point is 01:18:48 Makes sense. So another thing I think is maybe you can talk about is like, you know, we went from in 23, one data engineer to now more than 20 data engineers. And it sounds like we're continuing to be very hurry and are eagerly hiring. We love to hire more. And so I'm curious, like, what makes a good data engineer? Like, what are you looking for when you're trying to hire someone? And then how do we find them?
Starting point is 01:19:14 Yeah. So the most important characteristics, I think, are like curiosity and a good, investigative process. As part of the job, you'll be learning about some new domain, some industry, some kind of data, some kind of trading, and also you will have like messy, unfamiliar data that is useful, but also has lots of embedded problems in it. And I think to succeed, you kind of have to enjoy learning about all these new things. And you have to enjoy learning about these new things. You have to enjoy, like, getting into the details.
Starting point is 01:20:03 You have to be careful and think about, like, what you know and what you don't yet know and how you would support going from point A to point B. I think people with science or social science backgrounds who do, work with like messy real world data and try to make sense of it are well suited to this. We also, you know, look for engineering sales. I think the systems complexity of what data engineers are doing is, you know, lower than much of what we ask yourself for engineers to do. But like, the data and business complexity is still very high. Data pipelines are code. You still want them to be good code.
Starting point is 01:20:52 You're still software. You want them to be good code. You want it to be like, you know, clear, correct, maintainable. Mm-hmm. And we look for that too. And where do we find data engineers? Some of them have worked in the financial industry before. Some of them are people who are like the data person at their like small or mid-sized startup
Starting point is 01:21:19 where they, like, had to, like, kind of understand the data and the business context and deal with its messiness because there were only so many other people around to do that who also, you know, have developed, like, good software engineering skills. And so are you mostly looking for people who essentially kind of have data wrangling experience is kind of a thing they've done already professionally? Usually, though, like, that doesn't have to be, like, you know, at a company. It can be, like, as part of, like, research of doing science or some other way. All of the people we've hired so far have been experienced hires who, you know,
Starting point is 01:22:12 we're doing some thing beforehand. And next year, we're hoping to venture into hiring, like, interns. Yeah, we started talking in 2025 about, like, interviewing in 26 for an internship that runs in the summer of 2027, for people who start in 28 after graduating from school, probably, and, you know, are really ramped up and doing useful work in 2029. Right. Which is long-term planning. And I think we're looking for people who, like, care about the data first and, you know, maybe they have good engineering skills now. Maybe they, like, have the potential to you after a bunch of training.
Starting point is 01:23:01 Right. But, like, the kind of excitement about digging into the data is, like, the primary thing you're looking at. Yeah. While, on the other hand, there are people who work with a lot of data and are, like, I am excited about building, you know, fancy models and neural networks or about, like, building, like, building. scalable systems, the ML or the software engineering parts of it. And, you know, we want those people at Jane Street, too, but not for this role. Got it. And then just at the interview side, like you want, you know, to find people who are excited about and have good taste about how to think about and dig into data. Like, how do you use an interview to suss that out? It's hard. We,
Starting point is 01:23:47 have, you know, a few data investigation interviews or data wrangling interviews, where we give them some unfamiliar data set, and they investigate and build some model of how it works and, like, do something with it. Seeing how someone approaches that problem, are they careful and detail-oriented, and do they, like, think about what they know, what assumptions they're making, how they would test those assumptions, or do they write a bunch of code and say, hopefully this works? But I'm not sure. Right. And once again, kind of rounding back to the like, you know, kind of engaging with like the real world underneath the data and actually think about what's going on and seeing if that
Starting point is 01:24:43 affects what you should be doing. Yeah. Awesome. All right. Well, maybe that's a good place. to end it. Yeah. Thanks for joining me. Thanks, Ron. You'll find a complete transcript of the episode along with show notes and links at Signals and Threads.com.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.