Odd Lots - Why Corporate America Still Runs on Ancient Software That Breaks

Episode Date: January 26, 2023

Southwest Airlines had a disastrous holiday season, thanks in part to a software bug that left crews out of place and grounded thousands of flights. But Southwest isn't alone in having software in the... headlines lately. The New York Stock Exchange recently had a software error that caused weird pricing on stocks and the FAA had its own computer issue that grounded planes earlier this month. So what's the deal with corporate software? Why do these crashes happen? And why does the user experience typically leave something to be desired? On this episode of the podcast we speak with Patrick McKenzie, an expert on engineering and infrastructure, who writes the Bits About Money newsletter and recently left payments company Stripe after six years. We talked about the challenges of keeping any software system alive after years of upgrades and updates, the distribution of tech talent across industries, and whether non-tech companies can close the gap with Silicon Valley.See omnystudio.com/listener for privacy information.

Transcript
Discussion (0)
Starting point is 00:00:00 Thanks for listening to Odd Lots. Follow the show on Amazon Music for more future episodes or just ask Alexa, play the Odd Lots podcast on Amazon Music. Hello and welcome to another episode of the Odd Lots podcast. I'm Joe Wisenthall. And I'm Tracy Allaway. Tracy, I forgot to ask you, and it's kind of embarrassing like this late in January, but how is your, uh, how is your New Year's? Like, how is your holidays? Do you have a good Christmas and stuff? Oh, thanks. Thanks, Joe. I had an excellent Christmas. I stayed at home for a week with my husband and my dog, and we did hardly anything, and it was absolutely glorious. How about you? It was all right. So the thing was I was down in Texas visiting family, which is nice. But there was that huge cold blast, and it's worse in, when the weather gets really cold, it's worse in a place like Texas because like none of the buildings are insulated particularly well. them of that heat. So it's like when it's really cold, it's actually like better to be in a cold place where people are used to it. So it's a little uncomfortable. But the good news is somehow I managed
Starting point is 00:01:16 to travel back and forth without getting having any major like airline disruptions. Right. So this is the key thing that happened right before Christmas, which is we had that very big winter storm, the Arctic bomb blast. And it disrupted a ton of flights. First off, because of the weather, But then what happened is you had this sort of cascade effect because the weather event was so large. A number of airlines, but one airline in particular experienced a lot of problems with its software. And Southwest had to cancel, I think in the end it was something like 16,000 flights. You had millions of passengers affected. And you had disruptions that, you know, a weather-related disruption that lasted one or two days ended up lasting, I think, more than a week.
Starting point is 00:02:06 because of the impact of the computer glitches, I guess. Yeah, yeah, that's right. Like, you know, airline travel, like, it always sort of cascades and ripples out, right? Because a canceled flight is going to affect other flights and so forth. But it seemed like Southwest experienced something unique, which is that it turned into this major software problem. And it's sort of like a reminder that, like, okay, when we use the Internet, when we use, like, sort of like, modern consumer software,
Starting point is 00:02:33 it's all very zippy and quick, and it has nice. interfaces. And then when you use like sort of like back end corporate software, particularly at large legacy institutions, it's just like it's nothing like the consumer internet. It's like clunky. It's like we all know what it's like. Right. So this is something that came up quite a lot. I used to cover the big banks. And one of the crazy things that I learned relatively early on while I was doing this was just how much of their IT system was still these old. creaky, big iron mainframes. Some of them still running on cobal, which is the programming language that I think it dates back to the 1950s or 1960s. And I remember, you know, you hear this. You hear like, oh, I can't believe that these big banks, our entire financial system in some respects is still running on these legacy computer systems. But on the other hand, if you look at what a big bank is, it's structurally a series of mergers and acquisitions. There used to be that, No, that's a good point.
Starting point is 00:03:37 That's really well put. There used to be this great flow chart that showed the formation of like a JP Morgan or a Bank of America. And you can just see it's like a series of roll-ups of smaller banks. And you think every time they acquire a new bank, they have to integrate another system into their own system. And in the end, you kind of end up with this just incredibly complex and kind of patchy IT structure that in some ways very much resembles the amalgamation. of all these smaller banks into a larger bank. Yeah, that's totally right. Like you think of like banks and when they merge.
Starting point is 00:04:13 It's just like, okay, like, you know, the capitals all merge together and the assets, etc. But like, right, like none of them are going to have like IT systems that work perfectly together. So they're like glued together with like duct tape and just over time. I think technical debt is a term that software engineers. And so you accumulate all this technical debt. Good term. And I'm sure like nobody likes it inside the bank at any level. But I've always been curious, like, what are the economics such that it is just so impossible for these legacy institutions, banking, airlines, the public sector to, like, get with the times, you know?
Starting point is 00:04:51 Absolutely. And when it comes to Southwest, so first of all, let me declare a small self-interest in Southwest, which is my dad was a pilot for Southwest for a long time. And I remember when the outages were happening this Christmas, I sent him a text going like, well, what do you think about this? And he just said, well, all the crew are out of position and they need to get them back. So a very straightforward ex-pilot answer about what was going on. Not that helpful for this podcast. But I do think in the case of Southwest, this has been a long running issue. And you have had, you know, people talking about it at various times the need to upgrade the infrastructure, the technological infrastructure. And yet it hasn't happened. And so the question is why. The question is whether or not is it cheaper just to keep running these old. mainframes and assume that you are going to have these outages during major events, cheaper than it would cost to actually upgrade it? No, absolutely. Well, let's talk to somebody who knows about software,
Starting point is 00:05:48 knows about economics, and can walk us through and help us understand the problem. We're going to be speaking with Patrick McKenzie. He's an expert on software and infrastructure. He's a writer of the bits about money newsletter, knows a lot about finance. and he just left Stripe after six years. He's still an advisor there. I've been reading his stuff for a long time, one of these people I sort of trust on almost any topic.
Starting point is 00:06:13 Patrick, thank you so much for coming on Nodlots. Thanks very much for having me. What is the deal? Let's just start, like, straight up. Like, why don't you like... I mean, Tracy mentioned, like, all these institutions, they just, like, they swish it all together. They sort of, like, tie it together with duct tape and so forth.
Starting point is 00:06:31 You get these unwieldy things. Just like give us the high level view of like, why is it so difficult at an abstract level to modernize legacy software? So I'll start with a disclaimer, which is sort of mandatory in engineering culture. We have this thing that we've come up with over the last course of two decades or so called blameless postmortems, where when there is a failure within a company where, you know, planes cannot be up in the air for a week at a time, rather than trying to point the finger at someone and say it was your decision or your inaction that caused this event. We, as engineers want to look at the objective reality of the system, figure out what went wrong for the benefit of both that organization and for the larger community.
Starting point is 00:07:15 And so this isn't to grind their nose in it, but just, you know, as an engineering matter, what probably happened. In an ideal world, and if this had happened at a Google, for example, the engineering teams would push because of the culture of these things to do a very public post-demortem. of what the decisions were, what the background is, etc., etc. In more traditional industries, I don't think that culture is fully baked yet, as it were, although there very well might be a post-mortem by, you know, the FAA and federal regulators because, you know, liveness constraint is a real thing for extremely important economic systems like airlines.
Starting point is 00:07:52 Anyhow, what probably happened, it wouldn't surprise anyone in the bowels of an airline that if you put off maintenance on airplanes for decades at a time that eventually bad things would happen, and no one would countenance that. However, software systems are quite similar where they aren't a build once and run for the rest of eternity sort of things. There were some decisions made early in their lifetime, which are no longer accurate for the world we live in. They do suffer from something engineers euphemously called BitRot, where software which worked back in the past will tend to succumb to mtrip you over time and not work exactly perfectly for all time afterwards. And so you need to be
Starting point is 00:08:32 doing an ongoing program of maintenance for your software just like you would for your airplanes. That bluntly was not done. And it seems to be credibly reported that the sort of cultural factors at Southwest that caused that to not be done might have been caused by a sort of like overly accounting slash penny pinching focused management culture, which thought, well, it costs money in the short term to do maintenance. We can cram down our engineering costs by doing less of this and relying more on external vendors, etc., etc. And then when stuff hits the oscillating plaid, the right people have not done the years of work that are
Starting point is 00:09:10 required to get to a posture where you can quickly recover from failures. And so they were left in a part where we would use euphemistically in the industry called Heroics was required from folks in operations trying to, you know, contact thousands and tens of thousands of employees by following and pass around their information on probably spreadsheets to figure out where the crews actually were to be able to tick the boxes that are required due to regulation to allow people to get back and up in the air. And so that's probably the high level like root cause of what happened. But like a reason to conduct these postmortems is it's different one like one single
Starting point is 00:09:49 decision made by one person. It's a result of cultural factors, business decisions made over a of probably decades in this place. And we want to like tease out the various nuances there and like make both Southwest and other organizations aware of that so that they don't suffer critical systemic failures in their own places. So first of all, BitRot is a fantastic term that I am going to have to try to work into all my conversations going forward. But secondly, just to step back a bit, can you maybe explain, you know, when we talk about mainframe computer systems, what exactly are we talking about and like what is the counterpoint to mainframe i'm assuming it's more like cloud based applications and things like that but could you maybe define those that basic term very quickly
Starting point is 00:10:36 and then secondly my understanding of the southwest debacle was that they had this in-house software i think it was called sky solver but it was based on an application that g.e had been selling and then Southwest kind of customized it. And I guess my question is, how endemic is that type of software where you get something off the shelf, but then you customize it in such a way that it becomes, I guess, special to you and therefore your problem when it goes awry? So let's talk about the main friend first, and then we'll talk about the problem of who owns the problem, whether it's your problem or some of the vendor's problem, and the way that balls to end to be in the middle of people. and things get dropped. So mainframes, back in the day, many, many decades ago, computers were
Starting point is 00:11:28 approximately the size of a room. This is before the personal computer revolution. And banks, which were some of the earliest adopters of computers for sort of scaled usage in industry. And it's funny, like, the earliest users and computers tended to be either the financial industry or the military, either attempting to, like, move numbers which represented money or numbers which represented like literally artillery shells flying through the air. And both were very important to society. For better or worse, like banks, because they standardized on what was the best available technology at the time, ended up with a lot of mainframes, and they have kept those mainframes running for a good portion of 70 years now, in some cases, as you
Starting point is 00:12:06 mentioned earlier. The alternative to mainframes, cloud is a bit of a buzzword. So there's the personal computer form factor that you're familiar with where something sits on your desk. there are servers which would typically sit in a server rack somewhere and the difference between like servers and quote unquote the cloud is like the traditional way to manage servers would be you would have a data center that would be owned or leased by yourself and you would put hardware that you owned in that data center and in the cloud the cloud case the data center is owned by amazon or google or microsoft you have rented access to a machine there which sits on their balance sheet and
Starting point is 00:12:43 it's probably not one machine. It is probably like an awful lot of machines potentially with some virtualization layer and your engineers can cause systems to scale up or scale down based on how many machines you need on a minute by minute or second by second basis. It's sort of like the high level version of the sales pitch that cloud vendors will give you. So that is a quick run through between like four generations of the mainstay of how technology gets done at scale. And the question of where software gets written. So it is quite common for businesses in the traditional economy to not have an internal software engineering competence. And as a result, they'll go to vendors, like in this case, you probably know this better than me, but for example, GE, and the vendor will sell quote-unquote package software, which might not be fully responsive to the needs of the business, and then some customization happens. What could happen is the customization might happen within the business if they have some level of software engineers internally. what happens more frequently in a lot of places, particularly in Japan where I live, is that the business will contract with, we call them system integrators here,
Starting point is 00:13:50 they may be called consultancies in America, you know, a Deloitte or another large consulting firm like that, and Accenture, and say, okay, we have this package system, we have these business requirements, clearly some software engineers have to be involved, we don't have enough on our staff,
Starting point is 00:14:07 can you, like, figure out the missing link for us? And so there is an, sense of contract negotiation, consultancy goes off and does the customizations, and they deliver the work to the organization, and then there are various different ways to do maintenance, but often that will involve the original team standing down for a while, which that always sounds like a great idea when it's pitched because it's like, oh, great, I don't have to pay expensive engineers every week to just sit around waiting for something to happen. And then, when something actually happens, you sort of like realize that the other end of the option value there, where
Starting point is 00:14:42 oh, not so great. I don't have a team of like experts who understand the system ready to call up at a moment's notice and bring in to debug the problems that we're seeing right now. And so I have no specific knowledge of how it went in this individual instance. But a thing that happens during a lot of outages is a, you know, quick look into the history of the system to say, hey, wait, who actually built this for us? Are they still in business? Can we get on their calendar immediately? Is there an engineering team there that is ready to, you know, to hop in and work at this, like, goodness, they're going to charge us, I-bleeding rates to do it, but we're going to have to pay that. That's a tertiary consideration at this point. The bigger consideration is like, how quickly can we get our, you know, organization to organization, like legal paperwork, et cetera, spun up? And then how quickly can they get the engineering team spun up to address the system in real time? And this sort of thing is why the sort of like major-scaled software companies in the economy that Google and Microsoft, et cetera, et cetera, you all know the names.
Starting point is 00:15:45 Largely treat engineering as an internal competence rather than, you know, the software for the iPhone isn't built by Deloitte, it's built by Apple. And almost everything in the critical path will either be open source or something that there is a team on Apple that owns the sort of totality of the experience for Apple. Hello, I'm Michelle Hussein, and for more than 20 years, I was at the BBC. military withdrawal from Afghanistan. But all the time I was delivering the headlines, I wanted to go further than the news of the day,
Starting point is 00:16:30 to spend more time with the people shaping our world. And that's what I'm doing here on this podcast. Speaking to people from Nigel Farage, to Russia needs to be taught a lesson, to tech journalist Karaswisher. And the tech industry is running wild. You know, they've gotten what they wanted and they've seen a huge run-up in their stock prices.
Starting point is 00:16:50 This will be a place where every weekend you can count on one essential conversation to help make sense of the world. So please join me, listen and subscribe to the Michelle Hussein show from Bloomberg weekend, wherever you get your podcast. You certainly ask interesting questions. What separates good leaders from transformational ones? I'm Jessica Chen and in season two of Leading By Example, We'll sit down with executives like Grace Chen of Bertie Gray to find out.
Starting point is 00:17:28 It's important to understand where you spike, but also really acknowledge where you don't and find people who can fill those gaps. Listen to leading by example, executives making an impact on the IHeart radio app, Apple Podcast, or wherever you get your podcasts. So something that I sort of hinted at in the beginning, and we talked about this a little bit. We did an episode on Petroleum Engineering, but I'm really curious about, like, the distribution of engineering talent. Like, in my mind, I would have to imagine that a talented engineer would be more excited to work for Stripe than Google, and more excited to work for Google than Deloitte, and more excited to work for Deloitte than Southwest. And I'm just sort of, and I'm curious, like, A, if that's the case, and whether, like, this is a problem, and whether it would be better if there were just more high-tech experts, software experts,
Starting point is 00:18:31 who would want to work for a Southwest directly or work for a bank directly, whether this sort of distribution of engineering talent contributes to software bottlenecks. So I have nuanced thoughts at engineering talent. On the one hand, like, that is certainly a thing that exists in the world, and, like, talent is not necessarily distributed equally among different industries, individuals, etc. On the other hand, I think that Silicon Valley an ecosystem to which I owe a lot, occasionally has an overly enamored view of itself thinking that, oh yes,
Starting point is 00:19:02 we have the best engineers in the world, and all the engineers who made every other system that we rely on in our lives are sort of second rate, and clearly that's not true. Like, the phone system works, airplanes, like, what's the old Louis CK monologue? Like, it flies through the sky and then teleports beams up to space. Like, none of that happened by accident. And so, So it's important to not focus on like the, you know,
Starting point is 00:19:29 engineers working in traditional industry are worse engineers, that engineers that work in software companies. The larger problem that they have is one, they don't drive the bus. They have less ability to control the situations within their organizations and control, like, larger decisions made such that they have influence over decisions on like what the maintenance schedule would look like or who gets to make decisions with respect to whether something.
Starting point is 00:19:53 ships or not. And there is a little bit of sort of like the life cycle of engineers thing that happens where the original architects of a system did it 40 years ago. Careers are about as long as they are. Many of the original architects of the system will sort of be aging into the retirement years at this point. And so there is a question of did the organization put in the work over the years to recruit newer engineers to inculcate them into how the system is made, et cetera, et cetera, or do they just allow all the knowledge to walk out the door? And did they, like, a thing that comes up a lot is, do they put enough work on a day-to-day basis to maintain an engineering brand
Starting point is 00:20:30 so that they can get new talented engineers to join them in 2023, such that those engineers are the word used to soften graybeard. But, you know, the wisened veteran 30 years from now so that when something happens in 2053, there are people who have been around the block that know where the skeletons are buried in the system. Interestingly, traditional industry is getting better over the years at having an engineering brand sort of moving away from this world where engineering was largely seen as a cost center,
Starting point is 00:20:59 where the goal is just to cram down the amount of money you spend on it and improve your margins. I think there's a few things that played into that. One of them was particularly in the internet age. It became obvious that, you know, in finance you talk about like the front office and the back office. Engineering used to be in the back office and the front office where the salespeople, people live is the one that generates all the money for the bank. And increasingly because, like, experiences that were in the palm of the user's hands were the thing that were actually generating the sales, those experiences became sort of institutionally important within
Starting point is 00:21:33 banks and airlines and other firms. And then when those experiences became important, took a while, but gradually, the people and teams that built those experiences became more institutionally important than they'd previously been. And so, like, if you look at the large money center banks in the U.S., they certainly have no small number of, like, technical challenges. But a thing that is, like, largely true in 2023, which was not true in 2010, is that their mobile apps are actually kind of good these days. Like, if you download, you know, not to endorse anybody in particular, but like Chase or
Starting point is 00:22:05 Capable 1, when you play with their mobile app, it's like, oh, this kind of feels like a mobile app made it in Silicon Valley. And the reason is, well, yeah, they hired a lot of people who made the apps from Silicon Valley. And those folks brought their skills and sort of like level of competence with this and their taste now exercise it on behalf of old world companies. One hopes knock on wood that that seeps into the back end of these systems too where Apple, Google, Microsoft, they have very talented teams on the front end of their systems, but they also have very talented teams on the back end of the systems. And that's bluntly why you don't see like core systems at
Starting point is 00:22:39 Google going down for a week at a time. That is almost unimaginable, you know, how much of capitalism would break if Google Docs was just down for a week. Given that, there are many parts of the world where there is a liveliness constraint. We would, like all else being equal, if it is safe to fly airplanes, we would prefer that there would be airplanes flying versus all the airplanes being on the ground, because airplanes generate value for human society. If that is true, then it must be the case that the back-end systems of airlines that control whether airplanes are allowed to fly at a given time. That has to be, at least as a important as Google Dox is, which implies that the airlines need to put at least as much work
Starting point is 00:23:19 as Google does roughly into having true mastery of their own back-end systems and the various problems that can happen there. So just on this note, you know, there's another travel disruption that happened recently and we haven't even mentioned it, but that was the FAA experiencing some sort of computer event and grounding, I think, all the domestic departures for one morning. And it was recently reported that the proximate cause of that was because there were some software engineers who were trying to upgrade the system and they accidentally deleted a bunch of critical files as they were trying to do this. Can you talk a little bit more about the technical challenges when it comes to trying to fix some of these legacy systems? Why exactly is it so difficult? You sort of talked about it from a organizational perspective, but from a technical perspective, but from a tech.
Starting point is 00:24:11 perspective, why does this seem to be such a big challenge? Sure. So one thing is that that symptom, the underlying cause is that a problem happened during an upgrade is extremely well understood in the software engineering field. It's called out in, among other places, Google's book about site reliability engineers, which is essentially the subcategory of engineers that Google relies on keeping the world running. And it is often the case that the people who are attempting to make an incremental change. do not have full context of how the system came to be in the state that is currently in, and a change that they thought will have a limited sort of area of impact,
Starting point is 00:24:51 ends up having a larger area of impact. The term of art we use in the industry is blast radius. Like you hope to quantify the amount of blast radius of something that goes wrong, such that, you know, if I'm making a mistake, am I going to, like, bring down our blog, or am I going to bring down, like, credit card processing worldwide? And, you know, be much more careful if the blast radius includes, like, worldwide credit card processing.
Starting point is 00:25:11 plausibly, like, you should engineer your system such that there's no way to take down worldwide credit card process. That's actually harder to do than it just sells. So anyhow, how does one get to the point where it is difficult to understand, like, what the implications of the changes you're making truly are? These are, like, meat and potatoes questions, and then the meat and potatoes' answers are often things like, was the system adequately documented when it was made? Frequently the answer is no.
Starting point is 00:25:37 And a lot of information about, like, how systems are put together. survives this oral lore within the engineering team at various companies, which, you know, that's an uncomfortable bit of information to hold in your head when you start talking about, like, the lifecycle of engineers and the fact that the original architects of many of these systems are literally no longer with us, either because they have retired or they might be like beyond our ability to call up out of retirement at this point. And so you have to write down what you do. And that concept was not like, new to governments and bureaucracies as a result of software engineering happening in the last
Starting point is 00:26:17 70 years. It's sort of fundamental to the operation of large organizations. Software is just how people choose to do work with each other, but is a lesson that we keep relearning. There is often that issue where, because software is how people and organization choose to work with each other, often software will interface various systems together and problems will happen at the boundaries between systems, either between literal computer systems or between other breakages between organizations. So a thing that you will see frequently is, my software didn't fail, your software didn't fail.
Starting point is 00:26:51 We mutually fail together at that point where we are supposed to transfer information, and then both sides end up pointing at their counterparty. And so part of the discipline of software engineering is one, like creating a culture where you don't want to point fingers of the counterparty into creating structures and incentives such that, you know, like complex systems that involve multiple different parties with multiple different engineering teams who might not report to the same
Starting point is 00:27:18 payroll department will like converge on correct outcomes. And there are a variety of ways to do that in this industry. Some of them are better than others. So without like naming the particular company, there exists a credit card system, which is extremely, so credit card system as a baseline are extremely reliable. You probably don't remember the last time that you were unable to use a credit card for a week
Starting point is 00:27:43 because literally that does not happen. So like A plus plus for achieving that outcome. What one credit card company does to achieve that is they are willing to make changes to their system precisely twice a year after six months of testing every change that they make. That's an incredible amount of upfront work to do like relatively small amounts of engineering. And so the pace at which that ecosystem involves is much, much slower than more software-forward companies like Google Apple-Amazon, et cetera, et cetera, where they're shipping thousands of changes to their systems every day.
Starting point is 00:28:19 So one of the interesting bits about the remixing the skills and techniques of Silicon Valley is attempting to get people who are comfortable with the engineering practices that allow you to ship software thousands of times a day into positions of authority at old line companies such that they can gradually transition from the point that they're at where they might be able to ship software once or twice a year, maybe quarterly to the point where they will be shipping software. Like, let's start with five weeks and move up from there. I'm Matt Miller. And I'm Hannah Elliott, inviting you to join us for the Bloomberg Hot Pursuit podcast. Every week we bring you news and industry insight on everything cars.
Starting point is 00:29:14 And we do a whole lot more than just talk about cars, Matt. We actually get behind the wheel of basically. every latest model, especially the luxury ones and the sports cars, direct from the showroom floor. It really is remarkable how many cars we have access to. I feel a little bit guilty about it, but everything from $40,000 EVs to exotic half million dollar supercars. We also speak with the insiders who shape the automotive industry from the top CEOs and collectors to visionary designers and racing champions. Search for Bloomberg Hot Pursuit on YouTube, Apple, Spotify, or wherever you get your podcasts.
Starting point is 00:29:49 Maybe you listen while you're on your weekend drive, maybe go into cars and coffee. Listen to us. Talk about what we are driving this week. That's Bloomberg Hot Pursuit. I'm Matt Miller in New York. And I'm Hannah Elliott in Los Angeles. Subscribe today wherever you get your podcast. What separates good leaders from transformational ones?
Starting point is 00:30:09 I'm Jessica Chen, and in season two of Leading By Example, we'll sit down with executives like Grace Chen of Bertie Gray to find out. It's important to understand where you spike, but also really acknowledge where you don't and find people who can fill those gaps. Listen to leading by example, executives making an impact on the IHeart radio app, Apple Podcast, or wherever you get your podcasts. So we've been talking a lot about like failures or sort of like collapses. But the other thing that I like sort of associate with large business software is just like the user experience. is just not as good. And I think that was, you know, the sort of like the UX, I believe was part of the story with like the infamous like city error where they transmitted $900 million. They shouldn't have to some counterparties. And I think some of the users were confused by the
Starting point is 00:31:09 internal software, whether they were actually sending that or not. I heard a story from someone who worked at the VC arm once of a major bank about like the hoops that they have to go through just to share documents with each other, like PowerPoints, because they get flagged often internally. I assume that there's some sort of like regulatory issues or booking travel. Like I, you know, going to like booking.com is really easy when Tracy and I book travel for work here. It's not that bad, but like the usability of the software internally, it's not as smooth and snappy as consumer travel sites. Why is that? Like, what is, why is, why is, why is, why is, the sort of like business software internally just like is not as uh you know yeah easy to use and
Starting point is 00:31:57 sort of visually appealing as the consumer internet so this is getting better over time and we'll talk about that in a moment but broadly like your your observation is entirely accurate if you where the softwares are for the entire world you might like rank applications by their importance to the world and say okay if you were an online application on someone's phone that allows someone to share cat photos. That's like important in some sense, but probably not as important as sending a billion dollars outside of a bank. And so we would have a lot more talent and time spent on the question of, can you send a billion dollars out of a bank versus does a 13-year-old have an awesome experience when sending a cat photo? In actual fact, though, much, much more time and talent is
Starting point is 00:32:42 spent on the cap photo question than is spent on the like wiring billions of dollars out of banks question. That was a choice. It sounds silly to say, maybe we should stop choosing stupid things, but the, like, true answer is business software gets better when the, you know, people and organizations that cause that software to be built choose that the quality of that software is something that is very relevant to their interests. And so one way to, like, come to their realization is to lose a billion dollars. And then, you know, hopefully the next time you're at middle-level engineering management says we should spend a little more on maintenance. you will say, ah, yes, I agree with you. We should spend a little more on maintenance versus
Starting point is 00:33:20 taking another billion dollar charge at a time not of our choosing. Part of it is just the culture of quality coming back to these things. Part of it is also through sort of teaching the user inside of organizations that software doesn't have to be terrible because most software in the world exists inside of companies and runs business processes. That's something that is not broadly known, but like of all the lines of software in the world, most agree, exist inside of companies. And for a very long time, because most software people interacted with was at their employer, and it was generally kind of terrible. They just had an image like software is generally kind of terrible. And then the iPhone came around. Everyone has a powerful computer
Starting point is 00:34:01 in their hand for X number of hours a day. You've used applications, you've tapped three times, and, you know, interesting things happened in the world as a result of you tapping three times, and you broadly like the experience. And then you go back to work and say, wait, I've used software that doesn't suck. All the stuff I use at work sucks. Hey, IT department, hey, senior management, can you please make our expense tracking software not be terrible? And that is starting to happen. Built as a result of internal software producers advocating for change as a result of that user feedback within companies and also as a result of various startups happening to say, like, not to throw a particular expense solution under the bus, but the thing that most like old line economy companies
Starting point is 00:34:43 probably use is not a thing that people love using to book their travel. And if you, you know, use trip actions or something that is designed by a modern team with modern sort of UX affordances, it's a much nicer solution for the end user. And in some companies, end users are starting to have some level of ability to advocate for what software gets adopted, whereas previously that was made by processes that were not user-centric, where which team was better at doing at whining and dining the person in charge of the purchasing decision and not winning on the basis of product quality. In the last couple of years, even like some enterprise software is starting to win largely on the basis of product quality versus on sort of the more traditional sales motion, although goodness knows that the traditional sales motion is still very important to enterprise software companies. So I have a slightly weird question, but I'm thinking a lot about it as we have this conversation.
Starting point is 00:35:36 But it feels to me like software engineering and computer programming, it always seems to be in flux. Like if your job is a software engineer, it feels like there's always something to do. You're always trying to fix a problem or adapt a system. And I guess my question is why. You know, I fully admit my own programming experience is confined to like HTML, which I learned from that website HTML goodies in like 1998. But back then, you know, you program your website, you design it in HTML, you release it into the wild, and you're kind of done. And yet it seems with these large-scale systems that there's always change, something is always in motion, something is always in flux. Why is that?
Starting point is 00:36:23 So let me push back a tiny bit on this year. Yeah. Like, how many years have we had lawyers available? And does anyone ever go up to the lawyers and say, like, come on, guys, it's 2023. haven't you like figured out all the laws and all the contracts already? Haven't you finished the law yet? Yeah. And so why does the law change on a week-to-week basis? Well, it doesn't change per se. It's just the world is complicated. The number of commercial relationships between organizations is increasing all the time. We have increasing demands on what those relationships will do. And the job of lawyers is to adapt to that increasingly complex world every week and continue delivering like the law that society needs and the outcomes that come as a result of like competently execution. on the ability of organizations to collaborate internally with their employees and with other organizations. What's my answer for software engineering? Well, software engineers, they are, you know, working this week on a increasingly complex world where software has more leverage than it had even last week, where there are increasing demands on the world, etc, etc, etc.
Starting point is 00:37:22 Is there going to be a time where the last line of software is written? Probably not. There will never be a last bit of software written. There will never be a last contract written. There will never be a last book written because humans want more things. out of the world and we have kind of like infinite capacity for want at the margin. That was a compelling answer. Can you talk a little bit about, you know, we've been talking about banks, airlines, first startups. Can you tell us like how are the challenges for the public sector?
Starting point is 00:37:49 And I remember like Obama had a thing about like I want to like bring government websites or government tech in the modern age. It just seems like whatever problems exist for big companies seem to be even worse or more tricky when you're dealing with the public sector. I think at one point, at one point New Jersey was explicitly like begging the internet for COBOL programmers, wasn't it? In like 2020, I think that was a thing that happened.
Starting point is 00:38:14 Yeah. So full disclosure here, I did a nonprofit organization last year where a few of us in the tech industry banded together to work on the vaccine location information infrastructure for the United States because the public sector was having a great deal of difficulty creating websites. that would track where the vaccine was and route vaccine seekers to it. So have lots of thoughts here. So many issues.
Starting point is 00:38:37 Again, we did not wake up in 2023 with these issues magically. It's the result of decisions that we've collectively made as a society over many years. One decision that we've made in the United States in particular is that government pay scales are what they are. If you compare those government pay scales to what private industry pays for technologists, they are sharply out of whack. And so then, you know, if you look at GS whatever, the highest paid public sector employees in the United States make less than Google interns do self for the equilibrium. If you can get hired by Google, you know, it would require you to... Distribution of talent really does matter in this realm.
Starting point is 00:39:17 Right. And one of the things that the government has been attempting to do over the years is create things where there are groups of people who are like, officially their government employees, unofficially, think they're sort of doing an active service to the nation in places like the digital services agency, etc., etc., where they already made their money in tech. They're now on the GS whatever making a fraction of what they previously made, but are contributing software expertise to these various problems where, like, what the government needs is like some competent software written, and that requires having competent software people available in quantity. Another issue that governments have is, like,
Starting point is 00:39:54 what is the true goal you are solving for without getting too political about it in some parts of the government. Like, you know, an organization might exist as largely a jobs program, and IT modernization might sharply decrease the effectiveness of that organization at employing a large number of people to, like, repeatedly do a process that a machine could do in a faster fashion. And so sometimes, like, the powers that be within organizations are like, well, you know, I don't necessarily consider IT modernization. one of my top priorities at the moment because that would cause me to need to break faith with a number of people that I employ slash, you know, sometimes my own career trajectory as a bureaucrat is, and this is true within private industry as well, you know, there's a bit of empire building involved where you want your number of people that you manage and your budgets to go up every year and you don't want to say, okay, like I've solved my problem so I can deal with 5% as much budget next year, thank you. That is incentive incompatible. That's not great from the perspective.
Starting point is 00:40:58 of the parts of society which aren't employed by government, but which nonetheless depend on government for, you know, providing goods and services. And so this is ultimately a thing that we have to resolve through the political system on pushing back a little bit on saying like, hey, you kind of have to be good at what you do. And these days, that involves making software that is also good at what you do. I have no magic bullet for how to cause that to be a, you know,
Starting point is 00:41:22 a stunning rallying cry for political parties, but probably something that needs to get said in a lot of places for, enough decades until the message sinks in. Patrick, this has been an amazing conversation, and I feel like, A, I already want to have you back and be, like, almost each one of your answers could be, like, its own, like, have a full conversation. I have one last question. How does BitRot happen? And I mean, like, you know, even I, like, you know, you go away on vacation for two weeks. You come back to your office computer and, like, things are weird and sort of janky. They don't quite work the same. Like, what is that process? Because you would think
Starting point is 00:41:56 that just like words on a database like wouldn't rot. So what's actually going on there? So the sardonic but true answer that you have to think of at scale is like bits in a computer can literally flip by gamma rays coming from outer space that interact with like the physical manifestation of your your memory in the computer. And that's one cause of this. That's true. That does happen. That isn't like the dominant thing that happens. The the dominant thing that happens is like there exists change in the broader system that must happen on any given basis, change is a sort of risk. It does not always managed well. This gets back to that. It's like a commanding majority of systemic downtime at well managed software companies is caused
Starting point is 00:42:36 by attempts to upgrade the system that go less than optimally. It's the thing that is amenable to study. Like, BitRot happens in some cases because, you know, you had a constellation of software, et cetera, installed on your machine and installed on the other machines that your machine connected to, which was working, you might say exactly perfectly, exactly perfect because unknown in software, but like it was working right now. Something about the constellation changed as a result of a decision made about a machine that is not directly under your control, and that decision must be made at scale in the economy because software can't be, software can't allow to be static to deliver the things that we want from software as a society.
Starting point is 00:43:18 And then that changed caused some other part of the system to behave in a less, great manner, and then eventually, you know, you see the ripple effects of it in your daily life. That's the dominant way bid rot happens. It is not the bits actually getting corrupted over time. But again, a thing that does happen, and we have things in engineering to control against that. Well, Patrick, you are the perfect guest for this topic. Really appreciate you coming out on odd lot. Really appreciate you having me and would be glad to be back sometime. Definitely. Thanks so much, Patrick. That was great. I learned so many excellent new terms, like bit rot, blast radius, heroics. We mutually failed together. That one will come in handy, Joe. I'm sure.
Starting point is 00:43:57 Everything we do wrong is like, mutual failure. No, I love that, Patrick. Thank you so much. Thanks very much for having me. Tracy, I think really Patrick was like the perfect guess for that topic. Like we needed to do this for a while. And I'm glad we like did it with Patrick. Yeah. Well, I also feel like this is something that's going to keep coming up. And so we might have more opportunities from Patrick to do interesting post-mortems on various tech failures. Yeah, absolutely. I mean, there were so many interesting things. I really like some of the questions that you asked about, like, ownership of software. And I think, like, it's sort of like that really clicked to me because, like, if you're a software company and software is the main product. And, you know, I'm thinking about, like, in the manufacturing analogy, you know, it's like a Taiwan semiconductor, the sort of like institutional knowledge to build something exists within the firm and it just, you know, gets handed down, handed down. When software isn't your main product, like if you're a Southwest, like if you're a city group, etc., then you can sort of see why that process of like internal knowledge that works in manufacturing, you don't get that sort of ongoing feedback, you know, sort of distribution of knowledge in some of these large organizations.
Starting point is 00:45:26 Yeah, absolutely. And it feels like, I mean, Patrick kind of, I think he used the expression like dropping the ball in the middle of both of us, but it does seem. like that system kind of produces opportunities for, I guess, I'm trying to think how to phrase this, for people to sort of like, no one takes total responsibility for a systems failure like that, right? Because on the one hand, someone designed the software, but on the other hand, maybe it was customized by someone else. Maybe you have two different systems talking to each other and both of them mess up in some way or there's some sort of misunderstanding. It just seems like there's such a, like, gray area.
Starting point is 00:46:05 And maybe this is one of the reasons why it's so difficult to fix because you have all these different things that are sort of operating together. Well, even like, you know, like his last answer about how BitRot happens, right? Like somewhere, like all these interconnected computers, somewhere someone has to make a change because he pointed out in his answer to your question about why software is never like done or why it's never a solved problem. It's like we're always demanding more. So there will never be a time where someone like has the luxury of not making a change. And then everyone else has to interact with that software. The computer is done. We figured it out. Technology is solved. Well maybe that's
Starting point is 00:46:45 what the AI like what's it, you know, the singularity will be. Chat GPT. Yeah. It's like we've finished. We finished. It could be. But like, you know, that seems like it makes a lot of sense. Someone has to make a change because that's just how the world works. And then all these other interconnected systems, like maybe they're fine with the change, but something happens and then eventually they have to change too. And so it's just like this constant state of flux. Can I tell you my one internalized programming lesson? Tell me. So, you know, I mentioned HTML.
Starting point is 00:47:16 And then when I was in high school, part of our computer science class, we had to learn JavaScript and we had to create a program. And so I wrote this program, again, like keep in mind that this was the year 2000 or something like that. I wrote a program. It was like a digital fortune cookie and you know, you could click on it and it would give you a fortune. And then at the end of it, at the end of this module, we had to sign a contract signing over our program to our computer teacher. That's crazy. That's crazy. No. It was an extremely valuable lesson, which is all the coding that you're doing will ultimately belong to someone else. And they'll be able to monetize it. That's the downside is they'll monetize it for you and maybe you won't get it. as much. But the upside, I guess, is that you don't have to take responsibility for it. Once you write it, it goes out into the world, the computer professor owns it and he can do with it
Starting point is 00:48:07 what he will. I love that that was actually the lesson. Also, I feel like in another universe, like, you could have sold that startup for $100 million. I'm like 2010 to Facebook and it went like super viral. Like it seems like one of those things like, oh, how did this person make their fortune? Oh, they made like a digital fortune cookie. But that also like, I loved his answer about like the public sector because it's like there you really do have this problem of like salary disparities and it is pretty crazy that we sort of treat like the way we're sort of solving this problem in this country is kind of getting people to volunteer like going to work for the government in IT and tech. It's like what you do after you're like rich and you want to give
Starting point is 00:48:49 something back is like, okay, I'm going to like go work for the federal government and like try to help them, like, update their systems, which is great that people want to do that. And, like, I love that. But, like, that does not seem like a great sustainable solution to having, like, a modern government that can, like, communicate with people and provide services for people the way that they expect. No, absolutely. And it's something that we see again and again in various ways. Shall we leave it there? Let's leave it there. Okay. This has been another episode of the All Thoughts podcast. I'm Tracy Allo. You can follow me on Twitter at Tracy Allo. way. And I'm Joe Wisenthall. You can follow me on Twitter at the stalwart. Follow our guest,
Starting point is 00:49:29 Patrick McKenzie. He's on Twitter at patio 11 and check out his bits about money newsletter. Follow our producers, Carmen Rodriguez at Carmen Erman and Dash Bennett at Dashbot. And check out all of the podcasts here at Bloomberg under the handle at podcasts. And for more Odd Lots content, go to Bloomberg.com slash Odd Lots, where we post the transcripts of the episodes, Tracy and I blog, and we have a weekly newsletter that comes out every Friday. Go there, sign up, and get it to your inboxes. Thanks for listening. On April 4th, 23, around two in the morning, a man was found stabbed multiple times on a sidewalk in downtown San Francisco. Hey, who did this to you? What happened next turned the story into a political firestorm.
Starting point is 00:50:50 Reports have identified the victim as Bob Lee, the founder of Cash App. From Bloomberg Podcasts, this is Foundering, The Killing of Bob Lee, beginning April 16. What separates good leaders from transformational ones? I'm Jessica Chen, and in season two of Leading By Example, we'll sit down with executives like Grace Chen of Bertie Gray to find out. It's important to understand where you spike, but also really acknowledge where you don't and find people who can fill those gaps. Listen to leading. by example, executives making an impact on the IHeart Radio app, Apple Podcast, or wherever you get your podcasts.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.