Invest Like the Best with Patrick O'Shaughnessy - Ali Ghodsi – The Past, Present, and Future of Big Data – [Founder’s Field Guide, EP.18]

Episode Date: January 28, 2021

My Guest today is Ali Ghodsi, founder and CEO of Databricks, a data analytics platform for data scientists and developers. He's also the founder of Apache Spark, the open-source project that Databrick...s is built on, and is an accomplished researcher at UC Berkley's computer science department. Our conversation ranges from the origins of distributed computing to modern data infrastructure, how companies can leverage their massive datasets, and the transformation of Databricks through its phases of growth as a business. While technical, it's exactly the kind of conversation I like to have on this show. I hope you enjoy my conversation with Ali Ghodsi.    For the full show notes, transcript, and links to mentioned content check out https://www.joincolossus.com/episodes/4919706/ghodsi-the-past-present-and-future-of-big-data    This episode of Founder's Field Guide is sponsored by Klaviyo.  Klaviyo is the ultimate marketing platform for ecommerce. With targeted segmentation, email automation, SMS marketing, and more, Klaviyo helps you create your ideal customer experience. See why Klaviyo's trusted by more than 50,000 brands, like Living Proof, Solo Stove, and Nomad to help them grow their business. For a free trial check out https://www.klaviyo.com/founders.    This episode is also sponsored by Vanta.  Vanta has built software that makes it easier to both get and maintain your SOC 2 report, at a fraction of the normal cost. Founders Field Guide listeners can redeem a $1k off coupon at vanta.com/patrick.    Founder's Field Guide is a property of Colossus Inc. For more episodes of Founder's Field Guide go to https://www.joincolossus.com/episodes.    Stay up to date on all our podcasts by signing up to Colossus Weekly, our quick dive every Sunday highlighting the top business and investing concepts from our podcasts and the best of what we read that week.  Sign up here - https://www.joincolossus.com/newsletter. Follow Patrick on Twitter at @patrick_oshag Follow Colossus on Twitter at @JoinColossus   Show Notes [00:02:48] – [First question] – What is Databricks [00:03:34] – History of distributed computing [00:05:35] – Hardware that made this all possible [00:07:20] – Early challenges in building out these systems [00:09:43] – What has made networking technology better [00:10:35] – Doing something in storage vs with memory [00:11:45] – Origins of Hadoop [00:12:42] – Use cases of distributed data in 2010 that weren’t possible in 2000 [00:13:35] – Origins of Spark [00:15:25] – Early Spark and then the transformation into Databricks [00:16:50] – Early uses cases [00:17:37] – Their relationship to the open-source project [00:21:07] – What customers need in order to work with Databricks [00:23:11] – Their customer interaction [00:26:27] – How they think about making investments [00:28:24] – Their competitive advantage [00:30:13] – Other companies in moving the needle in building distributed computing industry [00:32:10] – Walls that need to be broken down today [00:34:02] – Best practices for companies when it comes to their data             [00:34:13] – Jeff Lawson Podcast Episode [00:38:47] – Lessons being a CEO [00:39:53] – Working at the University of Berkeley’s AMPLab [00:41:56] – What excites him about the future [00:43:29] – Kindest thing anyone has done for him

Transcript
Discussion (0)
Starting point is 00:00:00 This episode of Founders Field Guide is sponsored by Clavio. Want to deliver marketing moments that last a lifetime, Clavio is the ultimate marketing platform for e-commerce. With targeted segmentation, email automation, SMS marketing, and more, Clavio helps you create your ideal customer experience. See why more than 50,000 brands like Living Proof, solo stove, and Nomad, trust Clavio to grow their business. Keep your customers coming back.
Starting point is 00:00:23 Get a free trial at clavio.com slash founders. That's K-L-A-V-Y-O.com. slash founders. Stay tuned at the end of the episode where I talked to Clavio customer Nomad on their origin story and how they work with Clavio. This episode is also brought to you by Vanta. Does your startup need a SOC2 report to close big deals? Or do you already have a SOC2 report and want to make it easier to maintain? Vanta has built software that makes it easier to both get and renew your sock two. With Vantta's continuous monitoring solution, you avoid hosting auditors on site and taking hundreds of screenshots to prove that you are compliant so you can focus on building your
Starting point is 00:00:59 business. Vanta partners with audit firms who file your SOC2 report directly inside of Vanta at a fraction of the normal cost. Hundreds of companies, including more than 100 Y Combinator businesses, are leveraging Vantas today to streamline compliance and focus on building their businesses. Founders Field Guide listeners can redeem a $1,000 off coupon at Vanta.com forward slash Patrick. That's vanta.com forward slash Patrick. Hello and welcome, everyone. I'm Patrick O'Shaughnessy, and this is Founders Field Guide. Founders Field Guide is a series of conversations with founders, CEOs, and operators building great businesses. I believe we are all builders in our own way, and this series is dedicated to stories and lessons from builders of all types. You can find more episodes at investorfieldguide.com.
Starting point is 00:01:46 Patrick O'Shaughnessy is the CEO of O'Shaughnessy Asset Management. All opinions expressed by Patrick and podcast guests are solely their own opinions and do not reflect the opinion of O'Shaunacy Asset Management. This podcast is for informational purposes only and should not be relied upon as a basis for investment decisions. Clients of Oshonosi asset management may maintain positions and the securities discussed in this podcast. My guest today is Ali Goetze, founder and CEO of Databricks, a data analytics platform for data scientists and developers. He's also the founder of Apache Spark, the open source project that Databricks is built on, and is an accomplished researcher at UC Berkeley's Computer Science Department. Our conversation ranges from the origins of distributed computing to modern data infrastructure, how companies can leverage their massive data sets, and the transformation of data bricks
Starting point is 00:02:36 through its phases of growth as a business. While technical, it's exactly the kind of conversation I like to have on this show. I hope you enjoy my great conversation with Ollie Goatsy. So, Lee, I'd love to start our conversation at the end with what data bricks is today, to level set for the audience, exactly what you do, what your focus is, and what the business does for customers. Could you just walk us through as we sit here at the end of 2020, what the company looks like and the service or problem it solves for customers?
Starting point is 00:03:04 We're a 70-year-old company. We have about 1,700 employees. And we help enterprises take massive amounts of data and do machine learning, AI, and data send us on that data. Most enterprises, they've seen how Silicon Valley forward tech companies have used data in a really strategic way to disrupt industries. They want to do the same thing, but they don't have thousands of engineers. that can help them build data platform custom for their use case,
Starting point is 00:03:31 we've built that and we enable them to do that. I would love to go all the way back and sort of tell the history of distributed computing because everybody will have heard the term big data. This was a really popular term, I don't know, five, seven years ago. And I think that concept, that term, the fact that it was being talked about in normal business circles, was the result of progress in the world of distributed compute and storage. I'd love you to rewind however far back you think is appropriate to go.
Starting point is 00:03:56 maybe it's back to the 2006 Yahoo Days, tell the modern history of distributing computing what it means and why it's so interesting and important. I think what happened is that around 2000, we hit this wall, we call it Moore's Wall, because they didn't figure out how to make computers faster. So everything started moving into these data centers, a new computer and it was a new data center. In these data centers where you had hundreds or thousands of machines, people started collecting more and more data. And the reason for this was multiple. One was, the price of storage kept going down. So it became just cheaper and cheaper to store all this massive amounts of data and no one wanted to throw it away. And they had heard that there were
Starting point is 00:04:35 some for tech companies like Google that had gotten a lot of value out of that data. So they wanted to do the same thing. Secondly, more and more people were connected on the internet. There are these sites that had billions of users attached to it that were coming, visiting these sites. There was this aspiration that we collect all this data. Maybe we can do great things with it. I think around 2005, 2006, people still didn't know. I mean, the Ford tech companies knew. They had typically an ad business. They were collecting the data. They were optimizing how to show ads to people. The rest of the enterprises didn't really know what to do with it. So this big data revolution sort of entered its first phase, which is, let's collect all this data. They're amazing things
Starting point is 00:05:12 we can do it once we get there. It's cheap. Why not do it? So that's kind of what started around that time. And people started collecting these things into data lakes, massive, massive of datasets. And back then, the measurement of success was, how much data do you have? We were very successful. We have one petabyte to data. We were even more success. Our data lake has grown from one petabyte to two petabytes. This is amazing. So it was like kind of the first generation. What exactly was going on almost down to the hardware that was revolutionary back in the early to mid-2000s that made some of this possible? So people are used to having their own data on their own computer. Everyone's familiar with the term cloud. But I think early on, everyone would think
Starting point is 00:05:51 of cloud as someone else's computer. A computer is somewhere else, not under my desk, somewhere else. What literally was happening from the hardware and software standpoint in the early to mid-2000s to make some of this possible? Yeah, I mean, early 2000, you would buy a big supercomputer and solve a lot of our problems. In fact, I remember Berkeley, we visited Twitter in the very, very early days, just when they had one giant machine, we processed all the tweets. I don't remember how much memory it had, but it was some gigantic amount. And we're like, wow, how did you get a machine with that much memory in it? That was the way to solve problem. It would get a very expensive supercomputer. It would do all your computations for you.
Starting point is 00:06:26 But as we hit this Moore's Wall, and it couldn't scale these computers. Around 2005, the CPU speed stagnated. Computers were not getting any faster anymore. Computers are basically in 3 gigahertz since then. So that's kind of started meaning that we probably have to distribute. Our needs are not going down. The amount of data we have is not going down. The number of computing processing capacity we need is just increasing exponential. We probably need more machines. So that's when things started distributing out into data centers. It soon then became infeasible for every organization to manage thousands of thousands
Starting point is 00:06:58 of machines on their own on-prem. So the cloud revolution kind of took off where I said, hey, this is a utility. Why don't we manage all these thousands of machines for you in the cloud? You can just use them. This also accelerated things because now anyone could pretty much with a credit card, to start renting thousands of machines, this new computer in the cloud, and start storing things on it very cheaply and get started. What was the harder challenge back then between hardware and the software needed to coordinate putting stuff into those systems and then accessing and pulling it out
Starting point is 00:07:29 in a timely fashion? What was the most innovation that happened to make that possible? Well, actually, a few things changed. In the very early days, the problem was there were so much data and the networks were not fast enough. So whether it's in the cloud or whether it was your own data center that you had, the big innovation of MapReduce, the big thing if you go back to that time was we have to move the computations close to the data. You specify what you want to do with your data, but then you had this thing called MapReduce, which was really clever about let's move the code that runs on this data close to the data because there's so much data we can't transfer it over the network. If we try to transfer it over the network, the whole network will collapse.
Starting point is 00:08:07 That was the name of the game back then. And there was a lot of research on how do we avoid that. Turned up, you don't actually have to need to do that. anymore. Can you say a little bit more about that? So I just want to make sure I understand. So let's make this stupid simple. I've got a terabyte of data sitting somewhere on actual hardware and whatever it is, it's customer data or something. If I want to do some sort of compute on top of that data, run a query or produce some output, the network was the limiting factor. So just getting the data from there somewhere else to do the compute was hard. So some innovation was just do the compute in the same data center or something. Am I getting that right? Yeah. Let's take an example. Let's say
Starting point is 00:08:42 you have a petabyte of data on how people have clicked on a website. Since you can't buy a supercomputer anymore, you've probably distributed that over many machines. So you have 100 machines, they have hard drives, and you've stored it on the hard drives of those machines, the dataset, the petabyte of how people have clicked on your website. But now you want to compute something. Simple example, we just want to see what's the average number of clicks on a particular link. 20 years ago would be just write the code that takes all those numbers and averages them and then go over all of the data set. But here to do that, we would have to move all the data from these hundreds of machines to the machine that's doing the computation. And that would just crash the network.
Starting point is 00:09:19 In fact, we did that at UC Berkeley. We were developing the first versions of Spark. We crashed parts of that network. They called us up and said, what are you doing? We said, well, we're transferring all this data from all these machines to every other machine. So that was the name of the game in the early days. How do we avoid that? How do we actually just move pieces of the computation that computes the average, just to all these hundred machines, let them compute it locally, and then send some kind of aggregate. So you mentioned that originally that was a unique solve, but then today we no longer have to do it that way. So what's changed? Networking technology, unlike CPUs, have just gone faster and faster and faster. And they've also come up with techniques.
Starting point is 00:09:58 This is actually research that started at UCSD. They figured out how to configure the networks in these data centers in a way, such that any two machines can completely at full speed communicate with each other. You can no longer collapse the network with these heavy computations. And in some sense, these days, the network has been completely virtualized, and it's no longer in your way. So we no longer need to actually move the code close to the data. And that happened around 2009, 2010, 2011. And today, virtually every data center, every public cloud provides you these really, really fast networks that just won't get in your way. Can you describe the difference between doing something from storage like you first described
Starting point is 00:10:40 and doing it in memory? I think that's an important difference to describe for the audience. And then I want to talk about Hadoop, its origins, how that bleeds into the origins of Spark. Back then, memory was expensive. Again, that's something that's changed a lot. Memory's gotten way cheaper. Storage got way cheaper.
Starting point is 00:10:57 The CPUs are kind of the things that have stagnated in the last couple of decades. So with these big data sets, you would have to do the processing mostly on disk. So the data sits on disk, you load it in, you do some computation on it, you write it back to disk. Because you don't have enough memory to store all of it in memory. I think that some people, memory and hard disk might sound like the same thing. Can you just describe literally what the difference between those two things means? On the hard drive, you have persistent storage. So if you turn off the computer, the data remains there.
Starting point is 00:11:27 Memory, when you reboot, it goes away. But the difference is it's much, much, much faster. So it's orders of magnitude faster to access data that's in memory. And you typically have much less memory than you have disk. Perfect. So can you describe the origins of Hadoop and why if you're in the data world, everyone knows that this is a key milestone in the timeline of doing work with big datasets.
Starting point is 00:11:48 Can you describe what Hadoop did and why that was so interesting and important? Yeah. Hadoop enabled you to process larger data sets than ever before in parallel on hundreds or thousands of machines. Think of it. Early 2000s, this new machine or computer arrives, which is the data center with hundreds or thousands of machines in it. The kind of operating system for it that was developed, the first operating system was Hadoop and MapReduce, which was a way in which you could basically now process any amount of data you wanted to process. Unfortunately, though, it was a little bit complicated, just like early computers and
Starting point is 00:12:22 early operating systems. All the computations would have to be described in terms of just two functions, one called Map and one called Reduce, which made it very cumbersome and complicated to write programs for this. But it was amazing. Now you could crawl the whole worldwide web, process it, and do computations on it if you just were sophisticated enough. What were some examples of interesting use cases of this new distributed compute technology that wouldn't have been possible, maybe that were happening around 2010 or something like that, that wouldn't have been possible inside 2000? The first use cases of this were largely, it started with crawling the web and building indices for all the stuff that you crawl, all this large data you have.
Starting point is 00:13:03 Other use cases were click logs, logs of how people are clicking on these websites. You have billions of users on the website, and they're clicking around. How do you collect all of that that makes sense of it so that you can actually start doing more advanced things with it? For instance, maybe you can figure out what ads to show someone or what's more interesting to them. So you could now, for the first time, do these really, really large-scale computations, which you couldn't do before.
Starting point is 00:13:25 It was a major breakthrough, but it was very, very hard. those programs. It was not normal programmers. So now having laid a lot of great groundwork, hopefully people will have some context. We can get to the origins of Spark itself, the open source project around which Databricks is built. What was the very beginning of Spark? Who designed it? What was the intention? How did it get going? The real story is that there was actually a group in the lab. So we were sitting together at New Berkeley and the people that built these computer systems, which was our group and the people who were just doing pure machine learning. They were more sort of math background.
Starting point is 00:13:56 We're sitting right next to each other because the idea at Berkeley is all these people should work together and come up with really great things. And they were trying to do recommendation. So in particular, they were participating in the Netflix competition. And the Netflix competition was a competition in which Netflix showed you how people had rated movies. And you had to come up with a machine learning predictive model, which would recommend movies to other people based on their preferences and how movies have been rated in the past. and whoever got the most number of people clicking on their recommendation would win the contest and win, I think, half a million dollars, a million dollars. The machine learning team was saying that this was really, really hard to do for them.
Starting point is 00:14:33 And we started talking more closely than what's going on. And they said, well, we're using this Hadoop thing, and it's just so slow. And we were doing many iterations over the data because these machine learning algorithms are highly iterative in nature. You go over the data again and again and you keep refining it until you get good enough. And so every time we have to load all the stuff from disk, put it in memory, do some iteration on it, write it back to disk. It's just too slow.
Starting point is 00:14:56 It's taking too long. Can you guys help us? That's what it started with. And we said, yeah, memory is getting pretty cheap. Why don't we just figure out a way where we can just load all of this data into the memory? And then we can just do the iteration super fast in memory because that's faster than this. That was the origin of Spark. So in some sense, it was a tool created to be able to do really fast machine learning
Starting point is 00:15:15 to win the Netflix competition. Fantastic. And what were the first year or two of Spark like? what was happening in terms of the growth of its participants, its use cases, and how does that then translate into the creation of Databricks? You've got Spark as an Apache open source project that in many ways anyone can access. And then Databricks as a corporate sponsor or entity that sits next to it, on top of it, around it. I don't know how you would describe it. But describe those early days and same question around Databricks. Why did Databricks come to be in the same way
Starting point is 00:15:48 that Spark came to be? Popularity of Spark took a very long time. to take off. So it was started in 2009. People were excited, and as an academic project, he had a lot of impact and a lot of excitement on it. But out in the sort of, if you went to industry, no one knew about it. It took many years, 2009, 10, 11, 12, not much was happening. And we were really trying to get the industry to adopt it. We went to these companies that were at the time, they had this Hadoop technology. And we told them, please take this. This is 100 times faster. It's much, much, much easier to use. And it supports machine learning. It's a supports real-time computations.
Starting point is 00:16:23 And they just looked at it and said, no, this is just academic project. What if these students quit to start their own company or something? Then we'd be left here with the software. They just ignored us. 2012, we kind of had it. And it was like, if we want people to adopt this, we probably have to start a company around it. Because it's not going to work sitting here at UC Berkeley. They're not going to pick it up.
Starting point is 00:16:44 We just don't seem to be getting the attention that this project deserves. What was the kernel early customer use case that you latched on to? I'm always interested with new businesses like this, what the first commercial engagements are with people that need the service. What was the origin story there? Yeah, I mean, there was lots of lots of use cases of industry. They wanted to do machine learning, basically AI use cases. I think the exact first use case, Databricks was a company that's trying to understand
Starting point is 00:17:10 how video screens were being actually downloaded on the Internet and how they could increase the quality of those. Obviously, if you do that on the planet, we're talking massive amounts of data, massive amounts of viewers, inequality deteriorates and improves over time, being able to in real time making quick decisions on that. Maybe we should flow the traffic over there instead, and being able to do that with AI and predictive technologies. That was the first use case that our first customer had at Databricks. Can you describe how a business like yours relates to the open source project? How do you think about how those two things work together, interrelate, hand off to one another?
Starting point is 00:17:48 it's a very unique and increasingly common way that a lot of developer-facing technology projects and businesses are built is to have an open-source component. How do you think about the benefit of that piece of the business? We are a SaaS business. So we are a cloud company. So we manage and run software for you in the cloud. These open-source projects, they're really on-prem products. You download them. They have a particular version. They're not cloud software. Cloud software, like Google search doesn't have a version. When you go and search on Google, I don't ask you which version you search on. It's just there.
Starting point is 00:18:21 It just works. They keep upgrading it and cloud software. We do the same thing. We run Spark and lots of other things these days. Spark is a small portion of what we do today at Databricks in the cloud on behalf of our customers and we just automate that away. But that's not what we were doing the first two, three years of Databics. The first two, three years, we just wanted Spark to take off.
Starting point is 00:18:39 It seemed just impossible because it was all the rage about Hadoop. and people were just keep talking about Hadoop and how it's awesome. I remember Cloudera had just gone a $700 million investment from Intel, so it was one of the largest investments that year, I think the second after Uber's investment. So we were just trying to get this to take off. And it seemed it was impossible. Spark doesn't work if you don't have enough memory, which was not true.
Starting point is 00:19:02 Spark is just good for some machine use cases, not others. Or Spark only works if you have real-time computations, but it doesn't work on the others. So it just seemed impossible to educate them. markets. None of this is true. So that was our first struggle the first two, three years. As you think about it today, how much of Databricks the business is in your mind related to the open source world? I'd love to hear what the lineup is outside of Spark, the other services that you manage in the cloud on behalf of customers to sort of lay out the different methods or
Starting point is 00:19:33 products or services. But just as you think about it today, how key is open source, if at all? and do you think that will change in the future? I think open source is critical to enterprises. I think enterprises don't want to lock themselves in to proprietary software. They've gotten burned since the 80s. They've bet on vendors that are really, really good, who have great innovations, and they give them all their data, and they lock it in into those proprietary formats. And then those vendors, because of that lock-in or that moat, they become complacent after a while.
Starting point is 00:20:05 They don't need to innovate anymore. eventually the founders or the original folks move on. And at that point, it's just bloated old software, which is very costly. And they can just keep increasing the price. And you have to pay more for it because it's so hard for you to move off of it because you're locked in. So I think enterprises, if they have the choice, they prefer open source because it avoids that lock-in. So it's critical to what we do. So pretty much every element of the Databricks platform where there would be a lock-in,
Starting point is 00:20:32 we've opened it up as an open-source project. So today there's Spark to access. all your data sets and get all your data. There is Delta, which is the key project that makes your data really high quality and really performant for downstream use cases. There is a project called MLFlow, which is really how you operationalize end-to-end machine learning. Finally, there's a project called Redash, which is how you deal with all your visualizations and dashboards and things of that nature. So that's also at its core open source. We don't want anyone to get locked into us. We want them to pick us because we're providing so much value in our software.
Starting point is 00:21:06 that's in the cloud. So you've described sort of an interesting now platform, a set of tools that machine learning researchers could use to do their job. Let's boil this down to like an individual research team or something. They come to you and let's make it as simple as possible. They've got some big data set. They want to use that data set to build a prediction and that prediction is then used somehow in their business or for some use case. Just something really straightforward. Something like Netflix probably is a good example for people to think about this. Lots of data. They want to use that data to build a predictive model to do something interesting. What do they need to show up with of their own? What do they need to bring to the table to then start
Starting point is 00:21:41 working with Databricks? Is it just an initial data set? How do you think about what a team or a person or a researcher shows up with that then makes data bricks and its platform powerful? I think they need data sets, so they need actual data, which luckily almost every enterprise on the planet has been collecting since mid-2000s. And they need to have a clear understanding of what are the use cases that they really want to deploy or actually realize. that requires that they understand their business. What's the most important project to have impact on the business? What's the biggest business value to provide? Is it a prediction for something? So for a company like Shell, it's being able to predict its equipment breaking down in advance.
Starting point is 00:22:21 If they can do that, then they can replace those parks in advance. That saves them hundreds of millions of dollars. And it's actually better for the nature. It's better for the employees, it's an environment and so. For a company like Comcast, it's a use case where they have a remote control and it has a voice button. They press that and it can speak into it and say, hey, what's the weather today? And then that actualizes. For a company like a pharma company like Regeneron, the use case is finding that particular gene that is responsible for a disease. For instance, they found the gene that's responsible for chronic liver disease and
Starting point is 00:22:53 data breaks. And then they can do drug testing and develop drugs to cure those diseases. So it depends on. So understanding that business use case shouldn't be undervalued. If they're just showing up saying, hey, we have a bunch of data and we want to do some cool stuff, some AI, then that's usually what. when we say, well, what is it really you want to do? What are the use cases that really are pertinent for your business? How does the business itself work? One of these teams shows up,
Starting point is 00:23:14 Comcast shows up, it's a team of researchers. Are they self-serving, just starting to use the platform without really interacting with somebody at Databricks? Is it a higher touch? Is there a service model? I'm always interested by how these engagements work, and I'm sure it's different based on size, but what are the ways in which Databricks engages with its customers? I think it's different from many other sort of models. Because of this open-source nature. And because there is a completely free open version of this called Community Edition, there are hundreds of thousands of data scientists that come and use that every day. And they can just swipe a credit card, not talk to us. And there's a free version that they can
Starting point is 00:23:48 use Community Edition. What then happens is that typically our sales teams start engaging with those that have a lot of usage on the platform. And that's when we start getting more strategic. A really strategic project that can help a company save $100 million, dollars, that's typically not a project that some engineer swipes a credit card, and then soon they are saving $100 million for their business and paying Databix millions of dollars for that. That usually requires investment from the leadership of that company, go in and say, we want to do this, and we're going to put resources around it. So that's when we get engaged on the sales side.
Starting point is 00:24:22 And depending on if they need help, we have actually professional services to help them with it. If they need augmented services, there are SIs that we work with that can come in and help them with that. Then there's, of course, the platform that I can explain a lot of bit how they actually use it. Yeah, please do. I'd love to roll right into the platform. Typically, they have datasets. Getting access to that data typically happens with Spark. Spark is the technology that enables getting it all loaded into Lake, where they can actually start processing it. The next step is to make sure that that data actually has high quality and a structure and that it's organized in a way so that you can access it really fast. For that, the biggest innovation in the project that Databricks is
Starting point is 00:25:00 spending most of its resources on, it's the Delta project, the Delta Lake project. It's also an open source project. That's where you organize your data so that you just don't have a data swamp where you just dumped everything into a lake. So now you have your data structured, and it's really fast to access it. You might have even set it up in Delta so that it's in a streaming real-time fashion. So the data is getting updated in a real-time fashion. Now we start looking at, depending on your use case, if you want to do machine learning, you start actually exploring building machine learning models. So you start actually looking at the data and looking at using various machine learning models to do predictions.
Starting point is 00:25:34 It's a highly iterative task to come up with the machine learning model. Machine learning folks usually don't just sit there and then come up with the predictive model and they're done. They actually typically have to iterate on it hundreds of times. So they try a prediction. Maybe the accuracy is not that good. They go back to the data. They augment it with more data.
Starting point is 00:25:51 Maybe they buy some data sets to augment it to see if there is more signal they can get there. And then try it out again. So they iterate on this a lot. The open source project called ML Flow helps. them be productive when they do that. It tracks all the models they created, it helps them govern it, set access control on it. So you can actually, with ML flow, do what's called serving, which is the final product where you're actually using it. You're using it in the web page, you're using it on that remote control or calling in and saying, right now, give me that
Starting point is 00:26:20 prediction that I need. Is this machine going to break down or not? That's the platform end-to-end what it looks like. So it's really infrastructure for machine learning researchers. With that in mind and very general purpose. I love the three different examples you gave of Shell Regeneron Comcast. You know, everyone knows those businesses, but very different use cases or products or outcomes, but all using the same sort of data-driven research model somewhere in the middle to produce that product or service. With that in mind, how do you, as a company, think about making investments yourself that will earn good or high returns on capital to the future, especially because I imagine a lot of this is engineering, research,
Starting point is 00:26:58 development, et cetera, how do you think about as a capital allocator making investments internally? Well, first, we think that this ham is going to be absolutely gigantic. I gave you those three examples on purpose, picked completely polar opposite use cases. And I could just go on all day with use cases like that. You can find them in every industry. It's not just one or two interesting use cases for company. If you look at the original companies that started doing this, deploying these machine learning techniques, they have hundreds or thousands of use cases internally. A company like Uber predicts the price for you. It predicts the route for you.
Starting point is 00:27:32 It tells you when the food's going to be ready. Put more people in the same route so it can do carpooling. It's not just one data science or a research team. You actually enable your whole organization to be data-driven. Then you can actually compete in a different way and you can disrupt the industry that you're in. Based on that, we're looking at how can we actually enable the whole organization? How can democratize the AI? Any investments we can do into the platform that enables more people in the organization
Starting point is 00:27:57 to be able to use these techniques and be data driven. That's we think strategic because we think eventually 10 years, 15 years from now, every company will have every BU be using data and AI in a strategic way. Obviously, we're not there yet on the planet. We're going through the charity life cycles. But things to make it simpler, broaden the time for that. That's where we're investing our dollars. That's where we're going.
Starting point is 00:28:23 How do you think about your competitive advantage, versus other data companies with a TAM as big as this is already and is likely to become into the future. Big markets usually draw a lot of really smart, talented entrepreneurs and companies and use cases. Do you think about that much? Do you think about other companies going after a similar market? And if so, how you structure your business to sort of have an accumulating advantage as you get bigger versus competitors? Yeah. First thing is make sure that you have an innovative DNA in the company, that's remained in Database. So continuing cannibalizing yourself, continuing to figure out a better way to do things, even if you face innovator's dilemma
Starting point is 00:29:05 and in my current your existing revenue base, that's okay. Cannibalize it and come up with the next thing. That's why Database is a very small portion of what we do today is part. Today it happens to be Delta. I'm sure tomorrow there will be something else. Building in that innovation DNA into the company is essential and enabling the company take risks and be able to actually continuing to innovate, that's essential to compete. And one thing that really helps us there is open source. We don't have a lock-in mode that we can just sit there and say, well, we now have everybody's data.
Starting point is 00:29:34 We've locked it in. They can't leave us, close the doors, fire all the researchers. We'll have to continue innovating. That's one thing. But the second thing is the fact that we actually own the whole pipeline from the data coming in all the way to the use cases. So it's sort of vertically integrated that way. That helps us a lot because you get benefits when you're,
Starting point is 00:29:53 all the way from the data in Jess, all the way to the productionization. A lot of companies, they actually cut out just one sliver of that pipeline. They're not end to end. There's lots of room for optimizations that you can do if you own it end to end, which is what we have. That itself is also an advantage. Now, others could do that too. It's just other companies haven't done it for some reason.
Starting point is 00:30:13 What do you think the most interesting other companies outside of Databricks are in the modern data stack? As you think about the growing market of people that have a lot of data, want to use that data to do something good or productive for their business. I don't think anyone would argue that that's happening or is going to happen a lot more in the future. What other types of companies or even individual companies do you think are really critical and moving the needle on what's possible right now? I look at startups.
Starting point is 00:30:39 They're typically a series A or we startups that are interesting. The big company names that everybody knows about, I find them pretty uninteresting. I mean, they've had some innovation that's typically 10 years old. They're milking it with a go-to-market machine that's broadening, getting that tech out to bigger and bigger market. But the core innovations aren't that interesting to me. The really interesting ones that you find in startups that are working on make machine learning or AI more productive or simpler if you use, how to do that in real time. There's also a lot of interesting things happening on the visualization front. How do you build data apps that can actually have built-in visualizations that you can interact with?
Starting point is 00:31:12 There's lots of startups in that space too. Well, say a bit more about that second category. What do you mean by visualizations and why is that interesting? What might that enable? You have data-driven apps where you're actually interacting with the data in a UI. It's not classic BI. It's not just a normal application. This new category that you're seeing emerging in the field, that I think is very interesting.
Starting point is 00:31:33 Is there an example of a company that people could go to just to understand kind of what you mean by this? I'm still struggling to understand the middle interacting correctly with data visualizations. I mean, there are these technologies like bouquet, like Dash, where you can actually build data-driven applications, where your application interacts with the data and visualizes it. And it's much, much more rich in the types of things you can do than classic BI dashboards, which typically are histograms and, you know, drop downs and various types of graphs that you just click on. Here you can actually leverage machine learning under the hood if you wanted to. It's very interesting. What do you think is the biggest bottleneck right now in the technology or the development of computing, broadly speaking. So if distributed systems and compute and storage
Starting point is 00:32:22 were like a huge breakthrough that got us through Moore's wall. What are the walls today? What walls are we trying to break through? The biggest wall is between data and AI. And you see it both organizationally inside enterprises, but you also see it in the tech and just the preferences people have. So the teams that are responsible for managing all these massive data sets, typically it's an IT department. They have to make sure that the data is secure. They have to make sure that that data is reliable. They're going to own it for the next 10 years. It needs to be compliant, so they're very conservative in nature. And then you have the folks that are in line of business that are close to the use cases. They're the ones that are coming up with these
Starting point is 00:32:58 amazing AI use cases that is going to transform their business. But they can't actually succeed without the data. There are companies that focus on these users, data management companies that are awesome at data management and processing your data, making it fast, reliable, govern and so on. But they have zero AI built into that, to their technology. And on the other hand, in the line of business, you have companies that are focusing on just AI and machine learning, but they don't have actually any of that governance, data capabilities, data management, reliability, security capabilities that you need. Even the preferences of the practitioners different.
Starting point is 00:33:32 Typically, you use something like Java on the IT side. That's what they like. They want it to be production ready. On the other side, on the line of business, it's technologies like Python and so on. So there is this big divide. And typically also these two teams don't even report up to the same person in the organization. So this is slowing companies down. And this is not how it was done in the forward tech companies 10 years ago.
Starting point is 00:33:53 They were organized differently and they were using one tech stack for both. I think that's the biggest barrier that needs to be broken down to accelerate us getting value out of the data. When we had a conversation with Jeff Lawson at Twilio, one of the things that stood out was how Twilio and other companies helped change developers from sort of this back of house function to a very front of house, incredibly important function. in a business where now developers are impossible to get. It's just a really in-demand talent. I'm curious what your thoughts are on whether or not the sort of data world is going through that same transition where you see more chief data officers, more attention given to this part of any given business, and sort of what you see as the best practices,
Starting point is 00:34:37 whether it's a Comcast or a Shell or whatever the examples are, like the traditional firms that are doing this well, what do those firms share in common from your perspective? because I think that sort of serves as advice for those playing catch-up that want to use their data productively. If you look at it, there was a lot of collecting just data and how much data do we have on the management and sizing that up, one petabyte, two petabytes. Eventually, the business started saying, what value are we getting out of it? Things have changed a lot now. I see now that every enterprise I talk to, there is this awareness from the top that this is absolutely essential. We have to invest in this.
Starting point is 00:35:13 We have to get our AI and data strategy right. If we don't, just around the corner, there might be a startup, probably out of Silicon Valley, that's completely tech-driven. It has thousands of engineers. And they're just going to disrupt our business. We're going to be put out of business, just the way Uber did it with the cab medallions, just the way Airbnb did it with the hotels, just the way Netflix did it with Blockbuster and so on and so forth.
Starting point is 00:35:34 So we're next in line. We've got to do something. I mean, there's this urgency. That's great. But a lot of them struggle on how to do it. The ones that do it well, I think, have a few ingredients and a few things in common. One is they try to consolidate it under one leader or at least have some kind of center of excellence where you can get all these folks working together. The wrong way to do it is to
Starting point is 00:35:55 completely separate line of business and IT. IT is responsible for the data and line of business is responsible for these projects, these AI projects. That inevitably gets stuck in politics between the two departments that have different goals that the business is asking them from. Security, reliability on one hand side and the other is, you know, business impact and business, business value. So getting those two together, if you have a chief data officer, that's great. The title isn't important. If you have some org structure where they can actually work toward the same goal, that's really, really important.
Starting point is 00:36:26 So we see that in all of the companies that are doing that well. And of course, the leader of that organization is an important person. So picking the right leader, agents that understands how to sort of bring the organization along because these enterprises have 30, 40, 50 years of history. So how do you transform them and make them sort of? of adopt this new kind of data-native approach to things. That's going to be critical. Second, typically leveraging open source to avoid lock-in. Most of the development in the space is happening in open source. There are new machine learning models every week released by the universities.
Starting point is 00:37:00 So that's really, really critical. Three, building with AI in mind from the beginning, don't just collect a lot of data and brag about how much data you've collected and figure out later. Later we'll figure out these AI projects. You kind of have to do it from the get-go. And finally, to be able to move fast and be agile, bet on the cloud. Running your own on-prem data centers at this point, it's probably obvious. It's going to be slow. You're going to be stuck with old equipment. I think most people have realized that now with the pandemic.
Starting point is 00:37:28 But more importantly, having a multi-cloud strategy. Because they're right now three, four cloud vendors, and they're all big companies with a lot of capital. They're not going to go away. They have different strength and weaknesses. Make sure that you don't put all your eggs in one basket and have a multi-cloud strategy. Those are some of the ingredients we see really work well for the companies that are successful with data and AI.
Starting point is 00:37:47 You've been the CEO of the business now for a while. How would you sum up the major lessons that you've taken away on what it means to be a good CEO? Extra points if it's lessons learned by doing something the wrong way and figuring out how to build a large and an important company. First of all, the company goes through different phases. So you need to be good at different things at the different phases. And actually, sometimes the things that you had to be good at that one phase actually
Starting point is 00:38:10 hurts you now. So you have to kind of transform at each phase. First phase was really a product market fit phase, understanding really what enterprises need and just making sure that you're building the right thing for them. So that was the first phase. The second phase was, okay, we figured it out. This is what they need.
Starting point is 00:38:26 How do we scale the machine? So bringing in the pros that really could scale the machinery. And this third phase is, I would say, an optimization phase where you now have so many people in the organization with thousands of employees that you have to make sure that all the processes are smooth and they just work and you can just drop in new employees that come in every day and they'll become bricksters as we call them and they'll do things the way bricksters do things here they'll come up with new innovations and they do things the data bricks way
Starting point is 00:38:54 that requires much much more process if you actually compare the first phase in the third phase things you need to be good at in the first phase actually hurt you at this phase first phase we said be a co-founder, be an owner. If you see trash in the kitchen, even if you didn't put it there, you clean it up. It's your company. You're an owner. Actually, this phase, I would say, don't pick it up. Let's go back and figure out how did it end up there in the first place, and let's put in processes to make sure that that never happens again. It's just a little bit of a different sort of mindset you have to do at scale and a trade-offs. For me, probably the biggest lesson is how important trust is and building trust with leaders. Nothing you can do overnight.
Starting point is 00:39:37 It ultimately, especially as the organization grows, finding leaders that you can trust and you can work together to drive the company because you can't do this alone. You need other really strong leaders. That's probably the biggest lesson of how important that is and how time consuming that is and how difficult that is. To go back all the way back to where we started the origins of all this, the environment of Berkeley, not just Berkeley, but of research centers in the U.S. and around the world. AMP Lab specifically comes to mind. Maybe just say a few words about how that all works. And I'm sure most people haven't heard of AMP, for example, that are listening.
Starting point is 00:40:12 What kind of work happens there? Why is that such an important source of progress and innovation in the world of technology? Berkeley had a special way of doing things. I would give a lot of that credit to one particular professor who is now retired. his name is Dave Patterson, who's also now a touring award winner, which is kind of the Nobel Prize for us in tech, the closest we can come to it. He had this mentality that we need to really work together and collaborate. As evidence of that, at some point he said, let's have all the students and professors sit together and actually have students from different backgrounds work
Starting point is 00:40:46 together. Administration said, no, we don't have room for that. They're not enough rooms. and he said, hey, so me and my fellow professors, we'll just give up our rooms in. We don't need a room. Give those up and let's build an open space where we're all working together. That has been sort of the successful formula that Berkeley has applied over and over. And many of those students have gone off and become professors in other schools, and they're applying the same model. It's a very collaborative model.
Starting point is 00:41:10 Had that not existed, for instance, Spark would not have been developed because it happened because two people that were sitting next to each other, one was a statistics, machine learning mathematician and the other was Matejahara Systems researcher, so they actually worked together. It's this interdisciplinary way of working collaboratively. And also it was very sort of in close collaboration with industry and taking problems that exist in the industry and solving those in the research lab. So very pragmatic collaborative approach, both with industry and across the researchers and the professors and the different sort of seniority, I think that has had a huge impact on Berkeley. It's research over the many past decades, but also,
Starting point is 00:41:48 data breaks, frankly speaking. One of our four cultural principles is teamwork makes a dream work. So it's a very highly teamwork-oriented culture. What are you most excited about in the future of technology? What are kernels of potential change and innovation that you think are most interesting today? The most interesting thing is, I think the merger of two markets, two platforms that are going to merge. Technology for AI and data science is going to merge with technology for data warehousing and data management. We even call this paradigm, the lakehouse paradigm, because it's portmanteau two words, data lakes, which typically are used for AI and data warehousing, which is used for data management
Starting point is 00:42:27 and data. So data plus AI comes lake house. I'm excited as that develops. We've been saying it for many years. But in the last year or so, a year and a half, we've heard a lot of other really large companies also get behind this and talk about the lakehouse pattern. So realizing that vision, because I think, you know, you know, No one is fully there yet, including the otherrics.
Starting point is 00:42:47 Realizing that, I think, is going to simplify things a lot for enterprises and get closer to what the Silicon Valley forward tech companies had in the early 2000s. They were able to build it all from ground up with the specific use cases in mind of their business, proprietary to their business, and they had thousands of thousands of PhDs and engineers. The rest of the enterprises don't have that. So building this and providing this to the enterprises, I think actually will transform all these businesses over the next decade. Well, I really love learning about your business. And I think for people that have heard that term big data, maybe they hear a little less than they did five years ago. But this puts so much context around how distributed systems made that possible after Moore's Wall. I actually hadn't heard that phrase Moore's Wall. So it's great to learn about that. And just how Databricks as a company is playing into this ecosystem has been fascinating. I ask the same closing question to everyone that I interview. And so I'll ask you as well. That question is to ask what the kindest thing that anyone's ever done for you is. I think the kind of thing I'm really
Starting point is 00:43:42 really thankful for all the healthcare workers. On the personal side, it's been a very difficult year for me and my family because my 12-month-old son got diagnosed with cancer. During COVID, we've been in the hospital. And the interesting thing is we actually were able to detect this cancer by screening him every three months because we had done genetic testing on him. And we had found a genetic marker that actually says that he's highly predisposed to this particular cancer. This would not have been possible 10 years ago. So he would have probably had this and would have found out much, much later. So I'm thankful to all the healthcare workers. I'm also thankful to the technology that actually enabled us to be able to actually spot this gene in advance and screen him every three
Starting point is 00:44:21 months and then find the cancer so fast and hopefully make him live a long life. That's what I'm most grateful. It's an incredible answer. I'm thankful for you sharing it with us and obviously we'll be thinking about your son. Ali, I've learned so much from you today. Really appreciate the time and all the energy and what you're building. Thank you. Thank you so much, Patrick. This episode was brought to you by Clavio. In this four-part mini-series, I sit down with Clavio customer Nomad and discuss their origin story, why they chose Clavio for their business, and how your brand can grow online sales with Clavio's e-commerce marketing platform. In this week's episode, Nomad Marketing Director Chuck Melbur and I discuss how easy it is to get up, running, and growing
Starting point is 00:45:02 with Clavio's marketing platform. Plus, Chuck shares his advice for other marketers out there. Chuck, I'm curious what it feels like if you were setting some new product up on Clavio, some new e-commerce site, let's say, I launched a widget business tomorrow and I partnered with Clavio. What would it feel like that first week setting it up? What process do I go through? So their Shopify app is pretty easy to get set up and that automatically starts sending data back and forth, which is the most important part of this marketing puzzle. After that, Clavio actually has a ton of built-out flows or pre-populated flows for you to start working with. Like the logic's already there, the inspiration is there for you to work with.
Starting point is 00:45:39 you could be a marketer selling product X, jump on Clavio, not really know much about browser card abandonment, get it set up pretty easy. You know, they have a nice Wizzy Wig editor, drag and drop. If you're a really good graphic designer, you could, of course, you know, design your own stuff and bring it in or do your own HTML. But I assume most people that are getting into the e-commerce space are probably someone like me who's more into the Wizzy Wig editors and there just works quite nice, makes it easy. Across the time, Chuck, that you've used Clavio as a marketing person, what has it been
Starting point is 00:46:08 like to work with them? What does it feel like to work with clavio? In what ways do you tend to interact with a person or not? And just kind of a felt sense of being a Clavio customer. Yeah. So in my six years of Nomad, I've worked with a number of different companies. My experience with Clavio is one of the better ones out there for sure, working with reps on a one-to-one basis. But then also looking at their documentation that's available, as you guys will know, like documentation can get rather crazy and cumbersome very quickly. They've done a good job of distilling it to be super digestible for someone who's not super tech-heavy. Claibio does a great job of just making it easy for me to find my answer and then like
Starting point is 00:46:43 make my solution or fix the problem if I'm having one. Chuck, any closing advice that you would have for other people, obviously e-commerce has exploded during 2020 and not just existing companies, but the number of new companies selling something online. Any advice you would give to especially the people responsible for marketing and getting the word out and building a thoughtful marketing organization, especially to new and young companies. Experiment test by come up with a hypothesis and give it a shot. Don't be scared to try something out even if it seems kind of wild off the cuff. We're still experimenting with that
Starting point is 00:47:17 kind of stuff on a day-to-day basis. Email and social media are both two platforms that are somewhat fleeting in nature. Like you send out the letter, people see it or they don't. Because of that, you're able to experiment a lot and try out different things because it's not going to live on your website forever, you know, it's not going to be something people see every single day. So if you have a hairbrain marketing idea and you think it might work, give it a shot. I never know. To find more episodes or sign up for our weekly summary, visit investorfieldguide.com. Thanks for listening to Founders Field Guide.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.