The Data Stack Show - Re-Air: AI, Abstractions, and the Future of Data Engineering with Pete Hunt of Dagster
Episode Date: September 23, 2026This week on The Data Stack Show, Brooks and John welcome Pete Hunt, CEO of Dagster Labs. During this conversation, Pete takes listeners on a fascinating journey through the evolution of data platform...s, sharing insights from his experiences at Facebook and X (Twitter). Pete discusses the critical challenges of managing data complexity, emphasizing the importance of software engineering best practices and strong abstractions in data orchestration. He explores how emerging technologies like AI are transforming data workflows, breaking down traditional team boundaries, and enabling more collaborative, self-service data engineering. The conversation also highlights Dagster's innovative approach to building a data control plane, with a focus on empowering different stakeholders through flexible, composable tools. Listeners will gain valuable perspectives on the future of data engineering, the role of AI in development, the importance of creating adaptable, user-friendly data infrastructure, and so much more. Highlights from this week’s conversation include:Pete's Background and Journey in Data (1:36)Evolution of Data Practices (3:02)Integration Challenges with Acquired Companies (5:13)Trust and Safety as a Service (8:12)Transition to Dagster (11:26)Value Creation in Networking (14:42)Observability in Data Pipelines (18:44)The Era of Big Complexity (21:38)Abstraction as a Tool for Complexity (24:41)Composability and Workflow Engines (28:08)The Need for Guardrails (33:13)AI in Development Tools (36:24)Internal Components Marketplace (40:14)Reimagining Data Integration (43:03)Importance of Abstraction in Data Tools (46:17)Parting Advice for Listeners and Closing Thoughts (48:01)The Data Stack Show is a weekly podcast powered by RudderStack, customer data infrastructure that enables you to deliver real-time customer event data everywhere it’s needed to power smarter decisions and better customer experiences. Each week, we’ll talk to data engineers, analysts, and data scientists about their experience around building and maintaining data infrastructure, delivering data and data products, and driving better outcomes across their businesses with data.RudderStack helps businesses make the most out of their customer data while ensuring data privacy and security. To learn more about RudderStack visit rudderstack.com. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Transcript
Discussion (0)
Hey, everyone. Before we dive in, we wanted to take a moment to thank you for listening and being part of our community.
Today, we're revisiting one of our most popular episodes in the archives, a conversation full of insights worth hearing again.
We hope you enjoy it and remember you can stay up to date with the latest content and subscribe to the show at datastack show.
Hi, I'm Eric Dodds. And I'm John Wessel.
Welcome to The Datastack Show.
The Datastack Show is a podcast where we talk about the technical, business, and human challenges involved in data work.
Join our casual conversations with innovators and data professionals to learn about new data technologies and how data teams are run at top companies.
Before we dig into today's episode, we want to give a huge thanks to our presenting sponsor, RudderSack.
They give us the equipment and time to do this show week in, week out, and provide you the valuable content.
RudderSack provides customer data infrastructure and is used by the world's most innovative companies to collect, transform, and deliver their event data wherever it's needed, all in RutterSack.
real time. You can learn more at rudderstack.com.
Okay, so special episode here today. We're here with Pete Hunt from Dagster. Pete is actually
the fourth person from Dagster we have ever talked to on the show, which is, I think,
a show record. I think so. And also, if you're like, hey, this is an unfamiliar voice, what's the
deal? Eric is on a plane right now, so it couldn't make the recording. I'm Brooks, producer of the show. You
I've probably heard me here and there before, but here to kick things off today and excited to connect with Pete.
So Pete, what we always do first in our intro is, will you give us just like the quick, high-level version of your background?
We'll get more in depth later.
But yeah, tell us kind of where you started and what you're doing today.
Yeah, it's great to be here.
Thanks for having me.
I'm Pete.
I'm the CEO here at Dagster.
Come from an engineering background.
So kind of the first big thing I worked on was Reactjs at Facebook, which was a large successful open.
project. Then I really wanted to get an entrepreneurship. So I left and started a company called
Smite, which did. That's really where I got into data, the large scale stream processing to try
to find fake and compromised accounts on the internet. Ended up selling that to the company that was
known as Twitter back then. Stayed there for a couple of years. And then my old buddy from Facebook,
Nick Schrock, recruited me over to Daxter. And before I even knew what was happening, I was CEO.
So that was very exciting.
It was very cool.
So Pete, just so many things to talk about.
One of the things I want to talk about in regards to data teams, which we talked about
before the show, is this idea of data people starting, let's say, kind of more from an
analyst background.
They're not from a development background.
And we're seeing people kind of drifting that way.
We're seeing data practices drift that way.
So I want to dig into that with you.
And then what are you excited to talk about?
I mean, I love talking about that stuff.
I'm, as you can probably guess, I'm very into, like, dev tools and frameworks and infrastructure.
And in many ways, that's about enabling different personas to participate in, like, an engineering process.
And I think it's just a really exciting time to talk about, talking about that kind of thing, both because, like, those practices are evolving, obviously, but also, you know, there's a lot of use of, like, large language models to generate code.
I think that changes the math a little bit on who can do what on the stack.
So, you know, we could talk about maybe how, you know, DevTools best practices impact that.
Awesome.
So good.
Yeah, we, John and I have been talking a good bit about actually the kind of shifting ground underneath us all and excited to, yeah, get your take on how Gen A.I is kind of changing the landscape here.
So let's dig in.
Let's do it.
All right.
All right, Pete.
Again, we are so excited to have another person from Dagster here on the show.
you started at Facebook, didn't start in data, but I imagine even back then, like, just tell us
a little more about kind of working then. Like, did you think about data a lot? Or was it really
like at Smite? You're just like, hey, now I'm getting into this. Or was it kind of something
that you had kind of always maybe had an affinity for or kind of drifted towards?
Well, certainly back then, you know, I was originally working on a product team. And it's very, you know,
engineering and power type of organization. So back then, the latest and greatest technology was
hive. And so I was pulling, you know, I was pulling my own metrics to decide, you know,
number one, like what products should we focus on? Because again, like back then it was very much
like individual, small teams in many ways making their own decisions as to what priority.
We wanted to make data driven decisions. So we're pulling from, I think there were like weekly
snapshots back then. I think that was the best we could do. And it was the,
era when you would like tee up your hive query and then go get lunch and 20 minutes later would have
the wrong answer and then you'd have to run it again and then go get coffee. So I was always a user there
for both, you know, guiding product development and debugging stuff. And then, you know, over time as we
would do things like acquire Instagram, for example, they would have to get integrated with the data
systems at Facebook. And so there was a data integration problem that I was, you know, I was a part of. So I was
always, I kind of started on the periphery, but it was always, it was always around me, you know.
Yeah. Well, and I'm sure you got very familiar with the problems and kind of friction points for,
you know, actually working with the data. Yeah, yeah, I was around when they rolled out this thing.
I think it was called Peregrin originally, but that eventually became Presto and Outrino.
And it took those like 20 minute queries down to, you know, a minute or something.
And it was like, you could sit at your desk and still be in flow. And yeah, it was incredible.
I was shocked that you could see that sort of.
Was there like a sense of euphoria when everybody realized like, well, we have this now?
Yeah.
I mean, it's, you know, they're never going to attribute stock growth to, you know, the data platform team.
But like, think about it, right?
It's a big giant social network.
It's not like you can talk to your users at any sort of scale to figure out what they want.
You have to make the decisions based on data, right?
And if you're able to like, you know, make your data-driven decisions, it's like 20 times.
faster. That's a big deal, right? So I think a lot of that, the growth in that company comes down
to technologies like that that enable these like business and technical people to be able to make
quick decisions. Totally. That's really cool. So this is really funny. I just, I just thought of this.
So remember 2012, 2013 when the like, I think they were called MOOCs, massive online course,
like a udacity, right? Yeah, yeah. That was like around the time I was first getting more into some
data science stuff. So there was a course. I still remember the course. I mean,
it's been years now by I wish I could remember her name, one of the data scientists at Facebook
at the time. And I'm still, now I'm like thinking through in context of this conversation,
like some of the really interesting things. Like I remember, because, you know, like she's at
Facebook. A lot of her examples are like related to Facebook. But one of the really interesting things
was the was like the birthday problem, right? Because we were looking at outliers like as part of
the data science class. And there's this funny thing, which I just never thought of like,
oh yeah, like look at all those like January 1st birthday.
birthdays of like, you know, like some rounded off year is like, oh, yeah, that makes sense.
Like Facebook has that problem. Like somebody just filled in a random birthday. So there's like a lot of
interesting data problems that like you wouldn't think of, you know, in the social space.
And it's really interesting when you start to look at it in an adversarial context as well.
That was the thing that I did after Facebook was like trying to find fake accounts. And, you know,
there was common birthdays. To give, to give listeners context, the timing of this was,
Very crucial too, right? This was during kind of COVID and lots of unrest.
Yeah, we started the company actually, like end of 2014, early 2015.
But we ended up selling it to Twitter in 2018. And that was like, you know, I was there from
2018 to 2022. And so it was very much like, you know, elections, COVID. I think there was like
some global disaster too that I can't decide that. I can't remember.
That you blocked from your mind at this point.
Just really quickly, smart is the name of the company.
What did y'all do?
We called it kind of trust and safety as a service, but really what it was, we would ingest event data from marketplaces and social networks.
And then in like near real time, we would try to basically find fake accounts and compromised accounts.
This was really where I got into data.
I actually got in through like street processing mostly.
And what was kind of interesting about this problem is it's set for.
First of all, it sounds like a machine learning problem, right?
It's like, oh, you like do some feature engineering, you label the data and you, like,
throw it into like logistic regression or something and you get a classifier out.
That doesn't work.
And the reason why is because it's adversarial.
And so you actually don't have up-to-date labels because the patterns, the attacker patterns change
all the time.
So oftentimes, at least back then, I think now they're using like transformer models and they
work really well.
But back then there was like this combination of like,
anomaly detection, manually curated heuristics, and targeted machine learning at specific problems.
And the thing is, you had to respond and label the data like, you know, at very low latency.
Because, you know, you're talking about like if you wait five minutes, you know, that could be there was, you know, you can get a lot of spam into a system or compromise a lot of counts of five minutes, right?
So it was just a very interesting problem.
And that's really where I fell in love with the data space.
It was really cool.
Nice.
I imagine you think a lot, probably read a lot about similar problem, but I think today, right,
it's like trust and safety with all these foundational models.
I mean, is that something you're still super interested in and kind of in today's day and age with AI?
So to tell you the truth here, I found trust and safety to be a very interesting technical problem.
I thought it was like, I mean, all the problems that I just laid out to me as like somebody
that grew up writing code and was really interested in distributed systems and like data analysis
and stuff. It's like a mystery and you can apply many different types of techniques to get to the
air. And there's all sorts of like interesting tricks you can use. So it's just a very, I thought it's
very fascinating from an intellectual perspective. I thought it was, you know, really like obviously
fulfilling to that like it's, that technology is now primarily used for like child safety and cybercrime
over there. I haven't worked there for a long time now, but I believe it's still used over there
for those sorts of applications, which is obviously like a very fulfilling thing. But there are
lots of parts about that. It's also a very fraught, basically category, right? Like, you know,
I think that everybody wants to stop like child predators, right? Yeah. But there's, you know,
once you get beyond that, there gets to be a very big gray area around policy, you know, what's legal,
what's not, what's proper, what's not.
And to me, especially during that time,
it was just like, it's pretty messy.
And it wasn't really what was fulfilling for me.
And really, for me, what I'm excited about was like all these interesting data problems
and like, you know, making developers happy through like dev tools and infrastructure.
So I did end up leaving after like three and a half years or so.
But it was a good run there.
And so reconnected with a friend from Facebook and went to Daxter.
Yeah.
Tell us a little more about that.
Yeah, so let's see.
I had known Nick back in the day.
I was working on React.js and he was working on GraphQL.
He's both like open source projects that came out of Facebook at that time.
We like metaphorically and actually sat down the hall from each other.
And, you know, we always stayed in touch and we're friends.
And I knew he started Dagster and I put a little money in early at the seed round.
So I was always close to the company.
And, you know, when he was looking for a new head of engineering,
It was right around the time that I was kind of ready to wrap up at Twitter.
And he was like, hey, you know, you knew head of engineering, can you help me search?
And I helped him with a search for a little while.
And then I just decided, hey, you know what?
Like, I think it's a little bit of a change.
I'll come over and be head of engineering.
All right.
Like, if I have to.
Yeah, yeah.
And, you know, we had a really good first year.
And it was one of those things where I had done the CEO thing before.
And, you know, I knew it was like a job where you're really busy and you don't have time to do everything.
And so I would kind of try to find places where his attention was elsewhere and just try to
like help out there.
Right.
So like there are certain things, you know, that you kind of learn the first time around
and mistakes that you make that you don't want to make the second time around.
So I helped him not make those mistakes the second time around.
And by the end of the year, it's like, listen, man, like, I've been, you know, I've been a solo
founder for a long time.
It's a ton of work.
And frankly, I think he got into it because he wanted to like write code and work with
customers and like be a, you know, be a technical vision.
or whatever. And that's the CTO's job, not the CEO's job. The CEO's job is to clean the
toilets and make sure that there's money in the company bank account and stuff like that.
So we talked and we decided that it made sense for me to be CEO and he could step into the
CTO spot. And I think it's great, you know, like it's, for me, it's like stepping into
an old pair of shoes picking up right where I left off. And for Nick, I think he gets to work on
the stuff that he's really excited to work on and go deep on. Yeah. I just want to call out
something really quick there that I think it's cool and not a given.
like having that background of like somebody that you know and trust because like if he had brought
somebody else and as head of you know and that's unlikely to have happened it's possible it could
happen but I think it's cool like having that like kind of long term connection where you can have
that flexibility right where you both like kind of understand you know how to work together
then you can do some neat stuff like that that maybe you could normally do in other context
Yeah, you know, it's like people you work with, you know, like they always come back around in the future and you never know who you're going to work with in the future. And so I guess like always in my career, I was always kind of trying to have this like aura of like value around me. It's not about me. It's about actually other people. It's like, you know, somebody's like if they have some sort of interaction or they're working with me, like they come away like more successful. It's like how I was, how I was, how.
how I tried to think about my early career.
And I think that kind of like worked in a lot of ways and helped create like a good network
for me.
And like very concretely like what that means is like often I like wouldn't work on the cool
thing.
So like back at Facebook for example, you know, the transition to native mobile was like a big deal.
Right.
That was where all the best people were going.
They were retraining to go to mobile.
And I just like stayed on the website, you know, because that's kind of where people needed it.
Like they needed somebody who was good who was willing to be.
focused on maybe the thing that wasn't super hot right now.
And like, you know, that's where React came from, right?
And that was a really successful project.
And so like, I think for me, that strategy really worked and created a great network for me
that has certainly well in the future.
Yeah, really cool and just great, I think, advice for anyone and everyone.
I do also want to call out, I think he said, you know, Nick wanted to be the technical
visionary.
We have had him on the show before.
It's probably been about a year ago.
But if you want to hear a technical visionary, go back and listen to that episode.
I mean, the way he articulates his vision for orchestration in Dagster is, I mean, it is pretty incredible.
So, yeah, go back and check that one out.
Yeah, it's been a thinker for sure.
Yeah, for sure.
Yeah, for sure.
It want to get into the kind of nitty-gritty of orchestration, talk about Dagster.
But before we do that, can we just get like your definition of orchestration?
Just zoom in all the way out, kind of basic level, what is orchestration?
And then from there, I think we can, I may let John take over and go deep.
And I have no idea with orchestration.
I was going to say, everybody, I'm excited to hear your definition.
Well, I don't, you know, it's interesting because everybody does have a different definition, right?
And, you know, I mean, just to get really concrete really quickly, we're like the thing that schedules, runs,
and monitors your data pipelines.
But I think that when you frame it,
like that, there's a wide variety of technologies you can use. And I think that we're on kind of this
evolutionary path from a, like, you can imagine a spectrum or a timeline. There's like schedulers over
here and there's a control plane over here, right? And it's similar to kind of how you saw
like container orchestration and infrastructure like that can infrastructure evolve over time.
You started with like a single server and like Damon's running on the Unix box all the way to
something like Kubernetes, which does, which is really a control plan.
over all of your services. And so we kind of think of orchestration as going a similar route.
So you start with like Kron or something that looks like Kron, built-in scheduler into a product,
like a Control M or something. And, you know, that thing is very simple. It just runs your jobs
at a certain time. You quickly find that you wait, you know, you're over computing. You're running
every step at every, you know, time slice. Failure has become a big problem. Observability is like,
non-existent. So then you move to something that's like a workflow orchestrator, right? This would be
like an Apache Airflow or something like that. And then now you've got like a smarter Cron and you've got a
smartocrine that can retry and retry individual steps, right? So you've seen a major improvement over
something like Cron. You know, the problem is though that like if you're on a data team,
it's kind of an impedance mismatch between what those workflow orchestrators are doing and what you're
trying to do as a data team. It's like specifically like the data. The data.
Data team is thinking in terms of tables or machine learning models or files in a data warehouse.
And the workflow engine is thinking in terms of like these opaque steps, right?
So we kind of see like this step, this, you know, there was this move from like kind of
more of like workflow orchestration to like a data control plane.
And like fundamentally what you need there is a deep understanding of the data assets,
the lineage between them, the current state of them, all the metadata.
And then you get this rich system of record of every single data asset in the, you know,
in your organization.
And once you've got that information,
you can really build a bunch of like interesting observability stuff on top of that
and really help.
Like kind of,
to me,
that's the last piece of,
you know,
orchestration is like being able to observe what's happening
and let a human operator,
you know,
fix,
you know,
issues with your pipelines.
So I just want to pause to see that like made any degree of sense.
Yeah.
Yeah.
I mean,
definitely to me,
but Brooks is probably the better one to,
like,
respond to that.
No,
it was great.
No,
Yeah, love kind of breaking it, breaking down the fundamentals. It was great.
We're going to take a quick break from the episode to talk about our sponsor, Rudder Stack.
Now, I could say a bunch of nice things as if I found a fancy new tool.
But John has been implementing Rudder Stack for over half a decade.
John, you work with customer event data every day and you know how hard it can be
to make sure that data is clean and then to stream it everywhere it needs to go.
Yeah, Eric, as you know, customer data can get messy.
and if you've ever seen a tag manager, you know how messy it can get.
So RudderStack has really been one of my team's secret weapons.
We can collect and standardize data from anywhere, web, mobile, even server side, and then send it to our downstream tools.
Now, rumor has it that you have implemented the longest running production instance of Rudderstack at six years in going.
Yes, I can confirm that.
And one of the reasons we picked Rudderstack was that it does not store the data and we can lie.
stream data to our downstream tools.
One of the things about the implementation that has been so common over all the years and with
so many rudder stack customers is that it wasn't a wholesale replacement of your stack.
It fit right into your existing tool set.
Yeah.
And even with technical tools, Eric, things like Kafka or PubSub, but you don't have to have
all that complicated customer data infrastructure.
Well, if you need to stream clean customer data to your entire stack, including your data
infrastructure tools, head over to rudderstack.com to learn more.
So now it's my fun time.
So on the technical side, we talked about this a little bit in the intro.
I'd love to, well, let's start here.
Let's talk data stack.
Let's talk a little bit of evolution, modern data stack, and talk about tooling,
like how you've seen that evolve and then maybe where you see it headed.
Sure, yeah.
I mean, I think, you know, we were talking about running 20-minute queries on Hive back in the
day.
that was definitely the pre-modern data stack.
In many ways, big data was still a challenge, right?
I would say that even in the high of the era, like,
we had these big data tools,
but big data was not a solve problem yet.
We couldn't, like, easily compute over, like,
unbounded sets of data, you know, in a reasonable way
or, like, effectively unbounded anyway.
And so I would say with the arrival of tools,
like that Peregrine thing that I told you about,
but also, like, Snowflake, Databricks, Big Query,
you really got to like almost interactive query speeds, right?
And to me, you know, right around the time that like big query came out and I think the,
you know, it's like the Google Dremel paper, right?
Like that to me was when like big data kind of became a solved problem.
Like, okay, we know how to do this.
We can like run more computed a problem and solve most data challenges.
Then, you know, once you get this new capability, people start using it, right?
And they start to basically build a problem.
a bunch of stuff on top. And to me, that was like kind of like where we entered the modern
data stack era. And the way we think about it at Dagster is that created this era of big
complexity where suddenly you have all these different stakeholders building all this mission
critical stuff on top of this, you know, this new capability that they have. But, you know,
oftentimes they aren't, the tooling and infrastructure doesn't support the level of service
they need to provide for really like production system right so like specifically what I'm talking about
here is like the clicking buttons in a UI and pressing save and critical state is now in some system
that only one team knows about and is not version controlled right and you know everybody knows
it was like a big market correction 22 2023 maybe those people got laid off and now suddenly
there's this like you know you got this whole giant you know data estate or mansion that's built
on top of like one little wooden pillar that is, you know, not maintained by anybody and
termites are munching at it and eventually that thing's going to give out, right? So what we,
you know, at Dagster, we really believe that software engineering best practices are the way
to tame big complexity. This has been a trend like in every other part of engineering, right?
Like, again, you know, citing two examples. If you think about infrastructure management,
it started out as like a syshing into an individual box and like running the magic commands that only that person knew to make sure that like iNetD was running or whatever and now it's all done through tools like terraform and infrastructure is code right and now it's you know you can roll back you can like onboard a new person and they can actually understand what's going on and you look at the front end world it's a similar thing right there used to be there's a big giant hairy CSS file that nobody knew how to understand
or nobody could understand.
And today, there are tools like React and CSS modules and stuff like that,
like really enable this like kind of standardized way to build and operate, you know,
your applications and solve version controlled.
And so we have been trying to do that for data.
We're not the only people trying to do it for data.
Like we've seen DVT, for example, bring this style of development to that particular persona.
But we think it's, you know, we think that like a data platform control plane really like brings this,
this way to manage complexity
to like the whole data platform.
Yeah, for sure.
And then I think this leads perfectly
into kind of our next topic here
is the complexity topic
of like you have a lot of complexity.
You're managing a lot of complexity.
So what strategies are you guys
that Dax are thinking about
to make it more simple?
Because it is a complex problem, right?
There's a certain amount of complexity.
It's just a complex problem.
But I know you guys are working on
a lot of strategies to, you know,
to simplify.
Yep. So there's like kind of exactly one tool that we all have in our arsenal to address complexity of any kind, which is abstraction, right? Which is taking a, it's almost like taking this weird amorphous problem, finding the pattern, a common pattern and path to success. And then like putting it up, like wrapping it up in a box and making it like kind of a repeatable process. So you kind of present a clean, understandable interface.
to a really complex problem underneath.
So you can kind of solve the lower layer once,
and then the problems above it are a bit simpler.
So really it's about abstraction, right?
When we talk about complexity management,
we've seen this in all areas of engineering.
Like we started out writing assembler code with go-toes.
Then we abstracted that away into these structured programming languages
with like reusable functions, object-oriented programming, you know, etc.
And we think getting the abstraction rate in the data,
platform is key to managing the complexity of the data platform. And what we saw was,
can you articulate getting the abstraction, right? Like, how do you think about that? What does that mean to
you? And even some examples of maybe how kind of you and the team think about the problems you're
solving. Yeah. I mean, this is the art and science of building a framework, right? How do you get
the abstraction, right? Yeah. Yeah. And so I would say there are, you know, people write,
books on this stuff. So you can read like Martin Fowler and you can, you know, people talk about like
coupling and cohesion as principles here. You want to have the different parts of your system
have low coupling so they can be examined independently and high cohesion. When you read a single
module, they make sense. There's a wide body of computer science literature about this. But really what
I think it comes down to is how much power do you want to sacrifice or sacrifice
for the user in order to give them some new value is like fundamentally the first thing that
you think about. And then the second thing is like how do they pull the escape hatch when they
need to? Or do you even want them to be able to pull an escape hatch? And usually in most systems,
you do want to give them some sort of escape hatch that is reasonable. So that's how I really think about
it. And so oftentimes you're trading some amount of flexibility in order to get some property of
system that you want and oftentimes that is increased developer velocity or you know increased
observability or ability to debug your system does that make sense at all totally yeah that no that was
awesome thank you yeah yeah and i think one of the things too that because we talked about this so
there's the roles question that we were talking about before the show of like all right i'm more of an
analyst or now dbt's introduced this kind of like analytics engineer like all right we're going to
blur some lines here
piece and then like the other side of like I'm a, you know,
more traditionally trained engineer or maybe I'm a DevOps person or something.
So I think it'd be interesting to talk about even maybe specifically,
maybe generally with orchestrators, but specifically with Dagster,
like how do you see those coming together?
And then on top of that,
we've got this new component of the AI piece that also makes that a little bit more
complicated as far as like what the roles might even look like.
Yeah, so I would start by saying, like, you know, we've talked a bit about abstraction, right?
Very much related to that is this notion of composability, which is like, okay, I've abstracted away one component of this system.
And when I connected to a different component of the system, like, it works in a predictable way.
Right.
So the idea here is that, like, we've given you a set of abstractions.
You can put them together, like Lego blocks is like the analogy everybody uses, but a number of ways analogies you could use.
You combine them in ways that the abstraction author didn't think of.
before prescribed for you.
Right.
And the system still has the properties that you want or that we all agreed to, right?
And so I promise I'm going to get to answering your question is.
Like, the challenge with kind of like these workflow engines is that the task abstraction
is a very weak abstraction.
You don't trade very much power.
It can do like kind of anything.
But in exchange, you don't get very much benefit from it either.
Like you can't really arbitrarily compose them together.
And it's actually quite difficult.
to observe what's going on inside of an opaque task and let's, you know, manually instrumented
with some sort of observability system. And so you don't get very many composability benefits.
And so then when you want to onboard a bunch of different stakeholders, you either, like,
regardless of their persona, whether they're software engineers or data analysts or infrastructure
engineers, you're still like, if you have a really weak abstraction, it's very risky because
people can step on each other's toes and cause interactions between the components
as you don't expect.
So what often either happens with a workflow orchestration tool is you either just get a big mess
or you build, like the user builds their own abstractions, like a platform team that will
build their own abstractions on top.
And then their stakeholders will onboard onto that.
And I think what usually happens is both.
It's usually like a big mess.
And then the team is like, oh man, we got to go clean this up.
And the process of cleaning that up is like refactoring into these abstractions that then these
stakeholder teams can use.
So, you know, our abstraction, by the way, is this thing called a data asset, which can
represent like a table and a data warehouse or a file in an object store or something.
And that to us is like a really great way to enable different stakeholder teams to interact.
And so what we, you know, what we see today is we will get.
like machine learning engineering teams
that will be building stuff
in like, you know,
either notebooks or just like kind of stuff
using the Python scientific stack,
being able to integrate, like,
write and deploy data assets into the Dagster
right alongside the DBT engineer
that imports their DBT project into DaGster
and every DBT model gets represented
as an asset within DaGster.
And so this is what I mean by like, like,
you know, having those two teams like work together
because like, like I saw this at Twitter, right?
Like we're trying to find spam.
There's a team that's using machine learning.
Actually, it was in Scala, but there's a team using machine learning.
And then there's a team that's like hacking together like SQL queries that like work.
That have like the magic rejects that finds the spam campaign.
Right.
And you need both of these things to really deliver, right?
And they both depend on the same upstream data sets.
But they're using completely different like stacks and they can't build on each other's work.
And we often found them building parallel data sets that did exactly.
the same thing, except they were slightly different because they couldn't, like, there wasn't a good
abstraction for them to be able to, like, collaborate on one platform. So that's why we, like,
keep, like, hammering on this, like, notion of an abstraction. And so, you know, we think that over
time, you know, more stakeholders will be able to participate directly using this abstraction. And
with, you know, the rise of tools like Claude Code and cursor, you know, what I think is happening and is
going to happen is that, you know,
individuals are going to feel
like more empowered
to work in areas of the stack
that they were not previously, like, familiar
with. Like, we're seeing this today
at Dagster where, you know,
we would have engineers that
previously would only work
in the Python code base. And then when it came
time to deploy something, they'd like call up
somebody from our platform team and say, hey, can you help
me write the Terraform to do the
RDS configure whatever?
And today, they're like,
using LLMs to generate the Terraform config, getting it reviewed by the platform team,
and it's just much more efficient.
They're just able to do more, right?
So I do think that it is.
The boundaries between teams and stakeholders are changing for sure.
So we talked about guardrails.
In this new world where more people are able to do more things, I mean, guardrails kind of jumped in my mind.
And you talked about, okay, platform team is reviewing the stuff that they did.
but are you thinking about that kind of more, more critically and even maybe at a higher level,
right?
It's like, hey, as more people do more things, like we need guardrails, we want to have more freedom,
but we need guardrails.
And here's how we're doing that at Tagsure.
Yeah.
So in many ways, like abstractions are guardrails, right?
Yeah.
And we often hear, you know, you talk to a data platform team and they often say,
listen, like, you know, it's our job to build the central platform, a shared set of tools and
best practices to enable other teams to be successful, right? And the, like, their goal really is to give
those teams as much autonomy as possible, but like no more than that. You know what I mean?
So they want to put like guardrails that make sure the organization like stays compliant with the
obligations they made to their customers and regulators. Also, you know, make sure that everybody
stays within budget and leverages the best practices and tools to make those teams more successful.
And so very much like the platform team's job is to build these abstractions, build these
guardrails, you know, for these stakeholder teams. Now, the way that this has worked in the past
is, you know, with Dagstra in particular, is we have this notion of an asset as like our
kind of fundamental unit of composition. And it works really well for teams that are, you know,
Python forward, right? Python, they could take Daxstra out of the box and generally,
be pretty successful. It also works really well for organizations where we have a really good
out-of-the-box integration. So like, or technologies rather. So DBT, really good example.
Pointed to extra at your DBT project. Your DBT developers can be, you know, dexter users, no problem.
But there's a big world out there of like diverse stakeholders, you know, different tools.
And a lot of organizations have like their own tools that they built internally. And, you know,
they wanted a way to basically build those guardrails or build those abstraction layers on top of
Dagster for their stakeholder teams. And we saw customers doing that, you know, using, you know,
just using Python, right? Yeah. They maybe build a YAML, DSL domain specific language on top of
Dagster. Maybe they would build like a special Python library that that would kind of translate
their domain concepts into like Dagster concepts. But what we had found was everybody was doing
it in slightly different ways. The tooling was often like MVP status. So like they didn't have a
beautiful VS code auto complete extension for their thing. Right. It was right. You know,
whatever they were able to get done in the limited time that they had to work on this. And so we said,
listen, let's take all the stuff that users are already doing and build them great tooling to build
their own abstractions. And we'll like ship some out of the box abstractions too for like common
use cases like a DBT or data movement tool. And so that's a thing that we built called Daxra
components. And I think it's, I think it's very interesting developing new dev tools and new
abstractions in the age of AI because like it's part of our, we consider like Claude as like a
user just alongside like our design partners, right? So we'll test an API and we'll say, hey,
did Claude actually understand this? Was it able to like one shot what we wanted to do? Wow.
And, you know, it's very, you see this with tools like, you know, if you talk to the folks at VSEL working on V0, they do similar things, right?
And was that easier or harder than you expected it to be?
What's interesting is the stuff that makes it good for LLMs makes it good for humans to.
I was going to ask, are there parts that are a divergent pass you've seen there?
Or like, so far it's been like we can just kind of optimize duly for both at the same time.
It's 80% both at the same time and 20% different.
So there's like the way that you provide the documentation to the LLM is quite different than how you're providing to humans.
Right?
You want to like actually with humans you want to like give them a bunch of examples and context and stuff like that.
With an LLM you have a finite number of tokens that you really want to be able to burn on this sort of thing.
So the way you deliver the documentation is different.
And there's this thing called Model Context Protocol, which, you know, we built.
kind of an integration, let's do integrate with like all these LLMs and give them kind of
programmatic access to tools and documentation. So certainly like that's a thing specifically for
LLMs, but 80% of it is like, you know, if we were talking to an LLM, we would say we need to
reduce the number of tokens in the context window. Right. And if we're talking to a human, it's like,
we want you to only have to look at one file to know, to be able to like solve this problem or know
what's going on. And the code should be concise, right? These are things that humans like. And LLMs,
I think, also, you know, benefit from. Similarly, like, you know, human, like LLMs need feedback
from tooling that says, hey, did I write my code correctly? Does it pass the schema check? Does it,
you know, initialized correctly? And that needs to be as fast as possible so the LLM can work. And,
like, it's the same thing for human, right? So what's, what is interesting is like, as we
started to put this through the like LLM ringer, the framework just got a lot better for humans.
So it's, I don't know, we live in this weird age of like cybernetic program.
Yeah.
So I think the components thing, that conversation is really interesting.
And I just want to make sure that like that I and our list just kind of understand is the future state here where like components are going to be at like the like there's going to be a specific component, optimist?
for like a specific destination like Postgres or something and a specific source like
Salesforce I don't know is that how I should think about components I know there's it's kind of an
abstraction above that but is that you think that's one of the kind of practical use
let me so let me give you an example like I can give you a couple of examples because it is an
abstraction above that right right so we're going to ship with you know importing a dbt project
as a component so as a dbt project component we're a ship with various BI
tools and data movement tools as components. So, you know, you want to integrate with like,
you know, whatever ELT tool or whatever the I tool you want. Right. There's a component for that.
And we think that's going to actually like kind of reduce the time from like, you know,
not knowing Dexter at all to having something in production like by 10x, right? But really the value is like
you're going to have, we're shipping an internal like components marketplace for enterprises where like
they're not going to want their teams to take any SaaS data movement tool off of the shelf,
right? They're going to have their approved vendor that they use or their approved technologies
they use. And so they're going to build their own internal component that makes it very easy
for teams to spin up their own data movement pipeline and adheres to their best practices.
Okay. And then like bringing in like model context protocol, MCP, things like that. So I've got this like
marketplace. And it's like that abstraction layer up, which is higher leverage, right?
Like if you're not trying to connect directly to, you know, specific SaaS tools and you're
like up at the layer above where you're working with like all the common, you know, extraction
tools, for example, or extraction transformation tools. So then like if I'm kind of an ordinary
or kind of a less technical user, I theoretically could use whatever my company uses as far as like
an LLM. There's potentially like from the Dagster side, this.
model context protocol, which gives context to the LLM for what I'm doing here. And then I could
describe in English, hey, I want to move data from here, transform it in this way, and I want
it to land here, for example. That's right. Yeah. And it's, and like the way to think about it,
too, is like the person doing that, they probably spike really deep on some other technology.
Maybe they're really good financial analysts and maybe they're really good DBT development.
They're a really good machine learning person. And they're not a dad.
Dagster expert, right?
Yeah, right.
So there's like, integrate this thing with this thing.
And then like, do it for them.
And then there, but there is going to probably be a small team of Daxter experts at the company, right?
And their job is going to be to basically build those custom components.
And through the model context protocol integration between like Dagster and CloudCode or
cursor or whatever, those custom, like, those Dagster data platform engineers can like teach
the model how to be really a thing.
effective in their stack.
Right.
So it's actually like once you see the full development workflow, it's like, I think it's
going to really change how teams develop.
I think it's just going to empower like a lot more folks to participate in like a
self-service way without creating like a bunch of technical debt or having to block on,
you know, other teams to build part of their session.
So here's a follow up question then.
And it's an unfair question.
We talked a lot about how like data is like, you know, drifting toward a lot of these
workflows that are really kind of more.
mature, you know, what a front-end developer might do, even like the DevOps world.
What is something, because you guys are, because it's a little bit more greenfield,
what is something you're like, I think we can get better because we know about how
all these other workflows work. Does that make sense?
In terms of like, like, what's the value? Developer experience essentially, right?
Because you guys have worked in these other like, in these kind of, you know, front-end with
like React, for example, or like DevOps, Kubernetes, things like that.
is there, and maybe the answer is AI, but is there something like specific you're like,
we're going to like be able to replicate all the best practices or so the really good things.
And like here's like a couple things we're excited that actually may be better because we get to
kind of start over.
Oh yeah.
I mean, I think that everybody knows that there's like too many tools and the integration between
tools is like a big pain in the butt.
Yeah.
And so like, you know, when you kind of like, if you have it like an orchestrator that understands
the asset lineage in a very.
deep way and understands where the data is coming from where it's going to, where it's stored,
the current status of it, whether it's passing its quality checks or not.
Like a really great observability tool and data discovery tool just automatically.
Right.
And it's not like we did a process to document all our data.
We just wrote the code in this way and we got all these capabilities.
And this is like, again, the power of like a really good abstraction, right?
It's like, you know, we put some guardrails in place, which probably sacrifices a little bit
power. In exchange, though, you get like a data catalog, like out of the box and you get like,
you know, an understanding of the freshness of your data assets and like, if they fail their
freshness checks, we can like automatically remediated and stuff like that. So it's actually,
you know, it is a bit of a rethinking of the stack. When you start to go from like, hey,
orchestrator is just a fancy scheduler to know this is like a like a control plane across the whole
data platform. It actually does, you know, rethink what you can do.
in the shape of the stack.
You know, I kind of think of it as like, you know,
you've got your big data compute layer below here,
like snowflake data bricks and stuff like that.
You've got your like BI tools and data activation up here.
There's a bunch of messy stuff in the middle.
We can really help, like a control plane really tames the complexity of that messy middle.
Right.
Yeah, I think that makes a lot of sense.
This is like kind of a very specific question.
So there's like a lot,
there's kind of the general flow,
which we've talked through several times,
where we've got, you know,
what's call it sources,
we've got transformation,
steps that are happening in the middle,
we've got data landing.
What about some of the more like,
what I call it,
edge cases,
because they're very common,
but some of these like MLAI,
I've got like unstructured data
that I want to bring in,
or even like,
maybe even more edge case of like,
I have like fairly sophisticated like
security and like governance that I need to maintain.
And like I have all these SQL scripts
that like run to do things.
Or I have like auditors here today.
Like there's,
just all these interesting, like, you know, long, like, essentially like a long tail of people
that, you know, I think are going to be also users. So, you know, I've talked to all of those,
but maybe you pick one of those. I'd be curious for it more. I mean, I'll tell you that we target the,
you know, when you zoom out and we're really in the business of taming the complexity and we think
that software engineering best practices is the way to do that, it kind of implies like a technical
or semi-technical user, right? So there are teams, like I, you know,
you know, working on trust and safety at Twitter, right?
Like there's a ton of like data compliance and privacy things that you have to work through.
And there's, you know, giant legal orcs that you have to interface with.
I think generally like Daxter is not the tool for them.
Sure.
But for a lot of the kind of technical stakeholders that are writing SQL or doing data analysis or anything like that,
that's kind of really where we see like Daxter, you know, being kind of the tool for them.
I'm not sure if I answered your question.
No.
Yeah, no, I think that's helpful because essentially to do the abstractions well and make them useful,
you can't solve for every use case.
Or like the abstraction is kind of like bad or not really an abstraction, right?
Right, it's a kitchen sink, right?
Yeah, exactly.
Yeah, we don't want to do everything.
Like we design the asset abstraction and a couple of other abstractions around it.
That makes sense.
It's kind of like you define the asset abstraction.
You think about the life cycle of a data asset.
And then we can hook in with best of breed tools or, you know, custom code from the user and then like kind of bring it all together into one place.
Right.
We are almost at the buzzer here.
But Daxure components out now, if folks want to learn more, see it in action, what's the best thing you get at the website?
Yeah, this should go to dagsur.io.
And it's, you know, it's an open source framework so you can read the documentation and install yourself.
or if you request a demo from our team, we'll get you on with an engineer and they can
demo our commercial offering.
Super exciting. Last question before we wrap here, Pete, you have had an extremely interesting
and I think fair to say, prolific career, though. I don't think you strike me as pretty
humble. I might not say that about yourself. But you have learned a lot of interesting lessons
along the way. Parting piece of advice to our listeners, somebody working in data
today, maybe especially facing, you know, AI is changing a lot of things.
But jury is still out on exactly, you know, what things will look like.
But what would be kind of just a parting piece of advice that you'd give our listeners?
Yeah, I mean, it's a very interesting and broad question.
But I would say, just be like an empathetic person and try to help people out.
And like in the tech industry, it's the type of thing where, like, helping people out indirectly
equals success, you know? And so even if you're just totally selfish and you're totally looking
out for yourself, adopting a default strategy of being an empathetic person will be good for everybody.
So that's what I would leave people with. Yeah. That's great. We'll be been an awesome show.
Thank you so much for coming on. And yeah, I'm sure the way it's going, we'll have someone else from
Daxter. For sure. We'll look forward to it. But thank you, Pete.
Cool. Thanks, guys. Yeah. Thanks, Pete.
The Datastack show is brought to you by Rudderstock.
Learn more at rudder sack.com.
