Screaming in the Cloud - The Invisible Network Powering AWS with Matt Rehder
Episode Date: September 10, 2026AWS VP of Global Networking Matt Rehder joins Corey Quinn to pull back the curtain on the massive network infrastructure behind AWS. They explore resiliency at scale, AWS’s move toward flat...ter networks, the advantages of building custom hardware, and why AI is making networking exciting again.Show Highlights: (01:08) Meet AWS Networking Lead(02:01) Why AWS Avoids Global Outages(05:23) RNG Flat Network Explained(10:08) Overbuild Capacity And Custom Hardware(17:37) VPC Virtual Network Origins(19:50) Why TCP Still Wins(20:38) SRD Inside AWS(21:54) Opt In SRD Transport(25:31) Networks As Utilities (27:24) Learning And Growing Engineers(29:22) Training Talent In House(31:27) AI Rekindles Networking(34:26) Where To Learn Networking
Transcript
Discussion (0)
We build more network capacity than we need, both from a reliability perspective, so that we have a lot of redundancy and resiliency.
Many devices can fail. Many links can fail. We still have sufficient capacity.
Welcome to Screaming in the Cloud. I'm Corey Quinn. Somehow, I have managed to breach containment and go talk to people at Amazon.
My guest today is Matt Rader, who is the VP of global networking. Matt, thank you for joining me.
Yeah, thanks for having me.
This episode is sponsored in part by my day job.
Duck Bill. Do you have a horrifying AWS bill? That can mean a lot of things.
Predicting what it's going to be. Determining what it should be. Negotiating your next long-term
contract with AWS. Or just figuring out why it increasingly resembles a phone number, but
nobody seems to quite know why that is. To learn more, visit duckbillhq.com. Remember,
you can't duck the duck bill bill, which my CEO relies.
informs me is absolutely not our slogan. So what is it you do exactly? Networking. What's the point
of it all? That feels like something old people care about, said the old person. Yeah, I mean,
the simplest answers, we make all the servers talk to each other. Otherwise, they're just expensive
space heaters. Otherwise, they're very expensive space heaters, yeah. So it's really interconnecting all the
servers, interconnecting all the data centers around the world, and then connecting to all the people
around the world on the internet is the gist of it. You were kind enough to recently give me a tour
of the networking lab, one of the networking labs,
I'm guessing there might be more than what.
We have more than one.
We didn't show you the super secret one, but we can...
Oh, exactly, yeah, the real secret one.
Yeah, that's where the cables are longer
so we can fit more data into them before it comes back around.
Oh, yeah, that's how networking works, right?
My networking background myself is such that it's just enough to know
what I don't know, but it does guide me toward asking
some of the right questions.
There's a whole different sense of scale that I don't think most folk appreciate.
Yeah. One of the
enduring qualities of
AWS that has also led to
frustration, but I maintain is the right answer,
is your networking is phenomenal.
You have a harsh separation
between regions. You have
knock wood, never yet
seen a global networking
outage. There have never been a global rolling
outage of things going down.
Failures are region bound.
In virtually every case,
I am aware of. And that was a choice,
and it does mean that when you log in an AWS account,
Each region makes it look like they're now 31 AWS accounts.
Go hunt and find it down.
But the resilience and durability story are unparalleled by any other company.
I include every other hyperscaler in that list.
Resilience is great and it's been that way for a long time.
Now it seems that systems have gotten good enough across the board,
especially since modern AI systems are not that reliable.
We're not seeing five-nines of uptime.
anything, but companies are putting it in their critical path to a point where it feels like
resilience is something companies pay lip service to largely. It is an afterthought more than a lot of
other things. It just isn't an area of focus. My sense is that this is something that happens
when you've been too long without a really good, really notable outage. Sort of like a self-inflicted,
we've been too good. People forget this stuff can break. How do you see it? Well, first of all,
thank you for the kind words. We care a lot about resiliency. That is our number one guiding principle
especially for our network, is massive reliability.
And obviously, we want to contain outages to within a region,
but we never want to have them span to a region.
We go to the level of isolation,
even with inside of individual data centers,
we're building isolation at every level to get to that.
In terms of resiliency, I've definitely noticed a trend.
There's increased risk appetite, I would say, from a lot of customers,
who they want to move faster,
and they're willing to trade off some of the multiple nines of availability
that traditionally have been a primary focus for all customers.
But I think that's a smaller number of customers.
I think most customers still care deeply about resiliency.
I talk to lots of them all the time, and they're very focused on resiliency.
They're pushing harder and harder for higher availability, higher reliability on the network all of the time, which is great.
That helps us all be better.
I think for us, it's a balancing act of how do we build great services that have all the level of resiliency that some customers want,
but don't slow ourselves down with all the resiliency in order to move as fast.
we possibly can. So that's the challenge that we're facing right now. It's interesting. The resilience
demands are much higher in cloud than they are in the era when most companies ran their own data
centers just because of the correlation risk. Like, great, today, my bank is down tomorrow,
your grocery store is down. Well, you're bad at running websites and so am I at the end.
When AWS has a problem, suddenly lots of companies are suddenly impacted, even those who have
previously done a great job of planning for resilience, but, oh, no, we have a critical
path dependency on a third-party vendor who did not realize they themselves had that dependency
exposure. Everything's interconnected. Some of these are circular. And the fact that this is not
broadly appreciated, understood, et cetera, by, even by many engineers and technologists, again,
is testament to you folks getting this very right, very early on. Yeah. Which brings us to,
I guess, the showpiece thing that you wanted to show me in your networking lab,
which is the, I guess, the embodiment of the jellyfish paper.
Can you explain that in a way that, I mean, I could make an attempt at it,
but I'm going to sound way dumber than you will. Please take it.
Yeah, as you said, like when AWS was building the cloud,
we knew customers were trusting us with their business.
It's a tremendous responsibility.
And in order to convince them that this was a good idea,
we knew we had to have nearly perfect availability.
We had to be far superior to what they've been able to do.
And again, that's always been a guiding principle, especially at our infrastructure, network layers.
Network is underneath everything.
If you have a network issue, it affects many services, many customers.
And we've had a very, very reliable network for many, many years.
We're very proud of our network.
But we saw an opportunity to make it even better, and particularly in terms of its availability and resiliency.
And so the new network is called resilient network graph, resilience, or RNG, for short.
And the idea is the way networks have been built for really the past 20 or 30 years, their hierarchical, their factory network.
for people who know about networks.
And it's really stacked layers of switches.
And so there's a tier one, there's a tier two,
there's a tier three.
They're all interconnecting into each other.
And that's how you get big scale in networks.
You have aggregation switches that are on a leaf.
You have edge switches.
You have a core switch.
You've an out-of-band switch for the managing the servers that like to break.
Yes, all the things.
And so many layers of switches.
And for a long time, people have asked,
why do you need all those layers in your network, Matt?
Why can't you have a flatter network?
And mathematical theory will show you that is actually the most efficient
way to build a network, both from our cost perspective, but also from a reliability perspective.
The problem with building a flat network is all of the interconnectedness that needs to happen
to plug everything together. At minimal scale, you can take a switcher that might have like 32
ports on it, and if you have a small number of switches, you can connect them all together.
Easy, right? If you have many, many thousands of switches, you can't plug them all together.
You don't have enough ports to. So you need to effectively randomly interconnect them with some
links and then you have to figure out how to route through multiple different switchhops to get to
your destination. There's no structure like there is in a hierarchical network where everything is very
clean and orderly. So math has taught us for years. This is the best way to build a network.
Many people have tried to build networks like this. No one, to my knowledge, has ever actually
built a network like this at any sort of scale. We decided three or four years ago that we think
we can actually pull this off. So we started working down that path and doing a bunch of research,
doing a bunch of simulation. This is a very novel concept.
Most of the network engineers in my team thought this was a really bad idea. Don't do this, Matt.
This is not how we run networks. It's going to break. It's going to be weird.
We eventually convinced ourselves through a lot of simulation that, no, this will work, and it will work more reliably than the network that we have today.
And so we invested, built an engineering team, did a lot of work, actually started to build these things for real.
Learned a lot of lessons. That's one of the other keys to resiliency is learning a lot of lessons and actually taking the learnings from those lessons and then trying to, like, bake.
them into your product. And so in order to build this for real, we had to build this in
smaller scale, we had to test this out, we had to see how it was going to fail the problems
we were going to have, and then iterate quickly into that before we could deploy at scale.
And today we're in production in multiple data centers, and this is now the new default
network for all core services for AWS, for any new data center we're building.
We're basically we've switched to this new RNG flat network from the traditional networks
we used to build.
My take on this is not that it's not going to be possible for you to do. A real brave take
on my part, given you already done it and
proven it, yeah, with the benefit of hindsight, I sound real awesome.
Now, my question is, why do it at all?
Why do it at all?
Exactly, because at some point, it's no longer your bottleneck.
Your network is already ridiculously reliable.
So at some point, making it even more reliable seems like that's not the bottleneck.
That is no longer the area of focus, of contention that is customer exposed.
Why spend the investment, research, energy, hardware time, et cetera, on that rather than other areas
that are more, I guess, directly perceived as impactful to customer workloads.
Yeah, I see any outage in the network as a problem.
And while our network is extremely durable and customer applications are highly reliable,
we still see behind the curtain and we still see the opportunities and the risks.
And it's just we want to make this better for our customers.
Like, to me, the network is still in the way in the sense of you know it's there,
you feel it's there.
Sometimes you still have network degradation.
If we can make it even more reliable, it becomes more invisible and it enables more
customer workloads. And so it's part of our customer obsession, I guess, at the end of the day,
of like, it's just not good enough. And I think it never will be good enough, really. It's just
a constant work that we're doing to continually try and drive improvement in this. And not all
customers will care, but some of them do. And therefore, it's worth it for those customers.
Somewhere between 10 and 20 years ago, I wound up helping with a cluster buildout for more or less
a 256 node HPC cluster alike, where they wound up putting a whole bunch of
servers into racks in a data center cage that they had. Great. Awesome. And the idea was that customers
would then pay to basically load their data into this, do all kinds of number crunching. And it was a
multi-petabyte cluster, which at that time was reasonably impressive. Today, you're like, huh, that's
cute. I have S3 buckets like that, which different era. And a question I had for them during the
planning process was, okay, great. I see the aggregation switches at the top of each rack. I
see the core networking structure here. Great. Awesome. So how are these petabytes?
the data getting here exactly. Oh, customers will ship us a pile of this. Awesome, great. And then we
plugged them in over there on the bench in the rack. See, we thought about this. Stop with your
impertinent questions. Cool. Just one more. What is the sustained throughput rate between all of those
switches and when you saturate, even assuming line rate, how long does it take to fill all of those
servers throughout there? And the answer distilled down to, oh no, because it's a narrow pipe problem.
How do you do this?
And that was not well understood.
These are smart people.
I'm not trying to dunk on them.
It's the sort of thing that's really obvious the second time.
The first time you have questions on this.
And even now, companies that are born in AWS and are exploring, well, what if we move this workload to our own data center?
Let's go ahead and build that out.
They are misled by a very strange aspect of AWS that I confess I don't fully understand myself.
And I'm hoping you can shed some light.
In my experience, you can do this concurrently with basically everything you imagine, pick any two points in the same availability zone, in the same cluster, etc.
You can get damn near line rate network transfer between those points.
And then when I have a bunch of things doing that simultaneously, I don't see a subsequent degradation.
It is still right there at line rate.
And the only answer I have on that is witchcraft, which...
which, yeah, from a certain level of ignorance, everything seems like magic.
I get it. How do you do that?
Yeah, we overbuild the network. I mean, it is that simple.
We build more network capacity than we need, both from a reliability perspective,
so that we have a lot of redundancy and resiliency.
Many devices can fail. Many links can fail.
We still have sufficient capacity for all of the customer traffic.
And so it's really is that overbuilding and building as much capacity as we possibly can
to create that illusion of elasticity or that illusion that basically there is no network there.
these servers are all just directly wired together.
You can send as much data as you want between them all the time.
And that's really been our mission for the last 15 years is if you want to do that and you
want to do that at massive scale, you have to do it reliably, but you also have to have a cost
structure so that you can affordably do that.
A lot of the reasons there are network problems, it comes down to constraints or lack
of capacity, and then people get creative in that they want to like maximize the usage of
the network and you get into things like quality of service or how do I do traffic engineering
or how do I move this traffic around?
If you just had more network, you don't actually need that stuff
and your network runs much, much more reliably.
How do you even with QOS?
Well, I don't know if you have the capacity
to put it all through with the same latency target.
You don't need it.
Yeah, and that's been our mission for many, many years
is we really try and keep our network very simple.
I mean, that's another secret to our resiliency.
We try not to do creative or fun things in the network.
Keep it as simple as absolutely possible.
Have a lot of capacity.
Have a lot more capacity than you actually.
need so things can fail and you're still fine. And running a giant network becomes something
you can actually achieve. But the key there is it's the cost structure and the ability to have
that much capacity. Like how do you actually pull that off? For us, it's we started investing
and building our own hardware. And so taking control of our own hardware, we did this about 15 years
ago and started this journey, meant that we could be really prescriptive about just the things we wanted
in the hardware and in the software on our devices and not take along everything else that
when you buy from a vendor is going to come along.
So you'd have that section of the firmware in case you wanted to plug it into a
Melanox thing later.
Yeah, we know we're not going to do that.
Why bother taking up the RAM?
Delete that code.
It's one less thing that can fail.
Again, fewer features, fewer functionality, simple, tailored to your use case specifically
is really the key there.
And so over the last 15 years, we've again iteratively matured.
We keep making it better and better and better with every generation.
And you get to this point where today, 100% of the AWS network
is built with devices that we design ourselves.
And the other fun secret about the way we build our network,
we actually use the same switch everywhere.
So again, most you get these hierarchical networks we talked about before
with core and aggregation and edge and all these other layers,
but also each one of those layers would use different type of router
for different fit for purpose.
And there's good advantages and reasons to do that.
But then now you have more complexity,
you have more device types to manage.
They all have their own little nuances.
We said, well, why can't we just use the basically our top of Rex?
switch, and why don't we just use it everywhere? And we've achieved that at this stage. It took us
about 10 years to fully get to the point where we could do that, but we're bullheaded and we just
kind of continue to push forward and get to that stage. So now, like our entire internet network,
our backbone network runs on the same top of rack switch that sits in the top of rack switching
and connects to EC2 servers. I have memories misspent youth of driving a van with a core switch
in it that costs more than the band we had rented to do this. It's like, I'm pretty sure we're not
insured for this, but that's the boss's problem, certainly not mine. And we didn't hit anything,
so fortunately it was no one's problem. Yeah, it's the world has changed. The way you address
these things has changed, and it's wild. One caveat, I know I'm going to get comments on this,
otherwise if I don't say this. Like, well, yeah, you talk about economies of scale, that's why
data transfer is so expensive. I get it. But in the context of inside of a VPC, inside of a
subnet, you're getting that full magic line rate between two endpoints,
assume both EC2 instances, that is free.
There is no additional charge metered to customers until it starts crossing other boundaries.
So it's not just that you've overbuilt and made it super awesome.
It is not charged explicitly for the most common use cases where those things matter.
And I want to make sure that nuance is clear.
Because otherwise it's, well, yeah, if I was charging X dollars per gigabyte, I too would invest
in making sure you could shove as many things as possible through it.
but that's not what's going on here.
No, that's not charged.
And as availability zones get larger,
again, there's a magic of elasticity
that comes from AWS.
Behind the scenes,
those are actual data centers
filled with devices
that all have to be interconnected together
as we scale out these availability zones,
which means more and more and more network,
and that's also a driver for us
to drive cost efficiency,
and so we can keep it free effectively.
Like, we don't want to charge for that
because we want to get the network out of the way
and let customers just move their data
at whatever speed they possibly can.
And that's valuable.
As long as people can
that it's not going to cross those chargeable boundaries, which is where it's not even that it's too expensive.
It's that I thought it was free, and it's not. That scares people.
One of the bit of magic in here that I talk to folks, even folks who have networking backgrounds there,
right, you look at the typical, you start inspecting the traffic that's going over the wire.
You look and you see the physical, the back address on the physical layer, layer two.
You look at the IP logic around layer three.
You look at TCP and all the stuff that's being built in as those things happen.
the entire network that you see as a customer is fake.
It is an emulated, virtualized, imagining of a network designed to look like networks looked
25 years ago, and it is entirely living on top of what the reality actually is.
And that is wildly exciting.
Was it always that way?
This episode is sponsored by my own company, Duck Bill.
Having trouble with your AWS bill, perhaps it's time to renegotiate a
contract with them. Maybe you're just wondering how to predict what's going on in the wide world of
AWS. Well, that's where Duck Bill comes in to help. Remember, you can't duck the Duck Bill
Bill, which I am reliably informed by my business partner is absolutely not our motto. To learn
more, visit DuckbillHQ.com. Not in the very, very, very first days. When EC2 originally launched,
you could see the actual network within about a year we've
quickly realized this was going to be a bad idea and that's when we invested in building VPC.
And since it's really, I think, since about 2007 or 2008 when we introduced VPC, everything
from that point was virtualized.
And the whole idea was make a very simple virtual network for customers, abstract it and
separate it from the physical network.
That way your physical network can change and you can do whatever you want on the physical
world behind the scenes and you're not really disrupting customers, you're changing the
way customers perceive their network to function.
And it's been very powerful for us to have that separation.
I like to tell people I run the real network at AWS, because it's the physical network.
It's the network network.
I love the fact that it sounds like you're talking smack.
It's amazing.
Like, well, you know, those fake network things.
I don't know.
And we have amazing teams who build all of the network services the customers use.
And it's super powerful.
Like, I work with those teams, but my teams are not tightly coupled, right?
Like, we can build our network relatively separately from the services that are delivered on top of that networked customers, and that lets us all move faster.
One thing that I have always wondered about is the reason that the internet exists is the idea of interoperable standards.
And the one, there were a bunch of early things that came out like, oh, Apple Talk, sure, great, IP, ISP, what is it, IPX, ISPX, that, yeah, things that I don't even remember because they did not win.
TCP, on top of IP, did.
Very durable.
But if you look at the protocol definitions and see how they are structured, what they are built for,
it's the reason the internet works.
It is designed for a wide variety of network environments
with a wide variety of ever-changing conditions
in those environments.
Inside of a virtualized network,
like you have built,
a lot of the things that those services and protocols
have been built for historically will not happen.
You can state that with a certainty.
Like, we didn't build the ability
for the simulation to ask about the nature of itself
or whatever the challenge is.
Does that mean that there is a path
to a more,
efficient protocol for some workloads as long as it doesn't have to start dealing with the rest of the broader internet.
How do you think about that?
We already have that.
You do under the hood.
I know that much history.
There's SRD that you were talking about a few years back.
You announced this at re-invanced.
It was great.
We have this amazing TCP replacement protocol that we use at the time.
It's two power, I think, is EFA and a lot of the higher EBS.
Everything behind EFA runs on SRD as well as all EBS traffic within NADIS regions.
Yeah.
So all the storage traffic.
Yes, and it was like, wow, that's really interesting.
Can I learn more about it?
The answer was basically, no.
And cool, so why are you telling about it?
Honestly, it's really neat and we're proud of it, and it's how we do these things.
But outside of the context of how we run, what we do, how we see it,
it is not something a customer would even find useful, much less makes sense to even go too far down the path.
I'm talking about the other side of it.
When once you're inside of the virtualized network, once you have, I have an EC2 instance of talking to another,
and I want to send a bunch of this data over very quickly.
As the connection stands up, you start seeing TCP window scaling,
start slow, speeds up, et cetera.
If I have certain guarantees around this,
I could see, and these are famous disasters,
I could write my own protocol and put traffic over that.
I've been involved with the project once that did it,
and I was there.
It's hard enough to see them turn it off.
Yeah, it's really hard.
But I could see you folks putting something like that out.
We have.
So, SRD, you can go to an EC2 instance right now
on your E&I, and you can enable basically SRD transport under the covers.
It'll still look like TCP that you're communicating with,
but behind the scenes, we wrap your TCP packets in SRD,
which effectively makes them extremely reliable, lower latency, higher throughput,
and send the TCP bits.
And so from a TCP perspective, it just kind of looks like you're on a magic network
with infinite bandwidth and no packet loss.
But that way there's no...
That one is snuck completely past me. Fantastic.
But there's no change for the customer.
And that's the big thing, like, you can adopt DFA, right?
But if you are adopting EFA, there's a lot of work for you to do in your application so that you can make it EFA capable.
And a lot of customers just won't do that or can't do that.
And we wanted to do it.
Like, how do we make the network better for all of them?
And so right now this is an opt-in feature that you can go turn on.
And we're busily working behind the scenes.
How do we make this the default?
Like, how is this just the way it works for everyone?
And there's a bunch of technical challenges, but we're chipping away at that.
Today, at least, what are the workloads for which that makes sense?
And which is one of those...
Almost all workloads, honestly.
Okay.
Is there any exception cases?
Do not use it for that one thing?
The exception case is there is a tax of maximum packets per second.
So in order to do that encapsulation on the hardware layer,
the peak PPS you can get is reduced slightly.
That's the reason we haven't gone and turned it on by default for people yet.
Again, we're chipping away and reducing and kind of removing that bottleneck.
But for the vast majority of workloads,
unless you're very, very sensitive to peak PPS,
enabling that functionality gets you more reliability in your connectivity.
And the real way you see this is if you're looking at like tail latency,
for your application, you're looking at your P99, P99,
P99, latency, that's usually going to be dominated by some type of packet loss or some other
issue.
And if you enable this with SRD-based transport, mostly that disappears.
And that's the reason we've done it for EBS.
Tail latency matters a whole lot when you're talking about storage workloads.
And so we were on a mission to be like, how do we drive that tail latency down?
And that's where we came up with SRD and moved all of the EBS traffic over at SRD.
Okay, I've seen it all now.
All that's to do is talk to one more AWS customer and I see things that are new.
The challenge I always ask, okay, where's this not an appropriate fit for is because very often,
I will talk to folks who hear about these things like, great, we're going to go and implement
that everywhere.
Great.
What does your workload look like?
Oh, it calls out to an AI inference provider, and it waits for the response, and the response
comes back, and we're getting fast speeds.
We're getting almost 200 tokens a second on that.
It's, I promise, the network is not your bottleneck right now.
It's fine.
This is not the place to optimize.
There are better paths forward for you.
I'm excited for the day where this does become the default
with the obvious caveat then
that I can see a lot of workloads
that people are going to try deploying somewhere else
get disparate results as a result of this
once that becomes a default and start to wonder.
I mean, sure, the blame is going to accrue
in directions that are not able to you
because you're the good example in this.
How come it's so crappy in our data center,
probably our terrible networking team,
which is not true.
It's just this weird, almost under the hood
magic. Yeah, and really, like, something like SRD and our capability to do that, it's similar
to our network investment and we want to own our own devices, we want to own our own software,
and we've done on the server side for years with Nitro, we have our own hardware, we abstract
all of the customer bits with VPCs, which means we can do interesting thing behind the scenes
and we don't have to change, you know, many, many years-long standards like TCP. You want to
change TCP? Good luck. And you want to try and change everyone's applications. The way to make
meaningful changes to leave TCP alone, but just basically magically make it better by doing
embrace and extend. Yeah. Yeah. So as you talk to people about what you've been doing for the
last, well, 15 years at least, what do you find is the biggest misconception that customers have
about AWS's network? About AWS's network? Most people don't know what exists. I agree wholeheartedly.
Yes. Yes. And so it's just the lack of understanding of the sheer scale and all of the investment.
all the work that goes on behind the scenes to just make all of this function and work for customers.
And friends and family and everyone else, they just see it as, well, there's just, you know,
some magic happens to the Internet and things happen.
A while back in one of his iterations, Peter DeSantis, before he transitioned to his Amazon role now,
was the SVP of utility computing.
And I love the term because it encapsulated an experience, I think, lots of folks have, even if they didn't notice it.
When something is a utility, like electricity or water, you don't turn on the switch or turn on the faucet and then wonder if it's going to work this time. It is expected.
And we've all gone through that as consumers. There was a time at some point where when I go to google.com in my web browser and it doesn't work.
In the early days, Google must be having an issue. And at some point it switched for all of us, oh, my Wi-Fi must be having problems.
It's just my problem, yeah.
Exactly, to the point where on the rare occasion, I can't think of one in the last 10 years,
but toward the end of that period, when Google would actually be down and have an issue of serving.
I didn't realize that was even a possibility.
How could that happen?
And I really think the network at AWS site is in so many ways, almost a victim of its own success.
The electricity example is exactly what we have in our head.
We want the network to be like the light switch.
You flip it on and off.
You expect it to always work.
and it's extremely rare and surprising when it doesn't work.
And that's been our mission is make the network invisible, get it out of the way.
A question I often get from folks is, great.
How did you get started?
How did I become me?
And I started off as a Unix followed shortly by Linux, Systems Administrator.
And during the global financial crisis in 2008, suddenly salary freeze.
Everyone hates their job, but no one's hiring.
What do I do for the next year?
Well, I started learning networking because that had always been an area.
I hand waved over.
And I did dabble as a network engineer after that.
But by and large, it made me a much better systems person.
Because once I understood what was on the wire, I could reason about it.
I could understand what was going on.
The challenge now is that we have complexity on top of complexity, on top of complexity, and so on and so forth.
It is dizzingly high now.
And I worry on some level that networking as a whole is no longer as central and core to modern engineering understanding of this.
There are publicly traded companies built in the cloud that do very well and their internal networking teams.
Their entire professional expertise is more or less around configuring things like Transit Gateway and CloudWand and these are all abstractions on top of it.
They don't understand even the old days of MDIX AutoSense, which was a consumer feature then came to.
It's like, great.
When I plug the Ethernet cable in, it's not lighting up.
Why?
Oh, that's a rollover cable.
You need a straight through.
Wait, you mean there's different standards for the wide.
I have a mug that just has the colors of the Ethernet B spec, and every once in a while someone
sees it, like, yes, like, okay, I find my people. It's great. But now I sound like an old man,
talking about the Great War once upon a time. Where is the next generation of all of this come
from? Yeah, it's something we grapple with. And like, again, we've been very successful such that
most of our customers don't have to deal with the vagaries of networking. And you like it.
Most people look at that and they're like, that's terrible. I don't want to do that.
I can, I can watch nostalgic about this.
If I were on the phone this morning with Cisco TAC
because of a weird routing issue,
I would be considerably less charitable,
but that's 10 years in my past.
For us, it's really grow your own.
And so, like, I mean, I joined Amazon
as a very junior engineer back in 2008,
and my initial job was kind of doing monitoring of the network
and, you know, seeing things breaking and learning to fix it.
And I've stuck around long enough and learned enough
that managed to grow into this position.
It's kind of the same path we have
with all of our new hires. We're hiring lots of junior engineers, either from, we hire lots of
people from our data centers, actually. Some amazingly talented people who were there. I was just
talking to someone who's visiting one of our data centers and someone who's working at a gas
station three years ago, and they were just wowed with what they knew. And I'm like, we should go
hire that guy for the networking team because they actually touched the physical equipment.
They understand how this stuff works. And they often become our best engineers in the future.
But it is a deep investment in us in terms of like hiring younger talent, junior talent,
and like going through the process and training and growing them and teaching them how we build networks.
Because at the end of the day, there's no accelerator for this.
It is increasingly complicated, as you said.
It's much, much more complicated than it was 15 years ago when I was starting.
But you still need to know the same foundational things that I had to know 15 years ago.
But you also have to know this increasing mountain of new stuff over the top.
And that only comes from just learning the basics, learning the practical bits,
and then slowly adding the knowledge on top with experience.
Because you know that, customers don't have to.
I want to highlight just how rare that is.
It seems that the idea of developing talent internally
has almost gone out of fashion in the industry.
It's, oh, no, this person doesn't have experience.
We can't hire them for the role because it'll take six months
to get them up to speed.
Yeah, I get it.
Nine months later, the role's still open.
So what are we really doing here, friends?
The idea of teaching people and arming them to do this.
Because there was a time where every company
that wanted to be on the internet,
which, let's face it, is probably most of them,
needed to have a networking, if not person, team.
Where it was, you need people able to handle this stuff.
It gets really weird, really fast.
And now you don't need that.
I mean, I went through a parallel evolution
running email systems.
Now most companies do not run their own mail service.
Nor should you.
Thank God.
Yeah, networking is still continuing to evolve,
and I'm not naive enough to think
that we are at the end of history.
And now, final form, what's the next step?
What's the next phase?
Where do you see this going?
Where we see this going?
It's really more of the same.
I mean, there's so much, there's been an acceleration and investment in networking because of AI,
which is exciting as someone who does networking.
There's bandwidth demands are higher.
There's more capacity to build.
There's more things to connect.
And so I've seen more energy go into networking in the last three years than I had in the previous 10 years.
And so there's just accelerated investments both at the silicon and the hardware level,
but as well as the software, the protocols, everything else.
It feels like networking is exciting again in a way that maybe we made it fairly boring by the time you got to 2019.
I will say, when I misspent you, there were times I wish there had been a lot less exciting some nights.
Yes, yes.
Yes.
But yes.
And so I think it's just going to be more of that.
Like we're going to continue to find the world has an insatiable appetite for bandwidth.
And I think we're only scratching the surface of that.
And if we were able to 100x the internet and network capacity, it wouldn't take long for applications to find uses for that.
And I think there's still a lot of things that are not done or they're done in ineffective ways because of either not enough capacity of network and or like costs of networking.
And so the more you reduce that and eliminate that, I think you just open up new opportunities in innovation.
And so I think it's just going to be this constant march for more and more and more and more capacity.
We've seen that in real time.
The thing that drove that home to me more than almost anything else was the day that GDPR took effect.
Because a lot of U.S.-based websites suddenly were not fully compliant.
So they turned off all the tracking and all the rest.
And suddenly the internet was blindingly fast.
It was like, wow, you're loading an article
and like the total load was something like 50K.
It was wild to me.
And you look at the same article
with all the other stuff turned on.
It's why is loading that 50K article
taking 25 megabytes of nonsense?
As soon as you have the capacity,
people find ways to fill it.
I'm not opining on whether it's all used.
Or just look at streaming video, for example, right?
Like, I mean, you can get very good quality
with the latest codex, but it's not the best quality.
Like, real broadcast-level quality video
is like an order of magnitude or two higher bandwidth
than what stream to you, even if you're getting
like the 4K, ultra, whatever, on your fire TV.
And so the reason is bandwidth.
There's always need for more bandwidth,
and there always will be need for more bandwidth.
And so the more we can do to make it cheaper
and more reliable and just more prevalent,
we're only going to be benefiting our customers.
You have the news anchor trucks
that show up at the scene of something going on.
They have the big satellite dish on the roof.
It's like, do you think that's because they haven't quite figured out that you can tether to yourself all?
There are some very dedicated, very concerning bandwidth requirements around this.
And it's one of those areas networking more so than many, but complexity passes everywhere,
and it just slips below the baseline surface of awareness.
Thank you so much for being as generous with your time as you are at explaining how some of this magic all works.
If someone's watching this or listening to this, depending upon their point of view,
and they want to learn more about how networks work and how infrastructure happens,
Where should they start?
Where should they begin learning about this secret thing that still drives our entire world?
Yeah, it's again, it's going back to the raw material.
And there's a book written, I think, the 70s or 80s.
It's the TCPI-P-I-P illustrated book by Stevens.
It's where I'd suggest to go.
Like, that's where I send anyone who wants to get into networking and learn something new.
Because those same fundamentals, like you talked about, of all those standards,
they still matter.
And the entire world is built on, you know, versions of them that have evolved over time.
but that fundamental skill set will take you a long way,
and you have to have it to really understand how these things work.
And, of course, we're linked to that.
I love that book myself, though.
I will say some of the bandwidth references in there of like,
oh, that's a really fast 10-based e-connection.
That might not have aged super well.
It just shows you how far we've come.
But yeah, I just had someone, actually one of our data texts this morning,
was messaging me asking me, how do I learn about networking, Matt?
That's exactly what I sent him.
Go get this on Amazon.
I have no idea if it's still up, and I'll check it.
If not, my apologies.
But there was a great site,
Warriors of the dot net was always a video that explained in a five or six minute cartoon approach
of how packet switching routed networks work in an extraordinarily accessible way.
I hope that's still out there.
That was always a fantastic primer.
Yeah, some of their terminology choices you can argue with and like, well, there's a lot more to it there.
Yeah, but you're going from zero to one and you're giving people a chance to see what's next.
And huh, that's curious, I want to dig deeper.
The thing that actually tipped me over the edge was subnet mass.
Okay, it's this weird number string.
And all I know is that when I get it wrong, some things work and other things don't and I don't understand it.
Maybe it's time I stop hand waving over it.
Yes.
And for my sins, I learned how networks work because I didn't know how networks work and I didn't believe it was possible.
And now that I know better, I do not believe the networks work and I'm astounded that it's possible.
It feels like it works in spite of itself, but clearly it does.
It is tremendously complicated, but it gives us a lot of pleasure to make these things simple.
for everyone. Thank you so much. Matt Raider, VP of Global Networking at AWS. I'm Corey Quinn. Stick around.
