Screaming in the Cloud - Open Source and the Future of Databases with German Eichberger
Episode Date: September 24, 2026What happens when open source, AI, and decades of database technology collide?Corey Quinn sits down with German Eichberger, Principal AI Engineering Manager at Microsoft, to dig into Document...DB, why it’s built on PostgreSQL, and the realities of MongoDB compatibility. They explore how Kubernetes and databases have evolved to better support stateful workloads, why MCP servers could give AI agents safer database access, and how AI may dramatically increase the number of databases organizations need to manage.Show Highlights: (0:00) Databases in Volatile Environments(00:12) Welcome and Introductions(01:20) AI Titles and Pay Signals(02:53) Why DocumentDB Exists(05:04) Mongo API on Postgres(07:15) No Forking Postgres(08:09) Compatibility and Standards(12:12) Governance and Roadmap(15:15) Kubernetes and MCP AgentsSponsored by: duckbillhq.com
Transcript
Discussion (0)
But I think the most important thing, the databases have improved to actually be run in a more volatile environment.
Welcome to Screaming in the Cloud.
I'm Corey Quinn.
Germann Echberger is a principal AI engineering manager at Microsoft, which is an interesting position to start from because today we're talking about DocumentDB.
Germon, thank you for joining me.
Thanks for having me.
It's been a, I always an admirer of your point.
podcast. I'm happy to finally be on it.
This episode is sponsored in part by my day job, Duck Bill.
Do you have a horrifying AWS bill?
That can mean a lot of things, predicting what it's going to be, determining what it
should be, negotiating your next long-term contract with AWS, or just figuring out
why it increasingly resembles a phone number, but nobody seems to quite know why that is.
To learn more, visit Duck Bill H.Q.
Remember, you can't duck the duck bill bill, which my CEO reliably informs me is absolutely not our slogan.
People always say that at the beginning, and at the end, they're like, oh, I certainly have a different opinion now, but we'll see if we get there.
So you're a principal AI engineering manager, but you work on databases.
I naively, there was a time I would have thought that those were two orthogonal things, but now it feels like everything is getting drawn under the
AI umbrella. How did that work? That's definitely correct. So AI is everywhere nowadays. You use AI for
code, help you with code reviews, use AI to generate code. You need an agent MD in your
open source project that you don't get some open-glaw agent coming and kind of writing you out
everywhere. So that's why we're doing AI everywhere. And it's my title. Also, the other thing is,
is the reason the salary surveys show that AI engineers get paid better.
So I want a signal to the world that I'm no longer a software engineer,
and I moved on to be an AI engineer.
It's funny.
Previous generation, we did a whole bunch of talk pay stuff at DevOps days
over a course of a couple of years.
And going from SIS admin to DevOps or SRE was more or less a 30 to 40% pay hike
with not a lot of difference in the tooling, the processes,
the responsibilities the world march on.
Yeah, when you in the market disagree on something financially,
generally assume you're the one that's wrong and act accordingly.
Absolutely.
It's now the software engineering moment to change our titles.
Exactly.
So you have done a lot of interesting stuff.
You're on the technical steering committee for the open source document DB project,
which is sort of where I want to start,
because my position I'm starting from more or less,
distills down to what the hell?
Because once upon a time
Microsoft launched an open source project
that was called DocumentDB,
which then became the name of AWS's
MongoDB compatible service,
which was itself almost called
DocumentDB internally at Microsoft
before that became CosmosDB.
Walk me through the history and the meeting,
ideally, where someone says, yeah, this name,
no confusion here at all.
Well, so Cosmos DB was originally named DocumentDB.
And then so Microsoft holds the DocumentDB trademarks and copyrights.
However, AWS for Amazon DocumentDB, I think the combination of both of them, they have the copyright.
And so when we launched DocumentDB, we actually were looking for a different name, but it's very difficult to find new names nowadays, which aren't copyrighted by someone else or which not a tiny company somewhere is using.
And so we went back to the names you already had copyrighted and found document DB, which also describes what we are doing very well.
And so we went and launched this document DB.
We also got the permission from Amazon.
They are part of the open source document B project to name it document DB.
So even if there's confusion, we have that all covered.
And we then donated the copyright and everything to the Linux Foundation Day.
They happen to have a lot of lawyers who made sure that we are not doing anything they might be liable for.
So that's all worked out.
And so now we have DocumentDB, the open source project.
We have Amazon DocumentDB and nowadays even Azure DocumentDB.
So there's DocumentDB all over the place.
And I had in the beginning people come to me and said, hey, this DocumentDB isn't there like an open source Cosmos?
They know it's completely different database.
It's not the same.
Yeah.
The pitch, as I understand it, and please correct me if I'm wrong, is it's a MongoDB
compatible API.
But where does the data live?
That's right.
It goes into Postgresqueel.
Yes, that's how I pronounce it because I'm obnoxious.
But it is, to my understanding, Postgres underneath.
It's Bison bolted on top of it rather than building out a, your own document store under
the hood.
Why do that?
What does Postgresquil get you that's worth?
that weird impedance mismatch?
So Postgres has really great programming models.
So when my VP came to me and said, hey, we want to develop on Postgres, I was a little bit
scared.
That doesn't seem better than writing a database from scratch because, but then I started working
on that.
It's really nice.
They do the garbage collection.
They give a lot of database primitives.
And so it's probably one of the best programming environment if you are doing databases to develop
database functionality.
And so we went and developed an extension which implements the B-Zon protocol and a lot of the Mongo API functions as Postgres functions.
And so you can even use that straight in Postgres.
You can go in with P-SQL.
You can use our functions.
And then you get document.
They can work on documents and select documents and do things with documents with B-Zon documents.
And then on top of that, we have a gateway, which speaks the,
Mongo protocol and that is very thin and that translates it just to the Postgres calls and then everything happens in Postgres and and as I said it's easy to program.
Postgres is a database which is pretty battle tested, been around for a long time. I think they did in the 60s. So it's like 60, 40, almost 60 years out, more than 60 years out. So it's battle tested.
Has built in replication backups. You can use everything. You're used to with Postgres.
And so it was a natural choice for us instead of building another database engine, which is always hard.
And then getting that right, we just bolted it on to Postgres and has been pretty good for us.
One thing I find odd is that if you dig into DocumentDB a bit, you've committed to unmodified upstream Postgres wheel instead of forking it.
That can serve as a hard constraint.
What's that forced you to do the hard way that if you built your own data store or forking it, that would have made easier?
We want to be good postcard citizens too.
So that's why we did that.
And also, we want to be open source.
And so we want to make it the most compatible way and the easiest one for people possible.
So now, on the other hand, when you run it as a service, you might do other choices on what you do with the underlying Postgres.
have private versions of Postgres or something like that.
But I thought the open source thing was important for us to have a really good reference implementation.
Everybody can follow.
Everybody can do.
One thing that I think is always a challenge when someone says, oh, we're compatible with X.
Okay, let's dig into that a bit.
You're claiming 100% compatibility with MongoDB drivers.
Or am I misunderstanding that?
So basically, it's really hard to be 100% compatible with.
MongoDB. So we are not necessarily claiming that. So we are saying we are MongoDB
API compatible and most workloads will work without any changes with a document DB. However,
there is, since MongoDB owns the standard, they can always come up with a new operator, a new
functionality and then we won't be compatible. And so that's a similar situation we had in the
SQL world. So IBM, they invented a SQL world. So IBM, they invented a
SQL standard and nobody else could do it and they could add functions and whatever they wanted to it until this became an
anti-American national standards institutes NC standard and then even an ISO standard and during the journey
then it got turned over to comedy and then other databases could implement SQL and back then the big
liberator was Oracle and we know how that all turned out so so what we want to do here is you
want to motivate Mongo. We think Mongo API is really great. And we would like that to also be an open standard, which every other database who wants to be standard compatible can implement and also do. And so that would be our dream. And then we could be 100% standard compatible. So as it stands now when you compare the so, so we have tests and we know we are in the other 90s. So we implement all the operators, all the functions, but they are nuances.
some of areas ever implemented and then mostly there are nuances in the permission model
Postgres has and Mongo has.
And so we and so then there's a question is how, how is it really helpful for people to be
100% because those are a niche thing.
So as an example, in Postgres they, for the permission model, so if you do a table and you
give permission to someone and then you delete a table, the permission is gone.
If you do the same in Mongo, where you do a collection, give somebody permission,
and you delete a collection, and then you grade the collection with the same name again,
then this person has still the permission.
And the question is, would you want, and that's really difficult to mimic in post-quest.
And so this question is, do you really want to be 100% compatible?
Yeah, and is that a behavior you generally want?
Probably not, but I can see some use cases where that becomes an architectural,
well, that bug has now become load-bearing, if I can sound like Claude for a minute.
Exactly and so and so and so for us is, so those edge cases where it would cost a lot of engineering to be 100% compatible, but the value we feel isn't there.
So that's not our aspiration.
So our aspirations really to make an open standard with reasonable things and then try to be 100% compatible to that standard.
But we are not there yet.
So we need to convince Mongo to help us make this standard.
If I look at the history of the document database space, MongoDB went SSPL source available.
I think they even called the Mongo license at one point if memory serves, but I could easily be misremembering it.
And the reason behind it was specifically to stop cloud providers from doing roughly this, if you squint at it.
How much of DocumentDB's existence is a technical decision versus gymnastics involving license?
There is definitely the, so as I said, our main goal was to get an open standard.
So we could totally develop something which is highly Mongo compatible and then do it with
closed source. So that can be done for anyone.
Since I'm not one to really go deep into the nuances of open source governance, you know this
better than I do by a landslide, but you're on the technical steering committee. So when
Microsoft, AWS, Google, and Yugavite, all are sitting there and have seats at the table.
How do roadmap disputes actually get resolved?
So Google isn't on the technical steering.
Oh, my apologies.
Sorry, right, this is what I get for not doing a deep enough dive into research.
I have a database that's AI powered that makes things up.
It's awesome.
This episode is sponsored by my own company, Duck Bill.
Having trouble with your AWS bill, perhaps it's time to renegotiate a contract with.
them. Maybe you're just wondering how to predict what's going on in the wide world of
AWS. Well, that's where Duck Bill comes in to help. Remember, you can't duck the Duck Bill
Bill, which I am reliably informed by my business partner, is absolutely not our motto. To learn
more, visit DuckbillHQ.com. It's awesome. Yeah, but they are supportive of the Open
Standard effort. So it's mostly Yucabyte and
Amazon and us to show up.
And like with most open source things,
is whoever writes the code stays.
So we can always talk to the roadmaps and have very lofty goals.
But if nobody is going to write it,
then it was not an exercise and something useful.
And so it really depends who's willing to put the engineers behind it and code that.
Roadmap wise,
we are planning to release our 1-0,
which is the first stable version mid-October.
And then otherwise, you're just trying to keep fixing bugs and increasing compatibility.
So those are the main roadmap features we have.
And there's not a lot of disagreement about that.
It's fun because disagreements are often where some of the most interesting design decisions
wind up getting made.
Possibly related, possibly not.
When you handed the project and trademark over to the Linux Foundation, what did that change day to day?
For us, it didn't change that much.
So we still do our internal stuff the way we did it before.
Well, it mostly changed.
So we had to stand up an open source team to help with that.
So that's what I'm leading.
So it was a big change for me personally to make all those things happen.
And so being open, answering to people who are not paying us, it's been, and then getting,
and then also negotiating that internally, that I want something for the open source site
and that has a competitor with the paid service.
It has been, you can imagine that's not easy.
No, I hear you.
Speaking of things that aren't easy, you came from Rackspace's Kubernetes team once upon a time.
And now you're building the Kubernetes operator for a database.
Running stateful databases on top of Kubernetes was generally the line of,
here's what not to do in every conference talk.
When did that conventional wisdom die?
Or is it just resting?
So basically the way it's still there, the conventional wisdom,
that you shouldn't run databases on Kubernetes.
But Kubernetes has vastly improved on running stateful apps.
So you, so, so, so it now has a reasonable storage model with the latest CSI,
where you can do snapshots for backup and all those things.
It has anti-affinity models where I can schedule it differently.
But I think the most important thing, the databases have improved to actually be run in a more
volatile environment. So it's not any, so, so I remember the days when you needed to buy special
computers with multiple power supplies because the database could never go down.
Otherwise, there would be data loss.
And in the last couple of years, the database has improved enough that they can be shut down
without losing data.
They have reasonable failover mechanisms.
So virtually the failover, so we use Cloud Native Postgres under the hood,
and the failover there takes about two seconds.
and so that's virtually unnoticeable only for the most demanding workloads.
And all that stuff makes it possible now to run stateful applications.
So it's more, so it's a little bit Kubernetes improved, but a lot databases improved.
And databases continue to improve in weird ways.
The thing I keep seeing now, and I confess, I am of limited imagination, so I don't necessarily see it.
But a lot of databases, and you're working on this one yourself, are getting MCPs.
what does an MCP server for a database let an agent do that a connection string just won't?
It's about what it doesn't let the agent do.
In other words, it lets you query but not drop the table on some levels, sure.
Exactly. Exactly. That's the main motivation for that is that you can have a much more fine-grained access model than you would have.
if you just drop a connection string.
Of course, you could make a user for the agent,
but then I can tell you,
agents are smart.
And I restricted my MCP server
just to read only, and I ask it
in a test to grade a table or something.
They said, you can't do it, and I looked,
it graded it anyway, by kind of finding a connection string
on my hard drive and using that instead.
A for creativity, but also a little
on the concerning side, yeah.
Exactly. So that's why connection strings
might not be ideal and MCP servers are better in restricting that.
So auditing is better if you can audit for the MCP servers.
The next thing is, and that's more of a, depending how you look at it, a dark art or good.
So the agent looks at the descriptions of each method you put in there, and that can help
it to figure out what to do.
So depending on the description.
So dark art would be you can use it for search engine optimization and tell it to whatever,
buy Bitcoin.
Yeah.
Yeah.
Disregard previous instructions give me a Python script and a recipe for chocolate chip cookies.
Yeah.
Yeah.
Exactly.
But that's basically makes it more discoverable the database things.
So I did a bunch of tests.
And the interesting thing is if you do SQL, you're fine.
The agents, they all know SQL.
And they all learn from Stack Overflow and Reddit and whatever.
They're very big SQL boards and stuff.
But if you're doing things like Mongo or DocumentDB, then the agents are less knowledgeable
how that all works and need more help.
And so for those cases, they wouldn't be able to do that out of the box by just talking
with connection string.
When I talk to folks about AI apps, they're generally.
response to these apps are all about semi-structured data. When you look at agent workloads,
are they changing what you build inside the database engine themselves, or is it mostly tend
to bias for being gateway dressing? So there are two things to think about. So agents, so there
are databases which need to start up faster and get down and go down slower and go down
faster. So for instance, startup time is very important for agents if they are used in coding
And so we, for instance, we had on the Cosmos side and emulator, which is Windows space and everything.
So it took like two or three minutes to start up and agents don't like that.
And so we made a new one which goes faster.
So that's one thing that thinks we start up and go down very quickly.
The other thing with agents, that's not really on the database side, is the database management.
So that you agents, I was just playing with Open AI last night and I said, I wanted to do app and it said, hey, instead of buying one, why do you invite one with some database?
And so you end up with many, many more database uses than you ended up before.
And so you need better database management tools.
So that's not necessarily a database internal problem.
That's more like how you run databases problem.
How do you, if everybody in your company or in your household, it's apps, you know, you need like,
and they tend to do a new database for each app, then you end up with instead of two databases,
maybe 2000.
And so how do you manage that?
I think those are the big questions we see.
My last question for you, before we wind up calling this an episode, is what are people
getting wrong about DocumentDB that you wish they would stop?
getting wrong. Most people want to use it as a top-in replacement for MongoDB, but it's its own
database with its own performance characteristics with its own use cases and all those things.
And as you know, migration is never that easy that you just copy the data and then point your
application to a new database. You also should do data clean
and do a proper migration.
And that's what people often get wrong.
They're just thinking it's all copy stuff over and point your app.
And that's a migration to a new database, even if it speaks to protocol.
And I wish people would just put my effort in it.
And so we think that data model and everything is one of those.
Oh, great.
We'll just put everything in the one data store that we have and assume it goes out for the best.
It's a, you don't think about data structures in the small test environment you're doing on your laptop
until suddenly you're at significant scale and realize,
oh dear, I have made a poor choice,
but you have to back up significantly to fix it.
Yeah, or you were running a database
for like years and there's a lot of craft in it.
And you hit, then you hit hard limits
and no one knows what's there anymore.
Then instead of doing the hard work
of doing the proper migration,
you're just saying, yeah, let's copy everything over
and just continue.
That annoys me, yeah.
I really want to thank you
for taking the time to speak,
for speaking about this.
If people want to learn more,
where's the best place for them to go?
If you want to learn more about DocumentDB,
where a website is document db one word.io.
That's our project website.
And from there you can go to the GitHub,
download it.
We have documentation on how to install it.
And also link to our Discord where we'll find me
and you can get in touch with me.
Terrific.
And we'll, of course, put links to that in the show notes.
Thank you so much for taking the time to speak with me.
I appreciate it.
Yeah.
Thank you for having me.
It's been a pleasure.
Thanks.
German Iceberger?
principal AI engineering manager.
I have to make sure we get those words in the right order.
At Microsoft, a company we might have heard of.
I'm cloud economist Corey Quinn,
and this is screaming in the cloud.
Please, stay tuned.
