CppCast - Safety and Interoperability with Rust and C++
Episode Date: October 5, 2026Mathieu and Jason are joined by Jon Bauman to discuss safety and interoperability between Rust and C++. They dig into what "safety" really means (there are several categories of safety), what can and ...can't be guaranteed, how C++ could become a memory-safe language, and the challenges of making Rust and C++ work well together. Jon is currently looking for sponsorship to continue this important work. You can contact Jon via Joncppcast@shumi.org. Links Jon's talks - YouTube playlist P3874R1: Should C++ be a memory-safe language? - Jon Bauman, Timur Doumler, Nevin Liber, Ryan McDougall, Pablo Halpern, Jeff Garland, Jonathan Müller P3874 R1 Should C++ be a memory-safe language? - cplusplus/papers issue #2475 C++ Interoperability Initiative problem statement - Rust Foundation BorrowSanitizer - an LLVM sanitizer for finding aliasing violations across Rust, C and C++ C++/Rust Interop Problem Space Mapping - Rust Project Goals 2025H2 Miri example on the Rust Playground There is no memory safety without thread safety - Ralf Jung RustBelt: Logical Foundations for the Future of Safe Systems Programming - MPI-SWS Understanding and Evolving the Rust Programming Language - Ralf Jung's PhD thesis
Transcript
Discussion (0)
In this episode, Mattu and I are joined by John Bowman to discuss safety and interoperability
between C++ and Rust. John is currently looking for sponsors to help him continue in this work.
Check out the show description for more information on how to contact John.
Welcome to the 40012th episode of CBPCast, the first podcast for C++ developers by C++
developers. I'm your host, Jason Turner. Every fourth episode of C++ Weekly is a crossover episode
with CBPCast. You can choose to do.
watch the podcast on YouTube or listen to it on your favorite service.
Now, actually, I am considering splitting these things back up and having four regular episodes
of C++ weekly per month and then one also additional episode of CBPCast per month, if you
will.
So if our listeners have any comments, feel free to share that.
Otherwise, you can watch your favorite service.
I am joined by my co-host, Matu Ropère.
How are you doing, Mathieu?
Hello, I'm good.
I mean, it's been an interesting time.
The thing that I mentioned in the past couple episodes has happened.
If you have not been checking your LinkedIn, I have been officially made redundant from the company who hired me for a job four months ago.
Because for strategic realignment, they hired people for their consulting team and then decided that they didn't want to sell consulting anymore.
So that's been interesting to navigate on the plus side.
So are you actively looking for a job right now then?
Well, even for a job or like, you know, I have a severance payment and whatnot.
So I'm not in any rush.
I'm more looking like if you want any consulting.
Well, I guess I can still do that.
I guess now I can also give you consulting on Unity if you're interested in that
because I am a certified Unity consultant.
Certified.
From back when it existed and I worked for it.
And I guess you are going to have a harder time getting in from the company
because I assume they're going to have less of an offer of that service.
So I'm happy to provide it directly if anyone is interested.
I'll point you to some other folks who were working.
with me and are probably also looking for new opportunities right now. So it's been an interesting
time. I will just, you know, occasionally we bring up the differences between being an employee
in the US versus being an employee not in the US or more Europe, I would say, to the point.
And the idea that you have severance after only four months of working there is astounding to me
because I'm going to be honest, I don't think they had to and I cannot tell if it's because
they forgot or or it's because like I don't want to be too disparaging but I think the
the layout service is very optimized and it's very good at having like a template
layoff offer for people all the time and I think by by being very optimized and streamlined
on their process they might have forgotten what congratulations so you know I mean look
I would like to see that's just a way of saying we're sorry we didn't mean it like that
maybe that they forgot.
I honestly do not have the answer.
I did not rush to tell them that they could have done this differently.
Well, I mean, I don't blame you.
Then not asking for...
No, the third period in Sweden is six months.
It's usually six months and France is three months.
So, yes.
Technically, they could probably have done with less than a several deal.
But, you know, I assume that it's just because they like me so much and they were really sorry.
That's just how I...
That's probably the case, yeah.
If my background looks unusual to anyone paying...
attention. It's because I'm actually recording from a hotel room because this is the middle of CBP
Kong week right now. This week we are joined by John Bowman. John Bowman has worked in tech for over two
decades across various industries and programming disciplines. After several decades trying out
prescriptivism, they prefer descriptivism and not bundling gender metadata with third person pronouns,
but they will not mind at all if you call them, him, or even her. They are proud alumnus of the
University of Michigan, Etsy, Mozilla, and the Rust Foundation, working at every level of the tech stack
from operating systems to web applications, John brought their experience in Rust and standards
developments to bear as the implementer of AVIF support in Firefox, and the erstwhile lead for the
Rust Foundation's C++ Interoperability Initiative. As a result, John and the Rust Foundation joined
WG21. In less than a year of membership, John, along with Timor Dumler, which, yes, is the CBP cast,
host Emeritus and five other committee members authored P3874, should C++ be a memory-safe
language which received overwhelming consensus from EWG and the Croydon meeting this past March.
In their spare time, John enjoys the vibrant Seattle film community being a philosophy of language
dilettante. Delatante. Delatant. I thought I was doing well up to this point.
Villasant.
Spending time in the natural beauty of the Pacific Northwest and being a friend to all cats,
they are currently taking a well-deserved break from work,
but hope to return soon and help continue the effort to improve C++
and bring the benefits of memory safety to as many people as possible.
Welcome to the podcast, John.
Thanks. Thanks for having me.
So, should C++ be memory safe?
You're just jumping right to it there.
Well, I like to think,
So there's betteridge's law of headlines, which is like if you ever see like a newspaper.
Yeah, that's what I was thinking when I was asking.
That the answer is always no for those.
I mean, there isn't a law for, you know, paper titles with question marks yet.
But my hope is that it was like a genuine question.
Because I think when I joined the committee, I believe I was the first official representative of a different
programming language who had showed up there. And it was a really interesting experience on a lot of
levels. But the thing that was, I think, most useful that only occurred to me after attending
several meetings was that the concept of memory safety, though it may seem fairly straightforward,
like the more that you dig into it, the more it's really quite subtle. And the way that those
terms get used both by people who are sort of in the implementation space, people who are in the
standards writing space, and people who are in the sort of like pundit space of, you know, writing
articles targeted in industry are all using this term subtly differently. And so what the paper
is really about was trying to get clarity about what do we mean when we say that a programming
language is memory safe and what does it demand in terms of what a language must do in order to
provide memory safety guarantees so i didn't really make it you know a provocative title though i think
some people found it that way but it was more a question of let's really define our terms so that we
know what we are committing to or we know what we are ruling out when we declare that the
direction of the language should be to achieve memory safety or not
So this paper has had consensus approval.
Yeah, at DWG.
At DWG.
Can you give a description for our audience as to,
you just told us that there's many different definitions.
People subtly mean different things when they say memory safety.
So give us like, yeah, that was the hook.
Yeah, so to consumers of software,
which is pretty much everybody these days.
I don't know that it means much of anything other than things don't go wrong.
Like, I'm not as likely to lose my data.
The service is not going to go down.
But that's not who any of these conversations are really pitched at.
I try to keep, you know, the whole world in my mind when doing my job because they're the people who are, you know, all affected by it and have no voice at all in these conversations.
But I think the primary audiences for these discussions are people who are leaders in technology and policy.
And they are primarily consuming things like releases from national security organizations and the White House or big tech companies like Google and Microsoft or even academic papers.
There's a big one from ACM.
And they're all talking about memory safety and memory safe languages.
And sort of most critically, they do not consider C++ a memory safe language, but they do consider Rust a memory safe language.
And so what you have to sort of intuit from this meaning is that it's not a question of, is it possible to cause memory safety violations when using these languages.
of course, in rust, you have unsafe rust, and that can cause memory safety violations just as you can in C++.
What they're pointing to is whether the language is effective in guaranteeing against causing memory safety violations.
And the reason that rust is shown in practice to be successful in that is because of the separation between safe and unsafe rust.
The ability to write performant and useful applications in Rust without ever using any unsafe means that there's a guarantee that it is not the application code that is responsible for those memory safety violations.
That lack of separation in C++, I believe, is the thing behind the inability to make a similar guarantee because it's, it's,
It's certainly possible to write applications that have, you know, no memory safety defects in C++.
You just don't really know until, well, you never really know.
It's more of a like, you can't universally just prove a negative, right?
Yeah, right.
Yeah.
It's been running for 10 years with no issues, but there might still be a latent issue somewhere.
Yeah.
So in that sense, talking about what is a memory safe programming language.
is a little bit more of a pragmatic evaluation.
This is what we observe when they're used in the world.
But when you're talking to people who are writing code
and people who are implementing compilers and standards language,
it's a very much more a precise term,
where I think fundamentally what memory safety means
is that you can have a syntactically explicit subset of the language.
or in some cases the entire language,
but that's not really feasible
if you want to have a systems language.
But you can have a syntactically explicit subset of the language
where there is a guarantee
against being able to cause undefined behavior.
And I think that's the precise definition
that we found consensus for in EWG and Croydon.
And that was really sort of the contentious point
about whether there needs to be a full,
guarantee or merely reducing the instances of undefined behavior was sufficient. And that's kind of,
you know, with the news that you were mentioning in the last weekly episode was about that now
we've added this catalog of all the undefined behavior that's explicit in the standard, thanks
to the work of Timur and Josh in P3100. And like, that was a really amazing paper because it showed
both, you know, where the undefined behavior we know about is and did a really impressive
evaluation of, okay, given all this, what would it take and what is possible in terms of
reducing it or even diagnosing it in the language that we have now?
Yeah, and just for the sake of our listeners, I will have a link to that subsection of the
standard in the notes so that you can check that out.
I also think it's worth reading.
Yeah.
Yeah.
I don't know if it made it.
if it made it into like so that's like a list sort of like an appendix but does it include sort of
the analysis of like here's the undefined behavior that can be diagnosed either statically or or at
runtime not that i saw that made it into the standard but i could have missed up yeah well if yeah
if you also want to link to p3100 that has a really good discussion of that that's one of the things
that i found most useful in that paper okay let me put the pull that up 3100 i i've got a question and
And I think it's going to sound silly because I am moderately familiar with Rust and the borrow checker and those kinds of ideas.
But I found myself just now wondering during your discussion, how strong is the guarantee that you can't have undefined behavior and safe rust?
Like, is it ever discovered that someone said, oops, I just found this corner case that does actually let me do something unsafe and safe code?
Does that happen?
So that's an interesting question.
And I think what we're getting down to is something that Lisa Lippincott, who's been on the committee for ages and does great work, calls pointing the finger of blame.
And so the idea of safe rust is that there's nothing in that subset of the language that is undefined.
There are lots of examples of things that are undefined behavior in both C++ and Rust,
and a lot of them are the same.
For instance, like maybe the simplest one is dereferencing a null pointer.
And that would be undefined in either language.
The difference is you can't dereference a raw pointer in SafeRust.
That's simply not an allowed operation.
So in order to de-reference a raw pointer, you would have to use an unsyferenced.
block. And so the idea that you kind of separate the specific primitives that can lead to undefined
behavior and restrict them to the subset, sorry, the superset, which includes them is unsafe rest.
That's where you get your basic guarantee that, but then people say it's like, oh, well,
if I can't do anything unsafe, how am I ever going to like make an efficient data structure?
You can't have, you know, any aliasing at all. How do I make something like,
like graph structure.
And this is where we get into the kind of subtle part,
because it's not the case that all implementations of safe APIs
are only made of safe code.
You can encapsulate unsafe implementations with a safe API.
And the idea is that because callers don't have to use any unsafe to call them,
there is a guarantee that no undefined behavior can occur.
Now, because in the implementation there is unsafe code,
if that is incorrect, there's a possibility that undefined behavior may occur.
And so the question of whether that interface successfully and correctly encapsulates that unsafety,
that is basically the question of what in Rust and sort of like programming language formalism is called soundness.
And this is why we still want to minimize the amount of unsafe code in general and ideally have it mostly be in things like the standard library.
Like, you know, the Rust vector is full of unsafe code for performance reasons.
It has to deal with uninitialized memory.
It has to deal with raw pointers and things.
But there is, as far as we can tell, no way to basically expose.
any faults in that unsafe code from a safe consumer.
Now, the question of, well, how firm is that guarantee?
There's a whole project in programming language formalism called Rust Belt,
where basically they're doing formal verification of the type model of Rust
and trying to determine can we actually prove these safety guarantees.
And like, I'm not a programming language formalist, as far as I know.
Formalism engages some with C.
I don't think there's really any programming language formalists to operate at the level of C++.
But, yeah, it's a pretty well-accepted result that it has been shown that the model in general
and some specific modeling of Rust Standard Library types do, in fact, provide this guarantee.
but I mean we're all pragmatists here we live in the real world I think the more compelling thing is
showing like what have we seen in systems that use rust in terms of the existence of undefined
behavior being exploitable from user code and the answer is like we really haven't seen it
even though there are known compiler bugs you know so-called soundness issues all software has bugs
compilers have bugs hardware sometimes as bugs. So I don't think we can say that it's a
guarantee on the level of like nothing can ever happen. The question is more about this pointing the
finger of blame thing where it's like, okay, I've written a rush program and I've used only safe
code. Something undefined has happened. Like I know that it's not on me. More importantly,
if I've used a little unsafe code and something undefined has happened, I know
where my bug, you know, almost certainly is.
Right.
Okay.
It's a, I don't know, it got me thinking about one thing because, and I know it's partially
tangential, but not entirely.
Rust is this good ecosystem that people keep seeing we should have in C++, which is cool
cargo, which allows you to pull like packages very easily.
And if, if JavaScript told us anything, is that that's, that is kind of an unstoppable force,
which means, I don't know, give it five years and everything.
Everybody's using Rust left pad, we're having no idea that we're actually pulling it.
And Rust left pad actually has telemetrics to some, you know.
Yeah, and my question is there, yeah, but how, like, I understand the idea of like,
oh, but there is no, I'm only like, you know, all the API boundaries that I'm, that I'm calling or safe.
Right.
But how do you start verifying that at scale when, when again, the ecosystem makes it easy for people to pull dependencies?
Because I know, for example, in people who make, like, medical software in C++, they have, like, a list of third party packages.
they usually run with like very old version because the process of certifying new version is so painful that, well, too bad, I'm going to use 10 years boosts or something.
Yes, Jason, you might have consulted for some of these people.
It is, some of them are all living with a difficult world.
And how does that mash with like an ecosystem where it's very trivial to pool like a ton of fruit party code and kind of like an expected workflow to a degree?
I don't know.
Maybe it's not a real problem.
It's more like an engineering problem.
than a theoretical problem?
Yeah, well, it's the sort of thing where, like, in security in general, you know,
you don't get to control your adversary, and so they're naturally going to attack you
where you're weakest.
And so I think that's part of the reason why it's not just Russ, but like basically every
language has to deal with the supply chain security problem.
and the easier you make it to import sort of third-party code,
kind of the bigger, the surface area of that attack gets,
I certainly don't know anybody who's like solved this in an easier,
scalable way because it's always going to rely on verifying that, you know,
this code that, you know, if you had the time to go through with a five-tooth comb
and verify all its properties, like, you know,
maybe you could have just written it yourself.
The whole idea of depending on third-party libraries is saving time and effort.
But I do know of at least one project within the Rust community for having an extended standard library
where large players in the Rust space are basically contributing funds to try to add more code that has had its security reviewed and audited.
and so to make life easier for developers and being able to extend more trust.
I know that's something that's under way I haven't followed it closely.
I just don't see this as like a specific to rust problem because obviously you talk about left pad and yeah.
Yeah, no, it's not specific to rust.
I think I think about it because I understand that for example, we compare that to C++
where people are usually much more conservative with the number of third party they use, if any.
Yeah.
No, but I think I just had a thought, but it's not a problem specific to Rust for sure,
but it might be a problem that Rust at the moment is uniquely capable of solving in a sense.
If I can specify, like on my cargo command line or whatever, say,
only accept packages that do not have any unsafe code in them.
And then if a dependency were to change at some point, my build would just simply break.
That could be, I think, very valuable.
Yeah, if what your concern is, is undefined behavior, sure.
If your concern is simply malicious code, unfortunately not, because you don't need undefined behavior to just like, you know, mine crypto.
Yeah, read or read random files from your...
That's correct.
I guess it's because we were talking about memory safety.
and I guess having like malicious good doesn't count, right?
It's only if there is an actual like you be hidden somewhere.
And above because you have many third parties and they do use unsafe stuff that is maybe not as good quality.
Yeah.
Yeah, I like to think.
Sorry, go ahead.
Oh, I was just going to say like I like to think of the metaphor with security is that, you know, when you're when you're testing things for correctness, you can sort of like check all your doors and windows.
and especially make sure to double check the ones that get used a lot.
But when you have an adversary, it doesn't, they'll go for the ones that you're checking the least.
And so if you like, you know, leave anyone open, they are going to get in.
It doesn't matter if you've left it open by virtue of undefined behavior or by virtue of just allowing in some code that is, you know, just calling window open.
Yeah, I guess that's a fair point. That's a fair point.
Going back to what you mentioned, about like the whole idea is that to be safe, if I understand your description, is that you have a way to have a superset.
Is there like a kind of like a limit to how big or how small the superset of the language has to be to still like be considered like, you know, safe?
Because you could argue constex per code is cannot invoke you be.
So it is technically a superset of C++ that is safe.
Obviously, you're not going to make a very interesting program if you just do constex.
because by definition it cannot have input, but do you know what I mean?
Yeah.
Yeah, yeah.
Like how big or how small does it have to be?
To be just a little pedantic first, the safe part of Rust is the subset.
It's the superset that includes the things that can cause you be.
So similarly, ConstExper is a subset of C++.
And so the things that get extremely,
get excluded because it can cause you be.
I don't know that there's like any formal reason that it can't be larger than a certain size,
but in practical reasons,
you would like it to be as small as possible.
And in terms of rust,
it's down to like there's sort of like five major ones and like three more that are kind of like very,
obscure in terms of operations which are necessarily unsafe.
And let's put myself on the spot and see if I can remember the major one.
So obviously, dereferencing a raw pointer, accessing a field of an untagged union, mutating a static
variable.
That is like a global one.
Oh, interesting.
Okay.
Yeah.
I mean, that's a problem in C++ as well.
And then the two that are sort of like more of a blanket thing, you'd be like, oh, well, that includes everything is implementing an unsafe trait and calling unsafe code.
That's it.
Like those are the things that are unsafe.
And so that's not like a huge amount, but it basically provides all the things that you really can't do in safe rust.
Right.
Right.
Yeah.
So I'm guessing if you stuck to modern C++, I mean, you would still probably have the issue of pointers because, I mean, iterators are pointers in disguise, although they can be checked.
So I guess maybe in some ways you could argue that iterators may have a way of checking that they are actually like better than just a herbal pointer that you can't guarantee anything about.
But it does feel like a sum of it.
And I admit, I have watched your talk at ACCU about this, but I haven't.
by the entire paper.
So how feasible is that?
Because I could try to guess, like, how much do you need to add or change in C++ to get there?
Assuming obviously modern C++, I know the C people who sometimes I talk to in game,
they would have very strong different opinions.
But, you know, if we stuck to modern C++, how much do we have to change?
How much do we need to have to add to get somewhere there, basically, is my question.
Yeah, and honestly, I don't think I would have embarked on the effort within the committee that I did if Sean Baxter hadn't already done the work to show quite explicitly what you need to do.
Because what he did with Circle was basically implement all of Rust's safety features on top of C++.
Bjarn has talked for many years about the subset of superset strategy.
And this is actually even the language that I used in my WG21 paper and that's in the resolution.
And the idea is that in existing C++, you have the tools that you need to do the things that you want.
But you don't have the tools to do them in a way without causing undefined behavior.
So you create a superset.
You add new tools like Rust's checked references, which is how the barotecher does it work,
to give you the ability to have, you know, referential types without, you know, being subject to
undefined behavior? I mean, we still have raw pointers because there are things that references
can't do. But once you add those safe facilities, then you can have the subset where you exclude
the unsafe facilities. And so that's kind of how you get there. So I think there's quite a bit that would
need to be added to C++ if you wanted to have a safe subset akin to Rust. But I think we can see
quite precisely what that would look like if you went the Rust route by looking at the paper
that Sean Baxter wrote. And the fact that he has a working implementation of it, even if it's not
an industrial strength commercial compiler, it is, it lowers to LLVM. It's his own customer
front end.
And I think it goes a long way to have like a working proof of concept.
Okay.
Yeah.
You mentioned something about adding supersets, right?
But I'm thinking, would it be like if you need to have those things on top to be to be
considered safe, but also you kind of probably want to put them very deep in the language.
I'm thinking vector, you know, the STL in general, like for that thing.
to work, you're either going to need like a STL version 2 that is considered safe, and I don't
want to hear about the ABI nightmare that it would be. Or you only have one STL, and that means
that every compiler needs to be able to, you know, pass the safe annotation and whatnot, even
if some people are like, I don't care about safety, I don't, I will just discover all of that,
but it still need to be compatible. I want to, before John answers, if you don't mind, mention
that in C++-26, 29, 26, but C-plus-26 with contracts enabled,
we have something around the standard library without making STL-2
because you cannot, like you can choose to enable contract verification
that your index operation into a vector, for example,
is not past the end of the index.
And so we are, you know, kind of get a flavor of that,
But now I'm very curious from John's perspective.
And with Matthew's question in mind, do you think that that has gained us anything?
And what do you think we would need to do from there?
So I don't actually, like, I'm not saying anything pro or anti-contracts,
but I don't think it's contracts that have added that facility.
I think it's the hardened standard library.
Yes, but I thought they were implemented via contracts.
Yeah.
And that is sort of like the implementation vehicle.
that WG21 chose, and I think that's perfectly fine decision.
But the code to do that predated contracts.
That came from Louis Dion's paper at Apple,
and I believe that was voted in in Hagenburg.
That would have been like February 2025, I want to say.
And so the idea of checking for containers like Vector,
where you know the valid length, that's pretty straightforward.
Like the reason that wasn't done before, it was a question of like, can you eat the performance cost?
And I believe it was Chandler Carruth wrote a pretty interesting blog post about Google's experience with turning that stuff on and realizing actually like the cost is not so great.
And so I think that's actually super valuable.
And like, yeah, we should absolutely be turning on any of the additional safety.
Which is very curious, because coming from the game industry and our hatred of the STL,
like, the big thing that I have been even trying to fight for or against in cases,
is that if you use the debug STL on Windows, you know, like with the bounce checks and everything,
it does insert a check on operator's square bracket.
And when you do loops, you can work that around.
because if you do like a 4, you know, range 4, what it does is that it actually only does,
it leaves the check so it's done only once when you enter the loop.
And then by construction, you're actually not invoking square brackets.
So it doesn't need to check the look every iteration.
But I don't know, maybe if you do that with optimized compilation, it's fine.
Because what I've noticed is if you make a stupid like C-style loop with a square bracket operator
and it inserts an if test every time, they will kill any form of like SIMD optimization
and a bunch of stuff in the backend pipeline.
and we'll just stop working.
So maybe Google, like, you know,
I know they have some patches on Clang,
so I don't know if it's because of that,
or it's because they do something smarter.
I'm really curious about this.
Because I've been trying to convince people that the SCL is not that bad,
and they don't need to re-implement something.
And so hearing that we're going to add even more,
like contracts in the STL makes me a bit concerned
that I'm going to have even more pushback next time I suggest it.
Yeah, I think part of the reason that I try to separate it
from the contract discussion,
is that, like, the Harden Standard Library stuff is the least controversial thing I've ever seen in the committee.
Everybody was for it.
Like, it's just an obvious thing.
It's opt-in.
And, like, yeah, like, let's not have undefined behavior that we can easily check against at runtime.
There's no way it could be statically verified anyway.
So it's not like you're paying a runtime cost instead of a compile time cost.
but like you just decide if you want to you know have a you know what is it that's under tightrope walkers
if you want to if you want to go without a net oh like that's safe to net yeah no that's fair
my main question is are we going to go that route too for like trying to have safety in the
standard or are we going to start creating like basically binary incompatible builds
because one of them is built with the safe toggle on and the other one is built with the safe toggle
off. Yeah, as far as other kinds of safety, no, because other kinds of safety aren't checkable at runtime.
Yeah, no, everything in the compiler, I mean, people might complain that it makes compile time longer,
and that's another topic that is a recurring issue that I keep hearing about. But yeah, no, right now I'm trying
to focus about, like, is it going to make any binary incompatibility because it has to insert runtime
checks in case you ever link with someone who needs to compile it safe, or is it going to be
one of those things that is an ABI break and you need to have different tool chain settings.
I think it's all behind the ABI, Matthew.
So it's neither.
It's just one module might assert and the other one might not.
But I could be wrong.
Could result in an ODR violation, actually.
Yeah, that's basically what I'm going.
It's like, how do you handle the concept of having like extra checks?
In general, how do you have opt-in safety in shared components,
like for example, the standard library,
or even a third-party library that everybody loves to use?
Yeah, I think that is a particularly tricky C-plus-plus problem
with regards to ODR stuff.
And I would say that's, you know, kind of beyond my C-plus-knowledge,
there's definitely other committee members
or you could get on that could, you know,
speak much more thoughtfully to that.
But I know it's both a complicated problem,
but not a new problem because it's already the case that you might be like linking the even the same code
in different compiler invocations with different flags and then what does that mean, right?
Yeah, no, it absolutely is.
I guess I was just more like, is that one more that we will have to handle one way or the other?
It might be.
And how bad does it trickle down?
Yeah.
I feel like if I were in a class right now and a student asked me this,
I'd say, that's an interesting question.
Don't do that.
Yes, yeah, right.
Yes.
I think it's because, as you mentioned, your paper heard,
maybe I'm here to be the devil advocate of like the kind of people who are like
see plus plus because it's fast and like if you talk about safety,
you're like, yeah, I don't care.
Because that's so far what I've heard around me in the games industry.
You know, like they always care about like how fast does it compile and how much
does the compiler insert that I did not explicitly write.
And that, I think in some degree for some people would,
include safety checks.
This is my go-to...
Oh, sorry, John. Yeah, it's all right.
Oh, I was just going to say,
I also have a background in the games industry.
It's been a little while, but, yeah,
I worked on, like, some PS2 and Wii games back in the day.
And, like, with regards to safety,
what I would say is that, you know,
at least on a console game,
and in the era that I worked on console games,
you weren't having any DLC, so maybe we're not so concerned about that.
But we're not concerned about vulnerabilities in the same way,
but I'll tell you one of our chief concerns as a game developer
was that Sony and Nintendo didn't reject our submission.
And I can tell you that having your console game crash will fail you.
So something that can increase that kind of safety, absolutely.
important. I agree. And I guess going back on that topic, a thing that surprised me, because I've
worked with mobile a bunch recently, obviously, like Unity, one of the biggest, like the most
Unity games, a lot of Unity games are using mobile. And that's the thing. You can make as many
PC application that crashes all the time. I mean, it's bad for your reputation, but no one is
going to enforce that you fix it. But your crash rate is actually part of your app rating on, like,
I think both app stores. And you will either get delisted or start getting pushback on your
discoverability, which is basically like a kill switch.
People will go very far to avoid that problem, but I don't know.
I know it's slightly tangential to the discussion, but I kind of, it is interesting that
some ecosystem have managed to kind of enforce safety to be a concern, or at least
stability, which I guess you can say is downstream of safety or to some degree.
Some people may also argue that as long as it crashes, but it's not exploitable, it is still
considered safe, right, because of panic.
Yeah, I'm going to make an argument.
I'm not putting on my instructor hat right now.
If your primary goal is to keep the game running so it doesn't crash
so it doesn't get rejected by Sony,
I'm going to state with some level of confidence
that I believe your code actually has more vulnerabilities than if it had crashed.
Because you're going to do whatever it takes to keep running,
regardless of whether or not you're in an undefined state right now.
Absolutely.
Absolutely.
And yeah.
Yeah.
It's interesting because people often talk about like, oh, games don't care about safety.
And I always hesitate to reference something that I don't understand deeply.
But my understanding is that the way the Xbox 360 copy protection code was basically defeated was undefined behavior in a released game.
Yeah, no, no, this is, I use this in my training material.
I cover two different exploits that I'm like, the ones, the PS4-5 file system jailbreak
and the Xbox console jailbreak, both almost certainly would have been caught if they
had just enabled dash-W conversion in their builds.
Yeah, so even if a developer doesn't care so much, the, you know, the platform owners are
certainly going to care about these things a great deal. Yeah. So I think you're right. We will see that,
we will see that have a ripple effect or already has a ripple effect in some, in some, in something.
I think it might depend a lot in where you're publishing your apps and how much people care.
I do agree that, you know, there is this weird pernicious thing that technically not crashing
in being UB right now might be better than having a UB caught and crashing just if we look at your stability metrics.
And I don't know if that's the kind of thing that's going to...
I mean, again, I'm being deviled advocate.
I do agree with all of you, right?
Like, applications should not crash and they should not invoke you B.
It's just obviously bad.
But I just thought this is an interesting point that I'm sure someone will make.
That and how much more compilation time it has to do all those safety check at compile time
because I already heard people telling me, oh, yeah, I would like to use Russ,
but it's too slow for iteration on games,
so I'm not going to bother.
The compile time,
which is already the argument we hear from C++ versus C, right?
It's just one step further.
And so I'm wondering if adding safety extra compiled checks
will also bring that into the conversation at some point.
I have an interesting anecdote.
I have no idea how well this translates to drill code.
I'm guessing not very well.
But when you're working with GCC,
and they now have at least partially hardened version of the standard library.
If you do something that is trivially easily obvious,
like create a vector of one element and then try to access the second element in the vector
in a code block, the compile will fail.
Because now with whatever extra code they've added around the hardening,
it's now able to see that at compile time as well.
But again, trivial case, small localized case.
Maybe it won't translate at all.
But I just found it interesting.
I mean, that's good stuff.
Yeah.
And well, the thing about static analysis is there are a lot of obvious things to the compiler that may not be obvious to humans.
And that's great when we can check those and like avoid a problem that's either even a logic error or, you know, undefined behavior down the line.
I think a lot of the discussion and a lot of the confusion in the standards world and, you know, in this discourse about, you know, governments and industry is what are the guarantees?
And fundamentally, we know that with a language that allows you to do all the things that a systems language needs to do, you can't provide a guarantee against undefined behavior.
There is a particular computer science result that I often like to talk about called Rice's theorem.
Are either you familiar with that one?
I don't think so, no.
Not sure.
So it basically just says that in general, knowing whether any non-trivial property of a computer program is undecidable.
And in this case, a trivial property is one that is only sometimes true, like trivial properties you can show.
But this is sort of like a corollary of the halting problem, which I think a lot more people are familiar with.
And so the idea is you can't generate like a guarantee that a program is valid unless you are willing to reject some valid programs.
And so you sort of have the choice.
Do you overestimate or do you underestimate?
made. And that's what I think is kind of genius about what Rust has done, which is that they do both.
Unsafe Rust basically allows in things that are known to be invalid so that you can write
everything you need to, whereas Safe Rust rejects some things that are valid so that you can
have a guarantee that you haven't done anything which is invalid. And it's the marriage of those two,
which is like powerful but also fairly subtle.
If you don't have that separation,
if you don't have kind of like two languages in your language
and you want to be a systems language,
you have to do what C and C++ have done,
which is admit things which are invalid.
Right.
That's cool.
No, that's a nice way of putting it down, I think,
and explaining it.
Because again, there's so many talks around about safety.
It's a big topic at every conference.
So I think it's good that we can have like some interesting
an intuitive way of explaining what are we talking about and what all the boundaries and where it lands.
If I may ask to go ahead, Jason first, because before I go ahead.
I was checking my notes.
I wanted to follow up on one thing that you mentioned, though, Matthew, about the standard library.
Yeah.
Because I think the question of how much of the C++ standard library would need to
change for the sake of safety is probably the biggest unanswered question because while Sean
did implement, you know, basically all of Russ's safety features in his, you know, language circle,
he didn't implement full new standard libraries. He did a couple small things. And he was generally
approaching it as there would be like a like a stood two, which is all like the fully safe
versions, I think it would require actually auditing the standard library to know sort of the
extent of the interfaces that need to be totally reinvented versus which ones would need to be
augmented. As it stands now, like the best kind of comparison we have is to look at the Rust
standard library. And by and large, it contains like safe interfaces for things, but there are
some APIs that are unsafe to call. For instance, like, you know, instead vector, you can
call an unsafe function that basically exposes all the internals and takes it apart.
So for the purpose of doing things like FFI, because if you want to be able to transfer
things to like a wholly different, you know, system that may be using a different allocator
that has different semantics, sometimes you need to be able to do that stuff.
But that is inherently unsafe.
And so what that means is that you as the programmer are taking responsibility for the fact
that like you are doing operations that could cause you be and you're saying the compiler can't
verify this for me so you know i am taking responsibility for for the the proof obligation that this
code even though it has the tools to cause undefined behavior it does not it's sort of like
opening the drawer of very sharp knives and you know saying yeah i've i've taken my safety training
course yeah because i guess you can make sure that like square bracket operator is safe but you
can't guarantee that what you will do with like the data will be safe, for example, because
then you can just do whatever you want with what's inside or anything. Yeah.
We spoke in a bunch about Rust, and since you're mentioning FFI, I guess this is a good
segue to the question of, like, you used to work on interrupt from what you said. So I'm
curious about that, because I know he was also talked about, and all three of us actually
were at ACU, and I know it was a topic, like interrupt between languages. And I'm curious.
because you started on the bus side, I think,
and now you ended up in the C plus plus side.
So I'm a bit curious, like, how did that go?
How did that happen?
So the interesting thing about introp is, as far as I'm concerned,
you know, it doesn't really have sides.
Like, introp is tricky because it exists in this, like,
liminal space between two languages.
And I think that's also, like, a big part of the social reason
why, like, it's not that great,
because nobody is really focusing on the idea of like we just want to make introp better for everybody.
Plenty of individual organizations want to make introp better for their particular business case.
And so you see efforts like Krubit from Google do an amazing amount of work to make
interop between Rust and C++,
plus very fine-grained and high-performance and expressive.
And it works very well, as I understand, for what Google is doing with it
and automatically exposing their C++ interfaces to new Rust code.
But it doesn't really work if you can't control the compile of your whole world.
if you can't change your compiler to do special things.
And then there are totally different approaches,
which may require a lot more manual effort.
Something like CXX is very lightweight,
but somewhat opinionated and requires extra work on the part of the programmer
to be able to define their interfaces
and then generate sort of that glue code,
both on the rust and C++ side.
So when I came into this space and was leading the interoperability initiative,
I tried to look around and see like, okay, what is the state of the world
and what is useful for me to do that people are not already working on?
So after taking about six months to analyze things and to interview a ton of different people
at different big tech companies, small tech companies,
people involved in implementing introp libraries, people involved in the Rust language and standard
library, people on the C++ standard and standard library side, I came out with this sort of like
high level problem statement and strategy. Because one of the things that I believe really passionately
is if we don't have like a clear understanding of our problems before trying to think about
solutions, we will, you know, very quickly narrow our vision to like fitting the problem to the
solution we have. That's like, you know, it's like if you only have a hammer, then like every
problem is a nail. And the way that I characterize this, and this is a document that's, you know,
still out there is that there's kind of like three different tracks that that we can pursue and you
potentially pursue them in parallel. And the first one was what I called short term. And that is like,
there's a bunch of existing interop solutions that are all done in sort of library code.
And they are useful.
I used them when I was at Mozilla working on AVIF.
I used the sort of lowest level ones where you're not just using the pure language are,
now I'm blanking on the bind gen.
It's bind gen and C bind gen are the tools for just generating the binding.
on the opposite side.
And using those is fine if your surface area is small and simple.
But then you have other things like CXX is popular.
There's another one called Zengar,
which comes at it a little more from the C++ side.
And I know Adobe and David Sankill are a big advocate of that one.
But these are all things that people are already working on that are already useful.
And so I'm like, I think the thing that the Rust Foundation could do to help them is basically provide support and sort of like social connections when they get blocked by things.
Like a good example of something like that was there was a longstanding bug for 128 bit integer ABI where Clang and GCC weren't agreeing.
And so because, well, it's not really clang.
It's LLVM.
And because the Rust compiler lowers to LLVM, this was affecting Rust interoperability.
And the fix was in LLVM, but notably a different project from, you know, the Rust project.
And so getting that fixed through is not necessarily the same priority for LLVM as it might have been for Rust.
And so I feel like that was something that ended up getting resolved before I joined the Rust Foundation.
But it was the kind of thing that, like, if I were there at the time, I would have wanted to, you know, provide more support to hopefully get that kind of thing fixed faster.
And I think the main thing that we did during my time at the Interoperability Initiative on what I would call that short-term strategy is we helped provide funding both directly and indirectly for a project called borrow sanitizer, which is basically a tool that uses an,
LLVM interpreter to discover bugs that you might have in your interrupt code where the C++ side
violates the memory model assumptions that underlie unsafe rust.
That's interesting.
Yeah, so you can think of it's very tricky.
It's a kind of a minefield.
And there already exists a tool on the rust side for this called MIRI, which, so MIR is the
mid-level intermediate representation.
representation of Rust code. It's the level before it gets lowered to LLVM that still has sort of the flow control data that you need to do, you know, the barrow checking analysis. And it's an interpreter that you can use to run your unsafe Rust code and it will discover undefined behavior. And so this is like a really useful tool for helping you get your unsafe Rust code right. But, you know, before Barrow Sanitizer, you don't have a similar.
counterpart on the C++ side. And as we all know, Interop has two sides, or as I like to say,
no sides. So I was really proud of the work we did to support them. And like, you know, they just
gave a talk at RustCon, I think, last month. And I think that's great. The second strategy
is what I would call like longer term. And that's changes to the Rust language itself and
standard library for improving interoperability.
And we did some things there.
It's a longer incubation period, obviously, for those kind of changes.
And so there's still undergoing that work.
I think the primary thing to follow for that is a Rust Project goal,
which is the way that the Rust Project organizes their priorities,
which I wrote and got accepted before I left called the problem space mapping,
which is the idea of like, let's try to figure out what are the specific things about interoperability
that would require changes in the standard library or language level.
And let's try to get consensus around like, what does the shape of this look like?
Let's basically refine the problem statement to the more specific use cases
so that we can get input from different people.
And instead of like arguing endlessly about like my way to do intraopas right versus mine,
like you know, you have like Adobe's approach and Google's approach are distinctly different,
but that's just because they have different business cases, different models.
They're never going to convince one another, but they're definitely sub-problems that they
agree on.
And to the extent that they can agree on sub-problems, we can, you know, cooperate and share solutions.
But then the third strategy, what I called social interoperability was like, let's actually
have the Rust community and the C++ community like talk to one another directly.
And this is why I started going to the WG21 meetings.
Like to me, it's kind of wild that like, you know, we want to do interoperability,
but we did it from the standpoint of like, oh, all we can do is code and all we can do
is code like on our side of the divide.
And that's going to be extremely limiting.
Like it's almost a good metaphor for the.
idea that like we're doing interoperability, we're doing FFI, because we have linkers, right?
And linkers don't exist for the sake of interoperability.
They exist because you need to write programs that are too big to do in one compilation unit.
And the interoperability piece, like, you know, it almost falls out of it, but it wasn't
designed for that.
So the fact that it's like pretty fraught and awkward is not at all surprising.
like we're using something in an unintended way.
But like what kind of interoperability could we get if we actually designed for it?
And the thing that I noticed was that during my time working on AVIF at Mozilla,
the most fraught part of the interoperability was the fact that I had to use unsafe rust.
And, you know, I'm here writing this parser for untrusted image data from the
internet, and that's why we did it in rust. But if I mess up the piece where we're transiting to
the C++-based static image library in Firefox, like, I could potentially drop that on the ground.
And so my thing that occurred to me is like, well, what if we could actually do FFI but still
in safe rust? Like, and at first that sounded like a pipe dream because like, well, that would
require that, you know, you have some kind of safety guarantee on the C++ side of things.
But I think it was that plus seeing what Sean Baxter had done made me realize what's like,
well, maybe that's not totally impossible. And if we could have, you know, freedom from
undefined behavior on both sides of FFI, we could have like a very different kind of guarantee.
And even though that doesn't, you know, make all of the C++ code safe, it at least moves that
frontier of where you need to be concerned about safety across this basically uncrossable gap
that is interoperability. And then it basically becomes a question of like, well, how far do you want to
push it in your C++ code? And that is a consideration that I think is much better left to the
application developers than the language community. And so that's why I spent so much time
engaging on the C++ side of things, because that's where there was the most work to do that
nobody was really working on and that what I thought was the most for everybody to gain in the very long term.
Okay.
It's interesting that you says the most in the C++ because I would have thought that C will also have been there.
Because, you know, I'm assuming that Ross does not re-implement open SSL or over extremely common libraries.
And, you know.
You'd be wrong there.
Okay.
Okay.
Then maybe I am wrong.
But I keep thinking like a lot of like low-level system libraries that a lot of that I get you see in every app.
like curl or whatever.
Like, they are by definition and safe
because they're returned in C, which you could argue
is even less than C++.
Like, they would also probably be potentially interested
and be like, oh, I can expose like a safe API to this.
Well, the thing about it is that C is much less amenable
to creating a safe subset
because it lacks the encapsulation
necessary. What makes the ability for Rust and potentially for C++ to provide safe interfaces
is that you can't just reach in and muck with their internals. And, you know, C just doesn't provide
that same kind of encapsulation. And to your point, like, the Rust Foundation has sort of like an
innovation lab that they use to foster and support development of new, um,
or existing code that they see as valuable to the community.
And I believe the first project that they basically included was something called
Russell, which is Rust LS or Rust TLS.
And this is exactly what you're talking about.
So there's a lot of effort to re-implement those kinds of fundamental tools in Rust.
And I think they've largely been very successful.
Yeah, I suppose then you shift the problem because then do Sphinx would like to be
called from a C++ then.
Because ideally, I guess if you have the choice between two cryptography libraries,
you would rather call the one in Rust because it's the one that is least likely to have,
you know, critical vulnerabilities.
Well, I don't think it's quite that simple.
I mean, what you get, as soon as you rewrite something in Rust, you,
depending on how much, you know, or any unsafe code you use,
you get a guarantee against undefined behavior.
But like we're talking about before, there's still the concept of logic bugs.
If I had the, you know, to choose between like the brand new Rust thing or like the battle tested, you know, C or C++ implementation, like you have to be pragmatic in considering all those things.
But like if you're talking about like in the long term, you know, what is better?
What basically allows me to have more of my guarantees checked by machines?
Like that's the real advantage of Rust.
I don't think it means rewrite the world in rust.
If I thought rewriting the world in rust was the answer,
I don't think I would have taken the interoperability initiative role.
I took it because C++ is so pervasive and affects everybody.
So I want it to continue to be useful and evolve to be better because it's never going away.
Like, Fortran and Cobol haven't gone away.
They've just become incredibly rarefied spaces for expertise.
I agree. I think you make a very fair point. And I got a bit tunnel vision. And I assume it's kind of a common thing. You keep hearing about safety and you're like, okay, cool. So then we write it in the memory safe language because then obviously it's going to be better and safe. But as you mentioned, what about logic errors and how many, like how many of the bugs in, because I know people keep thinking like, you know, the top five problems that we keep running into our memory issues and blah, blah, blah. But to your point, like, yeah, but that doesn't guarantee.
Like if you make a rewrite in rust, I mean, if you can, maybe it's a good thing.
I'm not, I don't have like a personal opinion either way.
But that doesn't guarantee, even if that guarantees you that maybe you're not going to get
those top five errors, that doesn't guarantee that you still don't do like something stupid,
like have an issue with the entropy of your generator or like a logic error, basically, as you,
as you may know, like go to false, go to fail, right?
Like technically, like, it's a logic error.
I mean, maybe go-to is considered unsafe.
But, you know what I mean?
You could potentially just write something that is logically unsound, but passes all the checks.
And I think I got stupidly tunnel vision by hearing safety.
And I don't know how often this is a mix up that people do when they are like, oh, now it's going to be safe.
It's like, well, it's going to be safe from those kinds of errors.
But it's not going to be guaranteed to actually still do the thing.
Logic errors, as you mentioned, are still around.
Yeah.
I mean, this is something that's come up in the committee a lot, like the idea.
of like, well, what do we mean by safety?
And, you know, there's different kinds of safety.
And like, I think we both don't specify it enough
and sometimes subdivide it too much.
Because at a high level, we have the concepts of language safety
and functional safety, where, like, language safety is the things that, you know,
your compiler can actually be involved in checking.
And functional safety is, like, what sort of impact
does this have like on the world like is my plane going to crash is my toaster going to leak my credit card
information you know to the the dark web but that's a very different kind of safety from like do i
does my language guarantee against undefined behavior but then people will sort of like have the
concept of like oh well you know this is memory safe but i might actually actually
have integer overflow.
But like, that's not really a coherent distinction because once you have integer overflow,
that might lead to you indexing outside, you know, your memory allocation.
So, like, really, is that, is that still memory safety?
Which is why my view on it, which was highly influenced by sort of like the preeminent, you know,
language formalism research in the Rust community.
And in particular, someone named Ralph Young, was that,
Like, there's no meaningful division of language safety into, like, spatial or temporal safety or, like, you know, range safety.
It's all just a question of, like, undefined behavior is possible or not.
Because once you have undefined behavior, anything can happen.
Yeah, I think I thought this is, again, the magic word of C++ and of compiling in general.
It's all about UB, right?
When we talk about safety, we say, like, what you mean is we mean there is no UB.
you can have a very defined behavior
is unintended because you wrote the bug
but it's still there.
Sorry, Jason.
Yes, I'm going to say
thank you very much for joining us, John.
Unfortunately, we are out of time.
It has been, yes.
If you have any final thoughts
that you would like to share
before we wrap this up,
feel free to.
Yeah, I think the final thought
that I would like to share
is that I think there's so much
we have to learn
from one another, especially across communities when we approach things from a place of curiosity
and epistemic humility, which is just like admitting when like maybe you don't know something.
And I like to believe that our responsibility as much as possible should be to all the people
who are affected by our decisions who have no input on them.
we're always immediately answerable to our employers
and governments and stuff like that.
But I think in terms of being a good programmer
in the sense of moral good,
it's on us to make sure that we're thinking about
the things that are not in KPIs
that are not necessarily going to be
in our performance evaluations and things like that.
And I hope we can work together on that.
I think that is severely lacking in many cases.
So thank you for that comment.
Yeah, yeah.
I think I could talk about this for hours too.
So a very important thought for the listeners.
Speaking of listeners, thank you all for listening.
We are going to go every month, as usual.
We are always looking for new guests.
John, you had been selected.
We have this informal rule that a guest can nominate someone,
and then we will come and knocking at your door.
So you had been nominated by a past,
by a past interview guest.
If you want to nominate someone online or offline,
we can give you the chance to do that.
And of course, listeners, again,
feel free to contact us if you would like to be on the show.
Jason and I are always looking for people,
especially if you were not someone that is very well known,
that you have like a small project that maybe people don't know about.
We love to find new things that we think people should like hear about.
Yeah.
So I do have someone I would like to nominate.
someone who I got to know in my work in the interoperability initiative and actually mentioned
Krubit earlier. Taylor Kramer works at Google on Rust Interoperability. And actually, I believe,
was the closing keynote speaker at ACCU. And she does amazing work, has been involved in the Rust Project going way back,
but is also far more versed in the deep details of C++ as well.
And so I think she would be a great person to have on your show.
Awesome. Thank you very much.
Thank you.
All right.
And with that, I think we're done.
Thank you for listening.
And we'll see you next month.
See next month.
