PurePerformance - OpenTelemetry and the Reality of Vendor Choice with Adriana Villela and Josh Lee
Episode Date: July 6, 2026I still hear people say, “OpenTelemetry is vendor-neutral, so you can switch any time!”In this episode, Adriana Villela and Josh Lee (both active OpenTelemetry contributors) help bust that myth.Wh...ile OTel standardizes instrumentation and signal transport—and unlocks a rich ecosystem of tools—switching vendors isn’t as simple as it sounds. There’s real cost in retraining engineers, migrating dashboards, SLOs, and alerts, and reworking deep integrations across your delivery pipeline.We also dive into a key challenge the community is tackling: helping engineers instrument by value, not by default—making it easier to capture the right signals with high quality instead of just collecting everything.Here the links we discussed:Adriana's LinkedIn: https://www.linkedin.com/in/adrianavillela/Josh's LinkedIn: https://www.linkedin.com/in/joshuamlee/The blog article: https://thenewstack.io/opentelemetry-vendor-neutrality-guide/CND Austria Talk: https://www.youtube.com/watch?v=1gxLseuaTdMKCD Prague Talk: https://www.youtube.com/watch?v=pPXG20CXKxQOpenTelemetry Project Website: https://opentelemetry.io/
Transcript
Discussion (0)
It's time for pure performance.
Get your stopwatches ready.
It's time for Pure Performance with Andy Grabner and Brian Wilson.
Hello everybody.
Welcome to another episode of Pure Performance.
My name is Brian Wilson.
And I'd like to introduce my real pain in the butt, jerk of a co-host.
Andy Grabner.
Andy, how are you doing today?
Thank you for, are you using an AI later on to remove all of the nasty words?
No, you said, you say,
You said you were not to be like, you know, mean person or something before.
So I'm just letting people know what you really like.
Oh, yeah, of course, yeah.
Thank you.
Finally, the world knows.
Finally, the world knows behind that laughter is somebody as rude as Abraham Lincoln.
A little inside joke action going on there.
Anyway, we could talk about, you know, your attitude as much as we want today, Andy.
But I think people would rather...
Yeah, the show is not about us.
No, but even though you always say it's about me, like you're always asking me to make you
sound better, make you look better, even though it's a, you know, Andy is a grueling taskmaster,
everybody, and just don't let the smile fool you.
So, sir, should I call you, sir?
Okay, sir, what are we talking about today?
We invited two real celebrities when it comes to open source advocacy for open observability for open
telemetry.
We've got two people that have been touring the World Conference.
stages, and I will let them introduce themselves in a second, but we have Adriana and Josh,
who I think have not only presented well together at different conferences, but they've also
written a really cool blog post that kind of triggered this whole discussion today.
And the blog post is around vendor neutrality isn't magic, a hard look at the open telemetry
ecosystem.
I've been working and talking with a lot of organizations.
and sometimes there's this misconception out there
that you sprinkle a little bit open telemetry on it
and then all of the vendor locking problems are gone.
And the reality is a little bit different.
I really liked a couple of the quotes
that Adriana and Josh had in their blog post.
One of them is open telemetry really gives your choice
but you still need to make that decision
on what you really do with the data in the backend.
But yeah, without further ado,
thank you so much, Adriana and Josh to the show.
I would like to start with Adriana for a quick intro
and then also giving the chance to Josh before we dive into the topic.
Adriana, you go first.
All right, so my name is Adriana Vilela.
I am based out of Toronto, Canada, originally from Brazil.
I work with Andy at Dinah Trace and as a principal developer advocate.
And I've been in the Devrel space, I guess, since 2022.
And before that, just skipped around through various areas of tech,
did Java development for a number of years,
zigzagged between management and not management
and settled in the DevOps, reliability, observability space
is my kind of permanent forever home until the next thing.
Went what brought you before I let Josh introduce himself,
what brought you into Open Telemetry in particular?
It was a bit of a kind of a happy accident
where I found myself managing an observability team at two cows.
And I didn't really know enough about observability.
And so I decided to educate myself to really be able to run this team properly and give proper direction.
And so I was learning in public about it, asking lots of questions on various forums and got to learning about open telemetry.
And, you know, I was trying to get folks at two cows to use open telemetry back in 2020.
2022 when Traces was not even generally available.
And I'm like, trust me, this thing is going to be big.
So it's nice to see that I wasn't wrong.
Well, I think you were very right.
I mean, all of Open Telemet to talk off,
just made a major step in the graduation in the CNCF space.
So yeah, thank you so much.
Now, Josh, thank you also so much for being on the show.
As I mentioned, the two of you have been on different stages together,
who is Josh Lee? Can you give us a little bit of a background on your end?
Sure, absolutely. Thanks. Thanks for having me here.
I am an open source advocate at Altinity. So we do stuff with Clickhouse, but we're not affiliated
with Clickhouse Incorporated. And I am one of our observability SMEs because it's a very,
very important use case for Clickhouse. And before that, I was a product manager at an
observability platform, which is where I really got most of my hotel knowledge. And I try to
I try to contribute.
I say I speak about Otel more than I contribute these days,
but I'm still trying to contribute when I come.
And I think what I love about this whole community,
and I'm sure you see this a similar way,
is that although we're working for different vendors
that are trying to solve similar problems,
we get together,
and in a community where we try to understand
what are the real challenges that need to be solved
for the end users.
And then, you know, obviously we still need to all cater to the company
that pays our bills, but still it's really great
the collaboration that you see, especially in the
CNCF, and I'm sure in other
communities as well.
The talk that you gave,
remind me again, where did you
tour around? Where did you give the talk?
Vienna was the first one, right? And then
Warsaw and a couple of online
conferences. Then we gave it virtually
as well. Osad, I think, right?
Yeah. Open source. Observability
day.
Yeah.
And then
Portugal.
We used to
learn it solo a couple of times also at
like meetups and yeah, it's gotten around.
Yeah, I delivered it in
London this
year.
Yeah, yeah, it's been
definitely lots of, lots of
countries. It's gotten a lot of love, which
makes me really happy. I think we even did it
in Amsterdam for rejects, right?
Oh, yeah, yeah, yeah.
Let me ask you a question.
I assume the talk is online available somehow.
One of these talks is available, so we'll definitely link to it.
Oh, yes.
But do you still, as of today, end of June, 2026,
do you still run into people that have the misbelief
that open telemetry by instrumenting their app,
they can, free of choice, just switch the backend vendor
with a click of a button or with a switch of a feature flag,
Or is this still around us?
Do we still need to educate the market?
I still see this marketing message, for sure.
I think that part of the reason this talk has been so popular is because it resonates
with something that a lot of engineers already understood, but maybe didn't have the words for.
So yeah, I think that's part of the reason why it's resonating is that people have kind
of understood this to be a marketing myth, but I still see the messaging for sure.
Yeah, I think it's one of those cases where like when you hear it, you're like, oh, duh,
is so obvious but sometimes you just need someone to spell it out for you and that's essentially
what we've done yeah so now for those listeners that you know are not as deep into open telemetry
and in the ecosystem into the topic in general what is it that makes open telemetry a good
choice but also what is it that doesn't make it completely vendor neutral what are the things that are
not vendor neutral i think let's let's talk about this because people
may not think about it, right?
And maybe we cover the second topic.
I always bring up two questions and it doesn't make sense because then we're switching back and forth.
I do that all the time, you'll worry.
We have curious minds.
Yeah, so what is it?
What is it that people typically don't think about when you think about vendor neutrality.
What is not vendor neutral right now in our ecosystem?
I don't know what you take that one, Adriana.
Yeah, the U.S.
I mean, like, so all the vendors that ingest OTP, Open Telemetry, OTP data natively, they can ingest the data.
And they will render the data.
But it's what is done with the data at the end of the day is what's different.
So like you will see a lot of similarities.
Like you'll probably see different vendors will show you like, you know, the,
the trace waterfall graph that'll be pretty common across the board and the logs that's pretty
common across the board but then i think the the real shock comes with regards to like some like the
smaller things like the dashboards you spent all this time in your previous vendor building out
dashboards and now you're like oh i can't transfer my dashboards for from vendor x to vendor y not only
that, oh, it uses a completely different querying language. Oh, and by the way, those SLOs that I created,
I have to recreate them. And if you used any terraform, because a lot of, a lot of the observability
vendors have terraform providers that help. And the terraform providers will vary from
vendor to vendor, but I think a lot of them have like, you know, dashboard creation capabilities
and whatnot. And again, like, because vendor X is not the same as vendor Y,
those dashboards are going to look different.
That terraform's going to look different.
You can't have like a direct one-to-one translation.
And those are the things that people don't think about.
Like, it's awesome.
Yay, I don't have to re-instrument my code.
And that is huge.
Like, we don't want to, we don't want to downplay that in any way.
Because I think that was a thing that really held people back in the before times.
Because can you imagine you want to break up with your vendor, but you can't because
now you have to strip out all the existing instrumentation that you had.
and then put in new one.
And that means you're introducing technical debt by that very act.
And now we're saying, no, you don't have to do that.
You don't have to touch your coat.
So, yay, that's like a ton of work that's already, like, that you don't have to worry about.
But there is still work in getting to know your new vendor, your new back end.
Josh, anything from your end to it?
What do you see?
I mean, yes, it's exactly that, right?
It's all of those things that don't come with you because the OpenSlimatory Project explicitly does not include a back end.
But it's, and it's, yes, like the myth that we talk about, the myth that we're busting is this idea that like because of Open Cemetery Magic Sauce, you can just switch vendors.
I don't know that many people, right?
Like, people do switch vendors, but it's not a thing that happens often.
It's not something you're going to go through more than every couple of years.
Hopefully, if you do, you have other organizational problems.
Right?
But what's beautiful about open telemetry to me,
and what we talk about in our talk, right,
is that it enables these things to be interoperable all the time.
It enables these things to all work together better all the time.
If I'm Amazon, for example,
I can build open telemetry into my services
without picking a favorite vendor.
And then all of my customers get the benefit of being able to use that first
party telemetry without some vendor trying to like interpret
or reverse engineer what's going on inside my system.
And I use Amazon example, but same thing for anyone building an open source framework or library or any kind of tool that you're going to share with other people.
This vendor neutrality now enables you to build in the instrumentation as the first party and give that to all of your customers, which is amazing.
And it also enables this explosion of new tools that live on that common interoperability.
Like Adrienne, I think you mentioned O'Te DessCop viewer in one of your other talks.
That to me is something that would never could have existed without open telemetry
because who would build that for a specific vendor?
Maybe for the most prominent vendors.
But a universal tool that you can use with anything, that's, I mean, it's awesome.
You know, it's interesting.
You all brought up some really interesting points.
So in my role, I work with a lot of either customers or prospects for Dinotray's, right?
and yeah the most common thing we hear is we want to use open telemetry because we want to be vendor neutral
and sometimes it's because we want to be able to switch tools right and in the back of my mind I'm always kind of half laughing because I'm like
how many people switch tools that often no matter what the tool is right because it's a heavy lift
but I do think there is now I'm sure you're seeing people doing things a lot more advanced with open telemetry and all that
But I think for at least the community of people I see, a lot of it is based on the myth of being able to switch vendors.
You know, there's no mind taken into the heavier lift.
The heavy lift isn't getting instrumentation.
And no matter how you do it, it's about what you do with that data on the back end.
And all the processes you build around that within the organization, the learning of whatever tool it is, the dashboards, as you mentioned.
but there's people a lot of times are stuck on the idea of if I go open telemetry, at least I see a lot more of it than I am vendor neutral, period, right?
So I'm really glad to see you all reinforcing, you know, my counter argument, not against open telemetry, but the trying to dispel that myth.
Like I'm just like, yeah, you can use open telemetry if you want. That's fine.
I think the one thing that you said, Adriana, maybe I'm biased, right?
because I've been working at Dynetrae's for almost 15 years now,
and especially stuff we've done with instrumentation and all.
To me, at least, I've never, for most people I run into with Open Telemetry, right?
They're just doing the basics.
They're not doing customization.
They're doing the auto instrumenter and sending that in.
And to me, that versus most vendors,
it's not even that much of an advantage because it's not like they're going through,
in adding custom telemetry collections,
adding custom instrumentation, right?
If we look at open allelmetry, for instance, right?
That fulfilled, I was really excited when I saw that
because that fulfilled the promise of open telemetry,
right?
I remember the big promise of open telemetry,
at least the one I saw as the big one,
was that all these code vendors
were going to pre-instrument their code with open telemetry
so that when you pop out on, it's pre-instrumented,
you hook up your collector,
and you have vendor-specific instrumentation,
which is awesome.
And that's what we see with open-elometry, right?
It's all very, very specific.
But at least in my world, I haven't seen any vendors doing that with regular code.
And unless you're doing customizations to it, to me it's like it doesn't even give you that much of an advantage because now you have to go through, put all the open-elemetry in, manage it yourself.
and for nothing additional.
Now, of course, if you're going through and doing a lot of custom stuff,
if you are adding those additional components, fantastic,
because then you don't have, because you're not going to get that when you switch vendors, right?
That's your custom code, your everything else.
So if you're taking it to that level, now you're really using the promise of it.
But I think it's just, it's almost like Kubernetes, right?
When Kubernetes came out, people were like, oh, we're going to move to Kubernetes.
Like, okay, well, what are you moving to Kubernetes?
I've got this monolith.
We're moving it there.
Well, why?
Oh, because we want to be on Kubernetes.
And I think there's a lot of hype around going to Open Telemetry to go to it,
but without the thought of it.
Now, again, I don't begrudge anyone for going to Open Telemetry.
If you want to take that approach, absolutely fine.
Like, who cares?
But, like, it's, I think to me, what bothers me is mostly just the idea of,
I'm going to do this and I'm going to be free.
you know it's like you still have like to me vendor locking comes in in the back end right it comes in with that
visual layer because not only dashboards not only how you would analyze not only what you build around the organization
if you have tools that are doing automations based upon that like you know again i'm not not speaking
from a dinah trace advertising point of view but you know like we have at our platform like workflows
and everything else that's automatically built into that so if you were to say have all that stuff
leveraged dinah trace and then you wanted to switch to another vendor,
but you would need that vendor to have all that stuff in there as well.
So to me,
the lock-in always comes from that back-in and what are you doing with the data.
But it's,
I guess the point that I'm making here is a ramble.
I think it's really great that, like,
there's discussions now about, like, what is that lock-in?
What are the benefits?
And, again, I'm not going to knock the idea of having benefits of using open telemetry
because it's great, yeah, and you do feel empowered.
But it's, I think it's, it's, it's,
high time, people take that honest look, as
you did mention in your blog about, like,
where is that locking come in?
And what does that lock in really, really mean?
And it's not just, we'll use open telemetry
and we're going to be absolutely free to, you know,
have 20 dance partners
of one night, you know?
Yeah, one way I put it, right, is open telemetry
liberates your code artifacts and your code
itself, but it doesn't, it doesn't
liberate your entire organization, right?
Which is a completely different beast from your
individual code artifacts. And like you mentioned,
vendors a couple of times there. I think you were mentioning, like, I think you were referring
to non-observability vendors, right? But like tool vendors and library vendors, right?
And so Open Telemetry is absolutely liberating for them.
And you kind of also hinted that something that I think we missed in our talk in our blog post,
which is resume-driven development.
Yeah.
So, like, which is one reason I think that people, right, that's one reason I think people choose
Kubernetes sometimes is because like, right, even if it's not the best technical fit, well,
this is a skill that I need. And so I can.
take it with me to other jobs.
Same thing with Open Telemetry.
Maybe your organization isn't going to switch vendors,
but people switch organizations frequently.
Yeah. Yeah.
Yeah.
Yeah.
I mean, it's so much easier bringing on someone
who knows and understands Open Telemetry
than someone who's like very narrowly focused on vendor X.
I did want to bring up another point,
not necessarily related to the top.
of this talk, but I think an important thing that almost leads into our follow-up talk that we did
on this topic, which is like, you know, when a lot of people get into open telemetry and they're like,
auto-instrumentation is kind of like the gateway drug, right, to telemetry, right? Because it gives you,
it gives you that no-touch instrumentation, right? Like, it's magical. You're like, what? I don't have to
touch my code. And this thing's like, like,
You did the work.
Thank you.
Adding fans and stuff, which is great.
But then you kind of end up with, it's almost like you wake up with the hangover afterwards
and you're like, what have I done?
Because you're realizing that, first of all, the auto instrumentation, it's great,
but it doesn't cover everything.
Right.
So ultimately to expect that you're going to like cover everything with auto instrumentation
and not supplement with manual instrumentation, I think that is a misconception that we need
to bust.
And there's also another thing that a lot of organizations are grappling with, which, you know,
they thought, well, if we instrument everything, then it will be all good.
But something that I started seeing early on, even in my observability journey, the company
I was working at, I remember they were sending telemetry to a vendor.
And they were using open tracing before I was like convincing them to move over to open telemetry.
And they were like, we send all this data to this back.
in, but we don't know what to make of this.
It's like it's too much data.
We don't, it's like looking for a needle in a haystack.
So it's almost like, it's almost as bad as having no data because you don't even know
where to start looking.
So that's another thing that people have to be mindful of when it comes to their telemetry
data.
Like instrumenting isn't going to be the be all and all.
It's instrumenting intelligently and effectively.
Yeah, I think that's a great point because as you mentioned, if you're just going,
with out-of-the-box open telemetry, you're not getting a lot, right?
And my thought process would be if you're going to go with open telemetry,
do it well and add that customization to it.
It's not going to instrument your specific lines of code or your specific methods,
and those are the pieces you're going to need to see some information.
And so if you're going open telemetry, go in and put all that key stuff in that you do need
to be able to see because that's where you're going to get the real advantage from using it.
Because now if you do switch, as you say, everything you've done specifically for your organization
follows you from vendor to vendor.
And I don't want to sound like I'm anti-open telemetage because I'm not at all.
I'm more anti-like the idea that like, you know, we want open telemetry because it's vendor neutral
and just like that statement.
Yeah, it's not a silver bullet.
And I think it's sold this one sometimes.
It's more of like, which is why I appreciate your blog because it's like, okay, no,
the word has to get out there.
like, yes, it's great.
It's fantastic.
But understand the full scope of it and that if you do use it, use it and make it powerful
because it can be extremely powerful, right?
And again, I keep going back to the whole open-allimetry with the AAM models.
Like, that just blew me away when it's like, oh, yeah, you just, you know, add your collector
and this fantastic set of data comes in automatically.
I'm like, this is what it could be, you know?
So it's more of, I guess my side is like, encouraging.
people who use open telemetry to really take advantage of it.
I want to quickly add also something from my side and Adriana we we think
I forwarded you that email we are currently working with Autodesk and trying to
figure out what are the challenges of large organizations and what do they value and
what what did they run into and I want to quote here
Umar Khan hopefully I pronounce his name correctly but he made a very interesting
quote, instrument by value, not by default.
And what he meant with this is really instrument, what you know,
delivers the value to your team so that you know if your system is running properly as expected,
but not just accept the default, because the default could be just collecting data,
but you don't know why you collect it, and then in the end you're proud to decollect data,
but nobody uses it, so it then just becomes a cost factor.
So I really like that instrument by value, not by default.
That's great.
Yeah.
And I think this also means, and I think this is still something,
Josh and Adran, I would like to get your opinion.
When I work with people, most people still don't know what it really is
that they need from their services and applications to tell them whether it runs as expected or not.
I still often get questions like, hey, what is a good metric and what is a good SLO?
And I said, you know, you should know your system best.
What is it that makes your shareholders angry?
What is it that it makes you users angry?
What is it that gets you on the news and then try to figure out what is the telemetry that you need,
but it's a log and metric, a trace, so that you get alerted before there's a problem
that you have enough information to fix the issue.
But it feels even though, I mean, we've all been in the space for so long.
For us, observability is like, you know, we do it in our sleep.
But now we are, with open telemetry especially,
we are pushing this additional task on a group of engineers
that are, I think, still very new to this topic.
And I think something, and I think it feels like a lot of people are still lost
in what they should instrument.
And now they're rather instrumenting too much
than the right thing.
So going with the default versus the value
and then in the end, everybody's screaming,
and now it's very costly.
Some of my thoughts.
That's spot on.
That's spot on.
And I think the things that people need to remember
as fundamentals and Josh, feel free to chime in
and correct me,
if you disagree, but my, my thought is like, I think first and foremost, the trace has to be the
first class citizen of your observability story because it tells the story end to end, right?
And then you have your supporting characters, the logs and the traces, and they're still very
important because they help to add additional context to that.
And keeping that in mind, I think when instrumenting, people get like, especially if the application's
not been instrumented, they get like very overwhelmed.
And they're like, well, where do I start?
I got to start somewhere.
And maybe the temptation is like, you know, I've got a bunch of microservices.
I'm going to instrument my Java microservice, the service that does blah.
But like, think about like, we don't, we don't think about these microservices in isolation.
They're part of a greater whole, right?
So what does the application do?
What are the critical path processes?
like flows, what are the critical path workflows in your organization?
The ones where you get called in the middle of the night most often for, start there, right?
Because these are the ones where you see the issues, maybe the bottlenecks, and that's where you
instrument that flow, whether it's like, you know, and it'll be probably crossing multiple
services and that is what you're going to be focusing on on instrumenting first rather than
let's just instrument everything and get overwhelmed and freak out and then you're like you know
get you get the analysis paralysis what do you what do you what do you think josh no i absolutely
agree with that right like trade yeah start with tracing trace your critical paths from your
tracing you can derive your rd metrics which is your most important like early morning indicator
and health you know health indicator for the overall system once you've done to you've done
tracing properly, you also get a topology, which is really, really useful in debugging.
And then just adding to that, I would say, Ray, like, things that can be compared to other things,
right? Like a telemetry data point is usually meaningless in isolation, but it's when we compare
it to the context that it starts to become meaningful. This is maybe a tangent, but once upon
the time I was running like a mail order business, and I was new to all of this, and I was
doing the order fulfillment, and the order fulfillment system would print all of the packing
slips and then at the end it would print this manifest that had like the total number of packages
going out that day.
I'm like, why is this metric here?
I obviously I know that I have nine packing slips.
I don't need another piece of paper that tells me nine.
But it becomes a really good validation, right?
Then I don't have to look at every single label and make sure that it's correct.
I just look at my basket and say, okay, I've got nine packages.
The slip says nine.
I'm off to the post office.
So being able to compare things is really useful.
Right.
in the hotel world, an example that I come up with,
see all the time, right?
It's just looking at your hotel collector
and making sure that the telemetry out matches the telemetry in,
and that you don't have a bottleneck somewhere.
And then another thing is comparing over time, right?
So we talk about this a little bit in our follow-up talk,
but it's like you want to be, you want to be looking at trends, not moments.
So if my hotel collector is using 700 megabytes of RAM,
is that good? Is that bad?
I don't really know.
that on its own doesn't really tell me anything.
But if I deploy a new image and it goes from using 700 megabytes of RAM
to like 1.4 gigs, that tells me something.
Yeah, and also metrics, as you're talking about the collector,
one of my talks I tried to define also some of the metrics and SLOs
for the observability platform.
Also, what we do internally, we call them critical user journeys.
So, for instance, we measure internally.
How long does it take from,
let's say a log that comes in into our backend until that log is analyzed, stored,
and available on the dashboard, right?
And I think also for if whoever's listening and if you're building your observability platform,
there's a lot of critical indicators like, you know, what does the resource consumption, obviously,
but also what is the time of a metric from creation until it shows up in the dashboard?
How much data loss do you have?
Things like this, right?
I think these are all also critical metrics that we,
the people that are now responsible for operating an end-to-end
observability platform, they need to be aware of.
You mentioned two things there, Andy, that I think are interesting in contrast, right?
You mentioned resource usage, which is where a lot of people start,
and I even mention those metrics because they're easy.
We already get them out of our systems, right?
We can just ask the system.
whereas the things that are user-facing,
the symptoms are actually the things
that we want to be focusing on measuring first,
but they're harder to measure.
And, yeah, getting to that user experience is definitely important.
Yeah.
And it's also that metric that I just meant, right?
Looking at an possibility back and measuring the time
from a signal, from its creation,
until it's stored and available,
is not an easy metric to calculate.
This is something we need to put in thoughts.
What is easier is,
is the individual hops in the middle, as you said,
memory consumption, throughput in the queue,
latency between the systems.
But I think this is what we,
what also,
Uma meant, right,
with instrument by value,
not by default.
What is the value,
what is the stuff that you're promising
with your software to your end users?
And how can you ensure that you're fulfilling that promise?
Like with your nine packages,
you fulfill the promise that all of the packages,
packages have been delivered.
For that you need to know how many packages need to go out.
And to that effect, I think also emphasizing the importance of data correlation
and being able to have a place where you can visualize your telemetry in one spot.
And I think a lot of the telemetry backends out there provide that visualization in one spot.
And I think that's, which is great, but like the correlation is really where the magic sauce is.
Because, you know, in the early days of observability, we'd refer to traces, metrics, and logs as pillars.
And it was aptly named, right, because they were like literally just standing in isolation.
And you had different tool sets that specialized in each one, right?
And then when observability became a thing, then it was.
it's like, oh, these things are actually interrelated.
You know, as I mentioned, like, yes, the traces are the backbone of observability,
but they're partially useful because they still need that support of the metrics and logs.
And how can those metrics and logs be useful if you correlate them back to the traces?
So then your observability becomes more of a braid rather than, you know, a braid of these signals rather than the pillars.
And then it's like, oh, okay, this is how.
this is how I can derive like actual further insights into into my application right like a lot of the
times sometimes your your signal is like I have a huge amount of like memory consumption or latency or
whatever great okay so can we trace that metric back to a corresponding span and a corresponding
trace that that span is part of and that's that is gold for you yeah I love
I love the resource metadata and the semantic convention so much.
We talk about this a lot, but we now have this common language.
Even if your telemetry is stored in disparate systems,
at least you're joining across common,
because of the semantic conventions,
there are common keys and common values that you can join across.
And so you don't have to deal with this incredibly common problem
and observability of like, oh, well, over here, it's CPU underscore usage.
And, you know, an idea is a better one to use,
because that would be something you join across, right?
But over here it's like account dash ID,
and over here it's ACCT, uppercase ID,
and how do you even get that?
Developers don't care about logs metrics and signals, right?
We care about, and developers and operators, right?
We care about services and processes that we are responsible for
and the experiences that they provide.
And so being able to find all of the telemetry
about the entity that we care about
and about the entities that are related to it
is really the superpower.
You know, this makes me think, you know, we go back to the earlier days of what we were doing,
and it was trying to get organizations on board with the idea that performance and observability was important in the first place, right?
We saw that arc from the earlier days where it was like, oh, do we need this?
And then when it became critical.
But even back then, a lot, we were talking about, you know, performance as code or whatever the heck we were calling it, right?
you as a developer should know how your system should perform,
even the idea of checking in a performance metric with your new code base.
Here's my new function.
I'm going to check this into production,
and it's going to perform in 80 milliseconds or less with no errors and whatever it might be, right?
And the idea was trying to get, the challenge is always trying to get developers
to care about their performance.
and I think one of the amazing things about Open Telemetry,
which just during the course of this talk makes me want to embrace it more,
is if you're asking developers to know what to instrument,
know what additional pieces to put in,
not just using the auto instrumenter,
but saying these are the really key parts of the code,
so I'm going to add some instrumentation to it,
that drives them to be much more performance aware of their code.
it gets us to that end state that we've been talking about
for a long time of developers care about your code performance,
put some stuff in there.
Now, if they're doing open telemetry,
and they're not just doing the out-of-the-box open telemetry, so to say,
that's bringing everyone up to this idea, right?
It's going to make everyone say,
these are the important things for me to know,
like whether it's not it's memory or it's CPU
in conjunction with my code performance, right?
You know, one of the old things was like, yeah,
CPU's at 90%.
my code must be bad.
Well, no, your code is just under max load, right?
You need other indicators to understand if that 90% CPU.
So when they start thinking about that and saying,
these are all my indicators,
I want to make sure I have the telemetry for this
so that if it comes back to me, I can figure it out.
I think that's just like an awesome side benefit of open telemetry
because it makes everybody engage in performance even more.
You know?
I want to bring up one more topic
that especially knowing Josh
you mentioned you used to work
for an possibility vendor
I think it was in Stana
that I see on your LinkedIn profile
Adriana you also worked for a different vendor
before Dinah Trace
so we all have different backgrounds
do you see
where does open telemetries
still need to mature
in terms of capabilities
that let's say the long term
vendors have in their agents
is there still, and I know that obviously
profiling was just added or has been
added for multiple languages that's
coming, really user monitoring.
Is there anything else
where you
and I want to know
you also be critical because you have your
background and your history with the vendor
and you know what these agents can do?
Is there anything else that we want to make sure
people are aware of
that there are certain capabilities
where either open telemetry
is already on the right trajectory,
it's planned, or are there still
certain gaps where there might be a need even for a mixed setup between commercial vendors and open
telemetry? Yeah. The biggest gap I see is ergonomics, right? In terms of capabilities, as
you mentioned, the last few things, right, they were very, very close to parity with what I'm aware
of most vendor agents being able to do. But the experience of using it and the learning curve
is pretty brutal with open telemetry. Yeah. Yeah. And
And to be fair, I would say the hotel folks are very aware of that.
And that's why they have, like, there's so many initiatives out there to improve that experience.
Like, there's a developer experience, SIG.
I think there's just, I want to say the injector is a project or a SIG for making it a little bit easier for, like,
configuring open telemetry out of the box because that's that's another thing.
A lot of a lot of development teams will tend to create wrappers around open telemetry
just because that initial setup can be a little bit gnarly to begin with.
And so let's make that as easy as possible.
Like interestingly enough, like I worked at Lightspe up before this.
And lights up had, they had basically like wrapper libraries around O'Tell that had, that had some constructs where basically it already had some configuration stuff in place that made it easy to like, you know, you just set up some environment variables and you could configure your collector to send data lights up a lot more easily.
So there was a lot fewer configuration steps.
So I think the open telemetry folks are very aware of that.
are making steps to improve that.
Another one that I think is a good one,
a good project to watch out for is O'Tl Weaver
around semantic conventions
because I think as companies really start
expanding their use of open telemetry,
semantic conventions become more and more important.
And what do we often see in large organizations?
Everybody is doing their own damn thing.
So the left hand doesn't know what the right hand is doing,
And then all of a sudden we've lost a common language for our telemetry, which is the semantic conventions.
So, I mean, we have the O'Tel semantic conventions, but organizations have their own internal semantic conventions,
i.e. what are the things that are important for them to capture as attributes in their telemetry?
And so O'Tell Weaver is a way for you to define, not only define your semantic conventions, but codify them,
because then it'll, yeah, there's even a setting in Weaver that allows you to take, um, take,
those YAML definitions of your semantic conventions apply a JNJA template and generate code and
documentation like data structures in Go or Java classes to support that structure.
And then on top of that, that's all well and good, but like, how do we ensure that the code is
actually using those semantic conventions and Weaver has like basically a live checker that does
that? So I think those things are very important and getting people educated about those projects
and getting them excited and using them,
I think will help open telemetry as well.
Like there's, I think awareness is a huge piece
because there are so many moving parts.
Like this recognition that like, you know,
there are ergonomic problems.
So then there are new initiatives to address those ergonomic problems.
But is everyone aware?
Like, I can barely keep up.
I'm one of the hotel community managers.
I'm a maintainer of the hotel and user sick.
I can barely keep up.
So imagine, imagine other folks who aren't as, you know,
entrenched in hotel.
Yeah.
Definitely, I just want to say
we definitely going to add Hotel Weaver
and any other links that you have
on these projects to the podcast description.
Sorry, Josh, go ahead.
Oh, no worries.
Yeah, I think this is an opportunity
for the vendors as well, right?
Like, I think we've been asked questions
about this in some of the instances of our talk
and people ask, like, should I use
a vendor implementation
of the Open Telemetry Collector, right?
or a vendor wrapper around the open-sometrial collector.
And to me, right, like at the end of the day,
if you have open telemetry and then open-sometry out,
I'm not going to be a purist about the code that's actually running that, right?
I think OTP and the semantic conventions is really where the magic lies.
There is the challenge then, though, right?
Once you start wrapping those things,
this stuff tends to be moving really, really fast.
And so then you sort of create this currency challenge,
which even just, right, like as I mentioned,
even just keeping up to date with what is the latest version
and being aware of it,
let alone having it as part of your build.
and then disseminated across your entire organization.
It's, yeah, it's a challenge.
I mean, what I see, and Brian, correct me if I'm wrong,
but I think one of the reasons why people may choose
a vendor-specific collector is supportability
and who is responsible if things break.
Do you want to be responsible for yet another very critical component
or do you pay a vendor for being responsible for it?
And I think that's at least what I see.
Yeah, that's exactly it. I mean, enterprises, if you work at a large scale enterprise, or maybe it doesn't even have to be a large scale enterprise, if you have mission critical stuff in production that's using open telemetry and stuff breaks, you need to throw some money at someone that you can call to support you in the middle of the night. And that's extremely important. Like when I spent many years working in a bank and I remember, I remember it kind of blew my mind.
when I first learned about that, I'm like, oh my God, we can't like use anything.
We can't even use like open source tools because they want to like throw money at someone
to provide support.
But it's like, yeah, you work at a bank where there's like critical data at stake, like financial
information, tons of PII data.
You better damn well hope that someone is throwing money at someone else to fix the problem
if things go south, right?
And that's something that needs to be kept in mind.
And I think I think then, you know, if you are a vendor,
providing like your own version of a collector or wrap around the collector.
I think it's important, especially for vendors who are like, who are our hotel vendor,
like who support hotel.
I think it's really important then to ensure that whatever they provide is compatible, like with upstream.
So contributing back changes to upstream, I think is very beneficial for the community.
And, you know, just staying up to date.
staying up to date with hotel. And I think, again, if you are a vendor that that says you support
hotel, I think it's very important, then that you contribute to hotel, which is, you know,
something that I appreciate, like, where we're at Dinotrace. We do have a number of
hotel contributors. I mean, I see this across the board within Open Telemetry. So many big
observability back-ins have people, like, in the, um, uh, at, as, you know, as, you know,
CIG maintainers in the governance committee, technical committee, like representing their
organizations contributing to O'Tell so that we all benefit, right?
Yeah.
You know, it's interesting.
I never thought about the collector as, for lack of a better word, a bottleneck in the
hotel system, right?
Like where, like in my head it was always just, oh, connect the collector, right?
But then obviously you start increasing the load on it
and all this kind of
just other components, right?
It's another piece to think about.
I wonder, and maybe
Josh and Adriana,
you could tell us if it would make sense.
Andy, I'm just thinking, like, you know, we used to do
all of our performance anti-patterns and all.
I wonder if this warrants like a separate episode
on like what you need to
observe on your collector, or not necessarily even observed,
but like what are all the...
How do you scale?
How do you scale a collector, yeah.
I think that...
Up-amp.
Exactly.
I think we just, I think we have our friends from Bindplain.
That would be great to reach out to them and have an episode.
And I mean, the collector is usually critical.
It's not just the load, but it's also depending on which sampling strategy you apply.
That means you need to properly size the collector to be able to hold your traces in memory
if you're doing tail-based sampling and then large organizations,
they have hundreds and thousands of collectors, even more.
I think the Bimplane folks, they talked about a million collectors being deployed at a large
ointz.
Automobile company, right, because there's a collector running in every car.
That's wild.
And even just strategy, like, you know, we're starting with open telemetry, right?
We're setting up our first collector.
there's all these different options.
Like, what's going to set us up for success as we scale out?
Like, what choices should we be making?
What consideration should we be making if, as we grow with our open telemetry
and we do start scaling collectors?
And, yeah, I think that, yeah, I think that's a great conversation.
And one I frankly hadn't thought of, like, I knew there was options in collector,
but it never dawned on me that, like, oh, that's a very important decision to make at that point.
You could go down like a deep rabbit hole just on the hotel operator.
you see.
And also I do want to give like a shout
out also to like the importance
of monitoring your own collector
because the collector can emit its own telemetry.
But then do you use a collector to monitor that?
Sorry?
Then to use a collector to collect the telemetry
from the collector monitoring?
Yes, yes, you can.
So you can use the collector itself.
You can use the collector to collect
its own metrics but not recommended.
So you often do have like
a dedicated collector to collect
the metrics or so the
telemetry of other collectors.
Yeah, very meta, very meta.
But you could also do it in Prometheus exposition format if you don't like using the
same tool to monitor itself.
So you can use some other tool that scrapes the premises metrics, yeah.
Josh, Adriana, unfortunately we are almost at the end of the time.
I just have one last final question.
When will people see you again on stage?
Whether the next conference is coming up where you are talking about this topic?
I think October.
About this topic, I don't know if this topic specifically,
but in October, November, we are giving a couple of keynotes.
So one, I think I want to say November is, oh my God, Cloud Native Denmark.
Cloud Native Denmark, we're giving the keynote there.
And then we're giving a keynote also at Cloud Native Poland.
So cool.
Yeah.
That's awesome. Awesome.
And then also isn't there, what is in Prague?
Open source summit.
Open source summit is happening in Prague.
I think the CFP is still open for that.
Yes, it's open.
As is the CFP is open for observability summit.
So open source summit is October 8th and 9th.
Observability summit is October 5th, both taking place in Prague.
CFP is still open.
Submit your CFPs.
Prague is beautiful.
We are just there for KCD Czech in Slovak.
Oh my God, gorgeous, gorgeous city.
Cool.
All right, Brian, I think we need to close it up, unfortunately.
Yep, we do.
My last thought is just to say that it's amazing where Open Telemetry has come.
I remember what it came out.
It was like, okay, you know, especially having the cynical view of this vendor.
Like, oh, who's this kid who just moved to my neighborhood, right?
But it's really been awesome.
And I think, you know, again, addressing my complaints and negativity earlier, I think
It's more the idea that knowing where open telemetry is now and what you can do with it,
when people come to it with the attitude of, oh, this is just another plug-and-play agent type of thing,
but it's open source, right?
It's so much more than that.
And yeah, you have to do put a little work into it, but, like, you can do quite amazing things with it.
And I guess my frustration really comes down to, you know, people who don't take advantage of what it can bring to you.
Right?
they're going to just do this vanilla deploy and think the world's their oyster.
It's like, no, no, no, it can be your oyster.
But do a little research on it first, learn it first, and do it.
Or even just, you know, I imagine a future where you run your code in an environment
and you have some sort of hotel skill and it's going to analyze your code
and say these are the heavy hitters in your code to add custom instrumentation to
if you don't even want to understand your code.
I mean, there's so much you could do.
So either way, I'm rambling now and...
That's my last thought. Andy, anything from you?
All good.
Thank you.
I just want to thank, especially the two of you,
who are continuously getting the word out
because we need to educate more and more people.
We hope that this podcast contributes to educating more people.
And we will also link to Adriana.
I know you have your own kicking out podcast
and you have your, you know, the analytics can do that with open telemetries.
We'll make sure to eat all the links.
What?
And Josh, if you have any, you know, just send us links that we can add to the description
so that people can read up on the topics that you think are relevant for the community.
Thank you.
Absolutely.
Thanks.
Thank you everyone for listening.
Thank Josh and Adiriana for being on.
And we will see you on the next episode.
Thanks.
Bye-bye.
Bye.
