PurePerformance - OpenTelemetry and the Reality of Vendor Choice with Adriana Villela and Josh Lee

Episode Date: July 6, 2026

I still hear people say, “OpenTelemetry is vendor-neutral, so you can switch any time!”In this episode, Adriana Villela and Josh Lee (both active OpenTelemetry contributors) help bust that myth.Wh...ile OTel standardizes instrumentation and signal transport—and unlocks a rich ecosystem of tools—switching vendors isn’t as simple as it sounds. There’s real cost in retraining engineers, migrating dashboards, SLOs, and alerts, and reworking deep integrations across your delivery pipeline.We also dive into a key challenge the community is tackling: helping engineers instrument by value, not by default—making it easier to capture the right signals with high quality instead of just collecting everything.Here the links we discussed:Adriana's LinkedIn: https://www.linkedin.com/in/adrianavillela/Josh's LinkedIn: https://www.linkedin.com/in/joshuamlee/The blog article: https://thenewstack.io/opentelemetry-vendor-neutrality-guide/CND Austria Talk: https://www.youtube.com/watch?v=1gxLseuaTdMKCD Prague Talk: https://www.youtube.com/watch?v=pPXG20CXKxQOpenTelemetry Project Website: https://opentelemetry.io/

Transcript
Discussion (0)
Starting point is 00:00:00 It's time for pure performance. Get your stopwatches ready. It's time for Pure Performance with Andy Grabner and Brian Wilson. Hello everybody. Welcome to another episode of Pure Performance. My name is Brian Wilson. And I'd like to introduce my real pain in the butt, jerk of a co-host. Andy Grabner.
Starting point is 00:00:37 Andy, how are you doing today? Thank you for, are you using an AI later on to remove all of the nasty words? No, you said, you say, You said you were not to be like, you know, mean person or something before. So I'm just letting people know what you really like. Oh, yeah, of course, yeah. Thank you. Finally, the world knows.
Starting point is 00:00:57 Finally, the world knows behind that laughter is somebody as rude as Abraham Lincoln. A little inside joke action going on there. Anyway, we could talk about, you know, your attitude as much as we want today, Andy. But I think people would rather... Yeah, the show is not about us. No, but even though you always say it's about me, like you're always asking me to make you sound better, make you look better, even though it's a, you know, Andy is a grueling taskmaster, everybody, and just don't let the smile fool you.
Starting point is 00:01:25 So, sir, should I call you, sir? Okay, sir, what are we talking about today? We invited two real celebrities when it comes to open source advocacy for open observability for open telemetry. We've got two people that have been touring the World Conference. stages, and I will let them introduce themselves in a second, but we have Adriana and Josh, who I think have not only presented well together at different conferences, but they've also written a really cool blog post that kind of triggered this whole discussion today.
Starting point is 00:02:07 And the blog post is around vendor neutrality isn't magic, a hard look at the open telemetry ecosystem. I've been working and talking with a lot of organizations. and sometimes there's this misconception out there that you sprinkle a little bit open telemetry on it and then all of the vendor locking problems are gone. And the reality is a little bit different. I really liked a couple of the quotes
Starting point is 00:02:30 that Adriana and Josh had in their blog post. One of them is open telemetry really gives your choice but you still need to make that decision on what you really do with the data in the backend. But yeah, without further ado, thank you so much, Adriana and Josh to the show. I would like to start with Adriana for a quick intro and then also giving the chance to Josh before we dive into the topic.
Starting point is 00:02:52 Adriana, you go first. All right, so my name is Adriana Vilela. I am based out of Toronto, Canada, originally from Brazil. I work with Andy at Dinah Trace and as a principal developer advocate. And I've been in the Devrel space, I guess, since 2022. And before that, just skipped around through various areas of tech, did Java development for a number of years, zigzagged between management and not management
Starting point is 00:03:20 and settled in the DevOps, reliability, observability space is my kind of permanent forever home until the next thing. Went what brought you before I let Josh introduce himself, what brought you into Open Telemetry in particular? It was a bit of a kind of a happy accident where I found myself managing an observability team at two cows. And I didn't really know enough about observability. And so I decided to educate myself to really be able to run this team properly and give proper direction.
Starting point is 00:03:56 And so I was learning in public about it, asking lots of questions on various forums and got to learning about open telemetry. And, you know, I was trying to get folks at two cows to use open telemetry back in 2020. 2022 when Traces was not even generally available. And I'm like, trust me, this thing is going to be big. So it's nice to see that I wasn't wrong. Well, I think you were very right. I mean, all of Open Telemet to talk off, just made a major step in the graduation in the CNCF space.
Starting point is 00:04:32 So yeah, thank you so much. Now, Josh, thank you also so much for being on the show. As I mentioned, the two of you have been on different stages together, who is Josh Lee? Can you give us a little bit of a background on your end? Sure, absolutely. Thanks. Thanks for having me here. I am an open source advocate at Altinity. So we do stuff with Clickhouse, but we're not affiliated with Clickhouse Incorporated. And I am one of our observability SMEs because it's a very, very important use case for Clickhouse. And before that, I was a product manager at an
Starting point is 00:05:06 observability platform, which is where I really got most of my hotel knowledge. And I try to I try to contribute. I say I speak about Otel more than I contribute these days, but I'm still trying to contribute when I come. And I think what I love about this whole community, and I'm sure you see this a similar way, is that although we're working for different vendors that are trying to solve similar problems,
Starting point is 00:05:27 we get together, and in a community where we try to understand what are the real challenges that need to be solved for the end users. And then, you know, obviously we still need to all cater to the company that pays our bills, but still it's really great the collaboration that you see, especially in the CNCF, and I'm sure in other
Starting point is 00:05:46 communities as well. The talk that you gave, remind me again, where did you tour around? Where did you give the talk? Vienna was the first one, right? And then Warsaw and a couple of online conferences. Then we gave it virtually as well. Osad, I think, right?
Starting point is 00:06:05 Yeah. Open source. Observability day. Yeah. And then Portugal. We used to learn it solo a couple of times also at like meetups and yeah, it's gotten around.
Starting point is 00:06:17 Yeah, I delivered it in London this year. Yeah, yeah, it's been definitely lots of, lots of countries. It's gotten a lot of love, which makes me really happy. I think we even did it in Amsterdam for rejects, right?
Starting point is 00:06:34 Oh, yeah, yeah, yeah. Let me ask you a question. I assume the talk is online available somehow. One of these talks is available, so we'll definitely link to it. Oh, yes. But do you still, as of today, end of June, 2026, do you still run into people that have the misbelief that open telemetry by instrumenting their app,
Starting point is 00:07:01 they can, free of choice, just switch the backend vendor with a click of a button or with a switch of a feature flag, Or is this still around us? Do we still need to educate the market? I still see this marketing message, for sure. I think that part of the reason this talk has been so popular is because it resonates with something that a lot of engineers already understood, but maybe didn't have the words for. So yeah, I think that's part of the reason why it's resonating is that people have kind
Starting point is 00:07:31 of understood this to be a marketing myth, but I still see the messaging for sure. Yeah, I think it's one of those cases where like when you hear it, you're like, oh, duh, is so obvious but sometimes you just need someone to spell it out for you and that's essentially what we've done yeah so now for those listeners that you know are not as deep into open telemetry and in the ecosystem into the topic in general what is it that makes open telemetry a good choice but also what is it that doesn't make it completely vendor neutral what are the things that are not vendor neutral i think let's let's talk about this because people may not think about it, right?
Starting point is 00:08:11 And maybe we cover the second topic. I always bring up two questions and it doesn't make sense because then we're switching back and forth. I do that all the time, you'll worry. We have curious minds. Yeah, so what is it? What is it that people typically don't think about when you think about vendor neutrality. What is not vendor neutral right now in our ecosystem? I don't know what you take that one, Adriana.
Starting point is 00:08:36 Yeah, the U.S. I mean, like, so all the vendors that ingest OTP, Open Telemetry, OTP data natively, they can ingest the data. And they will render the data. But it's what is done with the data at the end of the day is what's different. So like you will see a lot of similarities. Like you'll probably see different vendors will show you like, you know, the, the trace waterfall graph that'll be pretty common across the board and the logs that's pretty common across the board but then i think the the real shock comes with regards to like some like the
Starting point is 00:09:23 smaller things like the dashboards you spent all this time in your previous vendor building out dashboards and now you're like oh i can't transfer my dashboards for from vendor x to vendor y not only that, oh, it uses a completely different querying language. Oh, and by the way, those SLOs that I created, I have to recreate them. And if you used any terraform, because a lot of, a lot of the observability vendors have terraform providers that help. And the terraform providers will vary from vendor to vendor, but I think a lot of them have like, you know, dashboard creation capabilities and whatnot. And again, like, because vendor X is not the same as vendor Y, those dashboards are going to look different.
Starting point is 00:10:08 That terraform's going to look different. You can't have like a direct one-to-one translation. And those are the things that people don't think about. Like, it's awesome. Yay, I don't have to re-instrument my code. And that is huge. Like, we don't want to, we don't want to downplay that in any way. Because I think that was a thing that really held people back in the before times.
Starting point is 00:10:27 Because can you imagine you want to break up with your vendor, but you can't because now you have to strip out all the existing instrumentation that you had. and then put in new one. And that means you're introducing technical debt by that very act. And now we're saying, no, you don't have to do that. You don't have to touch your coat. So, yay, that's like a ton of work that's already, like, that you don't have to worry about. But there is still work in getting to know your new vendor, your new back end.
Starting point is 00:11:00 Josh, anything from your end to it? What do you see? I mean, yes, it's exactly that, right? It's all of those things that don't come with you because the OpenSlimatory Project explicitly does not include a back end. But it's, and it's, yes, like the myth that we talk about, the myth that we're busting is this idea that like because of Open Cemetery Magic Sauce, you can just switch vendors. I don't know that many people, right? Like, people do switch vendors, but it's not a thing that happens often. It's not something you're going to go through more than every couple of years.
Starting point is 00:11:31 Hopefully, if you do, you have other organizational problems. Right? But what's beautiful about open telemetry to me, and what we talk about in our talk, right, is that it enables these things to be interoperable all the time. It enables these things to all work together better all the time. If I'm Amazon, for example, I can build open telemetry into my services
Starting point is 00:11:51 without picking a favorite vendor. And then all of my customers get the benefit of being able to use that first party telemetry without some vendor trying to like interpret or reverse engineer what's going on inside my system. And I use Amazon example, but same thing for anyone building an open source framework or library or any kind of tool that you're going to share with other people. This vendor neutrality now enables you to build in the instrumentation as the first party and give that to all of your customers, which is amazing. And it also enables this explosion of new tools that live on that common interoperability. Like Adrienne, I think you mentioned O'Te DessCop viewer in one of your other talks.
Starting point is 00:12:29 That to me is something that would never could have existed without open telemetry because who would build that for a specific vendor? Maybe for the most prominent vendors. But a universal tool that you can use with anything, that's, I mean, it's awesome. You know, it's interesting. You all brought up some really interesting points. So in my role, I work with a lot of either customers or prospects for Dinotray's, right? and yeah the most common thing we hear is we want to use open telemetry because we want to be vendor neutral
Starting point is 00:13:03 and sometimes it's because we want to be able to switch tools right and in the back of my mind I'm always kind of half laughing because I'm like how many people switch tools that often no matter what the tool is right because it's a heavy lift but I do think there is now I'm sure you're seeing people doing things a lot more advanced with open telemetry and all that But I think for at least the community of people I see, a lot of it is based on the myth of being able to switch vendors. You know, there's no mind taken into the heavier lift. The heavy lift isn't getting instrumentation. And no matter how you do it, it's about what you do with that data on the back end. And all the processes you build around that within the organization, the learning of whatever tool it is, the dashboards, as you mentioned.
Starting point is 00:13:51 but there's people a lot of times are stuck on the idea of if I go open telemetry, at least I see a lot more of it than I am vendor neutral, period, right? So I'm really glad to see you all reinforcing, you know, my counter argument, not against open telemetry, but the trying to dispel that myth. Like I'm just like, yeah, you can use open telemetry if you want. That's fine. I think the one thing that you said, Adriana, maybe I'm biased, right? because I've been working at Dynetrae's for almost 15 years now, and especially stuff we've done with instrumentation and all. To me, at least, I've never, for most people I run into with Open Telemetry, right? They're just doing the basics.
Starting point is 00:14:35 They're not doing customization. They're doing the auto instrumenter and sending that in. And to me, that versus most vendors, it's not even that much of an advantage because it's not like they're going through, in adding custom telemetry collections, adding custom instrumentation, right? If we look at open allelmetry, for instance, right? That fulfilled, I was really excited when I saw that
Starting point is 00:15:00 because that fulfilled the promise of open telemetry, right? I remember the big promise of open telemetry, at least the one I saw as the big one, was that all these code vendors were going to pre-instrument their code with open telemetry so that when you pop out on, it's pre-instrumented, you hook up your collector,
Starting point is 00:15:16 and you have vendor-specific instrumentation, which is awesome. And that's what we see with open-elometry, right? It's all very, very specific. But at least in my world, I haven't seen any vendors doing that with regular code. And unless you're doing customizations to it, to me it's like it doesn't even give you that much of an advantage because now you have to go through, put all the open-elemetry in, manage it yourself. and for nothing additional. Now, of course, if you're going through and doing a lot of custom stuff,
Starting point is 00:15:55 if you are adding those additional components, fantastic, because then you don't have, because you're not going to get that when you switch vendors, right? That's your custom code, your everything else. So if you're taking it to that level, now you're really using the promise of it. But I think it's just, it's almost like Kubernetes, right? When Kubernetes came out, people were like, oh, we're going to move to Kubernetes. Like, okay, well, what are you moving to Kubernetes? I've got this monolith.
Starting point is 00:16:18 We're moving it there. Well, why? Oh, because we want to be on Kubernetes. And I think there's a lot of hype around going to Open Telemetry to go to it, but without the thought of it. Now, again, I don't begrudge anyone for going to Open Telemetry. If you want to take that approach, absolutely fine. Like, who cares?
Starting point is 00:16:33 But, like, it's, I think to me, what bothers me is mostly just the idea of, I'm going to do this and I'm going to be free. you know it's like you still have like to me vendor locking comes in in the back end right it comes in with that visual layer because not only dashboards not only how you would analyze not only what you build around the organization if you have tools that are doing automations based upon that like you know again i'm not not speaking from a dinah trace advertising point of view but you know like we have at our platform like workflows and everything else that's automatically built into that so if you were to say have all that stuff leveraged dinah trace and then you wanted to switch to another vendor,
Starting point is 00:17:16 but you would need that vendor to have all that stuff in there as well. So to me, the lock-in always comes from that back-in and what are you doing with the data. But it's, I guess the point that I'm making here is a ramble. I think it's really great that, like, there's discussions now about, like, what is that lock-in? What are the benefits?
Starting point is 00:17:34 And, again, I'm not going to knock the idea of having benefits of using open telemetry because it's great, yeah, and you do feel empowered. But it's, I think it's, it's, it's, high time, people take that honest look, as you did mention in your blog about, like, where is that locking come in? And what does that lock in really, really mean? And it's not just, we'll use open telemetry
Starting point is 00:17:55 and we're going to be absolutely free to, you know, have 20 dance partners of one night, you know? Yeah, one way I put it, right, is open telemetry liberates your code artifacts and your code itself, but it doesn't, it doesn't liberate your entire organization, right? Which is a completely different beast from your
Starting point is 00:18:11 individual code artifacts. And like you mentioned, vendors a couple of times there. I think you were mentioning, like, I think you were referring to non-observability vendors, right? But like tool vendors and library vendors, right? And so Open Telemetry is absolutely liberating for them. And you kind of also hinted that something that I think we missed in our talk in our blog post, which is resume-driven development. Yeah. So, like, which is one reason I think that people, right, that's one reason I think people choose
Starting point is 00:18:37 Kubernetes sometimes is because like, right, even if it's not the best technical fit, well, this is a skill that I need. And so I can. take it with me to other jobs. Same thing with Open Telemetry. Maybe your organization isn't going to switch vendors, but people switch organizations frequently. Yeah. Yeah. Yeah.
Starting point is 00:18:54 Yeah. I mean, it's so much easier bringing on someone who knows and understands Open Telemetry than someone who's like very narrowly focused on vendor X. I did want to bring up another point, not necessarily related to the top. of this talk, but I think an important thing that almost leads into our follow-up talk that we did on this topic, which is like, you know, when a lot of people get into open telemetry and they're like,
Starting point is 00:19:29 auto-instrumentation is kind of like the gateway drug, right, to telemetry, right? Because it gives you, it gives you that no-touch instrumentation, right? Like, it's magical. You're like, what? I don't have to touch my code. And this thing's like, like, You did the work. Thank you. Adding fans and stuff, which is great. But then you kind of end up with, it's almost like you wake up with the hangover afterwards and you're like, what have I done?
Starting point is 00:19:53 Because you're realizing that, first of all, the auto instrumentation, it's great, but it doesn't cover everything. Right. So ultimately to expect that you're going to like cover everything with auto instrumentation and not supplement with manual instrumentation, I think that is a misconception that we need to bust. And there's also another thing that a lot of organizations are grappling with, which, you know, they thought, well, if we instrument everything, then it will be all good.
Starting point is 00:20:23 But something that I started seeing early on, even in my observability journey, the company I was working at, I remember they were sending telemetry to a vendor. And they were using open tracing before I was like convincing them to move over to open telemetry. And they were like, we send all this data to this back. in, but we don't know what to make of this. It's like it's too much data. We don't, it's like looking for a needle in a haystack. So it's almost like, it's almost as bad as having no data because you don't even know
Starting point is 00:20:54 where to start looking. So that's another thing that people have to be mindful of when it comes to their telemetry data. Like instrumenting isn't going to be the be all and all. It's instrumenting intelligently and effectively. Yeah, I think that's a great point because as you mentioned, if you're just going, with out-of-the-box open telemetry, you're not getting a lot, right? And my thought process would be if you're going to go with open telemetry,
Starting point is 00:21:25 do it well and add that customization to it. It's not going to instrument your specific lines of code or your specific methods, and those are the pieces you're going to need to see some information. And so if you're going open telemetry, go in and put all that key stuff in that you do need to be able to see because that's where you're going to get the real advantage from using it. Because now if you do switch, as you say, everything you've done specifically for your organization follows you from vendor to vendor. And I don't want to sound like I'm anti-open telemetage because I'm not at all.
Starting point is 00:21:58 I'm more anti-like the idea that like, you know, we want open telemetry because it's vendor neutral and just like that statement. Yeah, it's not a silver bullet. And I think it's sold this one sometimes. It's more of like, which is why I appreciate your blog because it's like, okay, no, the word has to get out there. like, yes, it's great. It's fantastic.
Starting point is 00:22:14 But understand the full scope of it and that if you do use it, use it and make it powerful because it can be extremely powerful, right? And again, I keep going back to the whole open-allimetry with the AAM models. Like, that just blew me away when it's like, oh, yeah, you just, you know, add your collector and this fantastic set of data comes in automatically. I'm like, this is what it could be, you know? So it's more of, I guess my side is like, encouraging. people who use open telemetry to really take advantage of it.
Starting point is 00:22:44 I want to quickly add also something from my side and Adriana we we think I forwarded you that email we are currently working with Autodesk and trying to figure out what are the challenges of large organizations and what do they value and what what did they run into and I want to quote here Umar Khan hopefully I pronounce his name correctly but he made a very interesting quote, instrument by value, not by default. And what he meant with this is really instrument, what you know, delivers the value to your team so that you know if your system is running properly as expected,
Starting point is 00:23:25 but not just accept the default, because the default could be just collecting data, but you don't know why you collect it, and then in the end you're proud to decollect data, but nobody uses it, so it then just becomes a cost factor. So I really like that instrument by value, not by default. That's great. Yeah. And I think this also means, and I think this is still something, Josh and Adran, I would like to get your opinion.
Starting point is 00:23:52 When I work with people, most people still don't know what it really is that they need from their services and applications to tell them whether it runs as expected or not. I still often get questions like, hey, what is a good metric and what is a good SLO? And I said, you know, you should know your system best. What is it that makes your shareholders angry? What is it that it makes you users angry? What is it that gets you on the news and then try to figure out what is the telemetry that you need, but it's a log and metric, a trace, so that you get alerted before there's a problem
Starting point is 00:24:32 that you have enough information to fix the issue. But it feels even though, I mean, we've all been in the space for so long. For us, observability is like, you know, we do it in our sleep. But now we are, with open telemetry especially, we are pushing this additional task on a group of engineers that are, I think, still very new to this topic. And I think something, and I think it feels like a lot of people are still lost in what they should instrument.
Starting point is 00:25:02 And now they're rather instrumenting too much than the right thing. So going with the default versus the value and then in the end, everybody's screaming, and now it's very costly. Some of my thoughts. That's spot on. That's spot on.
Starting point is 00:25:19 And I think the things that people need to remember as fundamentals and Josh, feel free to chime in and correct me, if you disagree, but my, my thought is like, I think first and foremost, the trace has to be the first class citizen of your observability story because it tells the story end to end, right? And then you have your supporting characters, the logs and the traces, and they're still very important because they help to add additional context to that. And keeping that in mind, I think when instrumenting, people get like, especially if the application's
Starting point is 00:26:00 not been instrumented, they get like very overwhelmed. And they're like, well, where do I start? I got to start somewhere. And maybe the temptation is like, you know, I've got a bunch of microservices. I'm going to instrument my Java microservice, the service that does blah. But like, think about like, we don't, we don't think about these microservices in isolation. They're part of a greater whole, right? So what does the application do?
Starting point is 00:26:25 What are the critical path processes? like flows, what are the critical path workflows in your organization? The ones where you get called in the middle of the night most often for, start there, right? Because these are the ones where you see the issues, maybe the bottlenecks, and that's where you instrument that flow, whether it's like, you know, and it'll be probably crossing multiple services and that is what you're going to be focusing on on instrumenting first rather than let's just instrument everything and get overwhelmed and freak out and then you're like you know get you get the analysis paralysis what do you what do you what do you think josh no i absolutely
Starting point is 00:27:12 agree with that right like trade yeah start with tracing trace your critical paths from your tracing you can derive your rd metrics which is your most important like early morning indicator and health you know health indicator for the overall system once you've done to you've done tracing properly, you also get a topology, which is really, really useful in debugging. And then just adding to that, I would say, Ray, like, things that can be compared to other things, right? Like a telemetry data point is usually meaningless in isolation, but it's when we compare it to the context that it starts to become meaningful. This is maybe a tangent, but once upon the time I was running like a mail order business, and I was new to all of this, and I was
Starting point is 00:27:49 doing the order fulfillment, and the order fulfillment system would print all of the packing slips and then at the end it would print this manifest that had like the total number of packages going out that day. I'm like, why is this metric here? I obviously I know that I have nine packing slips. I don't need another piece of paper that tells me nine. But it becomes a really good validation, right? Then I don't have to look at every single label and make sure that it's correct.
Starting point is 00:28:12 I just look at my basket and say, okay, I've got nine packages. The slip says nine. I'm off to the post office. So being able to compare things is really useful. Right. in the hotel world, an example that I come up with, see all the time, right? It's just looking at your hotel collector
Starting point is 00:28:27 and making sure that the telemetry out matches the telemetry in, and that you don't have a bottleneck somewhere. And then another thing is comparing over time, right? So we talk about this a little bit in our follow-up talk, but it's like you want to be, you want to be looking at trends, not moments. So if my hotel collector is using 700 megabytes of RAM, is that good? Is that bad? I don't really know.
Starting point is 00:28:51 that on its own doesn't really tell me anything. But if I deploy a new image and it goes from using 700 megabytes of RAM to like 1.4 gigs, that tells me something. Yeah, and also metrics, as you're talking about the collector, one of my talks I tried to define also some of the metrics and SLOs for the observability platform. Also, what we do internally, we call them critical user journeys. So, for instance, we measure internally.
Starting point is 00:29:21 How long does it take from, let's say a log that comes in into our backend until that log is analyzed, stored, and available on the dashboard, right? And I think also for if whoever's listening and if you're building your observability platform, there's a lot of critical indicators like, you know, what does the resource consumption, obviously, but also what is the time of a metric from creation until it shows up in the dashboard? How much data loss do you have? Things like this, right?
Starting point is 00:29:54 I think these are all also critical metrics that we, the people that are now responsible for operating an end-to-end observability platform, they need to be aware of. You mentioned two things there, Andy, that I think are interesting in contrast, right? You mentioned resource usage, which is where a lot of people start, and I even mention those metrics because they're easy. We already get them out of our systems, right? We can just ask the system.
Starting point is 00:30:20 whereas the things that are user-facing, the symptoms are actually the things that we want to be focusing on measuring first, but they're harder to measure. And, yeah, getting to that user experience is definitely important. Yeah. And it's also that metric that I just meant, right? Looking at an possibility back and measuring the time
Starting point is 00:30:39 from a signal, from its creation, until it's stored and available, is not an easy metric to calculate. This is something we need to put in thoughts. What is easier is, is the individual hops in the middle, as you said, memory consumption, throughput in the queue, latency between the systems.
Starting point is 00:30:59 But I think this is what we, what also, Uma meant, right, with instrument by value, not by default. What is the value, what is the stuff that you're promising with your software to your end users?
Starting point is 00:31:14 And how can you ensure that you're fulfilling that promise? Like with your nine packages, you fulfill the promise that all of the packages, packages have been delivered. For that you need to know how many packages need to go out. And to that effect, I think also emphasizing the importance of data correlation and being able to have a place where you can visualize your telemetry in one spot. And I think a lot of the telemetry backends out there provide that visualization in one spot.
Starting point is 00:31:50 And I think that's, which is great, but like the correlation is really where the magic sauce is. Because, you know, in the early days of observability, we'd refer to traces, metrics, and logs as pillars. And it was aptly named, right, because they were like literally just standing in isolation. And you had different tool sets that specialized in each one, right? And then when observability became a thing, then it was. it's like, oh, these things are actually interrelated. You know, as I mentioned, like, yes, the traces are the backbone of observability, but they're partially useful because they still need that support of the metrics and logs.
Starting point is 00:32:34 And how can those metrics and logs be useful if you correlate them back to the traces? So then your observability becomes more of a braid rather than, you know, a braid of these signals rather than the pillars. And then it's like, oh, okay, this is how. this is how I can derive like actual further insights into into my application right like a lot of the times sometimes your your signal is like I have a huge amount of like memory consumption or latency or whatever great okay so can we trace that metric back to a corresponding span and a corresponding trace that that span is part of and that's that is gold for you yeah I love I love the resource metadata and the semantic convention so much.
Starting point is 00:33:22 We talk about this a lot, but we now have this common language. Even if your telemetry is stored in disparate systems, at least you're joining across common, because of the semantic conventions, there are common keys and common values that you can join across. And so you don't have to deal with this incredibly common problem and observability of like, oh, well, over here, it's CPU underscore usage. And, you know, an idea is a better one to use,
Starting point is 00:33:49 because that would be something you join across, right? But over here it's like account dash ID, and over here it's ACCT, uppercase ID, and how do you even get that? Developers don't care about logs metrics and signals, right? We care about, and developers and operators, right? We care about services and processes that we are responsible for and the experiences that they provide.
Starting point is 00:34:11 And so being able to find all of the telemetry about the entity that we care about and about the entities that are related to it is really the superpower. You know, this makes me think, you know, we go back to the earlier days of what we were doing, and it was trying to get organizations on board with the idea that performance and observability was important in the first place, right? We saw that arc from the earlier days where it was like, oh, do we need this? And then when it became critical.
Starting point is 00:34:40 But even back then, a lot, we were talking about, you know, performance as code or whatever the heck we were calling it, right? you as a developer should know how your system should perform, even the idea of checking in a performance metric with your new code base. Here's my new function. I'm going to check this into production, and it's going to perform in 80 milliseconds or less with no errors and whatever it might be, right? And the idea was trying to get, the challenge is always trying to get developers to care about their performance.
Starting point is 00:35:15 and I think one of the amazing things about Open Telemetry, which just during the course of this talk makes me want to embrace it more, is if you're asking developers to know what to instrument, know what additional pieces to put in, not just using the auto instrumenter, but saying these are the really key parts of the code, so I'm going to add some instrumentation to it, that drives them to be much more performance aware of their code.
Starting point is 00:35:44 it gets us to that end state that we've been talking about for a long time of developers care about your code performance, put some stuff in there. Now, if they're doing open telemetry, and they're not just doing the out-of-the-box open telemetry, so to say, that's bringing everyone up to this idea, right? It's going to make everyone say, these are the important things for me to know,
Starting point is 00:36:06 like whether it's not it's memory or it's CPU in conjunction with my code performance, right? You know, one of the old things was like, yeah, CPU's at 90%. my code must be bad. Well, no, your code is just under max load, right? You need other indicators to understand if that 90% CPU. So when they start thinking about that and saying,
Starting point is 00:36:23 these are all my indicators, I want to make sure I have the telemetry for this so that if it comes back to me, I can figure it out. I think that's just like an awesome side benefit of open telemetry because it makes everybody engage in performance even more. You know? I want to bring up one more topic that especially knowing Josh
Starting point is 00:36:45 you mentioned you used to work for an possibility vendor I think it was in Stana that I see on your LinkedIn profile Adriana you also worked for a different vendor before Dinah Trace so we all have different backgrounds do you see
Starting point is 00:37:02 where does open telemetries still need to mature in terms of capabilities that let's say the long term vendors have in their agents is there still, and I know that obviously profiling was just added or has been added for multiple languages that's
Starting point is 00:37:19 coming, really user monitoring. Is there anything else where you and I want to know you also be critical because you have your background and your history with the vendor and you know what these agents can do? Is there anything else that we want to make sure
Starting point is 00:37:33 people are aware of that there are certain capabilities where either open telemetry is already on the right trajectory, it's planned, or are there still certain gaps where there might be a need even for a mixed setup between commercial vendors and open telemetry? Yeah. The biggest gap I see is ergonomics, right? In terms of capabilities, as you mentioned, the last few things, right, they were very, very close to parity with what I'm aware
Starting point is 00:38:03 of most vendor agents being able to do. But the experience of using it and the learning curve is pretty brutal with open telemetry. Yeah. Yeah. And And to be fair, I would say the hotel folks are very aware of that. And that's why they have, like, there's so many initiatives out there to improve that experience. Like, there's a developer experience, SIG. I think there's just, I want to say the injector is a project or a SIG for making it a little bit easier for, like, configuring open telemetry out of the box because that's that's another thing. A lot of a lot of development teams will tend to create wrappers around open telemetry
Starting point is 00:38:51 just because that initial setup can be a little bit gnarly to begin with. And so let's make that as easy as possible. Like interestingly enough, like I worked at Lightspe up before this. And lights up had, they had basically like wrapper libraries around O'Tell that had, that had some constructs where basically it already had some configuration stuff in place that made it easy to like, you know, you just set up some environment variables and you could configure your collector to send data lights up a lot more easily. So there was a lot fewer configuration steps. So I think the open telemetry folks are very aware of that. are making steps to improve that. Another one that I think is a good one,
Starting point is 00:39:42 a good project to watch out for is O'Tl Weaver around semantic conventions because I think as companies really start expanding their use of open telemetry, semantic conventions become more and more important. And what do we often see in large organizations? Everybody is doing their own damn thing. So the left hand doesn't know what the right hand is doing,
Starting point is 00:40:06 And then all of a sudden we've lost a common language for our telemetry, which is the semantic conventions. So, I mean, we have the O'Tel semantic conventions, but organizations have their own internal semantic conventions, i.e. what are the things that are important for them to capture as attributes in their telemetry? And so O'Tell Weaver is a way for you to define, not only define your semantic conventions, but codify them, because then it'll, yeah, there's even a setting in Weaver that allows you to take, um, take, those YAML definitions of your semantic conventions apply a JNJA template and generate code and documentation like data structures in Go or Java classes to support that structure. And then on top of that, that's all well and good, but like, how do we ensure that the code is
Starting point is 00:40:55 actually using those semantic conventions and Weaver has like basically a live checker that does that? So I think those things are very important and getting people educated about those projects and getting them excited and using them, I think will help open telemetry as well. Like there's, I think awareness is a huge piece because there are so many moving parts. Like this recognition that like, you know, there are ergonomic problems.
Starting point is 00:41:18 So then there are new initiatives to address those ergonomic problems. But is everyone aware? Like, I can barely keep up. I'm one of the hotel community managers. I'm a maintainer of the hotel and user sick. I can barely keep up. So imagine, imagine other folks who aren't as, you know, entrenched in hotel.
Starting point is 00:41:37 Yeah. Definitely, I just want to say we definitely going to add Hotel Weaver and any other links that you have on these projects to the podcast description. Sorry, Josh, go ahead. Oh, no worries. Yeah, I think this is an opportunity
Starting point is 00:41:53 for the vendors as well, right? Like, I think we've been asked questions about this in some of the instances of our talk and people ask, like, should I use a vendor implementation of the Open Telemetry Collector, right? or a vendor wrapper around the open-sometrial collector. And to me, right, like at the end of the day,
Starting point is 00:42:08 if you have open telemetry and then open-sometry out, I'm not going to be a purist about the code that's actually running that, right? I think OTP and the semantic conventions is really where the magic lies. There is the challenge then, though, right? Once you start wrapping those things, this stuff tends to be moving really, really fast. And so then you sort of create this currency challenge, which even just, right, like as I mentioned,
Starting point is 00:42:28 even just keeping up to date with what is the latest version and being aware of it, let alone having it as part of your build. and then disseminated across your entire organization. It's, yeah, it's a challenge. I mean, what I see, and Brian, correct me if I'm wrong, but I think one of the reasons why people may choose a vendor-specific collector is supportability
Starting point is 00:42:51 and who is responsible if things break. Do you want to be responsible for yet another very critical component or do you pay a vendor for being responsible for it? And I think that's at least what I see. Yeah, that's exactly it. I mean, enterprises, if you work at a large scale enterprise, or maybe it doesn't even have to be a large scale enterprise, if you have mission critical stuff in production that's using open telemetry and stuff breaks, you need to throw some money at someone that you can call to support you in the middle of the night. And that's extremely important. Like when I spent many years working in a bank and I remember, I remember it kind of blew my mind. when I first learned about that, I'm like, oh my God, we can't like use anything. We can't even use like open source tools because they want to like throw money at someone to provide support.
Starting point is 00:43:42 But it's like, yeah, you work at a bank where there's like critical data at stake, like financial information, tons of PII data. You better damn well hope that someone is throwing money at someone else to fix the problem if things go south, right? And that's something that needs to be kept in mind. And I think I think then, you know, if you are a vendor, providing like your own version of a collector or wrap around the collector. I think it's important, especially for vendors who are like, who are our hotel vendor,
Starting point is 00:44:15 like who support hotel. I think it's really important then to ensure that whatever they provide is compatible, like with upstream. So contributing back changes to upstream, I think is very beneficial for the community. And, you know, just staying up to date. staying up to date with hotel. And I think, again, if you are a vendor that that says you support hotel, I think it's very important, then that you contribute to hotel, which is, you know, something that I appreciate, like, where we're at Dinotrace. We do have a number of hotel contributors. I mean, I see this across the board within Open Telemetry. So many big
Starting point is 00:44:56 observability back-ins have people, like, in the, um, uh, at, as, you know, as, you know, CIG maintainers in the governance committee, technical committee, like representing their organizations contributing to O'Tell so that we all benefit, right? Yeah. You know, it's interesting. I never thought about the collector as, for lack of a better word, a bottleneck in the hotel system, right? Like where, like in my head it was always just, oh, connect the collector, right?
Starting point is 00:45:31 But then obviously you start increasing the load on it and all this kind of just other components, right? It's another piece to think about. I wonder, and maybe Josh and Adriana, you could tell us if it would make sense. Andy, I'm just thinking, like, you know, we used to do
Starting point is 00:45:47 all of our performance anti-patterns and all. I wonder if this warrants like a separate episode on like what you need to observe on your collector, or not necessarily even observed, but like what are all the... How do you scale? How do you scale a collector, yeah. I think that...
Starting point is 00:46:04 Up-amp. Exactly. I think we just, I think we have our friends from Bindplain. That would be great to reach out to them and have an episode. And I mean, the collector is usually critical. It's not just the load, but it's also depending on which sampling strategy you apply. That means you need to properly size the collector to be able to hold your traces in memory if you're doing tail-based sampling and then large organizations,
Starting point is 00:46:32 they have hundreds and thousands of collectors, even more. I think the Bimplane folks, they talked about a million collectors being deployed at a large ointz. Automobile company, right, because there's a collector running in every car. That's wild. And even just strategy, like, you know, we're starting with open telemetry, right? We're setting up our first collector. there's all these different options.
Starting point is 00:46:58 Like, what's going to set us up for success as we scale out? Like, what choices should we be making? What consideration should we be making if, as we grow with our open telemetry and we do start scaling collectors? And, yeah, I think that, yeah, I think that's a great conversation. And one I frankly hadn't thought of, like, I knew there was options in collector, but it never dawned on me that, like, oh, that's a very important decision to make at that point. You could go down like a deep rabbit hole just on the hotel operator.
Starting point is 00:47:25 you see. And also I do want to give like a shout out also to like the importance of monitoring your own collector because the collector can emit its own telemetry. But then do you use a collector to monitor that? Sorry? Then to use a collector to collect the telemetry
Starting point is 00:47:41 from the collector monitoring? Yes, yes, you can. So you can use the collector itself. You can use the collector to collect its own metrics but not recommended. So you often do have like a dedicated collector to collect the metrics or so the
Starting point is 00:47:55 telemetry of other collectors. Yeah, very meta, very meta. But you could also do it in Prometheus exposition format if you don't like using the same tool to monitor itself. So you can use some other tool that scrapes the premises metrics, yeah. Josh, Adriana, unfortunately we are almost at the end of the time. I just have one last final question. When will people see you again on stage?
Starting point is 00:48:20 Whether the next conference is coming up where you are talking about this topic? I think October. About this topic, I don't know if this topic specifically, but in October, November, we are giving a couple of keynotes. So one, I think I want to say November is, oh my God, Cloud Native Denmark. Cloud Native Denmark, we're giving the keynote there. And then we're giving a keynote also at Cloud Native Poland. So cool.
Starting point is 00:48:57 Yeah. That's awesome. Awesome. And then also isn't there, what is in Prague? Open source summit. Open source summit is happening in Prague. I think the CFP is still open for that. Yes, it's open. As is the CFP is open for observability summit.
Starting point is 00:49:12 So open source summit is October 8th and 9th. Observability summit is October 5th, both taking place in Prague. CFP is still open. Submit your CFPs. Prague is beautiful. We are just there for KCD Czech in Slovak. Oh my God, gorgeous, gorgeous city. Cool.
Starting point is 00:49:31 All right, Brian, I think we need to close it up, unfortunately. Yep, we do. My last thought is just to say that it's amazing where Open Telemetry has come. I remember what it came out. It was like, okay, you know, especially having the cynical view of this vendor. Like, oh, who's this kid who just moved to my neighborhood, right? But it's really been awesome. And I think, you know, again, addressing my complaints and negativity earlier, I think
Starting point is 00:49:56 It's more the idea that knowing where open telemetry is now and what you can do with it, when people come to it with the attitude of, oh, this is just another plug-and-play agent type of thing, but it's open source, right? It's so much more than that. And yeah, you have to do put a little work into it, but, like, you can do quite amazing things with it. And I guess my frustration really comes down to, you know, people who don't take advantage of what it can bring to you. Right? they're going to just do this vanilla deploy and think the world's their oyster.
Starting point is 00:50:28 It's like, no, no, no, it can be your oyster. But do a little research on it first, learn it first, and do it. Or even just, you know, I imagine a future where you run your code in an environment and you have some sort of hotel skill and it's going to analyze your code and say these are the heavy hitters in your code to add custom instrumentation to if you don't even want to understand your code. I mean, there's so much you could do. So either way, I'm rambling now and...
Starting point is 00:50:54 That's my last thought. Andy, anything from you? All good. Thank you. I just want to thank, especially the two of you, who are continuously getting the word out because we need to educate more and more people. We hope that this podcast contributes to educating more people. And we will also link to Adriana.
Starting point is 00:51:15 I know you have your own kicking out podcast and you have your, you know, the analytics can do that with open telemetries. We'll make sure to eat all the links. What? And Josh, if you have any, you know, just send us links that we can add to the description so that people can read up on the topics that you think are relevant for the community. Thank you. Absolutely.
Starting point is 00:51:37 Thanks. Thank you everyone for listening. Thank Josh and Adiriana for being on. And we will see you on the next episode. Thanks. Bye-bye. Bye.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.