PurePerformance - Blueprints for OTel Success: Standardizing Observability at Scale with Dan Gomez Blanco

Episode Date: July 20, 2026

"There is no single way to deploy OpenTelemetry at scale—and that’s exactly the challenge."As organizations adopt OTel across teams and environments, they face tough questions around standardizati...on, configuration, and operating resilient observability pipelines.To address these challenges, the OpenTelemetry community has introduced Blueprints and Reference Implementations—practical guidance on topics like data standards, consistent agent and collector configuration, pipeline resilience, and intelligent sampling.In this episode, we’re joined by Dan Gomez Blanco, maintainer of the OpenTelemetry End-User SIG, to explore real-world reference architectures from organizations like Skyscanner, Adobe, and Mastodon.Tune in to learn how the community is turning OTel complexity into shared best practices—and how you can contribute your own blueprint

Transcript
Discussion (0)
Starting point is 00:00:00 It's time for pure performance. Get your stopwatches ready. It's time for Pure Performance with Andy Grabner and Brian Wilson. Hello, everybody, and welcome to another episode of Pure Performance. My name is Brian Wilson. And as always, we have with me. We have with me, that's awesome. We have with me our co-hosts, Andy Grabner.
Starting point is 00:00:38 How are you doing today, Andy? Good. Have you been drinking again with your utter personality? I just had three shots of espresso. Okay. Yeah. So two for you, one for the other, Brian? Yeah, makes sense.
Starting point is 00:00:50 Yes, one for the entity, known as Brian. Yeah, yeah. You know, it's, we all have, as of this recording, which the World Cup will be over, we all still have teams in play. I have. Shall we make a prediction, though? I think this would be something for the opening once we introduce our case. We should mention who we're playing.
Starting point is 00:01:12 Right, so U.S. is playing Turkey. Next. Andy, Austria is playing, what did you say? Algeria. We just played Argentina and lost against... We didn't lose against Argentina, we lost against Messi.
Starting point is 00:01:25 So that's the first statement to make. Yeah, and we'll just need to draw against Algeria to advance into the knockout phase. And Dan, what's Scotland playing against? Scotland's playing easy, Brazil. Okay. Yeah, that's an easy, easy, easy game in a group phase. What do you need to win to advance
Starting point is 00:01:44 So what do you need a draw? What do you need? I think a draw would do it, but we're aiming for the third place and see if we can get that way in game. I have no idea what we need. I think the US is already in the next round because you won both of your first two games. Okay, okay.
Starting point is 00:02:05 If I'm not mistaken. But hey, you know what? That actually already introduced kind of our guest, so at least we'll let him speak already. then Gomez Blanco, thank you so much for being again on the show. And also thanks for the reminder. The last time we spoke was three years ago on peer performance. Folks, we will also link to the episodes.
Starting point is 00:02:27 Back then it was called adopting open observability across your organization. And back then you were at a different company. You switched. You were at SkyScanter? I did. I was an end user back then. And since then, I've been, I've been, I guess, really interested in continuing to talk to end users. And, yeah, from SkyScanner, from leading observability there, and writing practical open telemetry,
Starting point is 00:02:54 I think that's the reason why we in the previous podcast. Yeah, so I now work for New Relic as a principal observability architect, and I do still work with end users to adopt best practices in observability. best practices in adopting open telemetry. And yeah, and I've been during all these years, part of the community, part of the hotel community, been part of the governance committee, and maintaining or co-maintaining the end user special interest group in hotel
Starting point is 00:03:29 with now a renewed focus on hotel blueprints, which is what we're talking about today. Today, yeah. Hey, one comment and one question. First of all, the comment, similar to what we did in the previous recording where we had Josh Lee on and also Adriana. It's great that the open telemetry community doesn't care about,
Starting point is 00:03:48 let's say, organizational or competitive challenges that we have, right? So we're working all in the same field. Even though we are competitors on paper, we are a community in real life and we work together and really want to make sure we help our end users. So that's great. The question that I have with the end users seek,
Starting point is 00:04:08 that means you're also working with Adriana because she's also in there. Correct. Yeah. Adriana, Rees and Andrey, both are co-mentainers of end-us of SIG. And, you know, what we're doing in that special interest group, so far has been work with, we have Open Telemetry Live as live sessions within, you know, their YouTube and LinkedIn live with end users that tell us about their adoption challenges, their learnings.
Starting point is 00:04:37 also with maintainers that basically bring up like news of what's happened in Open Telemetry. Sometimes people that, you know, employees from different vendors come and talk about, you know, some best practice, all in a completely vendor-neutral way. I think this is, you know, what you were saying before. As maintainers of an end-user special interest group, we always have to keep the bar, you know, high in terms of like non-sales pitchy stuff. Yeah, right. They're always keeping an vendor neutral.
Starting point is 00:05:08 Yeah, cool. And the reason why I think it was Adriana who brought it up, or maybe I also happened to then see your LinkedIn post or your blog posts around the new blueprints and also the reference implementations. And this is really why we're here today. Because in my world, also Brian, both of us, we work with end users, organizations that are trying to implement observability and whether they're using agents or Open Telemetry,
Starting point is 00:05:37 most of them are now moving into the world of Open Telemetry, they always ask, so how do I do this right? I know the documentation. I've played around with Astroshop, with the different demo apps. I understand how to collect the log. I understand how to set up a single collector and send it to any type of backend. But how do I basically get this into an enterprise scale?
Starting point is 00:05:59 And how do I get beyond the honeymoon phase of, I've done a PUC and a demo towards scaling this for real. So I would like to learn from you a background about these blueprints, about these reference implementations, and how we can also encourage more people to share their lessons, learn, and their learnings with you and your community. Yeah, so I think I should probably say first that blueprints or O'Tail blueprints is what's called an initiative or a project within open telemetry.
Starting point is 00:06:27 So this is something that allows us as a community to bring together, there are people from multiple special interest groups, multiple sex, multiple backgrounds. Now we even got end users involved in this initiative. So it allows us to basically publicize some work that we want the community to feedback into. So with Ojo Blueprints, this is something that's been in my mind for a long time to basically put together
Starting point is 00:06:58 some common patterns, common. design patterns and common challenges that companies and organizations find when they adopt hotel. And also, you know, a way for the community to talk about these best practices. I think it can be useful as well for maintainers to understand what are the friction points and then what we can improve as a project, right? So yeah, that was opened, that was approved as a project. We kicked it off maybe, I think, two or three months ago and then we've had people like co-leads there in that project like Lucas and Tiffany that work with me in the OTHO Blueprints and we've made making some really good progress.
Starting point is 00:07:42 I got a question for you and when in your day work, right, at New Relic and we in our day work at Dinotrace it is sometimes hard to get our users to publicly talk about how they're adopting certain things. Do you think this is easier with OpenTalermen? because it's open source, or do you still have to go through similar hoops to get these organizations to publicly say,
Starting point is 00:08:10 because you have three, right, you have Adobe, you have Mestodon, you have SkyScanner, I'm sure there's more coming. Is it easier or more challenging? It's easier. Because we are very clear as well
Starting point is 00:08:21 when we, you know, by the way, if you're an end user and you're listening to this, we are calling out for end users to give us, you know, to share the reference architectures. And then we don't, in fact, we don't want to talk about anything that relates to the back end. So if you use dinotrays or if you use neuralic or if you use whatever, Grafana, it's okay.
Starting point is 00:08:42 But we don't want to talk about that in the reference implementation. It's not applicable to the scope of a hotel, right? So what we're interested is in the open telemetry scope. And because it's all open source, it makes it a little bit easier to share. Another aspect is that sometimes people don't even want to talk about their infrastructure or how they deploy workloads. That's a bit more challenging. Some companies may have some type of policy internally that forbids them from talking about their infrastructure.
Starting point is 00:09:15 You can't really get around that. But reference architectures do give us a framework. And what we've heard from end users is that we're not like, you know, when we ask someone to come and share their reference architecture, we give them a template. And this template explains how to share it well, basically, how to tell a story. What do we want to learn from you? And that's really important for them as well, for some of them, to share to their whatever like compliance or legal team.
Starting point is 00:09:44 And say, look, this is the type of information that I will be sharing. And here's some previous examples, and that helped, right? Them internal as well. And that also means if the template, I also saw this, I think it's well documented and folks, we will be sharing all the relevant links. I think you have links
Starting point is 00:10:03 to GitHub issues or GitHub repositories where you can open up a new issue that is based on a blueprint template or reference implementation and then you just fill out all the relevant fields. For those people that are listening in and they are contemplating, but they say, well, I don't have a lot of time
Starting point is 00:10:21 and I don't know how to write, Is it part of your role and your sixth role also to then guide people through the actual writing process and then tidying it up and making it presentable? Yeah, in fact, when you open an issue, there is an issue template. And one of the questions we ask is, what level of involvement do you want to have in this? Either you want to write it all yourself or you can write part of it, but we can help you write that. It's good. reference architecture. Of course, if we would prefer, we'll basically have limited bandwidth.
Starting point is 00:10:55 You come and write it. We can't really say, yeah, we'll be there for everyone that wants to come and we'll write it for you, but we will prioritize that accordingly as well. Well, maybe a shout-up for people to use some of your tokens for... Actually, that's the way that, and I should give a shout-out to the developer expedient SIG, which they've done a lot of the work already. Without this sort of like umbrella of a project, they interviewed a bunch of like, well, SkyScanner, Mastodon, Adobe. They interviewed them and then they brought the reference architectures or the reference implementations
Starting point is 00:11:34 for them as blockposts. So when we were starting this project for OTO blueprints and reference architectures, we almost like didn't know that was going to, that was happening in the background. happens sometimes in hotel. It's just very big project, right? So you don't really know everything that's happening. And then we found out they were already about to publish these. So then we said, okay, we've got these reference implementations already. Let's give it a framework so that the next ones, you know, just we can give it, we can do it a bit more self-serve, but also, you know, make sure that we can help people in the process as well.
Starting point is 00:12:11 I think it might be a good idea too for this to just take a step back for listeners and explain what the blueprints are and what problem they're addressing, right? Because there's, you know, as we were talking yesterday, well, in the last episode, but yesterday for us, there was, towards the end of the podcast, we started talking about, you know, beyond the, the quote-unquote complexities, let's say, of deploying open telemetry to your system. Now it's like, how do you start managing and configuring your collectors and all that, right? So let's just take a step back and talk about, like, what is this blueprint addressing,
Starting point is 00:12:50 what kind of problems do people typically run into where these are going to be really useful? And by the way, I want to say the fact that these are being created now really speaks as like a testament to the maturity of open telemetry. Right? So it's fantastic to hear that this is going on because it's like, yeah, this is just chugging along happily and it's awesome to see that this is coming up.
Starting point is 00:13:17 So let's take a couple steps back and I think also address the fact that blueprints, or the term blueprints was something that was raised during the graduation process for Open Telemetry. Now we're all celebrating that hotel graduated as a CNCF project.
Starting point is 00:13:38 But part of that process involves going through a bunch of checks, right, in terms of contributor health and security and so on. One of them is about end user feedback. And the feedback that we got from end users during that creation process
Starting point is 00:13:55 was that, you know, they needed a way to start to think about O'Tail strategy or like the way to deploy O'Tail. And this is where like define deploy O'Tail. And this is where I started to think, like, you know, when people say, oh, you know, I deployed open telemetry
Starting point is 00:14:13 and I don't get much value from it, for example, I don't know, if someone makes that assessment, what do they mean by I deployed hotel? Did they just pick up a collector, put it into a host, got some infrastructure metrics, that could be deployed in hotel, or did they, you know, added the Java agent or the Java instrumentation agent to their workloads and they, you know, called it a day, that could be as well. So, like, the term, like, boy in hotel could mean different things to different people. So with Blueprints, what we wanted to do is take these reference implementations and the experience that we have by working with end users and extract these common patterns,
Starting point is 00:14:56 these common challenges that people are trying to solve in specific environments. So the challenges that are team that is operating in a Kubernetes cluster and wants to monitor the workloads on that cluster will be different from a team that is the cloud ops team that is in charge of the underlying infrastructure or an application team that wants to connect browser telemetry
Starting point is 00:15:22 to their back end and have a full story there. But there's one common theme for blueprints, which is that we want them to be something that connects multiple components. So there is a part of O'Tail that is complex
Starting point is 00:15:37 and that is And the thing I mentioned that in the original blog post introducing blueprints is when we talk about complexity, we have two types, right? And this is coming from that Fred Brooks paper from 1986, which is the year that was bought, by the way, which is no silver bullet. You have like, you know, why is stuff so complex?
Starting point is 00:16:03 And you have essential complexity and you have accidental complexity. An hotel has a lot of accidental complexity, but also some essential complexity. O'Tail is very broad in the way that it goes from, as I said, from client-side, browser mobile, to infrastructure, to serverless, and Kubernetes and hose monitoring. So all these things, if you deploy them in a way that is not aligned
Starting point is 00:16:33 or between these workloads, you could end up with the very problem that O'Tail was trying to solve, which is uncontextualized data, low quality data, low value or ROI being quite low in terms of the data that you produce. So that is the accidental part. So what we're trying to do with blueprints is like acknowledge the OTA can be complex and then think about how do we make it simpler to string together a strategy to do it. I need to take notes
Starting point is 00:17:10 because I think some of the things you just said, first of all, yes, they're written in the blog, but it also will make a good soundbite for promoting this blog post because as you said, people are adopting open telemetron in the end, they want to see the value out of it, and the
Starting point is 00:17:25 pure fact that we acknowledge that it solves a complex problem and therefore it needs guidance is good, right? And hopefully this will get more people to think about this. You mentioned that it's a different scenario when somebody tries to monitor their Kubernetes cluster
Starting point is 00:17:44 that they have under control or whether you're monitoring non-Cubernetes workloads. And I think actually the first blueprint that I've seen out there is exactly talking about how do you use Open Telemetry to monitor your infrastructure and your processes. Yeah, yeah, which some people think or have said already, like, oh, that's interesting because, like, you know, we think about CNCF, cloud native, Kubernetes,
Starting point is 00:18:08 hotel works really well there, but the first blueprint that we released was non-C Kubernetes, you know, bare metal and containerized workloads outside of Kubernetes. And that actually made me quite happy because that's probably one of the areas in hotel where, like, people are a bit less supported by the tooling. So if you run on Kubernetes, it's going to be a little bit easier to deploy something that, you know, you have that orchestration layer. that will make it easier to deploy a centralized strategy. But if you're like outside of that and then, you know,
Starting point is 00:18:47 in bare metal infrastructure and so on, it becomes a little bit more difficult. And yeah, quite happy to see as well the collaboration that happened there to, you know, basically describe how to work with Opamp as the protocol for management of agents. of agents, including collectors, but also the fact that we're also contributing or collaborating with some of the maintainers to say, are we ready to recommend this at scale? Are people using Opampa scale?
Starting point is 00:19:23 So there is a caveat in that saying, okay, you know, Opamp is being used in production. We recommend it as part of a blueprint. But the specification is in BTAs, so like, you know, things could change, right? but it's okay. Blueprints are not supposed to be the Open Telemetry specification. We will change advising a blueprint maybe in the future to say
Starting point is 00:19:48 maybe it's better to use a different way of doing this. But yeah, that's part of the evolution of a project, right? For me it's interesting and I've also had so many conversations where I said, do you use Open Telemetry? And then I said, why would we use Open Telemetry but not on Kubernetes?
Starting point is 00:20:06 And then I said, ah, right, because we are, I think, obviously, it was born in that CNCF cloud native bubble. And obviously, within our bubble, we know that everybody knows open telemetry. But it's also interesting that outside of that bubble, people think it's just something that is born here and is constrained to Kubernetes and cloud native. And I think it's good. This is also why I like the fact that the first blueprint is around this topic. actually it has been a long run in seg that is coming up with tooling and guidance for open telemetry and mainframes so yeah
Starting point is 00:20:43 yeah we've also I remember that also we have I think some of our team colleagues you know mainframe is a big topic for our customers that we all have had new relic and data dog and we and we had a mainframe agent for a long long time.
Starting point is 00:21:04 And we did, didn't we, Brian, we did a podcast on open telemetry for mainframe. I believe so. Sounds familiar. Yeah. So I remember being like, wait, you can use open telemetry on mainframe. It makes complete sense, but you don't think of it that way because you're like mainframe, open telemetry. There's too much of a gap.
Starting point is 00:21:22 But it's when you go back and think about what is doing, it's like, oh, yeah, it's just, you know, doing what our agent's doing. Of course it can do it, you know. And they may. I was going to say the amount of people that are now using it in IoT and like, you know, different ways, you know, like open telemetry is really everywhere, really. Yeah, it is. And this is also wanted to say the, you know, collectors at scale and open that scale,
Starting point is 00:21:51 bind plane, who we, I think, all know, they, I remember one of their blog posts, I think it's with Mercedes or one of their reference implementations, right, where they're using million or like a million of collectors that they're managing because I think they have collectors installed in every car or every truck or whatever it is, right? So this definitely works at scale. Yeah, it does, yeah. So blueprints, any other blueprints maybe been worked on? Yeah, so there's one that I think will be published by the time this particular episode is released.
Starting point is 00:22:32 which is the one that I'm personally authoring on managed telemetry platforms in Kubernetes environments. So we go for the, this is perhaps the most common, the most common patterns, the most common design patterns. And specifically, this is aimed at a more of a platform engineering approach to managing a centralized telemetry platform. And what we mean by platform, as case you'll know this well, Andy as well, is not just the infrastructure, right? I think this is one of the things that we need to think about
Starting point is 00:23:04 as not just about the collector architecture either, is how you make it easy for the rest of your organization and that will be people in charge of applications to emit telemetry, that's high quality. So it touches on things like the operator, the Kubernetes operator for open telemetry to automatically injects instrumentation into work. but also acknowledges that not everyone is able to run an operator and that may make more sense
Starting point is 00:23:39 to integrate with all the CI tooling, perhaps like base Docker images, or any way that you know one can control that base layer of configuration for O'Tail and then allow these teams to emit or to control the telemetry that is on top. But it also goes into, of course, some of the collector architecture and gateways and how to build reliable pipelines as well. So yeah, that's something that will be published soon, and I think it's one of perhaps the most common patterns that we've seen in the industry on that.
Starting point is 00:24:14 So basically providing observability as self-service kind of through, you know, providing, as I said, I think making observability easy, accessible and implementable for everybody. and the key thing here is that you're enforcing standards, right? Because I think the biggest thing, what we also discussed yesterday, is around data quality. It doesn't make sense if we just say we're collecting all the data we can, but we actually don't know which data we really need.
Starting point is 00:24:45 We had a quote yesterday, actually from Uma or from Autodesk, and I need to put it up again because I thought this was such a nice quote from him, where he talked about instruments by value, not by default. So instrument really what provides you the value that they need out of observability and don't just go with the default because the default gives you a lot of data but the question is if it gives you what you really need.
Starting point is 00:25:12 Yeah, and there's an aspect as well of cost and it's not just cost, right? When you think about high-quality telemetry if you're like inundated by low-quality telemetry that is also affecting the way that you reason about a system, right? You have a lot of noise.
Starting point is 00:25:31 It will become more difficult for you to reason about it, and it will become more costly for your agents to reason about it too, right? So it's about storing what really matters as well. And even, I think I spoke to Anne Curry that has a podcast on, but she was the, she's one of the authors of Building Green Software. It's a really good book, by the way, I recommend reading. And it's about, like, you know, storing what might be. matters as well. If you want to
Starting point is 00:26:00 sort of like be more carbon neutral or be more like efficient in the way that you store data or the way that you emit carbon emissions, there is an aspect of like storing data that is unrelated
Starting point is 00:26:18 to the compute resources or how you may how you may generate carbon emissions from like power in the devices but also the embodied carbon in the storage devices that are used for that data. So anyway, so what this says is, like, if you store less data and the data that matters to you, your carbon emissions will be reduced as well.
Starting point is 00:26:42 So it's a matter of that as well. Yeah. Yeah, we'll definitely. So he said, building green software? Yes. The book? Perfect, yeah. So folks who will add this to the show notes.
Starting point is 00:26:51 To try to get. Yeah, another guest, exactly. Please then do an introduction and we'll get in on the podcast. And if we are, to be honest, if we're talking about, I wanted to mention another book that I would recommend as part of this Blueprints things because the format that we, the structure of blueprints is based on an old book that's called Good Strategy, Bad Strategy by Richard Rumelt. And I think that's a book that changed my life a little bit a few years ago
Starting point is 00:27:22 because it basically deals with how one thinks or one thinks or one, should think about strategy and how to scope a strategy to solve certain problems. So the way that blueprints, and this is, you know, what people can expect from hotel blueprints is that the way that we scope the area that we'll be tackling with these recommendations and these design patterns is by looking at the challenges, at the common challenges that need to be solved. So we will not recommend anything that doesn't solve a real problem for people. And so the idea. is that you go through first a diagnosis phase where you list the common problems, you
Starting point is 00:28:04 then go through a phase where you add some guidelines or some recommendations, and then you go through a set of like implementation actions that implement those guidelines. So it's a very well structured sort of like problem guideline actions. And I think we think that this is great to build a story of a strategy and a blueprint that actually solved problems. It's not just, you know, this is some cool tool in that you should adopt. Right. Yeah. Do you know?
Starting point is 00:28:36 Go on. I was going to say this brings two thoughts to mind that I think are really important with these blueprints, right? In the previous episode, Andy, we were talking about, you know, a lot of times people will say, I want to do open telemetry because I want to, you know, the whole vendor lock-in thing, right? But that's the reason. There's no thought behind that reason. It's just I hear vendor neutral. I want to do that, right?
Starting point is 00:28:58 So two things that the blueprint is on that side, number one, just as we always say with like, oh, I want to move to serverless, I want to move to Kubernetes. Why? What are you going to get out of it? Right? If you have these blueprints, it's like, we're going to, we want to do this. And looking at what the blueprint does and how this does, I can translate that into all the benefits that I can get from it and all the advantage of doing that. but also, you know, we see a lot that customers are starting on their, as you mentioned earlier, Dan, they want to start moving over to O'Tel. Where do I start? How do I do it and all, right? And it becomes a resume builder to go to Open Telemetry.
Starting point is 00:29:38 Sometimes that's the only reason is like, look, I did. And I think, to me, the goal would be, let's make it so O'Tel is no longer a resume builder because it becomes so easy. You have exactly what you need to do. You have the best practices. It's like, oh, we want to do hotel? Great, we know. It's not a big thing anymore, right?
Starting point is 00:29:56 Like the dial-up connection versus the always on with the cable modem, right? And then that extends into the idea if there are good blueprints and good practices as you're going through and using different forms of AI to assist you in all these things. If there are best practices, if there are blueprints, it's easier to give the instruction based on that, then, hey, I want to do this. how do you want to do it? Well, I have now an exact, you know, a manifest of what I want to do based on best practices. That'll, again, make this all easier because the end of the day, like, yeah, we all work at vendors and we want people to buy our software. But the end of the day, our passion is really like, make sure you've got observability on. Make sure you have a really good performing website, right?
Starting point is 00:30:38 That's why we love what we do is because we want to see all that. And everybody, obviously with the open telemetry community, especially the people maintaining the O-Tel bit, right? it just makes that entry point much more purposeful and in the end easier so that we could get great performing software and we're not getting angry to say why is it taking so long yeah you know and then one of the things that we
Starting point is 00:31:06 that we have in a roadmap unofficial roadmap I guess you know we we have ideas what happens after we've released this new first blueprint is making the output of blueprint a little bit more structured for agents, right? At the moment, you can feed the whole blueprint and it might be like 4,000, 6,000 words that, you know, an agent will take it in his context. I think it's fine. However, if you were to structure it in a different way, for
Starting point is 00:31:36 example, like using skills or using other ways that, you know, that agentic workflows may use it in a more optimal way, and that would also help massively, I think. So from blue If you're going to lose skills. If you think about a skill, all these, I was mentioning there, the fact that blueprints are, you know, common challenge and here's a guideline, here is the step.
Starting point is 00:32:06 That's almost what a skill is, right? You tell an agent, every time you see this, which is your challenge, apply this guideline, and then maybe you give it a set of steps to apply. So, I don't know. I think we have a lot of years of strategic thinking and whatnot, many books written. But it's very simple at the end, right? It's like problem, guideline, set of steps to implement it.
Starting point is 00:32:31 And it just works. Are you also hearing a request for blueprints, or maybe it's a reference implementation, but how to gradually convert an existing observability implementation, whether it's homegrown or through agents over to open telemetry. Do you have any, because this came up in the conversation I recently had, so really specifically asking how do I convert if I have, let's say I have a tool X right now and I want to convert this over to open telemetry,
Starting point is 00:33:03 what are the things I need to look out for? That specifically hasn't come up yet. Well, what has come up during conversations and we were talking about there's a cubecon in, this is before the project properly started, but we had a session to almost like to bootstrap the
Starting point is 00:33:25 Blueprints Initiative at KubeCon, Amsterdam. So that was in February March, something like that. And this relates to that because we were talking about the we're taking blueprints and then we have
Starting point is 00:33:40 specific vendors or specific solutions or backends that may have their own standard, right? How do these two layer up? And I think, you know, one of the things that we would consider is something like the open telemetry demo has something similar, right? There is the open telemetry demo like bare, like the vanilla one, and then as a vendor you can have a fork of the open telemetry demo that you may apply your, you know, your specific settings on top. So I think Blueprints could have a similar sort of impact, right? where like you can take a blueprint and then you can say,
Starting point is 00:34:20 well, if you want to do this with backend X and if you're running already, because it doesn't really answer your question specifically, but if you were to be here in this particular proprietary format and then you want to move to O'Dell into this blueprint, maybe there is like an extension to that or something. Good idea. Hey, Andy, that makes me think too.
Starting point is 00:34:45 And I guess this would apply it to you as well, Dan, that, like, it seems like it would be beneficial for vendors to create blueprints for data enrichment for their backend tools, right? So if you think about, you know, broad data comes in and if you just leave it all raw without anything, you'll have traces, logs, and metrics, and they'll be separated. But, like, you know, for instance, if you're going to attach a log to a trace, you want that trace ID in the log or things like that. but also there's other metadata data, like I know in some Dinotrace process is like if we're using our agent and the collector to grab the
Starting point is 00:35:21 hotel data, there's some natural metadata enrichment we do before we send it back to us. But if you're doing just pure hotel play, like I think yeah, this is the blueprint on getting the most out of your back-in platform
Starting point is 00:35:38 with the hotel data. Here's where you'll want to do whatever enrichment if it's not something that could be done on. I mean, just just thinking like blueprints, it wouldn't be obviously something that the hotel community would do, right? I mean more like, you know, it'd probably be a really cool idea for vendors to come out with, like,
Starting point is 00:35:54 you know, here's best practices for hotel for our platform. Not the design side, but the data feeds on. Yeah, I think, you know, part of that is, you know, the work on distributions, right? If you have an hotel distribution for a particular back end, then that encodes some of that, some of those standards. But yeah, to your point,
Starting point is 00:36:18 I do think that sometimes it makes sense to extend that. It's not just about the binary that you deploy. It's not just about the config. It's like, I don't know, maybe all the things that go with that, and I can see value in that as well. And ideally, obviously, the answer of the hotel community
Starting point is 00:36:34 should be vendor X if you have something specific and it's not just valuable for you, but it would be something valuable for everybody, I not work with the O-TIL community and extend open telemetry itself, right? I mean, the advice that, hopefully, the advice that we put in blueprints is applicable. Right, exactly, yeah. Hey, Dan, so we talked about blueprints where one is out there, non-Cubonitis environments, you're working with one on managed telemetry platforms for Kubernetes workloads.
Starting point is 00:37:05 Now, can we talk a little bit about those reference implementations? How do, I mean, these are those that organizations, end users are contributing where they basically talk about the implementation. Can you give us an example of one of those that have been part of the initial publish and then what people can learn out of it? Yep. I mean, I'll probably choose. If I were to choose one and choose a Skyscarner one. But it makes me really happy that it wasn't neat. I did it.
Starting point is 00:37:38 It makes me happy that, you know, after I left and, you know, like, I was part of leading that strategy, but someone else takes their baton. And now it's part of the hotel community, right? So, yeah, I think that's one of the cases where O'Tail was used for, you know, to basically, and I think they mentioned that in that reference implementation, is that one of the reasons to adopt O'Tail
Starting point is 00:38:03 was to actually move to a vendor, right, while remaining vendor neutral in their implementation, And that's something that we've seen from many end users that they want to perhaps move towards like a platform engineering type of organization where they want to empower the rest of the organization to move faster, but that perhaps means that they don't run their own OS stack
Starting point is 00:38:31 and they can still run open source internally without the, you know, without the vendor locking of a particular solution and remain more aligned. with the rest of the cloud native environments or ecosystem. But they do that with hotel. So, yes, some of the patterns that were listed there, I think, but also listed in other reference implementations. Like, for example, Adobe were running a central gateway
Starting point is 00:39:08 for for telemetry so they have their individual teams that publish telemetry into a central gateway and two of these I think the
Starting point is 00:39:19 mastodon one was a bit different while they were running one collector deployment per main space while Adobe and SkyScanner were running that central gateway
Starting point is 00:39:35 that central collector gateway there were difference in them And all this is actually mentioned in the blueprint about to manage telemetry platforms. The difference then was like SkyScanner runs one single deployment for the collector gateway from metrics, traces, and logs,
Starting point is 00:39:50 with different configuration for memory limiter configurations in each of the pipelines. So if you want to start applying back pressure earlier on logs or on spans that you do on metrics, for example, while Adobe went for a solution
Starting point is 00:40:08 where they have different gateways, different gateway deployments per signal. So they have one gateway deployment for metrics, one gateway for traces, one gateway for logs. Both options are equally valid, right? Yeah. It depends on,
Starting point is 00:40:22 there's always going to be a trade-off between. Yeah, you basically introduce deployment complexity, but you may gain some stability and resiliency for certain things. Yeah. Resource allocation, maybe it's easier. So, yeah, I think both are valid. and it's great to see those actually being used at a scale of those companies.
Starting point is 00:40:45 Cool. So folks, if you're listening in and if you are the first time you hear about this, please do my favor. Go to the links, check out these Adobe Mestodon and Sky scanner reference implementations. Also check out the blueprints by the time when this episode airs. There will be at least the second blueprint that was just released by then. and yeah, contribute, right? Contributions, we are always looking for contributions
Starting point is 00:41:10 if you have implemented open telemetry at small scale or at big scale. I think there's no minimum no minimum requirement. It's good. Please, please share with us. Is there any... I have one more question.
Starting point is 00:41:31 This might be a little bit unrelated to just a blueprint discussion, but coming back to this vendor vendor neutrality. Obviously, you are also working for a vendor, New Relic, you also have your agents. Obviously, as the world is moving towards open telemetry,
Starting point is 00:41:49 we are setting, we are agreeing that the data that can come in, the thing that we can do within an application that we're observing is set by what open telemetry can do. Over the years, you guys at New Relic, the folks at Datadog and AppD and Instan and how they're all called and Dinah Trace.
Starting point is 00:42:09 We've all invested in agents and we've also built our own capabilities that are beyond what Open Telemetry currently provides. So in our case, I think of life debugging, right? That's one capability. I'm sure you have something similar.
Starting point is 00:42:25 How can we... Is there anything, is there any blueprint or best practice on how you can balance both kind of getting the openness and the Yeah, and the open is the standard of open telemetry while still getting some of these capabilities
Starting point is 00:42:42 that vendors have put in, and can this still coexist? Is there any blueprint best practice? That would be a best practice for vendors, and I think I've seen two approaches to this, and one is like you go down the route of a distribution, but you can, in an hotel distribution, and I should explain what that is, because sometimes we don't always get to know that.
Starting point is 00:43:06 is like a distribution is a repackaging of open telemetry components in a way that works the best with a particular backend. So that could be taking components in particular versions, a configuration that's applied in a particular way. There are certain requirements for the collector distributions, for example. If you have a collector distribution, you need to make it extensible. It needs to read open telemetry collector config.
Starting point is 00:43:36 like the scheme that needs to be, so there's some requirements for you to be able to call it a distribution of a collector, right? But in general, that is the case. And most importantly, this is why I wanted to mention this, you are allowed to extend that distribution with your own stuff.
Starting point is 00:43:55 So if you're a vendor, like, you're allowed to do that. The other option is that you take your, and I think, you know, some vendors that are like Dinotra. we do it as well at New Relic that is integrating or add an hotel support within the existing agents, right? Within the existing
Starting point is 00:44:15 so like that interoperability there. Which is, I think, I would hope that is like perhaps more and more of a, you know, hopefully in the future, less of instrumentation is added in a proprietary way. It's so like picked up not just from open telemetry.
Starting point is 00:44:35 And I think this is where like we're starting to see this change now. And this makes me really excited because you mentioned like there are three faces, right? One is we had all vendors do their own instrumentation. Second, we moved to open telemetry. Now open telemetry provides a lot of the instrumentation. And now what we're seeing is like open telemetry
Starting point is 00:44:56 is not able to, so let's say, you know, proprietary agents provided some of the instrumentation, but they were finding it hard to scale to the number of open source libraries that are there. Instrumenting everything, it's impossible. Then Open Telemetry came in, and you had like 14,000 developers each year contributing to O'Tail. So then you might be able to have higher velocity
Starting point is 00:45:21 in terms of integrating with a lot of different open source libraries out there. And if you don't find it, you're the owner of that library. You can contribute to O'Tail and added. However, what we're seeing now more and more, which is the ultimate vision of open telemetry, is that open source libraries themselves come instrumental with O'Tail. This is one of the key differences here
Starting point is 00:45:41 with proprietary agents is that the O'Tail API is completely decoupled from its implementation. So like we're seeing, you know, especially in the JNI space, lots of tooling out there that's already built in with O'Tail natively or that you can just enable O'Tail export.
Starting point is 00:45:58 For example, if you take Claudecode, you can just say, here's my OTOP endpoint. And code will just send that data to a central place, right? So, yeah, I think we're seeing more and more of that. And I think that is where like the next phase will be like Hotel, just maybe have the base instrumentation for the most common things. But more and more libraries are there, we'll start to use Hotel Native.
Starting point is 00:46:24 And also, I mean, we are both working for vendors. So we both know that all of us vendors over many, many years, we had to reverse engineer the same libraries is to figure out how to instrument it. And this is obviously now done in a more economical, sustainable, better way by the people that actually know that code and not somebody that has not written the code
Starting point is 00:46:45 but needs to figure out how to instrument it. And also like, you may release a new feature and then as a library owner, right? You want that feature to have like observability with your feature. Yeah. Hey, one last question because this came also up yesterday in the discussion with Adriana
Starting point is 00:47:03 and Josh, will the blueprints also cover practices, best practices for developers for instrumenting the code? Or will they just focus on the deployment of kind of the data pipeline? Because this was one of the things that they brought up is that there still needs to be education and tooling for developers to make sure their instrumentation provides value and it's not just collecting any type of data. Yeah, I think right now we are focused on the, not just infrastructure, but on the X as a service type of way. So you have either an observability team or a platform team that's in charge of helping to the rest of the organization to adopt hotel in these patterns. However, there was, when I was writing this particular blueprint for managed telemetry platforms, one of the aspects that we want to extend
Starting point is 00:48:00 is related to telemetry quality. And that starts to go into that space, right? How do you actually write good telemetry? And we know that Weaver, as a tool in open telemetry to manage semantic conventions registries and to measure if your data is, let's say, compliant with the semantic conventions that you have, that will be another blueprint that we'll want to do in the future.
Starting point is 00:48:29 And I think that moves closer to the developer surface. And I think in that one, I think it would probably make, you know, it would be a good idea. It'll make sense to cover some of these aspects about, like, you know, when should you use metrics or when should you use logs or spans or, you know. Yeah, and also detecting things like, you know, like Brian, you mentioned earlier, logs that are emitted by an app should always have a trace ID. or there should not be duplicated spains, Spain's without the red attributes and things like data.
Starting point is 00:49:03 So there's an aspect where, like, at the moment, we've only got two blueprints and they're like a little bit, I wouldn't say tangential, they're like parallel. There's a little bit of overlapping between some of the recommended patterns between non-Cubernetes and Kubernetes environments, but they're generally parallel. however those will have extension points right and this one for example for managed telemetry platforms
Starting point is 00:49:29 I'm calling out like please come and help us to write these for Weaver for example somatic conventions or for compliance and regulatory like requirements in terms of maybe you have like FIFA or FIPS or HIPPA or requirements that you need your collectors to comply with and then this is not covered in this blueprint, right? We just sort of like call that out of scope for this one.
Starting point is 00:49:57 But in a future blueprint, we can cover that. Cool. Then, thank you so much. Thank you. Coming on the show, back on the show, after a couple of years now, talking and enlightening us about blueprints and reference architectures, hopefully also encouraging many of our listeners to also think about contributing,
Starting point is 00:50:22 not just consuming but also contributing and we'll make sure to have all of the links in the podcast description in the notes Brian any final thoughts words from you? I just think it's an exciting time for O'Tel right besides being graduate or
Starting point is 00:50:38 whatever you call the official repository right these ideas like blueprints I think seeing what the agentic models are doing with open LLLometry built in we're
Starting point is 00:50:54 now starting to see because I remember here and I mentioned this on another podcast right when I first heard Open Telemetry there was this idea that like all these code vendors and stuff are going to start baking it in right haven't seen it now it's you know at least on the AI model side it's all baked in and it's amazing to see
Starting point is 00:51:10 you just again export it and all this stuff lights up so I think there's a lot of real world proof for people to say yeah this is real especially for the open source, you know, code developers to start adding it in, right? Because that's where we start getting all the promise.
Starting point is 00:51:30 So I think it's basically an inflection point right now, and it's an exciting one to see where this will continue to grow and what everyone can do with it. So Dan, thank you so much for being on and sharing the awesome information. Thanks for having me. It was a pleasure. All the best for Scotland, for the game.
Starting point is 00:51:51 Yes. All the best for Austria. Thank you. For the U.S. too. The U.S. Here we go. We've won World Cups, just not men. Thank you all. Thanks everyone for listening. Enjoy. Bye.
Starting point is 00:52:06 Bye.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.