PurePerformance - Blueprints for OTel Success: Standardizing Observability at Scale with Dan Gomez Blanco
Episode Date: July 20, 2026"There is no single way to deploy OpenTelemetry at scale—and that’s exactly the challenge."As organizations adopt OTel across teams and environments, they face tough questions around standardizati...on, configuration, and operating resilient observability pipelines.To address these challenges, the OpenTelemetry community has introduced Blueprints and Reference Implementations—practical guidance on topics like data standards, consistent agent and collector configuration, pipeline resilience, and intelligent sampling.In this episode, we’re joined by Dan Gomez Blanco, maintainer of the OpenTelemetry End-User SIG, to explore real-world reference architectures from organizations like Skyscanner, Adobe, and Mastodon.Tune in to learn how the community is turning OTel complexity into shared best practices—and how you can contribute your own blueprint
Transcript
Discussion (0)
It's time for pure performance.
Get your stopwatches ready.
It's time for Pure Performance with Andy Grabner and Brian Wilson.
Hello, everybody, and welcome to another episode of Pure Performance.
My name is Brian Wilson.
And as always, we have with me.
We have with me, that's awesome.
We have with me our co-hosts, Andy Grabner.
How are you doing today, Andy?
Good.
Have you been drinking again with your utter personality?
I just had three shots of espresso.
Okay.
Yeah.
So two for you, one for the other, Brian?
Yeah, makes sense.
Yes, one for the entity, known as Brian.
Yeah, yeah.
You know, it's, we all have, as of this recording, which the World Cup will be over,
we all still have teams in play.
I have.
Shall we make a prediction, though?
I think this would be something for the opening once we introduce our case.
We should mention who we're playing.
Right, so U.S. is playing Turkey.
Next.
Andy, Austria is playing,
what did you say?
Algeria.
We just played Argentina and lost against...
We didn't lose against Argentina,
we lost against Messi.
So that's the first statement to make.
Yeah, and we'll just need to draw against
Algeria to advance into the knockout phase.
And Dan, what's Scotland playing against?
Scotland's playing easy, Brazil.
Okay.
Yeah, that's an easy, easy, easy game in a group phase.
What do you need to win to advance
So what do you need a draw?
What do you need?
I think a draw would do it, but we're aiming for the third place
and see if we can get that way in game.
I have no idea what we need.
I think the US is already in the next round
because you won both of your first two games.
Okay, okay.
If I'm not mistaken.
But hey, you know what?
That actually already introduced kind of our guest,
so at least we'll let him speak already.
then Gomez Blanco, thank you so much for being again on the show.
And also thanks for the reminder.
The last time we spoke was three years ago on peer performance.
Folks, we will also link to the episodes.
Back then it was called adopting open observability across your organization.
And back then you were at a different company.
You switched.
You were at SkyScanter?
I did.
I was an end user back then.
And since then, I've been, I've been, I guess, really interested in continuing to talk to end users.
And, yeah, from SkyScanner, from leading observability there, and writing practical open telemetry,
I think that's the reason why we in the previous podcast.
Yeah, so I now work for New Relic as a principal observability architect,
and I do still work with end users to adopt best practices in observability.
best practices in adopting open telemetry.
And yeah, and I've been during all these years,
part of the community, part of the hotel community,
been part of the governance committee,
and maintaining or co-maintaining the end user special interest group in hotel
with now a renewed focus on hotel blueprints,
which is what we're talking about today.
Today, yeah.
Hey, one comment and one question.
First of all, the comment, similar to what we did in the previous recording
where we had Josh Lee on and also Adriana.
It's great that the open telemetry community
doesn't care about,
let's say, organizational or competitive challenges that we have, right?
So we're working all in the same field.
Even though we are competitors on paper,
we are a community in real life
and we work together and really want to make sure
we help our end users.
So that's great.
The question that I have with the end users seek,
that means you're also working with Adriana
because she's also in there.
Correct.
Yeah.
Adriana, Rees and Andrey, both are co-mentainers of end-us of SIG.
And, you know, what we're doing in that special interest group, so far has been work with,
we have Open Telemetry Live as live sessions within, you know, their YouTube and LinkedIn
live with end users that tell us about their adoption challenges, their learnings.
also with maintainers that basically bring up like news of what's happened in Open Telemetry.
Sometimes people that, you know, employees from different vendors come and talk about, you know,
some best practice, all in a completely vendor-neutral way.
I think this is, you know, what you were saying before.
As maintainers of an end-user special interest group, we always have to keep the bar, you know, high
in terms of like non-sales pitchy stuff.
Yeah, right.
They're always keeping an vendor neutral.
Yeah, cool.
And the reason why I think it was Adriana who brought it up,
or maybe I also happened to then see your LinkedIn post or your blog posts
around the new blueprints and also the reference implementations.
And this is really why we're here today.
Because in my world, also Brian, both of us, we work with end users,
organizations that are trying to implement observability
and whether they're using agents or Open Telemetry,
most of them are now moving into the world of Open Telemetry,
they always ask, so how do I do this right?
I know the documentation.
I've played around with Astroshop, with the different demo apps.
I understand how to collect the log.
I understand how to set up a single collector
and send it to any type of backend.
But how do I basically get this into an enterprise scale?
And how do I get beyond the honeymoon phase of,
I've done a PUC and a demo towards scaling this for real.
So I would like to learn from you a background about these blueprints,
about these reference implementations,
and how we can also encourage more people to share their lessons,
learn, and their learnings with you and your community.
Yeah, so I think I should probably say first that blueprints or O'Tail blueprints
is what's called an initiative or a project within open telemetry.
So this is something that allows us as a community to bring together,
there are people from multiple special interest groups,
multiple sex, multiple backgrounds.
Now we even got end users involved in this initiative.
So it allows us to basically publicize some work
that we want the community to feedback into.
So with Ojo Blueprints, this is something that's been
in my mind for a long time to basically put together
some common patterns, common.
design patterns and common challenges that companies and organizations find when they adopt
hotel. And also, you know, a way for the community to talk about these best practices.
I think it can be useful as well for maintainers to understand what are the friction points
and then what we can improve as a project, right? So yeah, that was opened, that was approved
as a project. We kicked it off maybe, I think, two or three months ago and then we've had
people like co-leads there in that project like Lucas and Tiffany that work with me in
the OTHO Blueprints and we've made making some really good progress.
I got a question for you and when in your day work, right, at New Relic and we in our day
work at Dinotrace it is sometimes hard to get our users to publicly talk about how they're
adopting certain things. Do you think this is easier with OpenTalermen?
because it's open source,
or do you still have to go through
similar hoops
to get these organizations
to publicly say,
because you have three, right,
you have Adobe,
you have Mestodon,
you have SkyScanner,
I'm sure there's more coming.
Is it easier or more challenging?
It's easier.
Because we are very clear as well
when we, you know,
by the way, if you're an end user
and you're listening to this,
we are calling out for end users
to give us, you know,
to share the reference architectures.
And then we don't, in fact, we don't want to talk about anything that relates to the back end.
So if you use dinotrays or if you use neuralic or if you use whatever, Grafana, it's okay.
But we don't want to talk about that in the reference implementation.
It's not applicable to the scope of a hotel, right?
So what we're interested is in the open telemetry scope.
And because it's all open source, it makes it a little bit easier to share.
Another aspect is that sometimes people don't even want to talk about their infrastructure
or how they deploy workloads.
That's a bit more challenging.
Some companies may have some type of policy internally that forbids them from talking about their infrastructure.
You can't really get around that.
But reference architectures do give us a framework.
And what we've heard from end users is that we're not like, you know,
when we ask someone to come and share their reference architecture, we give them a template.
And this template explains how to share it well, basically, how to tell a story.
What do we want to learn from you?
And that's really important for them as well, for some of them, to share to their whatever
like compliance or legal team.
And say, look, this is the type of information that I will be sharing.
And here's some previous examples, and that helped, right?
Them internal as well.
And that also means
if the template, I also saw this,
I think it's well documented
and folks, we will be sharing all the relevant links.
I think you have links
to GitHub issues or GitHub repositories
where you can open up a new issue
that is based on a blueprint template
or reference implementation
and then you just fill out all the relevant fields.
For those people that are listening in
and they are contemplating,
but they say, well, I don't have a lot of time
and I don't know how to write,
Is it part of your role and your sixth role also to then guide people through the actual writing process
and then tidying it up and making it presentable?
Yeah, in fact, when you open an issue, there is an issue template.
And one of the questions we ask is, what level of involvement do you want to have in this?
Either you want to write it all yourself or you can write part of it, but we can help you write that.
It's good.
reference architecture. Of course, if we would prefer, we'll basically have limited bandwidth.
You come and write it. We can't really say, yeah, we'll be there for everyone that wants
to come and we'll write it for you, but we will prioritize that accordingly as well.
Well, maybe a shout-up for people to use some of your tokens for...
Actually, that's the way that, and I should give a shout-out to the developer expedient
SIG, which they've done a lot of the work already.
Without this sort of like umbrella of a project, they interviewed a bunch of like, well, SkyScanner,
Mastodon, Adobe.
They interviewed them and then they brought the reference architectures or the reference implementations
for them as blockposts.
So when we were starting this project for OTO blueprints and reference architectures, we
almost like didn't know that was going to, that was happening in the background.
happens sometimes in hotel. It's just very big project, right? So you don't really know
everything that's happening. And then we found out they were already about to publish these.
So then we said, okay, we've got these reference implementations already. Let's give it a framework
so that the next ones, you know, just we can give it, we can do it a bit more self-serve,
but also, you know, make sure that we can help people in the process as well.
I think it might be a good idea too for this to just take a
step back for listeners and explain what the blueprints are and what problem they're addressing, right?
Because there's, you know, as we were talking yesterday, well, in the last episode, but yesterday for us,
there was, towards the end of the podcast, we started talking about, you know, beyond the, the quote-unquote complexities,
let's say, of deploying open telemetry to your system. Now it's like, how do you start managing
and configuring your collectors and all that, right?
So let's just take a step back and talk about,
like, what is this blueprint addressing,
what kind of problems do people typically run into
where these are going to be really useful?
And by the way, I want to say the fact that these are being created now
really speaks as like a testament to the maturity of open telemetry.
Right?
So it's fantastic to hear that this is going on
because it's like, yeah, this is just chugging along happily
and it's awesome to see that this is coming up.
So let's take a couple steps back
and I think also address the fact that
blueprints, or the term blueprints
was something that was raised
during the graduation process
for Open Telemetry.
Now we're all celebrating that hotel
graduated as a CNCF project.
But part of that process involves going through
a bunch of
checks, right, in terms of
contributor health and security
and so on. One of them
is about end user feedback.
And the feedback that we got from end users
during that creation process
was that, you know, they needed
a way to start
to think about O'Tail strategy
or like the way to
deploy O'Tail. And this is where
like define
deploy O'Tail. And this is where
I started to think, like, you know, when people say, oh, you know, I deployed open telemetry
and I don't get much value from it, for example, I don't know, if someone makes that assessment,
what do they mean by I deployed hotel? Did they just pick up a collector, put it into a host,
got some infrastructure metrics, that could be deployed in hotel, or did they, you know,
added the Java agent or the Java instrumentation agent to their workloads and they, you know,
called it a day, that could be as well. So, like, the term, like,
boy in hotel could mean different things to different people.
So with Blueprints, what we wanted to do is take these reference implementations and
the experience that we have by working with end users and extract these common patterns,
these common challenges that people are trying to solve in specific environments.
So the challenges that are team that is operating in a Kubernetes cluster and wants to
monitor the workloads on that cluster
will be different from a team
that is the cloud ops team that is in charge
of the underlying infrastructure
or an application team that wants to
connect browser telemetry
to their back end and have
a full story there.
But there's one common theme
for blueprints, which is that
we want them to be
something that connects multiple
components. So there is
a part of O'Tail that is complex
and that is
And the thing I mentioned that in the original blog post
introducing blueprints is when we talk about complexity,
we have two types, right?
And this is coming from that Fred Brooks paper from 1986,
which is the year that was bought, by the way,
which is no silver bullet.
You have like, you know, why is stuff so complex?
And you have essential complexity
and you have accidental complexity.
An hotel has a lot of accidental complexity,
but also some essential complexity.
O'Tail is very broad in the way that it goes from, as I said,
from client-side, browser mobile, to infrastructure, to serverless,
and Kubernetes and hose monitoring.
So all these things, if you deploy them in a way that is not aligned
or between these workloads,
you could end up with the very problem that O'Tail was trying to solve,
which is uncontextualized data, low quality data, low value or ROI being quite low
in terms of the data that you produce.
So that is the accidental part.
So what we're trying to do with blueprints is like acknowledge the OTA can be complex
and then think about how do we make it simpler to string together a strategy to do it.
I need to take notes
because I think some of the things
you just said, first of all, yes, they're
written in the blog, but it also will make
a good soundbite for
promoting this blog post because
as you said, people are adopting
open telemetron in the end, they want to see the value
out of it, and the
pure fact that we acknowledge
that it solves a complex problem
and therefore it needs guidance
is good, right?
And hopefully this will get more people
to think about this.
You mentioned that it's a different scenario
when somebody tries to monitor their Kubernetes cluster
that they have under control
or whether you're monitoring non-Cubernetes workloads.
And I think actually the first blueprint that I've seen out there
is exactly talking about how do you use Open Telemetry
to monitor your infrastructure and your processes.
Yeah, yeah, which some people think or have said already,
like, oh, that's interesting because, like, you know,
we think about CNCF, cloud native, Kubernetes,
hotel works really well there, but the first blueprint that we released was non-C Kubernetes,
you know, bare metal and containerized workloads outside of Kubernetes.
And that actually made me quite happy because that's probably one of the areas in hotel
where, like, people are a bit less supported by the tooling.
So if you run on Kubernetes, it's going to be a little bit easier to deploy something that,
you know, you have that orchestration layer.
that will make it easier to deploy a centralized strategy.
But if you're like outside of that and then, you know,
in bare metal infrastructure and so on,
it becomes a little bit more difficult.
And yeah, quite happy to see as well the collaboration that happened there to,
you know, basically describe how to work with Opamp
as the protocol for management of agents.
of agents, including collectors, but also the fact that we're also contributing or collaborating
with some of the maintainers to say, are we ready to recommend this at scale?
Are people using Opampa scale?
So there is a caveat in that saying, okay, you know, Opamp is being used in production.
We recommend it as part of a blueprint.
But the specification is in BTAs, so like, you know, things could change, right?
but it's okay.
Blueprints are not supposed to be
the Open Telemetry specification.
We will change advising a blueprint
maybe in the future to say
maybe it's better to use a different way of doing this.
But yeah, that's part of the evolution
of a project, right?
For me it's interesting
and I've also had so many conversations
where I said, do you use Open Telemetry?
And then I said, why would we use Open Telemetry
but not on Kubernetes?
And then I said, ah, right, because we are, I think, obviously, it was born in that CNCF cloud native bubble.
And obviously, within our bubble, we know that everybody knows open telemetry.
But it's also interesting that outside of that bubble, people think it's just something that is born here and is constrained to Kubernetes and cloud native.
And I think it's good.
This is also why I like the fact that the first blueprint is around this topic.
actually it has been a long run in seg
that is coming up with tooling and guidance for open telemetry
and mainframes so yeah
yeah we've also
I remember that also we have
I think some of our team
colleagues you know mainframe is a big topic for
our customers that we all have had new relic and
data dog and we
and we had a mainframe agent for a long long
time.
And we did, didn't we, Brian, we did a podcast on open telemetry for mainframe.
I believe so.
Sounds familiar.
Yeah.
So I remember being like, wait, you can use open telemetry on mainframe.
It makes complete sense, but you don't think of it that way because you're like
mainframe, open telemetry.
There's too much of a gap.
But it's when you go back and think about what is doing, it's like, oh, yeah, it's
just, you know, doing what our agent's doing.
Of course it can do it, you know.
And they may.
I was going to say the amount of people that are now using it in IoT and like, you know,
different ways, you know, like open telemetry is really everywhere, really.
Yeah, it is.
And this is also wanted to say the, you know, collectors at scale and open that scale,
bind plane, who we, I think, all know, they, I remember one of their blog posts,
I think it's with Mercedes or one of their reference implementations, right,
where they're using million or like a million of collectors that they're managing
because I think they have collectors installed in every car or every truck or whatever it is, right?
So this definitely works at scale.
Yeah, it does, yeah.
So blueprints, any other blueprints maybe been worked on?
Yeah, so there's one that I think will be published by the time this particular episode is released.
which is the one that I'm personally authoring on managed telemetry platforms in Kubernetes environments.
So we go for the, this is perhaps the most common, the most common patterns,
the most common design patterns.
And specifically, this is aimed at a more of a platform engineering approach
to managing a centralized telemetry platform.
And what we mean by platform, as case you'll know this well, Andy as well,
is not just the infrastructure, right?
I think this is one of the things that we need to think about
as not just about the collector architecture either,
is how you make it easy for the rest of your organization
and that will be people in charge of applications
to emit telemetry, that's high quality.
So it touches on things like the operator,
the Kubernetes operator for open telemetry
to automatically injects instrumentation into work.
but also acknowledges that not everyone is able to run an operator and that may make more sense
to integrate with all the CI tooling, perhaps like base Docker images, or any way that
you know one can control that base layer of configuration for O'Tail and then allow these teams
to emit or to control the telemetry that is on top.
But it also goes into, of course, some of the collector architecture and gateways and
how to build reliable pipelines as well.
So yeah, that's something that will be published soon,
and I think it's one of perhaps the most common patterns
that we've seen in the industry on that.
So basically providing observability as self-service kind of through,
you know, providing, as I said,
I think making observability easy, accessible and implementable for everybody.
and the key thing here is that you're enforcing standards, right?
Because I think the biggest thing, what we also discussed yesterday,
is around data quality.
It doesn't make sense if we just say we're collecting all the data we can,
but we actually don't know which data we really need.
We had a quote yesterday, actually from Uma or from Autodesk,
and I need to put it up again because I thought this was such a nice quote from him,
where he talked about instruments by value, not by default.
So instrument really what provides you the value
that they need out of observability
and don't just go with the default
because the default gives you a lot of data
but the question is if it gives you what you really need.
Yeah, and there's an aspect as well
of cost
and it's not just cost, right?
When you think about high-quality telemetry
if you're like inundated by low-quality telemetry
that is also affecting the way
that you reason about a system, right?
You have a lot of noise.
It will become more difficult for you to reason about it,
and it will become more costly for your agents to reason about it too, right?
So it's about storing what really matters as well.
And even, I think I spoke to Anne Curry that has a podcast on,
but she was the, she's one of the authors of Building Green Software.
It's a really good book, by the way, I recommend reading.
And it's about, like, you know, storing what might be.
matters as well. If you want to
sort of like
be more
carbon neutral or be more like
efficient in the way that you store data
or the way that you emit
carbon emissions, there is
an aspect of like storing data
that is unrelated
to the compute resources
or how you may
how you may generate
carbon emissions from like
power in the devices but also the
embodied carbon in the storage devices that are used for that data.
So anyway, so what this says is, like, if you store less data and the data that matters
to you, your carbon emissions will be reduced as well.
So it's a matter of that as well.
Yeah.
Yeah, we'll definitely.
So he said, building green software?
Yes.
The book?
Perfect, yeah.
So folks who will add this to the show notes.
To try to get.
Yeah, another guest, exactly.
Please then do an introduction and we'll get in on the podcast.
And if we are, to be honest, if we're talking about,
I wanted to mention another book that I would recommend as part of this Blueprints
things because the format that we, the structure of blueprints is based on an old book
that's called Good Strategy, Bad Strategy by Richard Rumelt.
And I think that's a book that changed my life a little bit a few years ago
because it basically deals with how one thinks or one thinks or one,
should think about strategy and how to scope a strategy to solve certain problems.
So the way that blueprints, and this is, you know, what people can expect from hotel
blueprints is that the way that we scope the area that we'll be tackling with these
recommendations and these design patterns is by looking at the challenges, at the common
challenges that need to be solved. So we will not recommend anything that doesn't solve a real
problem for people. And so the idea.
is that you go through first a diagnosis phase where you list the common problems, you
then go through a phase where you add some guidelines or some recommendations, and then you go
through a set of like implementation actions that implement those guidelines. So it's a very well
structured sort of like problem guideline actions. And I think we think that this is great
to build a story of a strategy and a blueprint that actually solved problems.
It's not just, you know, this is some cool tool in that you should adopt.
Right.
Yeah.
Do you know?
Go on.
I was going to say this brings two thoughts to mind that I think are really important with these blueprints, right?
In the previous episode, Andy, we were talking about, you know, a lot of times people will say,
I want to do open telemetry because I want to, you know, the whole vendor lock-in thing, right?
But that's the reason.
There's no thought behind that reason.
It's just I hear vendor neutral.
I want to do that, right?
So two things that the blueprint is on that side, number one, just as we always say with like, oh, I want to move to serverless, I want to move to Kubernetes.
Why? What are you going to get out of it? Right? If you have these blueprints, it's like, we're going to, we want to do this.
And looking at what the blueprint does and how this does, I can translate that into all the benefits that I can get from it and all the advantage of doing that.
but also, you know, we see a lot that customers are starting on their, as you mentioned earlier, Dan,
they want to start moving over to O'Tel.
Where do I start?
How do I do it and all, right?
And it becomes a resume builder to go to Open Telemetry.
Sometimes that's the only reason is like, look, I did.
And I think, to me, the goal would be, let's make it so O'Tel is no longer a resume builder
because it becomes so easy.
You have exactly what you need to do.
You have the best practices.
It's like, oh, we want to do hotel?
Great, we know.
It's not a big thing anymore, right?
Like the dial-up connection versus the always on with the cable modem, right?
And then that extends into the idea if there are good blueprints and good practices
as you're going through and using different forms of AI to assist you in all these things.
If there are best practices, if there are blueprints, it's easier to give the instruction based on that, then, hey, I want to do this.
how do you want to do it? Well, I have now an exact, you know, a manifest of what I want to do based on best practices.
That'll, again, make this all easier because the end of the day, like, yeah, we all work at vendors and we want people to buy our software.
But the end of the day, our passion is really like, make sure you've got observability on.
Make sure you have a really good performing website, right?
That's why we love what we do is because we want to see all that.
And everybody, obviously with the open telemetry community, especially the people maintaining the O-Tel bit, right?
it just makes that entry point
much more purposeful and in the end
easier so that we could get
great performing software and we're not
getting angry to say why is it taking so long
yeah you know and then one of the things that we
that we have in a roadmap
unofficial roadmap I guess you know we we have ideas
what happens after we've released this new
first blueprint is
making the output of
blueprint a little bit more structured for agents, right? At the moment, you can feed the whole
blueprint and it might be like 4,000, 6,000 words that, you know, an agent will take it in
his context. I think it's fine. However, if you were to structure it in a different way, for
example, like using skills or using other ways that, you know, that agentic workflows may use
it in a more optimal way, and that would also help massively, I think. So from blue
If you're going to lose skills.
If you think about a skill,
all these, I was mentioning there,
the fact that blueprints are, you know,
common challenge and here's a guideline,
here is the step.
That's almost what a skill is, right?
You tell an agent, every time you see this,
which is your challenge, apply this guideline,
and then maybe you give it a set of steps to apply.
So, I don't know.
I think we have a lot of years of strategic thinking and whatnot, many books written.
But it's very simple at the end, right?
It's like problem, guideline, set of steps to implement it.
And it just works.
Are you also hearing a request for blueprints,
or maybe it's a reference implementation,
but how to gradually convert an existing observability implementation,
whether it's homegrown or through agents over to open telemetry.
Do you have any, because this came up in the conversation I recently had,
so really specifically asking how do I convert if I have, let's say I have a tool X right now
and I want to convert this over to open telemetry,
what are the things I need to look out for?
That specifically hasn't come up yet.
Well, what has come up during conversations
and we were talking about there's a cubecon in,
this is before the project
properly started, but we had a
session to almost like to
bootstrap the
Blueprints Initiative at KubeCon, Amsterdam.
So that was in February
March, something like that.
And
this relates to that
because we were talking about the
we're taking blueprints
and then we have
specific vendors or specific
solutions or backends that may have
their own standard, right? How do these two layer up? And I think, you know, one of the things that
we would consider is something like the open telemetry demo has something similar, right? There
is the open telemetry demo like bare, like the vanilla one, and then as a vendor you can have a fork
of the open telemetry demo that you may apply your, you know, your specific settings on top.
So I think Blueprints could have a similar sort of impact, right?
where like you can take a blueprint and then you can say,
well, if you want to do this with backend X
and if you're running already,
because it doesn't really answer your question specifically,
but if you were to be here in this particular proprietary format
and then you want to move to O'Dell into this blueprint,
maybe there is like an extension to that or something.
Good idea.
Hey, Andy, that makes me think too.
And I guess this would apply it to you as well, Dan, that, like, it seems like it would be beneficial for vendors to create blueprints for data enrichment for their backend tools, right?
So if you think about, you know, broad data comes in and if you just leave it all raw without anything, you'll have traces, logs, and metrics, and they'll be separated.
But, like, you know, for instance, if you're going to attach a log to a trace, you want that trace ID in the log or things like that.
but also there's other metadata
data, like I know in some
Dinotrace process is like
if we're using our agent
and the collector to grab the
hotel data, there's some
natural metadata enrichment
we do before we send it back to us.
But if you're doing just pure
hotel play, like I think
yeah, this is the blueprint
on getting the
most out of your back-in platform
with the hotel data.
Here's where you'll want to do
whatever enrichment
if it's not something that could be done on.
I mean, just just thinking like blueprints,
it wouldn't be obviously something that the hotel community would do, right?
I mean more like, you know,
it'd probably be a really cool idea for vendors to come out with, like,
you know, here's best practices for hotel for our platform.
Not the design side, but the data feeds on.
Yeah, I think, you know, part of that is, you know,
the work on distributions, right?
If you have an hotel distribution for a particular back end,
then that encodes some of that,
some of those standards.
But yeah, to your point,
I do think that sometimes it makes sense to extend that.
It's not just about the binary that you deploy.
It's not just about the config.
It's like, I don't know,
maybe all the things that go with that,
and I can see value in that as well.
And ideally, obviously,
the answer of the hotel community
should be vendor X if you have something specific
and it's not just valuable for you,
but it would be something valuable for everybody,
I not work with the O-TIL community and extend open telemetry itself, right?
I mean, the advice that, hopefully, the advice that we put in blueprints is applicable.
Right, exactly, yeah.
Hey, Dan, so we talked about blueprints where one is out there,
non-Cubonitis environments, you're working with one on managed telemetry platforms for Kubernetes workloads.
Now, can we talk a little bit about those reference implementations?
How do, I mean, these are those that organizations, end users are contributing where they basically talk about the implementation.
Can you give us an example of one of those that have been part of the initial publish and then what people can learn out of it?
Yep.
I mean, I'll probably choose.
If I were to choose one and choose a Skyscarner one.
But it makes me really happy that it wasn't neat.
I did it.
It makes me happy that, you know, after I left and, you know, like,
I was part of leading that strategy,
but someone else takes their baton.
And now it's part of the hotel community, right?
So, yeah, I think that's one of the cases
where O'Tail was used for, you know, to basically,
and I think they mentioned that in that reference implementation,
is that one of the reasons to adopt O'Tail
was to actually move to a vendor, right,
while remaining vendor neutral in their implementation,
And that's something that we've seen from many end users
that they want to perhaps move towards
like a platform engineering type of organization
where they want to empower the rest of the organization
to move faster,
but that perhaps means that they don't run their own OS stack
and they can still run open source internally
without the, you know, without the vendor locking
of a particular solution and remain more aligned.
with the rest of the cloud native environments or ecosystem.
But they do that with hotel.
So, yes, some of the patterns that were listed there, I think,
but also listed in other reference implementations.
Like, for example, Adobe were running a central gateway
for
for telemetry
so they have their
individual teams
that publish telemetry
into a central gateway
and two of these
I think the
mastodon one was a bit different
while they were running
one collector
deployment per main space
while
Adobe and SkyScanner
were running
that central gateway
that central collector gateway
there were difference in them
And all this is actually mentioned in the blueprint
about to manage telemetry platforms.
The difference then was like SkyScanner
runs one single deployment
for the collector gateway from metrics,
traces, and logs,
with different configuration for
memory limiter
configurations in each of the pipelines.
So if you want to start
applying back pressure earlier
on logs or on spans
that you do on metrics, for example,
while Adobe went for a solution
where they have different gateways,
different gateway deployments per signal.
So they have one gateway deployment for metrics,
one gateway for traces,
one gateway for logs.
Both options are equally valid, right?
Yeah.
It depends on,
there's always going to be a trade-off between.
Yeah, you basically introduce deployment complexity,
but you may gain some stability and resiliency
for certain things.
Yeah.
Resource allocation, maybe it's easier.
So, yeah, I think both are valid.
and it's great to see those actually being used at a scale of those companies.
Cool.
So folks, if you're listening in and if you are the first time you hear about this,
please do my favor.
Go to the links, check out these Adobe Mestodon and Sky scanner reference implementations.
Also check out the blueprints by the time when this episode airs.
There will be at least the second blueprint that was just released by then.
and yeah, contribute, right?
Contributions, we are always looking for contributions
if you have implemented open telemetry
at small scale or at big scale.
I think there's no minimum
no minimum requirement.
It's good.
Please, please share with us.
Is there any...
I have one more question.
This might be a little bit unrelated
to just a blueprint discussion,
but coming back to this vendor
vendor neutrality.
Obviously, you are also working for a vendor,
New Relic, you also have your agents.
Obviously, as the world is moving
towards open telemetry,
we are setting,
we are agreeing that the data
that can come in, the thing that we can do
within an application that we're observing
is set by what open telemetry
can do. Over the years, you guys at New Relic,
the folks at Datadog and AppD
and Instan and how they're all called and Dinah Trace.
We've all invested in agents
and we've also built our own
capabilities
that are beyond what Open Telemetry currently
provides. So in our
case, I think of
life debugging, right? That's one capability.
I'm sure you have something similar.
How can we...
Is there anything,
is there any blueprint or best practice
on how you can balance both
kind of getting the
openness and the
Yeah, and the open is the standard of open telemetry
while still getting some of these capabilities
that vendors have put in, and can this still coexist?
Is there any blueprint best practice?
That would be a best practice for vendors,
and I think I've seen two approaches to this,
and one is like you go down the route of a distribution,
but you can, in an hotel distribution,
and I should explain what that is,
because sometimes we don't always get to know that.
is like a distribution is a repackaging of open telemetry components
in a way that works the best with a particular backend.
So that could be taking components in particular versions,
a configuration that's applied in a particular way.
There are certain requirements for the collector distributions, for example.
If you have a collector distribution,
you need to make it extensible.
It needs to read open telemetry collector config.
like the scheme that needs to be,
so there's some requirements for you to be able to call it
a distribution of a collector, right?
But in general, that is the case.
And most importantly,
this is why I wanted to mention this,
you are allowed to extend
that distribution with your own stuff.
So if you're a vendor, like, you're allowed to do that.
The other option is that you take your,
and I think, you know,
some vendors that are like Dinotra.
we do it as well at New Relic
that is integrating or add an
hotel support within the existing agents, right?
Within the existing
so like that interoperability there.
Which is, I think, I would hope that is
like perhaps more and more of a, you know,
hopefully in the future,
less of instrumentation
is added in a
proprietary way.
It's so like picked up not just from open telemetry.
And I think this is where
like we're starting to see this change now.
And this makes me really excited
because you mentioned like there are three faces, right?
One is we had all vendors do their own instrumentation.
Second, we moved to open telemetry.
Now open telemetry provides a lot of the instrumentation.
And now what we're seeing is like open telemetry
is not able to, so let's say, you know,
proprietary agents provided some of the instrumentation,
but they were finding it hard to scale
to the number of open source libraries that are there.
Instrumenting everything, it's impossible.
Then Open Telemetry came in,
and you had like 14,000 developers each year contributing to O'Tail.
So then you might be able to have higher velocity
in terms of integrating with a lot of different open source libraries out there.
And if you don't find it, you're the owner of that library.
You can contribute to O'Tail and added.
However, what we're seeing now more and more,
which is the ultimate vision of open telemetry,
is that open source libraries themselves
come instrumental with O'Tail.
This is one of the key differences here
with proprietary agents
is that the O'Tail API is completely decoupled
from its implementation.
So like we're seeing, you know,
especially in the JNI space,
lots of tooling out there
that's already built in with O'Tail natively
or that you can just enable O'Tail export.
For example, if you take Claudecode,
you can just say,
here's my OTOP endpoint.
And code will just send that data to a central place, right?
So, yeah, I think we're seeing more and more of that.
And I think that is where like the next phase will be like Hotel,
just maybe have the base instrumentation for the most common things.
But more and more libraries are there, we'll start to use Hotel Native.
And also, I mean, we are both working for vendors.
So we both know that all of us vendors over many, many years,
we had to reverse engineer the same libraries
is to figure out how to instrument it.
And this is obviously now done in a more
economical, sustainable, better way
by the people that actually know that code
and not somebody that has not written the code
but needs to figure out how to instrument it.
And also like, you may release a new feature
and then as a library owner, right?
You want that feature to have like observability
with your feature.
Yeah.
Hey, one last question because this came also up yesterday
in the discussion with Adriana
and Josh, will the blueprints also cover practices, best practices for developers for instrumenting the code?
Or will they just focus on the deployment of kind of the data pipeline?
Because this was one of the things that they brought up is that there still needs to be education and tooling for developers
to make sure their instrumentation provides value and it's not just collecting any type of data.
Yeah, I think right now we are focused on the, not just infrastructure, but on the X as a service type of way.
So you have either an observability team or a platform team that's in charge of helping to the rest of the organization to adopt hotel in these patterns.
However, there was, when I was writing this particular blueprint for managed telemetry platforms,
one of the aspects that we want to extend
is related to telemetry quality.
And that starts to go into that space, right?
How do you actually write good telemetry?
And we know that Weaver, as a tool in open telemetry
to manage semantic conventions registries
and to measure if your data is, let's say,
compliant with the semantic conventions that you have,
that will be another blueprint that we'll want to do in the future.
And I think that moves closer to the developer surface.
And I think in that one, I think it would probably make, you know, it would be a good idea.
It'll make sense to cover some of these aspects about, like, you know,
when should you use metrics or when should you use logs or spans or, you know.
Yeah, and also detecting things like, you know, like Brian, you mentioned earlier,
logs that are emitted by an app should always have a trace ID.
or there should not be duplicated spains,
Spain's without the red attributes and things like data.
So there's an aspect where, like, at the moment,
we've only got two blueprints and they're like a little bit,
I wouldn't say tangential, they're like parallel.
There's a little bit of overlapping between some of the recommended patterns
between non-Cubernetes and Kubernetes environments,
but they're generally parallel.
however those will have extension points right
and this one for example for managed telemetry platforms
I'm calling out like please come and help us to write these
for Weaver for example somatic conventions
or for compliance and regulatory
like requirements in terms of
maybe you have like FIFA or FIPS or HIPPA
or requirements that you need your collectors to comply with
and then this is not covered in this blueprint, right?
We just sort of like call that out of scope for this one.
But in a future blueprint, we can cover that.
Cool.
Then, thank you so much.
Thank you.
Coming on the show, back on the show, after a couple of years now,
talking and enlightening us about blueprints and reference architectures,
hopefully also encouraging many of our listeners
to also think about contributing,
not just consuming but also contributing
and we'll make sure to have all of the
links in the podcast
description in the notes
Brian any final thoughts
words from you? I just think it's an
exciting time for O'Tel right besides
being graduate or
whatever you call the official
repository right
these ideas like
blueprints I think
seeing what the
agentic
models are doing with
open LLLometry built in we're
now starting to see because I remember here and I
mentioned this on another podcast right
when I first heard Open Telemetry there was this
idea that like all these code vendors and stuff
are going to start baking it in right
haven't seen it now it's you know
at least on the AI model side it's all
baked in and it's amazing to see
you just again
export it and all this stuff lights
up so I think there's a lot of
real world proof
for people to say yeah this is real
especially for the open source, you know,
code developers to start adding it in, right?
Because that's where we start getting all the promise.
So I think it's basically an inflection point right now,
and it's an exciting one to see where this will continue to grow
and what everyone can do with it.
So Dan, thank you so much for being on
and sharing the awesome information.
Thanks for having me.
It was a pleasure.
All the best for Scotland, for the game.
Yes. All the best for Austria.
Thank you. For the U.S. too.
The U.S.
Here we go.
We've won World Cups, just not men.
Thank you all. Thanks everyone for listening.
Enjoy.
Bye.
Bye.
