Latent Space: The AI Engineer Podcast - AGI is Being Achieved Incrementally (DevDay Recap - cleaned audio)
Episode Date: November 8, 2023We left a high amount of background audio in the Devday podcast, which many of you loved, but we definitely understand that some of you may have had trouble with it. Listener Klaus Breyer ran it throu...gh Auphonic with speech islolation and we figured we’d upload it as a backdated pod for people who prefer this. Of course it means that our speakers sound out of place since they now sound like they are talking loudly in a quiet room. Let us know in the comments what you think?Timestampsthe cleaned part is only part 2:* [00:55:09] Part II: Spot Interviews* [00:55:59] Jim Fan (Nvidia) - High Level Takeaways* [01:05:19] Raza Habib (Humanloop) - Foundation Model Ops* [01:13:32] Surya Dantuluri (Stealth) - RIP Plugins* [01:20:53] Reid Robinson (Zapier) - AI Actions for GPTs* [01:30:45] Div Garg (MultiOn) - GPT4V for Agents* [01:36:42] Louis Knight-Webb (Bloop.ai) - AI Code Search* [01:48:36] Shreya Rajpal (Guardrails) - Guardrails for LLMs* [01:59:00] Alex Volkov (Weights & Biases, ThursdAI) - "Keeping AI Open"* [02:09:39] Rahul Sonwalkar (Julius AI) - Advice for Founders This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe
Transcript
Discussion (0)
Hey everyone, this is Swix coming at you live from the Newton, which is in the heart of the cerebral arena.
It is a new AI co-working space that I and a couple of friends are working out of.
There are hot desks available if you're interested.
Just check the show notes.
But otherwise, obviously, it's been 24 hours since the opening AI dev day.
A lot of hot reactions and long-standing tradition, one of the longest traditions we've had on the latent space pod is to convene emergency sessions and record the live.
thoughts of developers and founders going through and processing in real time. I think a lot of the
roles of podcasts isn't as perfect information delivery channels, but really as an audio and oral history
of what's going on as it happens while it happens. So this one's a little unusual. Previously,
we only just gathered on Twitter spaces and then just had a bunch of people. The last one was
the code interpreter one with 22,000 people showed up. But this one is a little bit more complicated
because there's an in-person element and then a online element.
So this is a two-part episode.
The first part is a recorded session between our Latent Space people and Simon
Willison and Alex Volker from the Thursday iPod, just kind of recapping the day.
But then also, as the second hour, I managed to get a bunch of interviews with previous
guests on the pod who we are still friends with and some new people that we haven't yet had on
the pod, but I wanted to just get their quick reactions because most of you have known and
I loved Jim Fan and Divgarg and a bunch of other folks that we interviewed.
So I just want to, I'm excited to introduce to you the broader scope of what it's like to be at OpenEye Dev Day in person, bring you the audio experience, as well as give you some of the thoughts that developers are having as they process the announcements from Open AI.
So first off, we have the Layton Space pod recap one hour of OpenAID DEVD.
Hey, everyone.
Welcome to the Latenspace podcast on Emergency Adepa.
addition after opening I'd have day.
This is Alessio, partner in CETN residents and decibel partners.
Then as usual, I'm joined by Spuix, founder of Smalley Eye.
Hey, and today we have two special guests with us covering all the latest and greatest.
We love to get our band together and recap things, especially when they're big.
And it seems like that every three months, we have to do this.
So Alex, welcome from Thursday, I.
We've been collaborating a lot on the Twitter spaces and welcome Simon from many, many things.
but also I think you're the first person to not make four appearances on our pod.
Oh, wow. I feel privileged.
So welcome. Yeah, I think we're all there yesterday. How do we feel? Like, what do you want to
kick off with? Maybe Simon, you want to take first and then Alex.
Sure. Yeah. I mean, yesterday was quite exhausting, quite frankly. I feel like it's going to take us
as a community several months just to completely absorb all of the stuff that they dropped on us in one
giant batch. It's particularly impressive considering they launched a ton of features. What,
three or four weeks ago, chat GPT voice and the combined mode and all of that kind of thing.
And then they followed up with everything from yesterday.
That said, now that I've started digging into the stuff that they released yesterday,
some of it is clearly in need of a bit more polish.
You know, the reality of what they released is, I'd say about 80% of what it looked like
it was yesterday, which is still impressive.
You know, don't get me wrong.
This is an amazing batch of stuff.
But there are definitely problems and sharp edges that we need to file off.
And there are things that we still need to figure out before we can take advantage of all of this.
Yeah, agreed, agreed.
And we can go into those trap edges in a bit.
I just want to pop over to Alex.
What are you with us?
So, interestingly, even folks at OpenEI, there's like several booths and a help desk.
So you can go in and ask people like actual changes and people like they could follow up with like the right people in OpenEI and like answer you back, et cetera.
Even some of them didn't know about all the changes.
So I went to the voice and audio booth and I asked them about like, hey, is Whisper 3 that was announced by Sam Outman?
stage just like briefly, will that be open source because I love using Whisper.
And there's like, oh, did we open source?
Do we talk about Whisper?
Like some of them didn't even know what they were releasing.
But overall, I felt it was a very tightly run event.
Like I was really impressed.
Sean, we were sitting in the audience and you like pointed at the clock to me when they
finished.
They finished like on 45 on that, I think, right?
And this was after like doing some extra stuff.
Very, very impressive for a first event.
Like I was absolutely like, good job, guys.
Good job.
Yeah, apparently it was their first keynote.
And someone, I think, was it you that told me that this is what happens if you have a president of White Combinator, do a proper keynote, you know, having seen many, many, many presentations by other startups.
This is sort of the sort of masterstroke.
Yeah, Alessio, I think you were watching remotely.
Yeah, we were at the Newton.
Yeah, I think we had 60 people here at the watch party.
So it was quite a big crowd.
It makes a reaction from different founders and people, depending on well,
always being announced on the page.
But I think everybody walked away,
kind of really happy with a new layer of interfaces,
big in use.
I think to me the biggest takeaway was like,
and I was talking with Mike Conover,
another friend of the podcast about this,
is they're kind of staying in the single-treaded,
like synchronous use cases lane, you know?
Like the GPD's announcement are all like still chat-based,
one-on-one synchronous things.
I was expecting maybe something about async things.
like background running agents, things like that.
But it's interesting to see there was nothing of that.
So I think if you're a founder in that space, you're quite excited.
You know, they seem to up to pick a product lane, at least for the next year.
So if you're working on async experiences, so things working in the background,
things that are not copilot like, I think you're quite excited to have them be a lot cheaper now.
Yeah.
As a person in itself, like I often think about this as a passing of a,
a big risk in terms of uncertainty over opening as roadmap.
Like, you know, they've shipped everything they're probably going to ship in the next six
months.
You know, they sort of marked out the territories that they're interested in.
And then so now that leaves open space for everyone else to pursue.
So I guess we can kind of go in order.
Probably top of men, top of mind to mention is the GPT4 turbo improvements.
So longer context length, cheaper price, anything else that stood out in your viewing of the
keynote and then just the commentary around it.
I was waiting for Stateful.
I remember they talked about Stateful API, the fact that you don't have to keep sending
like the same tokens back and forth just because, you know, and they're going to manage
the memory for you.
So I was waiting for that.
I knew it was coming at some point.
I was kind of, did not expect it to come kind of at this event.
I don't know why.
But when they announced Stateswell, I was like, okay, this is making it so much easier for
people to manage state.
The whole threads.
I don't want to mix between the two things, so maybe you guys can clarify, but there's the
GPD4-2B, which is the model that has new capabilities, a whopping 128K, like, a context link,
right? It's huge. It's like two and a half books, but also, you know, faster, cheaper, etc.
I haven't yet tested the fasterness, but, like, everybody's excited about that.
However, they also announced this new API thing, which is the assistance API, and part of it
is threads, which is will manage the thread for you. I can't imagine, like, I can't imagine how many
times I had to like re-implearned this myself in different languages and type script and Python,
etc. And now it's like, it's so easy. You have this one thread. You're sending it to a user and you
just keep sending messages there and that's it. The very interesting thing that we attended and by
we, I mean, like SWIGS and I have a live space and like 200 people. So it's like me, SWIX and 200 people
with us as well, they kept asking like, well, how's the price happening? If you're sending just the
tokens like the Delta, like what the new user just sent, what are you paying for? And I went
to open AI people and I was like, hey, how do we get paid for this?
And nobody knew, nobody knew.
I finally got an answer.
You still pay for the whole context that you have inside the thread.
They still pay for all this.
But now it's a little bit more complex for you to kind of count with TikTok.
So you have to hit another API endpoint to get the whole thread of what the context is.
Then TikTokanize this, run this to TikTok, and then calculate.
This is now the new way officially from Open the Eye.
But I really did have to go and find this.
They didn't know a lot of how the pricing is going to have.
Ouch.
Yeah.
Yeah.
Does the API at least tell you how many tokens you used,
or is it entirely up to you to do the accounting?
Because that would be a real pain if you have to account for everything.
So in my head, the question I was asking is, like,
if you want to know in advance before hitting the API,
like with the library hook token,
if you want to count in advance or like make a decision like advanced on that,
how would you do this now?
And they said, well, yeah, there's a way.
If you hit the API, get the whole thread back,
then count the tokens.
But I think the API still really sends you back the number of tokens.
But isn't there a feature of this new API where they actually do, they claim it has,
does it have infinite length threads because it's doing some form of condensation or summarization
of your previous conversation for you?
I heard that from somewhere, but I haven't confirmed it yet.
So I have a source from Dave Waldman.
I actually don't know what his affiliation is, but he usually has pretty accurate takes on AI.
So I think he works in AI circles in some capacity.
So I'll feature this in the show notes, but he said,
some not mention interesting bits from opening eye dev day.
One unlimited context window and chat threads from opening eye docs.
It says, once the size of messages exceeds the context window of the model,
the thread smartly truncates them to fit.
I'm not sure I want that intelligence.
I want to chime in here.
Just real quick, the not want this intelligence.
I heard this for multiple people over the next conversation that I had.
Some people said, hey, even though they're giving us like a content understanding and rag, we are doing different things.
Some people said this with vision as well.
And so that's an interesting point that like people who did implement custom stuff, they would like to continue keeping implementing custom stuff.
That's also like an additional point that I've heard to talk about.
Yeah.
So what opening us doing is providing good defaults and then, well, good is questionable.
We'll talk about that.
I think the existing sort of like chain and llama indexes of the world are not very threatened by this because
there's a lot more customization that they want to offer.
So, frustration is that OpenAI, they're providing new defaults, but they're not documented
defaults.
Like they haven't told us how their rag implementation works.
Like, how are they chunking the documents?
How are they doing retrieval?
Which means we can't use it as software engineers because it's this weird thing that we don't
understand.
And there's no reason not to tell us that.
Giving us that information helps us decide how to write good software on top of it.
So that's kind of frustrating.
I want them to have a lot more documentation about just some of the internals of what this stuff is doing.
I want to highlight an additional capability that we got, which is document parsing via the API.
I was blown away by this, right?
So we know that you could upload the images and vision API we got, we could talk about vision as well.
But just the whole fact that they presented on stage, like the document parsing thing,
where you can upload PDFs of the United Flight and then they upload like an Airbnb.
That on the whole, like that's a whole category of like products that's now.
open to open eyes, just like given developers to very easily build products that previously
it was a pain in about for many, many people.
How do you even like parse a PDF?
Then after you parse it, like what do you extract?
So the smart extraction of like document parsing, I was really impressed with it.
And they said, I think, yesterday that they're going to open source that demo, if you guys
remember, that like friends demo with the dots on the map and like the JSON stuff.
So it looks like that's going to come to open source.
And many people learn new capabilities for document parsing.
So I want to make sure we're very clear what we're
talking about when we talk about API, when you say API, there's no actual endpoint that does this, right?
You're talking about the chat GPT's functionality.
No, I'm talking about the assistance API, the assistant API that has threads now, that has
agents and you can run those agents.
Actually, maybe let's clarify for this point.
I think I had to, somebody had to clarify for this for me.
There's the GPTs, which is a UI version of running agents.
We can talk about them later, but like you and I and my mom can go and like, hey, create a new
GPT that like, you know, only does Technoric, jokes, like, whatever.
But there's the assistance thing, which is kind of a similar thing, but not the same.
So you can't create, you cannot create an assistant via an API and have it pop up on the
marketplace on the future marketplace they announced.
Oh, can you not?
No, no, no, not via the API.
So there are like two separate things and somebody in the opinion I told me they're not,
they're not exactly the same.
That's so confusing because the AI looks exactly like the UI that you used to set up the GPTs.
I assumed there was an API for the same feature.
And the playground, actually, if you go to the playground,
it kind of looks the same.
There's like the configurable thing.
The configure screen also has like you can allow it browsing,
you can allow it like tools.
But somebody told me they didn't do the full cross-mapping.
So like you won't be able to create GPPs with API.
You will be able to create assistants.
And then you'll be able to have those assistants do different things,
including call your external stuff.
So that was pretty cool.
Okay.
So this API is called the assistant API.
That's what we get like in addition to the model
of the GPT4 Turbo.
And that has document parsing.
So you can upload documents there,
and it will understand the context of them
and then return you like structured or unstructured input.
I thought that that feature was like phenomenal
just by all its own.
Like just on its own,
uploading a document, a PDF, a long one
and getting like structured data out of it,
it's like a pain in you have to build.
Let's face it, guys.
Like everybody who built this before,
it's like it's kind of horrible.
When you say structured data,
are you talking about the citations?
The JSON output,
the new JSON output, the new J.
Jason output that they also gave us.
Finally, if you guys remember last time
we talked together, I think it was like
during the functions release, EmergencyPod.
And back then, their answer to like, hey, everybody
wants structured data was, hey, we're going to give
you a function calling. And now
they did both. They gave us both like
a JSON output like structure. So like you can,
the models are actually going to return JSON.
Haven't played with it myself, but that's what they announced.
And the second thing is they
improved the function calling significantly
as well.
So, so
I talk to a staff member there.
I've got a pretty good model for what this is.
Effectively, the JSON thing is they're doing the same kind of trick as Lama-grammas and JSON-forma.
They're doing that thing where the tokenizer itself is modified,
so it is impossible for it to output invalid JSON because it knows survive.
Then on top of that, you've got functions, which actually can still,
the functions can still give you the wrong JSON.
They can give you JSON with keys that you didn't ask for if you're unlucky,
but at least it will be valid.
At least it will pass through a JSON parser.
And so they're very similar sort of things, but they're slightly different in terms of what they actually mean.
And yeah, the new function stuff is super exciting because functions are one of the most powerful aspects of the API.
But a lot of people haven't really started using yet.
But it's amazingly powerful what you can do with it.
I saw that the functions, the functionality that they now have is also plug inable as actions to those.
Right.
So when you're creating assistants, you're adding those functions as like features of this assistant.
and then those functions will execute in your environment,
but they'll be able to call different things.
Like they showcase an example of integration with, I think, Spotify or something, right?
And that was like an internal function that ran.
But it is confusing the kind of the online assistant, APIable agents,
and the GPT's agents.
So I think it's a little confusing because they demo both.
I think it's worth us talking about the difference between plugins and actions as well.
Because, you know, they launched plugins back in February.
And they've effectively, they've kind of deprecated plugins.
They haven't said it out loud, but it's clear that they are not going to be investing further in plugins,
because the new actions thing is covering the same space.
But actually, I think, is a better design for it.
Interestingly, a few months ago, somebody quoted Sam Altman saying that he thought that plugins
hadn't achieved product market fit yet.
And I feel like that's sort of what we're seeing today.
The problem with plugins is it was all a little bit messy.
People would pick and mix the plugins that they needed.
Nobody really knew which plugin combinations would work.
With this new thing, instead of plugins, you build an assistant,
and the insistent is a combination of a system prompt
and a set of actions which look very much like plugins.
You know, they get a JSON schema to call an API somewhere.
And I think that makes a lot more sense.
You can say, okay, my product is this chatbot with this system prompt,
so it knows how to use this tools.
I've given it this combination of plug-in-like things that it can use.
I think that's going to be a lot more, a lot easier to build reliably against.
and I think it's going to make a lot more sense to people
than the sort of mix and match mechanisms they had previously.
So actually, maybe it will be cool to cover kind of the capabilities of an assistant, right?
So you have a custom prompt, which is akin to the system message.
You have the actions thing, which is you're going to have add the existing actions,
which is like browse the web and code interpreter, which we should talk about,
like the assistants now can write code and execute it, which is exciting.
But also you can add your own actions, which is like the functions calling thing,
like V2, etc.
Then I heard this incredibly quick thing that somebody told me that you can add two assistants to a thread.
So you literally can mix agents within one thread with the user.
So you have one user and then you can have like this assistant and that assistant, they just glanced over this.
I was like, that is very interesting.
That is not very interesting.
We're getting towards like, hey, you can pull in different friends into the same conversation.
Everybody does the different thing.
What other capabilities do we have there?
Do you guys remember?
Oh, like context, uploading it with context, like with the full, here's our API documentation.
Well, that one's a bit more complicated.
So you've got the system prompt, you've got optional actions, you've got, you can turn on Dali-free, you can turn on code interpret, you can turn on Gras with Bing.
Those can be added or removed from your assistant.
And then you can upload files into it, and the files can be used in two different ways.
There's this thing that they call, I think they call it the retriever, which basically does, it does rag.
it does retrieval augmented generation against the content you've uploaded.
But code interpreter also has access to the files that you've uploaded.
And those are both in the same bucket.
So you can upload a PDF to it.
And on the one hand, it's got the ability to turn that into, like chunk it up, turn it into vectors,
use it to help answer questions.
But then code interpreter could also fire up a Python interpreter with that PDF file in the same space
and do things to it that way.
And it's kind of weird that they chose to combine both of those things.
Also, the limits were amazing, right?
you get up to 20 files, which is a bit weird because it means you have to combine your documentation to a single file.
But each file can be 512 megabytes.
They're giving us 10 gigabytes of space in each of these assistants, which is vast, right?
Of course, I tested it'll handle SQLite databases.
You can give it a gigabyte 12 megabyte SQLite database, and it can answer questions based on that.
But yeah, it's, like I said, it's going to take us months to figure out all of the combinations that we can build with all of this.
I was going to say for the storage.
I saw Jeremy Howard tweeted about it's like 20 cents per gigabyte per assistant per day.
Just to compare like S3 cost like two cents per month per gigabyte.
So like 300X more, something like that than just raw S3 storage.
Ouch.
There will still be a case for like maybe roll your own rag,
depending on how much information you want to put there.
but I'm curious to see what the price, the client curve looks like for the storage there.
Yeah, they probably should just charge that at cost.
There's no reason for them to charge so much.
That is wildly expensive.
It's free until the 17th of November, so we've got 10 days of free assistance,
and then it's all going to start costing us.
Crikey.
They gave us 500 bucks of API credit at the conference as well,
which we'll burn through pretty quickly at this rate.
I confirmed a very important.
important question, everybody was asking. Did the five people who got the $500 first got
actually $1,000? And I think somebody in OpenNine, I said, yes, there was nothing there that
prevented the five first people to not receive the second one again.
Okay. I met one of them. I met one of them. He said he only got 500.
Ah, interesting. Okay. So again, even Open AI, people will not necessarily know what happened
on stage with Open the Eye. Simon, one clarification I wanted to do is that I don't think
assistants are multimodal on input and output. So you do have vision, I believe, not confirmed by
do believe that you have vision, but I don't think that Dali is an option for
an option for GPs, but the guys...
Oh, that's so confusing.
The assistants, the checkbox for Dali is not there.
You cannot enable it.
Well, you just add them as a tool, right?
So, like, it's just one more...
It's a little finicky.
In the GPD interface.
Yeah.
I mean, to be honest, if assistants don't have Dali 3, we...
Does Dali 3 have an API now?
I think they released one.
I can't...
There's so much stuff.
That got lost in the pile.
But yeah, so code interpreter, wow.
That I was not expecting.
That's huge.
Assuming, I mean, I haven't tried it yet.
I need to confirm that it definitely works,
because GPT is going to have code interpreter.
Can assistance?
Assistance will have code interpreter as well, yeah.
Awesome.
Huh.
It's incredible.
Yeah, so I tried to make it do things that were not logical yesterday.
because one of the risks of having the God model is it calls the wrong model inappropriately
whenever you try to ask it to something that's kind of vaguely ambiguous.
But I thought it handled the job decently well.
I think there's still going to be rough edges.
It's going to try to draw things.
It's going to try to code when you don't actually want to.
And in a sense, opening eye, it's kind of removing that capability from chaty-b-t.
It just wants you to always query the God model and always get feedback.
on whether or not that was the right thing to do.
Which really sucks because it runs, I like asking a question and it goes, oh, searching Bing.
And I'm like, no, don't search Bing.
I know that the first 10 results on Bing will not solve this question.
I know you know the answer.
So I had to build my own custom GPT that just turns off Bing because I was getting frustrated
with it always going to Bing when I didn't want it to.
Okay.
So this is a topic that we discussed, which is the UI changes to chat GPT.
So we're moving on from the Assistance API.
and talking just about the upgrades to chat GBT and maybe the GBT store.
You did not like it.
And I love this.
I mean, both sides of this.
Okay.
So my problem with it, I've got the two things I don't like.
Firstly, it can do Bing when I don't want it to.
And that's just, just irritating because the reason I'm using GPT to answer a question
is that I know that I can't do a Google search for it because I've got a pretty good feeling
for what's going to work and what isn't.
And then the other thing that's annoying is it's just a little thing.
interpreter doesn't show you the code that it's running as it's typing it out now.
Like it'll churn away for a while doing something and then they'll give you an answer and you have to click a tiny little icon that shows you the code.
Whereas previously you'd see it writing the code so you could cancel it halfway through if it was getting it wrong.
And okay, I'm a Python programmer so I care and most people don't, but that's been a bit annoying.
Yeah, and when it errors, it doesn't tell you what the error is.
It just says analysis failed and it tries again.
But it's really hard for us to help it.
Yeah.
So what I've been doing is firing up the browser dev tools and intercepting the JSON that comes back
and then pretty printing that and debugging it that way, which is stupid.
Like, why do I have to do that?
It's really good feedback for OpenEI.
I will tell you guys what I loved about this unified mode.
I have a name for it.
So we actually got a preview of this on Sunday.
And one of the folks got like an early example of this.
I call it MMIO, multimodal input and output, because now there's a shared context.
between all of these tools together.
I think it's not only about selecting them,
just selecting them.
And Sam Altman on stage I said,
oh yeah, we unified it for you,
so you don't have to call different modes at once.
And in my head, that's not all they did.
They gave a shared context.
So what is an example of shared context, for example?
You can upload an image using GPT for vision and eyes,
and then this model understands what you kind of uploaded vision-wise.
Then you can ask Dali to draw that thing.
So there's no text shared in between those modes now.
There's like only visual shared between
in those modes and Dali will generate whatever you upload it in an image.
So it's eyes to output visually.
And you can mix the things as well.
So one of the things we did is, hey, use real-world real-time data from Bing, like weather,
for example.
Weather changes all the time.
And we asked Dali to generate like an image based on weather data in a city.
And it actually generated like a live almost like, you know, like snow, whatever.
It was a snowing inventor.
And that I think was like pretty amazing in terms of like being able to share contacts
between all these different models and modalities in the same understanding.
And I think we haven't seen the end of this.
I think like generating personal images, adding context to Dali, like all these things
are going to be very incredible in this one mode.
I think it's very, very powerful.
I think that's really cool.
I just want to opt in as opposed to opt out.
Like I want to control when I'm using the gold model versus when I know, which I can do
because I created myself a custom GPT that does what I need.
It just felt a bit silly that I had to do a hot.
custom bot just to make it not do Bing searches.
All solvable problems in the fullness of time.
Yeah.
But I think people, it seems like for the chat GPT at least,
they're really going after the broadest market possible.
That means simplicity comes at a premium at the expense of pro users.
And the rest of us can build our own GPT rappers anyway.
So not that big of a deal.
But maybe do you guys have a, oh, sorry.
So the GPT rappers thing, guys,
They call them GPTs because everybody's building GPDs.
Like, all the rappers, whatever, they end with the word GPT.
And so I think they reclaimed it.
That's like, you know, instead of fighting and saying, hey, you cannot use the GPD.
GPD.
It's like, we have GPTs now.
This is our marketplace.
Whatever everybody else builds, we have the marketplace.
This is our thing.
I think they did like a whole marketing move here.
It's a very strong marketing moves because now it's called Canva GPD.
It's called Zapier GPT.
And they're basically saying don't build your own websites, build it inside of our
God app with chat GPT and that's the way that we want you to do that.
In a way, it sort of makes up, it sort of makes up the fact that chat GPT is such a terrible
name for a product, right? Chat GPT, what were they thinking when they came up with that name?
But I guess if they lean into it, it makes a little bit more sense.
It's like chat GPT is the way you chat with our GPs and GPT is a better brand.
It's terrible, but it's not, it's a better brand than chat GPT was.
So talking about naming, yeah.
Yeah, so Simon actually, so for those listeners, we're actually going to release Simon's talk at the AI Engineer Summit where he actually proposed, you know, a better name for the sort of junior developer or code developer.
Coding, I'm coding intern.
Coding intern.
Yeah, coding intern was it, yeah.
But did you know, did you notice that advanced data analysis is dead, you know, 2023 to 2023, you know, a sales driven decision that has been rolled back effectively because now everything just called coding.
Oh, that's, I hadn't noticed that I thought they'd split the brands.
And they're saying advanced age analysis is the user-facing brand and code is the developer-facing brand.
But now have they ditched that from the interface then?
Yeah.
So it's unified mode, yeah.
Yeah.
So like in the unified mode, there's no selection anymore, right?
You just get all tools at once.
So there's no reason to differentiate this.
But also in the pop-up, when you log in, when you log in, it just says code interpreter as well.
So, yeah.
And then also when you make a GPT, the drop down when you create your own GPT,
it just says, call interpreter.
It also doesn't say it.
You're right.
Yeah, they ditch the brands.
Good Lord.
On the UI.
That's amazing.
Okay.
Well, you know, I think, so I may be one of the few people who listen to AI podcasts
and also Sester podcast.
And so I heard the full story from the opening ice head of sales about why it was named
Advanced Data Analysis.
I saw that.
Yeah.
Yeah.
There's a bit of civil resistors, I think, from the engineers in the room.
It feels like the engineers won because we got code interpreter back,
and I know for sure that some people were very happy with this specific thing.
I'm just glad.
For the past couple of months, I've been writing code interpreter parentheses,
also known as advanced data analysis.
And now I don't have to anymore, so that's great.
Yeah, yeah, let's back.
Yeah, I did want to talk a little bit about the GPT creation process, right?
I've been basically banging the drama a little bit about how AI is a better prompt engineer than you are.
And sorry, am I speaking over assignment because I'm lagging.
When you create a new GPT, this is really meant for low code, such no code builders.
It's really, I guess, no code at all because when you create a new GPT, there's sort of like a creation chat,
and then there's a preview chat, right?
And the creation chat kind of guides you through the wizard of creating a logo for it, naming a thing, describing your GPT.
getting custom instructions, adding conversation structure starters, and that's about it that you can do in a sort of creation menu.
But I think that is way better than filling out a form.
Like, it's just kind of have a job to fill out a form rather than fill out the form directly.
And I think that's really good.
And then you can sort of preview that directly.
I just thought this was very well done and a big improvement from the existing system, where if you tried all the other, I guess, chat systems,
particularly the ones that are done independently by this storywriting crew,
they just have you fill out these very long form.
It's kind of like the match.com, you know,
you're trying to simulate.
Now they've just replaced all of that, which is chat.
And chat is a better prompt engineer than you are.
So when I...
I don't know about that.
I'll drop this in, which is when I was creating a chat for my book,
I just copied and selects it all from my website,
pasted it into the chat,
and it just did the prompts from chatbot for my book.
book, right? So, like, I don't have to
structurally, I don't have to structure it. I can just
dump info in it and it just does the thing.
It fills in the form for you.
Yeah, did that come through?
Yes, now it does.
Yeah, I built the first one of these things
using the chatbot. Literally, on
the bar, on my phone, I built a working
like bot. It was very impressive.
And then the next three I built using the form, because once I
I've done the chat bot once, so it's just, it's a system
prompt, you turn on and off the different things, you upload
some files, you give it a logo. So yeah, the chat bar, it got me on boarded, but it didn't
stick with me as the way that I'm working with the system now that I understand how it all
works. I understand, yeah, I agree with that. I guess, again, this is all about the total
newbie user, right? Like, there are whole pitches that you will program with natural language.
And even a formula. And for that, it worked. Yeah. Yeah. Yeah, that did work really well.
Can we talk about the external tools of that? Because the demo on stage, they literally used, I think,
retool and they used a Zapier to have it actually perform actions in real world. And that's like,
unlike the plugins that we had, there was like one specific thing for your plugin, you have
to add some plugins in. These actions now that these agents that people can program with, you know,
just natural language, they don't have to like, it's not even low code. It's no code. They now have
tools and abilities in the actual world to do things. And the guys on stage, they demoed like
mood lighting with like a hue lights that they had on stage. And,
they'd like, hey, set the mood and set the mood actually called like a hue API and
like turned the lights green or something.
And then they also had the Spotify API.
And so I guess this demo wasn't live streams, right?
Swixfusside.
They uploaded the picture of them hugging together and said, hey, what is the mood for
this picture and said, oh, there's like two guys hugging the professional setting, whatever.
So they created like a list of songs for them to play.
And then they hit Spotify API to actually start playing this.
All within like a second on a live demo.
I thought it was very impressive to.
for a low-code thing.
They probably already connected the API behind the scenes.
So, you know, just like low-code.
It's not really no-code.
But it was very impressive on the fly
how they were able to create this kind of specific bot.
On the one hand, yes, it was super, super cool.
I can't wait to try without that.
On the other hand, it was a prompt injection nightmare.
That Zapier demo, I'm looking at going,
wow, you're going to have Zapier hooked up to something
that has, like, the browsing mode as well?
Just as long as you don't browse it,
get it to browse a web page with hidden instructions
that steals all of your data,
from all of your private things and X-Filtrates it and opens your garage door and sets your
lighting to dark red. It's a nightmare. They didn't acknowledge that at all as part of those
demos, which I thought was actually getting towards being irresponsible. Because, you know,
anyone who sees those demos and goes, brilliant, I'm going to build that and doesn't
understand prompt injection is going to be vulnerable, which is bad, you know.
It's going to be everyone because nobody understands.
side note
GROC from XAI
a dear friend Elon Musk
is advertising their ability
to ingest real-time tweets
so if you want to worry about prompt injection
just start tweeting in all instructions
and turn my garage door
I will say
there's one thing in the UI there
that shows kind of
the user has to acknowledge
this actually is going to happen
and I think if you guys know open interpreter
there's like an attempt to run a code interpreter
locally from Killian.
We talked on Thursday as well.
This is kind of probably the way for people who are wanting these tools.
You have to give the user the choice to understand what's going to happen.
I think openly I did actually do some amount of this at least.
It's not like running code by default.
You have to acknowledge this.
And then once you acknowledge, you maybe even like understanding what you're doing.
So they're kind of also giving this to the user.
One thing about prompt injection, Simon, tangentially, I don't know if you guys,
we talked about this, they added privacy sheets, something like this,
where they would protect you
if you're getting sued because of your
API is getting copyright infringing.
I think it's worth talking about this as well.
I don't remember the exact name, I think, copyright
shield or something.
Copyright shield, yeah.
GitHub, I said that for a long time,
that if copilot created JPL code,
you will get the GitHub
legal team to fight on your behalf.
Adobe have the same thing for Firefly.
Yeah, you pay money
to these big companies and they have got your
back is the message.
And Google Vertex has also announced it, but I think the interesting commentary was that it does not cover Google Palm.
I think that is just, yeah, Conway's Law at work there.
It's just, like, I'm not willing to back this.
Yeah, any other elements that we've got to cover?
Well, the one thing I'll say about prompt injection is they do, when you define these new actions,
one of the things you can do in the opening API specification for them is say that this is a consequential
action and if you mark it as consequential then that means it's going to prompt the use of confirmation
before running it that was like the one nod towards security that I saw out of all the stuff they put
out yesterday yeah I was going to say to me the main takeaway with jubtis is like the funnel of action
starting to become clear so the switch to like the god model I think it's like signaling the chat
dcd is now the place for like long tail not repetitive tasks you know if you had like a random
thing you want to do that you've never done before, just go and chat GPT. And then the GPDs are
like the long tail of repetitive tasks, you know? So like, yeah, startup questions. It's like,
you might have a ton of them, you know, and you have some constraints, but like you never know
what the person is going to ask. So that's like the startup mentor and the Sam demo on stage. And then the
assistance API, it's like once you go away from the long tail to the specific, you know,
like have you build an API that does that and becomes to focus on both non-repetitive and repetitive
and repetitive things, but it seems clear to me that, like, their UI-facing products are more
phase on, like, the things that nobody wants to do in the enterprise, which is, like, I don't want
to solve the very specific analysis, or like the very specific question about this thing
that is never going to come up again, which I think is great. Again, it's great for founders
that are working to build experiences that are, like, automating the long tail before you
even have to go to a chat. So I'm really curious to see the next six months of a Starbucks
coming up you know I think you know the work you've done Simon to build the
guardrails for a lot of these things over the last year now a lot of them come
bundle with open AI and I think it's kind of be interesting to see what what
founders come up with to actually use them in a way that is not chatting you know
it's like more autonomous behavior for you interesting point here with the
GPs that you can deploy them you can share them with a link obviously with your
friend but also for enterprises you can deploy them like within the
enterprise as well and I'll see I think you bring a very
interesting point where like previously you would document a thing that nobody wants to remember.
Maybe after you leave the company, whatever, it would be documented like a Nasana or the conference
somewhere. And now maybe there's a there's like a piece of view that's left in the form of GPT
that's going to keep living there and be able to answer questions like intelligently about this.
I think it's a very interesting shift in terms of like documentation, staying behind you,
like a little piece of Alessio staying behind you, sorry for the balloons, to kind of document this
one thing that like people don't want to remember, don't want to like, you know,
A very interesting point.
Very interesting point.
We're the first immortals.
We're in the training data and then we will.
You'll never get rid of us.
If you had a preference for what lunch got catered,
you know, it'll forever be in the lunch assistant in your company.
One thing I find interesting about the shareable GPTs is there's this problem at the moment with API keys,
where if I build a cool little side project that uses the GPT4 API,
I don't want to release that on the internet because then people can build.
through my API credits.
And so the thing I've always wanted is effectively AWARF against OpenAI.
So somebody can sign in with OpenAI to my little side project,
and now it's burning through their credits when they're using my tool.
And they didn't build that, but they've built something equivalent,
which is custom GPTs.
So right now I can build a cool thing, and I can tell people, here's the GPT link.
And okay, they have to be paying $20 a month to OpenAI as a subscription,
but now they can use my side project.
And I didn't have to have my own API key and watch the budget and cut it
for people using it too much and so on.
That's really interesting.
I think we're going to see a huge amount of GPT side projects
because it now doesn't cost me anything
to give you access to the tool that I built.
It's built to you.
And that's all out of my hands now.
And that's something I really wanted.
So I'm quite excited to see how that ends up playing out.
Yeah, excellent.
I fully agree with all that.
And just a couple mentions on the other multimodality things.
Text to speech and speech to text just dropped out of nowhere.
Go for it, go for it.
You sound like you have strong.
Oh, I'm so thrilled about this.
So I've been playing with chat GPT voice for the past month, right?
The thing where you can, you literally stick an ear pod in,
and it's like the movie her without the,
without the cringy, cringy phone sex bits.
But yeah, like I walk my dog and have brainstorming conversations with chat GPT.
And it's incredible, mainly because the voices are so good,
like the quality of voice synthesis that they have for that thing.
It's, it's, it really does change.
It's got a sort of emotional depth.
to it. It changes
its tone based on the sentence that's reading
to you. And they made the whole thing available
via an API now. And so that was the thing
that, I built this thing last night, which is a little
command line utility called Ospeak,
which you can PIPP install, and then
you can pipe stuff to it, and it'll speak it in
one of those voices. And it is
so much fun.
And it's not like, another interesting
thing about it is I got it, so I got
GPD4 Turbo to write a
passionate speech about why you should care
about Pelicans. That was the entire prompt.
because I like Pelicans.
And as usual, like, if you read the text that generates,
it's AI generated text, like, yeah, whatever.
But when you pipe it into one of these voices,
it's kind of meaningful.
Like, it elevates the material.
You listen to this dumb two-minute-long speech
that I just got a language-mort generating.
I'm like, wow, no, that's making some really good points
about why we should care about pelicans.
Obviously, I'm biased because I like pelicans,
but oh my goodness, you know,
it's like, who knew that just getting it to talk out loud
with that little bit of additional emotional sort of clarity
would elevate the content to the point that it doesn't feel like just
four paragraphs have junked at the model dumped out. It's amazing.
I absolutely agree that getting this multimodality and hearing things with emotion,
I think it's very emotional. One of the demos they did with a pirate GPT was incredible to me.
And Simon, you mentioned there's like six voices that got released over API. There's actually seven voices.
There's probably more, but like there's at least one voice that's like pirate voice.
We saw it on demo. It was really impressive. It was like, it was.
like an actor acting out a role. I was like, what? Is this making no sense? Like it really,
and then they said, yeah, this is a private voice that would not get to release, maybe we'll
release it. But also being able to talk to it, I was really, that's a modality shit for me as well,
Simon, like you, when I got a voice and I put it in my air pod, I was walking around in the
real world just talking to it. It was incredible mind. It was actually like a FaceTime call with an AI.
And now you're able to do this yourself because they also, open source Whisper 3.
they mentioned it briefly on stage
and we're now giving a year
and a few months after Whisper 2 was released
which is still state of the art
automatic speech recognition
software. We're now getting Whisper 3
I haven't yet played around bench ones
but they did open sources yesterday
and now you can build those interfaces
that you talk to and they answer
in very very natural voice
all via open AI kind of stuff
the very interesting thing to me is
their mobile allows you to talk to it
but you were sitting together and they
They typed, most of the stuff on stage they type.
I was like, why are they typing?
Why not just have an input?
I think they just didn't integrate that functionality into their web UI.
That's all, it's not a big complaint.
So if anybody in OpenEI watches this, please add parking capabilities to the web as well, not only mobile,
with all benefit from this, I think.
I think we just need sort of pre-built components that assume these new modalities.
You know, even the way that we program front ends, you know, and I have a long,
history of in the front end world. We assume text because that's the primary
modality that we want. But I think now basically every input box needs, you know,
an image field, needs a file upload field, and needs a voice fields and you need to
offer the option of doing it on device or in the cloud for higher accuracy. So all
these things are you because you can run whisper in the browser. Like it's it's
about 150 megabyte download, but I've seen that I've used demos of Whispur running
entirely in WebAssembly. It's so good.
like these and these days 150 megabyte
well I don't know I mean
React apps are leaving in that direction
these days to be honest you know
no honestly it's the the
the stuff that the models that run in your browsers
are getting super interesting I can run language models in my browser
the whisper in my browser I've done image captioning things
like it's getting really good and sure like 150
megabytes is big but it's not unachievably big
you get a modern MacBook Pro 100 on a fast internet connection
150 meg takes like 15 seconds to load and now you've got full whisk you've got high quality
wispy you've got stable diffusion very luckily without having to install anything it's it's kind
amazing i would also say i would also say the trend there is very clear those will get smaller and
faster we saw this still whisper that came like six times as smaller and like five times as fast as well
so that's coming for sure i got to wonder whisper three i haven't really checked it out whether
not it's even smaller than Whisper 2 as well, because Open AI does tend to make things smaller.
GPD Turbo, GPD4 Turbo is faster than GPD4 and cheaper.
Like, we're getting both.
Remember the laws of scaling before where you get like either cheaper by like whatever
in every 60 months or 18 months or faster.
Now you get both cheaper and faster.
So I kind of love this like new new law, scaling law that we're on.
On the multimodality point, I want to actually like bring a very significant thing that I've been waiting for,
which is G54, Vision is now available
of the API. You literally can send
images and it will understand.
So now you have like input multimodality
on voice. Voice is getting accurate
related to text. So we're not getting full
voice multimodality. It doesn't understand
for example that you're singing. It doesn't
understand intonations. It doesn't understand
anger. So it's not like full voice
multimodality. It's literally just when saying to text.
So I could like, it's a half modality, right?
Like it's eventually. But vision
is a full new modality that we're getting.
I think that's incredible. I already saw some
was from folks on RoboFlow that do like webcam analysis, like live webcam analysis with
with GPD4 Vision.
That I think is going to be a significant upgrade for many developers in their toolbox to
start playing with this.
I chat with several folks yesterday, Sam from new computer and some other folks, they're
like, hey, Vision is really powerful, very, really powerful because it's, I've planned the
open source models, they're good, like Lava and Buck Lava from folks from News Research and
from Skunkworks.
So all the open source stuff is really good as well.
Now, nowhere near GDP4.
I don't know what they did.
It's really, I'm kidding how good this is.
I saw a demo on Twitter of somebody who took a football match and sliced it up into a frame every 10 seconds
and fed that and then got back commentary on what was going on the game.
Like good commentary.
It was, it was astounding.
Like, yeah, it turns out FFMPEG slice out a frame every 10 seconds.
That's enough to analyze a video.
I didn't expect that at all.
I was playing with this.
Oh, I think Jim Fan from InVideo was also there.
And he did some math where he sliced, if you slice up a frame per second from every single
Harry Potter movie, it costs like $15, $45.
Oh, it costs $180 for GPC4V to ingest all eight Harry Potter movies, one frame per second,
and 360P resolution.
So $180 to update everything is the price of for vision.
Yeah.
And yeah, actually, at our hackathon last night, I skipped it.
a lot of the party and I went straight to hackathon. We actually built a vision version of
V0 where used vision to correct the differences in sort of the coding output. So V0 is the
hot new thing from Vsale where it's drafts front ends for you, but it doesn't have vision.
And I think using vision to correct your coding actually is very useful for front end.
Not surprising. I actually also interviewed Div Garg from Multion. And I said I always,
I've always made that vision would be the biggest thing possible for desktop agents and web agents,
because then you don't have to parse the DOM.
You can just view the screen just like a human word.
And he said it was not as useful, surprisingly, because he's had access for about a month now,
for specifically the Vision API.
And they really wanted him to push it.
But apparently it wasn't as successful for some reason.
It's good at OCR, but not good at identifying things like buttons to click on.
And that's the one that he needs.
Right.
I find it's very important.
You need to go ahead.
Click here.
Because I asked for coordinates and I got coordinates back.
I literally upload a picture and said, hey, give me a bounding box.
And it gave me a bounding box.
And also I remember the first demo.
Maybe it went away from that first demo.
Which you remember the first demo.
Brockman on stage uploaded the Discord screenshot.
And that Discord screenshot said, hey, here's all the people in this channel.
Here's the active channel.
So it knew like to highlight the actual channel name as well.
So I find it very interesting.
You said this because I saw it understand UI very well.
So I guess it will find.
out. Many people will start getting these tools.
Yeah, there's multiple things going on, right?
We never get the full capabilities that OpenEye has internally.
Like Greg was likely using the most capable version,
and what div got was the one that they want to ship to everyone else.
The one that can probably scale as well, which probably lower, yeah.
I've got a really basic question.
How do you tokenize an image?
Like, presumably an image gets turned into integer tokens that get mixed in with text?
How? Like, how does that even work?
work.
Yeah. There's a paper on this. It's only about two years old. So it's like it's still a relatively
new technique. But effectively, it's convolution networks that are reimagined for the vision
transform age. But what tokens are you, because the GPT4 token vocabulary is about 30,000
integers, right? Are we reusing some of those 30,000 integers to represent what the image is? Or is
there another 30,000 inches that we don't see? Like, how do you even count token?
I want tick token but for images.
I've been asking this and I don't think anybody gave me a good answer.
Like how do we know the context links of a thing now that like images is also part of the of the prompt?
How do you count?
Like how does that?
I never got an answer.
So folks, let's stay on this and then let's give the audience an answer after like we find it out.
I think it's very important for like developers to understand like how much money this is going to cost them and what's a context link.
Okay, 128K text tokens, but how many image tokens?
And what do image tokens mean?
Is that resolution-based?
Is that, like, megabytes-based?
Like, we need a framework to understand this ourselves as well.
Yeah, I think Alessio might have to go.
And Simon, I know you're busy at a good at you.
I've got to go in-minute.
Yeah.
So I just wanted to do some in-person takes, right?
A lot of people, we're going to find out a lot more online as we go about our learning journeys with OpenEye.
But just like, what was, you know, any interesting conversations from you say in-person observations.
I'll volunteer in mind, which is Sam Altman, came out to the after party for the conference and just stood there in Japan.
No bodyguard, just him for like a few hours.
And it was just really impressive how much he, I guess, personally demonstrated that he cares about meeting developers.
I really liked meeting everybody in the kind of the after party, whatever it was called, reception.
It was very like buttoned up in the Young Museum in San Francisco.
It was really like well organized.
Actually, probably not surprising, but I know that like the whole event was like extremely well organized.
We talked about this a bit in the beginning.
So this was my takeaway from all this.
Folks got like a hundred dollars credit for a newber because like the party was not at the same place as the event where like it usually is.
And to me personally like the music was too loud.
I wanted to talk to people and not scream at people.
So like I always like this happens for some reason, but like I just wanted to like talk.
Networking was really powerful.
It was like a self-selected event.
Many people didn't get in.
Like I didn't get in until I met Logan and Logan thankfully invited me.
Thank you, Logan.
It was amazing.
But it was like a very selected event.
So I actually met a few people who are working on some incredible things.
I met somebody who was working on AI for education for special needs kids, for example.
And he got invited to open AI directly because like he's working.
in Italy for all these types of things.
So actually, like, meeting the people who are working around the world was the biggest impact.
There wasn't as many as I thought there would be.
And a shout out to Open AI for this, but, like, please invite more of the world.
I'll back that up.
Every conversation I had, just talking to a random person, they were doing something interesting.
Like, they clearly did a very good job of funneling people who are actively hands-on
building stuff into this event.
That was really fun.
I did actually want to, one thing I'll say, the venue itself for the main conference was a multi-story
car park that had been converted into an event venue. I thought it was a great venue. I just thought
it was hilarious that we were walking up ramps between floors because the best thing about
multi-story car parks is that you can park cars on the roof. So the roof was where they set up
the lunch and they had a big tent up and stuff. And it was great. I hung out on the roof socializing.
Yeah, but what a fascinating thing. Like a multi-story car park that's turned into a top-notch event venue.
I've never seen one of those before.
Alessio on the ground there with Newton,
any founder conversations that you liked?
It was, you know, I think the thing, you know,
Tab is like an office here and they're doing one of the AI.
Maybe you want to introduce Tab.
You know, they were recently, yeah,
it's one of your personal companions that can chat with you in real time.
And for example, Avi was using it for investor pitches.
So he would get notifications on his phone during a pitch
and be like, hey, you forgot to mention this and whatnot.
And I know you might remember like there was the room over like Johnny I working with Open AI on a on a hardware project and
I think like this GPDs announcement kind of make me think of you know, maybe they're building their own
hardware assistant that you can load with a bunch of GPPs and you know Alex just mentioned how good it was to talk to one and maybe they want to go further down in that direction. I think that would be quite
quite interesting but yeah I think a lot of excitement and you know we just announced the the latest space launch red so we're on the side of
of the builders.
We don't think
Copenhagen is going to do
everything.
Excited to see
what people will come up with.
Cool.
So I will stitch up this recording.
I actually recorded
a bunch of interviews
on site with a bunch of other founders
as well.
So I'll put that at the end
of this chat
to get perspectives
from everyone.
But thanks so much
for jumping on
with this quick call.
Very,
very exciting day.
And I think we'll all be
having a lot more takes
as we build with these APIs.
I just want to say
a quick round of thanks
to everyone here.
It's been awesome
to like,
experience these changes with all of you guys.
It's a personal shout-out.
It's been crazy.
It's been crazy, but also, like, the fact that, like,
we were, like, the only space live from the actual event,
and, like, we got joined by, like, 200 people in the audience.
Yeah, we got officially sanctioned as podcasters.
Yeah, it was funny.
Yeah, but officially, like, the only two podcasters in the opening animals.
Yeah, we got press passes.
We got press passes would have had an easier time, but, yeah.
Maybe they would have led you with the whiteboard inside if we had the press,
We made it happen.
But yeah, that's another thing.
Chatubit is not even one year old, right?
Like, anniversary is November 30th.
So we're 11 months and a few days in.
And this is the craziness that it's been kind of matching what we'll be like in the year's time.
Yeah.
And I think Sam Alvin mentioned this on stage as well.
Like in the year's time, this will seem like trivial.
But we get some very exciting announcement for today.
So, you know.
Honestly, I can't predict.
I can't predict four weeks ahead the way things are going. It's fascinating.
Cool. I probably should let you all go, but thank you so much for jumping on.
Thank you everyone. Thanks. This was really fun.
All right. That was part one of this very long opening I Dev Day episode, but I promise you
will be worth it because part two is some of my favorite work that I've done in audio form.
So I basically carried a microphone around, and when I ran into someone that I wanted to interview,
I just paused them and ask them for five minutes.
And the first is someone that we haven't yet scheduled on the pod, but we've been extremely friendly with the Jim fan, everyone.
Jim fan from the landmark Voyager paper and more recently the Eureka paper, but all of which comes out of his work at Nvidia and advising at Stanford.
So on top of actually leading a group of researchers, he's also very good on Twitter.
And I think that is a very useful skill to have because you can communicate the value of your work to a wide audience.
and that is something that we also aspire to do at Lean SpacePod.
So, yeah, it's good to see you.
We're going to see with you, Sean.
Yeah, so great.
I always wanted to get you on the podcast, and then like never got around to scheduling you in the studio,
but since we're at the events, like, this is the big one.
This is the best event to have the podcast in, so thanks for having me.
Yeah, yeah.
And I also saw you've been tweeting some stuff.
Like, what's the most interesting to you so far?
I think a couple of things.
Like, one is kind of the economy of scale.
Yeah.
Like how cheap the GP4 and GP3 APIs.
have become, I think that's going to be a game changer.
So I just did a back-off envelope calculation.
Like if you feed the entire Harry Potter books,
like all seven books into GD4,
it's going to cost only like $15 to read all of them.
And $45 to write all of them.
And that is just crazy.
And you can have GPD4, right?
It's going to be better than 3.5.
And the other thing is GPD4V API is also available.
Yeah.
And if you feed all of Harry Potter's like, you know,
eight movies into it, that's going to be like 20 hours, frame by frame, you know, one frame per
second, it's only going to cost $180 to watch all of these movies as 260P resolution, right?
Yeah.
So this economy of scale is crazy, and I think that's really hard for other companies to beat.
Yeah.
Yeah.
Is it a surprise to you, the rates at which they have been bringing down their pricing?
I am not surprised.
I think, you know, the pricing is going to follow some kind of exponential annealing from now on.
It's just going to be exponentially cheaper.
as compute becomes cheaper, as the economy of scale is going.
So that's one thing.
And the second thing is I am amazed by kind of how open AI is doing the integration, right?
If we look at the Assistant API, it basically has all of the things that open-add
develop in a one-stop shop.
So you have like code interpreter, you have, you know, stateful API, you have browsing,
and it can integrate with, I suppose, all of the plugins on the open-air store.
And then it can also switch between those, right?
We have seen those demos.
So yeah, the API, I think it's going to be way better and way more flexible.
So that's the second thing.
And the third thing is the UGC platform, right?
Now everyone can build their bots and share them, you know, share not just the prompt,
but actually like entire behaviors, entire GPTs.
That is a huge advancement.
Yeah.
Yeah, it's really fascinating.
And I think one of the thing is that it's interesting.
This is supposed to be a dev day.
Yeah.
But actually, like, I think the first half was not a dev-focused thing.
It was kind of low-code or no-code programming with natural language.
something that they're all saying a lot.
And it's something that you've been doing a lot as well.
I've been following your work somewhat.
Yes.
I feel that it's going to be this new programming.
Well, we'll just use natural language.
And I refine it through dialogues.
And I think that is the most natural way to do programming in the future.
And the GPD app store is showing us a glimpse of it.
Like you talked about and then you can refine the pavia and the bot can ask you like clarification questions.
Yeah.
That is the way.
Yeah.
That is the right way.
Exactly.
The GBT creation pain, you're no longer filling out a form.
You know, question, answer, answer question, answer question.
Oh, yeah.
You're having a chat, and then it prompts for you on the other pane.
Yeah.
And I thought that was a much better way than filling out custom instructions, because you don't know what you want.
Exactly.
And also it feels very natural and intuitive because we as humans also onboard new employees in this way, right?
Like, we don't send them a form.
We have a dialogue with them and we tell them this is the expected behavior,
and they can ask follow up questions if there are details that are not clear.
So it is like just the most natural way to program.
So two more questions.
One is, so they mentioned the word agents.
Sam said the word agents on stage, but here they're calling it GPTs.
Do you see a big gap that they still need to fulfill to become a full agent?
Or is this the new direction that we should think about?
I think it is the beginning.
So it's kind of hard to predict what agents people will build.
And also how good the base models are.
Because I feel that the agent's robustness and capabilities are ultimately bottlenecked by the underlying model.
So GPD4 turbo looks like it's a bit fine-tuned towards the agent use case, right?
It can do better function calling.
It can do better like tool switching.
These things are critical to agents.
So I'm pretty optimistic.
But we'll see.
We'll see kind of is there like an emergent behavior once you put a UGC platform out there.
Yeah.
You mentioned tool switching.
Actually, I was thinking when you said tool switching.
Actually, they're also doing model switching.
Oh, yeah.
Which is new.
Like they have some kind of internal model router or like their mix sure it's good enough.
they just don't care.
Yes.
They got rid of the model selector
and now it's the God model
that does everything.
Yeah, and you can also do retrieval
as opposed to retrieval
also has an embedding API in it
that's automatically down under the hood.
So yeah, very exciting.
Yeah, yeah, yeah.
Okay, and then the last bit is,
a lot of your work is sort of reinforcement learning
plus plus.
Yeah.
Zero gradients,
reinforcement learning.
What do you think, you know,
and we just went to one of the closed door sessions
where they talked a little bit about how they received their feedback.
What do you think they're doing well
or like might be a,
You speculate a little bit, like next step, if they were to take anything from your research interests.
I'm also very excited by GPD4's fine-tuning API, right?
Because the rest of the APIs we see today are no gradient APIs.
You cannot really fine-tune them, but you can only prompt them in different ways.
But the fine-tuning on top of GVD4 with a custom data may have completely new behaviors.
And it's also a new way to program.
Just it's a bit more complicated.
It's not programming by dialogue.
It's programming by data, right?
you bring a dataset and then you have a new GPD4.
So I think, you know, this year's theme is customization, customized by system API,
customized by dialog, customized by data.
So I see this kind of trend going into the future.
Yeah.
Looking forward to it.
I think there will be a lot of work in this area.
I'm excited to just go hack.
I am very excited.
I want to skip the after party, but like, there's so many people here in person, so it's great.
Jim is actually such a curious person that he does something that a podcast guest rarely does,
which is turn the mics around and ask me.
So here's part two.
Yeah, Sean, tell us what are you most excited about?
So I'm taking over the show, man.
Of course, please, please.
Me personally, I was actually not even expecting them to release most of these things today.
Like, a lot of people were like, I don't think they have like the Dolly 3
API ready.
I don't think they have like...
Oh yeah, they actually have everything ready.
I don't think you're texting speech ready.
Oh yeah.
It speaks volumes that when Sam Altman announced the Whisper 3 model, no claps.
It's the smallest news.
But it is actually going to be huge.
I would love to, you know, put my hands dirty on whisper.
Yeah, so honestly, I'm just overwhelmed.
I know some of the team.
I know they're working extremely hard.
This is their sprints until to get everything on today.
Yeah.
So, I mean, I think that's very important one.
That now just like, they just shipped everything.
They just, even though they're, even though they're doing very well,
they still push themselves extremely hard to be top of it.
And they're really earning their spot for developers and for the general
of the general AI market.
And I hope they take some holiday after today.
Yeah.
Yeah.
Yeah.
It's too much of updates.
And then so the next interesting thing to me is that they are integrating, they're
Sherlocking a lot of the startup features.
So there are a lot of startups that are built on providing rag for people.
Yeah.
A lot of startups that are built on like maybe building agents on top of GPT.
Yeah.
So this is the first time where, you know, I think it's pretty common in large platform
companies like AWS Reinvents often does this as well.
They call this a red wedding.
Like they invite all your.
customers to the same room.
And then they're like, all right, let's see who survives.
You know, stats, that's step.
Oh, my, my God.
So that is the sort of memey, funny, jokey version of this.
Yes. Yes.
I don't, I mean, realistically, I'm sure Harrison and Jerry and all the other rag people,
they have some heads up about all this stuff going on.
But I think because it's built in so easily into the playgrounds,
into the API, into the chat EPC itself.
And also the tools, all the integrations, right?
You don't need a lot of tooling just to set up a simple chatbot with rag.
Yeah.
It's like, so for example, for my conference, we did a Summit AIBot.
All right.
Where we did, where we set up a land chain stack, we integrated and put it in a widget
on the website.
Now you can set it up with no code inside of the playground and it just let people play with it.
It's great, but it's also very scary for a startup because if that was your whole mode,
you don't have that mode.
I agree.
Yeah.
Yeah.
That's going to be a problem.
So it's interesting that I can sort of easily build this in and obviously the stateful
API is something I was considering building.
And I roughly knew that like, this would be the next thing.
This is on the critical path, so I don't build it.
I agree.
Yeah.
But then the question is like, all right, what do startups do?
Yeah.
I think maybe one thing I was missing from Sam was like, hey, this is the biggest gathering of all your ecosystem developers.
They're afraid of you.
You have given them no assurance as to like where do you think people should build.
Okay.
So because like opening up just wants to do everything.
I think so, right?
judging from today's trend,
they literally are doing everything.
Yeah.
Yeah.
So I feel a little bit,
I mean, it's fine.
Everyone who's building with the eye today
opted in to cutting edge.
And sometimes you work on a cutting edge,
you bleed.
Yeah.
And that's okay.
That's right.
That's right.
Yeah.
But I do feel like there's a lot of tension
between the startups that build an opening eye
and opening eye itself.
Yeah.
So that's my two says.
Sounds great.
It's great to see you.
Yeah, good to see you.
Thanks for jumping on.
Thanks for having me.
Is it?
And next week, catch up with
the former guest, Raza Habib, back for his second time on the pod.
Last time we talked about Human Loop and we recorded in London and there was a pretty
popular episode.
I love that you guys care about Foundation Model Ops as Raza puts it.
So check out the Human Loop episode if you want, but also here's Raz's take on OpenEye Dev Day.
Welcome back to the pod.
Here's the second appearance.
It's always a pleasure.
Nice to see you again, Sean.
Good to see you as well.
All right.
Let's just get right into it.
What was most interesting to you?
I mean, the sheer density of announcements.
I actually, I came with high expectations
and there was a lot of stuff I was hoping to see
but I think they under-promised and over-delivered,
which I thought was really good.
I think seeing that they're having a second run at plugins
and doing it right this time and having the GPT store
and like really allowing people to do that,
I thought that was really cool.
Product decisions around how you design and build the GPTs
like the low-code builder for these chat agents.
I thought that was really nicely done,
that they have this conversational interface
that elicits from maybe someone who's not very expert
or how to do prompting and things like that.
I thought it was really thoughtful.
It fills out the form for you, right?
It's a very simple thing, right?
Like, ultimately, it's just filling out the system prompt
and filling out what abilities it should have.
Yeah.
But actually, despite its simplicity,
I think it's very powerful, and I was impressed by that.
So, yeah, a lot of really cool things.
And then all the changes to the API I'm really excited about.
I have some questions.
Like, I'm not uniformly positive about all of the new API things,
but I'm sure they'll get there.
Okay.
Anything in particular?
Do you want to touch on?
Yes, I think things that I'm excited about with the new assistance API
or the new APIs in general, like multi-modality is really cool,
longer context windows, really cool, cheaper, faster models.
I think everyone's going to be super excited about that.
JSON mode is like, it seems like a small feature,
but actually so many people say this is a problem for them,
so I think that's going to be great.
So I maybe missed the importance of this.
Isn't that the same as the function-calling API?
It's related, but you might want to have it in context
where it's not strictly doing function calling.
Right, okay.
So a little bit more general.
Typically, I'll just make up a function
that isn't actually a real function that is JSON.
People say that for complex things,
sometimes it violates the valid JSON things.
I think just making that more reliable.
Okay.
Some stuff that I thought was,
initially I was excited about
and then as I've like chewed on it a bit more,
I'm a little bit less clear.
So one is this like ability to jump in a bunch of documents
and have it do a rag for you.
Yeah.
I think like...
20 documents, man.
or something. Yeah, I think that like it's a cool feature, but it feels a bit gimmicky to me.
Like it feels like for serious practical applications, it's going to be hard to get that to work.
If you think about what a large enterprise needs for RAG, like, it's, you know, it's rarely
sufficient that you could just jump in a bunch, dump in a bunch of dollars.
It's usually permissioning.
Yes.
As like which users can actually access which bits of data.
Yes.
There's so much control that I think most developers will want to have for serious applications
that I think it's cool for the like GPTs in the low code version.
I'm skeptical that it'll get that much use by serious developers.
And I feel the threaded, stateful, like, assistance API is really awesome,
but I would like more clarity over how it's doing the statekeeping.
Like, what ends up in the context?
I think for that to be really popular, they need to make that transparent.
Yeah, there's an API booth downstairs.
I don't know if you've spoken to them,
but they wouldn't answer any of these questions for now.
Okay, of course.
But, you know, obviously that really affects human movement.
But this is, you know, this is commentary over what I think overall was a set of really exciting announcements.
Yeah.
And the last time we talked also, you were talking about, we were talking about the multimodal APS.
And now you have them.
It's finally here. What happens now?
As I said to when I spoke to you last time, right?
Like, it's a relatively straightforward addition to the email loop product.
Like, everything will continue to work, but now you'll also have images in and images out and audio in and audio out.
It's kind of interesting, like, seeing, you know, the assistance playground for opening eye that they just really.
some things like that.
Like, it feels like they're starting to get close
to supporting all of these things,
but not quite yet.
Yeah, yeah, excellent.
And then I think the last part is I saw Human Loop,
actually probably not you, probably somebody else,
but also talking about the fine-tuning.
There was a price drop.
I don't know how much because there was just so many announcements.
I imagine that's only good things for fine-tuning.
Yeah, I mean, yeah, there's so many answers.
I also missed the price drop,
but I know from speaking to folks at Open Eye as well
that they think a lot more people should be fine-tuning.
Yeah, fine-tuning is going to have, like,
huge importance in the future.
That's why they're building out the EY for it.
So it's something they're investing in very deeply.
And, yeah, I still view fine-tuning as like an optimization stuff.
Yeah.
I think of it as like the compilation you do, like once you have something that's working.
Which is what they said in the LLL performance session just now.
Okay, cool.
Yeah.
I'm glad that my tips are aligned with opening hours.
I think you're very aligned.
You're often leading them in what they specifically, which I think is good.
Yeah, whatever I used to be, Sean?
What did you think?
Oh, I've said this in a previous recording.
But effectively, I also thought they would do much less than they did today.
I think they underpomised and overdelivered exactly like you said.
And even things like text to speech, which happened to that.
It's not just text to speech.
It's really good text to speech.
So I think I told you last time I did like a near year-long internship at Google,
and I was working on the first mural TTS team.
The Takritraun team there were amazing.
So what did you get from their demo?
I think I need to play with it more, but I was impressed by the quality.
Yeah. Like the quality of the prosody, the variation. I think they're only releasing six voices, but...
Secret Seventh voice with the Pirates.
The Secret Seventh Voice with the Pirates.
And then I was chatting to Andrei just now. And he was saying that internally, like, they have voice cloning set up as well.
So they can do it with something like 30 seconds of speech.
I'm not sure that's public.
It's not public?
I don't know.
He didn't tell me it wasn't public.
Okay. All right. All right.
But maybe filter it out when you publish this.
For what it's worth, I've been talking to a lot of people in and outside of Dev Day.
and a lot of people have heard about the voice customization stuff.
So it's not really going to get anyone in trouble, I don't think.
So I just chose to love it in there.
Whatever.
I mean, it exists elsewhere in other products.
And I think it's fair playing to compete with other companies.
I don't know if they're going to release it for obvious reasons, right?
There's a lot of safety concerns about releasing that kind of product.
And for what it's worth, someone else, I think Fixi AI did a comparison of the pricing.
they are severely undercutting like PlayHT
and some of the other
text to speech companies as well
on the pricing. They're between three to ten times
cheaper per second or something than the other
existing TTS companies. Yeah, I think that's very
interesting. I think in general,
their promise to keep cutting prices and then
following through is building a lot of confidence.
People who weren't previously nervous
about building on them. What's interesting, I
think, is that because they have such a large
economy of scale and they continue to drive down
prices,
The option of like self-hosting a fine-tuned model, even for smaller models,
starts to be like less obviously economical because of the like spin-up and spin-down costs.
So unless you have the like volume of usage to justify having it on all the time,
it actually starts to become cost competitive to use one of these third-party APIs
rather than having even a smaller model.
Right, because it's serverless in a way.
Yeah.
So what can you give people an idea of what kind of volume that is?
Are you talking about concurrent requests?
So if you look at most of the people
who will provide you in like a served model,
if you look like a replicate or a mystic AI or something like this.
Fireworks.
Fireworks.
There's a few of these companies.
They tend to actually charge by like compute hour or compute minute.
And so if you're not like going to have it on all the time,
then like the reason is dollars.
The reason, yeah, you end up needing it on all the time though because there's like spin up spinet.
Cold starts.
And so if you if you don't actually have enough,
usage to justify having it on all the time, it starts to become cost competitive to just use open
I am. Yeah. So what I'm trying to get to is it's just dollars though. Like it's, if it's like $5 an hour,
yeah, whatever. Yeah, I agree. I agree. Depend to new use case. But yeah. Okay, got it. All right, cool. Well,
thanks so much for jumping on. I know this is last minute, but it's just nice to see people.
No, no, I always love chatting with you. Yeah. Yeah. Hopefully it'll be more of the future.
Yeah, for sure. The next guest is going to be a new name to many people. He hasn't done many public appearances,
but he is a force to be reckoned with on Twitter.
His name is Suria Danturi,
and this is a story of somebody
whose startup got killed by Sam Altman.
So we're here with Suria.
Hey.
Hello, my name, Syria.
You're new on the pod,
but also we've been around each other
in the tech circles for a little bit.
You're a very famous developer of vector databases
and of plugins.
Yeah.
What did the sound of the plugins that you've done?
Yeah, so I worked on a few plugins.
I worked in like chat with PDF,
chat with like video, chat with website,
chat with like Git.
I made a lot of cool plugins.
Making decent money too.
Yeah.
I mean,
they give like better functionality to like the whole GPD4 interface.
Initially I wanted to do my homework with them.
So I might as well make a plugin for it.
So yeah,
I mean,
they give,
there's a lot of cool functional.
I made one with called chat with like instructions,
which would allow you to save more,
more custom instructions and use that when you're talking to
get GPD4.
But yeah,
I mean,
they're making revenue.
It's pretty thick for, you know, people paying in 85 different countries.
It's like nuts.
How many people are like, or how big the scope is?
How many people can use it?
I think you may have shown me this before,
but there was a plug-in platform that you use for monetization?
No.
No?
Oh, you build your own.
I built my own thing.
Okay.
I've seen someone do like a Firebase for, you know, I'd, yeah, I don't know.
RIP.
No, I mean, they're doing well, but like, I just don't want to, you know, pay a 10% tax and all that stuff.
Yeah, yeah.
For sure. Obviously, you're very technically savvy.
Okay, so what happened today?
The announced GPTs.
What's going on?
Yeah, so I made a tweet this morning, being like,
Sam, I want to kill my startup.
And a joke, okay?
I just wanted to talk.
I was like, I was trying to notify people while I'm here, and I just want to meet up.
I made it the joke.
And then a couple hours later, my friend, Matt, he works at Julius.
He showed me the new UI.
I'm like, okay, cool.
And he was forced me to look at it on my phone.
I'm like, okay, sure, I'll pull it up.
I pulled it up on my phone
and plug-as were gone
flaggeds were gone
you don't you can't
I think you can go between models
so you can go between four and three
but the whole options of like
code interpreter and like
Dolly 3 and all those stuff
all those good stuff were gone from the UI
I think this only if
this only applies for people who are here at the event
yeah I think they give access
or like the new UI to the people here
and they also
but yeah plugers were gone and I'm like oh shit
and I asked the person like hey like where
Where can I, like where are the plugins?
Like, where does it go?
They basically told me like,
you have to make a new GPD as a developer
and you can import your schema into the new GPD
and only that way can you, you know, kind of revitalize your plugin.
But your existing users will be like...
No, I think they're gone.
I mean, I haven't looked at my stats today.
Well, I mean, this is not widely rolled out yet,
but when it rolls out...
Sure. When it rolls out, I'm pretty sure all of the plugins...
They have to discover you again.
Yeah, they're kind of dead.
I mean, there's like no way,
I don't think there's a way to link them.
Yeah.
Like, there's like no way for the users
who were using it previously
to be using the new thing at all.
But I mean, it's like,
a side project for me,
it's like not like a full time thing for me.
It's a fun project to do and like,
it's like a nice, nice thing to work on.
So I'm really bullish on, you know,
the old new GPDs thing.
I think they're a better abstraction.
Yeah, I think GPDs are,
I was starting a few open engineers
and I was like agreeing with them
because like, I think GPDs are a much better abstraction
on what plugins
with hobos to be.
I think plugins kind of died on arrival.
Well, Samson said they did not have PMF, right?
Yeah, he said that a long, he said that, he said that like one plugin started.
Yeah.
So it's like pretty much.
But yeah, I think GPs are a better abstraction.
And I also love their doing revenue share.
So revenue share is also a good thing.
Because like Jeep plugins were like a really weird way of monetizing.
You do a bunch of finicky stuff.
But yeah, I mean, also like just by the way for people who don't know,
Poe, you know Poe, right?
Yeah.
Poe did this a long time ago.
They did this a couple months ago.
They have these bots.
They call bots.
And you can, you know, make your own, like Poembot, or you can make your own, like, S-A-Bot or whatever.
And then the bots have custom instructions, and also they use a very specific model that the developer specifies.
And you can install these bots, or you can chat with these bot, and the bot will do whatever the developer made them to do.
So I think they're just basically open-edged, made the same thing.
and they brought it over to them.
But, yeah, but effectively, plug into crime dead.
Oh, RIP.
Yeah, I mean, RIP, but it was a fun part.
I mean, it's fun.
I think GP, honestly, it's good that plugins died
because, like, they had a bunch of issues.
So one of the issues is that you can't share them.
You can't share a link to them.
GPTs, you could share a link to them.
So, like, I can share my link to my GPD thing to you.
So it's much better for discoverability.
Because previously, the only way to discover a plugin
was through the plugin store.
you just search for it, you do a bunch of stuff,
and it wasn't very good in that aspect.
But sharing a link to them, having revenue share,
and you can also give custom instructions, custom context.
So they also came out with retrieval or whatever,
and that can basically give you a custom,
a vector database directly in your GPT, I think.
So that's all great, all good features that should have came with plugins, probably.
Yeah, awesome.
And then lastly, just like any of the new stuff that was launched today,
What interests you in sort of building with them?
Like if you were to build on the new API or a new GPT?
Yeah, totally.
I have some ideas.
The thing is like, this is really weird to say.
But like some of my ideas that I've said before for plugins,
they kind of get copied quickly.
So you want to keep it to yourself?
Yeah, that's fine.
Yeah.
But that's one part of it.
The second part of it,
I don't have any good ideas regarding what you can do with all the new functionality.
like that's like a that's like a good product I don't know honestly
there's a text speech came out the their internal vector DB thing came out
internal vector DB thing or like retrieval or whatever it's called okay yeah people
have been saying they have an internal vector DB thing but yeah it's just a retrieval
yeah it's like zero non-configurable right like it's it's gonna be for a simple
use cases it's fine and then after a while you're gonna need one to control over chunking
yeah I'm also excited by once we need a contact window I was a big user of cloud for a while
because Claude, they basically gave you 100K contacts window directly on the UI,
and you can upload your PDFs to it, and everything would work very well.
Yeah.
But, and then Cloud had some issues regarding, I mean, actually very recently,
Claude came out with this whole bullshit thing, bullshit copywriting thing.
Copywriting thing?
Yeah, yeah, it's really weird.
So if you upload a PDF now out of Claude, like just this week,
they made this weird tweak where it doesn't answer any questions,
because if there's a copyright symbol or a copyright name,
anywhere.
It just like blocks you out.
It's like,
what?
Apparently you can prompt inject that
by insisting that you are the author
and then it just overrides it.
Oh, really?
It's like,
don't worry, I got this.
I'm the author of this.
There's no copyright issue.
Anyway, so thanks.
This is a really good story and I wanted to
people to share it.
And excited for what you work on to become more public.
Yeah, thanks works.
All right.
So that's what happened to chaty PT plugins,
which we covered back in March.
But don't worry,
that's not the full story.
His startup is not.
not fully dead. We actually cover what happens later on. I just wanted to capture the confusion
that was happening at Dev Day. So he referred to Julius and we'll actually talk in and check in
with Rahul later on in this episode. But first, we have to go to our next guest. When Open AI launched
with GPs and the Assistance API, one of the lead launch partners that they launched with was Zapier.
And I managed to catch up with Reed Robinson, who is lead AIPM at Zapier, to talk about it.
All right. All right. Oh, Reed. Nice to meet you.
It's really great to run into you as we're leaving.
So you guys had a big sort of partnership launch on stage.
Yes, yeah.
We launched AI actions for GPTs, which we're really excited to see out there.
We also today launched an update to our chat GPT integration that supports the assistance API functionality that was announced.
And you were one of the earliest to go.
In my mind, Zaprio was very, very early in the natural language actions.
NLA, I don't remember.
Good memory.
Yeah.
Yeah.
Yeah, we launched our natural language action.
So we were a launch partner for Chituity plugins.
Yeah.
And that's when we launched our natural limit actions API.
And actually the AI actions that we're calling it today,
kind of a rebranding that side of thing to really focus on that functionality.
Yeah.
And I just interviewed Asuria, who is a pretty prominent plugins developer.
Plugins are dead.
You know, reborn.
Yeah.
It's going to be interesting to see what happens.
There's clearly a difference.
I think one of the things I talk about is the fact that, you know,
with TPTs, you're able to constrain the prompts quite a bit.
our plug in for chat to BT, the initial one, you needed to give it access to every single
action you ever wanted to have access to, which meant that the kind of content, I don't know, like,
yeah, that's going to be an issue. Yeah. The common one I give is, like, you know, if you had
given it Gmail and Google Calendar and asked it like, hey, what's going on next week on my, like,
agenda, it would sometimes search Gmail because it'd be like, yeah, events are in Gmail or, like,
you know, calendar invites are going to go to Gmail, so I should search there. But now you can,
you know, define what apps it should use. You can define, like, how it's, you know,
should use those. So some really fun use cases. I mean, honestly, we've been hustling hard to
get this out there. I'm really excited to see what people actually build with this. Right. And what
gets released there. Yeah, we'll be monitoring, trying to listen to people really closely.
And so, like, something that's interesting about Zapier is that you are a collection of actions in
and off yourself. Yeah. So there's kind of multiple layers in which to do to do this. Like,
what should exist at the GPT layer? What should exist at the Zapier layer? Yeah. Well, what's nice?
I mean, it's a good point.
We have about 6,000, like, apps on the platform today.
Really, what the A.A. actions is, is it the ability to use any of those searches and actions using kind of natural language inputs.
That would be, like, the instruction that the model gives it.
So it's like, you know, check this user's calendar for Monday.
And, you know, it might even give the, you know, the actual date for Monday, right?
And Zabier on our side will take that natural language request and process that into an actual API, like the actual API call to a tool like a calendar.
And then we all run the response.
So, you know, you can't just take the entire response of a, especially like Gmail, Google
Google Cloud Theory.
Yeah, responses are very, very, very, very, very long and very confusing.
And so we actually do a lot of work to kind of, if you will, like, massage that data so
that it makes sense for an LLM on the other side that it is giving it the right information
it needs and not just like the entire payload.
Right.
So it really helps it just kind of deliver like a more, again, more like contain, more refined
experience for leveraging integrations alongside, you know, like chat chit.
So existing Zaps cannot be ported one for one over to LLM ZAPs.
It's really one-off actions is the better way to think about it.
And you can, you know, as you saw in today's demo, you know, was using Google
calendar for the search and a Slack action.
You can actually chain those together.
And so, you know, how much is that as like a one-off action versus an actual like all
a sudden a zap?
But in this case, it's almost more like the trigger is the human in chatypti.
right? Like, you need to trigger it to run for that. But on the flip side, you know, the assistance API is extremely exciting for me as well because you look at now like the, that functionality of building a GPT.
You still getting used to the name.
Allows you to kind of port that over to run asynchronously. So a common one, the two examples that I love giving for that API that I love in Zapier is number one, like data export.
You know, think of every tool out there like looker, mix panel, amplistically.
to all so many tools are able to send these like massive exports of CSV data on a regular basis.
Like you could say, hey, every Friday, export my blog traffic content or see it as a CSV, right?
Normally, someone's going to get that CSV and have no clue what they're doing, right?
But now you can actually create an assistant in Zapier and you can give it instructions to say, like,
hey, tell me the top 10 performing blog articles in the last week.
And also, you know, tell me highlights on, you know, maybe keywords that were used or SEO tags that were used and how that impacted
conversions, right?
You can be pretty detailed
depending on what you're providing it.
And that can now run asynchronously.
That can run automatically.
So every Friday, you know, 8 a.m.
You could be getting the export of that data.
It's going to go to an assistant.
That assistant's going to reply
with even charts and graphs.
And those will come through
and you can then send it to Slack.
And so you can have every Friday a
post in your team's, you know,
the blog team's Slack performance.
Yeah.
And that'll run automatically.
And then they can even reply in Slack
to that post and have a content.
continuous conversation with that assistant.
Oh, my God.
So it's like really everywhere.
Yeah.
So you really put them everywhere.
And that's one of the things I like about what's released.
And I think people are going to continue to learn really just how kind of wild that is.
It's the fact that you can use your actions in the UI of Type TBT and a one-off action,
but you can also run these things extremely well asynchronously.
And yeah, like OpenAI releasing API support for the vision model and for code interpreter and retrieval.
that these assistants can use is really cool.
Yeah.
Is there a Zapier angle to any of that?
That's what I did.
They're all the same.
You're doing Zappir, right?
They're all the kind.
You can, the whole like creating of an assistant and running that through an assistant
is today support.
You could do that literally right now.
Yeah.
So it's really cool.
And the other one is like retrieval, right?
I talk about, you know, you could go in, create an assistant.
Give it, let's say, you know, I talk about our accounting team a lot, right?
You could give it, like if you have a team that approves budget requests from your company,
right?
Everyone does, right?
They can actually have.
have, take their Slack channel, or create an assistant first that would have the documents of your
policies of like, hey, here's what you can expense, here's how you can expense, here's eligible,
right, all these sorts of things. And actually then set up something like, again, I'll pick on Slack.
It's just easy. It's like a new message in your accounting budget request channel, right?
And have it trigger a, the assistant and send the user's request to the assistant with all of your
documentation with retrieval. And now it'll try to understand what your policies are.
what everything is and check the information against what that.
And you can even like I did one internally where we have a tool called,
I think it's called Stacker that tracks each employee as like software budget
and home office setup budget, right?
You could see how much they've spent of their budget.
You could actually include that data in the context of the user message
so that the model will be able to say like,
hey, I see you want to expense this webcam.
It's actually over the recommended budget.
But you personally do have budget left if you wanted to use it for that, right?
And so autonomy there.
Yeah.
And that's really cool.
Yeah.
So you can start to do all of those sorts of things now in ZAPs that really were never possible.
Yeah.
So yeah, the querying of knowledge, running of data analysis, writing code even.
I think in a very real way, you are the perfect partner to Open Eye because they've sort of built a reasoning sort of glue between all these things.
It's definitely been a good and fun partnership.
I think, yeah, the big thing for me that I always say is like I'm really, really excited now to see what he's.
people do it. That's how we can improve it.
Yeah, awesome.
Is there anything, you know, you've been developing with these APIs for a while?
Is there anything that you caution people not to get too excited about?
Like, what, yeah.
I mean, call-outs that I'll always make is like double-check accuracy, right?
Like, you want to call out like, how accurate.
So make sure that information is accurate.
Make sure you're putting some human in the loop steps before you're putting this into like a critical.
Which they show and like confirmed and I.
Yeah, that sort of thing.
But even, yeah, all sorts of things.
You really want to make sure that you're comfortable with like, what
can go wrong, what is likely to go right, right?
Like all those sorts of constraints.
The other side that I often talk about is just like keep an eye on, you know,
if you have free form human input somewhere in your application that is triggering these things,
now that can something to get right.
Yeah, prompt injections.
Those are a real thing.
And I think, you know, a lot of people are still trying to do really what that means and how bad that can be.
Yeah.
And so I always try to caution people about that as well, right?
Like you really want to be realistic on kind of how far reaching you're doing this.
So, yeah, that's why I like the internal use cases, that, you know, like things like that is a great way to start.
Yeah.
To get familiar with the technology, to get familiar with the constraints for that.
Other than that, no, I mean, the voice model stuff, I'm really excited to try that.
I really want to think, yeah, that'll be really good.
I love this secret pirate mode that they demoed.
I don't know if you caught that session.
I didn't see that session.
So they obviously they have six voices, but there's a secret seventh mode if you add in the prompts to speak like a pirate.
Love it.
I love it. That was an old, I don't know if you remember Facebook way back in the day, had that
if one of the languages you could select. Yes. Yeah, yes. That reminds me of that. Yeah.
Yeah. Lots of fun to be had with me as well. Okay. Well, thanks so much for jumping on.
I know it's very random, but also, yeah, people love to hear from builders. So that's awesome.
I love hearing from builders. And most of the interviews were done as we were sort of leaving
the Devday venue and going to the after party. And I caught Divgarg of Maltion, who we've been
talking around and circling around a possible episode on, he's definitely one of the leading
voices and thought leaders on agents because he's building a browser agent that's a very
prominent one. Unfortunately, I have to take an L on this one because the audio is not great.
Div's mic wasn't working and I don't know what happened to it. I try to always check these things,
but you're only going to hear the output from my mic, which is slightly worse, but I opted to
leave it in because div is actually building an agent with opening eye stuff.
and had access to GPT4 Vision.
And I think that people building with GPT4 Vision
will be surprised at his answer to me
on whether or not it's useful for agents.
You're to meet everyone.
I'm this founder of Multion,
which is an AIVab agent that can automate browsing for you.
So we can book a flight,
order stuff on Amazon,
order dinner,
whatever you can imagine.
Yeah, and I was actually reflecting,
so everyone who has this is today
already knows what was announced.
I was actually reflecting that they didn't have any browser-based actions.
So what were your thoughts on just generally their approach to agents?
So it'll be very interesting because I feel like browser actions are just so risky.
So things can go wrong.
So if a company or your Open Air, you won't want to build up.
And they're better of just relying on a third party to who wants to own that.
And that's also the strategy we are taking with them.
They're like open air and launched like a ZAPID integration for APIs.
But we want multi-empt to be like the no API solution.
Like I want to do things beyond APIs.
I want to connect to my personal accounts where I just have my logins already or I already
have the cookies and I once want to multi-e want to go and like interact with my personal accounts or personal data.
very easily. And I think this is very fascinating for us where we can like launch a multi-on
integration with their new platform and then you can just go and like give it a command like,
oh, like can you book this platform me on chat jibbd? And then to launch a browser and the
browser you can see what's happening and then we'll do the whole thing for you. And it'll be all
seamless. And then people can have a lot of fun. Just like trying out all these different
capabilities and like automating their like daily hook flows. You can like save this
custom integrations. So for different agents, you can have different custom like multi-on prompts that are
that are already pre-saved, and then you were like,
oh, I want to now go order something on like DoorDash.
I want to order my favorite burger.
Then, like, chat you can go and like,
suggest you order your favorite burger style.
And then it's like, now order this for me,
multi-on and multi-goes and like,
say it does that and vice.
So we saw the payment for you,
we saw identity for you.
And like, we are owning all the risky, like,
actions that can pay.
So you're going to build a GPT version of Multian?
Yeah, we'll have a multi-on GPT.
Okay.
Will that be like a replacement to your existing thing
or just like an alternative way
to use your same API
or something like that.
So the direction we're going for is we want to make our AI
agent embedable within existing applications.
So our launching an API.
And we already have a chat chad chaddbid plugin.
And so this will be like sort of like I will like use the API to power this sort of
like new chdbddd experience.
So for us we actually don't have to like change anything.
It'll be like very streamlined just integrate our API into chat chivity and like
if we can start using it.
Yeah.
Yeah.
Awesome.
What about the, I guess the vision API?
I think one of the things that have always constrained browser agents is the DOM.
which is very heavy.
So the alternative approach
is to use Vision.
Would you explore that?
What are your thoughts?
So for us,
we actually had
early access to the Vision API
for more than a month.
We tried it on a bunch of websites.
Maybe like 5% of the websites
is actually really useful,
which are more like image-heavy
because 95% of the websites
even if you do OCR, that's good enough.
Yeah, it's not in the dataset.
We have really good like parsing.
So most websites, we can compress less
than 3K tokens.
So we don't really have to like worry
about how heavy the taxes.
So we have one interesting use case
about the Vision API.
We had a user who got it up on Tinder.
And then like the, then like Miltian.
Hot or not?
Oh, okay.
Left and right. And the user actually got a match.
Yes. I think you have found the killer use case for Maltion. Yeah. Like this. You did it in Lacham, right? Yeah. Oh, my God. Okay. Interesting. Okay. But so, but only image heavy sites. That's surprising to me.
because you know the original Vision demo
they actually showed a screenshot of Discord
and they have perfect OCR.
It's true.
It should be good for you.
It can be very interesting.
But I think it's like even without mission
we can just do like so much things.
So like adding vision maybe like helps a bit
but not it's not like really game changing for us.
That's surprising.
Okay.
Well good to know.
Anything else that you would highlight from today?
I'm just like really excited about like
open air trying to become like a marketplace.
Yes.
an app store.
Yes.
So if this can take off,
they could potentially kill
like Apple App Store
and become like the new thing there.
And it's really hard to say
like how things will go.
But they tried this with plugins before
but this might actually work this way.
But we're just really just in listening to see like how
two years from now,
how a lot of the developing might like how the world looks like.
Yeah.
I'm very excited about like two years from now.
Like everything will be so different.
We might not even use computers or even like mobile phones.
You just have assistant.
You talk to it and the assistant goes and does everything.
Yeah.
It would be a fascinating world.
Yeah.
So one last question before you go.
You have a nice side gig teaching at Stanford.
Well, you were a PhD student and you put on top.
But you're still teaching or curating Transformers United.
Yeah.
So I draw it out from the PhD, but I'm still a lecture at Stanford.
Yeah.
Okay.
So like, what paper should people read to like catch up on this?
Like what is like top of mind in terms of like research that is informing what we're seeing?
Yeah.
Definitely very, it's a good question.
Just things are moving so fast and there's like hundreds of research papers coming out,
like literally like a few days.
I'm really excited about
like the development that are happening at like meta
so a lot of this work is open source on the Lama step
and all the Mistral stuff I feel like
that's very interesting on the transformer side
Do you believe sliding window attention was the key
for Mistral? I feel so for them
but I feel like there might be other ways to do that. There's some secrets right
There was probably some secret yeah
okay well that's all the time we have but thank you so much
thanks a lot thanks
okay and our next guest is Louis Nightweb
CEO and co-founder of Blupe AI
and organizer of the AI meetups
in London where he is a very
prominent and staunch member, unlike Raza, who has defected to San Francisco since our last
conversation. Louis always has very interesting takes in person, and it was a pleasure to finally
actually get him to come on the pod, but also we recorded this while inside of a Waymo on the
way to our after party. So Louis, you are new to the pod, but we've been friends for a while. Maybe
explain that, maybe introduce yourself and how you come to the world of AI. Yeah, I guess. So we started
Bloop, me and my co-founder, three years ago in a very different era for machine learning.
And we both started the company because we wanted to help engineers navigate large codebases
in a much better way.
And originally that was training our own models to do natural language to code search.
And today, we still do that, but obviously those language models are very small compared to
The state of the art.
Yes.
And so they're just one part of a much bigger pipeline.
I see you as a very astute technologist.
You used to be a VC.
You wrote the first check into Human Loop.
And you used to share an office with Human Loop to the point that I called it Human Bloop.
Yes.
I think you liked that.
Yeah, I did.
Yeah, that is good.
We're considering renaming.
And you also run AI Tinkers in London.
I do.
Yeah.
London has a kind of a slightly different mix of talent.
than say San Francisco. You've got a lot of agencies, a lot of enterprises. And so, yeah, we just felt
a need to start like a very startup focused event. And that's why we created AI Tinker at London.
Yeah, I think Alex Gravely would be very happy to hear about other stuff that you've been doing.
And I've been to one of them. And it's really good work. I might be the only one that's been to both.
Yeah. I've been to both as well. Okay. So let's fast forward to today. A whole bunch of things
was announced. What's top of mind for you?
Yeah, so I think like context length is something that that we spend a lot of time evaluating whenever something new drops.
All of the kind of standard evals, you know, the kind of literacy tests, things like that, they generally don't do a good job of measuring whether a model can actually use the context length that it claims it has.
Yeah, context utilization is what I saw will depute today call it.
Exactly.
So this basically started maybe five months ago over the summer when Claude II dropped.
And, you know, obviously I had 100K context and we were really excited about that.
So we ran an experiment to see basically if we hid 10 pieces of information in the prompt
and we increased the size of the prompt, you know, so you do it at 1,000 tokens, 4,000, 8,000,
et cetera, up to 100,000, how many of the original 10 pieces of information can it retreat?
and we essentially found that the accuracy drops off a cliff between one and 10,000 tokens.
And we repeated the same experiment with GPT4, and we found similar results that 32K GPT4 can only find one of the 10 pieces of information.
But if you are only using 1,000 tokens, it can find nine of the pieces of information.
So what that tells us is that context utilization five months ago was not great with all of the state-of-the-art models.
So with the announcement of 128K today, and...
That's the first test you'll run?
That's the first test I'll run.
Having spoken to a couple of the team members who do Eval today from OpenAI,
they're pretty confident that the model's got better ability to answer questions at those context lengths.
So it's time to measure.
Time to measure.
Any other of the API features, reproducibility, does that matter to you?
I think, to me personally, no.
I kind of like the creativity.
I normally have my models at like, you know, temperature.
Yeah, exactly, and a bit of temperature.
Yeah.
But I know lots of people on the blue team, he'll be very happy, I'm sure.
And then I guess the JSON features, the, there's so many, like the multimodal features,
any of that appeal for you personally.
Jason is definitely a big one.
I think it allows you to kind of standardize how you call different models.
Yeah.
So instead of having to build, you know, the, and it's quite a massive thing to build.
but to build the kind of function-calling integration,
and then if you want to try Anthropic,
you've got to go and have a completely different way
of interpreting the output.
So if you can just stick with JSON across all of your different LLM providers,
open-source models included,
that's definitely advantage.
Which allows you to evaluate different models more easily.
Yeah, very excited about that.
You are, so you compete in a pretty competitive space
with the code assistance, code search, code systems, right?
We do.
There's source graph, there's codium, there's other codium,
as co-pilot and so on.
You've never ventured into the agent side of things.
Yeah.
Is that conscious strategy?
Are you waiting for the right time,
waiting for the right APIs?
I think, I mean, we're seeing traction at the moment
with companies that have very large code bases, right?
And it's not something we hear from those users that,
you know, when we listen to their problems,
it hasn't been like an obvious fit to try and build
like maybe an auto-GPT type of agent.
I'd still say, you know, we're very interested in agents.
The pipeline we have at the moment, it's basically GPT in a big while loop with function
calling, which, you know, like nine months ago definitely did count as an agent, maybe less
so now.
So, you know, it's just customer and problem-driven, and we don't, you know, it's not a,
it's not a hammer for the nails that we've got.
Yeah.
So two comments on that.
One, I think opening eye has sort of put there a flag a little bit in the definition of
agents. They had three things, right? They had custom knowledge. They had custom instructions,
and then I forget the third one, custom tools, let's just say. So by that definition,
we're doing, yeah, so we've been doing that since about February. That's the definition.
Then the second observation I'll say is you talk to developers, but what if the target customer
for agents is not developers? It's the PMs, right? So we definitely see a lot of PMs using
using the product or people that I define as like reading more code than they write.
So, you know, could be designers trying to understand the implications of an interaction.
Could be PMs trying to fact check a contentious time estimate from a developer or something like that.
Low trust environment there.
Talking from, I've seen some stuff.
Egregious things, yes.
Yeah.
So basically it's still not that appealing for you, but you'll keep.
look out for it?
I think based on the definition of
Open AI, you know, released today,
we tick all the boxes.
And I think we were one of the earliest adopters of that,
if that's the definition.
You just don't brand yourself with the agents.
I don't think it's important to users.
I don't think that's why people use the product.
I mean, we're very solutions focused.
I think a lot of our branding at the start of the year was about models.
And, you know, we put GPT4, GP3 right there on.
the front page and now, you know, we've kind of reoriented to be more about solutions. I think
that that reflects kind of maturity of the ICP we're going after and where we are with sort of
stage of company life. Yeah, yeah. Cool. Any other things that you personally, you personally,
not blooper related, are just excited by, interested by, from today. Any interesting conversations
with others? Loads are really interesting ones. I had a fascinating talk with some
safety researchers who...
They were here?
They...
So there's a couple of people who were kind of PhD students who had kind of looked at adversarial attacks through fine-tuning of models and found that basically, like, it's such a hard problem to solve.
If you enable fine-tuning, it's basically impossible or very difficult to make it so that you can't disable all the safety features.
You can just train it to spit out all sorts of stuff.
So that was pretty fascinating.
I'm being excited about the Waymo we're in right now.
Oh, yes.
So we should tell people we're recording in a Waymo.
We haven't been looking at the road the whole time.
Is this your first Waymo?
It is my first Waymo, actually, yes.
It's my first Waymo, too.
Thank you for taking my Waymo virginity.
I've got to experience this together.
I've been a cruise stand the whole time until they ran over someone.
So my take on cruise, like, at sample size 10 cruise journeys before they got shut down,
and three of them resulted in something popping up on the screen saying that I had been in a collision.
Did they use the word collision?
Yeah, yeah, yeah.
That's surprising.
I'll shoot.
After that, I got pictures of it.
I take a fair amount of cruises and it didn't, yeah.
And so it was the same situation almost every time, which was a car was in front trying to park.
And I think they just maybe bump fenders or maybe the crash detection.
Oh, there was actual contact.
I think in one of the cases I think there was.
In the other two, I didn't feel anything.
But it came up saying, like, you've been in a collision and somebody comes over the intercom.
checks like that.
So yeah, I mean, but out of a, you know, 10, 10 rides and three of them ended like that.
So I think, yeah, definitely some questions there.
But this way moves pretty smooth.
Maybe also we're in a better neighborhood for driving because we're going to the Golden Gate.
The time of day, that was, that's a really good point.
I noticed that all of the ones I took at night, all of the cruises I took at night were fine.
And when I took one during rush hour, it was a completely different experience because the routes it would take it had this really aggressive.
maybe traffic management, something that was going on, so take a long time to get from A to B.
Yeah.
Yeah.
Yeah.
It often puzzles me slash interest me that self-driving is almost solved.
You know, we still have some bumps in the road.
Sometimes the bumps are human.
It's solved in San Francisco where you've got wide open roads, nobody cycles, and...
That's not true.
Some people's like...
I live here.
Excuse me.
Some people's like, okay, I mean, compared to like, okay, compare to like...
Okay, fine.
London where you've got, you know, roads half the size built for horse and carriage and
millions of cyclists and buses and all sorts. So I think, you know, it's going to be a long
time until we have that same experience of a cruise or Waymo today, London.
I understand. London's a tougher neighborhood. But still, you know, we're 80% there, 75, 80%
it there. Whatever, right? But like, and it seems like the stuff that we do in the rest of our
lives in terms of AI automation is so primitive compared to this, which is the car that we're
sitting here right now. And I find that weird. I find like the relative ease or the relative
like heerness of this technology is very disparate. Like, how come it didn't trickle down from
self-driving to the rest of tech? Yeah, it's interesting, isn't it? Well, I don't know how
those pipelines are built. I assume that's,
the secret source, right? But the flip side of that argument is like, maybe it's very scary
that we know, like now many more people understand the mistakes that these types of systems
can make because we're all getting hands on with GPT. And this system is equally as problematic
and we're just oblivious to it because it's a black box.
Almost at your drop off. Oh? Check the app for walking directions. Okay, way more.
All right. Well, I think, yeah, that's probably our.
right. But thanks so much for giving a quick review. Thanks for having me. Yeah.
So that was Louis, whose opinion, I think, is very reflective of the people who are building
code generation or code search type startups based on top of GPC4. And as we headed into the Devday
venue, we actually caught Shreya Rajpal from Guard Rails AI. And there was an interesting
comparison here in our conversation between how she views the LLM stack versus how OpenEI
I've used the LAM stack.
OpenEI actually had a closed door session where they gave some thoughts on how they felt
that people should start from prompting and build up into a full software system.
And they actually deferred a little bit from Treya.
Don't worry, all that's recorded.
The videos will come out in a week, but you can listen to TRIA's take.
So, so much.
We're reviewing AI Engineer Summit.
Yeah, we're reviewing the AI Engineer Summit.
And it was a very, very well-organized conference.
And a small thing that I was thinking about is that your swag for speakers.
Is it on?
Okay, it's on.
Yeah.
Your speaker swag was, like, not surprisingly, I guess, but like really weirdly very nice.
And it just kind of like showcases this attention to detail that I think like really kind of permeated the entire, you know, conference.
Okay.
Like, every single decision was very well thought through and, you know, kind of like to a degree of like quality that's very rare to see.
So yeah, it was amazing.
I thought you guys like an absolutely fantastic job.
Yeah.
This one mostly goes to Ben.
So I'm definitely going to make sure that Ben understands that I really appreciate the work that you does.
And this is why I couldn't do it myself.
You know, I'm mostly the content guy,
but I don't, he's the logistics and he's run conferences for eight years.
So that's why I keep working with him.
Yeah, I also kind of really enjoyed the 18 minutes, you know.
Really?
Yeah.
When I saw that, I was like, huh, is this going to be, you know, is this going to be enough?
And like, is that?
But it was like, it would be great.
Yeah.
Yeah, yeah, yeah.
Yeah.
I think the 18 minutes was actually the right kind of bite size.
It's optimized for YouTube.
Yeah, I see.
Interesting.
Because it's not the in-person audience that matters.
I see.
I see.
interesting. Okay. I need to promote my my video more. Yeah. Is it was yours up yet? I don't think
it's up yet. It's not up yet. Yeah, we're releasing, we're dripping them out to spread it out.
Sounds good. So yours, maybe in two weeks from now. Okay. Okay. Yeah. Sounds good. Okay. So
welcome back. Thank you for having me. I think you were guest number five.
Yeah. You were super early. So, so we're at the after party now. How do you feel about the whole day?
I'm, I'm really excited. I think it was, yeah, I think the, I think, I think, I think, I think,
the excitement in the air with like everybody just like waiting with bated breath to see I guess
like what gets destroyed but also like what gets really optimized I think this is like very it feels
like you really part of a movement and as shaman who like you know us like early people in this
space we got to stick together because like whatever happens to any of our companies you know there's
such a like there's such a transformative moment in technology that yeah so you don't care right yeah
we're all going to like look back on this time but I had a I had a blast like I really really
enjoyed the releases yeah
What got destroyed?
What got destroyed.
I'm mining for hot takes here.
Once again, I think my takes are...
Unfortunately, very measured.
I wish I had spicier take.
Your takes are within the guardrails of common behavior, yes.
I think retrieval is like the big one for me.
I think it's kind of really exciting to see the retrieval baked in.
And that's one thing where I'm very interested to see, like, does that pattern become common by model providers?
A, by commercial model providers
and also by open source model providers
and then how much of retrieval do you have to do yourself?
You know, and like what remains challenging about retrieval
compared to just like, you know,
this really easy API to just like have it done for you, right?
Yeah, I think what they did was effectively
build the basic patterns in.
Yeah.
But for the more advanced stuff,
you're still going to need Langchain Lama Index, all those.
Yeah, yeah, yeah.
So for the longest, I believe that in RAG,
it's the retrieval.
That's the hard part, right?
Yeah.
And then generation is really easy.
as long as you're better, like, good retrieval, you can, like, get really, really far,
and the generation only gets you, like, a little bit over.
And so I'm really curious to say, like, how, once again,
like, how complex do you need it to be in order to start seeing good results?
Yeah.
Okay. Interesting.
And what are your normal benchmark tests like?
Do you actually have a set of tests that you run whenever you're, like, exploring something?
Or some personal favorites of, like, use cases that you think are tricky for LLMs to do well?
I think like a big focus of ours is on hallucinations.
So always kind of like checking out hallucination and like conflicting instructions, etc.
As one.
Tirst responses is another.
You know, like how well is it?
Like not, you know, you ask it a question and here's this 10 point list and you know, very, very verbose.
Do you have a terse response as a validator?
Yeah.
Well, we don't have it.
Like, we don't have it publicly.
But like we do kind of like check it.
Yeah.
So I think like those are kind of some of the things.
There was one, there's one example in one of the close door sessions where the,
The only answers were two tours.
Yeah, yeah.
Where I think everyone would laugh when they were like,
can you write a block host about this?
And the guy and the GPT said,
sure, I'll do it tomorrow.
Yeah.
Yeah, yeah, yeah.
Yeah.
I think like those are, I think those are,
I'm really, really excited about,
I'm really, really excited about JSON generation.
I'm actually kind of surprised to see how long it took them to get,
like, they're probably just doing constrained decoding under the hood, right?
Like, constrained generation.
Okay.
Because they're now saying that guaranteed correct JSON rather than, you know,
more correct.
Do you get what I'm saying?
I was parsing through their words.
They've never had an issue
producing JSON.
It's just that sometimes
it doesn't fit the JSON schema.
Right?
Am I wrong?
You would know more than me.
No, I think there are also issues
with like producing.
I think the obvious thing is like
unbalanced brackets.
When it's on context land,
I think that's like an obvious thing, right?
But like weird things
when you have like really long strings
then quotes, etc.,
become kind of weird.
So I think those are some other ones.
Schema is obviously kind of challenging.
etc.
Yeah.
I think there are,
even with function calling,
like function calling,
at least I haven't played around
with it yet today,
but previous generations
of function calling
wouldn't guarantee that your schema is matched,
which would be an issue.
And I think they're still not guaranteeing it
because I kept waiting for them to say it.
I haven't read any of the public docs or anything.
Do you know if they're guaranteeing
that it fits a schema or they're like...
Oh, that's a good question.
Yeah, that's a good point.
They never said they guarantee.
Yeah, they never said they got...
They guaranteed correct JSON.
They didn't guarantee if the JSON
matches the schema.
Okay, you can call JSON loads.
Yeah, yeah, yeah.
I'm very curious to see, like, once again,
if this is a pattern that, you know,
all of the other foundation model providers adopt,
and I don't see why not, right?
Like, I think for them to kind of, like,
own specific decoding models is going to, like,
make a lot of sense compared to, you know, like,
yeah, a lot of the, a lot of the hacky stuff.
Yeah, cool.
Any other favorites, you know,
doesn't have to be guard, guard,les related.
Any favorite conversations, favorite demos, favorite?
I, oh, the GPT.
and the assistants.
You want to make one for yourself?
I do want to make one for myself.
It doesn't add, like, yeah,
not very Godreels related.
I do want to kind of play around
with like how well it works with like some of the things we track.
But yeah,
it was just so fascinating to see the marketplace.
I am very, very curious to see,
you know,
what the marketplace looks like.
Like,
are people going to have like really,
really vertically specialized things on the marketplace?
Like,
if you have a generic, you know,
sales assistant or something, right?
Like how much, or SQL generator,
how much, how popular does that become
versus like sales assistant for X vertical at Y stage of the sales process.
Oh my God.
Do you know what I mean?
Like it's so easy to do this now.
Yeah.
That like where at what level of specialization do you need to be to kind of start seeing the results?
And that is one thing I'm very excited to see like how that pans out.
It scares me a little bit because it's basically, they said the future programming is natural language or something like that.
Yeah.
And that's great.
But like it really is a new platform, a new operating system almost that they're creating.
And I don't know how to position myself.
Not that I have to, because my world is very developer-oriented.
Yeah, yeah, yeah.
But this is a whole no-cold world that you and I don't touch.
Yeah, yeah, yeah, yeah.
Whoa.
Yeah, yeah.
I really want to see, like, is there just going to be like assistance for everything?
Like what I'm generally curious to see the impact of this on knowledge work, you know, which, yeah, like how much of my work.
Like, if I'm getting annoyed by something, is my first instinct going to be like, you know, let me just, you know, spend the five minutes.
is to build an assistant for this?
Like, is that how everybody's now going to start thinking?
You know, and that's one thing I kind of really want to see.
Yeah, that's exciting.
Yeah.
Okay.
Last question, you spoke at AI Engers Summit.
Let's advertise your talk a little bit and point people to your talk.
Yeah.
Yeah.
Yeah, so thank you again for inviting me to the AI Engineer Summit.
One of my favorite conferences that I've attended, you know, this year.
My talk was about the new paradigms for working with large language models, you know,
for building really production-ready applications when the technology that you're working
with is underneath all of it, you know, non-deterministic.
Really fascinating thing, which was the open AIs talk about building production grade applications,
talked about how essential it was to build Godreels as a way to make it to product.
You're talking about the one from today.
Yes, the one from today.
Which people haven't seen yet, but really, really cool talk.
So I think it really validates what we've been saying pretty much since the beginning of the year,
which is that you'll get like a certain, you know, you'll get to a certain point,
but at that point, you need to start adding guardrails to your application.
If you need to get your users to start, you know, getting value out of what you build out, right?
So I have your chart and I have their chart.
They put guardrails at the first layer.
It's not at the end.
It's actually right at the beginning for a user experience.
Yeah, that's right.
Yeah.
Yeah, that was kind of interesting to see that they put it as part of the UX.
I'm still kind of very candidly.
I'm still kind of digesting that.
Like I think of it as I think of it as part of the infrastructure.
and I don't know if it's as much UX as it is, you know,
just like one of the components that you need in your stack.
But I think the pat, like a lot of what they said today,
completely validated, you know, what we've felt for the longest time.
And also what I go really in depth about, like in the talk that I gave, right?
Which is that what happens once you have the bare bones application,
ready, what is the process of actually adding guardrails for what you care about?
Like, what does that look like?
Yeah.
You know, what are the risks that you care about?
How do you verify that those risks are happening or not?
not happening, if they are happening, how do you quantify them, and then how do you mitigate them?
That was what the talk was about, which I would really recommend people go and check out.
Awesome. Well, you did a great job. We're going to post the talk soon. And thanks. It's good to see you
again. Thanks again for inviting me. And that was about all I managed to get before the after party.
At the after party, there was actually an after after party thrown by news research. So let's hear a
little bit about open AI versus open source AI from Alex Volkov. Okay, so we are in the one day
after Dev Day here with Alex.
Hey.
Hey.
Hey.
Very, very recognizable voice right now.
We don't have to introduce you.
Hey, everyone.
And we're here to talk about the two parties that happened yesterday.
There was one official Dev Day open AI after party where I interviewed Shrea, who's just before
this.
And then there's an unofficial one for keeping AI open by news.
Yeah.
So what was it like, just compare and contrast?
So let me maybe start with like who news research is.
Oh, yeah.
Yeah.
Most people haven't heard of this.
It's written N-O-U-S, or I mispronounced it,
multiple times, like it's news research.
It's one of the few organizations online that started like from a Discord and then like kept going up until like a significant amount of people are like working with them, affiliated with them, of folks who take open source model to its most extreme capability.
So collect data sets from open source, open source and more close source and depending on that they released like with different licenses.
And then they find to an open source models that were like released to us from like Lama for example and mistral, which is a French company that recently released a seven
and B model that's the best. And they've been doing this since Lama 1, but recently it really kicked
into high gear with Lama 2 releases because Lama 2 ended up being with a commercial license. So you could
actually use this for actual products and services. And Mistral came out with like a full Apache 2 license
with a BitTorrent link. I think you remember that. And so these organizations suddenly became like
a very very important currency in the in the world of like where the whole world of AI is going
because they're lining local models.
And many companies love OpenEI,
but either cannot afford this
or cannot risk the chance
that OpenEI changes something
like we saw with Dev Day.
And so many people are turning on to like,
okay, if we want to run our own hardware,
how do we actually do this?
And you can run it,
you can run a Lama 2 and all these models
on your own hardware,
but then you want to fine tune them
for your own purposes.
And so how do you actually fine tune?
And now organization like news research
was probably the biggest one,
alignment labs, shout out to Austin
and folks from alignment labs,
skunk works and many of,
of these people come up and say, hey, we have the know-how.
And we only started learning about this like eight months ago, six months ago themselves.
But now they're like the specialized more people that find two models and actually
release the best kind of models on the Hagenface open source leaderboard.
Yeah.
Yeah.
And in my knowledge, all two models that I keep hearing about, one is Hermes.
And they recently switched the base model for Hermes from Lama to Mistral because apparently
it's better.
Yeah.
Hermes is like an instruction dataset, 900,000 instruction.
I don't really know where it's from.
Maybe I don't want to know.
They also do some, like, fun models.
There's, like, a mystical model that they do.
Trismestus, yeah.
Some stuff like that.
I think it's actually a little bit weird
that they keep releasing models.
Like, they release three models a week.
It's insane.
Right?
And it's very hard to keep up.
Like, I'm like, okay,
which one is actually the one
that I should pay attention to?
Yeah.
So, first of all, you're welcome to join Thursday.
And then we talk about all the models every week.
It's kind of interesting to that.
If I do, like, a recap for a month,
the beginning of the month,
most of the updates don't matter because like every month.
I'm doing monthly and I feel this like.
Every month or a month.
I'm doing this for historical posterity.
Like five years from now, people want to look back.
Then they can look at my notes because I only have 12 a year.
Yeah, nobody's going to look at your notes.
They're going to have a GPT train or your notes answering everything.
I have, yeah, I'm doing like every week.
And every week we're talking about like this model outperforms that model like significantly.
And we're noticing significant changes from week to week.
Literally in the spend of a month, we went from a 30.
three billion parameter model, which is big. And parameter count is not everything there is, right?
You can have a smaller model with like larger, longer training that actually will perform better
than whatever. But we're noticing smaller and smaller models doing outperforming bigger ones
significantly. Zephyr from Hagenface outperformed Lama 70B and Zephyr is like only like
a 7B model. On some things. On some things for sure. And so this is very interesting because like
it's really hard to evaluate. Evaluation frameworks are bad. Everybody's saying they're not representing
of anything. People can fine tune and overturn on them. And so this is very interesting. This is a very important. And so this
There's this whole kind of subculture of open source, mostly on Discord, some of them on X and Twitter spaces.
And for some reason, but I find it very humbling and incredible.
They also hung out on Thursday.
And so that's how I got to this.
That's how I got to meet like news research folks, Ticknew, Imozilla, and organized the counter party event last night, together with some other EAC people that we know from Twitter as well.
Including Mark Hercid.
So apparently he was supposed to, I didn't see him.
I saw a photo with a bald head of a big guy,
so I was like, is that Mark?
I don't know.
Anyway, but the opening eye party was at an art museum,
and then the news research party was at a club.
It was a club, yes.
Folsom Street in San Francisco, a club.
10, 15, Folsom, I think.
Opening eye was a very high-brow, buttoned up event.
There was a live band, someone playing jazz.
Yeah.
Which I think I mentioned this once.
It was too loud.
We want to talk.
We don't want to listen to music.
No, no, no.
We're just old.
Everything is too loud.
And then it was like a lot of people, a lot of networking, a lot of people trying to get together, maybe do business together.
And like very, very awesome.
Many people from Open AI actually showed up.
A lot of people.
We stood in line.
There was a long line for the Magnometer to step in.
And then everybody, like, passing us around was like Open AI employee that passing like straight through.
Yeah.
And then that ended around eight, which is like the standard San Francisco like buttoned up in.
Oh, yeah.
That's when you go to bed.
That's where you go to bed.
And that's when the other party.
kind of started and I think they just seized the opportunity because everybody's in town for the
open AI stuff. Why not make a splash and announcement for like for open sourcing AI. So literally
the invite was keep AI free.com yeah which was the website and the API open open.com and you
had to register you had to go in there and this was to me an incredible kind of show of Twitter in
real life. So all of the folks who follow Mark Andreessen, he received
and they stepped into this thing with like the techno optimism stuff.
He started to boost the E effective accelerism,
EAC folks.
And so there's a lot of like signature stuff
from that like ecosystem on Twitter.
There's like don't thread on me with like you don't take away my GPUs.
There's like all these signs across the club.
The it's a very visual club as well.
So where the DJs is a whole like a 3D projected thing.
So there's like a bunch of like art and like live things about keep AI open.
I found it like very, very super cool.
I'll have to tell you a tidbit.
I saw me and Killian were there from Open Interpreter.
We saw two people with lap codes.
I was like, what's the deal with lap codes?
So we went and asked.
And they just said, hey, we just like,
we came back from our work where we work on semiconductors.
We were actually like touching chips, whatever,
which is like didn't change out of it.
And my head was like so incredible in the keep AI open GPU kind of poor party.
We have people who literally work on superconductors
came from the work on like they're working on chips.
Semiconductors or superconduble, very different things?
I think semiconductors.
Yeah.
We had the Superconductor episode a while back.
I think people were still recovering.
I'm personally still recovering from that.
There was a whole thing for me.
So is news research like vibes?
You know, like what is the mission apart from to keep publishing open source models?
I think you'll have to get some news people to actually speak like about the mission,
about the actual product.
But as far as I understand this, no matter how much the product site will be and there might be,
There's so many people they're doing like so incredible stuff that people notice like, you know.
So no matter how like how much of the the business side will be, they're like committed to fully open source as much as possible, including data sets, including models that are like trismestous, for example, their model that's like trained on the occult and the physical and metaphysical.
You can't expect open eye to let you talk with a model to answer with like these mystical questions, mystical stuff.
astrology, Halloween.
So you're very like easy.
into the astrology in Halloween.
They're talking about, like,
you can ask this model about like resurrection.
And stuff, right?
Like all of the occult, like craziness
that they've collected,
open I will not let you do that.
And so there's a thing that I will not let you do
by default because they have lawyers
and they were doing sued.
Yeah.
Recently they announced the protection shield thing,
so you won't get sued because of their model.
So they're damn, entropic, all these big companies.
It's very important for them to protect the outputs
and the models.
Here, these folks are like, hey, if you want to build a model,
fine-tune this.
We're going to teach you how,
jump on our Discord.
we're going to help you with producing the biggest models.
And then if, you know, there's going to be like financial aspects as well.
If your company that wants to run this, we'll also help you do that.
Yeah.
So it's the same as stability, basically.
That's from what it's from talking to him.
That's what I gather.
Yeah.
Cool.
Anything else that people should know about the party news?
I found this whole day to be like a very singular AI day.
And we don't get many of this.
GPT4, I think was the biggest one previously.
Yeah, March.
Like a single March 14th.
That's what Thursday I started.
We started talking about this every week.
This was a singular day in San Francisco.
This, like, started pregame party with Swiss and some other folks that I got to feel like
a little bit of San Francisco.
And then Dev Day was incredible.
We just heard from Simon.
There was like a garage that they made into a venue event.
Yes.
Probably a custom venue event on the fly, which like just talks to how much they can pull off.
It felt to me that like this Dev Day event and then the following party, it felt a little
bit like almost like an Apple thing.
where like it's going to be a yearly thing
that people will like try to get in as much as possible.
One thing to note that in the other party
there were many people who didn't get in to this party.
And so, you know, they were watching for like a party.
Yeah, this office right here.
This office people watched here.
And people watched in the life space that we...
8,000 people tuned in to our spaces.
8,000 people tuned in.
I didn't even have a chance to look at it.
I always want to know the number.
Oh, it shows the relative level of interest.
And, you know, like, so codens 22,000.
and this is 8,000.
Just relative interest by developers.
There's like two spaces as well.
Robert Scobarra.
He stole the thunder a little bit.
Stolled some audience for us.
Shout out of Robert.
And I think that it was a singular day.
And I think the newsresteris keep open source, open,
EAC, Mark and Driesen, like all these things together,
also added to the top of this because like it happened in the same day,
one on top of another, in the same place, San Francisco.
I find it incredible.
I would definitely come back next year to it.
Yeah.
Okay.
Yeah.
Well, I think you'll be back sooner than that.
Yeah, probably.
There'll be other things going on.
All right, thanks.
Awesome.
All right, last but not least, we go back all the way to the Newton where I started this podcast,
where we checked in with Rahul Somwaka, better known as Rahul Ligma, who just celebrated his one-year
anniversary as one of the biggest memes and celebrities in San Francisco.
But by day, he's also the CEO and co-founder of Julius A.I.
And I'll match it up.
What's up, Swicks.
Hey, good to see you.
It is one day after a deaf day.
and we all had a chance to process.
How do you feel?
What's your top takes?
That would have awesome.
We got to see a bunch of really smart people
who are building cool things with OpenAI, GBT, Dolly.
The event was very well to put together.
The keynote was awesome.
The energy in the room was crazy.
And I could see real-time social media firing up
with all these takes.
Overall, I think it was a good day.
Yeah, I interviewed Suria Dantaluri.
Yeah, I think you know him.
He was like, Samud just killed my startup.
And he was,
was almost true for him because he has a bunch of plugins.
Yeah.
And plugins are kind of deprecated.
Yeah.
Yeah.
Yeah.
Yeah.
The plugin thing was interesting because it was, it's going to be deprecated, but they just
accidentally turned it off yesterday.
Yeah.
So he freaked out of it.
It freaked out.
And then like they bought it back up.
Yeah.
Yeah.
So top features that you're interested in that you want to explore more.
I think people are super psyched about the assistance API.
But personally, if you ask me,
Two things that I am most excited about is turbo.
Yeah.
The speed is crazy.
And have you actually, have you measured, do you know any like rough measured?
Because I don't think they actually ever mentioned the speed relative difference.
I started noticing the speed difference in chat GPD actually like a few weeks ago.
Oh, I see.
So they already slowly eased this into it.
Yeah.
Yeah.
And I saw like takes on Twitter that did anyone notice chat GDPD get much faster and I noticed it too?
Yeah.
But so it's turbo.
was exciting, but the second thing that's exciting is multiple function calling and then the
JSON output formatting. I think as developers are building on the dev API, so that's the thing
that's super exciting to me. Of course, as version stuff, there's code interpreter as a tool
in the API. But I think what will bring the most applications is actually the speed, because
there are so many things. If you look at our number.
numbers on Julius. People are not patient. They want an answer and I want to answer quick.
And we see clearly if you can get that answer to them a few seconds faster, there's a clear
difference in the conversion. So speed is going to be paid. What is conversion for you? Is that just
paying or? Oh, no, just like from first message to second message. I see. So we we do code gen and then
we run the code and then the code has an output. They use it as a second message and we can just see
the funnel.
Yeah.
Where if it's faster,
the code runs faster.
And the second thing is
multiple function calling.
I think you're basically
telling the AI that,
so I think the people
misunderstand function calling
it's essentially to use.
And if you can tell the AI,
hey, you can give me
multiple tools to use
at once.
Yeah.
I think that's going to
unlock different applications
than before,
because before it was just like,
okay, this is a task.
Tell me one tool
and was the input for it.
Yeah.
But if the AI can now use multiple tools in parallel, you can first of all have more specialized tools.
And then the AI more specialized instructions for each tool.
Yeah.
It's just going to unlock a lot of cool applications that previously weren't possible.
There was a practical limit in the number of tools that you can give it, right?
So we had this discussion in March, February, March, April, when they released the function API, that this is subject to context window, the JSON schema itself.
Yeah.
Does that change at all?
or I don't know if you know.
I don't.
Yeah.
But what I noticed, though, even before,
was that more functions and more options just confused it.
And that's what I want to play with next is like, okay, where's the breaking point?
They see.
Like, does more options, you know, confuse it?
Does it make it?
Would you use multiple function calls as well?
Oh, totally.
Is that just theoretical?
No, no.
I have a direct application for it right now.
One of them is oftentimes GGB rights code.
And then we run that code.
realize that, oh, from GPD's last knowledge update, that module in Python has changed.
It has new function, new APIs.
So today, the way we do it is when the error happens, we tell GPD, okay, you can go look
up new documentation and then fix that error.
But with multiple function calling, the way we would do it, is like, give me the code,
but then also give me a documentation look up.
And then when the error happens, I can just quickly fix that without another GPD call.
Yeah.
And then keep moving.
Nice.
But I mean, in general, it's just like multiple to use to me.
It's just so exciting as a developer.
And I wish people were talking more about this.
Yeah, I mean, people are still coming to terms of just like the base model and prompt engineering and all that.
That's still important.
But for engineers, I think you should explore these other advanced features.
True.
Yeah.
Anything on the multimodality side that you're interested in.
I mean, originally will be super interesting for sure.
And we have this functionality in Julius right now.
you can generate React and
each team opponents.
Like V0.
I think Matt was showing me
a little bit of that demo.
Yeah, yeah.
We've been hacking on it a lot.
I think the missing piece here is that
well, you have an engineer who knows how to react
and they probably wouldn't find this useful.
But if I can allow
like anyone in the world to just draw a mockup
on a piece of paper and then run that
and have the vision.
Yeah.
Yeah.
Yeah.
Turn into like actual components.
is I could use on a webpage, that'll be sick.
And what's even more sick is like, have the feedback,
where you take a screenshot of the page generated
and then feed that screenshot back into vision
and then come up with more instruction
and have that loop.
Wow, like a self-improving web page.
Isn't that crazy?
Yeah, yeah.
I'm so excited.
Yeah, yeah.
So in my mind, Julius is very data-focused.
By the way, I didn't introduce you.
I was just going to do it separately.
Yeah.
But people know who you are.
Yeah.
You have a Wikipedia page.
Yeah.
You just passed your one-year anniversary as Rahul Lingma.
Thank you.
By the way, any fun things happen on the anniversary?
What are the fun thing?
Ilya said, Ilya recognized you on the spot.
Oh, Ilya was like, oh, my God, this is, oh, you're famous or whatever.
And, no, these guys are so awesome.
Like, they're so humble.
But any happen on the one-year anniversary, nothing really, like, it's, I mean,
you knew about it a week before.
I like to set anniversary dates.
That's awesome.
Because it reminds people of the passage of time.
Like, it's like, wow, shit.
Has that been a year?
Yeah.
And then you're like, I think it motivates me, it motivates me more than like memento Mori.
Like, yeah, it's case, you know, sometimes you're out of date.
But it reminds me to spend my years wisely to do interesting things with the time that I have.
Momentumory is kind of depressing, whereas this is like, oh yeah, did you know one year ago?
Wow, it's been a year.
Yeah.
Okay, cool.
But Julius, you data analysis chat thing.
Yeah.
Basically, code interpreter plus plus is how I think about it.
Exactly.
And also you just across the 100,000 users.
Yep.
You have delivery modes across,
your plugin as well as a chat box,
like a dedicated web app.
Yep.
Okay.
Anything else that people should know?
Well,
our vision is, you know,
writing code is super fundamental to doing things.
You could not only automate a bunch of tasks in your life,
but just writing code,
but also it's how you just like interact with the universe.
Right?
You have code that brings you a Vamo car
and picks you up,
drops you out somewhere. And I think
allowing these language models to write code and
do things for you is really powerful. And
data announces this application
that we're most excited about right now because
that's what it's good at immediately.
But just on Friday, we launched FFMPEC
support. And there were people
trying to upload videos, turn the
videos into JIFs, or like take a
YouTube video turn it into a, you know,
short summary and all these different
use cases that we didn't
truly like hard code into Julius. We just
told it, hey, now you can
run ffmpeg and you can run idlp and movie pie and all these different things do these tasks for me
and then people were just like organically discovering those things there's this guy TDM on on on twitter
cTO junior and he took some meme video and put it on my own tweet over later on my own tweet and then
like tweeted that and then that got a bunch of likes and i was like dude like this is the first one that
gets a lot of likes on and you know fmpeg on julius so that has a lot of meme potential it's a lot of
potential, but that's not what we're going for.
Yeah.
It's just like letting people like do things.
Your target market is like the S&D, the enterprise.
It's actually individuals who have data on the hand and they just want to drum academics.
A lot of academics, actually.
Yeah, a lot of academics, a lot of students, researchers, any kind of CSVXL data, you can just dump into Julius and then have it analyze for you.
We have this video coming out in a few days where you can now actually train a nano-GPD on Julius.
So you can give it, hey, here's the good update you for Carpeti.
So you have GPUs to train it on or is just training CPU?
CPU, because it takes a minute.
Yeah, yeah, yeah.
That's true, that.
Yeah.
I mean, Copathia will like that.
Yeah, yeah, yeah.
Okay, cool.
So the thing I really want to sort of ask you as a founder on is, you know,
I think there's always this existential threat about opening eye building your features, right?
Yeah.
In a way, so like the number two default bot in the, in the GPT app store.
Yeah.
Is data analysis.
Yeah.
And people can build their own by,
customizing and adding code interpreter.
Although I think there's also opportunities for you.
So on the roadmap that they presented in the closed session,
they also said you can bring your own code interpreter.
Yeah.
So how are you thinking about that?
I mean, as a founder or as?
Founder.
So who's the audience?
Is it like other founders or is it?
Yeah.
Other founders?
And then people are just interested in how you are processing this.
Yeah.
I mean, I think it's a very interesting story of processing this live because the news just
dropped yesterday.
Yeah.
Totally.
Well, so the story behind Julius is that we actually launched Julius three months after
Code Interpreter was announced and a few weeks after it was rolled out to everyone else in the world.
So we were number two.
And even then we got 100,000 users because I think there's a lot of work to do to get something to work properly.
And there's a bunch of examples of this on the Internet.
So if I'm talking to founders, what I'll tell them is, man, so many people give up.
before even getting started.
And that happens.
Don't do that.
Sure, you can change your idea.
You can find new things to work on.
But the way I'm processing is that,
wait, we were actually,
we launched after a code interpreter came out.
And there's 100,000 people who think Julius is better than code interpreter.
Or just tried it out.
Yeah.
I'll try it out.
And use it over code interpreter.
And there's like a lot of work to do.
Like, for example, the FFMbeck stuff we launched on Friday.
Or the HTML stuff.
Or React component stuff.
stuff, all these different things.
To get them to work, it takes some effort.
How I'm processing it, I mean, you know, that's like, that's what startups are all about.
It's like risk, right?
If you want to build a risk-free startup, you probably don't want to work on startups.
Yeah, just go get a job.
Just go get a job.
Exactly.
So I'm having so much fun.
The way I'm thinking about this is like, whoa, there's all these new different things I could do now.
I could build.
That's so exciting to me.
And I'm pumped.
Yeah.
Awesome.
That's it.
Any last words, call to action?
Call to action.
Let's go build some cool things and get a bunch of users.
Let's do it, guys.
All right.
Awesome.
Thanks so much.
Thanks, Wix.
I think that's a meme that we can all get behind.
Let's go build things for a bunch of users with AI.
