Limitless: An AI Podcast - The Government Banned GPT-5.6. OpenAI Released it Anyways

Episode Date: July 15, 2026

🌌 LIMITLESS HQ ⬇️EMAIL US:           info@limitless.fmNEWSLETTER:    https://limitlessft.substack.com/FOLLOW ON X:   https://x.com/LimitlessFTSPOTIFY:             https...://open.spotify.com/show/5oV29YUL8AzzwXkxEXlRMQAPPLE:                 https://podcasts.apple.com/us/podcast/limitless-podcast/id1813210890RSS FEED:           https://limitlessft.substack.com/------OpenAI has released GPT-5.6 with the new Sol, Terra, and Luna model tiers and an updated app experience. We also cover live demos, image generation and editing use cases, and concerns about autonomous access after reports of file deletion and database issues.------TIMESTAMPS0:00 GPT 5.6 Released to Everyone0:24 Safer Than It Looks?0:44 Three Models, Three Tiers2:28 Magic Bagel Game Demo4:40 Bunker Simulator and 3D Art6:50 Limitless Blender Intro8:39 Manhattan in Voxels10:54 Outfit Catalog from Photos14:01 Deletion Scare and Safety15:35 Choosing Between Fable and GPT18:49 Too Many Modes and Apps21:12 Final Verdict and Wrap-Up------RESOURCESJosh: https://x.com/JoshKaleEjaaz: https://x.com/cryptopunk7213------Not financial or tax advice. See our investment disclosures here:https://www.bankless.com/disclosures⁠Josh works with Anthropic as a contractor. All views expressed are his own and do not represent Anthropic, its leadership, or its affiliates. Nothing in this episode is investment advice.

Transcript
Discussion (0)
Starting point is 00:00:00 Three weeks ago, Open AI had an AI model that you couldn't access. The government completely banned it for being too dangerous. Then, just last week, they released it for everyone to use. GPD 5.6 is available to everyone. And the first thing it did was delete someone's entire desktop file. Someone gave it autonomous access to do a small task, and it did a cleanup, which ended up deleting everyone's code, which begs the question, is this model still safe enough to use? Now, putting this catastrophe aside, I want to know how GPD 5.6 compares to Fable,
Starting point is 00:00:30 The truth is, it's a really good reasoning model. It's fantastic at coding, but most importantly, it's cheaper to the tune of 50% cheaper than Fable 5. But the question on everyone's mind is, is it as good? The truth is, it's good at some things, and it's kind of terrible at other things. If it's long, agenetic work, it's fantastic. But if it's high-quality code, it's less so. So there are three different models here that we're going to talk about on today's episode. Sol, which is the premium tier, Terra, which is the mid, daily use tier,
Starting point is 00:00:58 and then Luna, which is the cheapest and fastest model. And we're going to put it to the test live on this show with three different demos. Yeah, I think that's the idea for this episode. As everyone kind of knows about GBT 5.6. Chances are, if you're listening to this, you're probably in the know enough to have been playing around with it. You've been using it in your chat GPT app or maybe your Codex app, which is now the chat GPT app.
Starting point is 00:01:17 There's a big release. There's a lot of things that happened around it. But I think the part that we're most excited to talk about is the demos, is the actual practical applications of what you could do. We've spent the last week kind of playing around with it, testing it out, seeing where those edge cases are on what the models are capable of, is that way we could come to you and give you an idea of the types of things that you can go off and try with GPT 5.6, Seoul, Terra, and Luna.
Starting point is 00:01:39 Now, Seoul is the flagship. Most of these demos are going to come from Seoul. We kind of wanted to test the best. We wanted to see what the model is capable of and what you can actually pull out of this thing. Is it fable level in terms of production, in terms of game design, in terms of all the demos that we like to do? So that's what we're going to go through this episode. It's just kind of testing things out, seeing where the edge cases are, and seeing how this model lands amongst the landscape of other tools that are at your disposal.
Starting point is 00:02:03 So, EJS, I have to ask first, are we looking at the GPD app, chat GPD app or the Codex app? Because I know they're separate now and the Codex app became ChatGPT. Well, you're actually looking at all of them. Recently, Open AI combined their chatbot interface with their coding interface with a series of other features into one super app. And that's what we're looking at on our screen today. I had to re-download an entire desktop app to get access to this thing. TBD on whether this is actually a good move,
Starting point is 00:02:31 but let's work through maybe some of the demos. Now, the first one that I have on my screen here is I asked Sol 5.6. Like this, I've been in New York for 10 years. I want to make a game. I love games. I've been playing games since I was a kid. I would like you to build a 3D explorer game where I can fly around Manhattan on a magical flying bagel
Starting point is 00:02:51 and collect various ingredients that go into said bagel. So it thought about it. And, you know, privy to it, it did this in one shot. So are you ready for this, Josh? Are you ready to see? Let's fly around Manhattan. So this is the loading screen. It's called Ride the Magic Bagel, very creative thinking over there.
Starting point is 00:03:10 And it says, take flight to start this game. Now, I noticed like there's a few little discrepancies on the design, but let's actually play the game. Okay, so this is me. I'm on a bagel. It doesn't look like me. I'm going through random assortments of circles. The buildings around me, Josh, are supposedly. meant to be Manhattan.
Starting point is 00:03:29 And if I press spacebar, I can use my cream cheese boost, as you can see on the left. Yeah. So it got very creative. If you notice, the physics is a lot different. So on previous demos that were done
Starting point is 00:03:42 using Fable or GPT, we have always been playing around with some kind of a 2D interface. I'm not a 3D graphics designer, but I know that if you spend a lot more time with this model, you'll be able to create something that is pretty AAA rated
Starting point is 00:03:55 or like close to. so that kinds of quality. I'm not talented enough to do that. But the fact that I did this and conjure this up in 10 minutes is pretty awesome. And it did some of the design and thinking around the game mechanics itself. Like if I crash into the ground, it's kind of like, okay, it's registering the spikes. And when I collect power-ups, it like gives me a boost up for whatever I want to do. Okay.
Starting point is 00:04:16 So I like how it says the everything portal. I'm assuming that that's an ode to the everything bagel, which is nice and tasteful. The New York City part I don't really see. So I guess in terms of like a one-shot demo, it does. okay, this is a fun demo. If you told me this was from GPT 5.5, I probably would have believed you. I'm not sure there's anything
Starting point is 00:04:32 exceptionally different about this particular demo. But there are some interesting demos that maybe you could not have done before that you can do now that we haven't tried. Okay, so for demo two, one thing that Open Air is kind of known for
Starting point is 00:04:45 is creating really aesthetically pleasing images from scratch. GPT Images 2 is one of my favorite image models to use. So I got it to create a very detailed floor plan of a aesthetically pleasing nuclear bunker that would be based in Denver, Colorado. Completely random, popped it out of my head, and it came up with this thing that you're seeing on your screen right here.
Starting point is 00:05:07 So it has an airlock entry. You can see it's quite central base for the lounge. You can see there's multiple rooms. There's a pantry, cold storage, general storage, et cetera, et cetera. And then I said, okay, now I want you to create a 3D interactive rendering of this building. Build it as a simulator, so it should allow me to walk around, enter it, enter different rooms,
Starting point is 00:05:25 and there should be signboards, basically, for me to see. So on the screen here, it created that visual. So I'm in the bunker, and I should be able to look around, and you'll see, there's the dining room. There is the kitchen. I can like go into the kitchen. A little dark. It's a little dark.
Starting point is 00:05:41 It's definitely a little dark. It's almost like too dark. But they've kind of like, they've nailed the kind of like central sphere side of things. I like that they have a map bottom left. It's very call of duty, dare I say. Nice little hood. It looks very futuristic. There's no remnants on whether this is in Denver, Colorado,
Starting point is 00:05:58 but I guess that is the entire point of a nuclear bunker. It's meant to be sealed. There's some kind of like locker system here as well. Again, in my opinion, it's very basic. And I wonder how much of that is just because I'm not a 3D artist or visual graphic designer. And I'm sure that if someone more talented than me had a little more time, more than 10 minutes, they'll be able to create something way more impressive. But it's cool to just one shot.
Starting point is 00:06:22 I think this is the dining hole. Everyone sits down. Yeah, I'm starting to see the kind of signature style of 5.6 just from this, which is, like, I've noticed both of these have similar color palettes. They both use green and purple. They both prefer, like, kind of shinier metallic objects. They both are very geometric and sharp edges. In a way, they're very GPT-esque, where it's kind of like very sleek and modern.
Starting point is 00:06:43 Like, it kind of looks what you would imagine the model to look like based on the design language of Open AI. This is a Minecraft walls. Yeah. Yeah, like this is interesting. I mean, it's cool. It's like cool that you could build this in one shot, but you have a third demo as well, right? What is number three? The last one is probably, I had such high hopes for it because this model is apparently amazing at visual editing, image generation, but also video creation from scratch.
Starting point is 00:07:09 So I attached it to Blender, which is a popular tool that you can use to kind of render and create videos. And I said, hey, we have this limitless podcast. It's one of the best AI podcasts in the world right. now. I want you to create an intro with me where there is a microphone, represent the brand itself, it's limitless, and throw in the logo in there as well. I ended up coming up with this. Now, what you see in your screen in the center is supposedly meant to be a microphone that looks more like an egg on top of a pedestal. You have limitless represented by the infinity symbol over here. And then finally at the bottom, the fadeaway is the limitless logo at the bottom. Now, creating this
Starting point is 00:07:48 subscribe took me like two minutes. I'm sure I could have spent more time on it and come up with a better graphic rendering, but I really wanted to give you guys an idea of what you can do right now on your desktop with zero experience. Okay, this is, I mean, it's cool. It's a very high quality blender element. I guess you could say. It's like it looks good. The lighting is cool. It looks professional. It just doesn't quite make sense. And I guess that's where we kind of are where the model has reasoning but lacks the nuance to really piece these things together in a way that a human would have. So it's getting good, perhaps better, but not great. Like, none of these demos were truly exceptional, but there are some that I've seen on X that are like actually pretty
Starting point is 00:08:27 impressive. In fact, one of them, this guy actually wanted to make Manhattan, New York City, and the model did it. So, EJS, to be fair, you only one shot at that prompt. You didn't give it an entire week like this person claims that they did. And over that week-long period, because as we know, there is backslash goal, which will allow the models to run for a very, very long time until it accomplishes a goal, it was actually able to generate this voxel-based Manhattan, which is basically just a low pixel count version of Manhattan. And I gotta say, it's, it's pretty good. It looks pretty good. It's definitely buggy. It's definitely glitch. You see the glitches happening. But you can also see that it is pretty geographically accurate when it comes to the buildings,
Starting point is 00:09:05 the topography, where all the parks are, what everything looks like. And this seems novel to me. I see this. I'm like, oh, that's kind of cool. Like this is a building block now for, there was a Spider-Man game that was based in New York City. Now you have the low-density pixel count version of Spider-Man, and you could build on top of this, and you can make more interesting things. And I think the games are always a fun way of testing these models because they kind of force you into this visual way of expressing them
Starting point is 00:09:29 that is easily accessible to anybody. It's like we can test it on code bases, but I'm not quite sure what the difference between 5.6 and 5.5.5 is on the edges of code, but you can see it in the visual outputs. And like, this is a pretty good demo. And there's a bunch of other features about this model, aside from just visual, like we have another example over here where someone actually way more talented than me gave it access to Blender and got it to render an entire highly detailed
Starting point is 00:09:53 visual of his MacBook and it did so, I mean, it looks pretty starting. It looks kind of like an Apple ad, dare I say, very high quality. But the other point is how quickly this model works. Now, the video I'm showing you on my screen right now looks like it's being sped up, but this is completely in real time. It gave access to someone. on's entire desktop and said, hey, I want you to build this 3D artifact for me. And it did so. It knew automatically how to use the tool and access the tool. And it's doing so in like a crazy amount of speed. Now, a lot of this is achieved because of Open AI's investment in this company called Cerebus. We covered this on an episode. It's got to be right. Yeah, an episode a few weeks ago,
Starting point is 00:10:33 which basically makes super fast chips for inference to the tune of 750 tokens per second, which is a pretty insane way. It just spits out prompts and outputs very, very quickly. Now, it is very expensive and it's not automatically available to the average retail user. You do need to get access to the API, but nevertheless, very impressive. Then there's this other one where people got really creative with their visual intelligence. Had someone uploaded their entire camera roll on their iPhone to GPT 5.6 and said, I want you to pull all the clothing items and accessories that I've worn over the years,
Starting point is 00:11:09 create a catalog for me and then create different combinations of outfits that maybe I could be wearing that I should be wearing that I haven't been. So it's a pretty awesome use case for this. This is my favorite demo because it relies, I'm sure, largely on the image generation model. And when I think about the things that are strongest when it comes to Open Aion chat GPT, I think of their image model, I think of their voice model. Those two are pretty exceptional. And this very clearly leans into the image model in that you can just feed it an entire catalog of pictures of you. It will extract it. It will assume what the rest of the clothing article will look like. And then I'll place it into a closet where you can kind of customize and mix and match your
Starting point is 00:11:44 clothes. And I think this is such a fun interactive use case. This would have been a like multi-million startup a like not too long ago where someone would have paid a good bit of money to download this app. They would have paid $20 a month for the subscription. Now you can just generate this in a few prompts on GPT 5.6 for probably the pro plan. I'm guessing like $100 a month. You could do this. And it's really, really cool. I think this of all the demos was my favorite because it really showcases the strong suits of GPT and 5.6 was able to crush this. Now, there is a demo in particular that we must cover because this demo, when I saw earlier when we were prepping for this podcast, I couldn't believe it actually happened in that GPT 5.6 sole, the big one, the smart one, it just accidentally
Starting point is 00:12:28 deleted all of this dude's files on his computer, which I was shocked by. It says, I caused a serious local data loss incident, a review subagents cleanup command expanded home incorrectly, and then ran this like command that killed a lot of the data that was on this guy's computer. This guy's name is Matt Schumer. He's involved with Grok. He's like a fairly prolific poster on X and he is one of two people that this actually happened to publicly at least. There was a second incident where GPT 5.6 deleted a lot of the code, except this time instead of on a local machine, it was an actual code base. And what he said is that, that GPT 5.6O just deleted my whole production database. That's it. Not a joke. This had never
Starting point is 00:13:11 happened to me before with any other model. Never. It's not safe. So this is like a little concerning. The fact that this has happened on multiple occasions to multiple people in varying degrees. One was a local machine. One was an entire database. You should be careful and use these models. They are very capable and we're giving them lots of access. But perhaps be careful with the access you give to these models because they can go ahead and actually do some kind of unrepairable damage. It turns out, you know, one of the fear-mongering things that the government was doing when they were banning Fable 5 5 and when they were preventing the release of GPD 5.6 was these models are too dangerous and, you know, put in the wrong hands, it can cause a lot of destruction. Now, given the examples that we've shown just now
Starting point is 00:13:50 aren't crazy feats of destruction, but removing your entire desktop, you know, you could lose important personal files, important personal pictures and stuff like that. That's what happened to Matt Schumer and there's no way of recovering it currently. But the other feature or demo that I've seen people use, which I'm honestly on the fence about, is people using GPT 5.6 to fine tune and in some cases train new AI models. Now, if you rewind literally about a month ago, Fable 5 was called out for its capability of doing exactly this. That's why they had to impose very strict guardrails on their model such that you couldn't do this type of thing. So it's very interesting for me to see this kind of like dichotomy between the two model providers and between these two models where
Starting point is 00:14:37 I guess 516 is getting a little more favorable treatment where they can still use this model, even though it's technically a quote unquote cybersecurity risk as deemed by the government itself to train and fine-tune other models. So what you're just seeing on your screen right now is the fact that someone says, it's an amazing researcher and you can use an entire prompt to get GPT 5.6 salt to post-trained 5.6 lunar. Now, granted, this isn't training a new model from scratch, but it's still involved in coercing a model to look very different from the existing model that you're using.
Starting point is 00:15:07 We have another example over here where someone trained a local model in a training pipeline from scratch end to end locally on his Mac. So technically that's a free model that can get access by anyone. It's kind of open source if you technically want to describe it as that.
Starting point is 00:15:20 So for the tinkers, for the hobbyists, for the builders out there that have always felt like building your own AI model has been out of reach and only kind of given to the expensive model labs. This might be a model
Starting point is 00:15:31 that you can somehow start to tinker with your own model and create something new. It's pretty cool. Yeah, so when do you use this? Like, as a user of ChatGPT, as if, let's say you have both subscriptions or let's say you're trying to choose one. Do you go with GPT? Do you go with Fable 5? I think the answer is probably dependent.
Starting point is 00:15:48 I know when, and we were talking about this before recording, that ChatGPT's membership goes a long way. If you pay even $20 a month, you can generate a lot of tokens through these models. Now, my understanding is that Seoul actually generates more tokens, each as you were mentioning this is that in order to accomplish the same goal, the tokens are cheaper, but it actually generates more of them to get to that goal. So it kind of offsets the costs more than you would imagine based on the paper cost per token outputs of these models. So that's something to keep note. But what I will say is that oftentimes they'll give you a lot of leash here. And if you actually want to build
Starting point is 00:16:20 really complex things, really long-form things, chat GPT is like pretty good at that. The allowances and the limits are like fairly high. If you are looking to do, I guess, more intellectual work, more planning. It feels like Fables is more of a high quality model. I know we still use that. We still prefer it. That is still the go-to model right now that we use to do the agenda prep. We used to like help build the artifacts. It just has this really just strong like subtle nuanced understanding of the world. And I find that it's very helpful for pretty much everything. And then for lower end tasks, you have Terra, you have Luna. Those are kind of comparable to perhaps opus and sonnet. So there is almost a one-to-one comparison. And I think a lot of it depends on
Starting point is 00:16:56 just kind of your general, the vibe you get from the models. A lot of bench marks now no longer work when it comes to helping me decide. I know that I have to actually get down there and test it and play around with it. And so far, I prefer the results of Fable. It feels like it's just generally smarter. It has this intuitive understanding that GPT 5.6 doesn't. But if you are looking to spend a lot of tokens and do a lot of work, chat GPD is going to take you a long way. I'm not entirely sure that this is Open AI's direct response to Fable 5. I think Sam even like alluded to it, that they're working on GPT6 and it should be released in under a month. So these training cycles are getting much, much quicker. You're basically getting three models for the price of one here.
Starting point is 00:17:35 So if you have a basic subscription tier, you have access to Sol, Terra, and Luna. And the kind of best way that I think about it is Sol is the kind of like Fabel 5-esque. It's their most powerful model to date. If you want to do hard work on complex tasks, use Sol. And then Terra is your day-to-date kind of model. It is kind of like the Opus 4.8 version model, if you want to compare it to Anthropic directly. And then there's Luna, which is kind of a model that is equivalent to haiku at Anthropics, so it's super cheap, it works super fast, and you can do it to do menial tasks that you don't really care about, kind of like the intelligence matter on that side.
Starting point is 00:18:11 Now, when it comes to cost, Sol is 50% of the cost of Fable, but as you mentioned earlier, Josh, it uses more tokens to think. So it's kind of like a way to cheat the metrics a little bit. Like it does more thinking, these models are known for spinning up a lot of different agents to do different biddings and works. Sometimes that annoys people because, like, you have agents to do. things that you've never asked it to do, but it just does. But on the flip side, it can do a lot of work for a longer time, a longer time horizon. And for tasks that you kind of want to just set
Starting point is 00:18:41 and forget and go to bed, you can do that. But with the adage that it might delete your entire production code base. So there are a lot of- Be careful with approvals. Please be extremely careful. And something really annoys me about this model released, Josh, which is, I have said time and again, I wish Open AI would just give me a model, maybe give me a few versions of it, and then leave me alone. But not only do we have three different models, but we have three different settings for three different models. So I'm just going to say this for the benefit of the audience's max mode, ultra mode, and terse mode are different versions of everything I just said for those different models, but for each individual model. So if you want the best of the best,
Starting point is 00:19:22 You use Sol in Max or Ultra mode for your most complex and hardest tasks. If you don't care about it and you want to use minimal costs, you want to use the Turs code mode for the cheapest model, which in this case would be Luna. Yeah, there was a weird rollout that happened here. It's like somewhat confusing in the sense that there's a bunch of different modes, there's a couple different models, and then the actual application and the way that you interface with these changed in a material way also,
Starting point is 00:19:47 where now the Codex app, which is the kind of the coding app, is now chat GPT. So Codex is rebranded to chat GPT. ChatGPT is kind of like depreciated and it's going away. And then Codex has chat GPT baked into it, but it also has a new codex and work feature. I guess the best way you can imagine work is kind of like a cloud co-work feature where it's just the only difference is the harness. So when you toggle the work mode, it gives you a less technical harness, I believe. And then when you toggle codex, it is more kind of catered towards creating code. And I think the co-work or the work feature is kind of based for knowledge workers. It's built for people who just want to do day-to-day tasks who want computer use.
Starting point is 00:20:29 They want to do kind of like automated spreadsheet or document creation. That type of thing is better for work. I will say that like I was somewhat confused when I tried to figure out how to use these tools and what is best to run when because it's not immediately clear. But I guess now I'm a codex guy. Like I don't use the GVT app anymore. I'm using codex. I'm running it.
Starting point is 00:20:49 I'm testing it there. And yeah, so far the results, been pretty cool. So yeah, that's pretty much it. Three different models. If you have a GPT subscription right now, please get on it. I'm curious what all of you folks end up building. My vibe take on this is it's very good. It's certainly competitive in some aspects,
Starting point is 00:21:07 but it's not good enough for me to switch over from using Claude or Fable 5 specifically. I just feel like the taste of Fable 5 is way better than GPD 5.6, but it is a good jump up. I don't like the fact that if I give it access to my desktop, that there is even the smallest chance that I would delete my entire desktop. I don't have that issue or concern with Fable 5, and that is one of the major reasons why people don't use the ability for an AI to take over your desktop. So I'm going to pause and wait for maybe GBT 5.7 or for GPD6,
Starting point is 00:21:40 which comes out in a month's time, but I think by that time, Fable 6 and successive models will come out. So it is extremely competitive, but it's nonetheless, a very good model from open air, and I like that I can use it relentlessly without worrying about my rate limits, getting it suited. Now, if you'll listen to this and you're wondering, hmm, okay, I'm convinced enough to use this model for A, B, and C. Let us know in the comments. I want to know, like, DM us, like what you're building, send us even demos. We would love to see it because we don't know what we don't know. Josh and I are in our podcast research world,
Starting point is 00:22:07 and I'm sure there are much more talented people out there in different professions that are using this for different ways, and maybe that can grow our idea of what this model can be useful. Yeah, so please don't forget to also share this episode if you enjoyed the episode with your friends, with your family, with anyone who might find this interesting, or who might be a user of GPT 5.6, or who you want to tell, don't use GPT 5.6 because it is a little scary and could be a little bit dangerous. But with that, you are caught up. That's a few fun demos. EJS. Thanks for prepping the demos. Those are pretty cool, pretty interesting. I'm hoping that next time we have a GPT demo, they're going to be maybe a little higher poly count, maybe like a little better visual graphics.
Starting point is 00:22:44 I mean, we'll see. Granted, I give him credit. It's only one shot. But hey, those one shot fable prompts, man, that was pretty good. But I think that is the episode. Don't forget to rate us on your favorite podcast player. And we are also opening up the doors to work with other people who would like to participate on Limitless the show. We are opening the door to sponsors and partners of all shapes and sizes.
Starting point is 00:23:05 So if you are interested in getting showcased on the show and becoming a character on the show, we are looking for amazing products, really interesting and compelling companies that are doing cool things to kind of showcase and highlight to help us keep the lights on and continue to publish this show every single day like we do at least four times a week. It's a lot. So if you made it to the end, if you watch this episode, thank you so much.
Starting point is 00:23:26 We have another one coming tomorrow, as always, and we will see you guys in the next one.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.