Everyday AI Podcast – An AI and ChatGPT Podcast - Ep 804: Open Source Surge? Does GLM-5.2 Make Open Source an Enterprise Priority? (Start Here Series Vol 29)

Episode Date: June 23, 2026

Is the open model GLM-5.2 really Opus 4.8 level? 🤯You mighta missed this, but over the past few weeks, three distinct forces have all converged at one: ↳ Chinese open models are near frontier SO...TA↳ Microsoft is reportedly considering open models to run Copilot↳ Enterprises everywhere are talking token efficiency as AI costs soarSo while many are watching GLM-5.2 as an isolated model, it's important we dive deeper on its wider implications.Open Source Surge? Does GLM-5.2 Make Open Source an Enterprise Priority? -- An Everyday AI Chat with Jordan WilsonNewsletter: Sign up for our free daily newsletterMore on this Episode: Episode PageToday's Episode on LinkedIn: Thoughts on this? Join the convo on LinkedIn and connect with other AI leaders.Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineupWebsite: YourEverydayAI.comEmail The Show: info@youreverydayai.comConnect with Jordan on LinkedInTopics Covered in This Episode:Open Source AI's "ChatGPT Moment"GLM 5.2 Model Benchmarks & PerformanceEnterprise Adoption Drivers for Open AIMicrosoft Evaluating DeepSeek for CopilotToken Maxing to Token Efficiency ShiftGLM 5.2 Infrastructure vs. Consumer UseAutonomous Workflow Overshoot ExplainedCapability Gap and Workflow ChallengesEnterprise Scenarios for Open Source ModelsFuture of Task-Specific SOTA AI ModelsTimestamps:00:00 Open source AI catching up04:52 Enterprise shift to DeepSeek models08:57 Comparing AI model performances12:46 Running AI models locally14:17 Open source model cost efficiency17:37 Cost challenges with AI models21:05 Agentic task token consumption25:05 Introducing the Start Here series27:58 Impact of AI on Job Roles32:29 Evaluating Open Source AI Models36:00 Considering open source models37:09 Future of open source AIKeywords: open source AI, open source AI models, GLM 5.2, z AI, Zhipu AI, Chinese open source models, DeepSeek, Microsoft, enterprise AI, token maxing, token efficiency, AI spend, AI deployment, open weight models, proprietary AI models, AI benchmarks, Artificial Analysis Intelligence Index, enterprise infrastructure, agentic workflows, coding tool use, autonomous agents, long context window, coding capabilities, API costs, AI privacy considerations, model distillation, data privacy, compute requirements, GPU infrastructure, AI hardware, API hosting, Hugging Face, AWS, AI cost reduction, Copilot Cowork, Azure security, Anthropic, OpenAI, Claude Opus, multimodal models, task-specific AI models, model capability gap, autonomous workflow overshoot, agentic tasks, non-agentic tasks, state of the art open models, model fine-tuning, small language models, AI adoption barriers, frontier models, AI job automation, workflow transformation, AI subsidies, token billing, Stanford AI study, AI industry trendsSend Everyday AI and Jordan a text message. (We can't reply back unless you leave contact info) Start Here ▶️Not sure where to start when it comes to AI? Start with our Start Here Series. You can listen to the first drop -- Episode 691 -- or get free access to our Inner Cricle community and all episodes: StartHereSeries.com Also, here's a link to the entire series on a Spotify playlist. 

Transcript
Discussion (0)
Starting point is 00:00:00 This is the Everyday AI show, the Everyday Podcast where we simplify AI and bring its power to your fingertips. Listen daily for practical advice to boost your career, business, and everyday life. There's three important things happening right now that make me think open source AI might be having its chat GPT moment. Number one, the models are actually pretty good with a recent splash from ZAI's GLM 5.2 leading the way. Number two, the era of token maxing is over as companies cut AI spend. And number three, one of the biggest and most influential companies in the world is looking at open source as a viable option. Granted, this doesn't mean that you'll have a frontier level model operating 24-7 on your
Starting point is 00:00:50 computer. That's not how any of this works. But for large enterprises, they will and do have that option today. But even if you're not a Fortune 100 company, with GPUs to spare, you two are going to have to start paying very close attention to open source models in 2026. Yes, the Chinese companies are distilling from American labs and there's privacy considerations, but that doesn't change the fact that U.S. tech companies are using these models in production
Starting point is 00:01:18 as AI costs are starting to skyrocket. So will models like GLM 5.2 thrust open models onto the streets of mainstream AI America, or will this just be another drop in the bucket until the next wave of U.S. lab models make the current open contenders look archaic in comparison? Well, let's find out on today's edition of Everyday AI as part of our Start Here series. All right, if you're new here, welcome, but let's talk about the big picture here. Open source AI has nearly caught the top proprietary models. So I think most even people who are bullish on open models would have admitted that,
Starting point is 00:01:59 for the most part, open source or open weight models are about six months behind. And I'd say now that gap is maybe only two months, two or three months, which is pretty incredible to see. And then a lot of benchmarks, which we're going to look at, open sources kind of caught the closed proprietary models. So they are now finally credible enough and powerful enough for serious enterprise evaluation. Also, Microsoft is reportedly looking at.
Starting point is 00:02:29 deep seek as it looks to lower its costs in co-pilot co-work. So that's huge. And the model pushing all of this, I think right now is ZAI's GLM 5.2. I know that's a mouthful. But we're going to look at some of the charts that show that this is now a big picture model. This is a big shakeup. And you have to be paying attention to GL52 and what comes after this.
Starting point is 00:02:56 And as companies now shift from token. maxing to token efficiency, open models may finally be having their chat GPT moment. So on today's show, here's what you're going to learn. You're going to learn why Microsoft reportedly looked at deep seek for lower cost co-pilot agents. You're going to see why GLM-5-2 is an enterprise infrastructure. Play not exactly a business laptop AI. You're going to know why autonomous workflow overshoot blocks adoption more than model quality. and I'm going to tell you about that secret issue that I think people aren't paying attention to
Starting point is 00:03:32 when it comes to close proprietary and open source AI, that autonomous workflow overshoot. All right, let's get into it. Welcome to the Start Here series. This is the Everyday AI Essential podcast series to both learn the AI basics and to double down on your knowledge. So if that's what you're trying to do, sweet, me too. So this is an ongoing series. we're actually on volume like 29 now. So make sure you go to start your series.com.
Starting point is 00:04:01 That's going to give you free access to our exclusive inner circle community. And there there's a playlist that has all of the start here series all in one spot on a Spotify playlist, all of the newsletters, everything all in one spot. And you can go connect with other business leaders that are trying to do the same. If you missed our last Start Here series episode, I think it was actually a really important one. We talked about AI super apps and why every. company is racing to create one in what they are. That was volume 28 or episode 799. And today, let's talk about the open source surge. So let's quickly recap these three different things happening all at once. So number one, Chinese open source models have kind of closed the gap.
Starting point is 00:04:46 They haven't closed the gap completely, but they've closed the gap to it's like teeny. All right. So obviously, I'm not going to talk a whole lot. on model distillation in this episode. But if you don't know what that is, essentially Chinese companies steal more or less from American companies, right? We've seen the U.S. government is working, you know, on this with the top labs. But Google, Open AI and Anthropic have all but said,
Starting point is 00:05:15 yes, that Chinese companies are stealing all of our work and making open source models. So we're going to look past that. And the reason why we're actually looking past that is, well, my. Microsoft, right? Microsoft being one of Anthropic and Open AI's biggest investors is reportedly looking at DeepSeek as a viable alternative to using the Anthropic and Open AI models. And that's one of the reasons why I think it's now finally time for enterprises to take a serious
Starting point is 00:05:45 look. So yes, there's obviously a lot of privacy data considerations when it comes to using open source models, all right, because you don't always know the weights. So, you know, you might be getting an output and blindly copying and pacing that, knowing there might be a geopolitical reason, you're getting a certain answer. So there's obviously a lot of considerations to take into account. But I think the fact that we're seeing, number one, the models are good enough. Number two, token maxing is going away. It's no longer about, oh, you know, everyone go burn five billion tokens. You can go climb up the internal company token leaderboard. That's over. Companies are cutting AI spend. And number three, Microsoft looking at open source.
Starting point is 00:06:24 models like Deepseek as a viable alternative. So let's talk a little bit more about this new model that is catching everyone's attention. This is from ZAI. I believe they're formally called Zifu AI, but they are called ZAI based in China. And GLM 5.2 is a 744 billion parameter MIT license open weight mixture of experts model. So it was developed by ZAI and it has a one million token context window and it is designed specifically for complex long horizon coding and autonomous Asian workflows. And it targets coding tool use long context and agentic work.
Starting point is 00:07:10 And here's the reason why we're talking about this, y'all. It is incredibly good. Okay. I use it a little bit. I've been impressed. I haven't had as much time as other people. and I'm going to be reading some thoughts from other kind of leaders in the AI space. But when you look at the artificial analysis intelligence index,
Starting point is 00:07:32 which was actually just updated to 4.1, a little detail there. But essentially, this takes about a dozen or so widely used and widely respected benchmarks, puts them all together and gives all of these models a score. So this is a good way to look at all the different kinds of benchmarks and to know about how smart or good a model is. So right now, obviously, we don't have Claude Fable, right? That could change it any minute. But the best model right now, it is one A and one B, at least what's generally commercially
Starting point is 00:08:08 available, Claude Obis 4.8 max and GBT 5.5 X high from opening I. And then not too far behind. Now you have a GLM 5.2. which surprisingly scores higher than any Google model. Again, that could also change because, you know, Google did say in June, they're coming out with their new model. So, you know,
Starting point is 00:08:33 that could change later today or later this week or next week. But regardless, this is the first time that, you know, in recent memory, at least, you know, since the AI race became more than just open AI.
Starting point is 00:08:49 This is the first time that an open. Open source or open weights company has cracked a top three company. Right. So it's now anthropic, open AI and ZAI. Yeah. So coming in ahead of Google, coming in ahead of GROC, coming in ahead of deep seek meta, the Kwan models, all of these other, you know, even the other Chinese open source models,
Starting point is 00:09:18 incredibly, incredibly well benchmarks. So a lot of accusations out there that maybe they're a little benchmarks or overfitting, you know, to make sure that they score really well on these benchmarks. But, I mean, let me just read for you all some reaction from some respected people in the AI community. So this from from Chris Somm, who's a partner at active capital, who said, I'm starting to feel like GLM 52 might be better than Opus. I was spending $300 a day on Claude, switch to GLM, spent $3.82 today, and it found and fixed a bug, a clawed bug from yesterday. I honestly can't tell which is better anymore. All right. Jeremy Howard, very well-known name.
Starting point is 00:10:08 He was the founding president at Kaggle. So he said, wow, ZAI's GLM 5.2 is a marvel. It is at least as good as Opus 4.8. in GVT-55. It is super fast, inexpensive, and not too verbose. It responds with nuance and judgment and handles long context very well. I've never experienced an open-weight model like this before. And then Guillermo Roch, who is the Vercel CEO.
Starting point is 00:10:36 So yeah, these are big names, right? You know Vercel. They're one of the leaders in the AI space. So he tweeted out, genuinely impressed, almost shocked at how good GLM-5-2 by ZAI is at coding. This changes things. And then last, but definitely not least, another prominent name in the AI space.
Starting point is 00:11:00 This is Matt Belasso, who's a former VP at Google DeepMind, also VP at Meta and worked at Microsoft as well. So he said, all day using GLM-5-2 didn't miss much. All right, so saying compared to using other models. So all-day using GLM-5-2, didn't miss much. First open model that passes the bar as a daily driver. Things are not going to be the same and then talking about how we need to get some serious hardware. All right. So now let's talk
Starting point is 00:11:30 about this because when people think open source models, what they default to is, well, oh, that means I can download this thing on my computer and I can run it 24-7 and I can plug it into OpenClaugh or Hermes agent. And all of a sudden, I have a, you know, a model that's the level. That's the of, you know, GPT 5, 5 and Opus 48 running for me 24, 7, and I don't have to pay any API bills. I don't have any subscriptions, et cetera. No, not at all, not at all. All right, no matter what anyone tells you,
Starting point is 00:12:01 because you're going to need at least $15, $20,000 in hardware, to get anywhere near that type of performance, right? And I'm just saying the outputs. So if you're just comparing outputs to outputs, yeah, you're going to need a machine that's at least $15,000. probably a little bit more, and it is going to be slow as a snail. So it is not economically feasible or economically reasonable, right, to, you know, run a model like GLM52 locally.
Starting point is 00:12:32 You will also be running a quant version. So you'll be running like a two-bit quant. So it's got to be a dumbed down, very slow version. So, yeah, good luck doing that. But, you know, there's people that argue out there like, oh, yeah, I'm doing this. So unless you have some very rare. special reason, you know, that you have to be able to run things locally, right? Because that's the other, obviously, aside from costs, you know, once models do become more powerful and smaller and efficient
Starting point is 00:13:03 enough to run on consumer hardware, obviously the privacy, right, being able to run things locally and not having to send things to the cloud. But that's not where you're at with GLM52, right? There's maybe, I don't know, a couple of people in the world that have a powerful enough machine to actually do that. But it's for most people, this is not going to replace your $200 subscription to Claude or Chad GPT. But this is a legitimate option for large enterprises who have access to compute. All right.
Starting point is 00:13:35 But small teams skip this, all right? I'm not saying skip GLM52 because obviously you can still use it on the API side. You know, there's plenty of hosts from, you know, AWS to hugging face, every, everyone in between that you can actually go and run this model using their servers, their GPUs, right? But this is not something you need to overlook. All right. And what do I mean by that?
Starting point is 00:14:02 GLM52 might not be a model that you're going to run locally. It may not be even something that you run on the API, even though it's a fraction of the cost, right? That's the other big thing. It depends on what level you're using, right? But if you're comparing. the highest output of GLM 52 to the highest, you know, GPD 55 or Opus 4Aid or Fable 5 whenever it comes back.
Starting point is 00:14:27 It is a fraction of the cost to run it on the API side. So yes, there still is, you know, if you're building modularly, which is smart to do and can, you know, swap out a set of API keys. Or if you're using something like open router that makes it even easier to do that, sure. You know, maybe you're running this model in the cloud. But this is not something, you know, I think most people think, oh, open weights, you know, I can, you know, download it, fine tune it.
Starting point is 00:14:55 You're not going to be doing this on a consumer piece of hardware. But so that's a factor one. Factor one is, well, the model is actually good enough. Why this might mean open sources having its chat chagipt moment. Number two, the Microsoft move. So according to Axios, and this is just not even a week old, we covered this on our Monday show. So Axios reported that Microsoft is actually looking at lower cost open source alternatives to Anthropic in GPT models. So right now, both in co-pilot, but specifically co-pilot, you know, they mainly have always relied on GPT models from OpenAI.
Starting point is 00:15:38 Over the past year, they've started to integrate and offer other models specifically from Anthropical. Claude models as well. But co-pilot co-work is obviously an agentic offering. This is based on Anthropics, very popular co-work technology. It's kind of Microsoft's version. Microsoft is obviously a big financial backer in both Anthropic and OpenAI. And they also well serve those models. So they make money on the cloud when their customers use these models.
Starting point is 00:16:12 So the fact that, and I do have to be, you know, you know, break this down and I won't go too much into it because we just talked about it on yesterday's show. But the fact that Microsoft is potentially looking at deep seek of all companies, right, because all the U.S. companies have called out deep seek by name for distilling their models. So I don't know if I'm in leadership at OpenA.I. and Anthropic, not super happy with my biggest financial backer by potentially, even though this is just report, the report could be all wrong, maybe it's true, maybe it's not. But the fact that this report is out there from a reliable source in Axios, that Microsoft, one of the biggest companies in the
Starting point is 00:16:58 world and one of the most trusted and respected companies when it comes to AI deployments, right, the fact that they're looking at a company like DeepSeek to potentially be an option under the hood for running co-pilot co-work to make it more affordable to customers is big, right? Because now that co-pilot co-work is generally available. That's new as well before it was just in beta. But now they're charging usage. So they're having to, instead of subsidizing this, like so many companies are, right, they're charging end users for usage. And well, what they're going to see very far is co-pilot co-work, no one's going to use it if they have to pay the APIs for, you know, Opus 4.8, which is one of the most, aside from Fable, from Anthropic, it's Opus
Starting point is 00:17:47 4.8 is one of the most, you know, expensive models out there. You know, depending on the task, it's actually twice, more than twice as expensive as GPD 5.5. So pretty big here. But the reported goal was just cheaper agents inside of Azure security protections. And that pressure exists because of the agentic work behaves just like it's it's nonstop right i think when we go back to the chatbot era we didn't have to worry as much about you know token usage uh but now when these agents run in loops right and it's it's easier to set off loops now uh in codex and claude code it's easier uh to you know put into this goal mode and you know have an agent work for not just hours but days Right. So now all of a sudden, we have to start thinking about things like token efficiency and, you know, the cost of compute.
Starting point is 00:18:45 All right. And that leads us to, well, factor number three. And that's the, the shift that we've gone. Well, now companies are not just saying let's burn tokens. They're saying let's save tokens. So there's been a lot of recent reporting over the last week or two. You know, it's New York Times article from this week that read tech workers max out their AI use. they're trying to minimize it. Business Insider Story titled Silicon Valley's AI token craze is facing a reality check. And then a tech crunch article that said the token bill comes due inside the industry scramble to manage AI's runaway costs. And I did cover this in depth on episode 789, also in the start here series when we talked about token maxing, the shift from token maxing to token efficiency. But in short, you had meta in other companies.
Starting point is 00:19:37 that had these internal leader boards where they were reportedly just rewarding employees for using tokens. And I think it started maybe with good intentions, right? Because AI leaders thought, well, hey, if our people are using AI a lot, that means that the business is going to grow and we're going to be saving time. But that backfired because people were just burning tokens. You know, a reported example is one engineer used 281. billion tokens in a month just to climb the rankings, right? Which is a pretty, pretty crazy number.
Starting point is 00:20:14 You know, you had reports that employees were just running agentic loops intentionally to keep running, even if they're doing personal projects, right? But they were just burning tokens just to burn tokens because, well, there was a thought for a brief period of time. I'll say from the end of 2025 to early 2026, well, people looked at employees all like for a brief period of time, right? They saw token usage as like a KPI as a key thing to evaluate employees on, right? Like, hey, if you're burning tokens, thumbs up, you're doing a great job in our book. And then there was the report that a company reportedly spent $500 million on Claude on accident in one month because they forgot to set limits.
Starting point is 00:21:02 And that shows you just the amount of tokens, right? So a Stanford study found that agenetic task can use up to a thousand times more tokens than a single chat. All right. Because chatbots answer once. Agents can call tools endlessly in a loop. Right. I actually, for fun, I have my Claude code, ultra code going on a simple task just to see how much usage it's going to burn when it shouldn't burn anything. Right.
Starting point is 00:21:31 But I'm looking at the chain of thought. And I'm seeing Opus 4.8. going in some silly loops, just burning tokens, you know, like, I don't know, burning marshmallows on accident. But this is going to a broader capability gap that companies are already struggling to close. And that capability gap is half the problem. And I'm going to go through this one quickly because I have covered this one in depth before on episodes 735 and episodes 755.
Starting point is 00:22:05 So if you want to know more about this, I did cover it in both of those episodes. But great study from Anthropic that talked a little bit about the capability gap. And this does get to the kind of closed source versus open source and GLM 52 stick with me here. But in that study, essentially, Anthropic looked at anonymized chats and they hit a ceiling for these different categories of work. And they said, here is the ceiling of a model's capabilities. and according to all these anonymized chats, here's what people are actually using it for. So one of the best categories that had the best usage
Starting point is 00:22:42 was only 33%. And that was computer math tests. But mini-tas was not even 10%. So you could have the best, as an example, business strategist and people were only using 10%. And that's the baseline kind of model capability gap. And that's the first bottleneck.
Starting point is 00:23:03 But there's kind of a new. This is kind of my secret term I teased in the beginning. I've been trying to put a name and a face on this. So I'm trying this out. Maybe I'll rename it down the line. Right. There's always these common concepts that come up over and over. And it takes me, you know, sometimes a month or two to put, you know, to put a label on it.
Starting point is 00:23:29 So you're not going to see this anywhere. This is something I'm trying on for size here. But I think aside from the model capability gap, I think the bigger problem maybe is autonomous workflow overshoot. And I think that's the next gap that you need to prepare for. And that might, you might see, stick with me here, you might see how that might actually make some of your day-to-day tasks looking at a model like GLM 52, even though it's text only. It's not multimodal.
Starting point is 00:24:00 it still might make it a feasible option in the future. So let me talk about this concept of autonomous workflow overshoot. So models right now, models by default, GPD-5, Opus 48, Gemini 35-35 Flash, they carry your company's context, they can plan, they can act, they can call tools, they can spin sub-agents, right? Essentially, I ask companies, what would you do if every single employee had at least one or many 24-7 agents? AI moves too fast to follow, but you're expected to keep up. Otherwise, your career or company might lag behind while AI native competitors leap ahead.
Starting point is 00:25:00 But you don't have 10 hours a day to understand it all. That's what I do for you. But after 700 plus episodes of everyday AI, the most common questions I get is, where do I start? That's why we created the Start Here series, an ongoing podcast series of more than a dozen episodes you can listen to in order. It covers the AI basics for beginners and sharpens the skills of AI champions pushing their companies forward. In the ongoing series, we explain complex trends in simple language that you can turn into action. There's three ways to jump in. Number one, go scroll back to the first one in episode 691.
Starting point is 00:25:38 Number two, tap the link in your show notes at any time for the start here series. Or you can just go to start here series.com, which also gives you free access to our inner circle community where you can connect with other business leaders doing the same. The start here series will slow down the pace of AI so you can get ahead. And it's not a rhetorical question because those capabilities are here now, right? I've had both Claude agents and Codex agents run for more than 24 hours. That's the autonomous workflow overshoot because the capability gap is humans are not fully understanding the model's capabilities or using the models to their fullest extent, right? This is a human not taking advantage of what's there. But this is different.
Starting point is 00:26:35 This is overshoot, right? So essentially, I'll argue that 99% of bleeding edge capabilities go unused. All right. And that's from a combination of the capability gap, but also workflow overshoots. So what is autonomous workflow overshoots? That's essentially that the capabilities of the agentic models are uncharted in your standard enterprise workflows because I will say that most workflows still need that human handoff, right?
Starting point is 00:27:14 If I told you at your company, hey, everyone has a 24-7 agent that can carry your company's context, it can plan, it can act, it can research, it can call tools, it can create PowerPoints, Excel, anything, websites, right? hardly anyone, hardly any business leader would be prepared to implement that. That is the overshoot, right? Today's models are more than most companies can handle. Not more than most humans can take advantage of the capabilities. That's two different things, right?
Starting point is 00:27:54 They can't handle this. That is the workflow overshoot. These models are capable of just too much because I think that even AI forward companies right now need multiple quarters a year or even more to completely rebuild job descriptions, their inputs and outputs, approvals, and even the type of work that you do, right? These things all have to change. This is also this concept of, yes, there is going to be all these new AI.
Starting point is 00:28:25 There's going to be all these new jobs that AI is going to create. I ultimately think AI will take away more traditional full-time roles than it will create. But AI is obviously going to create millions of jobs that we just don't know what they look like because it's going to take businesses a while to understand where all these autonomous agents' capabilities are headed so we can start creating both the context that the agents need, the expert driven loops that they need to do these workflows right, but also the lines of business, the streams of revenue that go along with those things. Those things are all going change and that is the autonomous workflow overshoot. So thank you for sticking with me because now I can
Starting point is 00:29:08 tie this together and answer the question, well, should open source be an enterprise priority? All right? And I'll say yes in three different scenarios. So let me lay those out for you. Number one scenario, if your company has access to compute and they also have a high API bill, right? They're shipped in a way. you know, from token maxing to efficiency, I think a general purpose model like GLM52 makes sense. Okay. Again, this is a very small sliver of companies, right? This is essentially your Fortune 100 companies. Not everyone has, you know, a rack of GPUs sitting on the shelf and can, you know, if they want to,
Starting point is 00:29:53 you know, roll out the type of compute that their enterprise needs. Like I'm saying, GMM52, you can't download this on your, every employee's laptop. It's not feasible. Doesn't make sense. It's not going to work. But if you have the compute, you can. Right. Yeah. And what's crazy is I've talked to and I've worked with plenty of companies that fit into this category. I've seen their, you know, server racks and all the, in, in video chips blazing and all the cooling, you know, water going underneath it. It's all above my head. But there's companies out there, yeah, they can, you know, find, they can, you know, find they can download g lm 52 it's open weights they can fine tune it and they can you know make a
Starting point is 00:30:37 portal um or a way that their employees can access this model and in theory they can use it 24 seven around the clock right obviously there's still bandwidth and other issues that you have it's not as easy as you know click download and you know click deploy but there are companies number one well maybe because of that what i said the autonomous workflow overshoot i'm literally thinking of one company in particular that's spending millions of dollars a year on Anthropic as an example, they could probably do this because I would say less than 1% of people using this, their current AI system that they're paying multiple seven figures for are less than 1% are using it to its full capabilities. So in most cases, the 99% of people, a GLM52 would probably be a
Starting point is 00:31:31 enough. Aside from the fact that it's not multimodal, that that is a huge downside, right? But for the most part, number one, the companies that should be using it are those that have access to compute. And for the most part, they're not needing 24-7 or can't take advantage of 24-7 autonomous coding agents. Number two, those for non-aetic tasks, all right? And this is just a different chunk, a different segment of work. So this is the non-agetic tasks. that can and should be chunked for future open models. All right. So like I said, so few people right now, their workflows require something like a Fable 5.
Starting point is 00:32:15 Yeah, it's fun to go in there and, you know, oh, let me make this 3.js World View game and, oh, let me go code up this website, right? Yes, there's obviously people in software development and dev roles that need that 24-7, But most people even using this, you know, you don't need Fable to write better emails, right? So I think when you start chunking your large enterprise companies, start chunking your non-agentic tasks, your non-frontier tasks. I think that that's a big group of people that can start looking at an open source model like this, maybe using it via the API. And then last but not least, it's preparing for the future. because I hope that intelligence just, you know, I hope we truly do get the intelligence too cheap to meter promise at some point soon.
Starting point is 00:33:11 But there's a reality that as the capabilities increase, right, at least for now, for the most part, just, well, that Stanford study showed that, well, agents just can burn a thousand times more tokens than this simple chatbot query, right? And that is one downside of GLM. too, it is, it is token inefficient. It burns through a lot of tokens to get that level of intelligence, although the level of intelligence is extremely high, the highest we've literally ever seen for an open model. But I do think that we're going to see smaller task soda models in early 2027. Let me tell you what I mean by that.
Starting point is 00:33:50 You know, task is soda. So, you know, task specific state of the art models. Right. So right now when we talk about models, They're just general, large language models, right? They're one model and people use it for everything. I've been a huge advocate and believer in the future, right? There's going to be a mixture of models technology.
Starting point is 00:34:11 Thank you, certain companies that made my crazy 2023 prediction a reality in 2026. I was only a couple of years too early. Right, but we've seen that, you know, open router just had their fusion technology, you know, perplexity with their computer, right? You put a prompt out and it will route it to whichever, model of things is best. It might put it through multiple models, right? But I think we're going to see that on a small, small language model platform here in 2027. Probably not this year, but essentially what's going to happen, right? The model distillation is both, you know, when you, I think most people
Starting point is 00:34:50 think about model distillation, they think, oh, you know, Chinese models distilling a big frontier trillion parameter model, right, which is what we're seeing. here with a lot of the, you know, the deep seeks and the quens, right, all these accusations flying around. But what about, well, just when, you know, frontier companies distill their own models, right? The legal and, you know, teacher student model scenario. Or, you know, we're obviously going to see a lot of these Chinese companies come out with small language models or task specific models. But I think those are going to be state of the art, right? So as an example, I think that whether it's through, you know, quote unquote, not exactly legal distillation or intentional legal distillation.
Starting point is 00:35:36 I think we're going to see state of the art open models for things like, you know, specific tasks, you know, front end coding, web search, text summarization, PDF parsing, you know, copywriting. I think we're eventually going to see dozens after that, probably hundreds of open models that are state of the art at one task because when you can create a model around one certain task, it can be smaller, it can be better, it can be more token efficient, which in turn makes it cost efficient, right, which is what people are wanting. So those are, I think, the three scenarios where enterprises should be considering open source models, whether it's GLM 52 or something else. So to quickly recap, Number one, if your company has access to compute and a high API bill.
Starting point is 00:36:33 Number two, if you're an enterprise company that can start chunking all of these different AI tasks into agentic versus non-agentic. And maybe you keep your current agentic options that are maybe a little bit more expensive. And then you take your non-agentic options to maybe on the API side, something like a GLM-5-2. And then last but not least, it's more, I think, for everyone else preparing for the future. And it is starting to categorize those tasks because not every task needs a fable five, right? Not every task is going to need a GPD 56 pro, although I'll still use it for every task, right? Like that's how you also have to start thinking because eventually these subsidies are going to go away.
Starting point is 00:37:15 We are going to have to become a token efficiency mindset. That is the future. And I think is this the Chad GPT moment for, uh, open source is GLM 52 that thing I don't know if it is it's still too early to tell however I do think
Starting point is 00:37:37 this is if nothing else the foundation for open source to have its chat gpt moment all right I hope this one was helpful if so please let me know about it go to start here series.com there you can sign up for free access to our
Starting point is 00:37:53 start here series kind of space inside of our inner circle community where every single episode is there ready for you to gobble up. You can listen to it on 2X. I'm not going to be mad at you. All right. I hope this is helpful. Thanks for tuning in. Hope to see you back tomorrow and every day for more Everyday AI. Thanks y'all. And that's a wrap for today's edition of Everyday AI. Thanks for joining us. If you enjoyed this episode, please subscribe and leave us a rating. It helps you helps keep us going. For a little more AI magic, visit Your EverydayAI.com and sign up to our daily newsletter so you don't get left behind. Go break some barriers and we'll see you next time.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.