Embedded - 535: The You in the Machine

Episode Date: October 1, 2026

Adam Smith from UC Santa Cruz joins us to discuss local Small Language Models (SLMs) and building open, autonomous tools. We explore his BayLeaf AI counterplatform, transagency (the "me-and-the-car" c...oncept of human-agent collaboration), context distillation, and running models locally without relying on data-harvesting corporate clouds.  Adam also shares practical strategies for working with coding agents, including why plain-text AGENTS.md files outperform model weight fine-tuning for long-term knowledge retention. He also describes how he uses open-source tools like OpenWebUI, OpenCode and OpenChamber to integrate small and local LLM integration into development environments and web chats.  Discussion Resources & Links Adam's Projects & Research: BayLeaf AI Counterplatform: Hyper-local, open-source AI infrastructure built for UC Santa Cruz. See also the BayLeaf Getting Started Guide and BayLeaf Blog. Adam Smith's UCSC Directory Page & Personal Website Adam Smith on YouTube & GitHub (@rndmcnlly) Open-Source Agent Harnesses & Wrappers: OpenCode – Open-source command-line coding agent harness with VS Code integration (comparable to Claude Code) OpenChamber – Graphical desktop wrapper for OpenCode with VS Code support (comparable to Claude Cowork) OpenWeb UI – Self-hosted web interface and desktop app for local and remote LLM models OpenRouter – Token broker/middleman routing requests across competitive LLM inference providers r/LocalLlama – Community subreddit for local LLMs, open weights, and hardware advice AGENTS.md Convention & Agent Skills Standard – Open standards for persistent context management Context Distillation Paper (arXiv) Adam's Guide to Building Transagent Capabilities: Scope by Project Roots: Harnesses like OpenCode and OpenChamber use project root folders to establish clear, safe boundaries for agent execution. Store Context in Plain Text: Keep global and project preferences in simple .md files (like AGENTS.md or GLOSSARY.md). Plain text files make your accumulated knowledge portable across different model front-ends and harnesses. Handle Errors by Rewinding: When your LLM or agent makes a mistake, don't yell at it or try to correct it inline—go back in time to before the error and alter your prompt or context to lead it away from the mistake. Refactor Preferences into Skills: As your AGENTS.md context grows, ask the agent to factor recurring preferences out into modular skill files (.agentskills). Teach Transagency: Distinguish between standalone agent capabilities (autonomous work) and transagent capabilities (how the agent works with you). Build Local Tools: Ask the agent to build desktop helpers (QuickLook triggers, push notifications, or fuzzy glossary lookup) to keep a shared orientation. Elecia: "Complete one project or start a dozen?" Adam: "Tens of thousands, and start hundreds more within the starting of the other projects recursively forever." Transcript

Transcript
Discussion (0)
Starting point is 00:00:06 Welcome to Embedded. I am in Lisa White alongside Christopher White. Do you remember that Think Geek t-shirt for anybody who doesn't, just don't, that said, I will replace you with a small shell script? Yeah, we're going to talk about that a lot today. And we have Professor Adam Smith to talk about it with us. Hi, Adam. Welcome to the show. Welcome to the studio. I'm happy to be here. I'm sort of local, but it's good to finally see the studio. Could you tell us about yourself as if we met at a tie-dye party that was populated mostly by software engineers? Oh man, this is a nonfiction scenario. But if I didn't know you all already, I'd say, hi, I'm Adam Smith. I work at UC Santa Cruz. And these days, a lot of what I'm focusing on is my Bayleaf AI counterplatform. But we'll probably talk more about that later.
Starting point is 00:01:00 Counterplatform sounds so cool. Yeah, it's a word that, well, the AI system showed it to me and I'm like, I'm going to adopt that. And that word didn't seem to be used much on the internet. So I'm, it kind of, we'll say more about it. That and the word hegemony, if anybody is reaching for their dictionary already. You'll need it later. But first, we're going to do lightning round. Are you ready?
Starting point is 00:01:26 I'm ready. You are not. I promise you are not. Yeah. But shall we? Great America or Santa Cruz Beach Boardwalk. Santa Cruz Beach Boardwalk. Best date format.
Starting point is 00:01:36 Month, day, day, year, year, year, or year, year, year, month, day. Y, Y, Y, M, M, M, D.D. Of course. With the capital, Y, Y, Y, M, M, D, D. Yeah. Best 90s land party game? Best 90s land party game. I'll go with the Santa Cruz local bridge builder game, which I was at a land, well, we
Starting point is 00:02:02 need to have a bunch of stories, but the story anyway, I was at a land party once in nearby Las Gatos, and people were playing games like Unreal Tournament and CounterStrike, and somebody started playing Bridge Builder, and somebody looked over and like, what game is that? I want to play that, too. And we all went from these loud, rock is violent, competitive games to quiet, contemplative bridge building, and people just did this until four in the morning. I mean, that's how I would, that's, those are all my games. Yeah. Are there dinosaur bones on the moon?
Starting point is 00:02:36 Oh, man. It's a yes or no question. Possibly. I know what that's a reference to and didn't think I was going to get to that. What research paper was the most fun to write? What research paper was most fun to write? I think there was one that was. trying to that I did while I was at University of Washington, where we were trying to set up what we want from generated puzzles at an educational game.
Starting point is 00:03:06 And we phrased it as, it's hard to do it without spinning out into all the details of the paper. But it came down to there was a very clear expression with, you know, there exists and for all quantifiers. And we're like, that's what we want from puzzles. We want it so that even if the puzzle has a zillion solutions, none of them let you escape, the educational concept we wanted you to be expressing in that particular experience. Forced education, in other words. Yeah, you could think of as forced practice where you could say, we're actually going to give you a huge space of ways to solve this puzzle so you can play around within that space,
Starting point is 00:03:47 but we'll gate keep it such that it doesn't count as complete until you've also practiced the thing that we want. in a pretty subtle way. Must not ask about Zork and physics simulators. Must not. Okay. Favorite LLM model? These days, you know, for the next two weeks,
Starting point is 00:04:05 I guess, because these things change so fast, it would be GLM 5.3 flash. That's been my favorite for the last two weeks, and it'll hold for another week and a half. Most overhyped LLM model. The most overhyped one, I think that would be
Starting point is 00:04:22 deep seek R1 which I guess people blamed it for a bunch of financial impacts and I thought it was a sort of interesting one but sort of people were scared about it for all the wrong reasons it was kind of interesting but also kind of forgettable but it was kind of the first
Starting point is 00:04:41 non-major American one right I'd say the one that Deepseek V3 that came out like a month before I was super interested about that one because it was, in my mind, the most significant open weight model at the time, which changed my mind a lot. And then they came out with the one that was trained to do the reasoning thing, the yapping before the answer. But that little thing, which I think is just sort of, in my mind, a stylistic way of talking that you get via fine tuning. That's the one that people were all scared about.
Starting point is 00:05:13 And I'm like, the thing that came out a month before that I had already been using was better. It's not that it was actually a little bit worse, but that was the bigger step up that made me think, yeah, this open weight stuff is we're going to have this. And there was a track off of the open AI stuff at that point. General artificial intelligence or augmented human intelligence. Definitely augmented human intelligence. I'm not always trying to put humans at the center of things, but I'm definitely not trying to make things that live on. on their own and do everything. I really like people having the right tools around them,
Starting point is 00:05:54 and when the right tools are not around them, they use some of the tools they have to make the tool that they want, and they just keep on doing it. And whether the tools can sort of speak subjective sentences in the meantime is sort of irrelevant to that. You're making fried tofu. Do you use cricket flour or acorn flour? Oh.
Starting point is 00:06:13 Oh. Oh. Fried tofu? at least in my house, we make it from homemade soy milk that you get from soybeans. You order on Amazon. The cricket flour is for the, what do we call those cookies? Bug cookies? No, no, there was a name.
Starting point is 00:06:38 High protein cookies? Cookies with legs. Oh, I got it. It was chocolate chirp cookies. We found a recipe online and made those. and the acorn flower, it's going to be another big masting year. So my family is probably actually going to collect acorns today.
Starting point is 00:06:55 Would you like some? I'll take a to-go bag later. You know that the good ones don't flow, right? So you don't have to open them in order to figure out which ones have bugs in them? It's not like a 100% test. But the ones that do float, you can just toss. Yes, yeah. But then that, because there's a later start.
Starting point is 00:07:15 of drying them, it slows down the drying process. So sometimes I just store the ones with bugs in them. So when you bash them open, at least the bug died a year ago. Bacon, erdush number. Oh, I haven't kept track of it. I think I have an erdos number of three. And if you count being an extra in a movie, I think I also have a bacon number of three. But I think I think. I think I also have a bacon number of three. But I think I I think you're not supposed to count just being an extra. We'll allow it. So together, it's not infinity. I think it's six. YouTube or Wikipedia as a learning resource. Historically, for me, it had been Wikipedia,
Starting point is 00:08:02 but there's so much stuff that I get from the medium of video these days, especially for cooking the texture of sauces and all sorts of stuff that things don't get written down. The sounds, not the smells, but the look and gloop. of things. The gloup of things. The gloup of things. Favorite fictional robot? A classic question.
Starting point is 00:08:23 I guess that would be data from Star Trek Next Generation, but I can't separate data from Wesley Crusher, who is not a robot, but they're both things on the threshold of becoming adult human. But there is this belief that, of course, Wesley is going to make it to adulthood, and data is not going to make it to full humanhood, but they're both on the threshold in the show, which makes me lump them together in one category. What science fiction have you found most moving recently?
Starting point is 00:08:56 Or ever, whatever. They're, in playing with LMs recently, I've occasionally had them generate short stories, and the fact that when they are stories from the perspective of AI agents, the fact that they are not entirely fictional, is moving. And the fact that I made the story and also things happen in it that I don't anticipate.
Starting point is 00:09:21 I'm like, whoa, I didn't know it was going to go there. But I made it. I don't think the idea that was put in there came from a specific other person, but also the story is about the AI's writing the stories. It's lost in layers and meta stuff. But those are the things that have been most moving to me recently. Do you have a tip everyone should know? people complain about the LMs being sycophantic.
Starting point is 00:09:49 So there's currently this trick of, if you want good feedback on something, you can tell your chatbot, hey, my colleague who uses ChatGBTGBT, way too much wrote this text. Let's tear it apart. And then the sycophantic robot will defend you against the terrible writing of your colleague,
Starting point is 00:10:05 which, of course, was just the text you wrote before with it. But that's a way of, it's weird that this trick works. it might not work in the future years. It makes it less sycophantic? It sort of leverages the sycophancy in your favor to help you find flaws that it would have maybe massed otherwise in the interest of pleasing you. I would like, I was playing with one recently, and the obsequiousness that has been removed
Starting point is 00:10:32 from some of the bigger models was in the small model I was playing with, and it drove me up the wall. I'm not used to that anymore. And I don't want it to tell me how amazing it finds my ideas, because they weren't. I just wanted to conjugate some Hungarian verbs. They weren't amazing ideas. In one of my projects, I sent a link to my project to someone else, and it was mostly a L.M created project, and they complained about the clot-ish in the read-me of it, because it was a personal project. I just wanted to chuck it out there to quickly share with other people. And it made a read-me I didn't look at, but when I looked at it, it had almost a certain smell. Like, I didn't quite remember when I had done the project, but I could identify from the style of read-me writing.
Starting point is 00:11:18 I think this was around March 2026. I went to look back to the commit history. And yeah, it's almost from the flavor of it. You could tell that it came from this particular not too many months ago, historical vintage. And the read-mys that come out now have a different taste. but those at the time had a... I'm going to gently flag this. Yeah.
Starting point is 00:11:39 No, sorry, no, I was being clogged. So, lighting around just over. Everyone knows my feelings about AI and LLMs, and so I don't need to rehash all that. I'm not going to be good today. So I'm not going to... You've already brought up one of the things that I find most boggling about some of the current LLMs,
Starting point is 00:12:02 which is they change. so fast and it's not my career. I don't want to know about everyone that comes out and have to weigh the pros and cons of, what do I do? I think for the most part, like, I've kind of made it my hobby and job and all these other things to be involved in this space. So I check on it multiple times a day sometimes. Right. Yeah, yeah.
Starting point is 00:12:31 It is your job. Yeah. But I think because these, in my mind, it's so much easier to catch up on things than to ride the wave of whatever the fashion things are and acknowledge that they change and just check in every six months. But as part of that, like, if you had a conviction, you're like, this is what I believe six months ago, just know that six months from now the things might be categorically different or which parts of them are categorically different is different each time. Like, be open to the new thing each time, but it's fine for you to snooze it for six months at a time. Yeah, I see people talking about, oh, X.3-3 underscore flash, you know, whatever, came out. And compared to, you know, three other models I've been doing. What are you doing?
Starting point is 00:13:18 Like, what is your job? Like, I mean, it's very hard to keep up with those things. And some of them, I've seen some people say, oh, well, this one's better for this and this one's better for that. And it seems like really difficult to keep up with. and at some point counterproductive, like you're spending more time, thinking about whether you're getting a better Python script out of 5.4 than 5.5.
Starting point is 00:13:39 Are you sure? No, did you do a test? There are sometimes when I have systematic evals, but then the things you want to evaluate for are what is changing. Things you wouldn't. Right. You would think that it's hopeless to do this.
Starting point is 00:13:51 I'm not even going to measure how often they do it, but now it's like they can. But at the same time, the smaller models, can run on your laptop and your phone, those things that a few months ago, you'd be like, well, the one that runs on my phone is really dumb.
Starting point is 00:14:06 It feels like the things from two years before, but it's now the case that the things that I can run on my phone feel like the ones that were, there was all the breathless hype about when that level of capability was a frontier thing six months ago, and now that's on my laptop. Right. So there is this almost trend downward
Starting point is 00:14:25 in size of many things. And it's a big surface to keep track of and I think most people can just ignore it for six months at a time. That was actually mostly what I want to talk to you about, which is the smaller LLMs. And the one that I used that was so obsequious was a pocket AI. Do you remember which model in, because Pocket AI is like the app that runs many models? Sage. Give your phone?
Starting point is 00:14:54 Oh, yeah, actually I have it right here. It was Gemma, three. 3-4B-I-T-I-KM. Okay, okay. I know what all those words mean, and I think of that one as being, I know, the exact date, a year old. There's a Gemma four line of ones that are almost the same size, but have a completely different feel to them in terms of what they can do.
Starting point is 00:15:22 Jim is a Google model? Yeah. Okay. Or Jim is a family of models that come from Google. relationship between what's the name of the organization, what is the name of their model family, is thoroughly confused. Some things name it after the company. Some like the Quinn family of models comes from Alibaba, but unless you looked it up, you wouldn't have, you'd be like, is there a company called Quinn? No, it's just, it's more of a product line or a model line.
Starting point is 00:15:49 But I think one of the differences with those, the, the Gemma three family ones, I think, came from a time when it was uncommon to have LLM's call tools. There were, mostly just about trying to be helpful chatbots. And more often, I want the thing to go off and look things up for me, prepare a prototype, generate a video of interacting with the prototype that involves the LM writing tiny little shell scripts and run them. And the things that were just good at talking, even if their talking ability has not changed very much, their ability to go use a computer to actually do something or to go out and sort of get the truth from the internet of something and bring that back to you, that's what's changed a lot maybe in the last year.
Starting point is 00:16:38 So I hadn't really played too much with this app, which is actually an iPad app, iPhone app, called Pocket Pal's. Pocket Pal. And I had chosen one that actually set a page. tutor. Oh, that's your first mistake. And that was actually my first mistake. That was why, and I can get a friendly general purpose pal that runs entirely on my phone or a real-time video analysis assistant that runs G-GML-L-Rg, S-M-O-L-V-L-L-M.
Starting point is 00:17:13 Small? Yes, small VM-500M instruct Q8 underscore 0. I guess Q is not the core of a project where you're just not. done with it and you keep adding final, final, final, final. I think that's what's happening. I think some of these things would be like if somebody had, I don't know, you've got a pirated movie and they have the name of the movie, then they have what codec it's in, then they have which resolution, then they name which subtitles and you're like,
Starting point is 00:17:42 well, did you see the new Jurassic Park? Did you see the new Jurassic Park, HD, H-EVC, HK resolution with Spanish subtitles that include the description of things that people are physically doing? and all that in the giant file name. It's sometimes shortening it to just, did you see Jurassic Park? You can say, yeah, I saw it in the 90s, and they're like, oh, but there's a new one.
Starting point is 00:18:01 So some of those details matter, and most of them don't know. Before you get all to sell the fan edits of Star Wars. There are effectively fan edits of LLMs, too, where people take an off-the-shelf one and fine-tune it for various things. And I think of those as being kind of like fan edits, and there's people who collect all the different fan edits
Starting point is 00:18:19 and do silly things with them. This pocket pal has a bunch of those. Yeah. I really didn't look at this. Yeah, I was telling you, I had locally, it was a different app, does the same thing, and it has seven or eight models that you can choose. There's one that's all about languages. There's one that wants to be my romance friend. I'm being good.
Starting point is 00:18:42 I'm being good. Well, I'll say, like, I have almost no personal use for the laptop scale ones. Sorry, the phone scale ones. when I say almost no, there is a particular place that about once per week I go in Santa Cruz where I don't have cell data coverage. And I asked for the Wi-Fi password once and I forgot it. So occasionally I have like a question and I don't need a super high confidence answer. I just want to like think through an idea and it is nice to have like an offline copy of all human knowledge
Starting point is 00:19:20 sort of mush together in my phone that I can be, I can effectively look up something in that one particular place where I don't have cellular data coverage. And it, in those times when you're offline and all you have in your hand is your phone, it's nice to have some capability, but that I spend, you know, 0.01% of my life in that particular mode. So it's cool that there's an app that can do those things, but it's certainly not something I do every day.
Starting point is 00:19:50 I think those things become a lot more useful when they can search the web. But if you're connected to the web enough that your local app can reach the web to do the research, you can often use a model that's much better than the one that's entirely on your phone. That's true to the extent that you are using it at home. The number of times I have been in the middle of nowhere in the bomb range of Southern California or visiting with rattlesnakes. I'm just going to say, if you're on a bomb range, you should not be asking questions of an LLM. You know, there are times, it's like, okay.
Starting point is 00:20:27 Is this a bomb? That is not what I was asking. That's on me. I'm sorry you exploded. But there are, and it's not just me. The show is nominally about devices. If we can make our devices this much smarter, I mean, that's really different than now. Yeah, and I think one sort of way of figure out, like, what would your phone be able to do next year is to look at, like, well, what can your laptop do this year?
Starting point is 00:21:03 And there have been times when I was on a plane and I wanted to continue a project that I had been doing on Wi-Fi before that used a remote model. And it involved manipulating some software. and the model that was on my laptop was useful enough to continue doing that task even though I was in the air. Again, I'm not in the air most often, but there might be more times when people wouldn't worry about being on Wi-Fi
Starting point is 00:21:31 or if you have a different lifestyle where you're more often disconnected, I think the models, when you plug them into harnesses so that they can call tools are very useful. having very resilient access to them by having it all download on your computer is, I think we will find more uses for those things when those things are in our hands, but I don't spend a lot of my time trying to speculate, like, what will I be able to do with my laptop?
Starting point is 00:22:00 I just sort of play with those of the next scale up now, and when they eventually sort of, I guess, trickle down in scale, it'll be a nice convenience, but it won't be a shocking new revelation because I had been playing with it six months before while connected to Wi-Fi. And this isn't because our computers are getting faster. This isn't Moore's Law where things are getting faster. This is the models getting smaller. Yeah, so you can think of this as almost like, let's just say, LLMs are a software. They're a weird, goopy kind of software, but it's ultimately it's bytes on disk that you download and you run.
Starting point is 00:22:34 So it's not that it would be one thing if the ability for these models to do more and more on your phone was a result of us throwing away old laptops and buying new ones. but with your same hardware, you can download new files and do new things. So think of it as just making the software much better over time. I think there had been on the sort of, I don't know, what's the word, the frontier scale ones,
Starting point is 00:23:04 people were putting more and more compute into them. The models were bigger and bigger, and they were doing more and more. So it made it seem like additional capability always comes from scaling up the energy and the hardware and all that stuff. And I think those things were highly correlated on this particular high-end fringe, but for a given scale of what runs on my laptop or what runs on my gaming PC or something like that,
Starting point is 00:23:27 the trend had been the other direction. It would be like if a new version of some productivity app or a new operating system came out, you're like, oh, there's new features in the app that still runs on my computer, or the latest version of the operating system is now more secure or something like that. but I didn't buy any new hardware, and the model itself might be of a similar number of bytes to an operating system or something like that, a couple of gigabytes. I remember when this used to happen with games where you could get a new game and it would be optimized and run better. But I've gotten so used to insidification that my software, I just expect that it will, continue to run worse every few months.
Starting point is 00:24:16 But we're going to go the other way for a little while. One of the things that makes me excited about kind of the open weight and open source and all the cluster of AI stuff in that less corporate, more autonomous local mode is that it does feel like it has a response to the insidification stuff. It doesn't work by collecting your data.
Starting point is 00:24:41 someone has already collected some data and they cook it down into a model. But at that point, it's a file on your computer that works online. You can share it with people. You can back it up. It'll work in the future, but maybe in the future,
Starting point is 00:24:53 you'll be disappointed that it only knows what happened in 2025. Whereas the sort of the received understanding of things like ChatGPT and Claude is that they get better by stealing your data. And I think those things are correlated, but not like we can make better and better models without collecting people's data.
Starting point is 00:25:16 I think more often when models gain new capabilities, it's because people thought of a new use for them, found that they were bad for them, and then generated a bunch of synthetic data that represents that use case. So in that sense, it was inspired by human use, but I don't think that the whole economy around these things could have been a lot different if it had come from organizations, that weren't already starting in this, like, let's build a data flywheel or whatever the term is.
Starting point is 00:25:47 And some of the ability to run these things on phones and iPads is that for years, CPU manufacturers and Apple and people like, to some extent, Qualcomm, I've been putting, you know, dedicated silicon in there to run, you know, neural networks before LLMs were even a thing. Apple was like, oh, we have this neural engine which can do. whatever and they didn't really do much with it except some graphic stuff, but now that's sitting there and they continue to improve that so you can actually do inference on those devices, but on like laptops, you still need a GPU, still need boatloads of RAM. Well, it's not that you need it. It's just that the GPU makes it goes a hundred times faster. Which is, but a secondary effect of the capability of a given model size going up, and maybe that there are devices that like don't have GPUs and you might say, well, it doesn't
Starting point is 00:26:39 makes sense to do CPU-only inference because that's too slow, but it might be, if a sufficiently smart, small model, it might still be useful at that slow speed. So I think it sure is convenient and power-efficient and, like, nice to your battery and so on if you have a GPU, but it's not, it's, I think the GPUs are only ever a convenience. Right. It's a 100x convenience, but not, not a qualitative requirement. Sure. Because just linear algebra at the end. And I've, uh, I've seen many of people writing on Reddit where some new pretty gigantic open weight model will come out and the weights are like half a terabyte and they've maybe somebody has some sort of gaming class GPU but they have for other reasons a ton of normal system RAM and they will run that
Starting point is 00:27:30 giant model in CPU mode and they're like I have the latest deep seeker whatever it generates one token every four seconds but I have it at home and for certain kinds of hobbyists like having the real thing at home is fine and they don't mind like it's like this they really got to touch it to have it on their computer they don't mind that for for that particular spectacle that it was a hundred X slower than the one that's half the size that actually does fit on their there in V-ram. But if I had half a terabyte of RAM I would sell it by a house right now. Yeah. Okay. I would like to replace my with a shell script. I have plenty of podcasts, blogs, Git repo, talks I've done at my book. So I want to
Starting point is 00:28:20 just become an expert system and then go to the beach. Ah, well, I would twist it a bit to be this is entirely fictional scenario, any clients who are listening. I'm still here. I am human right now. I am very, very human. We'll do something about that. There is, in my line of thinking in the last few months, I have this word that I've used, trans agency referring, mixing this idea. There's sort of human agency like, what can you do to achieve your purpose in life and be fulfilled and so on? Are you blocked from that or are you capable of doing it? And then there's this other related word like agents, like what do the autonomous agents do on their own?
Starting point is 00:29:06 and it feels like those things, even though they have a lot of letters in common, feel almost totally disconnected. And so this idea of trans agency that I use is the sort of human kind of agency you get through leveraging these things that, whether they're autonomous or not, I mostly don't care.
Starting point is 00:29:26 But I think of this as building tools that kind of externalize your thinking so that there are, suppose I want to upgrade some software in this service that I'm running for the university. The first time I did a software upgrade, there were a lot of steps I went through it, and then I captured that in a playbook that my agent can follow. The steps are a little bit different every time because new features land, and sometimes it's just a completely automated unattended upgrade. Sometimes it's like they added a new feature that partially replaces something I've made custom.
Starting point is 00:30:03 Sometimes it can be adapted. sometimes it's not. So it's not full automation like it's just a shell script updated, but I have a playbook that can run mostly without me, but if it gets to a state where it's like, this deserves your attention, it'll stop at that point and I'll see a little notification. And I'm like, oh, that's fine.
Starting point is 00:30:22 Let's switch to theirs now. And then it continues the rest of the process without me. So it's not replacing myself or yourself with a shell script, but having something that remembers your purpose, your process and can continue the boring parts, the routine parts without you, but then you're immediately present and involved when it was revealed not to be what your automation assumed was the case.
Starting point is 00:30:48 Okay, so building a tool that will help me do the things I need to do, I guess that's fine. I guess that's a good starting place. I do tend to use some of the bigger LLMs. I use them basically to make badge scripts, other shell scripts. But then, again, I expected to handle the one-offs because it's smart. But how would I go about making my own model? I remember I took Udocity's autonomous car I've done D&N training I understand
Starting point is 00:31:35 convolutional neural networks well enough to sketch it out and I know that you know I need a ton of data and then I need if I'm I may need the data to be labeled I may not depending on all that so I understand that part of AI but I don't really understand how I get to a model. Or an LLM type model. That's very different. I think the L and the LLM, the largeness,
Starting point is 00:32:04 which in my mind I associate with, a model is large if it doesn't fit in your computer. Whereas the ones that fit in your computer, I'd call them SLM's just almost like you might say, what's the difference between a computer and a PC? It's like, is it something I can carry around, then it's a PC or something? Oh, well, then I want a TLM, a tiny language model.
Starting point is 00:32:22 Yeah. Does that mean sort of a phone-sized one? how tiny is tiny. We could start with phone sized. Okay, okay. That's still pretty big, though. It is still pretty big. Well, let's say it's...
Starting point is 00:32:34 Raspberry pie is actually where I want to go, but... The phone is good. We could pick a number like two gigabytes and maybe at a reasonable quantization, maybe a $2 billion parameter model. And then we could say, what are the off-the-shelf base models that are of that size
Starting point is 00:32:53 that other people have found to have somewhat relevant capabilities, and then you can say, I want to fine-tune that from my purpose. Or maybe I don't specifically want to fine-tune, but I want to make a system
Starting point is 00:33:02 that has the base capabilities of this 2 billion parameter model that someone else has, and a lot of my project-specific or personal, cultural knowledge. And a lot of the way to immediately bring that knowledge to a model is just to, like,
Starting point is 00:33:19 have it in context or have text files that I can read. there is this idea of context distillation. This is the same sense of distillation when people talk about model distillation of like, I want the capabilities of the big model and the small model. Context distillation is saying, I want the capabilities of the small model plus context in just the small model. Essentially, I want to cram all the information that would have gone into maybe a 10-page long prompt or something like that.
Starting point is 00:33:48 And it works by you have your, model, let's say it the small one, do whatever task you wanted it to do with that 10 pages of extra context. You record the history of what it did with all that context available and use that as the training data for fine-tuning your model to say the same things without that prompt present, as if that knowledge had been internalized. So as a byproduct of using it when the knowledge is provided in just plain text form, you build up this very situated data set of what it looks like to use it
Starting point is 00:34:26 if this knowledge was built in, but at the time it's constructed, it's not built in. It's sort of given up front as a huge 10-page prompt up front. What about a thousand page prompt? Then it comes down to, I guess, the size of the context window
Starting point is 00:34:42 supported by the model. When models didn't have very, big context windows. The only way to get information into them was via various kinds of direct training. But now that some models, including some, you could probably run on your laptop, have 100,000 tokens or a million tokens of input. That can be hundreds and hundreds of pages of, if it was just plain text. I don't know how to convert tokens to pages.
Starting point is 00:35:13 Depends on the formatting and other things like that. But the way that many models in the last year or so have these huge context windows means that from most of the things that you would have reached for training in the past, you can now just shove it into context up front. Your model will run slower because every output token is potentially influenced by all those things. But then this creates a record of how the thing I wanted to, like, you have these logs of how. how it should behave as if that knowledge was built in.
Starting point is 00:35:48 And later, you can decide, do I want to fine-tune this thing to operate without the giant prompt? Or maybe it's fast enough to just never fine-tune it and just keep that giant prompt around. And a big reason to do that is that there's so much churn in the models that maybe it's better for you to carry your prompt to the next-generation model than to spend a lot of work fine-tuning this model that just won't feel interesting or ineffective in six months. It sounds to me just trying to make this into a metaphor I can think about. I can add the context and kind of like in Python, I can add the context and it's live.
Starting point is 00:36:37 And the more I have, the more it's going to get a little slower. or I can compile it. Yeah, you could think of it as, imagine you're using some interpreted programming language like Python. You could have a lot of your functionality implemented in high-level Python, and somebody could say, well, I'm going to go rewrite it and see and bake it into the actual interpreter core. From the outside, other than runtime,
Starting point is 00:37:06 people probably won't notice the difference between those. the analogy is not perfect because things like context rot and how things don't pay attention to everything in context but there are many things where in terms of avoiding premature optimization like let me go rewrite it all in C or in Rust is probably not the right choice for many projects
Starting point is 00:37:27 and you should have left it in JavaScript or in Python or in whatever interpreted mode, especially since if for some reason there were new Python interpreters coming out all the time your high-level source code is portable to those new versions in a way that if you had built it into the core, you'd have to redo
Starting point is 00:37:46 that work again when a new base model came out. But if I wanted to miniaturize it, I probably do need to bake it into the core. I don't think so. The, like, that, let's say that 100 pages of extra knowledge you wanted to shove in there
Starting point is 00:38:03 is actually super tiny compared to the size of the model itself. It might be 100 kilobytes versus your 2 billion byte model. So the size gains from effectively inlining that knowledge are not very big, just because the models themselves, if we're imagining this 2 billion parameter one, is much bigger. It's a little bit tricky because we're not just talking about the size of the text. We're talking about the size of the encoded version in the key value.
Starting point is 00:38:37 cash. But again, I think the size of those things compared to the size of the model make it that if your models are sufficiently big, it's usually not the best strategy to try to fine-tune things into the model unless you were confident you were going to have millions and millions of people using it. We really need to get that last 1% performance by not processing the prompt each time, which can also be cashed, which you would do in a high-volume serving environment anyway. You said the word token. Yeah. What is it token?
Starting point is 00:39:18 I guess I tell people like a token is a word. It's not money. You can't one way to think of it as when somebody says, you wouldn't ask people, how many words did you burn this week or something like that? Because they'd be like, wait, did you mean writing words or reading words? Those are totally different activities. Yes. Or reading a word that you've read before, like skimming a document, is much cheaper than reading a fresh new paper for the first time that these categories of, like, these days people use this word token as if it were a unit of money because sometimes you pay for them. But I really want to distinguish like output tokens, which are things generated by the model.
Starting point is 00:40:04 one word at a time, let's call those written words, like word as a unit of measure of text or whatever the thing it's processing. And then you can talk about input tokens, and it matters whether those inputs are cash or not. Think of that as familiar versus unfamiliar text. And input and output tokens cost different amounts of money by maybe a factor of 10 in most environments. And there's even different costs for cash versus uncashed input tokens.
Starting point is 00:40:32 So when people say, I've burned a billion tokens on this project, I imagine they're just adding all these things together. It would be like if you did a project and somebody said, yeah, that was a lot of work. That was like a 50,000 word project. And then you could say, well, maybe there's some sort of typical balance to reading and writing that you do in a project. And you could sort of back out of that an estimate of what was generated versus what was read. but collapsing it all into a number of tokens, a number of words is often pretty misleading. But it's not far off to just say a token is a word, even though there's not quite one-to-one conversion factor. But does that mean that a token I use on my mini-me model is equivalent to a token used on the latest Gemini model?
Starting point is 00:41:25 They're both words. If I ask them, what do I do next? Not all words are worth the same as other words. I guess this is where it comes down to like a word is a unit of text for the purposes of storage, but also if you use word as a unit of labor as a unit of labor and you're thinking about the laborer of reading a word, the labor of writing a word, it really matters who's the laborer. There might be people who read things and don't get much out of it, or they're very sloppy and they're writing. So they might have written a high volume of words, but you don't really find much value in the result of their writing labor. And I think that one way to say there might be better or more serious writers or writers who are more thoughtful about it. And you could say, well, maybe you could say somebody doesn't write very much, but what they wrote was really important.
Starting point is 00:42:23 It might make it look like there were fewer words or fewer tokens being processed, but there was a lot more valuable activity going on. And I find this useful for thinking of the models that just chat with you versus the ones that go and use tools. Like, imagine you're talking to someone on the phone, and when they're talking to you on the phone, they can't do anything else. You can talk to them and they can talk back versus, like, they can go use a computer to do other things. And you might on your conversation say, oh, I had a 500,000 word conversation. But did you spend those words just talking to someone?
Starting point is 00:42:59 Or did you spend some of those words calling a different person or looking things up on Wikipedia? Like, there's much more wise ways to spend those words than only talk. That makes me wonder how these models are tuned because you mentioned reasoning models, which I think you said something about yapping before an answer. Yeah, yeah, yeah. I use yap to refer to the self-talk. before the final response. It seems like it would be to the companies that are selling tokens to their benefit
Starting point is 00:43:28 to make the models as chatty as possible during those reasoning faces, which most people don't look at. Oh. Like, are they making it? Where is the efficiency? I mean, you can be more efficient in your language, right? And these are language models. So you want to know, like, what are the incentives to get them to be more concise,
Starting point is 00:43:48 especially if they currently get paid per token? And we can see this with chat models now. There's stock phrases they use when they make a mistake or when they're correcting you or something. And to my mind, many of those are just a waste of words. Like, you could use to say that in two words, but you said a whole sentence about, you know. How objectively sorry you are that you cannot help me. Yeah, something like that. I think even though the model providers get paid per token for outputs and the outputs are the tokens that they get paid more per, per
Starting point is 00:44:20 per unit of word. The fact that a lot of the yapping or the reasoning happens before the response, it feels like there's a strong incentive to, if you want your service to feel snappy, you should train it to be concise because it gets people to their answer experience, because in some sense, even though at the API level maybe people pay per token, they certainly don't want to consume per token. It's not that if your thing used twice as flowery language, you appreciate it twice as much. In fact, you maybe appreciate it less than the concise version.
Starting point is 00:44:59 And this is why I kind of enjoy the experience of when you run things on your laptop, you're not so carefully counting the tokens and thinking, like, was that one particular token worth it? It's like you might be thinking how much energy was going into this thing. And but then you have to think, well, it takes some energy to run my GPU to generate the tokens, but also it takes some energy to keep the bright backlight shining at me. And if I'm done with my thing sooner and I can close it, I'm actually okay with it, having spent more energy per word to get me to being done so I can close my laptop and do the next thing. And then the whole thing is asleep. But to that, I don't remember.
Starting point is 00:45:42 Did I answer your question about? Yeah, no, I think there's a tension between, responsiveness and just spending a lot of time being loquacious or whatever. Or maybe it re-highlights the value of saying, if we say one token is about one word, we don't really care about the words, we care about the labor or the benefit that comes from those words. Like, maybe you put a certain amount of work into writing things and I get more or less value out of it in the labor I do of reading them.
Starting point is 00:46:14 maybe we should have been talking about labor all along, but we've used these units that make it seem like it's a kind of money like a coin or that it's all about the text, not the effect the text had on the system or the user or something like that. So what's Bayleaf? Excellent transition. Abrupt. Oh, no, it's related.
Starting point is 00:46:36 No, it is totally related because I wanted to talk about the human dimension. So Bayleaf is the name for a... I guess it's an AI service that I run for the university. When I say I run it for the university, it's sort of like, I made this for you. It's like a gift, but maybe you didn't ask for it. Like a puppy. The university did not want an internal AI service. They probably imagined that eventually they would buy big name services like Gemini and OpenAI and those things.
Starting point is 00:47:10 And I thought there ought to be a different shape. but something that was credibly sort of feature-complete equivalent, so there could be a way for the university to say no to Gemini, to say, we have Gemini at home and we like ours better. But I had to make that. So I haven't trained any models for Bayleaf. Bayleaf is mostly just a cobbled-ogether collection of big chunks of open-source software, open-weight models, or inference provided through a competitive marketplace
Starting point is 00:47:44 of people who run open weight models. So it's a single person project, but for the people in my campus community, I can give them a link to something that looks like ChatT. They sign in with their UC Santa Cruz account. They get to a page where they can chat. Their old conversations are on the left.
Starting point is 00:48:02 They can find the past ones. Something it has that the big name providers don't is I have the ability to make course-specific models, or maybe models the wrong terms. It's the same underlying model, but there is a course-specific prompt and a bunch of course-specific tools to read from our learning management systems and so on. So when I give my students a link to the course-specific agent or character, it knows that they're a student, it knows what class they're in, it can read the syllabus for them.
Starting point is 00:48:34 So if they're like, what was the late policy? The student can go read the syllabus, but that doesn't often happen. The chatbot can read the syllabus and say, well, on the syllabus, it says, that the late policy is 5% off or whatever. And to the degree that I can plug it into things like GitHub, it can then go look at the code that they have on GitHub and compare it against assignment requirements and be like, it doesn't look like you're ready to turn it in. You didn't do this step. Here's a clue. And instead of it offering to say, let me do it for you to pass these conformance tests because it knows that they are a student.
Starting point is 00:49:08 It knows that I have written a generative AI policy for the class. It has all this context of the course. When it offers to help, it offers to help in a way that was influenced by me authoring that experience. Me as a teacher, separate from me as the system operator, although I am the biggest user of my system. Half of the courses that have used Bayleaf have been my own courses, but not all of them. But another important part of the Bayleaf platform is that it offers API-level services so that people can run desktop apps. And I don't do that just for convenience. Part of it is I want somebody to say, well, I really love using Claude Co-Work.
Starting point is 00:49:53 Do you have that? I want to say, yes, we have Claude Co-Work at home. It looks the same and it's even a little bit better in certain ways. The reason I have people do that is that the conversation data, instead of being stored in my database where I can read it, I want there to be a plausible setup where I can't read their data. I don't want to just reproduce in-house in shitification where I collect all their data and I gain insights from it. I make new features for them. I want it to be like to push more and more onto the individual users that they are consumers of pretty low-level APIs, a very interchangeable model. so that if I decided to turn off Bayleaf,
Starting point is 00:50:32 they'd be like, weird, the Bayleaf has turned off. How do I get something that's like Bayleaf? You just use the same upstream providers that I do, and all these things are sort of protocol interchangeable. I guess that's what happening at the user experience level, but there's maybe more of a political dimension to Bayleaf as well. I describe it as a counter platform,
Starting point is 00:50:54 and on the landing page for it, There's like a video of me giving a guest lecture for a class, and the headline is anarchist AI infrastructure. It's purpose. Bayleaf is being positioned as something that is purposely too politically charged for the university to ever just absorb and bring in-house. But it's also positioned to be something that is very clearly not just a clone or a renaming. of what the big name companies provide. So it's trying to offer people who are highly critical
Starting point is 00:51:35 of the big name offering things to say this one is qualitatively different. But if you're a fan of those things, it is from a user experience similar enough that you can use it as a credible replacement. And the hope would be that the people who had an experience
Starting point is 00:51:53 with generative AI two years ago be like, this one smells different enough that I'll come to it with fresh eyes and have a fresh experience with it so they can develop a new relationship to generative AI, but one that has a story, a true, incredible story for how it is not just clawed under a new name. And the hope would be that the university decides to adopt 80% of it. But because it's not an official university service, I can just add whatever experimental things I want. And I don't need to get a
Starting point is 00:52:29 approval for it on it. We don't need to have budget reviews and stuff like that. The budget is just what I'm willing to pay for myself, which is about $200 a month in various IT services. And that fits into my hobby of self-learning and sort of offering it to the campus community. It's still pretty cheap, especially if people don't use it. It's not that I need tons of engagement for my project to survive. In fact, the more people who engage with it, I lose money there. I want to change some minds and get the university to say, yeah, we don't need to make a deal with Open AI, or we don't need to make a deal with many of these other companies that would take us down the incitification route because we have something that is, for the people who demand it, it feels
Starting point is 00:53:16 equivalent. And for the people who reject it, it feels like a plausible rejection, a way of, like, if you hate Open AI, how do we actually make it so Open AI cannot be? sell to one more university, is trying to be an answer to that. But I don't attend UC Santa Cruz. Oh, yeah. How do I get my own Bayleaf? Yeah.
Starting point is 00:53:41 So the Bayleaf project is sort of hyper-targeted at UC Santa Cruz, but sort of only as an example. The hope would be that when people look, oh, what Adam is doing with Bailey for UC Santa Cruz, I could do that for, I don't know, my school. Google, my church, my company, my city, or whatever. My house. Or your house, yeah. At some level, Bayleaf is just sort of the campus scale version of what my family has at home.
Starting point is 00:54:10 My wife uses it. My daughter uses it. They don't go to the Bayleaf.dev website. There's a constellation of services that are almost the same. So I'm sort of dog-fooding all the same open source projects. And Bayleaf is sort of the cleaned up community scale version of what I have as my friends and family service. And so the building blocks are things like there's a package called Open WebUI, which is something that I say it's only about as difficult as WordPress is to set up, which is to say easy to get started. You're not selling it here.
Starting point is 00:54:44 I'm trying to give a very calibrated, like it is something you can self-host. Open WebUI, as the name suggests, is just the user interface. It does not provide the LM inference. You need to plug in that from some other service where you. plausibly pay per token, input and output for that. It is, if you dug into that project, it's maybe a million lines of Python and JavaScript code. It's maintained by community.
Starting point is 00:55:09 I've made contributions to it, particularly to make it more accessible to screen readers and so on so it's easier for universities to adopt or any other organization that has accessibility requirements. So Open WebUI is one of the components that I bring in that you could run Open WebUI. There's like a desktop app version of it, but I think. think of it as software that is sort of intended to be used by groups. Another thing in my constellation is the open code command line coding agent harness. And if you showed somebody a screenshot of open code and a screenshot of cloud code,
Starting point is 00:55:47 they'd be like, well, they both look like terminal coding agent things. They have slightly different colors. They do almost all the same things. but most people are not living in the terminal anyway. So yet another major building block is this desktop app called OpenChamber, which you can kind of think of as being the desktop wrapper around the OpenCode Commandline Core. And if people have ever used the Chat JupT desktop app or the Claude desktop app, Open Chamber is sort of the open source answer to those.
Starting point is 00:56:27 But it's possible to plug these things together so that you spend almost no time in the terminal after initial setup. And you have a desktop app experience where assuming you're using providers that are zero data retention providers, the only copy of your chat data is in a database on your computer, not being accumulated somewhere, whether it's used for training or evaluation or other things like that. One of the things that I like about both open code and open chamber, they're essentially the same thing when plugged together is that there's this idea of projects that instead of just having like a loose chat with all the material on your computer, you pick a particular folder, a particular path in your file system,
Starting point is 00:57:12 and say, this is the root of my project, and you can have many of them. And when you start conversations in that project folder, there's this implied permission thing. You can do what you want with what's in that folder, and you have to ask me for permission anytime you reach out of it. So it sort of respects the effort you've already done to organize your file system to write documentation down in text files and other formats. And if you're really doing some wild system administration stuff,
Starting point is 00:57:43 you could dangerously dare to have a conversation in your home directory where all that stuff is in the scope. There are sometimes when you're, When you're like, I want the agent to help me with this sort of systematic reorganization. I almost never do that, but there have been a few times when I liked that it wasn't always super narrowly scoped. But more often, I'll have projects attached to various Git repositories that I'm using. And then that's an encouragement to like have the Git repository not just contain source code, but a lot of my other like, where is this project going? Where is it going politically?
Starting point is 00:58:16 What is our messaging for different groups? If that's written down in text files in this directory, and this is what I've done for the Bayleaf software repository, I can then have conversations about political messaging in the context of that Get repository. It's able to participate both in the like, okay, let's debug this feature, but also let's reword the landing page to be to incorporate this blog that came out from someone else recently. Okay. How do I attach it to an LLM? How do I attach it to my mini-me? LLM. Oh, okay. So in either OpenWeMUI or in OpenCode and OpenCamber, there's often somewhere in the settings, a panel called Connections or providers where you give it the URL and the API key for your LLM inference service, which doesn't have to be a third party. It could be something running on your own computer. And so if you had custom trained or custom fine-tuned your model, maybe you took one of the Mof-4 models and you trained it to do whatever person-specific thing you wanted it to do,
Starting point is 00:59:24 you would then customize these front-ends to use that model when they need inference, and then in any of these surfaces, your model would be at work there. Most of these apps allow you to configure many different providers, so you can say, well, in my drop-down list, I have the ones that run on my laptop, ones that I have so many different providers plugged in to sort of taste test these different things, including local ones. There was, what was it? There was something I was doing with the BitTorrent protocol that one of the major corporate models didn't want to do because it was too movie piracy adjacent. So it was convenient to have an on-device model that had been fine-tuned to never refuse anything to be like, help me make the protocol parser for this thing related to the distributed hash table.
Starting point is 01:00:23 But I didn't even have to start a new conversation. When I got to the point where, let's say, Claude or whatever was refusing, I could say, use the on-device model that isn't afraid of the BitTorrent protocol. It will do some steps. then I can switch back to Claude. And at that point, the Claude models think that they had complied, and then they don't complain about the thing anymore. Yeah. Chris.
Starting point is 01:00:50 Chris? So, what's OpenRouter AI? So, OpenRouter is, like, there's a bunch of different companies that provide pay per token access to LLMs. Deep Infra is one. Nebius is another one. But the trouble is that each one of these providers has like a different collection of models that they offer when new models come out. Some providers are happy to launch those on day one. Some get around to it a month later.
Starting point is 01:01:20 Some are unreliable. And so Open Browder is a sort of a middleman that takes a 5% cut to say, we support all these models, send a request to us, and we'll send it to one of our backend providers that is online right now, that has it. that if there are many models for a given model, we'll send it to the cheapest one, the fastest one, or whatever preferences you have. So it's a middleman aggregation layer.
Starting point is 01:01:45 It's a token broker. Yeah, you can call it a token broker. So I could connect open web UI to open router AI, and then I don't have to worry about models so much directly. Then you wouldn't... You still want to choose a model. You still want to choose a model, but OpenRouter AI would let me choose a model
Starting point is 01:02:06 and I wouldn't have to worry who is hosting that. I guess it might be, I don't know, to make like a food delivery analogy or something like you could order from the restaurant specifically or you could order from Uber Eats or something and they'll offer you food from all the areas of restaurants or something.
Starting point is 01:02:23 This is a poor metaphor, but there are good and bad things about having companies like OpenRouter out there. It is yet another middleman that could be taking your data. they do take some extra money, but they do make it so you don't need to have agreements and accounts with each specific provider in order to get the model coverage that you want.
Starting point is 01:02:46 I think for most people getting an account. I don't mean like a paper month subscription, but I mean like a pay per million tokens usage account with a particular provider. Like I used to be a direct customer of Deep Infra for two years. They had a pretty wide model catalog, and they were pretty reliable. But then there would be sometimes when they didn't have a model I want or they weren't the reliable. And at the moment, I wanted to reach for a second provider, I'm like, I'll just switch to open a router.
Starting point is 01:03:19 And then I get these nice charts across all the different providers, all the different models, what has been my usage over time. So with convenience or like in exchange for a little bit more cost and a little bit less privacy, there's a lot more convenience. but you can always cut them out by making a direct protocol connection to whoever they were getting the token generation from or your own local computer. Assuming you are willing to go through that extra little bit of hassle, startup hassle. Oh, you can tell your agents to set it up directly. Just giving your credit card.
Starting point is 01:03:52 But 95% of all these services just end up at Amazon, right? At some point, it's all AWS. Oh, no, it's not all AWS. Some of them are Azure. That's what I mean. That's the 5%. I think there's like this category of the hyperscalular clouds versus the neoclouds. And there's actually by volume a lot of compute happening in the neocloud space,
Starting point is 01:04:14 which is like distinct companies that actually have different hardware. So it's not ultimately AWS. ADWS has a subpiece of it because Amazon does everything. That does just like, like there are many providers where you can, in the same way that you could get a virtual machine through EC2, there are ones where you can say, give me a virtual machine connected to a GPU and I'll run LLM inference myself.
Starting point is 01:04:40 Or if the only app you were going to run on that virtual machine was an LM inference engine, you can say, just give me that as a service and I don't want to do the system in for it. So it doesn't necessarily go to AWAS. It could go to another, I don't know how to say, IAS, infrastructure as a service out loud, IS or whatever. But I think most of
Starting point is 01:05:04 Most providers are themselves not owners of the GPUs. They are getting GPU cloud hosting from one layer below, and then they sell you token generation service above that. But as with all these things, you can sort of cut out the middleman many steps and say, well, my computer has a GPU. Surely I can do it. It's like, yes, you can. But to the limits of now your laptop gets really warm,
Starting point is 01:05:31 and it gets uncomfortable, and it is sometimes nice to be able to use these things without your laptop getting uncomfortably hot. A listener, Exploding Lemur, who's been on the show, picked up some old AMD V620s for semi-cheap, a five-slot PCI-E switch with slim SAS ports to uplinked to a computer. I really should have given you that this is a text. He bought a box to put some GPUs in. he bought some old GPUs that are good for inference with a lot of V-RAM. How much performance can he expect to get out of a hundred and twenty-eight gig V-ram and five-year-old GPUs? A lot.
Starting point is 01:06:11 Probably, yeah, having access to, it's almost like when people say, oh, you need GPUs for LMs, it's actually, you need V-RAM for fast-infference. And what's further complicated in that case is that, depending on how the different GPUs are wired together. It may be that they don't each have fast access to each other's V-RAM, in which case, even though in aggregate, there might be 128 gigabytes.
Starting point is 01:06:39 If each one is only 16, your maximum speed is limited by what you get on. There are complicated ways to spread it across devices, and I haven't done that because I'm not GPU-rich in that particular. I've got a Mac laptop that has the one integrated GPU and a lot of unified RAM, which has given me a sort of
Starting point is 01:06:55 a taste of what it's like to have a big, a weak GPU and a lot of V-Ram. And that has let me play with a lot of stuff on device that a lot of people who don't have that integrated RAM setup are much more limited by... Because the GPU with a lot of RAM is very, very expensive. Yeah, yeah. Whereas a Mac, it shares the RAM with GPU and the system,
Starting point is 01:07:17 and you can say, well, give me 48 gigs of your 64 gigs for the model. Yeah, these things are actually more I.O. intensive than they are compute-intensive, but you wouldn't guess that from like the public discourse. They would think that it's all about the compute. It's like, well, it's correlated with the compute, but it's actually bottlenecked on this other thing. But how many LLM seem to be graded by how many billion terms they have, which isn't really work because you can get smaller versions that are just as effective. How do you? I think at any given snapshot in time, the bigger ones are almost always better.
Starting point is 01:07:58 However, if you make a snapshot across or a slice at a given scale across time, the capability radically increases over time. Like last year's 4 billion parameter model is nothing compared to current, or what am I trying to say? At a given parameter size, capabilities have still been improving. Yeah, yeah. And that in getting lost in trying to gesture in 2D in the audio, medium, so I won't try with that. But I
Starting point is 01:08:32 think for a somebody has a particular machine that has many, probably many medium sized GPUs. They want to know what can they do with it. Yeah, or my following question to that was how do
Starting point is 01:08:47 we even identify what we're looking at? I mean, what scale are we talking about? Which one should he download? Oh, I think the heuristic would be figure out what is the
Starting point is 01:09:04 largest single chunk of V-RAM that is accessible. Let's say that's 16 gigabytes or something like that. And then you could say, okay, what is the highest number of parameter model that fits in 16 gigabytes without getting lost in quantization too much? Let's just say
Starting point is 01:09:22 that one parameter is one byte, which is true at the 8-bit quantization. level, roughly. So this is saying that they could run a 16 billion parameter model. And then you can say, okay, what is the newest 16 billion parameter model? That will give me sort of the best intelligence within the space that I'm allowed to give it. And it gets squishy in some ways.
Starting point is 01:09:48 It's like, well, is it better to run a highly quantized big model or a less quantized small one? It usually don't get to make that direct trade-off because usually one of those models is new, in which case you should just pick the newer one. And for comparison, like a frontier model would have how many parameters? That's the trickiness. They don't tell you. But people are guessing that they are in the low trillions, somewhere between 1 and 10. So that's a couple orders of magnitude.
Starting point is 01:10:19 Yeah. And to compare that with, like, for some of the things that when Chat 2D did it two years ago, people thought were super disruptive. of those things you can now reliably do with something that's multiple orders matters to smaller. Sometimes there's this metaphor I have of if it were the case that one billion
Starting point is 01:10:38 parameters weighed one pound and LLMs were vehicles, you could say there are some vehicles that are like it's a one ton work truck or something like that. Then there are other vehicles that might be like a motorcycle that's maybe 350 pounds, or I don't know how much old motorcycle weighs,
Starting point is 01:10:56 then there could be a light bike, or you can think of like a shoe, which might be like a one pound vehicle. And except for this thing where the vehicles get better over time at a given size, I think this saying, one billion parameters is one pound gives you a sense of these orders of magnitude to say there are things that some people are using giant heavy work trucks for, that you're like, you know what?
Starting point is 01:11:22 I don't work on a construction site. Do you have an e-bike? Do you have a skateboard? What can I do with the shoe that's in my pocket or that's on my foot or something like that? And the surprising thing that, like, the skateboard of this year outperforms the work truck of last year is sort of the fun part. But this gives you this heuristic of mostly shop by size and pick the biggest new thing. Like, for some reason, shoes are getting better and better over time. Just say, what is the size of vehicle that I want to go with?
Starting point is 01:11:53 And then say, what's the newest vehicle of that size? roughly. I think it's a good metaphor too because it's probably diminishing returns at some level of, okay, we have a trillion parameters. Well, great. Yeah. Is that better than $100 billion? I don't know that anybody can really quantify that.
Starting point is 01:12:08 And, you know, the efficiency, like you say, if somebody has a, you know, a Ford F450 that they drive around to the grocery store, they're burning a ton of gas. Yeah. For no. And I think a lot of the discussion of LMs and energy use come from this tendency to only talk about them. most extreme top end. Yeah. When if people are actually able to sort of shop among different sizes of models, people might be like in the same way that people are like, I'm proud to bike to work
Starting point is 01:12:37 instead of driving. Like, you could say, I'm proud to use 100 billion parameter model, 100 pound thing, instead of the 1,000 pound or the 10,000 pound one. But at least the way these things are currently sold, they sell you some... The biggest and the best. The biggest and the best. They don't say how big it is, and it's always changing. So you're
Starting point is 01:12:56 effectively, it's, I don't know, you're leasing the biggest heavy work truck rather than I like in the open-source, in the open-weight space where things are less convenient, but you get a lot of choice. It's very easy to make these intermediate choices and including
Starting point is 01:13:11 pulling down the usage, just multiple orders of magnitude, to say like, what I used to use the truck for, I legitimately can use the skateboard or something like that. Or another, sometimes I use that animal metaphors of like, the elephant could do this, but what can the kangaroo do? What can the mouse do or something like that? And the fact that the mouse can do something that the elephant, I don't want, the vehicle thing is a little bit easier because at least you ride vehicles to achieve reach destinations.
Starting point is 01:13:40 And the same way you can do a recreational drive in your car, you can have a recreational bike ride. There are times when I, I guess I can say recreational LLM use. Like, I want to learn about some new thing. I'm like, let's make a prototype of it. Somebody else had a YouTube video. Let's download the transcript of the YouTube video. Let's get it working on my computer. I wasn't really trying to achieve a task.
Starting point is 01:14:01 I'm like, let's go play and get new magical powers as a result of this playtime. And when I say get new magical powers, I often mean, let's get text files on my disk so that the next time I want to play with this, we've got like a playbook of how to do it the first time. Whereas, whenever you try out some new software, you could get lost for hours, getting the right version of the digital. dependencies and all the stuff that's not run down. I hate that. But if you can sort of stumble through it once with your agent and then you sort of save this recipe, the recipe could potentially guide you like maybe a new version of that thing comes out.
Starting point is 01:14:36 Maybe they change their build system. There's a chance the recipe will take you 80% of the way through that new version. Whereas if you would only like kept the binary on disk, that's much more brittle than this more natural language workflow version of it. But there are times that I do feel like I'm. going for a recreational thing, or it's not so much about just the enjoyment of it, but I want to gain a new skill.
Starting point is 01:15:01 Me and the machine, as one fused unit, are learning through experience, but it's not machine learning like gradient descent, not dataset. It's different from just me learning. It's me and the machine, as in me plus my text files, as a fused unit,
Starting point is 01:15:17 are learning these new things. Does using the open source, models and system that you've built make you feel less dirty than using the big models because I could use that. Oh, it does feel different to sort of run your own stack for these things.
Starting point is 01:15:41 Certainly in the circles that I'm in and shitification comes up all the time. So knowing that my thing does not depend on people's engagement to survive. It feels a lot different this way. And in fact, that better, more wholesome feeling often makes it easy to justify like, oh, my thing made a mistake. But I know that's because it's two orders of magnitude smaller than the other one.
Starting point is 01:16:12 So I'm willing to give up some kinds of comfort in order to get other kinds of comfort. Changing topics, which I'm going to do rapidly twice in a row. But first, aren't you a professor of computational media, like, games and music and stuff? I'd say we are trying for, since the department is about 10 years old now, we're trying to figure out how to communicate to people what even is computational media. And at the time the department was created, it almost felt like, Oh, computational media is just a euphemism for video games. And we've always been trying to tell people like, no, no, it's so much more broader.
Starting point is 01:16:57 But then each thing we pointed to often, what's like, well, that was an existing area. And my thought in this current moment with all this AI stuff is like, what's under the hype of AI stuff? Like, all these new capabilities and this Doom is like, it's just a new computer-mediated, a computer-based medium for expressing ideas between people or between. you write down notes this week and then a month later passed you as communicated something through the computer. And now you're like, this thing, you're like, I used to remember how to do this. Well, your agent with your notes from the past remembers how to do that now.
Starting point is 01:17:35 So I guess we're trying to say underneath all the AI hype and doom stuff is computational media and it's kind of convenient that all this stuff grew. People are taking it very seriously, too seriously in many ways. But it is a great example of a thing that's in. It's based on a computer. There's even audiovisual experience to it, but it's not a video game. And it should be some things that are authored by end users. There should be indie communities of them.
Starting point is 01:18:00 It shouldn't, like Hollywood movies, only come from big-name people through official channels. There should be – think of it as like – maybe you can make an analogy. Like, if OpenAI is like Netflix, the big-name proper licensed provider, I want like the long tail of weird stuff on YouTube. I kind of like the NPR and PBS. Oh, yeah, the sort of like, well, I guess NPR and PPS are getting... Or institutional ones. I also like institutional ones too, but we're in a time when those institutional options are getting less and less funding.
Starting point is 01:18:36 But a thing that I want to do, and I apparently didn't get enough time to do it this summer, is make contact with, like, city library systems to say that in the same. the same way that people who don't have internet access on their own can reliably go to a library and sit down at that library computer, use the library Wi-Fi from their phone. AI services could be provided there too that didn't steal your data, didn't cost $20 a month, that weren't overwhelming, that were supported by and understood by the staff. There could be whether it's national or state or city scale support for. going the sort of open source, open weight route on many of these things. And obviously, I want schools to be able to do that as well. Like, I like there to be abundance of different ways, the way that, like, you can get Wi-Fi at home.
Starting point is 01:19:32 You can get Wi-Fi at the grocery store. You can get Wi-Fi at the airport. You don't, it's not that we've figured out that there's one national-level Wi-Fi network and you sign in with your Social Security number. Like, there are so many intercompatible systems. that we feel that abundance, even though each organization provides it in a slightly different way. And then changing topics again rapidly. Do you have advice for people with ADHD?
Starting point is 01:20:00 Advice for people with ADHD. You talked about your struggles with it, some on your personal site. Yeah. I found that specifically, I'll connect it back to LMs and apps like Open Chamber. There are some things that I feel like it's hard for me to, like, I know it's important. I know I'm going to get in trouble if I don't do it. But I'll do anything. I'll do another layer of cleaning the kitchen.
Starting point is 01:20:27 I'll sort of the laundry. I'll scoop the cat litter instead of doing that thing that's important. Have you ever tried to learn five different languages on Duolingo at the same time? Yeah, yeah. And what were you procrastinating? I think I was procrastinating from some overdue grading. And I was trying to learn Mandarin, Vietnamese, Korean, and something else at the same. time.
Starting point is 01:20:49 Swedish. Swedish was in the mix because we had gone to Sweden recently for a trip, so that was just, yeah, that was a weird time. But specifically on the getting executive function things, there are sometimes when, like, I can, I can, there's lots of things I can do, but there's many things I feel like I can't start, but sometimes I can just hold down the voice to text button on my computer and be like, I have email my inbox from the thing about the, insurance thing we need to email, those other people, can you read our policy and figure out what is the right way to respond to them? Can you draft the email? It's like, don't send it, just leave it in drafts, yeah. And then I'm like, oh, okay, I started the project. Like, I had to dump enough of the task, but because I had already done the work to plug it into these things, like, if you just had a chatbot, it would be like, I would love to help you with this letter,
Starting point is 01:21:41 but I don't know anything else about what you're doing. So it really matters like, it can go read the insurance policy. It can go read the email. I can go write the drafts. Like, I can start it in a folder where the right constellation of things are pre-given, and that has been a way for certain things that I know I would have put off for months to sometimes just, and for some reason, the voice-to-text part of it as well, I just want to be like, when I finally get that willingness to do the first three sentences that has unblocked me, a bunch of times. And there's many more dimensions to ADHD other than inability to start things.
Starting point is 01:22:22 And many people face inability to start things for many other reasons. But I'm glad I now have a constellation of cool open source software that lets me make some of these super important but annoying things like not, it doesn't do the whole thing for me. I wouldn't want it to negotiate with the insurance person directly or something like that. but it suddenly became a much more tangible thing when I'm like, I just need to tweak this email. It has already, like, done, gathered all the important stuff into one place. And even if I end up, like, deleting it, redoing it again. It's because, like, oh, I wouldn't want the agent's email to go here.
Starting point is 01:23:02 I'm just going to fix it where we had our e-bikes stolen recently. So we learned about insurance reimbursement stuff. That's on my mind. So before we let you go, do you have any resources for how to start my own open source LLM projects? How do I get involved with this? How do you get involved with it? I guess part of it is gaining some personal practice with these tools. And some of that comes from if you have, even if you're using the big name corporate ones,
Starting point is 01:23:38 you can ask them to help you set up these things and be able to. like, they are perfectly happy to help you build their replacements and debug them. And when they fall apart, they fall apart many times, get you back on board. So you can take risk to build things that aren't resilient, that are flaky and personal. So it's like if you have a little bit help, you have the ability to jump into new spaces and make mistakes. There is a subreddit called Local Lama, where there's a huge range of expertise there. There are people who just showed up yesterday. There are people who have been involved for years.
Starting point is 01:24:11 But one of the things I get from that is like enthusiasm to try new things, many people's opinions on new models when they come out. So you don't need to try every one of these things. You can see other people reviewing it and throwing in their opinions. I check that like twice a day. It's actually the only subreddit that I check. But those people will at least filter down the things that might be worth caring about. And do you have any thoughts you'd like to leave us with?
Starting point is 01:24:44 This phrase that I use, the me in the machine or the you in the car, this thinking of you and the tools around you is a thing that could do more than just you can. And it's not because the tools did it all for you. But thinking of the you in the machine as a useful unit that you can go out in the world, gain experience. and all the particular software that you're using, which models, which harnesses, those things can sort of totally change around,
Starting point is 01:25:13 but the sort of you plus the notes you've accumulated from experience is very forward compatible. So you can sort of not bother to update what model and harness you're using for six months. Your text files will go with you when you eventually decide to update those. But it is useful to be using tools that get things to be written down on your computer as a result of the experience. And that's like valuable text that is expressed in words. You can measure in tokens.
Starting point is 01:25:45 And importantly, that text, I don't know, it's a crystallized form of your labor that you can carry with you around over time that is actually somewhat independent of, largely independent of the model that it came from. Sorry, I'm relating this to a book that it sounds, I mean, this accumulation of notes that become part of your identity, history, media? It seems like you could make a book from that.
Starting point is 01:26:20 You could. Or there's many people who say, like, oh, it's really important to have a journaling practice. And you're like, well, I don't want to. I'm too lazy. I've got other stuff. But now it's like really easy to have a journaling practice. You can write in your system prompts, hey, I want to have written evidence of anything important that we learn together, and you go do something, and you'll see it just like go right to files. Like, today we figured out how to build that particular project from source. Today, we learn that this driver needs to be this way. And just as a byproduct of use, you, the extended version of you, will take notes on a lot of this stuff.
Starting point is 01:26:58 And then after a while, you can say, let's reorganize all that stuff. Or wasn't there a time three months ago when I did this? and it will go look through the text files that in that case are almost entirely agent written, but they are written as a direct byproduct of these learning activities that you go on. But it comes down to, did you ever tell your agent that you wanted to write down the experiences, that cost two sentences to write into a preferences file? I feel like there are a lot of things I don't know about how to use LLMs effectively. I think someday we'll teach it in schools.
Starting point is 01:27:34 or implicitly the way that most people know how to operate web browsers, and then secondarily, most people know how to operate the cloud services that you access through the web browsers, even though we don't have a how to use the web class in schools. There is, like, middle school age, you just get enough experience being expected to do stuff on the web, and I think we'll get that sort of middle school age exposure to use these things, but we're all new to it now.
Starting point is 01:28:05 And some people are much deeper into it than other ones, only because they're two years in instead of one year in. But 10 years from now, that's going to be a difference between are you 10 years in versus 12 or something like that. Our guest has been Adam Smith. You see Santa Cruz Associate Professor of Computational Media. Thanks, I'm happy to be here. Thank you to Christopher for producing and co-hosting. Thank you to Kathleen and Joelle for their help with lightning round.
Starting point is 01:28:36 And thank you to our Patreon listener Slack group for their questions. Thank you for listening. You can always contact us at show atembedded.fm or at the contact link on Embedded FM. Although I have to admit if you do right now, well, Adam's point about emails was very relevant to my current email box. And if you are requesting to be a guest on the show and you're not an AI, go ahead and send me a another email because I don't really know about this set. There's sure a lot of them. And I guess now I'm supposed to have a quote to leave you with, but I don't. So, Adam, complete one project or start a dozen.
Starting point is 01:29:16 Tens of thousands and start hundreds more within the starting of the other projects recursively forever.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.