How I AI - I tested Grok Bot, Grok 4.6, and Cursor Origin - here’s my honest take

Episode Date: August 18, 2026

This week I’m doing a solo breakdown of everything xAI and Cursor have shipped recently, including Grok Bot, Cursor Origin, and the Grok 4.6 model. I set up five Grok Bots, ran Grok 4.6 through my C...laire Weighted Index against GPT-5.6 Sol, Claude Sonnet 5, and Opus 5, and spent time actually using Origin as a GitHub replacement. Here’s what’s worth your attention, what’s overhyped, and where I’m personally putting my time.What you’ll learn:The one Grok Bot feature no other agent platform has shipped yet, and why it made me actually use the productWhat a week of real Grok Bot use revealed, and why I still reach for my OpenClawsWhether Cursor Origin is a GitHub replacement or just a pretty redesignWhere Grok 4.6 landed on the Claire Index, and the one category where it genuinely surprised me—Brought to you by:Bolt.new—Turn your idea into a real productJira AI SDLC—Get your tokens’ worth with Jira—In this episode, we cover:(00:00) Why everyone’s quietly switching to Grok(01:52) Grok Bot overview and setup(03:22) My 5 Grok Bots(04:30) The killer feature: multi-account connectors(06:07) Grok Bot’s virtual machine and how it actually works(06:41) Experience overview(07:35) What I don’t love about Grok Bot(10:08) Grok Bot use cases and my honest verdict(12:20) Cursor Origin: the agent-native GitHub replacement(13:47) What Origin actually looks like in practice(14:59) Why I’m not switching from GitHub yet(17:42) What would get me to move over(18:52) Grok 4.6 and the How I AI Vibe bench(20:41) Claire Index results: where Grok 4.6 ranked(23:03) Design evals: where Grok surprised me(25:00) My conclusion and how I’m splitting my time now—Tools referenced:• Grok Bot: https://x.ai/bot• Cursor: https://cursor.com/home• Cursor Origin: https://cursor.com/origin• OpenClaw: https://openclaw.ai/• GitHub: https://github.com—Where to find Claire Vo:ChatPRD: https://www.chatprd.ai/Website: https://clairevo.com/LinkedIn: https://www.linkedin.com/in/clairevo/X: https://x.com/clairevo—Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Transcript
Discussion (0)
Starting point is 00:00:00 I don't know if you know this, but the whole open AI versus anthropic thing is old news. Everybody I know is secretly becoming a grok boy. Yep, ever since Cursor was acquired by SpaceX 4, I believe it was $60 billion American dollars. Everybody's been really excited about the new products released from the Cursor team, and in fact, I'm hearing more and more positive things about the GROC models for coding and the GROC models in general. In today, episode of Howa AI, we're going to go through all the releases from the cursor, XAI, SpaceX, whatever, the Elon's cinematic universe of code builders and model builders. We're going to go through all their products. And I'm going to tell you what I think of Grockbot,
Starting point is 00:00:47 origin, the new GitHub competitor from Cursor, as well as the GROC model. Let's get to it. This episode is brought to you by Bolt.new, the AI app builder for people who have ideas, and want to ship them. Most AI tools spit out code that looks great in a demo and falls apart the second you try to do anything real with it. Or they lock you into their own platform with no real way out. Bolt is different. You describe what you want to build, a startup MVP, a landing page, an internal tool, a side project.
Starting point is 00:01:20 And Bolt generates production ready code in minutes. Connect Stripe or any other MCP, hook up your domain and deploy it live. Founders are using Bolt to build businesses doing real revenue. Product managers are shipping prototypes their teams actually use. Designers and marketers are launching campaigns without waiting in line. Anyone can build. Engineering can ship. Everyone wins. You just need an idea and a weekend.
Starting point is 00:01:47 Check it out at bolt.new slash how I AI. I'm going to start with the most accessible product on our list today, Grockbot. In case you missed it, the cursor slash XAI team released GROCBOT, which is a chat style agent in a desktop in mobile app that you can use to do knowledge work for you. It's very similar to an open claw, but it's simpler, hosted, and easier to use. If you look at GROCBOT just from a UI perspective, and I'm just pulling up the marketing site right now, we'll look at my GROCB in a minute. It is definitely pulling from the IMessage terminal streamlined chat experience. And it's doing something I really love. Not everybody loves this, but I love a multi-agent experience.
Starting point is 00:02:35 I want each of my agents to have a job. I want them to have a name. I don't want one agent to rule them all. And the team has really leaned into this ethos with Grockbot. Now, I've been testing Grogbought. I had a couple days early access. And then I've been testing it, earnest. over about the past week. And I think there are some pros to Grockbot and I think there's some
Starting point is 00:02:57 cons to GrogBot, but I do suspect that GrockBot is going to catch the attention of cursor customers who want an all-purpose agent for their employees. So let's go into my Grockbot and show you a little bit about how it works. Now, as I said, this has this very like I message style experience. It's a left bubble and right bubble and you just chitty chat with your Grockbots. I have set up a couple of GrockBots in my own account. There is Praddy McProdd, my product manager, GrockBot that's hooked up to chat PRD. I have one of the out-of-the-box GrockBots commitment tracker that just keeps me honest with things I've told people that I will do for them, basically emails I need to reply to.
Starting point is 00:03:41 I have a moneymaker bot that chases invoices and payments and sales deals. I have a case study buddy because I'm doing a bunch of case studies. and I have a Grockbot that monitors all my data for chat parity and tells me trends that I should pay attention to. So I think the magic of Grockbot, if I had to tell you one thing that is the magic of Grockbot, it is plugins. So I've always thought that Cursor was the best MCP client. I still, to this day, when I want to connect to different MCPs in my stack with still default generally to Cursor, I thought the connector experience was really nice and I thought the harness did a really nice job of talking to all these systems. They have brought that great plug-in MCP experience into
Starting point is 00:04:26 Grockbot, but there's one killer feature, which is for any one of these connectors, whether it's Gmail or Slack or an MCP, you can connect multiple accounts to Grockbot. So if you have two g-mails, you can connect two g-mails. If you're in seven different slacks, you can connect the different slacks. This is something that Codex has not figured out. Claude has not figured out. No one seems to figure out that I have approximately a dozen email address and Slack accounts and different things that I need to log in, even if it's the same service. This experience of setting up multiple accounts per connector. Huge, huge, huge benefit. You could like package it up and send it off. It would have product market fit with me. As you can see,
Starting point is 00:05:17 here. I have already four email addresses connected to Gmail. I have my chat parity one and a couple other business ones that I'm working on. Having all four of these connected and being able to traverse all those accounts in Grokbot is a killer feature. Just thank you to the cursor slash XAI team for doing this. If you did this and this alone, I would be very happy. But they did not do that and that alone. Gropbot ships with tons and tons of out-of-the-box plugins. These are generally the ones that were available in Cursor as well. And so I've installed your basic productivity ones. I've installed GitHub and I've installed the chat PRD plugin. And really, I've been using Grockbot as a way to just traverse and use my data. So again, it's like a little bit of a fancy MCP client, but it does
Starting point is 00:06:08 come with something special, which is every Grockbot comes with, a computer and so you can actually go into this virtual machine and this is proddy mcprod's computer it's not very exciting but it can use chrome it can use terminal and it has some files and so again this is like a little bit of a open claw light where it has a virtual machine that virtual machine can use the web and it has connectors and mcps now What do I think of the experience using Grockbot? Well, I will say setting up a Grockbot super simple. You can simply create a bot. It will start up that machine. It will start up a bot and then you can just tell it what to do. So this one just got stood up. What's the first thing I want you on? And I say, I want you to be a family manager bot that keeps track of my personal email and personal calendar. Can't spell count. And again, it's going to go ahead and say, great, I'm going to find all your connectors. I'm going to set it up and I'm going to learn myself.
Starting point is 00:07:19 Really, really simple, streamlined onboarding experience. And it's going to go ahead and give me pretty good responses back and do work on my behalf. It's pretty chat-based, nothing fancy, but it's genius is in its simplicity. Now, what I don't like about Grokbot, well, I don't like. like that it's so simple. And I don't like that I can't hack it. If you look at my open clause, they are chaotic and technical and impossible to use and high maintenance. And I love them. I love them. Um, GrapBot, I have a little bit less control over. I manage them a little bit less tightly. And so I love them less. You know, like my children.
Starting point is 00:08:11 the joy is in the challenge. And I will tell you, Grockbot has not challenged me because it works simply and it works out of, out of the box. The other thing I don't like about Gropbot is you don't get to pick what model you're working with. You really actually don't get to have that sort of like sole.md control over what Grockbot does, how it's configured, etc. And you know me. I just like to craft out as if out of clay my agents so I like that experience I don't think everybody else does so if you're looking for that like highly tunable highly hackable highly transparent set up for your agent grokbot is not that open claw or hermes or any of those is a little bit more like that and then of course it doesn't run on my machine I don't have control over it it's a third party system some good some bad that being said
Starting point is 00:09:06 it works and the connector and kind of like ux affordances are quite nice i would say just related to that the last thing that i don't love about grok bot is whatever model it's using and maybe it's grok maybe it sounds a lot like clotslop is i just no vibes the vibes are bad it's like i don't want to hang out with this person doesn't really make me cheerful. I have tuned my open clause to be exactly who I want to talk to, what I want to talk to them. This, I'm getting like too much of that like classic slop, not this, not that slop, terrible names for products like a scope knife. Just stuff that I would never say. And so I think the personality tuning and the voice tuning across all these harnesses are really important. But they're even more important if you are going to do this multi-agent strategy where you're
Starting point is 00:10:03 and have people name their agents and have them do work for you. A couple of use cases that I think you can use GropBot for again. I've made, I like to make my agents have roles. So I've made a PM one. I made a data analyst one. I made a deal desk one. I made a finance one. I'm making a family one.
Starting point is 00:10:20 I think just thinking that way, what kind of teammate do I want? And then giving that teammate the tools it needs and giving it the instructions it needs is a great way to design your agents. Now, I would demo more. but I actually think this is all about workflow. So the product itself is super simple. I think the use cases are what make it interesting. And then the connector experience being top tier, I think is going to help the cursor XAI, SpaceX team,
Starting point is 00:10:48 get their claws in the enterprise use case. So I'm really curious, drop in the comments. Tell me what your use cases are. I will continue to share ones. And I will let you know if I decide I'm going to migrate any of my open clause over to Grockbot. if you want me to move over SpaceX team, just make Grockbot way harder to manage. And then apparently I will fall in love with it and never use anything else.
Starting point is 00:11:12 I'm going to keep eye on this, but TLDR, super simple, great connectors. Use your agents as employees. And if you like a bad time, stay with OpenClaw. This episode is brought to you by Jira by Atlassian. The teamwork graph in Jira delivers 44% more accurate agent results with 48. percent less token usage. That's a huge difference when working with AI coding agents like Claude, Cod, Cursor, Codex, or Copilot. The hardest part of shipping with AI isn't the code. It's the context. What's the right ticket? What did the specs say? What got decided in Slack?
Starting point is 00:11:49 The teamwork graph pulls all of that from across your entire stack, from Jira and Compliance to GitHub, and feeds it directly to your agents before they write a single line. You assign the work, The agent gets everything it needs and a PR surfaces when it's ready. No digging, no contact switching, no broken flow. Same team, smarter agents. Try them free at jira.dev. That's jira.d-ev. Okay, the next release from the cursor team is Origin,
Starting point is 00:12:26 the GitHub replacement announced at Cursors conferences here and finally released in Early Access today, which will be a couple days ago when we get this episode live. So Origin is in Early Access beta, and they've called it kind of code base inside the app. So if you're in the cursor web app or you're in the desktop app, you're really talking about code base, which is, of course, the right primitive for something that is code hosting. Now, the whole value proposition that cursor has put in front of us in terms of why they should build a GitHub replacement is they are basically building a agent-native GitHub replacement, which means that it's got all the Git primitives.
Starting point is 00:13:10 It's got code. It's got diffs. It's got pull requests. It's got all that same Git primitive stuff. But it's built in a UI way. It's built in a U.S. way for agents to collaborate with. And in particular, of course, the cursor cloud agents and the cursor desktop app. and the cursor or CLI. So it's it's Git. We're all stuck with Git. We've agreed on Git. Now have we
Starting point is 00:13:35 agreed on GitHub? And it seems like the tides are turning. But cursor's bet is that we are going to want a agent native GitHub and that GitHub is not going to get there fast enough. So what does that look like tactically? Well, Origin has repos. You can of course import GitHub repos. I'll tell you a little bit about my experience there. It has pull requests, which you can see. Agents can reply to comments. They can be assigned as reviewers, all those sorts of things. And there's a small set of extensions, CICD extensions, for example, like build preview branches in VERSL for those who want it. Okay, so what does that actually look like? Well, you go into the cursor web app and you click code base and then you basically import your GitHub repose very smart.
Starting point is 00:14:28 Now was not smart to release this the day that GitHub had a major outage or maybe it was genius to release this the day that GitHub had a major outage. So I had some trouble with this GitHub import, but all you do is click sync from GitHub. You authorize your GitHub account, pick which repositories you want to sync in, and then they are here. As you can see, you have kind of the same primitives. You have code. So this is your code view. You have pull request. so you have pull requests and then you have some settings. What I would say is with the GitHub integration, it's really, it seems like it's just a wrapper on the GitHub API.
Starting point is 00:15:08 So yes, in theory, this has been kind of redesigned. You can see here like there's some nice things about how bug bot feedback has been shown or how cursor feedback has been shown. It has some like smart little suggestions at the top that. are nice. But at the end of the day, this is just like GitHub with maybe less features, maybe less integrated, much more integrated with cursor. So I appreciate that. And I can see the vision. I can see the vision. I promise. But right now, if I'm using a GitHub hosted repo and I'm not ready to go like all in on the cursor ecosystem, there's just not enough here in this early launch
Starting point is 00:15:53 to get me to pull over. And so, yes, I can like at cursor here and ask it to fix the Vercel preview French issue. Please, why is it failing? Because this is failing. I can do that. Nope, we got an error. So maybe I can't do that.
Starting point is 00:16:19 GitHub's having an outage right now. So maybe this is a GitHub issue. Maybe this cursor is. issue. But again, like today's today origin, I'm seeing that they're setting the groundwork for getting you to get your repos in a cursor. They're setting the groundwork for Bugbot and cursor to be first class citizens in your repo. I'm sure there are some affordances on like the CLI side that make it easier for cloud agents to open PRs and it's just going to be just like a nice buttoned up experience. But again, if you have invested deeply in GitHub, which everybody has,
Starting point is 00:17:00 and you have your automations and you have your actions and you have your code owners and you have all this stuff, we're going to have to see more from this experience, more from why AI native, why agent native matters for us all to move over to this new product. Now, literally came out today. They're not even calling it beta. They're calling it, early beta. So I have faith. I'm excited to see what this looks like. But again, just the current GitHub integrated sync. It's like a little slower. I get the redesign, but it's not giving me life. And I'm just too embedded in the GitHub ecosystem. I need to see something wow from this in order for me to click unlock. You know, and I've spent an hour with it, two hours with it,
Starting point is 00:17:45 not that much. I've messed around with opening PRs inside cursor. Again, like, it's redesigned. It's synced to GitHub, though, so it kind of feels like I'm being slow migrated over to cursor and that they just want me to get, like, that they just want to get their claws in me and then really wow me with this experience. But out of the box, I'm not yet feeling the vision, although I'm excited to see how this plays out. It's worth experimenting with. They think it's worth keeping your eye on. it's interesting. No one is happy with GitHub right now. Both stability is like terrible and they really haven't yet given us that like AI native completely reimagined Git experience that we're hoping for. That being said, it is super early on cursor origin. I think we are mostly in the very,
Starting point is 00:18:35 very, very, very early stages of a very, very, very long migration. And it will be interesting to see if cursor becomes the new centralized source of truth for code or if it continues to just operate in the writing code agents space. All right. So those are our two products. I would say I like Grockbot. I have not yet seen the light on origin. The last thing we're going to talk about is GROC 4.6. This is the model that has everybody I know texting me saying, I think I kind of like the GROC models. Now, I think we're all trying this because cursor has defaulted, I've noticed some of their models to GROC. I'm not sure if I'd love it, but that means we give it a try. And, you know, we're going to put benchmarks aside. Everybody picks the benchmarks that they like. I want to talk about the How IAI Viber View for anybody who is new. Whenever a new model comes out, I do a couple things. I test PRDs. I test prototypes. I test design, wireframes, technical changes. And if I like talking to. the model. I run all these evals blind, so I run them against models, and then I actually go through
Starting point is 00:19:49 and grade all of these designs live myself. I've made some changes to the Howie AI Bench, in particular, this design one, where I let the model decide how it wants to redesign the page. And I've made a eval for a very complicated claims adjudication wireframe that has a lot of complex interactions. So those were two new things that we added to the Howie AI Vibech. And then we ran Grok 4.6 through these models. I graded them. AI judges them. I judged them. And then we put them together in a presentation that I show with you before I even know what the scores are. So let's get to the Howa AI clear weighted index of Groch 4.6 compared to other models. all right so here are the results now remember this is my benchmark so i get to decide how we grade
Starting point is 00:20:47 things and so it is 70% my taste 30% are l l l lm as a judge taste and interestingly enough groc 46 is right up there with my my favorite 5.6 soul beating out both sonnet 5 and opus 5 on the clear index i am not surprised by this i do love love GPT 56 whole. It's my default, but I am surprised to see 4-6 up here. I didn't hate it. Now, if you look at the overall recommendation by task, you'll see that 5.6 is my favorite direct writer on PRDs. I just like the clean way of writing. It's very matter of the fact. It's very comprehensive, a little bit technical, but not overly complex. I also like a GPT 5.6 prototype. And so I rated those quite highly.
Starting point is 00:21:45 Opus 5 got graded by the LLM as a judge as the best implementer for a technical like bug triage problem. And then absolutely zero surprise. I continue to like Sonnet 5 for an OpenClaw agent chit chat experience. This always wins. I know Sonnet 5 and for very pithy, AI back and forth, open claw model I just have not found. anything better. Now, if we look at my opinions about prototypes, it's really interesting to see that when following art direction, I like 5.6, but when given broad decisions to make its own design decisions, I like Grock 4.6. Now, I suspect that this is because I can just spot GPT and
Starting point is 00:22:37 Claude Slop a mile away, a mile away. GPD 5.6. Love forest green. Clod loves a brown, tan, orange combo. So I think I was just enjoying this breath of fresh air. That was Groff 4.6. That was not those two things. But in terms of executing complex UIs, I still really like 5'6 soul. And when following our direction, I still like 5.6 soul. Okay, so let's see what it actually shipped. So first question I was trying to answer in this benchmark is can it follow explicit design instruction? And GPT56 Sol and GROC both followed my instructions well and did a relatively good non-slop job of following instructions. GPD 56 just continues to crush at this dock routing app that I have it make where it has to give a very
Starting point is 00:23:32 complex but easy to parse at a glance UI for managing I guess like container ships. Crushes it always does better than the competition. But I think that Grok did a nice job with my technical incident triage app as well as this editorial page. Now, what it's not given any design, we saw that every model did pretty good. We saw a couple good designs out of 5'6. We saw some clean designs out of Claude Sonnet 5. And my favorite one was actually this, design from 4-6, which is clodsop adjacent. We see an orange, we see a brown, but was actually quite cute and a good interactive kind of ordering system for a coffee shop. So I was really pleased with that one. Dense information architecture. I, of course, loved 5-6. I think it's the
Starting point is 00:24:31 best at complex UI and designing things that are not overwhelming, but are technically complete. And And then in terms of just like taking a generic design and making it little better, I saw some good results from opus and sonnet. But when you aggregated all this up, all the ones I scored, good and bad, the two that bubbled up to the top were again, 5-6 soul as well as 4-6. Now, what's really funny about this, again, this is 70% My Taste, 30% the LLM as a judge. if you take my taste out of it, the LLM hates GROC and it loves Claude Opus, it loves Sonnet, and it does not like GPT-56 soul. Again, this is not weighted by an anthropic model.
Starting point is 00:25:21 I actually use GPT 5.5 to grade all these things because it's actually the harshest judge. So while the models still like Claude, I Claire like GPD 56 and I like GROC446 for design. Now, this isn't my experience talking to the model, although you all know I love to do that. This is pure outputs. So all I will say is my conclusion here is the Groch model is a competitor. None of it was terrible. I liked a lot of it compared to the models that you would use out of the box and combined with cursors, harness, its investment in new ways of thinking about things like Git and fun products
Starting point is 00:26:02 like Grokbot. I would say it's time for all of us to maybe become some level of a grok boy. I am so excited about all this new stuff. It's really fun to look at. Again, true transparency. I'm still spending a lot of my time in Codex. I'm spending a lot of time with the five, six models. But given this experience, I think about spinning up cursor a little bit more for coding and see if I can continue to like and get value out of grok, grok, grok, grok, grok. Thanks for joining this mini episode of HowAAI. I want to hear from you whether you love GrapBot, if you think cursor is going to replace
Starting point is 00:26:39 GitHub, and whether or not you've switched to GROC for any of your coding tasks. See you soon. Thanks so much for watching. If you enjoyed this show, please like and subscribe here on YouTube or even better, leave us a comment with your thoughts. You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app. Please consider leaving us a rating and review, which will help others find the show. You can see all our episodes and learn more about the show at how-I-A-I-I-I-I-I-Pod.com.
Starting point is 00:27:11 See you next time.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.