This Week in Startups - 90% of AI prototypes never reach production (w/ Temporal's Samar Abbas) | AI Basics

Episode Date: September 15, 2026

The TWiST AI Basics series is made possible by Google for Startups! Build context-aware AI workflows using DeepMind models and orchestration tools with Google for Startups' new Startup technical gui...de on generative media. Get the guide here: https://goo.gle/technical_guide   Today's show: The demo works, the vibes are immaculate, but then the app falls apart the second real users touch it. On today's AI Basics, Temporal co-founder and CEO Samar Abbas says the problem isn't with the model at all. You're missing a harness, the layer that keeps an AI agent's work durable, secure, and recoverable if/when something breaks. Temporal's durable execution platform is used by major companies from OpenAI to Stripe to Netflix, and helps to make their long-running software reliable and interruption-free. On AI Basics, he explains why so many AI-coded prototypes die in the imagineering stage, before ever hitting production. Plus he walks Jason through Temporal's live dashboard, to show how it gives developers full visibility into what their agents are doing at every step. Guest and Relevant Links: Samar Abbas on X: https://x.com/SamarAtTemporal Temporal: https://temporal.io/ Temporal raises $550M at a $12.55B valuation: https://temporal.io/blog/temporal-raises-usd550m-series-e-at-usd12-55b-valuation-ai GeekWire: Samar Abbas on the "massive platform shift" in AI: https://www.geekwire.com/2026/temporal-ceo-samar-abbas-on-the-massive-platform-shift-in-ai-fueling-the-startups-5b-valuation/ Temporal expands Google Cloud partnership: https://temporal.io/blog/temporal-expands-google-cloud-partnership-with-pay-as-you-go-pricing Timestamps: 0:00 Why every startup needs an AI harness 2:12 Why 90% of AI prototypes never reach production 4:16 What happens when an agent crashes mid-task 7:00 A look at Temporal's dashboard 13:02 OK but what is a "harness" exactly? 15:30 When to graduate an agent from laptop to production 17:44 Organizing agents into teams Subscribe to This Week in Startups on Apple: https://rb.gy/v19fcp Follow Jason: X: https://twitter.com/Jason LinkedIn: https://www.linkedin.com/in/jasoncalacanis   Thank you to our partner! Build context-aware AI workflows using DeepMind models and orchestration tools with Google for Startups' new Startup technical guide on generative media. Get the guide here: https://goo.gle/technical_guide Check out all our partner offers: https://partners.launch.co/   Great TWIST interviews: Will Guidara, Eoghan McCabe, Steve Huffman, Brian Chesky, Bob Moesta, Aaron Levie, Sophia Amoruso, Reid Hoffman, Frank Slootman, Billy McFarland   Check out our suite of newsletters: https://substack.com/@thisweekininc   Follow TWiST: Twitter: https://twitter.com/TWiStartups YouTube: https://www.youtube.com/thisweekin Instagram: https://www.instagram.com/thisweekinstartups TikTok: https://www.tiktok.com/@thisweekinstartups Substack: https://twistartups.substack.com

Transcript
Discussion (0)
Starting point is 00:00:00 All right, everybody, welcome back to our basic series. What is our basic series? Well, on this week in startups, we get asked the same basic questions over and over again. Sometimes it's a legal question for founders. Sometimes it's finance. Sometimes it's about product market fit or getting customers. And increasingly, and almost all the time, it's about basic things in AI. Every startup is not just building AI products, but they're using AI to build those products, right?
Starting point is 00:00:26 It's, you can only be competitive today if you're using AI inside your startup. So today on AI basics, our guest is Samar Abbas. He is the CEO of Temporal, T-E-M-P-O-R-A-L. Dot I-O if you want to go take a look at the website. And he's a 20-year engineering veteran worked at AWS, Microsoft, and Uber. Today, we're going to talk about why agents desperately need harnesses and how enterprises can safely dip their toes. to the agentic era and why he's building temporal on Google Cloud. Welcome, Samar.
Starting point is 00:01:01 Thanks, Jason, for having me here. Excited. Yeah. And so you've been at this for a while. You've been an engineer for a while. How have things, just as a warm up here, how have things change from when we were building products and services 10 years ago to how products and services have been built in the last 10 months or even the last 10 weeks? Because it feels like it's changed even between those two time frames, yeah? Yeah. So, Jason, like, I get this question a lot. Like, a lot, everything looks different, especially in this world of,
Starting point is 00:01:33 there's clearly a new platform shift happening with the rise of AI. Essentially, these models get smarter and smarter. One of the things that I see a lot is, although the way people are thinking about applications is very different, like, especially with the rise of now coding agents coming in the mix, everyone suddenly is a software developer on the planet. Because software is now more approachable to a wider class of people out there, not just like people with computer science backgrounds. Like everyone can build a play.
Starting point is 00:02:06 If you have an idea, coding agents are making those more accessible to everyone on the planet. So on one side, it feels like the way application, the entire new application platform is emerging on one end. But on the other side, what I see is we are seeing a very, similar set of problems is that's why you hear a lot that, oh, like, I have an awesome idea.
Starting point is 00:02:31 I got started and built an application through a coding agent, but like 90% of those ideas die after a POC, essentially. They never see light of the day. And why is that? And they die at that proof of concept stage. People get super excited. But then all of a sudden there's like this despair. It's not stable.
Starting point is 00:02:50 It's brittle. It's not replicable. What's the series of problems here? I see that all the time. People get excited inside my organization. They show me something, and then I'm like, okay, is it ready for production? And the answer is inevitably, I don't know. It doesn't feel like it, but maybe.
Starting point is 00:03:06 And that's really the difference between production level software and a proof of concept. There's a big gap there, isn't there? There is a big gap there. And I think this is where the worlds are starting to look more and more similar to what we are seeing in this platform shift, essentially. is as more and more of AI applications and especially these agents are starting to hit production, what we are seeing is they get longer and longer durations in time. The more scale, the more asynchronous they get, the more longer they run, the more actions they take in real world, that more value they create.
Starting point is 00:03:42 It's so clear for every organization that that's the pattern. But the moment they go there, you start running into the same set of reliability. scalability, durability, challenges, infrastructure level, which feels very similar to when the cloud platform shift happening. Because suddenly all of these agents are now starting to graduate
Starting point is 00:04:03 from running from your laptops into a real distributed environments and these problems starts to look more and more similar to like distributed in nature rather than a completely new AI problem. And one of the problems is the durability of the execution You, everybody's had the experience no matter what AI product you're using, that it's working, it's working, it's working, and then it peters out, it stops, and you have to start over. And you might have to start over, you know, a very large job. This happens in generative AI, whether you're making videos or images or sounds. This happens when you're doing research reports or you're doing big web research, whatever the product or service you're building is, is durability keeps happening. And then it doesn't
Starting point is 00:04:51 seem to know because this is not, you know, perfect software yet. And these large language models in some cases are making their best prediction. It doesn't know where it left off. Why is that? And how do you solve that problem? This is durability when running. I think, Jason, this is basically, I feel this is becoming table stakes. As these AI systems are maturing from like BOCs to powering real production.
Starting point is 00:05:21 construction systems. Imagine an agent which is kind of building, processing a refund. And literally, immediately after processing a refund, a failure happens. A machine restarts our process crashes. And it is before even notifying the customer, like, it just dies. And I think this is where a lot of these POCs are getting stuck is solved that durability problem because now what are your options? Either you reprocess, either that give another refund, a second refund. Yeah, you wish to are like- Could get costly. Exactly. And so handling those failures are exactly the place where all of those approach and systems are having challenges as this AI starts taking action in the real world. And this is where a platform like temporal is kind
Starting point is 00:06:15 of becoming super critical to just provide like at the end of the day, what we, give you is we are we describe this thing as durable execution where during an execution of a code if a failure happens we remember all of that state without you as a software developer writing a single line of code for it it means that if you process a refund and a failure happens or process crashes we will you don't have to manage that state yourself it's durability stored for you underneath the cover by the platform and we will guarantee your application can keep on making forward progress in the presence of all sorts of failures. So this is exactly the kind of thing, which is becoming the foundation, which is powering
Starting point is 00:06:58 these AI systems out there now. Which makes a lot of sense. It speaks to the fact that people are starting to move AI into production. They're moving it from, you know, just little projects inside an organization to actually doing the thing. And if you were to just, when you were talking, I was just visualizing like going to the counter where you have to change something with your flight. there's a young person who's new in the job and they're trying to figure out your baggage and what
Starting point is 00:07:23 seat you're in. And then this, you know, person comes with a lot of wisdom and like, let me show you. And then they walk through it step by step and they fix it. We have that in the human world. We have people who are trainees. We have people who are mentors. And they say, okay, walk me through what you did. And then they fix, okay, yeah, we're going to get these bags checked into Paris, no problem. And I was on your website earlier. And it all came together for me when I was looking at, you know, know, this beautiful Agentic AI example or the subscription example. And here you have, I guess, in the top left corner, your workflow, your code, and then you have this temporal event timeline that plays. And you can kind of watch it work. I don't know what the console is doing on the
Starting point is 00:08:06 bottom left. What is that part of the screen doing here? Yeah. So I can quickly explain. As I told you earlier, we are the platform which ensures durable execution of your code. So the top left window is actually showing your real application. This could be code written by a developer. This could be actually code written by one of the coding agents, which we have temporal skills, which is actually, you can just express. I want to build an application which helps me research a bunch of things on the internet and then consolidate the response. And then it can emit a code which looks similar to what you are seeing in the top left window here essentially. And the right window is kind of showing you actual runtime when you actually run that application. And,
Starting point is 00:08:47 it's doing a query, it's actually showing you progress of how the application is making progress in real time. You can see it's executing tools, it's calling LLMs, and then getting responses back, and then it's basically showing you how it's making forward progress. And I think the window in the bottom is the actual output. It's generating that application. Yes. And this really gets to the core of it.
Starting point is 00:09:16 I've never seen somebody show me what's going on behind the scenes. And most people who were just asking a question or where should I go on vacation, I guess I got vacation on my mind here as we end the summer, but, you know, it could be shopping application. It could be a research application. We don't actually get to see too much of what's happening now. We're abstracting that away from the client, the consumer using AI or even the enterprise. And that's for good reason, I guess. You want to make it simple. But it is. more reassuring when I see it playing out like this. I wish I had this heads-up display available to me in every consumer product. It would make me feel a bit more confident of where I left off.
Starting point is 00:09:56 And Jason, you are making a really great point. And I think that's one of the bigger powers that you get out of temporal is, first of all, it gives you a transactional engine, which guarantees not only durability and reliability, but at the same time, it's actually giving you complete visibility into especially these AI, agentic applications, which is mostly like what is an agent these days, an agent is an LLM. You pick an LLM, you take a prompt and give it a set of tools, and these LLMs are smart enough where they are deciding what they want to do and how to make forward progress. So exactly to your point is what temporal not only give you a transactional engine, which give you durability guarantees,
Starting point is 00:10:43 but it gives you complete visibility into what your agent is actually doing to execute your task to get the right outcome. I think would assume that the person could then go in and tweak some parts of that. I've seen some really interesting products emerging, you know, as I invest in companies where they say,
Starting point is 00:11:01 oh, when we make you this short 10-second video or this movie poster you wanted to make, we're going to actually show you how it's being built and then, hey, here's where you can make a change to something. And so that's the, I think, the ultimate possibility here, yeah, is that it'll give us the ability to change certain steps in the process because we have that insight into where it's breaking down or where it might be hallucinating or not performing up to our expectation, yeah?
Starting point is 00:11:30 It even goes beyond that tweaking. First of all, absolutely. It gives you visibility. It can see what are the things it's doing correctly. What are the things you need to adjust, essentially? and it gives you a way to go and make those adjustments to your business logic to kind of be more closer to the real outcome you want. But I think imagine the vast majority of these agents being written, they are now, there is this thing called code mode where an agent on the fly solving a complex problem by building more code essentially.
Starting point is 00:12:05 And it's actually step in the process where it creates, it emits more code. and then executes that code, and imagine a larger enterprise who is kind of now starting to transition these large business processes to these code mode agents. It's a very unnerving world from a security lens because suddenly you are literally like running code emitted at runtime from an LLM into a business environment, essentially. So what temporal gives you, it has two constructs called workflows and activities. Workflows is all about generating commands. So even those code mode is actually allows you to generate the code which needs to execute in the next step. But then it gives you an
Starting point is 00:12:51 ability we can intercept all of those commands and provide the necessary guardrails to secure that environment before it gets executed in a real sandbox or a runtime. Got it. So talk to me a little bit about the role of the harness versus the model. We've heard. a lot of people saying these days, hey, models are great, obviously, but you got to get the harness right. Explain to people in plain English what that means to get the harness right and that the power and this next phase of executing on the potential of the models is the harness. What does it mean for somebody who's listening who maybe can imagine it but doesn't understand practically what it means? So at least the way I describe this world is I feel right now we are transitioning.
Starting point is 00:13:40 from an MS DOS era of agents to a real cloud environment, essentially. Because imagine mass majority of the people today, how they are running agents, if they install something on their laptops and like a coding agent, and then they are running those coding agents, which is actually solving complex problems and adding value to their day-to-day work, essentially. But clearly the value we are getting from these LLMs is now people have started to build loops or even there is a thing loop engineering or graph engineering which has been talked about a lot out there which means that people are building these agents which runs more independently
Starting point is 00:14:20 for longer periods of time. So which means these agents are going to now move away from your laptop to a distributed environment. This is where a hardness becomes a very core component of driving execution is it's the brain where you people are separating out the brain is if you give an agent a very complex task which requires to talks to dozens of agents underneath the cover it needs to coordinate and it runs to do multiple tool invocations which themselves can run for minutes hours or even days to get some complex tasks done you need to separate out a brain so when an agent goes down the round path, it doesn't break the entire process. So harnesses are these brains, which is separated outside of the agentic loop, to drive these distributed architectures.
Starting point is 00:15:20 How do people know when it's time to take that agent they've been playing with on their desktop and then say, okay, I want this thing to run in a loop. I am confident enough. I've worked with it. It's my researcher. It finds me, let's say it's the sales department, finds me my next set of leads and then I wanted to put them into the CRM and then I wanted to draft me a custom message and then I want to get them on a Zoom call or whatever. How do we know when to move it off of our laptops and move it into this loop where it's been given, hey, your goal is to find me the great leads based on our CRM that are going to, you know, result in more sales and everything's going to be hunky dory? How do they make that jump? What?
Starting point is 00:16:07 needs to get involved when you're in an enterprise in making that kind of jump where it goes from my desktop. I'm just playing around trying to make myself more efficient, but now I want this to be a corporate process. I think it's for enterprises. I don't think enterprise can even adopt these agentic architectures until they can run with all of the guardrails which are needed to execute like these business processes in those environments, which means like none of the enterprises will ever allow these agents to run on a laptop, essentially. So for them, it's core table stakes, in my opinion. And I think it is basically more of a question for,
Starting point is 00:16:47 I think as a software developer, I think what's happening is, like in smaller organizations, people, you know, like people started with using AI by giving their prompts to chat GPT, and they're getting intelligence answers back. Now we are in the next phase, where people are building these agentic loops, which is giving it a bunch of tools also
Starting point is 00:17:09 and let the LLM invoke those tools to get their right outcomes. I think at this point, running agents on your laptop is no longer even an option. So that's why you are already seeing these architectures emerging with sandboxes, memory, and like all sorts of architecture evolving
Starting point is 00:17:29 where how people are recommending how you should be running these agents in a safe and secure manner. security alone is pushing and making, like making it a requirement that people should be running these agents outside of your laptop. But I think the other indication is at some point you graduate from just a single individual working on a problem to team of people. And now these team of peoples are coordinating with an army of agents running to solve complex business problems. Once you are in a world where it's no longer a single individual working on a task,
Starting point is 00:18:06 it's a team solving on complex business problems. The only viable path is through running those things through an agent in a distributed environment. And that's where a forward deployed engineer has to get involved. They've got to embed themselves in that business unit and use the proper tools and build it properly. So it is not brittle. And that's, I think, what a lot of organizations are experiencing. now, huh? We're going to have the paradigm shift from, I'm just going to do things on my laptop and have fun to, hey, how does this affect the entire organization?
Starting point is 00:18:39 Exactly. It's a really interesting moment in time. The value is amazing. The opportunities are unlimited, but you're going to have to be thoughtful, folks, and definitely check out temporal.io. Samar, thank you so much for coming on the program and sharing these important perspectives. Thanks a lot, Jason. Really enjoyed the conversation. Find out more about 10. Temporal, as I mentioned, by visiting temporal.io. And all of our AI Basic episodes are available at this week in startups.com slash basics. Thank you so much to our friends at Google Cloud.
Starting point is 00:19:13 We're making great products for all of my startups. We really appreciate that. And for supporting independent media link this week in startups.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.