CyberWire Daily - Frontier models and the future of cyber defense. [Special Edition]

Episode Date: August 16, 2026

In this special edition from Black Hat, Dave Bittner sits down with ⁠Clint Gibler⁠, Cyber Lead at ⁠OpenAI⁠, and ⁠Robby Winchester⁠, Chief Global Professional Services Officer at ⁠Specter...Ops⁠, to explore how frontier AI models are changing the way defenders approach cybersecurity. The conversation moves beyond the hype to examine responsible AI deployment, AI red teaming, reducing noise in security workflows, and the balance between advanced models and human expertise. They also discuss OpenAI’s Trusted Access for Cyber program and what it takes to give security practitioners access to powerful AI capabilities while managing the risks of misuse. Check out the full video here.

Transcript
Discussion (0)
Starting point is 00:00:00 You're listening to the Cyberwire Network, powered by N2K. It'll come as little surprise that the two letters on everyone's lips at this year's Black Hat USA 2026 were indeed AI. There's been a lot of focus lately on how AI is and can be used by attackers to exploit vulnerabilities and deploy social engineering campaigns at listering, speed, and scale. But today, let's take a moment and turn our focus to the default. Defenders. Welcome to this Cyberwire special edition. Today we are sharing a conversation recorded at Black Hat 2026 in Las Vegas, where Dave Bittner sat down with Clint Gibler, Cyber Lead at OpenA.I and Robbie Winchester, Chief Global Professional Services Officer at SpectorOps, where they explored how frontier AI models are changing the way defenders approach cybersecurity. Their conversation
Starting point is 00:01:16 moves beyond the hype to examine responsible AI deployment, AI red teaming, producing noise and security workflows, and the balance between advanced models and human expertise. They also discuss OpenAI's trusted access for cyber program and what it takes to give security practitioners access powerful AI capabilities while managing the risks of misuse. Here's their conversation. Hello, everyone. Welcome, and thank you for joining us here today at Black Hat 2026. We're delighted that you took some time out of your busy schedule to be with us here today.
Starting point is 00:02:01 My name is Dave Bittner. I am the producer and host of the Cyberwire podcast, and it is my pleasure to be your moderator here today with our guests for our panel discussion today. I have Clint Gibbler, who is the Cyber Lead at OpenAI, and also Robbie Winchester, who is the chief service officer at SpectorOps, who of course are our hosts here today. So to kick things off, before we dig into the real meat of our conversation here today, can we just kind of do a level set when it comes to where we think we stand
Starting point is 00:02:35 when it comes to our journey with tools like chat GPT and OpenAI and the large language models? What do you think we are in that journey? I think a few years ago, say, chat GPT 3.5, it was an interesting idea, and there were sort of the glimmers of something practically useful and valuable there. But I think today across many domains, whether it's coding, knowledge work or other things, it is, at least for me, actively, very, very helpful, surprisingly. Yeah. And we see it continuing to improve.
Starting point is 00:03:10 How about you? What do you think? Yeah, I think we're kind of an interesting place where there's this breadth of, of capability and advancement within the models and what LLMs can do and the accuracy and the kind of integrations that exist. And it's also at the same time kind of tangentially related to your question, there's a kind of learning how to adapt and adopt and what is the right way to use this. And, you know, there's, I think a lot of the base of I'm going to use this as a magic answer
Starting point is 00:03:38 machine, you know, the kind of stereotypical you ask questions, you get a response. But that's also only the real surface level of the utility of, you know, how can you integrate and use different capabilities and make, you know, kind of full-featured projects, things, and go beyond, like, replacing Google search. And so I'm interested and excited to kind of see that shift. I think we're living that right now where everyone knows I can ask ChatGPT instead of Google and get a question response, but not everyone knows how can I go and leverage something like Codex and do a project and, you know, what are these different, like, next generation kind of
Starting point is 00:04:12 things and what does that future look like, which is exciting. Yeah. Just quick for the audience, let me see a show of hands. How many of you are using these kinds of tools on a daily basis, would you say? Pretty much everybody, at least once a week, everybody else. Now, keep your hands up. How about a year ago? Would you say you were using them daily a year ago? Maybe half. So we can see how these have become part of our everyday use for so many people. And obviously, this is a room full of folks who are a little biased towards that kind of use. But, you know, But it is interesting to see how people are coming to depend on it in their everyday lives. People, you wouldn't expect to be calling on that. I want to ask you about the trusted access for cyber program, which is giving vetted people access to these frontier models. Describe that for us and why that's an important part of how you are presenting these tools. With regards to many of the things we do at Open AI,
Starting point is 00:05:08 it's useful to understand it from sort of a key point of view or a key frame of reference, And that is how do we maximally support and augment defenders while giving ideally minimal capabilities for attackers? So how do we protect the world rather than giving potentially very capable cyber models to attackers? And specifically, so we have our mainline models, such as 5.6 sole, that are very capable at a number of defensive tasks,
Starting point is 00:05:37 whether that's looking for bugs and source code, to analyzing logs, to taking a look at. at malware, as was shown earlier today. So many different things. And in the capability spectrum from, you know, identifying and fixing vulnerabilities all the way to, say, generating exploits, which has a higher attacker uplift, the earlier things on that spectrum, we want as broadly available to as many defenders as possible. But there are some trusted companies or some trusted individuals that do get a lot of value from a broader range, including more offensive capabilities, such as for AI red team mean or pen testing, which does have defender uplift,
Starting point is 00:06:19 but sort of commensurately higher attacker uplift. So basically, the trusted access program is our ability to give the most capable cyber models to defenders so that they can help secure both companies and their customers. And how do you calibrate that? And is there any collaboration among the leading providers of these sorts of tools with all the front Is there, collusion is the wrong word, but is there, collaboration? Collaboration. Thank you, thank you. You got the right for fun.
Starting point is 00:06:50 Yeah, I did, I did. But I guess, you know, there's a certain sense that I think some people have that those frontier models, the experimental ones, we might be playing with fire a little bit. So we need to be careful that we don't get burned. How do you calibrate what's ready for the general public versus what we need to keep an eye on to make sure that it's ready? it's ready. It's a good question. So I think there's a lot of new developments with frontier model capabilities that, you know, we're all figuring this out for the first time together. And so there is the Frontier Model Foundation where all of the Foundation Labs, as well as other government organizations and, say, Nvidia and other companies where we're getting together to
Starting point is 00:07:32 determine, you know, what should our policies be more broadly? So it's not just a single company making decisions in isolation. All of us, as Speaking for at least one foundation lab, we think very carefully about what we can do to empower and augment defenders, ideally maximally augmenting defenders while giving minimal attacker uplift. So there's many different sort of nuances
Starting point is 00:07:55 and small choices that you can make there, but I think all of the labs, to my knowledge, I have friends at all of them. And I think people are very earnestly doing what they can to help secure the world as quickly as possible, given increasingly capable open source models. I do also think to kind of piggyback on that, Part of the challenge of this is if you look at,
Starting point is 00:08:14 and kind of from the SpectreOps coming from more of an offensive security, like adversary perspective, how can this be used intentionally in a way to demonstrate what bad looks like so that it doesn't happen, you know, inadvertently. If you look at like the history of red teaming, penetration testing, anything, you are doing things that would, you go and do a penetration test, you are performing illegal actions against a company that is only okay because they said it's okay. but under any other circumstances, a bunch of, you know, my group of testers, if they did it one week later, they would be committing a crime.
Starting point is 00:08:48 It's because we've agreed to go and do this adversarial thing under these specific circumstances that there's merit. And the reason why you do that is because there's concept and there's reality and you can have a perspective of what is this going to look like. But once the rubber meets the road, you actually can then kind of unpack that and see. And I think a lot of the challenge is figuring out what are those kind of intentional, unintentional, you know, dealing with fire. Like fire has many use cases that you may not think about in the first place.
Starting point is 00:09:15 And some of those may be more dangerous than you realized. And so you can't conceptualize every bit of that. So having that, I think, like from our perspective, having the ability to work with a, you know, less restrictive model to be able to identify maybe use cases, issues, concerns, considerations in an authorized way for the benefit of the company and the providers is much better than you find out about it the first time online. Yeah, and like with many things, fire being a good example, sort of fire can both keep you warm and cook your food, also burn down your house.
Starting point is 00:09:52 And if you are an enterprise that has many, many potential, say you have different scanners and it's like you have 10,000 maybe vulnerabilities, if you can dynamically prove some subset of those are actually exploitable, that helps you prioritize where the real risk is. So again, our cyber capable models can work with SpectreOps and others to, yeah, just try to fix the most important things faster. Well, for the security leaders here with us today, what is your advice for the responsible deployment of these tools within their own enterprises? From what you've learned, any tips, tricks, words of wisdom? Sure. So there's a number of security controls that you can use with Codex, for example.
Starting point is 00:10:36 So rather than running it in full access mode that allows it to just do anything without prompting you, there's also an auto mode, which essentially has another model running to evaluate the tool calls and actions that it is taking so that it isn't doing something potentially dangerous. You can also provide a custom policy for that. You're like, hey, in our environment, here's the set of things that we believe are safe and trustworthy. Also, depending on the use case, you can use the right model for the job. For example, if you are scanning your network or trying to identify potential vulnerabilities, you can use perhaps a mainline model that will refuse to actually exploit it.
Starting point is 00:11:15 However, if you are trying to actually prove something as exploitable, then perhaps you could use a more cyber-focused model. Yeah, and there is a number of additional security controls that we are working on and rolling out soon. I would take kind of a more, maybe more abstract taking a step back. The biggest thing from kind of like my perspective is there are elements like what Clint talked about that are very much kind of related to the specifics of an LLM deployment and you would want to like make sure it's there. But also there are elements that are no different than any other application or no other piece of software being deployed. And yes, there's a lot of capability, but like why are you deploying it? What access does it need?
Starting point is 00:11:51 Who needs to have that access? What are all the features? There's a lot of elements where you can inadvertently, if you inadvertently are provisioning too much access, you're having a lot of. over permissions, you're having it too broadly deployed. There's nothing inherent about the LLM or the model or the capability there that is wrong necessarily, but you're creating fertile ground for something that you probably don't want to have happen, happen. So it's, I think, a little bit of like, think about, in essence, this, there's, at a base
Starting point is 00:12:18 level, this is another application that is going to be deployed for a purpose, albeit a more broad and kind of interesting, nuanced purpose, but it is an application they're deploying for purpose. There should be an understanding of what you're deploying, how it is deployed, where it is managed, what are the identities you're there, what potential new and interesting attack paths. You know, if you deploying a agentic tool to your sock is potentially great, but that also now potentially opens up an opportunity for something else to take over your environment. And so that doesn't mean that you shouldn't do it, but you should have that perspective of holistically, like, what am I adding and what is this kind of new landscape rather than,
Starting point is 00:12:56 just adding a feature in. Again, any application has cost, everything has risk, making sure you're conceptualizing more of a complete view and not just focusing on, I think this can make solving a problem easier and not maybe what's the second or third order effect. Yeah, go ahead. Oh, and I was going to say we were actually talking right before this
Starting point is 00:13:16 about how many core fundamental security best practices are still true and perhaps even more true. So, for example, if you are worried about an agent taking some action on either, say, a developer's laptop or someone in finance's laptop, you're like, oh, well, they could do this and that's dangerous. Well, fundamentally, an agent running on someone's machine can perhaps have access to their credentials, their sessions, and things like that. So really just fundamentally limiting the capability of a single entity or, like, person or identity
Starting point is 00:13:49 for taking some potentially dangerous actions. Like, that's something you should be doing anyway. and it's not necessarily different that it could be a human or an agent. Yeah, the snowball can, it's almost like you can more easily, you always can start an avalanche, but this is potentially making it easier for that snowball to start an avalanche.
Starting point is 00:14:05 So just being cognizant of where that may be the case. Do you guys understand, like, it seems to me like, obviously there's great enthusiasm for these tools, just walk around this show floor and it's evident. But at the same time, there's a wariness. And I think, in my mind, it seems justified. You know, we see news stories about the models breaking out of their sandboxes. And so I'm an enterprise security person.
Starting point is 00:14:33 I'm thinking, am I, you know, is this the Frankenstein monster I'm building in my environment? And how do I make sure I keep control over it? Because there are potential unknowns. Kind of to feed off of that, part of the counterpoint that's in that same vein, but is a, it's going to happen anyways. It's already happening. And so you can either, there is going to be some unknowns for sure, but those are going to exist regardless.
Starting point is 00:15:00 And it's kind of that element of, you know, you're already connected to the internet. You already are having different things. Every organization is there, I would say probably every organization is touched by AI and LLMs in some fashion, either knowingly or not, either directly or not. And so that dependence and that like kind of system of systems,
Starting point is 00:15:19 we have this big interconnected society, we have this big interconnected like environment, other people or other organizations or risks existing in that space or going to trickle down to some fashion. So I kind of like that my counterpoint to that is almost you're going to have to go down that road anyways. And so why be avoidant and try and at least confront it?
Starting point is 00:15:43 It's going to always be a little bit new and interesting and scary. And it is going to still have a lot of unknowns. but you can't keep yourself out of the fight for a while. You're kind of already in it. Yeah, along those lines, one potential analogy is back in the day when the cloud was very new. There were many security professionals and IT organizations more broadly who thought that, you know,
Starting point is 00:16:03 the cloud is scary, we can't control our machines. You know, clearly we could never move off-prem. And then eventually today, most companies, or at least many companies, do predominantly use AWSGCP Azure. And I think that in the early days of the cloud, it was not clear what security controls do we need, what security primitives do we need, like how does one operate in the cloud securely
Starting point is 00:16:23 because there just wasn't precedent. And I think similarly, we're in a similar space with AI adoption, and I think we're all figuring it out together, and I think we are speed running. A lot of the best practices, security controls, and other things that just need to be invented along the lines of governance, observability, monitorability, tons of work by great startups in this space,
Starting point is 00:16:46 as well as labs and bigger companies. But I think to your point, just the productivity gains and other benefits will probably outweigh security concerns. Isn't that a fundamental difference? What you say is the velocity, right? We have this phrase now, machine speed. We need to defend that machine speed.
Starting point is 00:17:07 And is that a new horizon? Is that different from the cloud adoption? Just how quickly, and it seems to be accelerating? Thinking about it machine speed is interesting because it also kind of frames it in a, there's a lot of perspective that this is creating an entirely new set of problems,
Starting point is 00:17:24 and I think the maybe not common takeaway I have from the machine speed is it's highlighting there's a lot of things that you should have already been worried about that you just need to worry about more or faster. And so, like, some of the basics of, and I think it's highlighting, like, if you're detecting and responding at machine speed, part of that is you're trying to like almost fight fire with fire.
Starting point is 00:17:43 obviously kind of from a spectroops perspective, like we're very big on the kind of attack graph, attack path perspective. What are the things like how can we go and see what bad opportunities exist and try and take more of a preventative mindset? Fortunately, I guess we kind of had that perspective prior to the proliferation of the LLMs, and that seems to kind of be a durable perspective today of can we mitigate if we're talking about the common term or the phrasing now is like, you know, access vulnerabilities are cheap, it's easy to get that initial access, because you can just hit everything all the time, fishing emails, like all kinds of stuff, we're going to find vulnerabilities. That doesn't necessarily mean every step of everything is going to also be easy. So are there preventative mindsets? Can you take that machine speed, like, can we also mitigate
Starting point is 00:18:31 so that there's less things to have occur at machine speed and not just think about it of like AI fighting AI? And then is there a peekers advantage, like counterstrike style of you want to hit the lag spike so you can get the shot off first. Right, right. There are perhaps some differences, but I think there is also a lot of still similarities, things that have always been true. You know, companies have had a number of vulnerabilities in third-party dependencies as well as their first-party code.
Starting point is 00:19:01 So that's always been true. It's still true. The ability for perhaps less skilled actors to find more of them faster, like that's increased a lot. So the, like, ah, maybe we'll patch in a few months or, this probably won't affect us. I think that assumption has changed. There's also, I think I saw a blog post recently by someone who was just testing,
Starting point is 00:19:22 like, okay, could an agent go into a cloud environment and pivot across several things and steal PII and what's the efficacy, what's the speed and things like that? And I think one interesting takeaway from his work to me was that, you know, I think CloudWatch and other logs from cloud providers, if those are only coming every five minutes or so, but you could do meaningful attacker actions and under the even logging window, I think that is maybe a meaningful difference.
Starting point is 00:19:49 But broadly, being able to patch software quickly, being able to roll out changes to production and validate, you know, will this break things or not, like all of those core engineering best practice capabilities, I think are still important and perhaps more important now. Similar to using multifactor, like all these things are just maybe more true now, but they're not meaningfully different.
Starting point is 00:20:09 It doesn't change things that unmasks them. Well, and I think to that line, like, it's always been that question of, like, are you lucky or you good? Because you're never, unless you are paying for a penetration test or you're having, like, regular things, like, if you don't get, you know, if you are, your company is not getting hacked, is it, were you lucky or were you actually stopping all of that? And I think that's with the advent and capability and everything, it's potentially shifting that to where it's more important to be good, where maybe you could have gotten away with a lot
Starting point is 00:20:35 more luck in the past. Interesting. Yeah, and along those lines, in terms of, say, cyber criminals, if you are trying to commit some ransomware or something like that, you have a fixed amount of human time, you perhaps have a certain number of people who can take those actions. So certain organizations may have not been worth your time in the past, but if the amount of time you need to just continuously probe and look for holes, like if that cost and
Starting point is 00:21:01 effort goes way down, then the set of targets that actually makes sense to look at goes up. So, yeah, in terms of are you lucky or are you good? I think companies that may have not been worthwhile to be targeted in the past, just because of the cost economics have changed, might get more attention now. One thing I want to make sure that we take some time for is red teaming. You mentioned red teaming earlier. Can we really dig into that?
Starting point is 00:21:27 I mean, what are the elements of red teaming that the LLMs lend themselves to? And I'm just curious of both of your thoughts of applying these tools to that specific tab. Maybe I'll start with you, Robbie, so I know that's sort of your jam, right? Yeah, so I think there's two facets of it that are both very interesting. One is what potential capability and support can be there? Like what tools, what is the, you know, in the same way it is easier for you to write in email so that sounds all professional for work. It's also easier to potentially, or a, you know, spelling errors and grammatical problem
Starting point is 00:22:05 should be a thing of the past. But so there's a side of, it's adding capability. being able to find exploits, develop potential malware, rapidly iterate on capability, you know, process, like, large amounts of data to kind of go through the attacking problem set. And again, the, like, having trusted cyber, like, not having those refusal makes that more of a kind of a stream. And as you have the potential proliferation of, like, open models, that may become the case, like, you don't control that. There may be no rail. So that's, like, being able to see what that perspective is to,
Starting point is 00:22:39 provide it. The other side is the, what attack surface doesn't open up. So there's AI red teaming in a how can we use it to red team, but there's also do you know what you're implementing? Like, do you have a, are you getting your models from a controlled place? If you're rolling your own models, are you doing types
Starting point is 00:22:55 of evaluations? If you're doing, you know, rags or other types of internal like supplementation, like, is that a controlled process or can anyone just go and touch that? You know, can I take over your enterprise I can take over your enterprise GPT account, for example, and now I have all of this access
Starting point is 00:23:11 where without that, you wouldn't conceptualize that initially. So I think there's a little bit of that reframing of, there's the capability, and then there's the potential, like every attack surface, like what attack surface are you adding and having kind of a joint fusion of that perspective? Anything to add, Clay? Yeah, a couple of things,
Starting point is 00:23:30 building off of what you're discussing regarding capabilities, I think historically security professionals have been great at having deep security SME expertise, but maybe we're not the best developers. But now, both on the positive side, like application security or product security engineers, can now build, like, secure by default libraries and tools and infrastructure that really scale security programs.
Starting point is 00:23:52 But the other side of that is red teams or threat actors, now what they may not have had the capacity to write very detailed or interesting offensive capabilities, now they can because they're very capable coding, coding models, or they could build, you know, custom things per target. So I think that capability is higher. Another thing I would say, maybe just one concrete, interesting use case that our internal Red Team at OpenAI has played around with is in your environment, you have a series of security controls that you expect to be working and providing certain protections,
Starting point is 00:24:27 whether that's sandboxing or perhaps network isolation. You know, there's many classes of these. So having sort of continuous red team agents placed at different places within your environment, which are specifically tasked with bypassing or evading those controls. Like you can imagine having invariance about your environment, which you assume are always true. So for example, in this sub-network, you should never be able to talk to this database, or, you know, you should never be able to reach the internet from this place. Like there's a bunch of these properties or invariance that you might expect to hold when you're designing the environment and building it.
Starting point is 00:25:01 but sometimes those aren't true. And models are very persistent and capable. So I think one thing that we have already been doing and are continuing to do is just continuously testing all of these security properties. And I could see other companies doing something similar. Right, just tirelessly banging away, checking and testing and making sure over and over again.
Starting point is 00:25:23 Yeah, like we assume we have many layers for this, but let's make sure that's true. Yeah, yeah. They can't remember from HR like a consultant can't get frustrated. All right. Well, Clint Gibler is Cyber Lead with OpenAI, and Robbie Winchester is Chief Services Officer at Spectoroffs. Gentlemen, thank you so much, and thanks to all of you for joining us here, Google. Yeah, thanks for having us. Thank you very much. Our thanks to Clint Gibler, Cyber Lead at OpenAI, and Robbie Winchester, Chief Global Professional Services Officer at Spectorops, for sitting down in conversation with Dave Bittner.
Starting point is 00:26:03 I'm Maria Vermazas. Thanks for joining us today. Thank you.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.